Complex product process route flexible planning system based on deep cycle Q network, planning method and application

By adopting a flexible process route planning method based on DRQN ​​in complex manufacturing environments, combined with LSTM and DQN technologies, the problem of lack of efficient strategy adjustment and knowledge update in the existing technology is solved, intelligent planning and rapid adaptation of process paths are achieved, and the intelligence and robustness of the manufacturing system are improved.

CN120197991APending Publication Date: 2025-06-24SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510366140.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When the existing technology faces dynamic process requirements and environmental disturbances in complex manufacturing environments, it lacks efficient strategy adjustment mechanisms and knowledge update mechanisms, resulting in model response lag, strategy migration difficulties, and even complete failure.

Method used

A flexible process route planning method based on deep cyclic Q network (DRQN) is adopted, and combined with the advantages of long and short-term memory network (LSTM) and deep Q network (DQN) to achieve dynamic decision-making on processing state. Introduce an adaptive reward mechanism and selective forgetting strategy to improve the model's ability to adjust under the needs of variable process and handle old knowledge.

Benefits of technology

It realizes intelligent planning and rapid adaptation of process paths, improves the intelligence level and robustness of the manufacturing system, reduces computing overhead, and can respond faster to process route changes caused by process changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197991A_ABST
    Figure CN120197991A_ABST
Patent Text Reader

Abstract

The invention discloses a complex product process route flexible planning method based on a deep cycle Q network. The method is suitable for efficient process route optimization in a dynamic manufacturing environment. The method comprises the following steps: firstly, modeling process characteristics, processing operation, resources and constraint relationships of a part, and constructing a state space and reward mechanism based on a partial observable Markov decision process; and then constructing a long short-term memory network to extract time sequence characteristics in a processing state, and constructing a DRQN on the basis to realize an optimal decision of processing operation. In order to adapt to variable process targets and environments, a self-adaptive reward function adjustment strategy is introduced, and dynamic optimization is realized by adjusting weights such as processing quality, time and energy consumption; and meanwhile, a selective forgetting mechanism is constructed, and model parameters are suppressed and updated by utilizing a Fisher information matrix, so that efficient forgetting of outdated process knowledge is realized, and the migration ability and generalization performance of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent manufacturing, and particularly relates to a flexible process route planning system, a planning method and an application of a complex product based on a deep recurrent Q-network. Background Art

[0002] Process route planning is a key link in the manufacturing process of complex products, which directly affects production efficiency, product quality and the allocation and utilization of manufacturing resources. Especially in high-end manufacturing fields such as aerospace and equipment manufacturing, the product structure is complex, the manufacturing processes are numerous, and the process route planning faces difficulties such as diverse process characteristics, complex operation constraints, and highly coupled resource utilization.

[0003] Although existing methods have introduced artificial intelligence algorithms, most of them still assume a static environment as a premise and lack the ability to cope with process changes and processing environment disturbances. In the manufacturing site, once process feature changes, processing resource replacements or production target adjustments occur, the original model often needs to be retrained from scratch, resulting in system response lags, difficult strategy migration, and even complete failure. In addition, the current methods lack a mechanism to identify and forget invalid knowledge, resulting in the interference of old strategies on the optimization effect in the new environment, seriously affecting the stability and practical deployment value of the system.

[0004] With the development of deep reinforcement learning technology, its application in the field of intelligent manufacturing has become a current research hotspot. Among them, the deep recurrent Q-network (DRQN) constructed by combining the long short-term memory network (LSTM) and the deep Q-network (DQN) shows strong advantages in capturing sequence features and modeling time-series decisions. LSTM can effectively model the long-term dependencies between operations, and DQN realizes the learning of the state-action value function, thereby improving the decision-making quality. Nevertheless, the DRQN framework still lacks an efficient strategy adjustment mechanism and knowledge update mechanism in the face of dynamic process requirements and environmental disturbances. Especially in the scenario of frequent process changes, the model often requires high-cost retraining and lacks flexibility and real-time performance.

[0005] Currently, the existing technologies cannot fully meet the actual needs of "high dynamic response, strong adaptability and sustainable learning" in complex manufacturing environments. There is an urgent need for a flexible process route planning method that can support online strategy adjustment, dynamically adapt to changes and achieve selective forgetting of knowledge, so as to improve the intelligent level and robustness of the manufacturing system. Summary of the Invention

[0006] The present invention provides a flexible process route planning method based on DRQN, which combines the time series modeling ability of LSTM and the strategy optimization ability of DQN to achieve dynamic decision-making on machining states; an adaptive reward mechanism and a selective forgetting strategy are introduced to improve the adjustment ability of the model under changing process requirements and the processing ability of old knowledge, so as to realize the intelligent planning and rapid adaptation of process routes.

[0007] The technical solution adopted by the present invention is as follows: a flexible process route planning method for complex product based on DRQN, including the following steps:

[0008] (1) Model information such as process features, process operations, machining resources, and constraint matrices of complex product parts;

[0009] (2) Based on the partially observable Markov decision process (POMDP), model the process route of complex products, define the state vector, action space, state transition probability, observation space, observation probability, and reward function, and establish a process constraint matrix to describe the dependency relationship between process features;

[0010] (3) Construct a long short-term memory network (LSTM) structure to perform sequence modeling on machining states and extract the time-dependent features hidden in the process route;

[0011] (4) Based on LSTM, construct a deep recurrent Q-network (DRQN), and train the intelligent agent to select machining operations according to the current state to achieve the optimal decision-making of the process route;

[0012] (5) Combine the adaptive adjustment strategy, and according to different machining scenarios and goals, adjust the weight parameters of machining quality, time, and energy consumption in the reward function to achieve a dynamic reward mechanism to adapt to environmental changes and process goal adjustments;

[0013] (6) Introduce a selective forgetting mechanism, evaluate the importance of model parameters by calculating the Fisher information matrix, and suppress the parameters related to the changed features to achieve the forgetting of specific information;

[0014] (7) Combine the Boltzmann exploration strategy and the dynamic forgetting mechanism to adaptively train the SADDQN model under process and environmental changes to optimize the part machining strategy.

[0015] Further, the specific implementation method of the step (1) is as follows: each complex product part is composed of several process features, each feature is completed by multiple machining operations, and different operations correspond to specific machining resources, including a machine tool set and a tool set, as well as machining steps such as the feed direction. The process relationships such as the dependency, adjacency, pattern, and datum between features are represented by a constraint matrix.

[0016] A more specific implementation of step (1) is as follows:

[0017] Model the process characteristics, process operations, processing resources, and constraint matrix depth cyclic Q-network information of complex product parts.

[0018] Each part consists of n process characteristics F, where F i is the i-th process characteristic; each F i is completed by a series of process operations Op i . Different operations need to match corresponding processing resources and processing steps, including the machine set MS and the tool set TS, and processing steps such as the feed direction DS depth cyclic Q-network. To represent the processing constraints between process characteristics, a constraint matrix CM is constructed, where the element in the i-th row and j-th column of CM represents the process relationship between feature F i and F j . The process relationship includes the dependency relationship DR, the adjacent relationship AR, the pattern relationship PR, and the datum relationship TR, which are used to constrain the process sequence and path.

[0019] Furthermore, the specific implementation of step (2) is as follows: Process route modeling based on POMDP, where: The state vector S represents the execution state of each processing operation. S i = 0 means to be executed, and S i = 1 means completed; the action space is the set of optional processing operations OpS avl that conform to the current state and constraint relationships; the state transition probability is 1 / |OpS avl |; the observation space includes the combination of processing resources and steps required for the current operation; the observation probability is calculated according to the number of combinations of processing resources and steps; the reward function is defined as R.

[0020] A more specific implementation of step (2) is as follows:

[0021] Model the process route of complex products based on the partially observable Markov decision process, define the state vector, action space, state transition probability, observation space, observation probability, and reward function, and establish a process constraint matrix to describe the dependency relationship between process characteristics, as defined below:

[0022] State vector: Map the execution status of part processing operations into the state vector S, defined as S = [S1, S2,... S i ,... S m , where m is the total number of operations of the part, and S i represents the processing state of operation Op i . The value of S i has the following two cases:

[0023] S i= 0: This operation can be executed;

[0024] S i = 1: This operation has been completed;

[0025] Action space: According to the current machining state s of part P t , under the constraint restrictions between process features, the executable operations a that can be selected at time t t Set OpS avl ={Op i ∈OpS^s t +Op i ~CM(P)}, where + represents performing a certain machining operation in a certain state; ~ represents conforming to the constraint relationship;

[0026] Transition probability: The next processing operation is a possible operation selected according to the current state and the constraint restrictions between process features. During the state transition process, the state transition probability is 1 / |OpS avl |, where |OpS avl | represents the number of executable operations;

[0027] Observation space: The currently selected action is completed through a series of steps of the selected machining resources and the machining step depth cyclic Q-network, including the available machine tools, available cutting tools, and feed directions for a certain machined part;

[0028] Observation probability: According to the currently selected action Op∈OpSavl, the corresponding available machining resources and machining steps are combined; During the state transition process, the state transition probability is defined as:

[0029] 1 / (|MS(Op i )|·|TS(Op i )|·|DS(Op i )|)

[0030] where MS(Op) and TS(Op) respectively represent the number of available machining resources and the number of machining steps corresponding to the execution of operation Op;

[0031] Reward function: After the agent selects an action, according to the change in the machining state of the part, in order to obtain the optimal process route, the reward function R is defined as:

[0032] R = w q ·R q -w t ·R t -w e ·R e

[0033] where R represents when the part is in state S t , performing action a tThe rewards obtained; R q , R t , R e respectively represent the processing quality, processing time, and processing energy consumption required for the selection action a t , that is, the corresponding operations; w q , w t , w e are normalized weight coefficients, and their values can be determined according to process requirements.

[0034] Furthermore, the specific implementation method of step (3) is as follows: By constructing a long short-term memory network, the current processing state vector and the hidden state of the previous moment are used as inputs, and are processed through the input gate, forget gate, and output gate in sequence to extract the time-dependent features in the process state sequence. Among them, the forget gate controls the retention degree of historical information, the input gate adjusts the update ratio of current information, and the output gate combines long-term memory to generate the current output state, so as to realize the modeling and expression of the time sequence relationship in the process flow.

[0035] A more specific implementation method of step (3) is as follows:

[0036] Construct a long short-term memory network structure to perform sequence modeling on the processing state, extract the time-dependent features hidden in the process route, and input the processing state information s at time t t and the output h of the long short-term memory network at the previous moment t-1 , and pass them through a linear combination to be processed by 3 gates, including the input gate, forget gate, and output gate. The linear combination of s t and h t-1 is compressed to the interval (0,1) through the sigmoid activation function to obtain the forgetting factor f t . Multiply the long-term memory c t-1 by the forgetting factor f t , which determines the forgetting ratio of the long-term memory c at the previous moment t-1 ; Multiply the memory factor by the short-term memory at the current time t, and then add it to the c t-1 that has completed the forgetting process to obtain a new long-term memory c t ; Calculate the output factor out t through the sigmoid activation function. Use the tanh activation function for the new long-term memory C t , multiply the output factor out t by the long-term memory c processed by the tanh activation function t to obtain the result output h t .

[0037] Further, the specific implementation of step (4) is as follows: Based on the time series modeling ability of LSTM, a deep recurrent Q-network is constructed. Multiple adjacent processing states are used as inputs, and the LSTM network extracts time-dependent features, which are used as the inputs of the deep network. During the training process, the agent predicts the Q-value of the current state, selects processing operations in combination with the reinforcement learning strategy, and realizes the optimal decision-making for the process route.

[0038] A more specific implementation of step (4) is as follows:

[0039] Based on the long short-term memory network, a deep recurrent Q-network is constructed to train the agent to select processing operations according to the current state and realize the optimal decision-making for the process route. Based on the ability of the long short-term memory network to model the processing state sequence, its output is used as the input feature of the Q-network to estimate the action value of each processing operation in the current state. By constructing a double Q-network structure as the policy network and the target network respectively, the experience replay mechanism is used to optimize the policy during the training process, and the exploration strategy is combined to select actions in the state space, enabling the agent to achieve the adaptive optimization and decision-making ability of the processing path in a dynamic manufacturing environment.

[0040] Further, the specific implementation of step (5) is as follows: By constructing an adaptive adjustment strategy, rapid response to different processing scenarios and process objectives is achieved. First, weight coefficients are set according to key factors such as processing energy consumption, processing time, and precision, and an adjustable reward function is constructed, enabling policy adjustment without retraining the model when external requirements change. Second, the double Q-network model in reinforcement learning is used to train and update the policy network and the target network respectively, and the exploration mechanism is combined to select the optimal processing operation in different states, endowing the model with strong generalization ability and dynamic adaptability. In addition, a local forgetting mechanism is introduced to update the experience buffer pool, only replacing the expired experience samples in the local area of the new data and retaining the global experience, effectively reducing the interference of old data and enhancing the robustness and response efficiency of the model to environmental changes and process updates.

[0041] A more specific implementation of step (5) is as follows:

[0042] Combined with the adaptive adjustment strategy, according to different processing scenarios and objectives, by adjusting the weight parameters of processing quality, time, and energy consumption in the reward function, a dynamic reward mechanism is realized to adapt to environmental changes and process objective adjustments;

[0043] First, a dynamic reward function is designed to adapt to different process demand change scenarios. Multiple evaluation indicators such as processing quality, processing time, and processing energy consumption of the deep recurrent Q-network are introduced in the reward construction, and adjustable weight coefficients are assigned. Dynamic regulation is carried out by constructing a reward function in the following form:

[0044]

[0045] Among them, w i is the weight of each process target, and cost i represents the cost corresponding to the i-th index, such as energy consumption, time, or precision. The deep recurrent Q-network can flexibly switch and dynamically adapt the strategy by adjusting the weight coefficients without retraining the model, so as to meet the optimization requirements of different processing objectives such as energy conservation priority, high-quality priority, or high-efficiency priority;

[0046] Secondly, to cope with the dynamic changes in the processing environment, a double-network structure in reinforcement learning is used to construct a policy model. The policy network and the target network are respectively set, and the online learning mechanism is used to update the parameters of the policy network in real time;

[0047] At each time step, the Boltzmann exploration strategy is used to select the processing operation, and according to the state transition sample set <s, a, r, s'>, by minimizing the mean absolute error loss between the Q value calculated by the target network and the Q T value calculated by the policy network, the model parameters θ - of the policy network are updated using gradient descent. The loss function is calculated and defined as:

[0048]

[0049] Among them, γ represents the discount factor, which takes values in the interval (0, 1) and is used to balance the emphasis of the agent on immediate rewards and future rewards; represents the value with the maximum expected return among all possible actions taken in the next state s z ′, a z ′ represents the optimal action taken in the state s z ′; w z is the weight of s z ′. The parameters θ - of the target network and the parameters θ of the policy network are synchronized every K steps and remain unchanged at other time steps. During the training process, the current policy is a trade-off between the prediction results of reinforcement learning and random results;

[0050] Finally, to improve the response efficiency of the model to environmental changes, a local forgetting experience buffer pool update mechanism is introduced;

[0051] When the process conditions or external resources change, the agent adds the newly collected samples to the experience pool, and only replaces the expired samples in the local neighborhood of the new data instead of global replacement, so as to retain the experience data with reference value in the old environment;

[0052] This strategy allows the experience samples to maintain a uniform distribution in the state space, reduces the interference of outdated data on the model strategy, and enhances the robustness and adaptability to environmental changes.

[0053] Furthermore, the specific implementation of step (6) is as follows: introduce a selective forgetting mechanism to address the impact of process changes and machining environment changes on model performance. By calculating the Fisher information matrix of the model parameters on the retained data and the forgotten data, evaluate the importance of each parameter, and specifically suppress the parameters with a higher correlation with the change to achieve selective forgetting of specific old knowledge and avoid affecting the model's performance on other data. This mechanism can achieve rapid adaptation without retraining the model. When the machining resources change or the environment is adjusted, further combine the decreasing reinforcement learning and environmental perturbation methods to guide the agent to reduce the effectiveness of the original strategy in the new environment while retaining the learning effect of the original environment, so as to update the model strategy without losing the existing performance and achieve active adaptation to environmental changes and knowledge update.

[0054] A more specific implementation of step (6) is as follows:

[0055] Introduce a selective forgetting mechanism, evaluate the importance of model parameters by calculating the Fisher information matrix, and suppress the parameters related to the change characteristics to achieve the forgetting of specific information;

[0056] Process change forgetting: First, calculate the FIM of the training set D and the forgetting set D f respectively represented by [D] and [D f . Then traverse each parameter θ of the deep recurrent Q-network model i , and the parameter selection function is

[0057]

[0058] Suppress the contribution of the selected parameters to the multi-model output, and the suppression function is

[0059]

[0060] where α is a hyperparameter in the interval (0,1) used to control the strictness of parameter selection, which is the ratio of the diagonal elements of [D] and [D f . λ is another hyperparameter used to control the performance of protecting the retained set D r . After completing parameter selection and suppression, return the updated model

[0061] Environmental change forgetting: To address the policy failure caused by changes in machining resources or manufacturing environment perturbations, further introduce decreasing reinforcement learning and the environment poisoning mechanism;

[0062] First, establish the following forgetting loss function to guide the new policy to deliberately weaken the knowledge of the original environment

[0063]

[0064] The first term encourages the new policy π′ to work inadequately in the unlearned environment u; the second term incentivizes the new policy π′ to have the same performance as the existing policy π in other environments. The first term guides the agent to search for and try different policies to fully explore the state space of environment u; the second term motivates the agent to strategically modify the policy;

[0065] Subsequently, introduce the "environment poisoning" mechanism to guide and optimize the policy by perturbing the environment transition function. In each poisoning cycle i, define the reward function

[0066]

[0067] where π i and π′ represent the current policy and the updated policy respectively; π i (s i ) is the probability distribution of available actions in state s i under policy π i ; Δ(π i (s i )||π′ i (s i )) represents the difference between π i (s i ) and π′ i (s i ), and the KL divergence is used to measure it; the second term is used to calculate the reward for executing the states of all environments except for changing the environment state u in each poisoning period, and λ1 and λ2 are balance coefficients; by the collaborative work of the above two mechanisms, it is possible to effectively forget obsolete knowledge and retain useful policies without retraining the model, improving the stability and adaptability of the model in dynamic manufacturing scenarios.

[0068] Furthermore, the specific implementation method of step (7) is as follows: First, randomly select a part from the training set and initialize its processing state. In each training episode, use the Boltzmann exploration strategy to select the optimal action (processing step) according to the current state, process drawing, and energy consumption table. After executing this action, obtain the new state and the corresponding reward value (considering energy consumption and processing quality), form an experience sample and store it in the experience pool. By continuously sampling the samples in the experience pool, evaluate the mean squared error (MAE) between the Q value and the target Q value (QT), and optimize the policy network Q through gradient descent to make the model gradually converge. This training process is continuously iterated until all processing steps are completed.

[0069] A more specific implementation of step (7) is as follows:

[0070] Combining the Boltzmann exploration strategy and the dynamic forgetting mechanism, adaptively train the SADDQN model under process and environmental changes to optimize the part processing strategy;

[0071] Randomly select a part from the training set and initialize its processing state;

[0072] In each training episode, use the Boltzmann exploration strategy to select the optimal processing action, process or resource selection according to the current state. After execution, record the energy consumption and quality to calculate the reward value, and transfer to the next state, forming an experience and storing it in the experience pool;

[0073] Continuously sample the training samples in the experience pool, optimize the policy network Q by minimizing the mean square error between the Q value and the target Q value. During the training process, periodically update the parameters of the target network QT until the model converges;

[0074] In the face of process requirements or environmental changes, the SADDQN algorithm introduces a dynamic reward function and a forgetting strategy;

[0075] If the process changes, adjust the reward function and forget the obsolete knowledge; if the processing environment changes, combine the LOFO mechanism to update the experience buffer pool and adjust the strategy to make the model have good adaptability.

[0076] The present invention proposes a flexible process route planning method based on DRQN, that is, a deep reinforcement learning algorithm combining DQN and LSTM. Combining the structural advantages of LSTM, it mines the sequence characteristics in process data to improve the accuracy and stability of process route planning. With the powerful dynamic decision-making ability of DQN combined with the adaptive strategy, it solves the challenges brought by the dynamically changing process requirements and processing environment in the complex product assembly process. It proposes a "selective forgetting" mechanism. Compared with the existing retraining-based methods, it does not require retraining of unchanged data, greatly reducing the computational overhead and enabling faster response to the process route changes caused by process changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is a schematic diagram of the overall process of the flexible process route planning method for complex products based on DRQN of the present invention;

[0078] Figure 2 It is a schematic diagram of the horizontal comparison of the performance between the present invention and other existing traditional process route planning algorithms. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0079] In order to describe the present invention more specifically and clearly, the technical solutions of the present invention will be described in a comprehensive and detailed manner below in conjunction with the accompanying drawings and specific embodiments.

[0080] Example 1:

[0081] A flexible planning method for the process route of complex products based on DRQN, comprising the following steps:

[0082] (1) Model information such as process features, process operations, processing resources, and constraint matrices of complex product parts;

[0083] (2) Based on the partially observable Markov decision process (POMDP), model the process route of complex products, define the state vector, action space, state transition probability, observation space, observation probability, and reward function, and establish a process constraint matrix to describe the dependency relationship between process features;

[0084] (3) Construct a long short-term memory network (LSTM) structure to perform sequence modeling on the processing state and extract the time-dependent features hidden in the process route;

[0085] (4) Based on the LSTM, construct a deep recurrent Q-network (DRQN), and train the agent to select processing operations according to the current state to achieve the optimal decision of the process path;

[0086] (5) Combine the adaptive adjustment strategy, and according to different processing scenarios and goals, adjust the weight parameters of processing quality, time, and energy consumption in the reward function to implement a dynamic reward mechanism to adapt to environmental changes and process goal adjustments;

[0087] (6) Introduce a selective forgetting mechanism, evaluate the importance of model parameters by calculating the Fisher information matrix, and suppress the parameters related to the changed features to achieve the forgetting of specific information;

[0088] (7) Combine the Boltzmann exploration strategy and the dynamic forgetting mechanism to adaptively train the SADDQN model under process and environmental changes to optimize the part processing strategy.

[0089] A flexible planning system for the process route of complex products based on DRQN, which uses the above method.

[0090] Example 2:

[0091] A flexible planning method for the process route of complex products based on DRQN, comprising the following steps:

[0092] (1) Each complex product part is composed of several process features, each feature is completed by multiple processing operations, different operations correspond to specific processing resources, including a machine tool set and a tool set, and processing steps such as the feed direction. The process relationships such as dependency, adjacency, pattern, and datum between features are represented by a constraint matrix.

[0093] (2) Process route modeling based on POMDP, where: the state vector S represents the execution states of each processing operation, S i = 0 means to be executed, S i = 1 means completed; the action space is the set of optional processing operations OpS avl that conforms to the current state and constraint relationships; the state transition probability is 1 / |OpS avl |; the observation space includes the combination of processing resources and steps required for the current operation; the observation probability is calculated according to the number of combinations of processing resources and steps; the reward function is defined as R.

[0094] (3) By constructing a long short-term memory network, taking the current processing state vector and the hidden state at the previous moment as inputs, and processing them through the input gate, forget gate, and output gate in sequence to extract the time-dependent features in the process state sequence. Among them, the forget gate controls the retention degree of historical information, the input gate adjusts the update ratio of current information, and the output gate combines long-term memory to generate the current output state, so as to realize the modeling and expression of the temporal relationship in the process flow.

[0095] (4) Based on the temporal modeling ability of LSTM, construct a deep recurrent Q-network, take adjacent multiple processing states as inputs, extract time-dependent features through the LSTM network, and use them as the inputs of the deep network. During the training process, the agent selects processing operations by predicting the Q value of the current state and combining with the reinforcement learning strategy, so as to realize the optimal decision-making of the process path.

[0096] (5) By constructing an adaptive adjustment strategy, realize the rapid response to different processing scenarios and process objectives. First, set weight coefficients according to key factors such as processing energy consumption, processing time, and precision, and construct an adjustable reward function, so that the policy can be adjusted without retraining the model when external requirements change; second, adopt the double Q-network model in reinforcement learning to train and update the policy network and the target network respectively, and combine the exploration mechanism to select the optimal processing operation in different states, so that the model has strong generalization ability and dynamic adaptability; in addition, introduce a local forgetting mechanism to update the experience buffer pool, only replace the expired experience samples in the local area of the new data, and retain the global experience, effectively reducing the interference of old data and improving the robustness and response efficiency of the model to environmental changes and process updates.

[0097] (6) Introduce a selective forgetting mechanism to address the impact of process changes and machining environment variations on model performance. By calculating the Fisher information matrix of model parameters on the retained data and the forgotten data, evaluate the importance of each parameter, and specifically suppress the parameters with a higher correlation with the changes, achieving selective forgetting of specific old knowledge and avoiding affecting the model's performance on other data. This mechanism can achieve rapid adaptation without retraining the model. When the machining resources change or the environment is adjusted, further combine the decreasing reinforcement learning and environmental perturbation methods to guide the agent to reduce the effectiveness of the original strategy in the new environment while retaining the learning effect of the original environment, thereby updating the model strategy without losing the existing performance and achieving active adaptation to environmental changes and knowledge update.

[0098] (7) First, randomly select a part from the training set and initialize its machining state. In each training episode, use the Boltzmann exploration strategy to select the optimal action (machining step) according to the current state, process diagram, and energy consumption table. After executing this action, obtain the new state and the corresponding reward value (considering energy consumption and machining quality), form an experience sample, and store it in the experience pool. By continuously sampling the samples in the experience pool, evaluate the mean squared error (MAE) between the Q value and the target Q value (QT), and optimize the policy network Q through gradient descent to make the model gradually converge. This training process is iterated until all machining steps are completed.

[0099] Example 3:

[0100] An application of a flexible process route planning method for complex products based on DRQN

[0101] Taking the typical machining process in a flexible machining system as an example, set parts with multiple process features, and set attributes such as machining time and accuracy for each process feature. At the same time, define the constraint relationship between operations according to the actual process flow. The machining resources (such as machine tools, cutting tools) and energy consumption parameters are assumed and configured according to the actual manufacturing conditions, and the energy consumption is calculated in combination with the equipment operation state and the process execution process.

[0102] A typical application scenario of the present invention can be abstracted as: in the manufacturing process of parts with multiple process features, according to the constraint relationship between process features and the machining resource configuration, dynamically plan the machining path, and when the process features are added, deleted, or adjusted, it can quickly adapt to the changes through the reinforcement learning strategy and regenerate the optimal process route. At the same time, combined with the real-time changes of the machining resource state and the process target, the model has the ability to update online. Without retraining, it can respond to new process requirements and environmental perturbations through strategy adjustment and forgetting mechanism, thereby improving the flexibility and intelligence level of the manufacturing system.

[0103] The specific implementation process of the technical solution of the present invention is as follows Figure 1 shown, including the following steps:

[0104] (1) Model information such as process characteristics, process operations, processing resources, and constraint matrices of complex product parts.

[0105] Each part is composed of n process characteristics F, where F i is the i-th process characteristic; each F i is completed by a series of process operations Op i . Different operations need to match corresponding processing resources and processing steps, including machine set MS and tool set TS, as well as processing steps such as feed direction DS, etc. To represent the processing constraints between process characteristics, a constraint matrix CM is constructed, where the element in the i-th row and j-th column of CM represents the process relationship between characteristic F i and F j . The process relationships include dependency relationship (DR), adjacency relationship (AR), pattern relationship (PR), and datum relationship (TR), which are used to constrain the process sequence and path.

[0106] (2) Model the process route of complex products based on the partially observable Markov decision process (POMDP).

[0107] Define the state vector, action space, state transition probability, observation space, observation probability, and reward function, and establish a process constraint matrix to describe the dependency relationship between process characteristics, as defined below:

[0108] State vector: Map the execution status of part processing operations to the state vector S, defined as S = [S1, S2,... S i ,... S m , where m is the total number of operations of the part, and S i represents the processing status of operation Op i . The value of S i has the following two cases:

[0109] S i = 0: This operation can be executed;

[0110] S i = 1: This operation has been completed;

[0111] Action space: According to the current processing status s t of part P, under the constraint of the process characteristics, the set of executable operations a t at time t, OpS avl = {Op i ∈ OpS ∧ s t + Op i~CM(P)}, where + indicates performing a certain processing operation in a certain state; ~ indicates compliance with the constraint relationship;

[0112] Transition probability: The next processing operation is a possible operation selected based on the constraints between the current state and process characteristics. During the state transition process, the state transition probability is 1 / |OpS avl |. Among them, |OpS avl | represents the number of executable operations;

[0113] Observation space: The currently selected action is completed through a series of steps such as the selected processing resources and processing steps. For example, the available machine tools, available cutting tools, and feed direction for a certain processed part;

[0114] Observation probability: According to the currently selected action Op∈OpSavl, the corresponding available processing resources (such as machine tools and cutting tools) and processing steps (feed direction) are combined. During the state transition process, the state transition probability is defined as:

[0115] 1 / (|MS(Op i )|·|TS(Op i )|·|DS(Op i )|)

[0116] Among them, MS(Op) and TS(Op) respectively represent the number of available processing resources and the number of processing steps corresponding to the execution of operation Op;

[0117] Reward function: After the agent selects an action, according to the change in the part processing state, in order to obtain the optimal process route, the reward function R is defined as:

[0118] R = wq·R a -w t ·R t -w e ·R e

[0119] Among them, R represents the reward obtained when the part state is S t and the action a t is executed; R q 、R t 、R e respectively represent the processing quality, processing time, and processing energy consumption required for selecting the action a t i.e., executing the corresponding operation; w q 、w t 、w e are normalized weight coefficients, and their values can be determined according to process requirements.

[0120] (3) Construct a Long Short-Term Memory (LSTM) structure to perform sequence modeling on the processing state and extract the time-dependent features implicit in the process route.

[0121] Input the processing state information s at time t t and the output h of the Long Short-Term Memory network at the previous time t-1 , and pass them through a linear combination into three gates (input gate, forget gate, output gate) for processing.

[0122] Compress the linear combination of s t and h t-1 to the interval (0, 1) through the sigmoid activation function to obtain the forgetting factor f. t ;

[0123] Multiply the long-term memory c t-1 by the forgetting factor f t , which determines the forgetting ratio of the long-term memory c t-1 at the previous time;

[0124] Multiply the memory factor by the short-term memory at the current time t, and then add it to c t-1 that has completed the forgetting process to obtain the new long-term memory c t ;

[0125] Calculate the output factor out through the sigmoid activation function t , and use the tanh activation function for the new long-term memory C t ;

[0126] Multiply the output factor out t by the long-term memory c t processed by the tanh activation function to obtain the result output h t .

[0127] (4) Build a Deep Recurrent Q-Network (DRQN) based on LSTM to train the agent to select processing operations according to the current state and achieve the optimal decision-making of the process route.

[0128] Based on the ability of LSTM to perform sequence modeling on the processing state sequence, use its output as the input feature of the Q-network to estimate the action value of each processing operation in the current state;

[0129] By constructing a double Q-network structure as the policy network and the target network respectively, optimize the policy using the experience replay mechanism during the training process, and combine the exploration strategy to select actions in the state space, enabling the agent to achieve adaptive optimization and decision-making capabilities for the processing path in a dynamic manufacturing environment.

[0130] (5) Combine with the adaptive adjustment strategy. According to different processing scenarios and objectives, by adjusting the weight parameters of processing quality, time, and energy consumption in the reward function, a dynamic reward mechanism is realized to adapt to environmental changes and process objective adjustments.

[0131] First, design a dynamic reward function to adapt to different process requirement change scenarios. Introduce multiple evaluation indicators such as processing quality, processing time, and processing energy consumption in the reward construction, and assign adjustable weight coefficients. Perform dynamic regulation by constructing a reward function in the following form:

[0132]

[0133] Among them, w i is the weight of each process objective, and cost i represents the cost corresponding to the i-th indicator, such as energy consumption, time, or precision, etc. By adjusting the weight coefficients, flexible switching and dynamic adaptation of the strategy can be achieved without retraining the model, so as to meet the optimization requirements of different processing objectives such as energy conservation priority, high-quality priority, or high-efficiency priority.

[0134] Secondly, to cope with the dynamic changes in the processing environment, adopt a double-network structure in reinforcement learning to construct a policy model. Set the policy network and the target network respectively, and use the online learning mechanism to update the parameters of the policy network in real time.

[0135] At each time step, use the Boltzmann exploration strategy to select processing operations, and according to the state transition sample set <s, a, r, s'>, by minimizing the mean absolute error loss between the Q value calculated by the target network and the Q T value calculated by the policy network, use gradient descent to update the model parameters θ - of the policy network. The loss function calculation is defined as:

[0136]

[0137] Among them, γ represents the discount factor, which takes values in the interval (0, 1) and is used to balance the emphasis of the agent on immediate rewards and future rewards;

[0138] represents the maximum value of the expected return among all possible actions taken in the next state s z ′, and a z ′ represents the optimal action taken in the state s z ′;

[0139] w z is the weight of s z ′. Synchronize the parameters θ - of the target network every K steps.with the policy network parameters θ and remain unchanged at other time steps. During the training process, the current policy is a trade-off between the reinforcement learning prediction results and random results.

[0140] Finally, to improve the response efficiency of the model to environmental changes, a local forgetting (LOFO) experience buffer pool update mechanism is introduced.

[0141] When the process conditions or external resources change, the agent adds the newly collected samples to the experience pool and only replaces the expired samples in the local neighborhood of the new data instead of global replacement, thus retaining the experience data that is still valuable under the old environment.

[0142] This policy allows the experience samples to maintain a uniform distribution in the state space, reduces the interference of outdated data on the model policy, and enhances the robustness and adaptability to environmental changes.

[0143] (6) Introduce a selective forgetting mechanism, evaluate the importance of model parameters by calculating the Fisher information matrix, and suppress the parameters related to the changed features to achieve the forgetting of specific information.

[0144] First, calculate the FIM of the training set D and the forgetting set D f denoted by [D] and [D f respectively; then traverse each parameter θ i of the DRQN model, and the parameter selection function is:

[0145]

[0146] Suppress the contribution of the selected parameters to the multi-model output, and the suppression function is

[0147]

[0148] where α is a hyperparameter in the interval (0,1) used to control the strictness of parameter selection, which is the ratio of the diagonal elements of [D] and [D f ; the parameter λ is another hyperparameter used to control the performance of the protected retention set D r . After completing parameter selection and suppression, return the updated model

[0149] Forgetting in a changing environment: To address the policy failure caused by changes in processing resources or manufacturing environment disturbances, a decreasing reinforcement learning and environment poisoning mechanism is further introduced.

[0150] First, establish the following forgetting loss function to guide the new policy to deliberately weaken the knowledge of the original environment:

[0151]

[0152] Item 1 encourages that the new policy π does not work well in the unlearned environment u, and Item 2 motivates that the new policy π′ has the same performance as the existing policy π in other environments. Item 1 guides the agent to search for and try different policies to fully explore the state space of environment u. Item 2 encourages the agent to strategically modify the policy.

[0153] Subsequently, the "environment poisoning" mechanism is introduced to guide and optimize the policy by perturbing the environment transition function. In each poisoning cycle i, the reward function is defined as follows:

[0154]

[0155] where π i and π′ represent the current policy and the updated policy respectively; π i (s i ) is the probability distribution of available actions in state s i under policy π i ; Δ(π i (s i )||π′ i (s i )) represents the difference between π i (s i ) and π′ i (s i ), and the KL divergence is used to measure it;

[0156] The second term is used to calculate the reward for executing the states of all environments except for changing the environment state u in each poisoning period; λ1 and λ2 are balance coefficients.

[0157] By the collaborative work of the above two mechanisms, without retraining the model, it is possible to effectively forget obsolete knowledge and retain useful policies, improving the stability and adaptability of the model in dynamic manufacturing scenarios.

[0158] (7) Combine the Boltzmann exploration strategy and the dynamic forgetting mechanism to adaptively train the SADDQN model under process and environment changes to optimize the part processing strategy.

[0159] Randomly select a part from the training set and initialize its processing state;

[0160] In each training episode, use the Boltzmann exploration strategy to select the optimal processing action (such as operation or resource selection) according to the current state, record the energy consumption and quality after execution to calculate the reward value, and transfer to the next state, forming an experience and storing it in the experience pool;

[0161] Continuously sample the training samples in the experience pool, and optimize the policy network Q by minimizing the mean absolute error (MAE) between the Q value and the target Q value (QT). During the training process, the parameters of the target network QT are updated periodically until the model converges;

[0162] In the face of process requirements or environmental changes, the SADDRQN algorithm introduces a dynamic reward function and a forgetting strategy.

[0163] If a process change occurs, the reward function is adjusted and the obsolete knowledge is forgotten; if the processing environment changes, the experience buffer pool is updated in combination with the LOFO mechanism and the strategy is adjusted, so that the model has good adaptability.

[0164] The experimental results of the present invention were carried out on the Ubuntu virtual machine in the Database Technology Research Office of Shenyang Aerospace University. In this experiment, a model part in actual processing was used as the experimental object, and the time-consuming and precision attribute values were supplemented according to the actual situation.

[0165] The experimental effects of the present invention were simultaneously compared with the effects obtained by four other algorithms, including the improved DPPO algorithm (GDPPO), genetic algorithm (GA), simulated annealing algorithm (SA), and ant colony algorithm (ACO).

[0166] Figure 2 It is a comparison of the effects of two cases of the present invention and several other algorithms in the process environment change experiment. Among them, the number of cycles represents the number of cycles required for a certain algorithm to obtain the optimal process route. The fewer the number of cycles, the better the effect. The energy consumption represents the actual energy consumption required for the optimal process route obtained by a certain algorithm. The lower the energy consumption, the better the effect.

[0167] The above description of the embodiments is for those of ordinary skill in the art in this technical field to understand and apply the present invention. It is obvious that those who are familiar with the technology in this field can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by those skilled in the art to the present invention according to the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. A flexible planning method for complex product process routes based on deep recurrent Q network, comprising the following steps: (1) Modeling the process characteristics, process operations, processing resources, and constraint matrix of complex product parts through deep recurrent Q network information; (2) Model the process route of complex products based on partially observable Markov decision processes, define the state vector, action space, state transition probability, observation space, observation probability and reward function, and establish a process constraint matrix to describe the dependencies between process features; (3) Construct a long short-term memory network structure to model the processing state sequence and extract the time-dependent features implicit in the process route; (4) Building a deep recurrent Q network based on the long short-term memory network, training the intelligent agent to select processing operations based on the current state and achieve the optimal decision of the process path; (5) Combined with the adaptive adjustment strategy, according to different processing scenarios and goals, the weight parameters of processing quality, time and energy consumption in the reward function are adjusted to realize a dynamic reward mechanism to adapt to environmental changes and process goal adjustments; (6) Introducing a selective forgetting mechanism, evaluating the importance of model parameters by calculating the Fisher information matrix and suppressing parameters related to the changed features to achieve forgetting of specific information; (7) Combining the Boltzmann exploration strategy and dynamic forgetting mechanism, the SADDQN model is adaptively trained to optimize the part processing strategy under process and environment changes.

2. The flexible planning method according to claim 1, characterized in that: The specific implementation of step (1) is as follows: each complex product part is composed of a number of process features, each feature is completed by a plurality of processing operations, different operations correspond to specific processing resources and processing steps, the processing resources include a machine tool set and a tool set, the processing steps include a feed direction, and the dependency, adjacency, mode and reference deep recurrent Q network process relationships between features are represented by a constraint matrix; Preferably, a more specific implementation of step (1) is: Model the process characteristics, process operations, processing resources, and constraint matrix deep loop Q network information of complex product parts. Each part consists of n process features F, where F i is the i-th process feature; each F i A series of processing operations Op i To complete, different operations need to match corresponding processing resources and processing steps, including machine set MS and tool set TS, as well as processing steps such as feed direction DS. Deep recurrent Q network, to represent the processing constraints between process features, construct the constraint matrix CM, where the i-th row and j-th column of CM represent the feature F i With F j The process relationships between them include dependency relationship DR, adjacent relationship AR, pattern relationship PR and reference relationship TR, which are used to constrain the process sequence and path.

3. The flexible planning method according to claim 1, characterized in that: The specific implementation of step (2) is as follows: process route modeling based on partially observable Markov decision process, where: the state vector S represents the execution state of each processing operation, S i =0 means to be executed, S i =1 means completed; the action space is the set of optional processing operations OpS that conform to the current state and constraint relationship avl ; The state transition probability is 1 / |OpS avl |; The observation space includes the processing resources and step combinations required for the current operation; The observation probability is calculated according to the number of processing resources and step combinations; The reward function is defined as R; Preferably, a more specific implementation of step (2) is: Based on the partially observable Markov decision process, the process route of complex products is modeled, the state vector, action space, state transition probability, observation space, observation probability and reward function are defined, and the process constraint matrix is ​​established to describe the dependency relationship between process features, which is defined as follows: State vector: Map the execution status of the part processing operation into the state vector S, which is defined as S = [S1, S2, ... S i ,...S m ], m is the total number of operations of the part, S i Indicates the operation Op i Processing status, S i The value is divided into the following two cases: S i =0: the operation can be performed; S i =1: The operation has been completed; Action space: According to the current processing state s of part P t , under the constraints between process features, the optional execution operation a at time t t Collection OpS avl = {Op i ∈OpS∧s t +Op i ~CM(P)}, where + means a certain processing operation is performed under a certain state; ~ means compliance with the constraint relationship; Transition probability: The next processing operation is a possible operation selected based on the constraints between the current state and the process characteristics. During the state transition process, the state transition probability is 1 / |OpS avl |, where |OpS avl | indicates executable operands; Observation space: The currently selected action is completed through a series of steps of the selected processing resources and processing steps deep recurrent Q network, the available machine tools and available tools for a certain processing part, and the feed direction; Observation probability: According to the currently selected action Op∈OpSavl, the corresponding available processing resources and processing steps are combined; during the state transition process, the state transition probability is defined as: 1 / (|MS(On i )|·|TS(On i )||DS(On i )|) Wherein, MS(Op) and TS(Op) represent the number of available processing resources and the number of processing steps corresponding to the execution of the operation Op, respectively; Reward function: After the agent selects an action, according to the change of the part processing state, in order to obtain the optimal process route, the reward function R is defined as: R=w q R q -w t ·R t -w e ·R e Among them, R means that when the part state is S t When , execute action a t Rewards received; R q , R t , R e Respectively represent the selection action a t That is, the processing quality, processing time, and processing energy consumption required to perform the corresponding operation; w q 、w t 、w e It is the normalized weight coefficient, and its value can be determined according to the process requirements.

4. The flexible planning method according to claim 1, characterized in that: The specific implementation method of step (3) is as follows: by constructing a long short-term memory network, taking the current processing state vector and the hidden state at the previous moment as input, and sequentially processing through an input gate, a forget gate, and an output gate, extracting the time-dependent features in the process state sequence, wherein the forget gate controls the degree of retention of historical information, the input gate adjusts the update ratio of current information, and the output gate generates the current output state in combination with the long-term memory, thereby realizing the modeling and expression of the temporal relationship in the process flow; Preferably, a more specific implementation of step (3) is: Construct a long short-term memory network structure, perform sequence modeling on the processing state, extract the time-dependent features implicit in the process route, and convert the processing state information s input at time t into t The output h of the LSTM network at the previous moment t-1 , is passed into three gates for processing through linear combination, including input gate, forget gate, and output gate, and s is activated by the sigmoid function. t With h t-1 The linear combination of is compressed to the interval (0,1) to obtain the forgetting factor f t , the long-term memory c t-1 Multiply by the forgetting factor f t , determines the long-term memory c of the previous moment t-1 The forgetting ratio; multiply the memory factor by the short-term memory at the current time t, and then multiply it by c to complete the forgetting process t-1 Add together to get the new long-term memory c t ; Calculate the output factor out through the sigmoid activation function t , for the new long-term memory C t Using the tanh activation function, the output factor out t Compared with the long-term memory c after processing with the tanh activation function t Multiply and get the result output h t .

5. The flexible planning method according to claim 1, characterized in that: The specific implementation method of step (4) is as follows: based on the time series modeling capability of the long short-term memory network, a deep recurrent Q network is constructed, multiple adjacent processing states are used as input, time-dependent features are extracted through the long short-term memory network, and used as input of the deep network. During the training process, the intelligent agent predicts the Q value of the current state and selects the processing operation in combination with the reinforcement learning strategy to achieve the optimal decision on the process path; Preferably, a more specific implementation of step (4) is: A deep recurrent Q network is constructed based on the long short-term memory network to train the agent to select processing operations according to the current state and achieve the optimal decision of the process path. Based on the long short-term memory network's ability to model the processing state sequence, its output is used as the input feature of the Q network to estimate the action value of each processing operation in the current state. By constructing a dual Q network structure as the policy network and the target network respectively, the experience replay mechanism is used to optimize the strategy during the training process, and the exploration strategy is combined to select actions in the state space, so that the intelligent agent can achieve adaptive optimization and decision-making capabilities of the processing path in a dynamic manufacturing environment.

6. The flexible planning method according to claim 1, characterized in that: The specific implementation method of step (5) is: by constructing an adaptive adjustment strategy, a rapid response to different processing scenarios and process objectives is achieved; First, the weight coefficients are set according to the key factors of the deep recurrent Q network, such as processing energy consumption, processing time and accuracy, to build a flexibly adjustable reward function, so that the strategy can be adjusted without retraining the model when external demands change; Secondly, the double Q network model in reinforcement learning is used to train and update the policy network and the target network respectively, and the exploration mechanism is combined to select the optimal processing operation under different states, so that the model has strong generalization ability and dynamic adaptability; In addition, a local forgetting mechanism is introduced to update the experience buffer pool, replacing only expired experience samples in the local area of ​​new data and retaining the rest of the global experience, effectively reducing the interference of old data and improving the robustness and response efficiency of the model to environmental changes and process updates; Preferably, a more specific implementation of step (5) is: Combined with the adaptive adjustment strategy, according to different processing scenarios and goals, the weight parameters of processing quality, time and energy consumption in the reward function are adjusted to realize a dynamic reward mechanism to adapt to environmental changes and process goal adjustments; First, a dynamic reward function is designed to adapt to different scenarios of process demand changes. Multiple evaluation indicators such as processing quality, processing time, and processing energy consumption deep cycle Q network are introduced into the reward construction, and adjustable weight coefficients are assigned. Dynamic regulation is performed by constructing a reward function in the following form: Among them, w i is the weight of each process target, cost i Represents the cost corresponding to the i-th indicator, such as energy consumption, time or accuracy. By adjusting the weight coefficient, the strategy can be flexibly switched and dynamically adapted without retraining the model, thereby meeting the optimization requirements of different processing goals of deep recurrent Q networks with energy saving priority, high quality priority or high efficiency priority. Secondly, in order to cope with the dynamic changes of the processing environment, the dual network structure in reinforcement learning is used to build a strategy model, setting the strategy network and the target network separately, and using the online learning mechanism to update the strategy network parameters in real time; At each time step, the Boltzmann exploration strategy is used to select the processing operation and the sample set is transferred according to the state<s,a,r,s'> , by minimizing the Q value calculated by the target network and the Q value calculated by the policy network T The mean absolute error loss between the values ​​is used to update the model parameters θ of the policy network using gradient descent. - , the loss function calculation is defined as: Among them, γ represents the discount factor, which takes values ​​in the interval (0,1) and is used to balance the agent's emphasis on immediate rewards and future rewards; Indicates that in the next state s z Among all possible actions taken under ′, the value with the maximum expected return is a z ′ means in state s z The optimal action to take under ′; w z Yes z ′, and synchronize the target network parameters θ every K steps - The policy network parameters θ remain unchanged at other time steps. During training, the current policy is a trade-off between the reinforcement learning prediction results and the random results. Finally, in order to improve the model's response efficiency to environmental changes, a local forgotten experience buffer pool update mechanism is introduced; When process conditions or external resources change, the agent adds newly collected samples to the experience pool and replaces only expired samples in the local neighborhood of the new data, rather than replacing them globally, thereby retaining the experience data that is still valuable for reference in the old environment. This strategy allows experience samples to remain evenly distributed in the state space, reduces the interference of outdated data on the model strategy, and improves robustness and adaptability to environmental changes.

7. The flexible planning method according to claim 1, characterized in that: The specific implementation method of step (6) is as follows: introducing a selective forgetting mechanism to cope with the impact of process changes and processing environment changes on model performance, by calculating the Fisher information matrix of model parameters on retained data and forgotten data, evaluating the importance of each parameter, and suppressing parameters with high correlation with the change in a targeted manner, thereby achieving selective forgetting of specific old knowledge and avoiding affecting the performance of the model on other data. This mechanism can achieve rapid adaptation without retraining the model; When processing resources change or the environment is adjusted, the agent is guided to reduce the effectiveness of the original strategy in the new environment by combining decremental reinforcement learning with environmental perturbation methods, while retaining the learning effect of the original environment. This allows the agent to update the model strategy without losing existing performance, thus achieving active adaptation to environmental changes and knowledge updating. Preferably, a more specific implementation of step (6) is: A selective forgetting mechanism is introduced to evaluate the importance of model parameters by calculating the Fisher information matrix and suppress the parameters related to the changed features to achieve the forgetting of specific information. Change process forgetting: First calculate the training set D and the forgetting set D f The FIM is represented by [D] and [D f ] represents; then traverse each parameter θ of the deep recurrent Q network model i , the parameter selection function is Suppress the contribution of the selected parameters to the multi-model output. The suppression function is Among them, α is a hyperparameter on the (0,1) interval, which is used to control the strictness of parameter selection, and [D] and [D f ] is the ratio of the diagonal elements; λ is another hyperparameter used to control the protection of the reserved set D r After completing parameter selection and suppression, the updated model is returned. Change environment forgetting: In order to deal with the failure of strategies caused by changes in processing resources or disturbances in the manufacturing environment, we further introduce decreasing reinforcement learning and environmental poisoning mechanisms; First, the following forgetting loss function is established to guide the new strategy to intentionally weaken the knowledge of the original environment The first item encourages the new policy π′ to work insufficiently in the unlearned environment u; the second item encourages the new policy π′ to have the same performance as the existing policy π in other environments. The first item guides the agent to search and try different strategies to fully explore the state space of the environment u; the second item encourages the agent to modify the strategy strategically; Subsequently, the "environment poisoning" mechanism is introduced to guide the optimization of the strategy by perturbing the environment transfer function. In each poisoning cycle i, the reward function is defined Among them, π i and π′ represent the current strategy and the updated strategy respectively; π i (s i ) is in strategy π i Next state i The probability distribution of available actions in ; Δ(π i (s i )||π′ i (s i )) represents π i (s i ) and π′ i (s i ), measured using KL divergence; the second term is used to calculate the reward for executing all the states of the environment except the changed environment state u in each poisoning period, λ1 and λ2 are the balance coefficients; Through the collaborative work of the above two mechanisms, it is possible to effectively forget outdated knowledge and retain useful strategies without retraining the model, thereby improving the stability and adaptability of the model in dynamic manufacturing scenarios.

8. The flexible planning method according to claim 1, characterized in that: The specific implementation method of step (7) is as follows: first, a part is randomly selected from the training set and its processing state is initialized. In each training round, the Boltzmann exploration strategy is used to select the optimal action according to the current state, process diagram and energy consumption table. After executing the action, a new state and a corresponding reward value are obtained, and an experience sample is formed and stored in the experience pool. By continuously sampling samples in the experience pool, the mean square error between the Q value and the target Q value is evaluated, and the strategy network Q is optimized by gradient descent to gradually converge the model. The training process is continuously iterated until all processing steps are completed; Preferably, a more specific implementation of step (7) is: Combining the Boltzmann exploration strategy and dynamic forgetting mechanism, the SADDQN model is adaptively trained to optimize the part processing strategy under process and environment changes; Randomly select a part from the training set and initialize its processing state; In each training round, the Boltzmann exploration strategy is used to select the optimal processing action, process or resource selection according to the current state. After execution, the energy consumption and quality are recorded to calculate the reward value, and then transferred to the next state to form an experience and store it in the experience pool; Continuously sample training samples from the experience pool, optimize the policy network Q by minimizing the mean square error between the Q value and the target Q value, and periodically update the target network QT parameters during training until the model converges; In the face of process requirements or environmental changes, the SADDQN algorithm introduces a dynamic reward function and forgetting strategy; If the process changes, the reward function is adjusted and outdated knowledge is forgotten; if the processing environment changes, the experience buffer pool is updated and the strategy is adjusted in combination with the LOFO mechanism, so that the model has good adaptive capabilities.

9. A complex product process route flexible planning system based on deep cyclic Q network, characterized in that: Use the flexible planning method described in any one of claims 1 to 8.

10. An application of a complex product process route flexible planning method based on a deep recurrent Q network according to any one of claims 1 to 8, characterized in that: Flexible planning methods are used to make dynamic decisions on processing status and realize intelligent planning and rapid adaptation of process paths.

Citation Information

Cited By

  • Vertical field large model privacy forgetting method and terminal based on double Fisher matrix guidance

    CN121389189A

  • A privacy-forgetting method for large-scale vertical domain models guided by dual Fisher matrices, and its application in terminals.

    CN121389189B

  • An aero-engine mixed-line pulsating assembly adaptive scheduling method and system

    CN122546947A

  • An aero-engine mixed-line pulsating assembly adaptive scheduling method and system

    CN122546947B