Intelligent decision analysis method and system of manufacturing system based on digital twinning
By using meta-learning, adversarial training and generative data augmentation technologies in the digital twin manufacturing system, the problems of scarcity and distribution differences in emerging markets are solved, the adaptability and prediction accuracy of the model are improved, and dynamic optimization of production strategies is achieved.
Patent Information
- Application Number
- CN202510687618.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The number of target domain data in emerging markets is small and has a huge difference from the source domain distribution, resulting in negative migration of the digital twin model on the test set, with the prediction error increasing by 25%.
By extracting multiple sets of simulation tasks from the source domain data, a basic model with fast adaptability is trained, the source domain and target domain features are aligned with adversarial training strategies, and the conditional diffusion model is used to generate synthetic data to build an enhanced data set. Finally, the multi-objective deep reinforcement learning algorithm is run in the digital twin and the scheduling strategy is dynamically adjusted.
It has improved the generalization ability and adaptability of the digital twin model in small sample scenarios in emerging markets, alleviated the prediction distortion problem caused by data scarcity, and realized dynamic optimization and feasibility verification of production strategies.
Smart Images

Figure CN120218679A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital twins, and more specifically, to an intelligent decision-making analysis method and system for a manufacturing system based on digital twins. Background Art
[0002] With the exponential growth of the global e-commerce market scale, multinational manufacturing enterprises are facing numerous challenges. Due to defects such as data islands, static models, and decision-making delays in traditional production scheduling systems, it is difficult to meet the cross-border e-commerce requirements of "fast order fulfillment and precise supply".
[0003] Especially when exploring emerging markets, due to the small number of target domain data and the extremely large distribution difference from the source domain, the domain adversarial adaptive loss in the digital twin model will lead to gradient conflicts due to insufficient samples, and the model after migration will show negative transfer on the test set, with the prediction error increasing by 25% instead. Summary of the Invention
[0004] The present invention provides an intelligent decision-making analysis method and system for a manufacturing system based on digital twins, which solves the technical problems in the related art that the number of target domain data in emerging markets is small and the distribution difference from the source domain is extremely large.
[0005] The present invention provides an intelligent decision-making analysis method for a manufacturing system based on digital twins, including the following steps: S100, by extracting multiple groups of simulated tasks from source domain data, training a base model with fast adaptation ability. The base model first fine-tunes the parameters on the support set of each small task, and then reversely updates the initial parameters according to the performance on the query set of all tasks, and finally obtains a meta-knowledge base that can quickly adapt to the new market; S200, on the target domain small samples, adopting an adversarial training strategy to align the source domain and target domain features. The feature extractor is forced to learn domain-invariant representations, while the domain classifier tries to distinguish the data sources, and at the same time dynamically adjusts the weights of the prediction loss and the adversarial loss to prevent negative transfer caused by overfitting of small samples; S300, based on the scarce target domain data and the domain knowledge graph, using a conditional diffusion model to generate a large number of synthetic data, and constructing an enhanced data set through distribution consistency testing and pre-trained model screening; S400, loading the enhanced data and the adapted model in the digital twin, running a multi-objective deep reinforcement learning algorithm, dynamically adjusting the weights of the order fulfillment rate and the inventory cost in combination with real-time feedback, and finally selecting the optimal scheduling strategy from the Pareto optimal solution set.
[0006] Further, in S100, it includes the following steps: S110, meta-task construction: sampling meta-tasks from the source domain data, and each task contains a support set and a query set; S120, Internal parameter adaptation: For each meta-task, calculate the support set loss and update it based on the current meta-parameters; S130, External meta-parameter optimization: Evaluate the performance of the adapted model on the query sets of all tasks and update the meta-parameters through backpropagation; S140, Convergence determination: Repeat S110 - S130 until the stopping condition is met.
[0007] Furthermore, in S140, the stopping condition is as follows: ; where is the convergence threshold, is the maximum number of iterations, , is the current iteration number, is the meta-loss value at the -th iteration, is the meta-loss value at the previous iteration.
[0008] Furthermore, in S200, the following steps are included: S210, Model initialization: Load the meta-pre-trained parameters as the initial parameters; S220, Domain adversarial feature alignment: Achieve domain invariance in the feature space through the gradient reversal layer; S230, Joint optimization: Synchronously optimize the prediction task loss and the domain adversarial loss; S240, Parameter update: Alternately update the parameters of the feature extractor and the domain classifier: S250, Few-shot early stopping strategy: Monitor the loss of the target domain validation set and terminate the training when there is no decrease for 5 consecutive iterations.
[0009] Furthermore, in S240, the calculation formula for parameter update is as follows: Fix and update simultaneously: ; Fix and update simultaneously: ; where are the parameters of the feature extractor, are the parameters of the domain classifier, is the learning rate for adversarial learning, , is the gradient with respect to the parameter , is the gradient with respect to the parameter The gradient, is the loss function of domain adversarial, is the total loss function.
[0010] Furthermore, the calculation formula of the few-shot early stopping strategy is as follows: ; where, is the loss value on the target domain validation set, is the current training round, represents the training termination signal, which is sent when the early stopping condition is triggered, represents taking the minimum value among all historical rounds less than the current round t, represents the historical training round index, which is used to compare the current loss with the historical optimal loss.
[0011] Furthermore, in S300, it includes the following steps: S310, Domain knowledge conditional encoding: Encode the target domain prior knowledge into a conditional vector; S320, Conditional diffusion model training: Train the denoising diffusion probability model with the target domain samples and conditions as inputs; S330, Diversity controllable sample generation: Sample and generate augmented data from the diffusion model, and constrain it to be consistent with the target domain distribution; S340, Generated sample quality verification: Screen valid samples through a pre-trained validator; S350, Augmented dataset construction: Merge real and synthetic data to construct the final training set.
[0012] Furthermore, in S400, it includes the following steps: S410, Environment initialization: Load the virtual factory model and inject augmented data ; S420, Multi-objective deep reinforcement learning policy training: Construct an Actor-Critic architecture and optimize three objectives: Actor network : Input the state and output the action ; Critic network : Evaluate the action value, including three output heads: ; Objective function: ; where, is the Actr network with parameters , is the Critic network with parameters , is the state vector, is the action vector, is the state distribution generated by the behavioral policy, is the variance penalty coefficient, is the variance of the action, is the cost evaluation value, is the quality of service evaluation value, is the feasibility evaluation value, represents the objective function, is the expectation operator, represents the variance of the action; S430, dynamic weight adjustment mechanism: adjust the objective weights according to the real-time feedback of the digital twin; S440, virtual-physical closed-loop verification: execute the policy in the digital twin to verify the feasibility; S450, Pareto front policy selection: select the optimal policy from the non-dominated solution set.
[0013] Furthermore, the calculation formula of the optimal policy is as follows: ; where, is the optimal policy, is the Pareto optimal policy set, is the weight of the th objective, is the Q value of the th objective, is the feasibility score of the policy in the state ; The output of the optimal policy is as follows: ; where is the action of the production line start / stop decision, is the resource allocation ratio, is the production capacity allocation weight.
[0014] The present invention also proposes an intelligent decision-making analysis system for a manufacturing system based on digital twins, which executes the steps in the foregoing intelligent decision-making analysis method for a manufacturing system based on digital twins, including: Meta-learning base model construction module: automatically construct multiple groups of simulation tasks from the historical data of the mature market, and obtain a meta-learning model with fast adaptation ability through collaborative training of internal and external loops, so that it can quickly capture the common laws across markets under a small amount of target data; Adversarial Domain Adaptation Module: Introduce an adversarial training mechanism on a very small amount of data in the target domain, dynamically balance the prediction accuracy and domain-invariant feature learning, and eliminate the distribution differences between the source domain and the target domain through the gradient reversal and weight adaptation strategies; Data Augmentation Generation Module: Generate synthetic data that conforms to the characteristics of the target domain based on the conditional diffusion model and the domain knowledge graph, and combine distribution consistency verification and pre-training model screening to expand high-quality training samples; Scheduling Optimization Decision Module: Simulate multi-objective scheduling strategies in the digital twin environment, dynamically optimize the order fulfillment rate, inventory cost, and feasibility score through reinforcement learning, and finally output a flexible production plan verified by virtualization.
[0015] The beneficial effects of the present invention are as follows: Through the deep integration of meta-learning and adversarial training, the present invention improves the generalization ability and adaptability of the digital twin model in the small-sample scenario of emerging markets, alleviates the prediction distortion problem caused by data scarcity, combines generative augmentation technology to expand high-quality training samples, and synchronously constructs a multi-objective flexible scheduling mechanism to realize the dynamic optimization and feasibility verification of production strategies in the digital twin environment, and finally forms a closed-loop optimized intelligent decision-making system. Brief Description of the Drawings
[0016] Figure 1 is a flowchart of an intelligent decision-making analysis method for a digital twin-based manufacturing system proposed by the present invention; Figure 2 is of the present invention Figure 1 flowchart of the sub-steps of S100; Figure 3 is of the present invention Figure 1 flowchart of the sub-steps of S200; Figure 4 is of the present invention Figure 1 flowchart of the sub-steps of S300; Figure 5 is of the present invention Figure 1 flowchart of the sub-steps of S400; Figure 6 is a structural block diagram of an intelligent decision-making analysis system for a digital twin-based manufacturing system proposed by the present invention.
[0017] In the figure: 101, Meta-Learning Base Model Construction Module; 102, Adversarial Domain Adaptation Module; 103, Data Augmentation Generation Module; 104, Scheduling Optimization Decision Module. Detailed Embodiments
[0018] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and thus implement the subject matter described herein, and that changes may be made to the functions and arrangements of the elements discussed without departing from the scope of protection of the content of this specification. Each example may omit, substitute, or add various processes or components as needed. Additionally, the features described for some examples may be combined in other examples.
[0019] As Figures 1-5 shown, an intelligent decision-making analysis method for a digital twin-based manufacturing system includes the following steps: S100, Meta-learning pre-training (offline): By extracting multiple sets of simulated tasks from the source domain (mature market) data, train a base model with fast adaptation ability. The base model first fine-tunes the parameters on the support set of each small task, and then reversely updates the initial parameters according to the performance of the query set of all tasks, and finally obtains a meta-knowledge base that can quickly adapt to the new market; It should be added that the mature market refers to the cross-border e-commerce target area with the following characteristics: Data completeness: Having more than 100,000 historical order data, covering the complete product life cycle (R & D, production, delisting); Demand stability: The demand fluctuation coefficient (standard deviation / mean) is less than 0.3, and the seasonal index deviation is less than 15%; Supply chain maturity: A multi-level warehousing network has been established, and the average order fulfillment time limit ≤ 5 days; In an embodiment of the present invention, it specifically includes the following steps: S110, Meta-task construction: Sampling from the source domain data meta-tasks , and each task includes: Support set: ; where the support set is used for internal parameter optimization; Query set: ; where the query set is used for external parameter optimization where, is the th meta-task, is the number of support set samples ( ), is the number of query set samples ( ), is the number of meta-tasks, is the input feature vector, is the corresponding label value, is the source domain dataset; S120, Internal Parameter Adaptation: For each meta-task , based on the current meta-parameters , calculate the support set loss and update; ; Task Loss Calculation: ; Parameter Update: ; where, is the internal learning rate( ), is the meta-parameter vector, is the adaptation parameter of task , is the model function with parameter , is the loss function on the support set, is the gradient operator with respect to parameter ; S130, External Meta-parameter Optimization: Evaluate the performance of the adapted model on the query sets of all tasks, and update the meta-parameters by backpropagation; ; Meta-loss Calculation: ; Meta-gradient Update: ; where, is the external learning rate( ), is the loss function on the query set, is the total meta-learning loss, is the adapted model with parameter , represents the meta-parameter vector to be optimized, is the partial derivative of the meta-loss with respect to the meta-parameters, is the k-th input sample in the query set, is the k-th label value in the query set; S140, Convergence Criterion: Repeat S110 - S130 until the stopping condition is satisfied; where the stopping condition is as follows: ; where, is the convergence threshold( ) is the maximum number of iterations ( ) is the current iteration number, is the meta-loss value for the th iteration, is the meta-loss value for the previous iteration; S200, Adversarial Domain Adaptation (Online): On a small number of samples in the target domain (emerging markets), an adversarial training strategy is adopted to align the source domain and target domain features. The feature extractor is forced to learn domain-invariant representations, while the domain classifier tries to distinguish the data sources. At the same time, the weights of the prediction loss and the adversarial loss are dynamically adjusted to prevent negative transfer caused by overfitting of small samples; It should be noted that emerging markets refer to cross-border e-commerce expansion regions with the following characteristics: Data scarcity: The amount of available effective historical data is less than 1000, and key fields are missing (such as user portraits, reasons for returns); High demand volatility: The demand volatility coefficient exceeds 0.5, and the daily order volume can vary by up to 300% due to the influence of social media; Incomplete supply chain: Dependence on a single logistics channel, with an average customs clearance time of ≥7 days, and door-to-door delivery cannot be achieved in more than 30% of the regions; In an embodiment of the present invention, it specifically includes the following steps: S210, Model initialization: Load the meta-pre-trained parameters as the initial parameters; ; Among them, is the initialized model parameter, is the optimal meta-parameter obtained by meta-learning; S220, Domain adversarial feature alignment: Achieve domain invariance in the feature space through the Gradient Reversal Layer (GRL); Feature extraction: ; Domain prediction: ; Domain classification loss: ; Among them, is the feature extractor network with parameter , is the extracted 256-dimensional feature vector, is the domain label probability predicted by the domain classifier, is the true domain label, 0 represents the source domain, and 1 represents the target domain, The number of samples in each batch ( ), is the domain classifier parameter, following a normal distribution , is the domain classifier network with parameter , is the loss function of domain adversarial, is the true domain label of the i-th sample, is the predicted domain label probability of the i-th sample; S230, Joint optimization: Synchronously optimize the prediction task loss and the domain adversarial loss: ; Prediction task loss (only the source domain has labels): ; Weight adjustment: ; Among them, is the total loss function, is the loss function of the prediction task, is the loss function of domain adversarial, is the source domain sample batch size ( ), is the domain adversarial loss weight coefficient (initial value 0.1), is the weight adjustment coefficient ( ), is the current target domain mean absolute error, is the initial source domain mean absolute error; S240, Parameter update: Alternately update parameters: Fix while updating : ; Fix while updating : ; Among them, is the parameter of the feature extractor, is the parameter of the domain classifier, is the learning rate of adversarial learning ( ), is the gradient with respect to the parameter , is the gradient with respect to the parameter ; S250, Small sample early stopping strategy: Monitor the loss of the target domain validation set , and terminate the training when there is no decrease for 5 consecutive iterations; ; The monitoring metrics are as follows: Validation set: Randomly divide 20% of the samples from ( = 20); Loss calculation: ; Among them, is the loss value on the target domain validation set, is the current training round, is the number of validation set samples ( , which is 15% of the target domain data ), is the input feature of the validation set, is the true label of the validation set, represents the training termination signal, which is sent when the early stopping condition is triggered, represents taking the minimum value among all historical rounds less than the current round t, represents the historical training round index, which is used to compare the current loss with the historical optimal loss, represents the input feature of the target domain sample, represents the true label of the target domain; S300, Synthetic data augmentation: Based on the scarce data of the target domain and the domain knowledge graph, use the conditional diffusion model to generate a large amount of synthetic data, and construct an augmented dataset through distribution consistency testing and pre-training model screening; In one embodiment of the present invention, it specifically includes the following steps: S310, Domain knowledge conditional encoding: Encode the prior knowledge of the target domain (such as regional consumption characteristics, product category tree) into a conditional vector; ; Among them, is the conditional vector, is the knowledge embedding layer (parameters ), is the regional encoding, is the category encoding, is the embedding layer parameter matrix; S320, Conditional diffusion model training: Train the denoising diffusion probabilistic model (DDPM) with the target domain sample and the condition as the input; Forward process: ; Reverse process: ; Among them Implemented by UNet, the optimization objective is: ; Among them, is the noise coefficient ( ), is the noise prediction network, is the diffusion step number ( ), is the original input data, is the data after adding noise at the t-th step, is the mean prediction network, is the variance prediction network, is the identity matrix, represents the Gaussian distribution, represents the data state after t-step diffusion, represents the data state after (t - 1)-step diffusion, is the conditional probability distribution of the forward diffusion process, represents the transition probability of single-step forward diffusion, represents the probability distribution of the conditional reverse diffusion process, represents the loss function of the DDPM model, represents the expectation with respect to the time step , the initial data and the random noise , represents the square of the Euclidean norm, represents the product from t = 1 to T; S330, Diversity-Controllable Sample Generation: Sampling from the diffusion model to generate augmented data, constrained to be consistent with the target domain distribution; ; Implementation method: Using classifier guidance sampling, injecting conditional gradients in the reverse process: ; Among them, is the k-th generated sample, is the conditional generation distribution, is the maximum mean discrepancy metric, is the real data sample, is the guidance coefficient, is the MMD threshold, represents the gradient operator with respect to , is the unconditional generation distribution, representing the generation model distribution at time t, represents the conditional probability distribution, representing the condition given the generation state The posterior probability, represents the conditional generation distribution, which is the generation model distribution at time t under the condition c, represents the set of generated samples, represents the set of real samples; S340, Verification of the quality of generated samples: Screening effective samples through a pre-trained verifier ; ; Among them, is the verification model, is the similarity threshold, , is the augmented dataset, is the KL divergence, is the verification model parameter, represents the delimiter symbol in the KL divergence calculation, represents the feature representation of the generated samples by the verification model, represents the feature representation of the real samples by the verification model; S350, Construction of the augmented dataset: Merging real and synthetic data to construct the final training set: ; The required balancing strategy is constructed as follows: Expand to , Select ; Among them, is the finally constructed augmented dataset, is the target domain dataset, is the final dataset size; S400, Elastic scheduling optimization: Loading the augmented data and the adaptation model in the digital twin, running the multi-objective deep reinforcement learning algorithm, dynamically adjusting the weights of the order fulfillment rate, inventory cost, etc. in combination with real-time feedback, and finally selecting the optimal scheduling strategy from the Pareto optimal solution set; In an embodiment of the present invention, it specifically includes the following steps: S410, Environment initialization: Loading the virtual factory model and injecting the augmented data ; ; Among them, is the digital twin environment, is the Unity3D factory model, is the finally constructed augmented dataset; S420, Multi-objective Deep Reinforcement Learning (MO-DDPG) Policy Training: Construct an Actor-Critic architecture and optimize three objectives: Actor Network : Input the state , and output the action (equipment start / stop, logistics path); Critic Network : Evaluate the action value, including three output heads: ; Objective function: ; Among them, is the Actr network, and the parameter is , is the Critic network, and the parameter is , is the state vector, is the action vector, is the state distribution generated by the behavioral policy, is the variance penalty coefficient, , is the variance of the action, is the cost evaluation value, is the quality of service evaluation value, is the feasibility evaluation value, represents the objective function, which is used to optimize the Actor network parameters, is the expectation operator, which is used to calculate the average return, represents the variance of the action, which is used to encourage action diversity; S430, Dynamic Weight Adjustment Mechanism: Adjust the objective weights according to the real-time feedback of the digital twin; ; Among them, is the weight of the th objective at the th moment, is the normalized score of the th objective at the th moment, is the temperature parameter ([[]] ), is the current value of the th objective, is the target value of the th objective; Specific threshold values: Service rate ≥ 0.95, cost ≤ 0.3, feasibility ≥ 0.9; S440, Virtual-Physical Closed-Loop Verification: Execute the policy in the digital twin , verify the feasibility; ; Among them, is the feasibility score of the policy , is the number of verification rounds, is the indicator function, is the result of the verification in the round; The constraint conditions include: the upper limit of the load is 80%; the upper limit of the delay is 36 hours; the carbon emission does not exceed the standard value; S450, Pareto front policy selection: select the optimal policy from the non-dominated solution set : ; Among them, is the optimal policy, is the Pareto optimal policy set, is the weight of the th objective, is the Q value of the th objective, is the feasibility score of the policy , is the system state, is the action output of the policy in the state ; Optimal policy output: ; Among them is the action of the production line start / stop decision, is the resource allocation ratio, is the production capacity allocation weight; As Figure 6 shown, based on the above method, an intelligent decision-making analysis system for a digital twin-based manufacturing system is proposed, including the following modules: Meta-learning base model construction module 101: Automatically construct multiple groups of simulation tasks from the historical data of the mature market, and obtain a meta-learning model with fast adaptation ability through collaborative training of internal and external loops, so that it can quickly capture the common laws across markets under a small amount of target data, laying a foundation for small-sample migration; Adversarial domain adaptation module 102: Introduce an adversarial training mechanism on a very small amount of data in the target domain, dynamically balance the prediction accuracy and domain-invariant feature learning, and eliminate the distribution difference between the source domain and the target domain through gradient reversal and weight adaptive strategies to avoid performance collapse of the model in emerging markets; Data augmentation generation module 103: Generate synthetic data that conforms to the characteristics of the target domain based on the conditional diffusion model and the domain knowledge graph. Combine distribution consistency verification and pre-trained model screening to expand high-quality training samples and solve the problem of model overfitting caused by data scarcity. Scheduling optimization decision-making module 104: Simulate multi-objective scheduling strategies in the digital twin environment. Dynamically optimize the order fulfillment rate, inventory cost, and feasibility score through reinforcement learning, and finally output a virtual-verified flexible production plan to ensure the reliable execution of the strategy in the physical system.
[0020] Based on the above methods and systems, the following is an example: A multinational enterprise "Global Tech" plans to enter the East African smartphone market but faces: Data scarcity: Only 3 months of local sales data (about 100 records) have been collected. Demand differences: African users prefer long-battery-life mobile phones (battery capacity greater than 5000mAh), which is significantly different from the European and American markets. Supply chain restrictions: A new assembly plant needs to be set up in Kenya, but there is no experience in capacity allocation. System application process: Meta-learning base model construction: Load 100,000 historical data from the European and American markets, automatically generate 1200 groups of simulation tasks (such as "festival promotions", "chip shortage responses"), and after training, the model can identify cross-market common laws (such as "the correlation between price sensitivity and GDP"). Adversarial domain adaptation: Input 100 pieces of data from Africa (including features such as battery capacity and local payment methods). The system detects a distribution shift in the "battery life demand" feature (KL divergence = 7.8), automatically strengthens the alignment of this dimension feature. After 8 hours of training, the prediction error drops from the initial 42% to 15%. Data augmentation generation: Generate 10,000 virtual orders based on the cooperation data of African operators. The generated data includes special scenarios: purchase delays caused by electricity price fluctuations during the dry season, and mobile payment accounting for 85%. Scheduling optimization decision: Digital twin simulation shows that directly copying the production scheduling plan of the Vietnam factory will result in 30% equipment idleness.
[0021] System recommended strategies: Priority deployment of production lines for large-battery models in the Kenya factory; Establish a direct link with the logistics center in Rwanda to shorten the delivery time; Dynamically reserve 15% of the production capacity to handle sudden community group-buying orders; By implementing the recommended strategies, the following improvement effects are obtained, as shown in Table 1: Table 1: Improvement effects
[0022] Through this system, Global Tech achieved a breakthrough in the East African market share from 0 to 17% within 6 months, verifying the commercial value of the digital twin decision-making system in small-sample scenarios.
[0023] The embodiments of the present invention have been described above. However, the present invention is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of the present invention.
Claims
1. An intelligent decision-making analysis method for a manufacturing system based on digital twins, characterized in that, It includes the following steps: S100. Train a base model with fast adaptation ability by extracting multiple groups of simulation tasks from the source domain data. The base model first fine-tunes the parameters on the support set of each small task, and then updates the initial parameters in reverse according to the performance on the query set of all tasks, and finally obtains a meta-knowledge base that can quickly adapt to the new market; S200. On the target domain small samples, adopt an adversarial training strategy to align the source domain and target domain features. The feature extractor is forced to learn domain-invariant representations, while the domain classifier tries to distinguish the data sources. At the same time, dynamically adjust the weights of the prediction loss and the adversarial loss to prevent negative transfer caused by overfitting of small samples; S300. Based on the scarce data in the target domain and the domain knowledge graph, use a conditional diffusion model to generate a large amount of synthetic data, and construct an enhanced data set through distribution consistency testing and pre-trained model screening; S400. Load the enhanced data and the adapted model in the digital twin, run the multi-objective deep reinforcement learning algorithm, dynamically adjust the weights of the order fulfillment rate and the inventory cost in combination with real-time feedback, and finally select the optimal scheduling strategy from the Pareto optimal solution set.
2. The intelligent decision-making analysis method of a manufacturing system based on digital twin according to claim 1, wherein, In S100, it includes the following steps: S110, Meta-task Construction: Sample meta-tasks from the source domain data, where each task contains a support set and a query set; S120. Internal parameter adaptation: For each meta-task, calculate the support set loss and update it based on the current meta-parameters; S130. External meta-parameter optimization: Evaluate the performance of the adapted model on the query set of all tasks and update the meta-parameters by backpropagation; S140. Convergence determination: Repeat S110 - S130 until the stop condition is met.
3. The intelligent decision-making analysis method of a manufacturing system based on digital twin according to claim 2, wherein In S140, the stop conditions are as follows: ; wherein, is the convergence threshold, is the maximum number of iterations, , is the current iteration number, is the -th iteration's meta-loss value, is the meta-loss value of the previous iteration.
4. The intelligent decision-making analysis method of a manufacturing system based on digital twin according to claim 1, wherein In S200, it includes the following steps: S210. Model initialization: Load the meta-pre-trained parameters as the initial parameters; S220. Domain adversarial feature alignment: Achieve domain invariance in the feature space through the gradient reversal layer; S230. Joint optimization: Synchronously optimize the prediction task loss and the domain adversarial loss; S240. Parameter update: Alternately update the parameters of the feature extractor and the domain classifier: S250. Small sample early stopping strategy: Monitor the loss of the target domain validation set and terminate the training when there is no decrease for 5 consecutive iterations.
5. The intelligent decision-making analysis method of a manufacturing system based on digital twin according to claim 4, characterized in that, In S240, the calculation formula for parameter update is as follows: Fixed Update simultaneously : ; Fixed Update simultaneously : ; Among them, are the parameters of the feature extractor, are the parameters of the domain classifier, is the learning rate of adversarial learning, , is the gradient with respect to the parameter , is the gradient with respect to the parameter , is the loss function of domain adversarial, is the total loss function.
6. The intelligent decision-making analysis method of a manufacturing system based on digital twin according to claim 5, characterized in that The calculation formula for the small sample early stopping strategy is as follows: wherein, is the loss value on the target domain validation set, is the current training round, indicates the training termination signal, which is issued when the early stopping condition is triggered, means taking the minimum value among all historical rounds less than the current round t, represents the historical training round index, which is used to compare the current loss with the historical optimal loss.
7. An intelligent decision-making analysis method for a manufacturing system based on digital twin according to claim 1, characterized in that, In S300, it includes the following steps: S310. Domain knowledge conditional encoding: Encode the prior knowledge of the target domain into a conditional vector; S320. Conditional diffusion model training: Train a denoising diffusion probability model with the target domain samples and conditions as inputs; S330. Diversity controllable sample generation: Sample from the diffusion model to generate enhanced data, and constrain it to be consistent with the target domain distribution; S340. Generated sample quality verification: Screen valid samples through a pre-trained validator; S350. Enhanced data set construction: Merge real and synthetic data to construct the final training set.
8. An intelligent decision-making analysis method for a manufacturing system based on digital twin according to claim 7, characterized in that In S400, it includes the following steps: S410. Environment initialization: Load the virtual factory model and inject enhanced data; S420. Multi-objective deep reinforcement learning strategy training: Construct an Actor-Critic architecture and optimize three objectives: Actor network: input state , output action ; Critic network: Evaluate the action value, with three output heads: ; Objective function: ; Among them, is the Actr network, and the parameter is , is the Critic network, and the parameter is , is the state vector, is the action vector, is the state distribution generated by the behavioral policy, is the variance penalty coefficient, is the variance of the action, is the cost evaluation value, is the quality of service evaluation value, is the feasibility evaluation value, represents the objective function, is the expectation operator, represents the variance of the action; S430, Dynamic weight adjustment mechanism: Adjust the target weight according to the real-time feedback of the digital twin; S440, Virtual-physical closed-loop verification: Execute the strategy in the digital twin to verify its feasibility; S450, Pareto front strategy selection: Select the optimal strategy from the non-dominated solution set.
9. The intelligent decision-making analysis method of a manufacturing system based on digital twin according to claim 1, characterized in that The calculation formula for the optimal strategy is as follows: ; Among them, is the optimal strategy, is the Pareto optimal strategy set, is the weight of the th objective, is the Q value of the th objective, is the feasibility score of strategy is the system state, is the action output of strategy in state ; The output of the optimal strategy is as follows: ; Among them is the action for production line start / stop decision-making is the resource allocation ratio is the production capacity allocation weight 10. An intelligent decision-making analysis system for a manufacturing system based on digital twins, characterized in that, Execute the steps in an intelligent decision-making analysis method for a digital twin-based manufacturing system as described in any one of claims 1-9, including: Meta-learning base model construction module: Automatically construct multiple groups of simulation tasks from the historical data of the mature market, and obtain a meta-learning model with fast adaptation ability through collaborative training of the inner and outer loops, enabling it to quickly capture the common laws across markets with a small amount of target data; Adversarial domain adaptation module: Introduce an adversarial training mechanism on a very small amount of data in the target domain, dynamically balance the prediction accuracy and domain-invariant feature learning, and eliminate the distribution differences between the source domain and the target domain through the gradient reversal and weight adaptation strategies; Data augmentation generation module: Generate synthetic data that conforms to the characteristics of the target domain based on the conditional diffusion model and the domain knowledge graph, and expand high-quality training samples by combining distribution consistency verification and pre-trained model screening; Scheduling optimization decision module: Simulate multi-objective scheduling strategies in the digital twin environment, dynamically optimize the order fulfillment rate, inventory cost, and feasibility score through reinforcement learning, and finally output a virtual-verified flexible production plan.
Citation Information
Patent Citations
Decision control method and system for digital twin information of intelligent factory based on 5G driving
CN114637262A
Intelligent clothing industry production regulation and control method and system based on data analysis
CN118798494A
Flexible production plan prediction system and method based on knowledge graph and language model
CN119623911A
Artificial intelligence based data-driven interconnected digital twins
US20240419154A1
Cited By
Production decision adaptive optimization method and system
CN120406165A
Adaptive Optimization Methods and Systems for Production Decision Making
CN120406165B
Project management system development method based on low code configuration
CN121455533A
Visual detection method and system for mobile phone battery by intelligent edge device
CN121582907A
Method, system and device for optimizing technological parameters of FCBGA chip carrier plate and medium
CN121809382A