Cross-working-condition evolution and migration method and system for intelligent installation and adjustment model of complex equipment body
By using a hybrid expert architecture and an online self-supervised evolution module, the problems of insufficient model generalization and adaptive lag in complex equipment assembly and adjustment are solved, enabling rapid and reliable cross-task skill transfer and improving the flexibility and intelligence of the assembly and adjustment system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHONGQING UNIV
- Filing Date
- 2026-01-20
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies suffer from insufficient model generalization ability, lagging online adaptive ability, and low cross-task knowledge transfer efficiency during the assembly and commissioning of complex equipment, resulting in high assembly failure rate, frequent downtime for debugging, long new product introduction cycle, and high data acquisition cost.
An embodied intelligent assembly model based on a hybrid expert architecture is adopted, which combines an online self-supervised evolution module and a model transfer module. By sharing a feature extraction layer, a parallel expert network library, and a dynamic gating network, the model can achieve lightweight incremental updates and safe and fast transfer. Residual initialization and knowledge distillation techniques are used to ensure the safety and controllability of new task strategies.
It enhances the model's ability to handle diverse tasks in complex assembly and adjustment, reduces downtime caused by performance degradation, shortens the new product introduction cycle, reduces data acquisition costs, and improves the flexibility and intelligence of the assembly and adjustment system.
Smart Images

Figure CN121900177A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent manufacturing and embodied intelligence technology, specifically a method and system for cross-working-condition evolution and migration of embodied intelligent assembly and adjustment models for complex equipment. Background Technology
[0002] As the manufacturing industry moves towards greater flexibility and customization, industrial sites place extremely high demands on the adaptability and rapid migration capabilities of automated assembly and debugging systems. Traditional automated assembly and debugging lines are typically designed for a single, specific product and rely heavily on precision tooling fixtures and pre-set fixed teaching trajectories. Once a product needs to be changed, or if there are batch differences in parts or tool wear in the production environment, traditional rigid automated systems often cannot adapt effectively. This leads to a significant increase in assembly failure rates and may even require prolonged downtime for reprogramming and mechanical debugging, resulting in substantial losses in human and time costs.
[0003] Embodied intelligence technology offers a new approach to solving the aforementioned problems. This technology emphasizes that intelligent agents (such as robots) engage in real-time physical interaction with their environment through their own sensory modules, learning and optimizing control strategies in the process, thus enabling them to adapt to uncertain environments. However, applying embodied intelligence models to the actual assembly and adjustment process of complex equipment still faces significant challenges.
[0004] First, the models lack generalization and multi-task processing capabilities. Many models employ a single, large neural network structure to handle all tasks. When faced with significant changes across products or extreme disturbances across operating conditions, the capacity and generalization ability of a single network are limited, easily leading to "catastrophic forgetting," where the ability to reliably handle older products is lost after learning new product skills. While existing technologies have used hybrid expert architectures to improve model capacity and processing diversity in perception or prediction tasks such as anomaly detection and parameter prediction, these efforts focus on open-loop perception or prediction problems and do not address how to apply hybrid expert architectures to closed-loop robot action sequence generation and decision-making to cope with complex physical interactions and variable task sequences in assembly and debugging.
[0005] Secondly, the online adaptive and evolutionary capabilities of models lag behind. Industrial field conditions (such as tool wear and part tolerances) drift slowly over time, requiring models to be able to self-adjust online and safely. Current model updates typically rely on retraining after collecting large-scale offline data, making it impossible to leverage immediate experience from successes and failures during actual production line operation for lightweight, incremental self-evolution. Some existing technologies adapt to changes by identifying working conditions and switching between different models. This strategy essentially involves storing and selecting multiple models, rather than the online parameter evolution of a single model. It fails to achieve continuous accumulation and iteration of knowledge within the model, making it difficult to keep up with the dynamic drift of working conditions.
[0006] Furthermore, cross-task knowledge transfer is inefficient and lacks security. When a new product or process is introduced to the production line, it is often necessary to re-collect a large amount of demonstration data to train the model from scratch, failing to effectively reuse existing mature assembly and adjustment skill modules, resulting in long new product introduction cycles and high data collection costs. Although some existing technologies, such as pre-training fine-tuning and domain adversarial learning, can be used to transfer knowledge between different domains, these methods mainly focus on the alignment of feature representations. For physically interactive tasks with strong security requirements, such as embodied intelligent assembly and adjustment, existing methods do not provide effective solutions for ensuring that the transferred strategy has safe and controllable behavioral boundaries in the initial stage, avoiding equipment collisions or product damage due to policy mutations. Directly applying transfer techniques from prediction tasks to action policy transfer lacks specific design for initial security and behavioral stability.
[0007] In summary, existing technologies lack a solution that can simultaneously address online adaptive evolution and safe, efficient cross-task skill transfer in complex physical interaction tasks. Therefore, there is an urgent need for an innovative method specifically designed for complex equipment-embedded intelligent assembly scenarios to achieve self-optimization of the model during continuous operation and rapid, reliable adaptation to task changes. Summary of the Invention
[0008] In view of this, the purpose of this invention is to provide a method and system for cross-condition evolution and transfer of intelligent assembly and adjustment models for complex equipment. The method adapts to condition drift through online self-supervised evolution of the model and achieves safe and rapid cross-task skill transfer by utilizing residual initialization and knowledge distillation, thereby solving the problems of insufficient model generalization, slow evolution and low transfer efficiency.
[0009] To achieve the above objectives, the present invention provides the following technical solution: This invention first proposes a method for cross-condition evolution and transfer of a complex equipment embodied intelligent assembly and adjustment model, including the following steps: Step 1: Construct and deploy an embodied intelligent assembly and adjustment model network based on a hybrid expert architecture. The embodied intelligent assembly and adjustment model network receives a composite input state vector consisting of real-time physical state and task semantics. and output robot motion control commands. ; Step 2: When the embodied intelligent assembly and adjustment model network is running at the actual assembly and debugging station, the embodied intelligent assembly and adjustment model network is dynamically and incrementally updated through the online model evolution module to adapt to the working condition drift in the production process in real time. Step 3: When the work task changes due to product replacement, the model migration module quickly generates a model adapted to the new task based on the skill modules in the existing embodied intelligent assembly and adjustment model network.
[0010] Furthermore, in step one, the embodied intelligent assembly and adjustment model network based on a hybrid expert architecture includes: A shared feature extraction layer is used to process the input multimodal sensing data and extract common low-level physical features; The parallel expert network library consists of multiple expert networks with the same structure but independent parameters. Each expert network is used to encode the assembly and adjustment skills of a specific type of product or working condition. Dynamic gating network, used to control the composite input state vector It outputs a normalized weight distribution to dynamically select and activate one or more expert networks, and merges the outputs of the selected expert networks to form the final action instruction.
[0011] Furthermore, in step two, the online model evolution module includes: The self-supervised evaluation mechanism is used to compare the real-time collected multimodal data and ontological state data with preset process thresholds to generate reward signals that characterize the success or failure of the assembly and adjustment task. The experience replay buffer is used to store experience samples consisting of state, action and evaluation signals in time sequence, and manage the samples according to a preset elimination strategy. An online evolution triggering mechanism is used to monitor in real time the statistical characteristics of the evaluation signals in the experience replay buffer and the entropy value of the output distribution of the dynamic gating network, and triggers model updates when the preset evolution evaluation index is met. The selective update unit is used to identify the active expert set and generate a parameter update mask based on the weight distribution output by the dynamic gating network, perform lightweight incremental optimization only on the parameters of the active expert set, and freeze the parameters of inactive experts.
[0012] Furthermore, the self-monitoring evaluation mechanism is as follows: based on product model... With process type Determine process thresholds, including maximum permissible contact force. Maximum permissible torque Maximum error threshold With the maximum allowable beat threshold At the end of each control cycle or process stage, calculate the assembly and adjustment accuracy error. Contact force With rhythm And an evaluation signal is generated based on a self-supervised scoring function; the self-supervised scoring function is expressed as: in: , , and It is a non-negative weighting coefficient used to assess the contribution of three types of indicators: balance accuracy, force control, and cycle time.
[0013] Furthermore, the online evolution triggering mechanism is as follows: Real-time statistics of the experience replay buffer China recently The mean of the evaluation of each sample: in: The function value of the self-supervised scoring function; Calculate the entropy of the output distribution of the dynamic gating network: in: The output distribution entropy is gated; The mean of the distribution entropy; For gating networks to the first The output weights of each expert; The total number of experts available; When satisfied or The model is triggered to evolve online at specific times, where: and The preset threshold is determined adaptively based on the statistical characteristics of historical production data, or set according to process specifications and safety limits.
[0014] Furthermore, in the selective update unit, the weights output by the current dynamic gating network are updated accordingly. Determine the set of active experts: in: The preset activation threshold; These are the weight coefficients of the gating network; Generate parameter update mask only for the set of active experts. The expert network parameters are optimized using lightweight incremental optimization, while other expert network parameters are frozen. In the selective update unit, lightweight incremental optimization is implemented using a low-rank adaptation method, and its objective function is: in: Current model parameter set; This represents the parameter increment relative to the previous version; The coefficient is a non-negative positive regularity coefficient; Indicates an experience-based replay buffer The reinforcement learning loss term is calculated from the sampled data; the learning rate is adaptively adjusted based on the recent evaluation improvement to improve the response speed to operating condition drift.
[0015] Furthermore, in step three, the method for quickly generating a model adapted to the new task through the model transfer module includes the following steps: 31) Task decomposition: Decompose old tasks and new tasks hierarchically to obtain the corresponding atomic skill sequences; 32) Expert Matching and Creation: Based on the atomic skill sequence of the new task, match the most similar old experts in the parallel expert network library. Create a new expert network corresponding to the new task based on a preset template. ; 33) Residual initialization: Let the old expert parameters be... The new expert parameters are initialized to : Incremental parameters Initialize to zero or near-zero values so that the following conditions are met during the initial migration: in: This represents the output increment of the new expert relative to the old expert. 34) Knowledge Distillation and Fine-tuning: Minimizing the incremental parameters of new experts by minimizing the loss function. Make minor adjustments: in: This is the distillation loss item; This is a task loss item; For incremental regularization terms; , and These are the combination coefficients for the three losses; and These are the action vectors output by the new expert and the action vectors output by the old expert, respectively. For error; The maximum allowable error threshold; For contact force; Maximum permissible contact force; For rhythm; The maximum permissible beat; , and The weights for the three types of penalties.
[0016] Furthermore, in step 34), the constraint loss term based on the new task process threshold... Construction method and self-supervised scoring function The construction method is consistent, but the variables are replaced with actions based on the new expert output. After execution, the accuracy error, contact force, and cycle time data obtained from prediction or simulation are compared and calculated with the process threshold of the new task.
[0017] Furthermore, the composite input state vector The physical state data comes from the RGB-D camera fixed to the robot's wrist, the six-dimensional force / torque sensor at the end effector, and the robot's encoder. The task semantic data comes from the work order information issued by the manufacturing execution system.
[0018] This invention also proposes a cross-condition evolution and migration system for a complex equipment embodied intelligent assembly and adjustment model, comprising: The model deployment and execution module is used to deploy the embodied intelligent assembly and debugging model network based on the hybrid expert architecture constructed as described above, and to execute assembly and debugging tasks. The online model evolution module dynamically and incrementally updates the embodied intelligent assembly and adjustment model network to adapt to the working condition drift during the production process in real time. The model migration module can quickly generate a model adapted to the new task based on the skill modules in the existing model network when the job task changes due to product replacement. The data management and sensing unit is used to collect multimodal sensing data, manage process thresholds, and establish an experience playback buffer.
[0019] The beneficial effects of this invention are as follows: The present invention provides a cross-condition evolution and migration method for a complex equipment-embedded intelligent assembly and adjustment model, which has the following technical advantages: (1) By constructing a model based on a hybrid expert architecture, the complex assembly and adjustment task is decomposed into sub-skills that can be processed by different expert networks, which effectively improves the internal capacity and structured representation ability of the model to handle diverse tasks, and lays the foundation for solving the generalization deficiency of a single network. (2) The unique online evolution module enables the model to perceive performance degradation or changes in operating conditions in real time during production operation through self-supervised evaluation, and trigger lightweight incremental updates for active experts. This mechanism enables the model to autonomously adapt to slow operating condition drift (such as tool wear), significantly reducing downtime due to performance degradation and improving production continuity; (3) For the model migration module for product transformation, by reusing the existing expert network as the skill base, combined with residual initialization and knowledge distillation, the behavioral safety and controllability of the new task strategy in the early stage of migration are ensured, and the dependence on the new task demonstration data is greatly reduced. This significantly shortens the new product introduction cycle and realizes rapid and reliable production line reconstruction.
[0020] Overall, the method of this invention forms a complete solution from "adaptive operation in daily working conditions" to "rapid switching to sudden tasks", which systematically improves the flexibility and intelligence level of the assembly and adjustment system. Attached Figure Description
[0021] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration: Figure 1 This is a flowchart of the cross-working-condition evolution and migration method of the intelligent assembly and adjustment model of complex equipment according to the present invention. Figure 2 This is a schematic diagram of the online evolution module of the model; Figure 3 This is a schematic diagram of the online evolution of the model caused by changing working conditions during bolt tightening; Figure 4 A schematic diagram illustrating the model migration from bolt tightening to wire harness insertion; Figure 5 This is a schematic diagram of the model transfer module. Detailed Implementation
[0022] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.
[0023] This embodiment proposes a method and system for cross-condition evolution and migration of embodied intelligent assembly and adjustment models for complex equipment. By building an embodied intelligent assembly and adjustment model network based on a hybrid expert architecture, and developing an online model evolution module and a model migration module, it solves the problems of insufficient generalization ability of existing embodied intelligent models in complex assembly and adjustment environments, inability to achieve online evolution, and long migration cycles for new products.
[0024] Specifically, this embodiment takes the tightening of bolts on the front subframe of the car body in the car assembly line as an example, while the model migration takes the replacement of the wiring harness connector plug and the engine controller by tightening the bolts as an example.
[0025] I. Workstation Hardware The workstation hardware in this embodiment: The six-axis industrial robot has an end effector consisting of a servo tightening gun for tightening bolts and an end gripper for gripping wire harness connector plugs. An external vision sensor, an RGB-D camera fixed to the robot's wrist, is used to identify the target's posture; A six-dimensional force and torque sensor at the end of the device is used for contact detection, force control input, and collision determination. Robot controller, providing real-time control and joint position and speed.
[0026] II. Threshold Configuration After the vehicle body arrives at the station, the manufacturing execution system at the workstation issues work order information. The system reads the product model identifier and the current process type, and loads the corresponding process thresholds, as well as the maximum allowable contact force and torque thresholds, cycle time thresholds, and assembly and adjustment accuracy error thresholds.
[0027] III. Cross-condition evolution and migration methods like Figure 1 As shown in the figure, the cross-condition evolution and migration method of the intelligent assembly and adjustment model of complex equipment in this embodiment includes the following steps.
[0028] Step 1: Construct and deploy an embodied intelligent assembly and adjustment model network based on a hybrid expert architecture. The embodied intelligent assembly and adjustment model network receives a composite input state vector consisting of real-time physical state and task semantics. and output robot motion control commands. .
[0029] In this embodiment, the composite input state vector The physical state data comes from the RGB-D camera fixed to the robot's wrist, the six-dimensional force / torque sensor at the end effector, and the robot's encoder. The task semantic data comes from the work order information issued by the manufacturing execution system.
[0030] In this embodiment, the embodied intelligent assembly and adjustment model network based on a hybrid expert architecture includes a shared feature extraction layer, a parallel expert network library, and a dynamic gating network.
[0031] The shared feature extraction layer processes the input multimodal sensing data to extract common underlying physical features. Specifically, the shared feature extraction layer receives multimodal data from vision sensors, end effector six-dimensional force and torque sensors, and the robot's body state, and extracts common features relevant to the tightening task.
[0032] The parallel expert network library consists of multiple expert networks with the same structure but independent parameters (e.g., each specializing in coarse hole alignment, compliant insertion, constant force tightening, anomaly detection, etc.). Each expert network is used to encode assembly and adjustment skills for a specific type of product or working condition.
[0033] Dynamic gating networks are used to control the composite input state vector. It outputs a normalized weight distribution to dynamically select and activate one or more expert networks, and merges the outputs of the selected expert networks to form the final action instruction.
[0034] After the robot aligns the tightening gun with the bolt holes, external vision scans the holes on the front subframe to acquire the bolt hole orientation, key positioning surfaces, and target pose. The system generates an approach path and guides the bolt to the pre-tightening area in front of the bolt hole. In each control cycle, the collected real-time data is input into the deployed embodied intelligent assembly and adjustment model based on hybrid experts, expert weights are derived through reasoning, and the output action commands of each expert are fused to form the final action command.
[0035] Step 2: When the embodied intelligent assembly and adjustment model network is running at the actual assembly and debugging station, the embodied intelligent assembly and adjustment model network is dynamically and incrementally updated through the online model evolution module to adapt to the working condition drift in the production process in real time.
[0036] like Figure 2 As shown, in this embodiment, the online model evolution module includes a self-supervised evaluation mechanism, an experience replay buffer, an online evolution triggering mechanism, and a selective update unit.
[0037] (1) Self-monitoring and evaluation mechanism The self-supervised evaluation mechanism compares real-time collected multimodal data and robot status data with preset process thresholds to generate reward signals indicating the success or failure of the assembly task. Specifically, during execution, the system collects real-time data on the monitoring accuracy error, peak force and torque, and cycle time generated by the industrial robot when performing assembly tasks such as screw hole positioning and tightening. This data is then compared with preset process thresholds to automatically generate reward signals indicating the success or failure of a single assembly task. If the tightening is in place, the contact force is within limits, and the cycle time meets the requirements, the task is considered successful and a positive reward is given. If the force exceeds limits, there is a collision, or the hole alignment error exceeds limits, the task is considered a failure and a negative reward is given. The failure type is also recorded for subsequent updates and traceability.
[0038] In this embodiment, the principle of the self-supervised evaluation mechanism is as follows: based on the product model... With process type Determine process thresholds, including maximum permissible contact force. Maximum permissible torque Maximum error threshold With the maximum allowable beat threshold Thresholds can be derived from process specifications, safety limits, or statistical analysis of historical production line data, and can be set separately for each product or process. At the end of each control cycle or process stage, the assembly and adjustment accuracy error is calculated. Contact force With rhythm These data are compared with the aforementioned process thresholds to obtain a reward signal for the success or failure of a single assembly and adjustment task, and an evaluation signal is generated based on a self-supervised scoring function. The self-supervised scoring function is expressed as: in: , , and It is a non-negative weighting coefficient used to assess the contribution of three types of indicators: balance accuracy, force control, and cycle time.
[0039] The obtained function value It is not a success / failure arbiter. It uses a composite input state vector. Input a shared feature extraction layer and a dynamic gating network to drive expert selection and action policy generation; output time step... Action instructions and the status -action -evaluate Triples are written to the experience replay buffer. : (2) Experience replay buffer The experience replay buffer is used to store experience samples consisting of state, action, and evaluation signals in chronological order, and manages these samples according to a preset elimination strategy. Specifically, the state-action-evaluation records of the industrial robot during assembly and adjustment operations are written into the experience replay buffer as experience samples. When the buffer capacity reaches a set threshold, the earliest and most stable samples are eliminated first to maintain adaptability to recent operating conditions.
[0040] (3) Online evolution triggering mechanism An online evolution triggering mechanism is used to monitor the statistical characteristics of the evaluation signals in the experience replay buffer and the entropy value of the output distribution of the dynamic gating network in real time. When preset evolution evaluation indicators are met, model updates are triggered. Specifically, the system calculates the mean of the self-supervised evaluation values in the experience replay buffer and the entropy of the gating output distribution in real time. When changes in operating conditions such as tool wear, product batch differences, or loose frame bolts cause anomalies in the collected data, online model evolution is triggered if any one of the online evolution indicators is met.
[0041] In this embodiment, the principle of the online evolution triggering mechanism is as follows.
[0042] The system extracts cached data from the experience replay buffer in real time. Valid samples constitute the statistical set This is used for calculating evolutionary evaluation metrics. The experience replay buffer is statistically analyzed in real time. China recently The mean of the evaluation of each sample: in: This represents the function value of the self-supervised scoring function.
[0043] The average evaluation value is consistently below the preset threshold. Triggering online evolution: Suppose a gating network for the first The output weights of each expert are: , Define the gating output distribution entropy to represent the total number of available experts: in: The output distribution entropy is gated; The mean of the distribution entropy; For gating networks to the first The output weights of each expert; This represents the total number of experts available for selection.
[0044] A sustained increase in gating entropy indicates that expert choice is uncertain and exceeds a preset threshold. : When satisfied or The model is triggered to evolve online at specific times, where: and The preset threshold is determined adaptively based on the statistical characteristics of historical production data, or set according to process specifications and safety limits.
[0045] (4) Selective update unit The selective update unit is used to identify the active expert set and generate a parameter update mask based on the weight distribution output by the dynamic gating network, perform lightweight incremental optimization only on the parameters of the active expert set, and freeze the parameters of inactive experts.
[0046] After triggering online evolution, the system constructs training batches by hierarchically sampling from the experience replay buffer. The sampling scale of boundary samples and anchor samples can be tens to hundreds, and the total number of training samples can be hundreds to thousands. The selective update module identifies the set of active experts whose contribution to the current task is higher than a threshold based on the weight distribution output by the dynamic gating network, and then generates a parameter update mask to limit the range of updatable parameters in the online evolution stage. Parameters are frozen for inactive experts, and only lightweight incremental optimization is performed on the parameters of active experts to suppress catastrophic forgetting. If the evaluation index recovers slowly, the exploration step size is increased; if performance oscillations occur, the step size is reduced to maintain the stability of the production line operation.
[0047] In the selective update unit of this embodiment, a gated selective parameter update strategy is adopted. This is based on the expert weights output by the dynamic gating network. , subscript In At this moment, identify the active expert group: in: The preset activation threshold; represents the weight coefficients of the gating network.
[0048] Generate parameter update mask only for the set of active experts. The expert network parameters are optimized using lightweight incremental optimization, while other expert network parameters are frozen.
[0049] In the selective update unit, lightweight incremental updates can be implemented using a low-rank adapted LoRA parameter fine-tuning method: periodically replaying data from the experience buffer. Sample small batches of data and minimize the incremental objective: in: Current model parameter set; This represents the parameter increment relative to the previous version; The coefficient is a non-negative positive regularity coefficient; Indicates an experience-based replay buffer The reinforcement learning loss term is calculated from the sampled data. The learning rate is adaptively adjusted based on the improvement in recent evaluations to improve the response speed to operating condition drift.
[0050] Step 3: When the work task changes due to product replacement, the model migration module quickly generates a model adapted to the new task based on the skill modules in the existing embodied intelligent assembly and adjustment model network.
[0051] When the manufacturing execution system issues a command to connect the wiring harness connector to the engine controller, causing a change in operating conditions, the system first hierarchically decomposes the bolt tightening task into a sub-task sequence of "positioning identification - coarse alignment - compliant input - screwing in and tightening - confirmation of position," and stores the corresponding expert network in the expert database. Then, it hierarchically decomposes the new task into a sub-task sequence of "positioning identification - coarse alignment - compliant input - insertion and locking - confirmation of position," such as... Figure 3-4 As shown, the system creates new expert network branches based on preset expert templates as needed, allocates independent incremental parameters, and registers them to the expert database to obtain new experts. Based on the atomic skill sequence, the system selects the most similar old expert from the parallel expert network database as the initialization source, performs residual initialization on the new expert, and makes its initial parameters satisfy the incremental initialization of zero, thereby ensuring that the initial output of the new expert is approximately consistent with that of the old expert, and has a safe and controllable starting point.
[0052] Building upon this foundation, the system uses the strategy output of the old expert under the same input conditions as the soft objective to perform knowledge distillation on the new expert. It also incorporates process constraint tasks constructed from new product trial run data to efficiently fine-tune incremental parameters. This allows the new expert to quickly learn the differences in the new product and acquire usable assembly strategies while maintaining reasonable behavioral boundaries, thus achieving rapid model transfer. If the transferred model encounters further changes in operating conditions on the production line, it can be further adapted to the new conditions through online evolution methods.
[0053] like Figure 5 As shown in the figure, in this embodiment, the method steps for quickly generating a model adapted to the new task through the model migration module are as follows.
[0054] 31) Task decomposition: Decompose old tasks and new tasks hierarchically to obtain the corresponding atomic skill sequences.
[0055] 32) Expert Matching and Creation: Based on the atomic skill sequence of the new task, match the most similar old experts in the parallel expert network library. Create a new expert network corresponding to the new task based on a preset template. .
[0056] 33) Residual Initialization: Let the new expert parameters (weights, biases, normalization parameters, etc.) be... The old expert parameters were Incremental parameters The residual initialization formula is: Incremental parameters Initialize to zero or near-zero values to ensure that the output of the new expert is approximately consistent with that of the old expert in the early stages of migration, providing a safe and controllable initial behavioral boundary. This ensures that the following conditions are met in the early stages of migration: in: This represents the output increment of the new expert relative to the old expert.
[0057] 34) Knowledge Distillation and Fine-tuning: Set a small-scale trial operation data set for the new product Each sample must contain at least one state. The action vector output by the old expert The new expert's output action vector Maximum allowable error threshold Maximum permissible contact force Maximum allowed beat . , and The weights for the three types of penalties, , and This is the combination coefficient of the three losses.
[0058] (1) Distillation loss term: Let the new expert imitate the output of the old model, i.e., soft target.
[0059] (2) Task loss item: to make the new expert meet the process constraints of the new product.
[0060] (3) Incremental regularization: Only small-step updates are performed.
[0061] (4) Total function: Minimize the incremental parameters of the new expert by minimizing the loss function. Make minor adjustments: in: This is the distillation loss item; This is a task loss item; For incremental regularization terms; , and These are the combination coefficients for the three losses; and These are the action vectors output by the new expert and the action vectors output by the old expert, respectively. For error; The maximum allowable error threshold; For contact force; Maximum permissible contact force; For rhythm; The maximum permissible beat; , and The weights for the three types of penalties.
[0062] Specifically, the constraint loss term based on the new task process threshold. Construction method and self-supervised scoring function The construction method is consistent, but the variables are replaced with actions based on the new expert output. After execution, the accuracy error, contact force, and cycle time data obtained from prediction or simulation are compared and calculated with the process threshold of the new task.
[0063] IV. Cross-condition evolution and migration systems This embodiment also proposes a cross-condition evolution and migration system for an embodied intelligent assembly and adjustment model of complex equipment, including a model deployment and execution module, a model online evolution module, a model migration module, and a data management and sensing unit. The model deployment and execution module is used to deploy the embodied intelligent assembly and adjustment model network based on a hybrid expert architecture constructed as described above in this embodiment, and to execute assembly and adjustment tasks; the model online evolution module dynamically and incrementally updates the embodied intelligent assembly and adjustment model network to adapt to condition drift during production in real time; when the work task changes due to product replacement, the model migration module quickly generates a model adapted to the new task based on the skill modules in the existing model network; the data management and sensing unit is used to collect multimodal sensing data, manage process thresholds, and manage the experience playback buffer.
[0064] (1) Embodied Intelligent Assembly Model Network Based on Hybrid Expert Architecture This embodiment constructs an embodied intelligent assembly and adjustment model network that receives a composite input state vector and outputs robot motion control commands. It adopts a hybrid expert architecture, consisting of a shared feature extraction layer, a parallel expert network library, and a dynamic gating network. The shared feature extraction layer extracts common low-level physical features from multimodal data. The parallel expert network library comprises several parallel expert networks, which are structurally identical but parameter-independent sub-neural networks designed to differentiate into operation modules focused on processing specific product features or specific working condition skills. The dynamic gating network outputs a normalized probability weight distribution based on the input task semantics and physical state information, identifying active experts whose parameters need updating. In each control cycle, the composite input state vector is input into the deployed hybrid expert-based embodied intelligent assembly and adjustment model, expert weights are inferred, and the output motion commands of each expert are fused to form the final motion command. The composite input state vector consists of real-time acquired physical states and task semantics.
[0065] (2) Online model evolution module Establish a self-supervised evaluation mechanism: This mechanism compares real-time collected multimodal data and robot status data with preset process thresholds during industrial robot operations to generate reward signals indicating the success or failure of the assembly and adjustment task. The process thresholds are preset based on prior process knowledge and rules.
[0066] Establish an experience playback buffer: This buffer stores experience samples generated during robot assembly and adjustment. New samples are continuously written to the buffer in chronological order, and old or low-value samples are deleted according to a preset elimination strategy when the buffer capacity reaches a set threshold, thus maintaining adaptability to the current working conditions. Experience samples are state-action-evaluation records collected by the robot during assembly and adjustment operations, including at least the current state vector, executed actions, and quality indicators obtained from self-supervised evaluation.
[0067] Online model evolution triggering mechanism: The mean of self-supervised evaluation values cached in the experience replay buffer is statistically analyzed in real time, and the entropy of the gated output distribution is calculated; when any one of the evaluation indicators is met, online model evolution is triggered. Evolution evaluation indicators: (1) The mean evaluation value is consistently lower than the preset threshold; (2) The gated entropy increases over a long period of time, indicating that the expert selection is uncertain. The threshold is adaptively determined by process specifications, safety limits or historical statistics.
[0068] The online model evolution module comprises three steps: selective updating, lightweight incremental optimization, and dynamic learning rate adjustment. It collects data cached in the experience replay buffer while running the model. During controlled periods, it selectively generates parameter update masks based on statistical results of gating weights, and then performs incremental optimization on the updatable masks. The parameter update mask limits the range of gradient calculation and parameter write-back. Dynamic learning rate adjustment scales the step size to ensure the stability of the production line operation.
[0069] (3) Model transfer module The model transfer module comprises three steps: task decomposition, residual initialization, and knowledge distillation. When product changes lead to changes in operating conditions, the new task is decomposed, and a model adapted to the new operating conditions is quickly obtained by combining residual initialization and knowledge distillation.
[0070] V. Beneficial Effects Compared with the prior art, the present invention has the following beneficial effects: (1) By using a self-supervised evaluation and reward mechanism, combined with clear online evolution trigger criteria, the model evolution has an objective basis, avoiding reliance on offline large-scale retraining or human experience judgment.
[0071] (2) A gated selective update mechanism is adopted, combined with lightweight incremental optimization, to update only the experts involved in the decision-making and freeze the rest of the experts, thereby reducing the risk of old skills being covered and maintaining the stability of the original capabilities; in conjunction with dynamic learning rate adjustment, online evolution of working condition drift is achieved.
[0072] (3) In the face of product changes, the new working condition adaptation process is transformed from zero learning to the combination and optimization of mature skills through hierarchical task decomposition and knowledge distillation. This enables new experts to quickly obtain usable strategies under trial operation data, without disturbing the capabilities of old experts, significantly shortening the new product introduction cycle and reducing data collection costs.
[0073] The above-described embodiments are merely preferred embodiments provided to fully illustrate the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are all within the scope of protection of the present invention. The scope of protection of the present invention is defined by the claims.
Claims
1. A method for cross-condition evolution and transfer of a complex equipment embodied intelligent assembly and adjustment model, characterized in that: Includes the following steps: Step 1: Construct and deploy an embodied intelligent assembly and adjustment model network based on a hybrid expert architecture. The embodied intelligent assembly and adjustment model network receives a composite input state vector consisting of real-time physical state and task semantics. and output robot motion control commands. ; Step 2: When the embodied intelligent assembly and adjustment model network is running at the actual assembly and debugging station, the embodied intelligent assembly and adjustment model network is dynamically and incrementally updated through the online model evolution module to adapt to the working condition drift in the production process in real time. Step 3: When the work task changes due to product replacement, the model migration module quickly generates a model adapted to the new task based on the skill modules in the existing embodied intelligent assembly and adjustment model network.
2. The method for cross-condition evolution and migration of the intelligent assembly and adjustment model of complex equipment according to claim 1, characterized in that: In step one, the embodied intelligent assembly and adjustment model network based on a hybrid expert architecture includes: A shared feature extraction layer is used to process the input multimodal sensing data and extract common low-level physical features; The parallel expert network library consists of multiple expert networks with the same structure but independent parameters. Each expert network is used to encode the assembly and adjustment skills of a specific type of product or working condition. Dynamic gating network, used to control the composite input state vector It outputs a normalized weight distribution to dynamically select and activate one or more expert networks, and merges the outputs of the selected expert networks to form the final action instruction.
3. The cross-condition evolution and migration method for the intelligent assembly and adjustment model of complex equipment according to claim 1, characterized in that: In step two, the online model evolution module includes: The self-supervised evaluation mechanism is used to compare the real-time collected multimodal data and ontological state data with preset process thresholds to generate reward signals that characterize the success or failure of the assembly and adjustment task. The experience replay buffer is used to store experience samples consisting of state, action and evaluation signals in time sequence, and manage the samples according to a preset elimination strategy. An online evolution triggering mechanism is used to monitor in real time the statistical characteristics of the evaluation signals in the experience replay buffer and the entropy value of the output distribution of the dynamic gating network, and triggers model updates when the preset evolution evaluation index is met. The selective update unit is used to identify the active expert set and generate a parameter update mask based on the weight distribution output by the dynamic gating network, perform lightweight incremental optimization only on the parameters of the active expert set, and freeze the parameters of inactive experts.
4. The method for cross-condition evolution and migration of the intelligent assembly and adjustment model of complex equipment according to claim 3, characterized in that: The self-monitoring evaluation mechanism is as follows: based on product model... With process type Determine process thresholds, including maximum permissible contact force. Maximum permissible torque Maximum error threshold With the maximum allowable beat threshold At the end of each control cycle or process stage, calculate the assembly and adjustment accuracy error. Contact force With rhythm And an evaluation signal is generated based on a self-supervised scoring function; the self-supervised scoring function is expressed as: in: , , and It is a non-negative weighting coefficient used to assess the contribution of three types of indicators: balance accuracy, force control, and cycle time.
5. The method for cross-condition evolution and migration of the intelligent assembly and adjustment model of complex equipment according to claim 3, characterized in that: The online evolution triggering mechanism is as follows: Real-time statistics of the experience replay buffer China recently The mean of the evaluation of each sample: in: The function value of the self-supervised scoring function; Calculate the entropy of the output distribution of the dynamic gating network: in: The output distribution entropy is gated; The mean of the distribution entropy; For gating networks to the first The output weights of each expert; The total number of experts available; When satisfied or The model is triggered to evolve online at specific times, where: and The preset threshold is determined adaptively based on the statistical characteristics of historical production data, or set according to process specifications and safety limits.
6. The method for cross-condition evolution and migration of the intelligent assembly and adjustment model of complex equipment according to claim 3, characterized in that: In the selective update unit, the weights output by the current dynamic gating network are used as the basis for the update. Determine the set of active experts: in: The preset activation threshold; These are the weight coefficients of the gating network; Generate parameter update mask only for the set of active experts. The expert network parameters are optimized using lightweight incremental optimization, while other expert network parameters are frozen. In the selective update unit, lightweight incremental optimization is implemented using a low-rank adaptation method, and its objective function is: in: Current model parameter set; This represents the parameter increment relative to the previous version; The coefficient is a non-negative positive regularity coefficient; Indicates an experience-based replay buffer The reinforcement learning loss term is calculated from the sampled data; the learning rate is adaptively adjusted based on the recent evaluation improvement to improve the response speed to operating condition drift.
7. The method for cross-condition evolution and migration of the intelligent assembly and adjustment model of complex equipment according to claim 1, characterized in that: In step three, the method for quickly generating a model adapted to the new task through the model transfer module consists of the following steps: 31) Task decomposition: Decompose old tasks and new tasks hierarchically to obtain the corresponding atomic skill sequences; 32) Expert Matching and Creation: Based on the atomic skill sequence of the new task, match the most similar old experts in the parallel expert network library. Create a new expert network corresponding to the new task based on a preset template. ; 33) Residual initialization: Let the old expert parameters be... The new expert parameters are initialized to : Incremental parameters Initialize to zero or near-zero values so that the following conditions are met during the initial migration: in: This represents the output increment of the new expert relative to the old expert. 34) Knowledge Distillation and Fine-tuning: Minimizing the incremental parameters of new experts by minimizing the loss function. Make minor adjustments: in: This is the distillation loss item; This is a task loss item; For incremental regularization terms; , and These are the combination coefficients for the three losses; and These are the action vectors output by the new expert and the action vectors output by the old expert, respectively. For error; The maximum allowable error threshold; For contact force; Maximum permissible contact force; For rhythm; The maximum permissible beat; , and The weights for the three types of penalties.
8. The method for cross-condition evolution and migration of the intelligent assembly and adjustment model of complex equipment according to claim 7, characterized in that: In step 34), the constraint loss term based on the new task process threshold... Construction method and self-supervised scoring function The construction method is consistent, but the variables are replaced with actions based on the new expert output. After execution, the accuracy error, contact force, and cycle time data obtained from prediction or simulation are compared and calculated with the process threshold of the new task.
9. The method for cross-condition evolution and migration of the intelligent assembly and adjustment model of complex equipment according to claim 1, characterized in that: The composite input state vector The physical state data comes from the RGB-D camera fixed to the robot's wrist, the six-dimensional force / torque sensor at the end effector, and the robot's encoder. The task semantic data comes from the work order information issued by the manufacturing execution system.
10. A cross-condition evolution and migration system for a complex equipment embodied intelligent assembly and adjustment model, characterized in that: include: The model deployment and execution module is used to deploy the embodied intelligent assembly and debugging model network based on the hybrid expert architecture constructed by the method of any one of claims 1-9, and to execute the assembly and debugging tasks; The online model evolution module dynamically and incrementally updates the embodied intelligent assembly and adjustment model network to adapt to the working condition drift during the production process in real time. The model migration module can quickly generate a model adapted to the new task based on the skill modules in the existing model network when the job task changes due to product replacement. The data management and sensing unit is used to collect multimodal sensing data, manage process thresholds, and establish an experience playback buffer.