A dual-arm collaborative training method and control method based on multi-modal sensor fusion, a collaborative iterative training method and related devices
Patent Information
- Application Number
- CN202611083400.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-18
AI Technical Summary
[0003]然而,上述常规方案存在以下缺陷:依赖视觉引导时,接触工件后视线易受遮挡或反光干扰,无法感知接触力动态;依赖力反馈控制时,远距离阶段缺乏目标位置信息,缺乏导航能力,碰撞风险高;以固定阈值作为任务阶段切换条件时,阶段边界处控制参数突变,容易因工件位置偏移或外力扰动而发生末端抖动或误切换,导致装配成功率低;无法基于执行的任务对双臂角色进行灵活分工
[0013] Compared to existing technologies, the present application provides a dual-arm collaborative training method, control method, collaborative iterative training method, and related apparatus based on multimodal sensor fusion. In the offline phase, multimodal sensor data is segmented and each modal encoder is trained to obtain a unified state representation. This is then used to train a decision model and contact strategy, which are then frozen in a model library. In the online control phase, each module calls the model library, encodes and fuses real-time multimodal data into a current unified state representation, and the decision model outputs a decision result including role assignment and relative pose constraints. In the contact phase, an end-effector correction is output. The control and solution modules calculate the control error, generate end-effector control quantities to drive execution, record logs, and iteratively retrain. Thus, multimodal fusion overcomes the limitations of single-sensor operation, dual-arm collaboration and internal force coordination achieve dynamic balance, phase switching improves continuity, fine-tuning of the contact strategy enhances end-effector assembly accuracy, and log feedback and iterative retraining establish a closed-loop update mechanism.
Smart Images

Figure CN122769980A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control, and more specifically, to a dual-arm collaborative training method, control method, collaborative iterative training method, and related apparatus based on multimodal sensor fusion. Background Technology
[0002] In existing technologies, conventional solutions for controlling dual-armed robots often employ single-modal sensor data. For example, environmental or wrist cameras are used for visual servo guidance, aligning the arms with a target for insertion. Another approach is to install a six-dimensional torque sensor at the end of the robotic arm, using a preset force threshold to achieve compliant control during the contact phase. Furthermore, during arm movement, static switching conditions are used to determine whether the control process enters different phases. For instance, distance or force thresholds are used as phase switching conditions. The division of labor between the two arms is often preset.
[0003] However, the above conventional solutions have the following drawbacks: when relying on visual guidance, the line of sight is easily obstructed or interfered with by reflections after contact with the workpiece, making it impossible to perceive the dynamics of the contact force; when relying on force feedback control, there is a lack of target position information and navigation capability at long distances, resulting in a high risk of collision; when using fixed thresholds as task stage switching conditions, the control parameters change abruptly at the stage boundary, which can easily cause end jitter or erroneous switching due to workpiece position shift or external force disturbance, leading to a low assembly success rate; and it is impossible to flexibly divide the roles of the two arms based on the tasks being performed.
[0004] In summary, existing technologies cannot achieve more flexible and precise control of robots with two arms, thus reducing the efficiency and safety of task execution for these robots. Summary of the Invention
[0005] The purpose of this application is to provide a dual-arm collaborative training method, control method, collaborative iterative training method and related device based on multimodal sensor fusion, which can improve the accuracy of dual-arm control and the continuity of stage switching.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, embodiments of this application provide a dual-arm collaborative training method based on multimodal sensor fusion, including: The acquisition and calibration module acquires multimodal sensing data corresponding to the demonstration operation and calibrates the multimodal sensing data to obtain multimodal demonstration trajectory data. The data synchronization and preprocessing module performs stage segmentation on the multimodal demonstration trajectory data to obtain corresponding trajectory segments; each trajectory segment corresponds to a work stage in the multimodal demonstration trajectory data; The multimodal coding module trains the encoder for each type of modal data separately to obtain the encoded feature vector of the corresponding type of modal data encoder; The multimodal coding module performs weighted fusion of all the encoded feature vectors based on the weights corresponding to each of the encoded feature vectors to obtain a unified state representation; the unified state representation is used to characterize the multimodal fusion features of the dual-arm robot. The sub-task sequence decision module trains the sub-task sequence decision model based on the unified state representation; the sub-task sequence decision model is used to define the control strategy for both arms based on the unified state representation; The contact phase micro-action strategy module trains the contact phase micro-action strategy based on the unified state representation, so that the contact phase micro-action strategy outputs end control commands when in the contact-related phase. The multimodal coding module freezes the coding parameters corresponding to all the trained modal data encoders into the modal data encoder library.
[0007] Secondly, embodiments of this application provide a dual-arm cooperative control method based on multimodal sensor fusion, comprising: The subtask sequence decision module calls the subtask sequence decision model in the decision model library, inputs the main task and the current unified state representation into the subtask sequence decision model, and obtains the subtask decision result; the main task is used to indicate the assembly task of the dual-arm robot; the current unified state represents the current multimodal fusion feature of the dual-arm robot; When the stage label indicates that the contact stage is in the contact stage, the micro-action strategy module obtains the end-effector control correction amount based on the current unified state representation; the contact stage includes the pre-insertion stage, the insertion stage, and the pressing stage; The dual-arm collaborative control and calculation module calculates the dual-arm control error based on the sub-task decision results and real-time multimodal data; The dual-arm collaborative control and calculation module generates end-of-arm control quantities based on the current unified state representation, the sub-task decision results, the end-of-arm control correction quantity, and the dual-arm control error. The dual-arm collaborative control and calculation module applies amplitude limits to the end-effector control quantity and converts the limited end-effector control quantity into joint control commands for both arms. The dual-arm collaborative control and calculation module sends the joint control commands to the dual-arm robot, driving the dual arms to perform assembly operations. The online assessment and rollback module monitors abnormal indicators in real time and performs exception handling operations when abnormal indicators are triggered.
[0008] Thirdly, embodiments of this application provide a dual-arm collaborative iterative training method based on multimodal sensor fusion, including: The acquisition and calibration module acquires multimodal sensing data corresponding to the demonstration operation and calibrates the multimodal sensing data to obtain multimodal demonstration trajectory data. The data synchronization and preprocessing module performs stage segmentation on the multimodal demonstration trajectory data to obtain corresponding trajectory segments; each trajectory segment corresponds to a work stage in the multimodal demonstration trajectory data; The multimodal coding module trains the encoder for each type of modal data separately to obtain the encoded feature vector of the corresponding type of modal data encoder; The multimodal coding module performs weighted fusion of all the encoded feature vectors based on the weights corresponding to each encoded feature vector to obtain a unified state representation; the unified state representation is used to characterize the multimodal fusion features of the dual-arm robot. The sub-task sequence decision module trains the sub-task sequence decision model based on the unified state representation; the sub-task sequence decision model is used to define the control strategy for both arms based on the unified state representation; The contact phase micro-action strategy module trains the contact phase micro-action strategy based on the unified state representation, so that the contact phase micro-action strategy outputs end control commands when in the contact-related phase. The multimodal coding module freezes all the trained modal data encoders into the modal data encoder library.
[0009] The subtask sequence decision module calls the subtask sequence decision model in the decision model library, inputs the main task and the current unified state representation into the subtask sequence decision model, and obtains the subtask decision result; the main task is used to indicate the assembly task of the dual-arm robot; the current unified state represents the current multimodal fusion feature of the dual-arm robot. When the stage label indicates that the contact stage is in the contact stage, the micro-action strategy module obtains the end-effector control correction amount based on the current unified state representation; the contact stage includes the pre-insertion stage, the insertion stage, and the pressing stage; The dual-arm collaborative control and calculation module calculates the dual-arm control error based on the sub-task decision results and real-time multimodal data; The dual-arm collaborative control and calculation module generates end-of-arm control quantities based on the current unified state representation, the sub-task decision results, the end-of-arm control correction quantity, and the dual-arm control error. The dual-arm collaborative control and calculation module applies amplitude limits to the end-effector control quantity and converts the limited end-effector control quantity into joint control commands for both arms. The dual-arm collaborative control and calculation module sends the joint control commands to the dual-arm robot, driving the dual arms to perform assembly operations. The online evaluation and rollback module records execution logs, updates the training sample pool based on the execution logs, and iteratively retrains the modal data encoder, the subtask sequence decision model, and / or the contact phase micro-action strategy based on the updated training sample pool.
[0010] Fourthly, embodiments of this application provide a dual-arm collaborative training device based on multimodal sensor fusion, comprising: The acquisition and calibration module is used to acquire multimodal sensing data corresponding to the demonstration operation, and to calibrate the multimodal sensing data to obtain multimodal demonstration trajectory data. The data synchronization and preprocessing module is used to segment the multimodal demonstration trajectory data into stages to obtain corresponding trajectory segments; each trajectory segment corresponds to a work stage in the multimodal demonstration trajectory data. A multimodal coding module is used to train each type of modal data encoder separately to obtain the coding feature vector of the corresponding type of modal data encoder; based on the weights corresponding to each coding feature vector, all coding feature vectors are weighted and fused to obtain a unified state representation; the unified state representation is used to characterize the multimodal fusion features of the dual-arm robot; and the coding parameters corresponding to all trained modal data encoders are frozen to the modal data encoder library. The subtask sequence decision module is used to train the subtask sequence decision model based on the unified state representation; the subtask sequence decision model is used to define the control strategy for both arms based on the unified state representation. The contact phase micro-action strategy module is used to train the contact phase micro-action strategy based on the unified state representation, so that the contact phase micro-action strategy outputs end control commands when in the contact-related phase.
[0011] Fifthly, embodiments of this application provide a dual-arm collaborative control device based on multimodal sensor fusion, comprising: The subtask sequence decision module is used to call the subtask sequence decision model in the decision model library, input the main task and the current unified state representation into the subtask sequence decision model, and obtain the subtask decision result; the main task is used to indicate the assembly task of the dual-arm robot; the current unified state represents the current multimodal fusion feature of the dual-arm robot. The contact phase micro-action strategy module is used to obtain the end-effector control correction amount based on the current unified state representation when the phase label indicates that it is in the contact phase; the contact phase includes the pre-insertion phase, the insertion phase, and the pressing phase; The dual-arm collaborative control and calculation module is used to calculate the dual-arm control error based on the sub-task decision results and real-time multimodal data; generate an end-effector control quantity based on the current unified state representation, the sub-task decision results, the end-effector control correction amount, and the dual-arm control error; apply an amplitude limit to the end-effector control quantity and convert the limited end-effector control quantity into joint control commands for the dual arms; and send the joint control commands to the dual-arm robot to drive the dual arms to perform assembly operations. The online assessment and rollback module is used to monitor abnormal indicators in real time and perform abnormal handling operations when abnormal indicators are triggered.
[0012] Sixthly, embodiments of this application provide a program product that, when executed by a processor, implements the method as described in any one of the first, second, or third aspects above.
[0013] Compared to existing technologies, the present application provides a dual-arm collaborative training method, control method, collaborative iterative training method, and related apparatus based on multimodal sensor fusion. In the offline phase, multimodal sensor data is segmented and each modal encoder is trained to obtain a unified state representation. This is then used to train a decision model and contact strategy, which are then frozen in a model library. In the online control phase, each module calls the model library, encodes and fuses real-time multimodal data into a current unified state representation, and the decision model outputs a decision result including role assignment and relative pose constraints. In the contact phase, an end-effector correction is output. The control and solution modules calculate the control error, generate end-effector control quantities to drive execution, record logs, and iteratively retrain. Thus, multimodal fusion overcomes the limitations of single-sensor operation, dual-arm collaboration and internal force coordination achieve dynamic balance, phase switching improves continuity, fine-tuning of the contact strategy enhances end-effector assembly accuracy, and log feedback and iterative retraining establish a closed-loop update mechanism.
[0014] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A flowchart illustrating a dual-arm collaborative training method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 2A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 3 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 4 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 5 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 6 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 7 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 8 A flowchart illustrating a dual-arm collaborative control method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 9 A flowchart illustrating another dual-arm collaborative control method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 10 A flowchart illustrating another dual-arm collaborative control method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 11 A flowchart illustrating another dual-arm collaborative control method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 12 A flowchart illustrating another dual-arm collaborative control method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 13 A flowchart illustrating another dual-arm collaborative control method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 14 A flowchart illustrating another dual-arm collaborative control method based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 15 This invention provides a dual-arm collaborative iterative training method based on multimodal sensor fusion. Figure 16 A schematic diagram of a dual-arm collaborative device based on multimodal sensor fusion provided in an embodiment of the present invention; Figure 17 This is a schematic diagram of another dual-arm collaborative device based on multimodal sensor fusion provided in an embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0018] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0019] In the description of this application, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set" and "connection" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0020] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0021] In the field of industrial automation, robotic assembly technology is one of the core components for achieving precision manufacturing. Robots with dual arms, in particular, are more complex to control. Currently, existing technological solutions for robotic arm assembly operations mainly fall into the following categories: The first category is assembly methods based on fixed trajectory planning. This method generates a fixed end-effector trajectory through teaching or path planning algorithms in the offline phase, and the robot executes the pre-set trajectory point by point in the online phase. This type of method relies on high-precision tooling and environmental calibration, assuming that the workpiece pose remains strictly consistent in each assembly. When there are initial pose deviations, workpiece tolerance changes, or environmental disturbances, the robot cannot perceive the deviation between the actual state and the expected trajectory, leading to collisions, interference fit failures, or damage to assembled parts.
[0022] The second category is assembly methods based on single vision positioning. This method uses an environmental camera or wrist camera to acquire workpiece images and adjusts the end effector pose in real time through visual servoing or template matching algorithms to align the robotic arm end effector with the target assembly hole or shaft. During the approach and alignment phases, visual positioning provides effective guidance information; however, once the robotic arm end effector enters the insertion or pressing phase, the vision sensor is affected by factors such as end effector occlusion, workpiece surface reflection, and insufficient depth resolution due to small assembly gaps, failing to provide reliable pose feedback. At this point, the visual servoing system loses effective input, control accuracy drops sharply, leading to insertion failure or repeated collisions.
[0023] The third category: Assembly methods based on single-arm force control. This method installs a six-dimensional force / torque sensor at the end of a single arm. Through force-position hybrid control or impedance control strategies, the robotic arm adjusts its direction and speed based on force feedback during contact. Single-arm force control can adapt to workpiece posture deviations to a certain extent, achieving compliant insertion. However, in dual-arm collaborative assembly scenarios, there is mechanical coupling between the two robotic arms: the insertion movement of the active arm changes the force state of the supporting arm, and the constraint reaction force of the supporting arm also affects the contact force distribution of the active arm. Existing single-arm force control methods simplify the other arm as a rigid environment or fixed support, ignoring the internal force transmission and collaborative constraint relationship between the two arms, leading to unreasonable force distribution, mutual restraint, or even jamming during the two-arm cooperation process.
[0024] The fourth category: Assembly methods based on imitation learning. This method records the demonstration trajectory of a human operator and uses imitation learning algorithms such as behavior cloning to train a policy network, enabling the robot to reproduce the expert's actions. This type of method can learn action patterns such as approach and alignment at a macroscopic level, but in the final few millimeters of insertion, mating, and pressing stages, there is a highly nonlinear coupling relationship between contact force and pose, making it difficult for simply imitating the demonstration trajectory to cover all possible deviations. Furthermore, traditional imitation learning often separates teaching, training, and execution, failing to continuously feed back successful experiences and failure samples accumulated during online execution into the model training, resulting in insufficient long-term generalization ability. When facing new workpieces or changing working conditions, it is necessary to re-collect demonstration data.
[0025] Clearly, the existing technologies described above present the following problems when controlling robotic equipment with two arms: Problem 1: Single-modal perception is difficult to cover the entire assembly process. Visual modality is suitable for long-distance guidance and alignment stages, but it is limited by occlusion and resolution during contact, insertion, and pressing stages; force modality is suitable for post-contact perception, but it lacks long-distance guidance capabilities, and using it alone will result in a long search path and a high probability of collision.
[0026] Problem 2: Insufficient description of the cooperative relationship between the two arms. Existing methods typically simplify the assembly task of the two arms into a "active arm operation and support arm fixation" model, lacking an explicit description of the relative pose constraints between the two arms, and also failing to consider the internal force coordination mechanism of the two arms after contact. When the active arm advances and inserts, the support arm needs to provide stable constraints and appropriate compliant yielding at the same time. Existing technology cannot achieve a dynamic balance between these two, resulting in mutual restraint between the two arms, uncontrolled contact forces, or assembly jamming.
[0027] Problem 3: Poor continuity of stage switching. In existing assembly control methods, the switching between stages such as approach, alignment, and insertion often relies on preset rules based on distance thresholds, force thresholds, or pose errors. When there are environmental disturbances or changes in workpiece tolerances, the preset thresholds may trigger stage switching prematurely or delayed, causing abrupt changes in control parameters and action commands at stage boundaries, resulting in end-effector jitter, incorrect switching, or insertion failure. Furthermore, stage switching rules are usually hard-coded and fixed, making it difficult to adapt to different assembly objects and working conditions.
[0028] Question 4: Insufficient high-precision assembly capability at the end effector. Traditional imitation learning or trajectory planning methods can reproduce macroscopic motion trajectories, but lack targeted optimization for the last few millimeters of insertion, mating, and pressing actions. During the contact phase, even slight pose deviations can lead to a sharp increase in contact force. Existing technologies lack the ability to finely adjust the end effector pose and compliance control parameters, easily resulting in problems such as post-contact jitter, unstable search direction, or low assembly success rate.
[0029] Question 5: Lack of a closed-loop update mechanism for training and execution. Existing solutions often separate offline teaching, model training, and online execution into independent processes. Successful corrections, abnormal events, and failure samples generated during online execution are not effectively recorded and utilized, preventing the model from continuously improving with the accumulation of real-world conditions. When faced with new workpiece deviations or environmental changes, the system's generalization ability is insufficient, requiring manual re-teaching and training, resulting in high maintenance costs.
[0030] To address the aforementioned technical deficiencies and problems, this invention proposes a dual-arm collaborative assembly control method that utilizes offline training, online control, and iterative optimization. The core of this method lies in the decoupling of offline model training and online real-time inference, and the realization of closed-loop iteration through execution log feedback.
[0031] In the offline training phase: visual, force, and state encoders are trained to map heterogeneous sensor data into a unified state representation; a sub-task sequence decision model is trained to dynamically output sub-objectives, dual-arm role assignments, and stage labels; and a contact stage micro-motion strategy is trained using a "imitation learning pre-training + reinforcement learning fine-tuning" strategy to enable it to master end-effector pose fine-tuning and compliant control parameters. After training, each model is frozen in the model library.
[0032] During the online control phase: Real-time multimodal data is synchronously collected in each control cycle to construct a unified state representation; the frozen decision model and contact strategy are invoked to output sub-task decision results and end-point control corrections; based on relative pose error and lateral internal force error, end-point control quantities are generated through a cooperative control law and converted into joint commands using damped least squares Jacobian inverse conversion; simultaneously, abnormal indicators are monitored in real time to trigger backoff or realignment. Online execution logs are fed back into the offline training sample pool to trigger periodic iterative retraining, enabling continuous improvement of system capabilities.
[0033] Optionally, this application involves two stages: offline training and online control. The following section provides a possible implementation method for offline training. Specifically... Figure 1 This is a flowchart illustrating a dual-arm collaborative training method based on multimodal sensor fusion provided in an embodiment of the present invention. (See attached diagram.) Figure 1 The method includes: Step 100: The acquisition and calibration module acquires the multimodal sensing data corresponding to the demonstration operation and calibrates the multimodal sensing data to obtain multimodal demonstration trajectory data.
[0034] Step 101: The data synchronization and preprocessing module performs stage segmentation on the multimodal demonstration trajectory data to obtain the corresponding trajectory segments.
[0035] Each trajectory segment corresponds to a work phase in the multimodal demonstration trajectory data.
[0036] Step 102: The multimodal coding module trains the encoder for each type of modal data to obtain the coding feature vector of the corresponding type of modal data encoder.
[0037] Step 103: The multimodal coding module performs weighted fusion of all coding feature vectors based on the weights corresponding to each coding feature vector to obtain a unified state representation.
[0038] Among them, the unified state representation is used to characterize the multimodal fusion features of the dual-arm robot; Step 104: The subtask sequence decision module trains the subtask sequence decision model based on the unified state representation.
[0039] The subtask sequence decision model is used to define the control strategy for both arms based on a unified state representation.
[0040] Step 105: The contact phase micro-action strategy module trains the contact phase micro-action strategy based on the unified state representation, so that the contact phase micro-action strategy outputs end control commands when in the contact-related phase.
[0041] The end-effector control command may include: end-effector position fine-tuning amount, attitude fine-tuning amount, and compliance control parameters.
[0042] Step 106: The multimodal coding module freezes the coding parameters corresponding to all trained modal data encoders into the modal data encoder library.
[0043] The dual-arm collaborative training method based on multimodal sensor fusion provided in this invention solves the problem that a single modality cannot cover the entire assembly process by collecting and calibrating multimodal sensor data, including environmental vision, wrist vision, six-dimensional torque at the end of both arms, and joint states. By segmenting the demonstration trajectory into stages and training a sub-task sequence decision model, sequential decision-making for each operation stage is achieved in a data-driven manner, avoiding abrupt stage switching caused by fixed thresholds. By training various modal encoders separately and weighting and fusing them into a unified state representation, the micro-motion strategy in the contact stage outputs end-position fine-tuning, attitude fine-tuning, and compliance control parameters based on multimodal fusion features, improving the compliance and success rate of high-precision assembly at the end. By freezing the encoder parameters into an encoder library, a multimodal feature extraction capability that can be directly called for online control is provided.
[0044] Optionally, the following provides a possible implementation method for the acquisition and coordinate system calibration of multimodal sensing data. Specifically, in Figure 1 On this basis, Figure 2 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 2 Step 100 includes: Step 100-1: The data acquisition and calibration module acquires multimodal sensor data corresponding to the dual-arm collaborative assembly demonstration operation.
[0045] The multimodal sensing data may include: environmental visual data, wrist visual data, six-dimensional torque data of the ends of both arms, and joint status data of both arms.
[0046] Optionally, for a demonstration of dual-arm collaborative assembly, one possible implementation is as follows: the operator completes the demonstration operation using dual master remote control devices, a force feedback handle, or a digital twin teaching interface. (At the sampling time...) Record visual data (This may include, but is not limited to, environmental visual data and wrist visual data), six-dimensional torque data at the ends of both arms, and joint state data of both arms. Arm joint status data should include at least joint position. Joint velocity Drive current , terminal pose and gripper state .
[0047] Step 100-2: The acquisition and calibration module performs coordinate system calibration on each sensor and robotic arm involved in the multimodal sensing data to obtain multimodal demonstration trajectory data.
[0048] One possible implementation of this coordinate system calibration is as follows: The transformation relationships between the environmental camera coordinate system, wrist camera coordinate system, left arm base coordinate system, right arm base coordinate system, tool center point (TCP) coordinate system, and world coordinate system are determined through calibration operations. The calibration method can employ hand-eye calibration or Zhang Zhengyou calibration based on a calibration board to obtain the homogeneous transformation matrix between each coordinate system.
[0049] Optionally, the following process parameters are also set according to the assembly object and the controller cycle: Target assembly cycle limit, contact determination threshold Locking depth threshold Threshold for determining jamming , rewind distance and safety force threshold These parameters will serve as the baseline values for subsequent contact determination, anomaly detection, and rollback operations.
[0050] The target assembly cycle limit is used to set the maximum number of control cycles allowed for a single assembly task, serving as a benchmark for assembly timeout. If assembly is not completed before reaching this limit, subsequent exception handling (such as requesting manual intervention) is triggered.
[0051] This contact determination threshold This is used to determine that the end has made contact with the workpiece and generate a contact tag when the component force in at least one direction of the six-dimensional torque data at the end of the dual arms exceeds the threshold. =1, which determines whether to proceed to the contact-related stage.
[0052] The lock depth threshold This is used to define the target insertion depth upon completion of assembly. Subsequent step 204 uses this to analyze whether the current sub-target has been achieved; subsequent step 211-1 uses this to determine whether the insertion depth has failed to effectively advance towards the threshold within multiple cycles, thereby identifying bottlenecks.
[0053] The threshold for determining the jamming This is used when the lateral internal force error calculated in subsequent step 206-2 exceeds this threshold, and continues... During each control cycle, the "lateral internal force exceeds limit" abnormal flag is triggered.
[0054] The rollback distance Used after an exception is triggered, the end effector moves in the reverse direction of insertion. The fixed retraction distance is used to release abnormal contact force, disengage from the stuck state, and create space for re-attempting insertion.
[0055] This safety force threshold This is the absolute safety limit for the six-dimensional torque at the end of the dual arms. When the force component in any direction exceeds this value, it is directly determined to be a dangerous contact state, triggering the emergency protection operation in step 210-2 (such as immediately stopping the advance or manual intervention) to prevent damage to the robotic arm or workpiece.
[0056] Subsequently, after the data acquisition is completed, the acquisition and calibration module uses the homogeneous transformation matrix obtained from the calibration to unify the above data into the world coordinate system or the assembly coordinate system, thereby obtaining multimodal demonstration trajectory data.
[0057] Optionally, in order to achieve more precise control over the multimodal demonstration trajectory data, this application also introduces the concept of "stage labels" to identify the approach stage, alignment stage, pre-insertion stage, insertion stage, compression stage, and completion stage involved in the multimodal demonstration trajectory data.
[0058] Optionally, the stage label can be constructed as follows: Each multimodal demonstration trajectory data is represented as a structured label vector. .in Indicates the task result label. The dual-arm role label indicates whether either of the two robotic arms is the active arm or the support arm. Represents the stage label sequence. This represents an exception type label. The set of stage labels is uniformly defined as follows:
[0059] In this example, by calibrating a unified coordinate system, the coordinate transformation relationship between the environmental camera, wrist camera, and dual-arm base is explicitly determined, enabling the subsequent fusion of multimodal data under the same spatial reference and eliminating the impact of dual-arm base installation deviation on data consistency. At the same time, it provides a spatial mapping reference for the spatiotemporal alignment of force and vision, allowing observations of the same physical contact event in different sensors to be directly correlated, thus improving the accuracy of contact determination.
[0060] Furthermore, by configuring process parameters, the operator's domain experience is transformed into quantifiable judgment thresholds, providing a clear benchmark for subsequent anomaly detection and rollback operations.
[0061] In addition, the construction of stage labels enables each multimodal demonstration trajectory data to come with the metadata required for supervised training. It can be directly used for supervised training of subsequent subtask sequence decision models and contact stage micro-action strategies without additional manual annotation, which significantly reduces data preparation costs.
[0062] Optionally, to improve data quality and adapt to the needs of subsequent staged model training, this application provides a data preprocessing mechanism for time alignment, noise suppression, contact determination, and stage segmentation. Specifically, in Figure 1 On this basis, Figure 3 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 3 Step 101 includes: Step 101-1: The data synchronization and preprocessing module performs time alignment and noise suppression on the multimodal demonstration trajectory data.
[0063] Optionally, for the environmental visual data, wrist visual data, six-dimensional torque data of the two arm ends, and the joint state data of the two arms in the multimodal demonstration trajectory data, a unified time axis is used. Align them.
[0064] The visual frames of the visual data (environmental visual data, wrist visual data) can be aligned to the training sampling frequency using linear interpolation or nearest neighbor preservation.
[0065] In this study, the six-dimensional force / torque data at the ends of both arms and the components of the joint state data of both arms were subjected to first-order low-pass filtering or moving average filtering, respectively. Let the filtered six-dimensional force / torque vector be denoted as... .
[0066] Step 101-2: The data synchronization and preprocessing module performs contact determination on the time-aligned and noise-suppressed multimodal demonstration trajectory data.
[0067] Optionally, the contact determination is used to determine whether the ends of the two arms have made contact with the workpiece.
[0068] Specifically, referring to the previous example, based on the continuous The force threshold, relative pose error, and insertion depth change rate within each sampling period are used to determine the contact tag. The trajectory is then segmented into a set of stage labels. middle.
[0069] Optionally, to reduce false positives, the contact determination preferably satisfies the following conditions simultaneously: the axial force exceeds a threshold, the rate of change of the lateral force exceeds a threshold, or the derivative of the insertion depth changes significantly.
[0070] Step 101-3: The data synchronization and preprocessing module divides the multimodal demonstration trajectory data into trajectory segments corresponding to multiple operation stages based on the contact determination results.
[0071] The operation phase includes at least the approach phase, alignment phase, pre-insertion phase, insertion phase, and pressing phase.
[0072] Optionally, trajectory segments corresponding to multiple operation stages are used as training samples. These training samples undergo coordinate normalization, image enhancement, minor attitude perturbation, sensor noise simulation, and failure segment retention. Furthermore, failed trajectories are not deleted but are retained in the sample pool as outlier samples.
[0073] In this example, by unifying the time axis alignment and using linear interpolation, multimodal demonstration trajectory data with different sampling rates are synchronized to the same time-series reference, eliminating timing deviations caused by differences in sensor sampling. Low-pass filtering or moving average filtering is applied to force / torque and joint states to remove high-frequency environmental noise while retaining valid contact signals, improving the accuracy of subsequent contact determination. Based on the contact determination, the continuous trajectory is segmented into different trajectory segments, enabling the subsequent multimodal coding module and sub-task sequence decision model to obtain targeted training samples.
[0074] In addition, data augmentation operations such as coordinate normalization, mirror enhancement, attitude perturbation, and sensor noise simulation expand the sample diversity, while the strategy of retaining failed segments allows anomalous samples to participate in training, which significantly improves the ability of the sub-task sequence decision model and the contact strategy model to identify and handle rare anomalous working conditions.
[0075] Optionally, this application provides a multimodal encoder training mechanism that trains dedicated encoders for three types of data: visual, force, and state. This compresses high-dimensional images, temporal force / torque, and robotic arm configuration information into low-dimensional feature vectors, providing basic input for the construction of a unified state representation and stage-adaptive fusion. Specifically, in Figure 1 On this basis, Figure 4 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 4 Step 102 includes: Step 102-1: The multimodal coding module trains the visual encoder so that the visual encoder maps visual data into visual feature vectors.
[0076] The visual feature vector is used to extract visual spatial information.
[0077] Step 102-2: The multimodal coding module trains the force encoder so that the force encoder maps the six-dimensional torque data at the ends of the two arms into force feature vectors.
[0078] The force-feel feature vector is used to extract dynamic information about contact force.
[0079] Step 102-3: The multimodal coding module trains the state encoder so that the state encoder maps the joint state data of the two arms into state feature vectors.
[0080] The state feature vector is used to extract the configuration information of the dual-arm robotic arm.
[0081] Optionally, for visual encoders Image features are extracted using convolutional networks or visual Transformers to construct visual feature vectors. For force encoders... For length of One-dimensional temporal convolutional encoding is performed on the time window and its difference sequence of the six-dimensional torque data at the ends of the two arms to obtain the force feature vector. For the state encoder... Input the joint state data of the two arms (such as joint position, joint velocity, drive current, TCP pose and gripper state) into the multilayer perceptron to obtain the state feature vector.
[0082] Optionally, in combination with the three types of encoded feature vectors obtained in step 102 above, in order to solve the problem of the reliability difference of modal information in different assembly stages, this application provides a stage-adaptive weighted fusion mechanism. The weights of visual, force, and state feature vectors are dynamically generated according to the current operation stage and the contact determination result, and weighted fusion is performed. This makes visual information dominant in the approach and alignment stages and force information enhanced in the insertion and pressing stages, thereby obtaining a unified state representation that can completely represent the current working condition and providing fusion input for subsequent sub-task decisions and contact strategy invocation.
[0083] Optionally, this application provides a stage-adaptive weighted fusion mechanism to obtain a unified state representation that can fully characterize the current working condition, providing fusion input for subsequent sub-task decisions and contact strategy invocation. Specifically, in Figure 3 On this basis, Figure 5 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 5 Step 103 includes: Step 103-1: The multimodal coding module generates the weights corresponding to each coding feature vector based on the current operation stage and the contact determination result.
[0084] Step 103-2: The multimodal coding module fuses all encoded feature vectors according to their corresponding weights to obtain a unified state representation.
[0085] Optionally, the current operation stage is taken as the approach stage ( =Close), the contact determination result is: the contact label is not in contact ( =0).
[0086] The weight generation calculation example (I) in step 103-1 is as follows: Currently, we are in the approach phase, dominated by visual feature vectors. Assume... =0.70, =0.15, =0.15.
[0087] in, The weights corresponding to the visual feature vectors. The weights corresponding to the force-feel feature vectors, This is the state feature vector, and it satisfies the normalization constraint: .
[0088] At this point, the contact label is not in contact, and the weight corresponding to the force feature vector remains low.
[0089] The unified state representation expression for the weighted summation corresponding to step 103-2 is as follows:
[0090] in, These are visual features. The visual feature transformation matrix, This is the visual feature vector (512-dimensional) output by the visual encoder. This term maps image features to a unified state representation space.
[0091] This is a characteristic term of force perception. It is the force characteristic transformation matrix. This is the force feature vector (256-dimensional) output by the force encoder. This term maps the temporal features of force / torque to a unified state representation space.
[0092] State characteristic terms. It is the state characteristic transformation matrix. This is the state feature vector (128-dimensional) output by the state encoder. This term maps the robotic arm configuration information to a unified state representation space.
[0093] Stage tag embedded item. It is a stage-embedded transformation matrix. This is the current stage label. The embedding vector (64-dimensional) for each stage (e.g., approach stage, insertion stage). This term injects semantic information about "which assembly stage is currently being performed" into the unified state representation.
[0094] Contact tag embedded items. It is a contact embedding transformation matrix. It is a contact label .
[0095] For the relative pose coding of the two arms, Let be the relative pose transformation matrix of the two arms. At the current time k, the relative spatial relationship between the left and right arms is a 4*4 homogeneous transformation matrix (including relative translation and relative rotation). Pose encoding function. Through multilayer perceptron or linear transformation, the 4*4 matrix is flattened and mapped to a low-dimensional feature vector (e.g., 128-dimensional). Let be the pose feature transformation matrix. The output 128-dimensional pose features are further mapped to the dimensions of the unified state representation (such as 512-dimensional or 1024-dimensional).
[0096] In Example (1), after mapping by each transformation matrix, it is assumed that each encoded feature is... The component contribution is: The mean value is approximately 0.35 (the visual features are active, and the environmental camera can clearly observe the global pose of the workpiece).
[0097] The average value is approximately 0.05 (weak force characteristics, the end has not yet contacted the workpiece).
[0098] The mean value is approximately 0.12 (the state characteristics reflect the initial configuration of the robotic arm).
[0099] Additional semantic items: For the "approach phase", inject the semantic bias of the "distant guidance phase".
[0100] Inject "contactless" semantic bias.
[0101] When the two arms are far apart (e.g., 200mm) and their posture needs to be aligned, this item is injected with the spatial semantic offset of "the two arms are in a dispersed state and need to move closer to each other and align".
[0102] In Example (1), the unified state representation The visual spatial information dimension dominates, and the sub-task sequence decision model outputs the sub-objective "continue to approach the workpiece" based on this, while the micro-action strategy module remains silent during the contact phase.
[0103] Optionally, the current operation stage can be used as the insertion stage. =Insert), the contact determination result is: the contact label has been contacted ( =1). The corresponding weight generation calculation example (II) is as follows: Currently in the insertion phase, dominated by force sensory feature vectors. Assume... =0.25, =0.65, =0.10. Furthermore, contact label =1, the weight corresponding to the force sensory feature vector Further increase to ≥0.6.
[0104] Correspondingly, each encoded feature in step 103-2 is... The component contribution is: The mean value is approximately 0.10 (weak visual characteristics, the end has penetrated deep into the hole, and the wrist camera is obscured).
[0105] The mean value is approximately 0.42 (strong force characteristics, with significant changes in axial and lateral forces).
[0106] The mean value is approximately 0.08 (strong force characteristics, with significant changes in axial and lateral forces).
[0107] Additional semantic items: For the "insertion phase", inject the semantic bias of the "deep contact phase".
[0108] Inject semantic bias of "already touched".
[0109] The two arms have entered a state of coordinated cooperation and their relative poses are stable (e.g., the pin is deeply inserted into the hole and the hole seat is fixedly supported). This item injects the spatial semantic bias of "the two arms are in a close cooperative state and their relative poses have converged".
[0110] The comparison between Example (I) and Example (II) above is shown in Table 1 below.
[0111] Table 1
[0112] Referring to Table 1, it can be seen that from the approach stage to the insertion stage, It decreased from 0.70 to 0.25. It increases from 0.15 to 0.65. This change is determined by step 103-1 based on the stage label. and contact labels Dynamically generated in each control cycle, rather than through hard switching, to ensure a consistent state representation. The internal feature distribution transitions smoothly with the assembly process, ensuring that the downstream model always obtains the most reliable perception information at the moment.
[0113] Optionally, based on the example above in step 103, step 104 is illustrated below: The subtask sequence decision module uses the latest... A sequence of unified state representations at each time step As a sub-task sequence decision model The input, where See the example above for the information included.
[0114] The output of the subtask sequence decision model is illustrated below using two scenarios.
[0115] Scene 1: Normal assembly and advancement of the two arms: When the stage embedding in the unified state representation is emb ("proximity stage"), visual feature components dominate, and the relative pose encoding of the two arms shows that the two arms are far apart, the output of this subtask sequence decision module is: Sub-target Move closer to the assembly target location.
[0116] Role allocation for both arms Left arm holding assembly (active arm), right arm holding assembly tooling (support arm) Next stage tags Alignment phase.
[0117] Relative pose constraints The spatial deviation between the end of the mounting component and the target assembly position is within the allowable range.
[0118] As the stage embedding switches to emb ("insertion stage"), force sensing features are significantly enhanced, and the relative pose encoding of the two arms shows that they have entered a close cooperation state, the sub-task sequence decision model maintains the sub-objective as "inserting the assembly part into the assembly target", the role assignment remains unchanged, the next stage label is updated to "pressing", and the relative pose constraint is tightened so that the angle deviation between the assembly part axis and the assembly target baseline does not exceed the preset threshold.
[0119] Scenario 2: Dynamic role switching when both arms encounter an anomaly: When an aberrant change occurs in the force perception feature of the unified state representation (significant increase in lateral force) and the aberration type label indicates a risk of stagnation, the subtask sequence decision model re-evaluates and outputs: Sub-target : Resume insertion after eliminating jamming.
[0120] Role allocation for both arms The right arm transforms into the active arm (for fine-tuning and avoiding obstacles while holding the assembly tool), and the left arm transforms into the support arm (for pausing the advancement of the assembly parts).
[0121] Next stage tags Alignment phase (revert to the alignment phase for readjustment).
[0122] Relative pose constraints (Revert to the alignment phase and readjust).
[0123] The aforementioned role allocation dynamically switches from "left main, right branch" to "right main, left branch," entirely determined autonomously by the model based on the current unified state representation, rather than by a preset fixed left / right arm rule.
[0124] Optionally, to address the problem of a surge in contact force caused by minute pose deviations during the insertion and pressing stages, this application provides a two-stage contact strategy training mechanism: first, imitation learning initializes the behavior pattern from a demonstration trajectory; then, reinforcement learning fine-tunes and optimizes it in a simulation environment. This enables the strategy to master the fine-tuning ability of the end-effector pose and the fine-tuning of compliance control parameters, suitable for high-precision compliant assembly scenarios in the pre-insertion, insertion, and pressing stages. Specifically, in Figure 3 On this basis, Figure 6 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 6 Step 105 includes: Step 105-1: The contact phase micro-motion strategy module is based on the imitation learning method and uses trajectory segments of the contact phase in the demonstration trajectory data to initially train the contact phase micro-motion strategy.
[0125] Optionally, the contact phase may include the pre-insertion phase, insertion phase, and pressing phase as shown in the example above. Then, in the demonstration trajectory data, all segments labeled as pre-insertion, insertion, and pressing phases are extracted to form a contact phase training sample set. This contact phase training sample set is used to train the contact phase micro-motion strategy. This contact phase micro-motion strategy is used to output end-effector pose fine-tuning, attitude fine-tuning, and compliance control parameters.
[0126] Among them, the end-effector pose fine-tuning amount represents the small translational correction amount for the dual-arm end effector in the XYZ directions. The attitude fine-tuning amount represents the small rotational correction amount (represented by quaternions or rotation matrices) for the end-effector attitude. The compliance control parameter represents the adjustment amount of the stiffness coefficient and damping coefficient.
[0127] Furthermore, the goal of the initial training mentioned above is to ensure that the output of the micro-motion strategy during the contact phase can reproduce the standard compliant operation mode of humans during the contact phase: correcting lateral pose deviation when slightly obstructed, reducing stiffness to pass compliantly during continuous advancement, and adjusting damping to avoid overshoot during the pressing phase.
[0128] Step 105-2: The contact phase micro-action strategy module is based on reinforcement learning and fine-tunes and optimizes the contact phase micro-action strategy in a simulation environment.
[0129] Optionally, the contact phase micro-action strategy (such as a policy network) trained in step 105-1 is used as the initial model and placed in the simulation environment for fine-tuning and optimization. During each round of initialization, random initial pose deviations are injected into the assembly and the contact environment parameters are changed, allowing the strategy to learn through trial and error under various perturbation conditions.
[0130] Therefore, reinforcement learning methods can be used to achieve the above fine-tuning optimization. Below is one possible implementation method, within the framework of reinforcement learning, where the state space... Action space and reward function These constitute the core elements of a Markov decision-making process. The three elements form an optimized closed loop through the "micro-action strategy of the contact phase."
[0131] The state space defines the environmental information that the reinforcement learning agent can observe at each time step. In the refined contact phase of the dual-arm force-coordinated assembly, the state mainly includes the relative pose, contact force, and torque between the robot's end effector and the target assembly part, as well as the current impedance control parameters. This information collectively forms the basis for the "contact phase micro-motion strategy" decision-making. This state space... The expression is as follows:
[0132] in, : The relative position vector of the end effector of the dual-arm robot with respect to the target assembly.
[0133] in, This refers to the relative positional deviation of the end effector relative to the target assembly in the X-axis direction. The relative positional deviation in the Y-axis direction. This represents the relative positional deviation in the Z-axis direction.
[0134] The relative attitude of the end effector of a dual-arm robot with respect to the target assembly is typically represented by a quaternion. This is the real part (scalar part) of the quaternion, which is related to the rotation angle. The imaginary part is the X-axis component. The Y-axis component is the imaginary part. The Z-axis component represents the imaginary part. A quaternion uses four numbers to describe rotation in 3D space. Furthermore, the norm of this quaternion is 1; optionally, q_w ≥ 0 is taken to eliminate the ambiguity that q and -q represent the same pose. In a dual-arm scenario, the relative position and relative pose are preferably normalized according to the current active arm / support arm role, or jointly given by the relative pose encoding of both arms.
[0135] : The contact force vector measured at the contact point. Wherein, This represents the contact force component (lateral force) in the X-axis direction. This represents the contact force component (lateral force) in the Y-axis direction. This represents the contact force component (axial force / propulsion direction) along the Z-axis. (The above...) This refers to the force component measured by a six-dimensional force / torque sensor. Furthermore, and It can reflect the lateral friction or deflection resistance during insertion. It reflects the resistance to axial propulsion.
[0136] : The contact torque vector measured at the contact point. Where, The contact torque component about the X-axis, The contact torque component about the Y-axis, This represents the contact torque component about the Z-axis. (The above...) This refers to the torque component measured by a six-dimensional force / torque sensor. , This reflects the deflection torque generated on the hole wall when the assembly is inserted at an angle. It reflects the torsional moment about the insertion axis.
[0137] : The set of current impedance control parameters, such as stiffness and damping coefficients in each degree of freedom. The number of parameters.
[0138] The motion space defines the operations that the "contact phase micro-motion strategy" can perform at each time step. In the refined contact phase, the actions of the "contact phase micro-motion strategy" mainly include adjusting the speed commands of the end effectors of the drive arm and support arm, as well as adjusting the impedance control parameters to achieve compliant contact and smooth insertion. This motion space... The expression is as follows:
[0139] in, : The velocity command for the i-th robot end effector, including linear velocity and angular velocity, is used to integrate within the control period Δt to obtain the end effector pose fine-tuning compensation. Wherein, Let be the linear velocity of the i-th end effector along the X-axis. Let be the linear velocity of the i-th end effector along the Y-axis. Let be the linear velocity of the i-th end effector along the Z-axis. Let be the angular velocity of the i-th end effector about the X-axis. Let be the angular velocity of the i-th end effector about the Y-axis. Let be the angular velocity of the i-th end effector about the Z-axis. The first three components represent the translation of the i-th end effector, and the last three components represent the rotation of the i-th end effector.
[0140] The adjustment amount of impedance control parameters: the "contact phase micro-motion strategy" dynamically changes the robot's compliance by adjusting these parameters to adapt to different contact situations.
[0141] reward function This is the core of reinforcement learning, used to guide the "micro-motion strategy during the contact phase" to learn how to maximize cumulative rewards, thereby achieving the preset assembly goal. In the refined contact phase, the reward function aims to encourage successful insertion, minimize contact impact, avoid jamming, and penalize unnecessary movements and excessive forces. This reward function... The expression is as follows:
[0142] in, Successful insertion reward. A large positive reward is given when the assembly is successfully inserted into place. This is typically determined by detecting the assembly depth or the relative position between the end effector and the target. The expression is as follows:
[0143] in It is a relatively large positive number.
[0144] Lateral contact force deviation penalty. This penalizes the lateral component of the deviation from the desired contact force, encouraging compliant contact and avoiding the use of all necessary axial thrust as a penalty. The expression is as follows:
[0145] Among them, the force penalty weight is positive, the lateral projection matrix is used to extract the force component orthogonal to the insertion direction, and the desired contact force is preferably given along the insertion direction, with its lateral component being zero or close to zero.
[0146] Contact torque penalty. Penalizes excessive contact torque to prevent assembly tilting or jamming. The expression is as follows:
[0147] in, It is a positive weighting coefficient.
[0148] : Motion efficiency reward / penalty. Encourage the agent to complete tasks with minimal movement, or penalize unnecessary strenuous activity. The expression is as follows:
[0149] in and It is a positive weighting coefficient used to penalize excessively large end-effector speed commands and impedance parameter adjustments.
[0150] Other penalties may include time penalties (to encourage quick task completion), collision penalties (if an unexpected collision occurs), or stall penalties (if an assembly process is detected to be stalled). The expression is as follows:
[0151] in It is a positive weighting coefficient. and It is an indicator function that takes the value 1 when a collision or jam occurs, and 0 otherwise.
[0152] Through the comprehensive design of the above reward functions, the "contact phase micro-motion strategy" will be guided to learn a strategy that minimizes contact impact and force while ensuring assembly success rate, thereby achieving compliant and efficient dual-arm collaborative assembly.
[0153] Optionally, for steps 104 and 105 in the above example, which respectively complete the training of the subtask sequence decision model and the micro-action strategy for the contact phase, to ensure that the offline learning results can be stably reused and not accidentally tampered with in the subsequent online control phase, this application provides a model freezing and versioning management mechanism to solidify the parameters of each trained model into the corresponding model library. This design ensures that only model inference and log recording are performed in the online phase, without directly modifying the model parameters, thereby guaranteeing the real-time performance and stability of the control cycle and providing a traceable baseline version for periodic iterative retraining.
[0154] Specifically, in Figure 6 On this basis, Figure 7 A flowchart illustrating another dual-arm collaborative training method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 7 It also includes: Step 107: The subtask sequence decision module freezes the trained subtask sequence decision model into the decision model library.
[0155] Optionally, the decision model library can store decision models for subtask sequences. All encoded parameters after training and freezing include: the weight matrices and bias vectors of each network layer of the model; the normalized parameters of the unified state representation sequence input; and the time window length. The network structure configuration information, such as the number of network layers and the dimension of hidden layers, as well as the version identifier and offline verification performance indicators at the time of freezing, are included. The above encoding parameters are generated by training in step 104 and frozen in step 107 for use in the online control phase.
[0156] Step 108: The contact phase micro-action strategy module freezes the trained contact phase micro-action strategy into the contact strategy library.
[0157] Optionally, the contact strategy library stores all policy network parameters after the micro-action policies for the contact phase have been trained and frozen, including: the policy backbone network weights initialized during the imitation learning phase; the policy network weights and value network weights fine-tuned during reinforcement learning; the normalized parameters of each dimension of the state space; the weight coefficients of each component of the reward function and the success determination threshold; the action space boundary constraints of the end-effector velocity command and impedance parameter adjustment; and the version identifier and offline verification performance indicators at the time of freezing. The above policy network parameters are generated by training in steps 105-1 and 105-2, and frozen in step 108, for use in the online control phase when activated in the contact-related phase.
[0158] Optionally, for online control, a possible implementation method is provided below, specifically... Figure 8 This is a flowchart illustrating a dual-arm cooperative control method based on multimodal sensor fusion provided in an embodiment of the present invention. (See attached diagram.) Figure 8 The method includes: Step 204: The subtask sequence decision module calls the subtask sequence decision model in the decision model library, inputs the main task and the current unified state representation into the subtask sequence decision model, and obtains the subtask decision result.
[0159] The sub-task decision results may include, but are not limited to: sub-objectives, dual-arm role assignments, next-stage labels, and relative pose constraints. The main task is used to instruct the assembly task of the dual-arm robot; the current unified state represents the current multimodal fusion characteristics of the dual-arm robot.
[0160] Step 205: When the stage label indicates that the contact stage is in the contact stage, the micro-action strategy module obtains the end-point control correction amount based on the current unified state representation.
[0161] The contact phase includes a pre-insertion phase, an insertion phase, and a pressing phase.
[0162] Step 206: The dual-arm collaborative control and solution module calculates the dual-arm control error based on the sub-task decision results and real-time multimodal data.
[0163] Step 207: The dual-arm collaborative control and solution module generates the end control quantity based on the current unified state representation, sub-task decision results, end control correction quantity, and dual-arm control error.
[0164] Step 208: The dual-arm collaborative control and calculation module applies amplitude limits to the end-effector control quantity and converts the limited end-effector control quantity into joint control commands for both arms.
[0165] Step 209: The dual-arm collaborative control and calculation module sends joint control commands to the dual-arm robot, driving the dual arms to perform assembly operations.
[0166] Step 210: The online evaluation and rollback module monitors abnormal indicators in real time and performs abnormal handling operations when abnormal indicators are triggered.
[0167] The dual-arm cooperative control method based on multimodal sensor fusion provided in this invention outputs sub-task decision results containing relative pose constraints through a sub-task sequence decision model. The dual-arm cooperative control and solution module calculates the dual-arm control error and generates end-effector control quantities based on these results, realizing explicit expression and dynamic coordination of the dual-arm cooperative relationship. By replacing fixed threshold switching rules with data-driven next-stage label output, abrupt changes in control parameters at stage boundaries are avoided, improving the continuity and adaptability of stage switching. The contact stage micro-motion strategy module outputs end-effector control correction quantities based on a unified state representation during the contact stage. The control and solution module incorporates these correction quantities into the generation of end-effector control quantities, realizing fine-tuning and compliant control of the end-effector pose. The online evaluation and rollback module monitors abnormal indicators in real time and performs abnormal handling operations to ensure the safety and stability of the assembly process.
[0168] Optionally, this application provides a dual-arm coordination error calculation mechanism: first, the relative pose error (including position and attitude deviation) between the two arms is calculated based on role assignment and relative pose constraints; then, the lateral internal force error generated by the interaction of the two arms is calculated based on real-time force sensing data. Both mechanisms together provide closed-loop feedback input for the kinematic calculation and force / position hybrid control in step 207, ensuring that the two arms maintain spatial coordination accuracy while avoiding excessive internal forces caused by mutual pulling during collaborative assembly tasks.
[0169] Specifically, in Figure 8 On this basis, Figure 9 A flowchart illustrating another dual-arm cooperative control method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 9 Step 206 includes: Step 206-1: The dual-arm collaborative control and calculation module calculates the relative pose error of the two arms based on the sub-task decision results.
[0170] Optionally, for example, the subtask decision model output: dual-arm role assignment. The left arm is determined to be the active arm (holding component), and the right arm to be the support arm (holding assembly tooling); relative pose constraints. The relative spatial relationship that must be maintained between the active arm and the support arm is specified.
[0171] Real-time sensor feedback: End-effector pose ( , ), support arm end position ( , ).
[0172] Sub-target reference value: by and Relative position reference obtained from joint analysis (Desired positional offset of the active arm relative to the support arm) and relative attitude reference quantity (Desired pose alignment relationship).
[0173] The calculation process for step 206-1 is as follows: The formula for calculating relative position error is as follows:
[0174] The reference offset is given by the sub-target, and the measured relative position is obtained by subtracting the positions of the two arm ends. The position error used for closed-loop correction is obtained by subtracting the measured relative position from the reference offset. If the position of the active arm relative to the support arm deviates from the preset cooperative working distance, a non-zero component is generated in the direction to be corrected.
[0175] The formula for calculating relative attitude error is as follows: .
[0176] The process involves first calculating the measured relative attitude matrix at the ends of both arms, then multiplying the transpose of the measured matrix with the reference attitude matrix to obtain the deviation rotation from the measured attitude correction to the desired attitude; Log represents the matrix logarithmic mapping, and the superscript ∨ indicates vectorization. If the assembly axis is angularly skewed from the assembly baseline, a non-zero component is generated in the direction of that skew axis. This relative attitude error is represented in the tangent space of the unified assembly coordinate system and uses the same coordinate convention as the subsequent attitude increments of the active arm and support arm.
[0177] Step 206-2: The dual-arm collaborative control and calculation module calculates the lateral internal force error of the dual arms based on real-time multimodal data.
[0178] Optionally, the input source for step 206-2 can be: real-time multimodal data: force measured by the force / torque sensor at the end of the active arm. Force measured by force / torque sensor at the end of support arm The decision result of the above subtask: Insertion direction unit vector. (by sub-targets) and relative pose constraints (Analysis obtained). Coordinate transformation relationship: from the end-effector coordinate system to the unified assembly coordinate system. rotation matrix From the coordinate system at the end of the support arm to the unified assembly coordinate system rotation matrix .
[0179] Calculation process of lateral internal force error: 1) Insertion direction and lateral projection: in the unified assembly coordinate system In this context, the unit vector is defined by the insertion direction specified by the current sub-target. Construct the formula for calculating the transverse projection matrix:
[0180] This matrix projects any three-dimensional vector onto a transverse plane orthogonal to the insertion direction, filters out the axial component, and retains only the transverse component for internal force evaluation.
[0181] 2) Force vector coordinate unification: The force vectors measured by the force / torque sensors at the ends of both arms are transformed from their respective end coordinate systems to a unified assembly coordinate system. Active arm: Support arm: .
[0182] 3) Lateral internal force error: After unifying the direction signs of the force vectors of the two arms in the assembly coordinate system, superimpose them, project them onto the lateral plane, and then subtract the lateral component of the target contact force reference. The expression for the lateral internal force error of the two arms is as follows:
[0183] in, The target contact force reference value is preferably a reference force along the insertion direction, and its lateral component is preferably zero or close to zero. Here, σA and σB are force direction sign coefficients, determined by the force sensor installation direction and force definition conventions.
[0184] It should be noted that, It reflects the deviation between the lateral internal force of the two arms and the expected value. If there is mutual pulling or pushing at the ends of the two arms in the lateral plane (for example, the assembly is stuck by the side wall of the assembly target due to asynchronous posture), the lateral components of the force sensors of the two arms will be superimposed to produce a non-zero value, which will be extracted by the projection matrix to form the lateral internal force error.
[0185] Optionally, this application provides an end-effector control quantity generation mechanism: by combining kinematic calculation with force / position hybrid control, and integrating spatial coordination error, internal force deviation and contact phase micro-motion correction, the end-effector position increment and attitude increment of the active arm and the support arm are calculated respectively, thereby transforming the macro-coordination strategy into the precise end-effector control quantity required for the underlying joint drive.
[0186] Specifically, in Figure 9 On this basis, Figure 10 A flowchart illustrating another dual-arm cooperative control method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 10 Step 207 includes: Step 207-1: The dual-arm collaborative control and calculation module determines the end position increment and attitude increment of the active arm and support arm of the dual arms based on the unified state representation, sub-target, relative pose constraint, end control correction, relative pose error and lateral internal force error, and generates end control quantity.
[0187] The active arm and the supporting arm of the two arms are determined by the role allocation of the two arms.
[0188] Optionally, an example of control law calculation is provided below.
[0189] 1) Expression for the position increment of the active arm (holding accessory):
[0190] in, Its function is to increase the proportional gain. Amplify the relative position deviation to drive the active arm to approach the desired relative pose. Its function is to increase the proportional gain. Amplify the attitude deviation and correct the attitude skew of the assembly parts relative to the assembly tooling. Its function is to act along the insertion direction The feedforward propulsion term enables the active arm to continuously descend at the desired pace of the sub-target. The negative sign indicates that when the lateral internal force is too large, the active arm will reverse to correct and weaken the pulling force. Its function is to make fine corrections during the contact phase (such as slight retreat or lateral correction when obstructed). Its function is to perform closed-loop correction of minor pose deviations in the non-contact stage based on visual feedback.
[0191] 2) Expression for the position increment of the support arm (holding assembly tool):
[0192] in, The negative sign causes the support arm and the active arm to move in opposite directions, working together to reduce the relative positional deviation between the two arms. The negative sign indicates that the support arm is adjusted to maintain the stability of the tooling reference. The negative sign causes the support arm to respond in the opposite direction to the lateral internal force, which together with the active arm weakens the mutual pulling. Its function is to make fine adjustments to the support arm (such as slight tooling avoidance or damping adjustment).
[0193] Where sat(·) is the amplitude limiting function. , This represents the fine-tuning amount of the policy output. This represents the visual servo correction amount. The attitude increment is calculated using the same method as the position increment.
[0194] Optionally, since the robotic arm joint actuators directly respond to joint angle or joint speed commands, rather than end-effector spatial quantities, it is necessary to map the end-effector control quantities to the joint space through inverse kinematics. One possible implementation is provided below: The end-effector control variable is converted into joint control commands using inverse kinematics or damped least squares Jacobian inverse. For the i-th robotic arm, let [denote...]. Its joint increment can be denoted as... And set the joint commands as follows: If the controller supports speed mode or impedance mode, the end-effector control quantity is further converted into joint speed command or impedance parameter. The least-squares damping factor in the above equation is distinct from the impedance control parameter.
[0195] Optionally, to provide a complete overview of the online control process, the following section offers a possible implementation method for the perception and coding mechanisms. Specifically, in Figure 8 On this basis, Figure 11 A flowchart illustrating another dual-arm cooperative control method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 11 It also includes: Step 200: The acquisition and calibration module synchronously acquires real-time multimodal sensing data in each control cycle and calibrates the real-time multimodal sensing data to obtain real-time multimodal data.
[0196] The selected location acquires real-time multimodal data within the control period Δt, including: environmental visual data, wrist visual data, six-dimensional torque data of the two arm ends, and joint state data of the two arms. The six-dimensional forces / torques of the six-dimensional torque data of the two arm ends are denoted as follows: The joint status data of both arms are .
[0197] Step 201: The data synchronization and preprocessing module performs contact determination on the real-time multimodal data and generates stage labels and contact labels.
[0198] Step 202: The multimodal coding module calls various modal data encoders in the modal data encoder library, inputs real-time multimodal data into various modal data encoders, and obtains the corresponding encoded feature vectors.
[0199] Step 203: The multimodal coding module generates weights corresponding to each coding feature vector based on stage labels, contact labels, and relative pose features. It then performs weighted fusion of all coding feature vectors to obtain the current unified state representation.
[0200] Optionally, referring to the example above, real-time multimodal data can be input into various modal data encoders to obtain corresponding encoded feature vectors, such as visual feature vectors. Force feature vector State feature vector .
[0201] Furthermore, based on stage labels and Generate normalized weights , , And satisfy Therefore, the current unified state is represented as follows (see above):
[0202] Optionally, the following provides a possible implementation method for calling various encoders during the online control phase. Specifically, in... Figure 11 On this basis, Figure 12 A flowchart illustrating another dual-arm cooperative control method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 12 Step 202 includes: Step 202-1: The multimodal coding module calls the visual encoder in the modal data encoder library, inputs the visual data into the visual encoder, and obtains the visual feature vector.
[0203] Step 202-2: The multimodal coding module calls the force encoder in the modal data encoder library, inputs the six-dimensional torque data of the ends of the two arms into the force encoder, and obtains the force feature vector.
[0204] Step 202-3: The multimodal coding module calls the state encoder in the modal data encoder library, inputs the state data of the two arm joints into the state encoder, and obtains the state feature vector.
[0205] Optionally, to enable fine-grained compliant control during the contact-related phase and avoid excessive policy intervention during the non-contact phase, this application provides a conditional contact policy invocation mechanism. Specifically, in Figure 8 On this basis, Figure 13A flowchart illustrating another dual-arm cooperative control method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 13 Step 205 includes: Step 205-1: The micro-action strategy module in the contact phase determines whether the phase label indicates that it is in the contact phase.
[0206] If yes, then proceed to step 205-2; otherwise, proceed to step 205-3.
[0207] Step 205-2: During the contact phase, the micro-action strategy module obtains the end-point control correction amount based on the current unified state representation.
[0208] Step 205-3: During the contact phase, the micro-motion strategy module does not output end-point control correction values.
[0209] The specific calculation method in this example is described in the example above and will not be repeated here. Optionally, to avoid the continuous accumulation of abnormal contact force, insertion jamming, or joint overload caused by sensor noise, assembly pose deviation, or environmental disturbances, this application provides an online evaluation and rollback mechanism. Specifically, in Figure 8 On this basis, Figure 14 A flowchart illustrating another dual-arm cooperative control method based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 14 Step 210 includes: Step 210-1: The online evaluation and rollback module monitors in real time at least one of the following abnormal indicators: lateral internal force exceeding the threshold for multiple control cycles, insertion depth not effectively increasing within multiple control cycles, visual confidence level below the threshold for multiple control cycles, and joint current or drive load exceeding the upper limit.
[0210] Step 210-2: When an abnormal indicator is triggered, the online evaluation and rollback module shall perform at least one of the following abnormal handling operations: rollback a preset rollback distance in the opposite direction of insertion, return to alignment or realign during the pre-insertion stage, reduce the advance speed, change the search direction, or request manual intervention.
[0211] Examples of how to determine the above abnormal indicators are as follows: When the anomaly type is excessive lateral internal force, the judgment conditions are as follows: like and continue If a certain number of cycles are reached, an abnormal condition will be triggered.
[0212] When the exception type is insertion depth stagnation, the judgment conditions are as follows: like ,recent If the depth growth is less than the threshold within a certain period, an abnormal condition is triggered.
[0213] When the anomaly type is decreased visual confidence, the judgment criteria are as follows: If the visual confidence level is below the threshold and is continuous If all cycles are below the threshold, an abnormal condition is triggered.
[0214] When the exception type is "drive load exceeds limit", the judgment conditions are as follows: If the current of any joint exceeds the upper limit or the total drive load exceeds the limit, an abnormal condition will be triggered.
[0215] The above four indicators are monitored independently. Meeting any one of these conditions determines an anomaly and proceeds to step 210-2. After an anomaly is triggered, the module selects and executes at least one anomaly handling operation based on the anomaly type: When the anomaly type is excessive lateral internal force, first proceed along the opposite direction of insertion. Back distance If the lateral internal force is still higher than the threshold after retraction, return to the "align" or "pre-insertion" stage to realign; reduce the advance speed if necessary.
[0216] When the exception type is "insertion depth stalled", first backtrack. Then, try to fine-tune the search direction by changing it in the horizontal plane; if two consecutive searches are ineffective, request manual intervention.
[0217] When the anomaly type is decreased visual confidence, reduce the propulsion speed and trigger the visual servoing enhancement mode; if it continues... If the confidence level has not recovered after one cycle, manual intervention is requested.
[0218] When the abnormality type is "drive load over limit", immediately stop propulsion (speed set to zero) and reverse. They then directly requested manual intervention.
[0219] Log recording: After each exception handling is completed, a structured execution log expression is recorded as follows:
[0220] Where s_k represents the current unified state representation at the time the exception is triggered; Indicates the control action being performed (reverse / realign / decelerate / change direction / manual intervention); Indicates the current stage label; Indicate the result label (Recovery successful / Intervention still required / Taken over); Labels indicating abnormal types (internal force exceeding limits / depth stagnation / visual impairment / drive overload); This indicates the cumulative number of rollbacks in this assembly task; This indicates a manual correction flag (Boolean value).
[0221] Furthermore, the aforementioned logs are periodically added to the sample pool for the next round of offline training, used for iterative retraining of the contact strategy. Optionally, to fully illustrate how the offline training results are continuously optimized through online execution feedback, achieving a closed-loop evolution of the dual-arm collaborative assembly system from initial deployment to iterative enhancement, the offline training link (steps 100-106) and the online control link (steps 204-211) are integrated into a unified flowchart for explanation. Through the execution log feedback and sample pool update in step 212, the system can transform abnormal working conditions encountered in actual assembly into retraining samples, periodically iteratively optimizing the modal encoder, decision model, and contact strategy, thereby continuously improving the assembly success rate and robustness under complex working conditions.
[0222] Specifically, Figure 15 A dual-arm collaborative iterative training method based on multimodal sensor fusion is provided in this embodiment of the invention. See [link to relevant documentation]. Figure 15 The method includes: Step 100: The acquisition and calibration module acquires the multimodal sensing data corresponding to the demonstration operation and calibrates the multimodal sensing data to obtain multimodal demonstration trajectory data.
[0223] Step 101: The data synchronization and preprocessing module performs stage segmentation on the multimodal demonstration trajectory data to obtain the corresponding trajectory segments.
[0224] Step 102: The multimodal coding module trains the encoder for each type of modal data to obtain the coding feature vector of the corresponding type of modal data encoder.
[0225] Optionally, as mentioned above, the multimodal coding module freezes the coding parameters corresponding to all trained modal data encoders into the modal data encoder library.
[0226] Step 103: The multimodal coding module performs weighted fusion of all coding feature vectors based on the weights corresponding to each coding feature vector to obtain a unified state representation.
[0227] Step 104: The subtask sequence decision module trains the subtask sequence decision model based on the unified state representation.
[0228] Optionally, as mentioned above, the subtask sequence decision module freezes the trained subtask sequence decision model into the decision model library.
[0229] Step 105: The contact phase micro-action strategy module trains the contact phase micro-action strategy based on the unified state representation, so that the contact phase micro-action strategy outputs end control commands when in the contact-related phase.
[0230] Optionally, as mentioned above, the contact phase micro-action strategy module freezes the trained contact phase micro-action strategy into the contact strategy library.
[0231] Step 200: The acquisition and calibration module synchronously acquires real-time multimodal sensing data in each control cycle and calibrates the real-time multimodal sensing data to obtain real-time multimodal data.
[0232] Step 201: The data synchronization and preprocessing module performs contact determination on the real-time multimodal data and generates stage labels and contact labels.
[0233] Step 202: The multimodal coding module calls various modal data encoders in the modal data encoder library, inputs real-time multimodal data into various modal data encoders, and obtains the corresponding encoded feature vectors.
[0234] Step 203: The multimodal coding module generates weights corresponding to each coding feature vector based on stage labels, contact labels, and relative pose features. It then performs weighted fusion of all coding feature vectors to obtain the current unified state representation.
[0235] Step 204: The subtask sequence decision module calls the subtask sequence decision model in the decision model library, inputs the main task and the current unified state representation into the subtask sequence decision model, and obtains the subtask decision result.
[0236] Step 205: When the stage label indicates that the contact stage is in the contact stage, the micro-action strategy module obtains the end-point control correction amount based on the current unified state representation.
[0237] Step 206: The dual-arm collaborative control and solution module calculates the dual-arm control error based on the sub-task decision results and real-time multimodal data.
[0238] Step 207: The dual-arm collaborative control and solution module generates the end control quantity based on the current unified state representation, sub-task decision results, end control correction quantity, and dual-arm control error.
[0239] Step 208: The dual-arm collaborative control and calculation module applies amplitude limits to the end-effector control quantity and converts the limited end-effector control quantity into joint control commands for both arms.
[0240] Step 209: The dual-arm collaborative control and calculation module sends joint control commands to the dual-arm robot, driving the dual arms to perform assembly operations.
[0241] Step 211: The online evaluation and rollback module records the execution log, updates the training sample pool based on the execution log, and iteratively retrains the modal data encoder, subtask sequence decision model, and / or contact phase micro-action strategy based on the updated training sample pool.
[0242] To better illustrate the superiority of the present application's example compared to existing technical solutions, a comparative experimental data example is provided below. The experimental object (main task) is a dual-arm connector assembly task. Experimental groups are as follows: Group A is fixed rules and visual servoing; Group B is behavior cloning; Group C is multimodal fusion (no sub-tasks); and Group D is the above-mentioned example solution of this invention. Under conditions of initial pose deviation, workpiece tolerance variation, and local visual occlusion, each group completed 60 dual-arm assembly tests. The complete task time from the approach phase to assembly completion or anomaly handling was recorded, along with the success rate, average cycle time, peak lateral force, average number of retries, and number of manual interventions. The results are shown in Table 2.
[0243] Table 2
[0244] As can be seen, the solution provided in this application achieves the following technical effects through a hierarchical control strategy and a multimodal fusion mechanism: 1) Improved assembly robustness in complex environments: By dynamically adjusting the sensing weights, the limitations of a single sensing mode in a specific assembly stage are effectively overcome. 2) Reduced the risk of jamming in dual-arm coordination: The introduction of relative posture constraints and internal force compensation logic enables the dual arms to form a compliant fit during assembly, reducing mechanical overload. 3) Enhanced system adaptability and fault tolerance: The hierarchical inference architecture and anomaly rollback mechanism enable the system to autonomously cope with part tolerances and positional drift, significantly improving the assembly success rate.
[0245] Optionally, to achieve the various steps and corresponding technical effects of the above embodiments, a dual-arm collaborative device based on multimodal sensor fusion is provided below. This device can call upon the required internal modules during offline training and online control phases to achieve corresponding functions. Specifically, Figure 16 This is a schematic diagram of a dual-arm collaborative device based on multimodal sensor fusion, provided in an embodiment of the present invention. Figure 16 (a) shows the modules invoked by the device 20 during offline training; Figure 16 (b) shows the modules invoked when the device 20 is in online control.
[0246] Below, firstly... Figure 16 (a) As can be seen, during offline training, the modules called by the device 20 include: acquisition and calibration module 200, data synchronization and preprocessing module 201, multimodal coding module 202, subtask sequence decision module 203, and contact phase micro-action strategy module 204.
[0247] The acquisition and calibration module 200 is used to acquire multimodal sensing data corresponding to the demonstration operation, and to calibrate the multimodal sensing data to obtain multimodal demonstration trajectory data.
[0248] The data synchronization and preprocessing module 201 is used to perform stage segmentation on the multimodal demonstration trajectory data to obtain the corresponding trajectory segments.
[0249] The multimodal coding module 202 is used to train the encoder for each type of modal data separately to obtain the encoding feature vector of the corresponding type of modal data encoder.
[0250] The multimodal coding module 202 is used to weight and fuse all coding feature vectors based on the weights corresponding to each coding feature vector to obtain a unified state representation.
[0251] The subtask sequence decision module 203 is used to train the subtask sequence decision model based on the unified state representation.
[0252] The contact phase micro-action strategy module 204 is used to train the contact phase micro-action strategy based on a unified state representation, so that the contact phase micro-action strategy outputs end control commands when it is in the contact-related phase.
[0253] The multimodal coding module 202 is used to freeze the coding parameters corresponding to all trained modal data encoders into the modal data encoder library.
[0254] Optionally, the acquisition and calibration module 200 is specifically used to acquire multimodal sensing data corresponding to the dual-arm collaborative assembly demonstration operation; to perform coordinate system calibration on each sensor and robotic arm involved in the multimodal sensing data, and to obtain multimodal demonstration trajectory data.
[0255] Optionally, the data synchronization and preprocessing module 201 is specifically used to perform time alignment and noise suppression on the multimodal demonstration trajectory data; and to perform contact determination on the time-aligned and noise-suppressed multimodal demonstration trajectory data. Based on the contact determination results, the multimodal demonstration trajectory data is divided into trajectory segments corresponding to multiple operation stages.
[0256] Optionally, the multimodal coding module 202 is specifically used to train the visual encoder so that the visual encoder maps visual data into visual feature vectors; train the force encoder so that the force encoder maps the six-dimensional torque data of the two arm ends into force feature vectors; and train the state encoder so that the state encoder maps the joint state data of the two arms into state feature vectors.
[0257] Optionally, the multimodal coding module 202 is specifically used to generate weights corresponding to each coding feature vector based on the current operation stage and the contact determination result; and to fuse all coding feature vectors according to their corresponding weights to obtain a unified state representation.
[0258] Optionally, the contact phase micro-motion strategy module 204 is specifically used to: initially train the contact phase micro-motion strategy using trajectory segments of the contact phase in the demonstration trajectory data based on imitation learning; and fine-tune and optimize the contact phase micro-motion strategy in a simulation environment based on reinforcement learning.
[0259] Optionally, the subtask sequence decision module 203 is also used to freeze the trained subtask sequence decision model to the decision model library.
[0260] The contact phase micro-action strategy module 204 is also used to freeze the trained contact phase micro-action strategy to the contact strategy library.
[0261] Below, firstly... Figure 16 (b) As explained, during online control, the modules invoked by the device 20 include: acquisition and calibration module 200, data synchronization and preprocessing module 201, multimodal coding module 202, subtask sequence decision module 203, contact phase micro-motion strategy module 204, dual-arm collaborative control and calculation module 205, and online evaluation and rollback module 206.
[0262] The subtask sequence decision module 203 is used to call the subtask sequence decision model in the decision model library, input the main task and the current unified state representation into the subtask sequence decision model, and obtain the subtask decision result.
[0263] The contact phase micro-action strategy module 204 is used to obtain the end-point control correction amount based on the current unified state representation when the phase label indicates that it is in the contact phase.
[0264] The dual-arm collaborative control and calculation module 205 is used to calculate the dual-arm control error based on the sub-task decision results and real-time multimodal data. Based on the current unified state representation, sub-task decision results, end-effector control correction, and dual-arm control error, it generates end-effector control quantities; it applies amplitude limits to the end-effector control quantities and converts the limited end-effector control quantities into joint control commands for the dual arms; it sends the joint control commands to the dual-arm robot to drive the dual arms to perform assembly operations.
[0265] The online assessment and rollback module 206 is used to monitor abnormal indicators in real time and perform abnormal handling operations when abnormal indicators are triggered.
[0266] Optionally, the dual-arm collaborative control and calculation module 205 is specifically used to calculate the relative pose error of the two arms based on the sub-task decision results, and to calculate the lateral internal force error of the two arms based on real-time multimodal data.
[0267] Optionally, the sub-task decision results include: sub-objectives, dual-arm role allocation and relative pose constraints; the dual-arm collaborative control and solution module 205 is specifically used to determine the end position increments and attitude increments of the active arm and the support arm of the dual arms based on the current unified state representation, sub-objectives, relative pose constraints, end-effector control correction, relative pose error and lateral internal force error, and generate end-effector control quantities.
[0268] Optionally, the acquisition and calibration module 200 is also used to synchronously acquire real-time multimodal sensing data in each control cycle and calibrate the real-time multimodal sensing data to obtain real-time multimodal data. The data synchronization and preprocessing module 201 is also used to determine contact on real-time multimodal data and generate stage labels and contact labels.
[0269] The multimodal coding module 202 is also used to call various modal data encoders in the modal data encoder library, input real-time multimodal data into various modal data encoders, and obtain the corresponding encoded feature vectors; The multimodal coding module 202 is also used to generate weights corresponding to each coding feature vector based on stage labels, contact labels and relative pose features, and to perform weighted fusion of all coding feature vectors to obtain the current unified state representation.
[0270] Optionally, the multimodal coding module 202 is specifically used to call the visual encoder in the modal data encoder library, input visual data into the visual encoder to obtain visual feature vectors; call the force encoder in the modal data encoder library, input the six-dimensional torque data of the two arm ends into the force encoder to obtain force feature vectors; and call the state encoder in the modal data encoder library, input the joint state data of the two arms into the state encoder to obtain state feature vectors.
[0271] Optionally, the contact phase micro-action strategy module 204 is specifically used to determine whether the phase label indicates that it is in the contact phase; if so, it obtains the end-point control correction amount based on the current unified state representation.
[0272] Optionally, the online evaluation and rollback module 206 is specifically used to monitor in real time at least one of the following abnormal indicators: lateral internal force exceeding a threshold for multiple control cycles, insertion depth not effectively increasing within multiple control cycles, visual confidence level below a threshold for multiple control cycles, and joint current or drive load exceeding the upper limit; when an abnormal indicator is triggered, at least one of the following abnormal handling operations is performed: rollback a preset rollback distance in the opposite direction of insertion, return to alignment or realign during the pre-insertion stage, reduce the propulsion speed, change the search direction, or request manual intervention.
[0273] Furthermore, in the above example where offline training and online control undergo multiple iterations, an implementation mechanism for the aforementioned device is provided, specifically, Figure 17 A schematic diagram of another dual-arm collaborative device based on multimodal sensor fusion provided in this embodiment of the invention is shown below. Figure 17 .
[0274] The acquisition and calibration module 200 is used to acquire multimodal sensing data corresponding to the demonstration operation, and to calibrate the multimodal sensing data to obtain multimodal demonstration trajectory data.
[0275] The data synchronization and preprocessing module 201 is used to perform stage segmentation on the multimodal demonstration trajectory data to obtain the corresponding trajectory segments.
[0276] The multimodal coding module 202 is used to train the encoder for each type of modality data separately to obtain the encoded feature vector of the corresponding type of modality data encoder. Based on the weights corresponding to each encoded feature vector, all encoded feature vectors are weighted and fused to obtain a unified state representation. All trained modality data encoders are then frozen into the modality data encoder library.
[0277] The subtask sequence decision module 203 is used to train the subtask sequence decision model based on the unified state representation.
[0278] The contact phase micro-action strategy module 204 is used to train the contact phase micro-action strategy based on a unified state representation, so that the contact phase micro-action strategy outputs end control commands when it is in the contact-related phase.
[0279] The subtask sequence decision module 203 is used to call the subtask sequence decision model in the decision model library, input the main task and the current unified state representation into the subtask sequence decision model, and obtain the subtask decision result.
[0280] The contact phase micro-action strategy module 204 is used to obtain the end-point control correction amount based on the current unified state representation when the phase label indicates that it is in the contact phase.
[0281] The dual-arm collaborative control and calculation module 205 is used to calculate the dual-arm control error based on the sub-task decision results and real-time multimodal data.
[0282] The dual-arm collaborative control and calculation module 205 is used to generate end control quantities based on the current unified state representation, sub-task decision results, end control correction quantities, and dual-arm control errors.
[0283] The dual-arm collaborative control and calculation module 205 is used to apply amplitude limits to the end control quantity and convert the limited end control quantity into joint control commands for both arms.
[0284] The dual-arm collaborative control and calculation module 205 is used to send joint control commands to the dual-arm robot to drive the dual arms to perform assembly operations.
[0285] The online evaluation and rollback module 206 is used to record execution logs, update the training sample pool based on the execution logs, and iteratively retrain the modal data encoder, subtask sequence decision model, and / or contact phase micro-action strategy based on the updated training sample pool.
[0286] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0287] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0288] If a function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0289] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0290] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within this application. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A dual-arm collaborative training method based on multi-modal sensor fusion, characterized in that, include: The acquisition and calibration module acquires multimodal sensing data corresponding to the demonstration operation and calibrates the multimodal sensing data to obtain multimodal demonstration trajectory data. The data synchronization and preprocessing module performs stage segmentation on the multimodal demonstration trajectory data to obtain corresponding trajectory segments; each trajectory segment corresponds to a work stage in the multimodal demonstration trajectory data; The multimodal coding module trains the encoder for each type of modal data separately to obtain the encoded feature vector of the corresponding type of modal data encoder; The multimodal coding module performs weighted fusion of all the encoded feature vectors based on the weights corresponding to each of the encoded feature vectors to obtain a unified state representation; the unified state representation is used to characterize the multimodal fusion features of the dual-arm robot. The sub-task sequence decision module trains the sub-task sequence decision model based on the unified state representation; the sub-task sequence decision model is used to define the control strategy for both arms based on the unified state representation; The contact phase micro-action strategy module trains the contact phase micro-action strategy based on the unified state representation, so that the contact phase micro-action strategy outputs end control commands when in the contact-related phase. The multimodal coding module freezes the coding parameters corresponding to all the trained modal data encoders into the modal data encoder library.
2. The method of claim 1, wherein, The steps for the acquisition and calibration module to acquire multimodal sensing data corresponding to the demonstration operation include: The acquisition and calibration module acquires multimodal sensing data corresponding to the dual-arm collaborative assembly demonstration operation; the multimodal sensing data includes: environmental visual data, wrist visual data, six-dimensional torque data at the ends of both arms, and joint status data of both arms; The step of calibrating the multimodal sensing data to obtain multimodal demonstration trajectory data includes: The acquisition and calibration module performs coordinate system calibration on each sensor and robotic arm involved in the multimodal sensing data to obtain the multimodal demonstration trajectory data.
3. The method of claim 1, wherein, The data synchronization and preprocessing module performs stage segmentation on the multimodal demonstration trajectory data to obtain corresponding trajectory segments, including: The data synchronization and preprocessing module performs time alignment and noise suppression on the multimodal demonstration trajectory data; The data synchronization and preprocessing module performs contact determination on the time-aligned and noise-suppressed multimodal demonstration trajectory data; Based on the contact determination result, the data synchronization and preprocessing module divides the multimodal demonstration trajectory data into multiple trajectory segments corresponding to the operation stages. The operation stages include at least the approach stage, alignment stage, pre-insertion stage, insertion stage, and pressing stage.
4. The method of claim 1, wherein, The multimodal coding module trains the encoder for each type of modal data to obtain the encoded feature vector of the corresponding type of modal data encoder, including: The multimodal coding module trains the visual encoder so that the visual encoder maps visual data into visual feature vectors; The multimodal coding module trains the force encoder so that the force encoder maps the six-dimensional torque data at the ends of both arms into force feature vectors; The multimodal coding module trains the state encoder so that the state encoder maps the joint state data of both arms into state feature vectors.
5. The method according to claim 3, characterized in that, The step of the multimodal coding module weighted and fused all the encoded feature vectors based on the weights corresponding to each encoded feature vector to obtain a unified state representation includes: The multimodal coding module generates weights corresponding to each of the coding feature vectors based on the current operation stage and the contact determination result; The multimodal coding module weights and fuses all the encoded feature vectors according to their corresponding weights to obtain the unified state representation.
6. The method of claim 3, wherein, The steps of training the contact phase micro-action policy module based on the unified state representation include: The contact phase micro-motion strategy module is based on the imitation learning method and uses trajectory segments of the contact phase in the demonstration trajectory data to initially train the contact phase micro-motion strategy; the contact phase includes the pre-insertion phase, the insertion phase, and the pressing phase; The contact phase micro-action strategy module is based on reinforcement learning and fine-tunes and optimizes the contact phase micro-action strategy in a simulation environment.
7. The method according to claim 6, characterized in that, Also includes: The sub-task sequence decision module freezes the trained sub-task sequence decision model into the decision model library; The contact phase micro-action strategy module freezes the trained contact phase micro-action strategy into the contact strategy library.
8. A dual-arm cooperative control method based on multimodal sensor fusion, characterized in that, include: The subtask sequence decision module calls the subtask sequence decision model in the decision model library, inputs the main task and the current unified state representation into the subtask sequence decision model, and obtains the subtask decision result. The main task is used to indicate the assembly task of the dual-arm robot; the current unified state represents the current multimodal fusion characteristics of the dual-arm robot. When the stage label indicates that the contact stage is in the contact stage, the micro-action strategy module obtains the end-effector control correction amount based on the current unified state representation; the contact stage includes the pre-insertion stage, the insertion stage, and the pressing stage; The dual-arm collaborative control and calculation module calculates the dual-arm control error based on the sub-task decision results and real-time multimodal data; The dual-arm collaborative control and calculation module generates end-of-arm control quantities based on the current unified state representation, the sub-task decision results, the end-of-arm control correction quantity, and the dual-arm control error. The dual-arm collaborative control and calculation module applies amplitude limits to the end-effector control quantity and converts the limited end-effector control quantity into joint control commands for both arms. The dual-arm collaborative control and calculation module sends the joint control commands to the dual-arm robot, driving the dual arms to perform assembly operations. The online assessment and rollback module monitors abnormal indicators in real time and performs exception handling operations when abnormal indicators are triggered.
9. The method according to claim 8, characterized in that, The step of the dual-arm collaborative control and calculation module calculating the dual-arm control error based on the sub-task decision results and the real-time multimodal data includes: The dual-arm collaborative control and calculation module calculates the relative pose error of the two arms based on the sub-task decision results. The dual-arm collaborative control and calculation module calculates the lateral internal force error of the dual arms based on the real-time multimodal data.
10. The method according to claim 9, characterized in that, The sub-task decision results include: sub-objectives, dual-arm role assignments, and relative pose constraints; the step of the dual-arm collaborative control and solution module generating end-of-arm control quantities based on the current unified state representation, the sub-task decision results, the end-of-arm control correction, and the dual-arm control error includes: The dual-arm collaborative control and calculation module determines the end-position increments and attitude increments of the active arm and the support arm of the dual arms based on the current unified state representation, the sub-target, the relative pose constraint, the end-effector control correction, the relative pose error, and the lateral internal force error, and generates the end-effector control quantity; wherein, the active arm and the support arm of the dual arms are determined by the dual-arm role allocation.
11. The method according to claim 8, characterized in that, Before the step of the subtask sequence decision module calling the subtask sequence decision model in the decision model library, inputting the main task and the current unified state representation into the subtask sequence decision model, and obtaining the subtask decision result, the following is also included: The acquisition and calibration module synchronously acquires real-time multimodal sensing data in each control cycle and calibrates the real-time multimodal sensing data to obtain real-time multimodal data. The data synchronization and preprocessing module performs contact determination on the real-time multimodal data and generates stage labels and contact labels; The multimodal coding module calls various modal data encoders in the modal data encoder library, inputs the real-time multimodal data into the various modal data encoders, and obtains the corresponding encoded feature vectors; The multimodal coding module generates weights corresponding to each of the encoded feature vectors based on the stage label, the contact label, and the relative pose features, and then performs weighted fusion of all the encoded feature vectors to obtain the current unified state representation.
12. The method according to claim 11, characterized in that, The multimodal coding module calls various modal data encoders from the modal data encoder library, inputs the real-time multimodal data into each modal data encoder, and obtains the corresponding encoded feature vectors. The steps include: The multimodal coding module calls the visual encoder in the modal data encoder library, inputs visual data into the visual encoder, and obtains visual feature vectors; The multimodal coding module calls the force encoder in the modal data encoder library, inputs the six-dimensional torque data of the ends of the two arms into the force encoder, and obtains the force feature vector; The multimodal coding module calls the state encoder in the modal data encoder library, inputs the joint state data of both arms into the state encoder, and obtains the state feature vector.
13. The method according to claim 8, characterized in that, The step of the contact phase micro-action strategy module obtaining the end-point control correction amount based on the current unified state representation when the phase label indicates that it is in the contact phase includes: The contact phase micro-action strategy module determines whether the phase label indicates that it is in the contact phase; If so, the contact phase micro-action strategy module obtains the end-point control correction amount based on the current unified state representation.
14. The method according to claim 8, characterized in that, The online evaluation and rollback module monitors abnormal indicators in real time and performs anomaly handling operations when an abnormal indicator is triggered, including: The online evaluation and rollback module monitors in real time at least one of the following abnormal indicators: lateral internal force exceeding the threshold for multiple control cycles, insertion depth not effectively increasing within multiple control cycles, visual confidence level below the threshold for multiple control cycles, and joint current or drive load exceeding the upper limit. When the abnormal indicator is triggered, the online evaluation and rollback module performs at least one of the following abnormal handling operations: rollback a preset rollback distance in the opposite direction of insertion, return to alignment or realign during the pre-insertion stage, reduce the advance speed, change the search direction, or request manual intervention.
15. A dual-arm collaborative iterative training method based on multimodal sensor fusion, characterized in that, include: The acquisition and calibration module acquires multimodal sensing data corresponding to the demonstration operation and calibrates the multimodal sensing data to obtain multimodal demonstration trajectory data. The data synchronization and preprocessing module performs stage segmentation on the multimodal demonstration trajectory data to obtain corresponding trajectory segments; each trajectory segment corresponds to a work stage in the multimodal demonstration trajectory data; The multimodal coding module trains the encoder for each type of modal data separately to obtain the encoded feature vector of the corresponding type of modal data encoder; The multimodal coding module performs weighted fusion of all the encoded feature vectors based on the weights corresponding to each encoded feature vector to obtain a unified state representation; the unified state representation is used to characterize the multimodal fusion features of the dual-arm robot. The sub-task sequence decision module trains the sub-task sequence decision model based on the unified state representation; the sub-task sequence decision model is used to define the control strategy for both arms based on the unified state representation; The contact phase micro-action strategy module trains the contact phase micro-action strategy based on the unified state representation, so that the contact phase micro-action strategy outputs end control commands when in the contact-related phase. The multimodal coding module freezes all the trained modal data encoders into the modal data encoder library. The subtask sequence decision module calls the subtask sequence decision model in the decision model library, inputs the main task and the current unified state representation into the subtask sequence decision model, and obtains the subtask decision result. The main task is used to indicate the assembly task of the dual-arm robot; the current unified state represents the current multimodal fusion characteristics of the dual-arm robot. When the stage label indicates that the contact stage is in the contact stage, the micro-action strategy module obtains the end-effector control correction amount based on the current unified state representation; the contact stage includes the pre-insertion stage, the insertion stage, and the pressing stage; The dual-arm collaborative control and calculation module calculates the dual-arm control error based on the sub-task decision results and real-time multimodal data; The dual-arm collaborative control and calculation module generates end-of-arm control quantities based on the current unified state representation, the sub-task decision results, the end-of-arm control correction quantity, and the dual-arm control error. The dual-arm collaborative control and calculation module applies amplitude limits to the end-effector control quantity and converts the limited end-effector control quantity into joint control commands for both arms. The dual-arm collaborative control and calculation module sends the joint control commands to the dual-arm robot, driving the dual arms to perform assembly operations. The online evaluation and rollback module records execution logs, updates the training sample pool based on the execution logs, and iteratively retrains the modal data encoder, the subtask sequence decision model, and / or the contact phase micro-action strategy based on the updated training sample pool.
16. A dual-arm collaborative training device based on multimodal sensor fusion, characterized in that, include: The acquisition and calibration module is used to acquire multimodal sensing data corresponding to the demonstration operation, and to calibrate the multimodal sensing data to obtain multimodal demonstration trajectory data. The data synchronization and preprocessing module is used to segment the multimodal demonstration trajectory data into stages to obtain corresponding trajectory segments; each trajectory segment corresponds to a work stage in the multimodal demonstration trajectory data. A multimodal coding module is used to train each type of modal data encoder separately to obtain the coding feature vector of the corresponding type of modal data encoder; based on the weights corresponding to each coding feature vector, all coding feature vectors are weighted and fused to obtain a unified state representation; the unified state representation is used to characterize the multimodal fusion features of the dual-arm robot; Freeze the encoding parameters corresponding to all the trained modal data encoders into the modal data encoder library; The subtask sequence decision module is used to train the subtask sequence decision model based on the unified state representation; the subtask sequence decision model is used to define the control strategy for both arms based on the unified state representation. The contact phase micro-action strategy module is used to train the contact phase micro-action strategy based on the unified state representation, so that the contact phase micro-action strategy outputs end control commands when in the contact-related phase.
17. A dual-arm collaborative control device based on multimodal sensor fusion, characterized in that, include: The subtask sequence decision module is used to call the subtask sequence decision model in the decision model library, input the main task and the current unified state representation into the subtask sequence decision model, and obtain the subtask decision result; the main task is used to indicate the assembly task of the dual-arm robot; the current unified state represents the current multimodal fusion feature of the dual-arm robot. The contact phase micro-action strategy module is used to obtain the end-effector control correction amount based on the current unified state representation when the phase label indicates that it is in the contact phase; the contact phase includes the pre-insertion phase, the insertion phase, and the pressing phase; The dual-arm collaborative control and calculation module is used to calculate the dual-arm control error based on the sub-task decision results and real-time multimodal data; and to generate the end-of-arm control quantity based on the current unified state representation, the sub-task decision results, the end-of-arm control correction amount, and the dual-arm control error. An amplitude limit is applied to the end-effector control quantity, and the limited end-effector control quantity is converted into joint control commands for the dual arms; the joint control commands are sent to the dual-arm robot to drive the dual arms to perform assembly operations; The online assessment and rollback module is used to monitor abnormal indicators in real time and perform abnormal handling operations when abnormal indicators are triggered.
18. A program product, characterized in that, When the program product is executed by the processor, it implements the method as described in any one of claims 1-7, 8-14, or 15.