Digital twin and augmented reality guided human-machine collaborative assembly method and system
Patent Information
- Application Number
- CN202510516119.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-04-23
AI Technical Summary
[0005]针对现有自动化装配系统柔性不足、人工装配缺乏实时指引,导致多品种切换效率低、装配精度波动大,且装配状态检测依赖事后验证,工艺控制滞后的技术问题,本申请提供一种数字孪生和增强现实引导的人机协作装配方法及系统,通过动态数字孪生模型与AR实时指引融合,实现机械臂高精度协同与人工操作标准化,预存工艺参数快速适配多品种生产,并通过实时反馈数据提升装配过程的一致性控制
1. 本申请通过多视角图像数据与机械臂传感数据的实时融合处理,动态生成适配不同产品的机械臂轨迹规划指令,结合AR眼镜中实时叠加的装配路径三维指引,无需依赖固定编程逻辑即可快速切换装配任务,显著提升产线适应多品种生产的灵活性。
Smart Images

Figure CN120395821B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of workpiece assembly technology, specifically to a human-machine collaborative assembly method and system guided by digital twins and augmented reality. Background Technology
[0002] With the rapid development of industrial automation technology, the manufacturing industry is increasingly demanding higher assembly efficiency, precision, and flexible production. In high-value-added industries such as aerospace and precision instruments, assembly processes often involve complex procedures and high-precision requirements, making traditional manual operations insufficient to meet the needs of large-scale customized production. Meanwhile, the introduction of technologies such as artificial intelligence and machine vision is driving the development of assembly systems towards intelligence, and overcoming existing technological bottlenecks has become a key focus of the industry.
[0003] In existing technologies, fully automated assembly lines complete standardized assembly tasks using high-precision robotic arms and pre-programmed procedures, suitable for mass production scenarios. For example, vision positioning systems guide the robotic arms to grasp workpieces, or force control sensors enable the fitting and assembly of precision components. On the other hand, in purely manual assembly scenarios, operators rely on two-dimensional drawings or electronic process cards to complete assembly steps step by step, supplemented by measuring tools for accuracy verification. Some solutions also attempt to pre-simulate the assembly process using offline simulation software to optimize the robotic arm's motion path.
[0004] However, in existing technologies, fully automated systems rely on fixed programming logic, making it difficult to adapt to frequent process changes in multi-variety, small-batch production, and the cost of equipment reconfiguration is high; purely manual operation lacks real-time visual guidance, is easily affected by the operator's experience level, and the assembly accuracy fluctuates greatly; in addition, the status detection during the assembly process usually relies on post-event sampling inspection, which cannot provide real-time feedback on workpiece position deviation and contact mechanical parameters, resulting in a lag in process consistency control. Summary of the Invention
[0005] To address the technical problems of insufficient flexibility in existing automated assembly systems, lack of real-time guidance in manual assembly leading to low efficiency in switching between multiple product types, large fluctuations in assembly accuracy, reliance on post-processing verification for assembly status detection, and lagging process control, this application provides a human-machine collaborative assembly method and system guided by digital twins and augmented reality. By integrating a dynamic digital twin model with real-time AR guidance, it achieves high-precision collaboration of robotic arms and standardization of manual operations, allows for rapid adaptation of pre-stored process parameters to multi-product production, and improves consistency control of the assembly process through real-time feedback data.
[0006] In a first aspect, this application provides a human-computer collaborative assembly method guided by digital twins and augmented reality, comprising the following steps: S1. Real-time acquisition of multi-view image data and collaborative robotic arm sensor data, and acquisition of first-person image data of the assembly workbench using AR glasses worn by the operator; S2. Extract features from multi-view image data and first-person image data, fuse the extracted image features with the sensor data of the collaborative robotic arm to generate scene assembly information, and build a digital twin model of the assembly workstation based on the scene assembly information. S3. Configure the standard assembly unit space status and standard hand gesture templates for each assembly process stage; Based on the digital twin model, the spatial state of the actual assembly unit is identified and the matching degree with the spatial state of the standard assembly unit is calculated to determine the current assembly process stage. Identify operator hand gestures in first-person image data and calculate the matching degree with standard hand gesture templates for the current assembly process stage to analyze the operator's intent; The spatial state of the assembly unit includes the position and orientation of the collaborative robotic arm, workpiece, and tools; S4. Based on the current assembly process stage and operating standards, predict the next assembly process, generate manual assembly action instructions, collaborative robotic arm action instructions, and assembly effect simulation data. The assembly effect simulation data includes robotic arm motion trajectory prediction, workpiece assembly three-dimensional morphology simulation, and prediction simulation animation of the next assembly effect. S5. Display the manual assembly action instructions in the AR glasses' field of view in the form of a 3D model, arrow indicators, and / or text prompts. At the same time, the collaborative robotic arm executes the collaborative robotic arm action instructions, and displays the current assembly status, assembly accuracy detection data, force control feedback parameters, and prediction simulation animation of the next assembly effect in real time through the visualization interface of the digital twin model.
[0007] It should be further noted that in step S1, the sensor data of the collaborative robotic arm includes the pose data of each joint of the collaborative robotic arm and the state data of the end effector.
[0008] It should be further noted that the multi-view image data includes images of the collaborative robotic arm, operator, and assembly table from different angles. It should be further noted that in step S2, a deep convolutional neural network is used to extract features from the multi-view image data, and a lightweight convolutional neural network is used to extract features from the first-person image data. Among them, the deep convolutional neural network is selected from one of the ResNet series, VGG series, and EfficientNet series; the lightweight convolutional neural network is selected from one of the MobileNet series and ShuffleNet series.
[0009] It should be further noted that in step S2, the digital twin model includes a three-dimensional dynamic mapping of the collaborative robotic arm, workpiece, tools, and assembly workbench.
[0010] It should be further noted that in step S2, the scene assembly information includes the spatial coordinates of the end effector of the collaborative robotic arm, the three-dimensional pose matrix of the workpiece and the tool, and the assembly contact mechanical parameters. Data fusion employs state estimation algorithms to achieve spatiotemporal registration of multi-source data. The state estimation algorithms include Kalman filtering and particle filtering algorithms. The assembly contact mechanical parameters include at least one of pressure distribution and torque vector.
[0011] It should be further noted that the spatial coordinates of the end effector of the collaborative robotic arm are obtained by forward kinematics calculation of the pose data of each joint of the collaborative robotic arm; The three-dimensional pose matrix of the workpiece and tool is obtained by image feature triangulation. Assembly contact mechanical parameters are generated by fusing end effector state data with image contact area features.
[0012] It should be further noted that in step S2, the digital twin model is constructed using a 3D modeling engine, including: Based on the spatial coordinates of the end effector of the collaborative robot arm, real-time pose data of each joint of the collaborative robot arm are generated through inverse kinematics calculation, and a kinematic simulation model of the collaborative robot arm is established. Based on the three-dimensional pose matrix of the workpiece and tool, the preset CAD model of the workpiece and tool is aligned with the real-time pose to generate a three-dimensional dynamic model of the assembly position of the workpiece and tool. By updating the contact state model between the tool and the workpiece using assembly contact mechanics parameters, a visual mapping of assembly contact forces can be achieved.
[0013] It should be further explained that in step S3, a geometric pose calculation algorithm is used to identify the position and posture of the collaborative robotic arm, workpiece and tool, and a hand key point detection algorithm is used to identify the operator's hand movements in the first-person image data. Among them, the geometric pose calculation algorithm includes the PnP algorithm or the ICP algorithm; The hand keypoint detection algorithm includes at least one of the following: a real-time detection model based on deep learning and a feature matching algorithm based on traditional image processing. The real-time detection model based on deep learning is selected from the MediaPipe hand detection framework, the OpenPose hand detection module, and the monocular 3D hand pose estimation model.
[0014] It should be further explained that the process of analyzing the operational intent in step S3 includes: The three-dimensional coordinates of the operator's finger joints in consecutive frame images are extracted using a hand key point detection algorithm to generate a temporal sequence of hand movement trajectories; Input the time sequence of hand movement trajectory into a pre-trained action classification model to identify at least one operation type among grasping, translation, and rotation; The matching degree is calculated based on the identified operation type and the standard hand action template of the current assembly process stage. When the matching degree is lower than the preset threshold, it is determined that the operation intention is abnormal.
[0015] It should be further explained that in step S4, a pre-trained large model is used to predict the next assembly process. The large model is a sequence prediction model. It generates manual assembly action instructions and collaborative robotic arm action instructions through a feature association mechanism. The collaborative robotic arm action instructions include instruction data of posture constraints and trajectory planning parameters. The sequence prediction model is selected from the Transformer architecture or LSTM network, and the feature association mechanism includes attention mechanism or graph neural network.
[0016] It should be further noted that in step S4, the collaborative robotic arm motion command includes motion type, joint motion parameters and trajectory planning method. The motion type includes at least one of grasping, moving and installing motions, and the trajectory planning method includes B-spline interpolation or Bézier curve. When the action type is an installation action, the collaborative robotic arm action command also includes a force interaction control strategy, which includes admittance control or impedance control.
[0017] It should be further explained that the process of generating the assembly effect simulation data in step S4 includes: Based on the real-time pose data of the collaborative robotic arm in the digital twin model, the motion path of the end effector of the collaborative robotic arm in the next 5-10 seconds is predicted by the kinematic trajectory planning algorithm, and a trajectory coordinate sequence containing timestamps is generated. Based on the assembly tolerance requirements in the preset process standards, a physics engine is used to simulate the deformation of the assembled workpiece, generating a stress distribution cloud map of the workpiece contact surface and a three-dimensional model of the assembly gap. By fusing trajectory coordinate sequences with deformation simulation data, a lightweight rendering engine is used to generate real-time visualization simulation results of collaborative robotic arm motion trajectory prediction animation and workpiece assembly form.
[0018] It should be further explained that in step S5, the manual assembly action instructions are converted into three-dimensional models, arrow indicators and / or text prompts, a spatial positioning algorithm is used to establish a virtual and real coordinate system mapping relationship, and visual-assisted alignment is achieved through an optical perspective AR device; Spatial positioning algorithms include SLAM algorithms or marker-based positioning algorithms.
[0019] It should be further noted that when the 3D model is overlaid and displayed, the assembly completion level is distinguished by a status indicator scheme; Status identification schemes include color coding or texture mapping.
[0020] It should be further explained that during the execution of collaborative robotic arm action commands, the motion trajectory is adjusted in real time based on a safety collision detection algorithm. The safety collision detection algorithm uses a collision detection model to calculate the dynamic distance between the collaborative robotic arm and other objects. When the dynamic distance is less than a set threshold, a safety response strategy is triggered. The collision detection model includes a hierarchical bounding box algorithm or a convex hull decomposition algorithm, and the safety response strategy includes joint torque limiting or emergency braking.
[0021] It should be further noted that in step S5, the visualization interface of the digital twin model uses cross-platform rendering technology to simultaneously display the visualization content of trajectory error and tolerance distribution. Cross-platform rendering technologies include WebGL or WebGPU. Trajectory error visualization includes coordinate deviation curves between the actual trajectory and the planned trajectory, and motion envelope. Tolerance distribution visualization includes assembly gap heatmaps and tolerance over-limit warning signs.
[0022] Secondly, this application provides a digital twin and augmented reality-guided human-computer collaborative assembly system for implementing the aforementioned human-computer collaborative assembly method, comprising: The visual perception module includes AR glasses and multiple vision cameras set around the workstation, wherein the AR glasses are used to acquire first-person image data of the assembly workbench, and the vision cameras are used to acquire multi-view image data in real time. The mechanical equipment module includes a collaborative robotic arm control submodule and a collaborative robotic arm. The collaborative robotic arm is equipped with a sensing unit. The collaborative robotic arm control submodule is used to receive motion commands from the collaborative robotic arm and execute the motion commands through the collaborative robotic arm. The sensing unit is used to acquire sensing data from the collaborative robotic arm. The scene assembly information generation module is used to extract features from multi-view image data and first-person image data, and then fuse the extracted image features with the sensor data of the collaborative robotic arm to generate scene assembly information. The digital twin module includes a mapping submodule and a data visualization interface. The mapping submodule is used to build a digital twin model of the assembly workstation based on the scene assembly information. The data visualization interface is used to display the current assembly status, assembly accuracy detection data, force control feedback parameters and prediction simulation animation of the next assembly effect in real time. The identification and prediction module is used to configure the operation standards, standard assembly unit spatial states, and standard hand action templates for each assembly process stage. Based on the digital twin model, it identifies the actual assembly unit spatial state and calculates the matching degree with the standard assembly unit spatial state to determine the current assembly process stage; it identifies the operator's hand actions in the first-person image data and calculates the matching degree with the standard hand action template for the current assembly process stage to analyze the operation intention; and it combines the current assembly process stage and operation standards to predict the next assembly process, generating manual assembly action instructions, collaborative robotic arm action instructions, and assembly effect simulation data. The AR interactive display module is used to overlay manual assembly instructions in the field of view as 3D models, arrow indicators and / or text prompts.
[0023] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described digital twin and augmented reality-guided human-computer collaborative assembly method.
[0024] Fourthly, this application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the aforementioned digital twin and augmented reality-guided human-computer collaborative assembly method.
[0025] As can be seen from the above technical solutions, this application has the following advantages: 1. This application dynamically generates robotic arm trajectory planning instructions adapted to different products through real-time fusion processing of multi-view image data and robotic arm sensor data. Combined with the real-time superimposed 3D assembly path guidance in AR glasses, assembly tasks can be quickly switched without relying on fixed programming logic, significantly improving the flexibility of the production line to adapt to multi-variety production.
[0026] 2. This application uses real-time 3D motion guidance from AR glasses and a digital twin model to accurately map the workpiece assembly posture, simultaneously constraining the manual operation path and the robotic arm's movement trajectory. At the same time, it corrects operational deviations in real time based on hand motion recognition in the first-person image, ensuring standardized execution of human-machine collaborative actions and improving the consistency of assembly accuracy.
[0027] 3. This application utilizes the visualization interface of the digital twin model to provide real-time feedback on the trajectory error of the robotic arm, the workpiece pose tolerance, and predictive simulation animation, enabling the operator and the robotic arm to synchronously adjust motion parameters during the assembly process. Through dynamic closed-loop control, real-time verification of process consistency is achieved, avoiding rework losses caused by the lag in traditional detection. Attached Figure Description
[0028] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is a flowchart of a human-computer collaborative assembly method guided by digital twins and augmented reality in one embodiment of this application.
[0030] Figure 2 This is a schematic block diagram of a human-machine collaborative assembly system guided by digital twins and augmented reality in one embodiment of this application.
[0031] Figure 3 This is a schematic diagram illustrating an application scenario of a human-machine collaborative assembly system guided by digital twins and augmented reality in one embodiment of this application.
[0032] Figure 4 This is a flowchart illustrating the operation of a human-machine collaborative assembly system guided by digital twins and augmented reality in one embodiment of this application.
[0033] Figure 5 This is a schematic diagram of the hardware structure of an electronic device in one embodiment of this application.
[0034] In the diagram, 1-AR glasses, 2-vision camera, 3-collaborative robotic arm, 4-data visualization interface, 5-computing host. Detailed Implementation
[0035] To make the purpose, features, and advantages of this application more apparent and understandable, specific embodiments and accompanying drawings will be used to clearly and completely describe the technical solution protected by this application. Obviously, the embodiments described below are only some embodiments of this application, and not all embodiments. Based on the embodiments in this patent, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this patent.
[0036] This application relates to a digital twin and augmented reality-guided human-machine collaborative assembly method primarily targeting the field of workpiece assembly technology. Through real-time fusion processing of multi-view image data and robotic arm sensor data, it dynamically generates robotic arm trajectory planning instructions adapted to different products. Combined with real-time overlaid 3D assembly path guidance in AR glasses, it allows for rapid switching of assembly tasks without relying on fixed programming logic, significantly improving the flexibility of production lines to adapt to multi-variety production. By using real-time 3D motion guidance from AR glasses and precise mapping of workpiece assembly poses to the digital twin model, it synchronously constrains the manual operation path and the robotic arm's motion trajectory. Simultaneously, based on hand motion recognition in first-person images, it corrects operational deviations in real time, ensuring standardized execution of human-machine collaborative actions and improving assembly accuracy consistency. Utilizing the visualization interface of the digital twin model, it provides real-time feedback on robotic arm trajectory errors, workpiece pose tolerances, and predictive simulation animations, enabling operators and robotic arms to synchronously adjust motion parameters during assembly. Through dynamic closed-loop control, it achieves real-time verification of process consistency, avoiding rework losses caused by traditional detection lag.
[0037] The digital twin and augmented reality-guided human-machine collaborative assembly method involved in this application mainly addresses the technical problems of insufficient flexibility in existing automated assembly systems, lack of real-time guidance for manual assembly, resulting in low efficiency in switching between multiple product types, large fluctuations in assembly accuracy, and reliance on post-process verification for assembly status detection, leading to lagging process control.
[0038] The following describes in detail the human-machine collaborative assembly method guided by digital twins and augmented reality related to this application. Specific details, such as particular system structures and technologies, are presented for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application can also be implemented in other embodiments without these specific details.
[0039] In the digital twin and augmented reality-guided human-machine collaborative assembly methods involved in this application, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or sets thereof. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0040] To facilitate a clear description of the technical solutions of this application, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" do not necessarily imply that they are different.
[0041] The terms "one embodiment" or "some embodiments" used in this application mean that one or more embodiments of this application include the specific features, structures, or characteristics described in that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this application do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0043] The digital twin and augmented reality-guided human-computer collaborative assembly method provided in this application embodiment is executed by a computer device, and correspondingly, the digital twin and augmented reality-guided human-computer collaborative assembly system runs in the computer device.
[0044] The following is a definition of some terms used in this plan to facilitate a better understanding of the plan: Digital twin (DT) is a technology for digitally modeling physical entities or systems throughout their entire lifecycle. It utilizes the Internet of Things (IoT), sensor networks, big data analytics, or artificial intelligence algorithms to construct a high-precision virtual mapping that is dynamically synchronized with the real object. Its core mechanism involves continuous interaction between the physical world and the digital space through real-time data acquisition and two-way communication: changes in the physical object's state (such as temperature, pressure, and motion trajectory) are transmitted to the digital twin, which then uses simulation to predict potential behaviors or optimization strategies and feeds the results back to the physical system to guide decision-making. This technology relies on interdisciplinary integration, including cloud computing, edge computing, and machine learning, and is widely used in industrial manufacturing, energy management, smart cities, and healthcare, supporting predictive maintenance, process optimization, and complex system management.
[0045] Augmented Reality (AR) is an interactive paradigm that seamlessly overlays virtual information (such as 3D models, text, and images) onto the real environment using computer vision, spatial positioning, and real-time rendering technologies. Its technological framework includes environmental perception (capturing physical space through cameras, LiDAR, or depth sensors), spatial registration (achieving geometric alignment between virtual objects and real-world scenes using SLAM (Simultaneous Localization and Mapping) technology), and display output (presenting the blended image through head-mounted displays, smart glasses, or mobile device screens). The core characteristics of AR lie in the real-time and interactive nature of this virtual-real fusion, emphasizing expanding information dimensions while preserving the user's perception of the real world. This can be achieved through methods such as overlaying data visualization layers to assist operation or using spatial anchoring to persist virtual objects.
[0046] Figure 1 This is a flowchart of a digital twin and augmented reality-guided human-computer collaborative assembly method according to an embodiment of this application. Figure 1 The implementing entity can be a human-machine collaborative assembly system guided by digital twins and augmented reality. Depending on different needs, the order of steps in this flowchart can be changed, and some can be omitted.
[0047] like Figure 1 As shown, this digital twin and augmented reality-guided human-computer collaborative assembly method includes: Step S1: Real-time acquisition of multi-view image data and collaborative robotic arm sensor data, and acquisition of first-person image data of the assembly workbench using AR glasses worn by the operator.
[0048] In some specific embodiments, the multi-view image data includes images of the collaborative robotic arm, operator, and assembly table from different angles. Through a collaborative data acquisition system using a multi-view camera array and AR glasses, the system achieves global observation of the assembly scene and capture of local details from the operator's perspective. It can construct a full-element perception data pool covering robotic arm movement, workpiece status, and personnel operation, providing spatiotemporally aligned multimodal raw data input for digital twin modeling.
[0049] In some specific embodiments, the collaborative robotic arm sensing data includes the pose data of each joint of the collaborative robotic arm and the state data of the end effector.
[0050] By defining the specific dimensions and parameter types of the robotic arm's sensing data and clarifying the data acquisition standards for the correlation between joint pose and end effector state, a complete kinematic modeling foundation for the robotic arm's kinematic chain is realized. This provides accurate data input for subsequent inverse kinematics calculation, trajectory planning, and contact force analysis, and can enhance the consistency of the digital twin system's representation of the robotic arm's dynamic behavior and the realism of the physical simulation.
[0051] Step S2: Feature extraction is performed on multi-view image data and first-person image data. The extracted image features are fused with the sensor data of the collaborative robotic arm to generate scene assembly information. A digital twin model of the assembly workstation is then constructed based on the scene assembly information.
[0052] In some specific embodiments, the digital twin model includes a three-dimensional dynamic mapping of a collaborative robotic arm, workpiece, tools, and assembly table.
[0053] Through feature extraction from deep neural networks and multi-source data fusion algorithms, the ability to jointly model the spatial relationships and physical properties of assembly elements is realized. It can generate dynamic scene description models that include the kinematic state of the robotic arm, the pose of the workpiece and tool, and contact mechanics, thus establishing a high-precision digital twin benchmark for virtual-real interaction and process analysis.
[0054] In some specific embodiments, deep convolutional neural networks are used to extract features from multi-view image data, and lightweight convolutional neural networks are used to extract features from first-person image data. Among them, the deep convolutional neural network is selected from one of the ResNet series, VGG series, and EfficientNet series; the lightweight convolutional neural network is selected from one of the MobileNet series and ShuffleNet series.
[0055] By designing a multimodal image processing architecture that combines deep convolutional networks and lightweight networks with differentiated design, the computing power is optimized and allocated to meet the needs of high-dimensional feature extraction from multi-view images and real-time processing of first-person images. This enables the collaborative processing capability of feature capture in complex scenes and low-latency AR display, ensuring efficient fusion and feature alignment of multi-source visual data under spatiotemporal consistency constraints.
[0056] In some specific embodiments, the scene assembly information includes the spatial coordinates of the end effector of the collaborative robotic arm, the three-dimensional pose matrix of the workpiece and the tool, and the assembly contact mechanical parameters; Data fusion employs state estimation algorithms to achieve spatiotemporal registration of multi-source data. State estimation algorithms include Kalman filtering and particle filtering algorithms. Assembly contact mechanical parameters include at least one of pressure distribution and torque vector.
[0057] By using the multi-source data spatiotemporal registration technology of state estimation algorithm, visual features, robotic arm pose and mechanical parameters are cross-modal correlated and fused, realizing unified quantitative modeling of spatial coordinates, contact state and dynamic mechanical properties of assembly elements. This can establish a multi-dimensional data fusion decision benchmark for quality prediction, anomaly detection and process optimization in the assembly process.
[0058] In some specific embodiments, the spatial coordinates of the end effector of the collaborative robotic arm are obtained by forward kinematics calculation of the pose data of each joint of the collaborative robotic arm; The three-dimensional pose matrix of the workpiece and tool is obtained by image feature triangulation. Assembly contact mechanical parameters are generated by fusing end effector state data with image contact area features.
[0059] By employing a collaborative solution mechanism combining forward kinematics calculation and image feature triangulation, and integrating feature complementarity verification of the end effector state and visual contact area, cross-sensor data verification of the robotic arm's end-effector positioning accuracy and workpiece pose calculation is achieved. This eliminates the error accumulation problem from a single data source and improves the robustness of spatial relationship modeling of assembly elements.
[0060] In some specific embodiments, a 3D modeling engine is used to construct a digital twin model, including: Based on the spatial coordinates of the end effector of the collaborative robot arm, real-time pose data of each joint of the collaborative robot arm are generated through inverse kinematics calculation, and a kinematic simulation model of the collaborative robot arm is established. Based on the three-dimensional pose matrix of the workpiece and tool, the preset CAD model of the workpiece and tool is aligned with the real-time pose to generate a three-dimensional dynamic model of the assembly position of the workpiece and tool. By updating the contact state model between the tool and the workpiece using assembly contact mechanics parameters, a visual mapping of assembly contact forces can be achieved.
[0061] By using inverse kinematics calculation and dynamic alignment technology with a preset CAD model, combined with state model updates driven by contact mechanics parameters, a high-fidelity visualization mapping of robotic arm motion simulation, workpiece assembly process and contact force evolution is achieved. This provides operators with a 3D interactive interface that combines physical realism and real-time performance to support precise human-machine collaborative decision-making.
[0062] Step S3: Configure the operation standards, standard assembly unit space status, and standard hand action templates for each assembly process stage; Based on the digital twin model, the spatial state of the actual assembly unit is identified and the matching degree with the spatial state of the standard assembly unit is calculated to determine the current assembly process stage. Identify operator hand gestures in first-person image data and calculate the matching degree with standard hand gesture templates for the current assembly process stage to analyze the operator's intent; The spatial state of the assembly unit includes the position and orientation of the collaborative robotic arm, workpiece, and tools.
[0063] By using the parallel analysis technology of mechanical system pose recognition and operator hand movements, the bidirectional state perception and intent matching capability of the assembly process stage is realized. This can establish a contextual understanding framework for human-machine task collaboration and provide a basis for process compliance verification for dynamic instruction generation.
[0064] In some specific embodiments, a geometric pose calculation algorithm is used to identify the position and posture of the collaborative robotic arm, workpiece and tool, and a hand key point detection algorithm is used to identify the operator's hand movements in the first-person image data; Among them, the geometric pose calculation algorithm includes the PnP algorithm or the ICP algorithm; The hand keypoint detection algorithm includes at least one of the following: a real-time detection model based on deep learning and a feature matching algorithm based on traditional image processing. The real-time detection model based on deep learning is selected from the MediaPipe hand detection framework, the OpenPose hand detection module, and the monocular 3D hand pose estimation model.
[0065] By employing a parallel processing architecture combining geometric pose calculation algorithms and deep learning hand keypoint detection technology, a two-way real-time monitoring capability for mechanical system state perception and operator behavior analysis is achieved. This enables the establishment of a rapid identification channel for human-computer interaction intentions and a mechanism for verifying action compliance, providing two-way data support for dynamic task allocation and safe collaborative control.
[0066] In some specific embodiments, the process of analyzing operational intent includes: The three-dimensional coordinates of the operator's finger joints in consecutive frame images are extracted using a hand key point detection algorithm to generate a temporal sequence of hand movement trajectories; Input the time sequence of hand movement trajectory into a pre-trained action classification model to identify at least one operation type among grasping, translation, and rotation; Configure standard hand gesture templates for each assembly process stage. Calculate the matching degree between the identified operation type and the standard hand gesture template for the current assembly process stage. If the matching degree is lower than a preset threshold, it is determined that the operation intention is abnormal.
[0067] By combining the temporal analysis of hand motion trajectories with the pre-trained motion classification model, and integrating it with the standard hand motion template matching mechanism in the assembly process, the system achieves continuous frame parsing of operational intentions and real-time early warning of abnormal operation modes. This enhances the granularity and response timeliness of operational compliance detection during human-machine collaboration.
[0068] Step S4: Based on the current assembly process stage and operating standards, predict the next assembly process, generate manual assembly action instructions, collaborative robotic arm action instructions, and assembly effect simulation data. The assembly effect simulation data includes the robotic arm motion trajectory prediction, the three-dimensional morphology simulation of the workpiece after assembly, and the prediction simulation animation of the next assembly effect.
[0069] By dynamically matching the current assembly process stage identification with the preset process stage operation standards, continuous deduction and process prediction of the assembly process chain are realized. Based on the spatial state analysis and process logic constraints of the digital twin model, standardized manual assembly action instructions, collaborative robotic arm action instructions, and assembly effect simulation data are generated simultaneously. The robotic arm motion trajectory prediction in the assembly effect simulation data strictly follows the spatial path constraints of the process specifications. The three-dimensional morphology simulation of the assembled workpiece accurately reflects the tolerance fit relationship. The prediction simulation animation fully presents the temporal logic of tool-workpiece interaction, ensuring that the human-machine collaborative instructions achieve closed-loop verification between the assembly sequence, spatial trajectory, and process standards, forming an executable, predictable, and traceable process advancement plan.
[0070] In some specific embodiments, a pre-trained large model is used to predict the next assembly process. The large model is a sequence prediction model. It generates manual assembly action instructions and collaborative robotic arm action instructions through a feature association mechanism. The collaborative robotic arm action instructions include instruction data of posture constraints and trajectory planning parameters. The sequence prediction model is selected from the Transformer architecture or LSTM network, and the feature association mechanism includes attention mechanism or graph neural network.
[0071] By leveraging the feature association mechanism of the sequence prediction model and the dynamic allocation strategy of attention weights, the system achieves context awareness and collaborative planning capabilities for multi-step actions in the assembly process stage. This optimizes the logical coherence and task adaptability of human-machine instruction generation, and supports the dynamic adjustment of complex assembly processes and the needs of nonlinear process reasoning.
[0072] In some specific embodiments, the collaborative robotic arm motion instructions include motion type, joint motion parameters, and trajectory planning method. The motion type includes at least one of grasping, moving, and installing motions, and the trajectory planning method includes B-spline interpolation or Bézier curve. When the action type is an installation action, the collaborative robotic arm action command also includes a force interaction control strategy, which includes admittance control or impedance control.
[0073] By using a multi-level parameterized description architecture for robotic arm motion commands and an adaptive selection mechanism that integrates force interaction control strategies and trajectory planning methods, the system achieves compliant control and contact stability assurance for robotic arm motion in precision assembly scenarios, balancing the dynamic constraints of high-precision positioning requirements and safe human-machine interaction.
[0074] In some specific embodiments, the process of generating assembly effect simulation data includes: Based on the real-time pose data of the collaborative robotic arm in the digital twin model, the motion path of the end effector of the collaborative robotic arm in the next 5-10 seconds is predicted by the kinematic trajectory planning algorithm, and a trajectory coordinate sequence containing timestamps is generated. Based on the assembly tolerance requirements in the preset process standards, a physics engine is used to simulate the deformation of the assembled workpiece, generating a stress distribution cloud map of the workpiece contact surface and a three-dimensional model of the assembly gap. By fusing trajectory coordinate sequences with deformation simulation data, a lightweight rendering engine is used to generate real-time visualization simulation results of collaborative robotic arm motion trajectory prediction animation and workpiece assembly form.
[0075] By integrating kinematic trajectory planning and physical engine deformation simulation, a multi-dimensional data generation capability is achieved for predicting robotic arm motion and visualizing assembly results. This enables the construction of a forward-looking human-machine collaborative guidance system to enhance the operator's initiative in spatial prediction and process correction of assembly results.
[0076] Step S5: The manual assembly action instructions are overlaid and displayed in the AR glasses' field of view in the form of a 3D model, arrow indicators, and / or text prompts. At the same time, the collaborative robotic arm executes the collaborative robotic arm action instructions, and the current assembly status, assembly accuracy detection data, force control feedback parameters, and prediction simulation animation of the next assembly effect are displayed in real time through the visualization interface of the digital twin model.
[0077] By using a closed-loop design of virtual and real interaction through AR guidance information and digital twin feedback, a visual monitoring and real-time correction mechanism for the operation process is realized. This can simultaneously improve the intuitiveness of human operation and the controllability of robotic arm movement to achieve the goal of efficient and safe human-machine collaboration.
[0078] In some specific embodiments, manual assembly instructions are transformed into three-dimensional models, arrow indicators and / or text prompts, a spatial positioning algorithm is used to establish a mapping relationship between virtual and real coordinate systems, and visual-assisted alignment is achieved through an optical perspective AR device; Spatial positioning algorithms include SLAM algorithms or marker-based positioning algorithms.
[0079] By establishing a precise mapping relationship between virtual and real coordinate systems through spatial positioning algorithms and combining it with the dynamic overlay presentation of multimodal AR prompts, the assembly guidance information from the operator's perspective is seamlessly integrated with the physical scene. This can reduce the cognitive conversion cost in the human-computer interaction process and improve the spatial consistency of operation steps.
[0080] In some specific embodiments, when the 3D model is overlaid and displayed, a status identification scheme is used to distinguish the degree of assembly completion; Status identification schemes include color coding or texture mapping.
[0081] Through a dynamic status identification design scheme using color coding and texture mapping, a multi-dimensional visual representation of assembly progress and quality indicators is achieved. This optimizes the operator's efficiency in extracting information from complex assembly states and enhances their ability to focus attention, thereby promoting precise control of key process nodes.
[0082] In some specific embodiments, the collaborative robotic arm adjusts its motion trajectory in real time based on a safety collision detection algorithm during the execution of collaborative robotic arm action commands. The safety collision detection algorithm uses a collision detection model to calculate the dynamic distance between the collaborative robotic arm and other objects. When the dynamic distance is less than a set threshold, a safety response strategy is triggered. The collision detection model includes a hierarchical bounding box algorithm or a convex hull decomposition algorithm, and the safety response strategy includes joint torque limiting or emergency braking.
[0083] Through a collaborative protection mechanism of dynamic collision detection model and multi-level safety response strategy, the robot arm's motion trajectory can be predicted in real time and autonomously avoid obstacles. This can build a safety protection closed loop in the human-machine collaborative space to reduce the probability of sudden collision events and ensure the continuity of collaborative operations.
[0084] In some specific embodiments, the visualization interface of the digital twin model uses cross-platform rendering technology to simultaneously display the visualization content of trajectory error and tolerance distribution; Cross-platform rendering technologies include WebGL or WebGPU. Trajectory error visualization includes coordinate deviation curves between the actual trajectory and the planned trajectory, and motion envelope. Tolerance distribution visualization includes assembly gap heatmaps and tolerance over-limit warning signs.
[0085] By integrating a visualization analysis module for trajectory error and tolerance distribution through cross-platform rendering technology, it achieves multi-dimensional synchronous presentation of assembly accuracy data and panoramic monitoring capabilities for process compliance. This enhances the efficiency of rapid location and source tracing analysis of quality issues to support continuous optimization and iteration of the assembly process.
[0086] In one specific embodiment, the digital twin and augmented reality-guided human-computer collaborative assembly method includes: Step S1: Real-time acquisition of multi-view image data and collaborative robotic arm sensor data, and acquisition of first-person image data of the assembly workbench using AR glasses worn by the operator. The multi-view image data includes collaborative robotic arm images, operator images and assembly workbench images from different angles. The collaborative robotic arm sensor data includes pose data of each joint of the collaborative robotic arm and end effector status data. Step S2: Use a deep convolutional neural network to extract features from multi-view image data, use a lightweight convolutional neural network to extract features from first-person image data, fuse the extracted image features with the sensor data of the collaborative robotic arm to generate scene assembly information, and build a digital twin model of the assembly workstation based on the scene assembly information, including the three-dimensional dynamic mapping of the collaborative robotic arm, workpiece, tools and assembly workbench. Among them, the data fusion uses a particle filter algorithm to achieve spatiotemporal registration of multi-source data and generate scene assembly information, including the spatial coordinates of the end effector of the collaborative robotic arm, the three-dimensional pose matrix of the workpiece and the tool, and the assembly contact mechanical parameters, including at least one of pressure distribution and torque vector. The spatial coordinates of the end effector of the collaborative robotic arm are obtained by forward kinematics calculation of the pose data of each joint of the collaborative robotic arm. The three-dimensional pose matrix of the workpiece and tool is obtained by image feature triangulation. The assembly contact mechanical parameters are generated by fusing end effector state data with image contact area features; A digital twin model is constructed using a 3D modeling engine, where: Based on the spatial coordinates of the end effector of the collaborative robot arm, real-time pose data of each joint of the collaborative robot arm are generated through inverse kinematics calculation, and a kinematic simulation model of the collaborative robot arm is established. Based on the three-dimensional pose matrix of the workpiece and tool, the preset CAD model of the workpiece and tool is aligned with the real-time pose to generate a three-dimensional dynamic model of the assembly position of the workpiece and tool. By updating the contact state model between the tool and the workpiece through assembly contact mechanics parameters, a visual mapping of assembly contact forces can be achieved. Step S3: Configure the operation standards, standard assembly unit space status, and standard hand action templates for each assembly process stage; Based on the digital twin model, the spatial state of the actual assembly unit is identified and the matching degree with the spatial state of the standard assembly unit is calculated to determine the current assembly process stage. Identify operator hand gestures in first-person image data and calculate the matching degree with standard hand gesture templates for the current assembly process stage to analyze the operator's intent; The spatial state of the assembly unit includes the position and orientation of the collaborative robotic arm, workpiece, and tools; Among them, the geometric pose calculation algorithm is the ICP algorithm, and the hand key point detection algorithm is the OpenPose hand detection module; The process of analyzing operational intent includes: The three-dimensional coordinates of the operator's finger joints in consecutive frame images are extracted using a hand key point detection algorithm to generate a temporal sequence of hand movement trajectories; Input the time sequence of hand movement trajectory into a pre-trained action classification model to identify at least one operation type among grasping, translation, and rotation; The matching degree is calculated based on the identified operation type and the standard hand action template of the current assembly process stage. When the matching degree is lower than the preset threshold, it is determined that the operation intention is abnormal. Step S4: Combining the current assembly process stage and operating standards, use a pre-trained large model to predict the next assembly process, generate manual assembly action instructions, collaborative robotic arm action instructions, and assembly effect simulation data. The assembly effect simulation data includes robotic arm motion trajectory prediction, three-dimensional morphology simulation of the workpiece after assembly, and prediction simulation animation of the next assembly effect. Among them, the large model is an LSTM network model, which generates manual assembly action instructions and collaborative robotic arm action instructions through a feature association mechanism. The collaborative robotic arm action instructions include instruction data of posture constraints and trajectory planning parameters. The feature association mechanism is an attention mechanism. The collaborative robotic arm motion instructions include motion type, joint motion parameters and trajectory planning method. The motion type includes at least one of grasping, moving and installing motions. The trajectory planning method is Bézier curve. When the motion type is an installing motion, the collaborative robotic arm motion instructions also include force interaction control strategy, which includes admittance control or impedance control. The process of generating assembly effect simulation data includes: Based on the real-time pose data of the collaborative robotic arm in the digital twin model, the motion path of the end effector of the collaborative robotic arm in the next 5-10 seconds is predicted by the kinematic trajectory planning algorithm, and a trajectory coordinate sequence containing timestamps is generated. Based on the assembly tolerance requirements in the preset process standards, a physics engine is used to simulate the deformation of the assembled workpiece, generating a stress distribution cloud map of the workpiece contact surface and a three-dimensional model of the assembly gap. By fusing trajectory coordinate sequences with deformation simulation data, a lightweight rendering engine is used to generate a collaborative robotic arm motion trajectory prediction animation and a real-time visual simulation result of the workpiece assembly shape. Step S5: The manual assembly action instructions are presented in the form of a 3D model, arrow indicators and / or text prompts. A virtual and real coordinate system mapping relationship is established using a spatial positioning algorithm. Visual alignment is achieved through an optical perspective AR device. The images are then overlaid and displayed in the field of view of the AR glasses. The spatial positioning algorithm is a marker point positioning algorithm. When the 3D model is overlaid and displayed, the assembly completion level is distinguished by a status identification scheme, which includes color coding or texture mapping. At the same time, the collaborative robotic arm executes the collaborative robotic arm action commands, and displays the current assembly status, assembly accuracy detection data, force control feedback parameters and the prediction simulation animation of the next assembly effect in real time through the visualization interface of the digital twin model. During the execution of collaborative robotic arm action commands, the motion trajectory is adjusted in real time based on a safety collision detection algorithm. The safety collision detection algorithm uses a collision detection model to calculate the dynamic distance between the collaborative robotic arm and other objects. When the dynamic distance is less than a set threshold, a safety response strategy is triggered. Collision detection models include hierarchical bounding box algorithms or convex hull decomposition algorithms, and safety response strategies include joint torque limiting or emergency braking. The visualization interface of the digital twin model uses cross-platform rendering technology to simultaneously display the trajectory error visualization and tolerance distribution visualization. The cross-platform rendering technology uses WebGL. The trajectory error visualization content includes the coordinate deviation curve between the actual trajectory and the planned trajectory, and the motion envelope. The tolerance distribution visualization content includes the assembly gap heat map and tolerance over-limit warning sign.
[0087] The following are embodiments of a digital twin and augmented reality-guided human-computer collaborative assembly system provided in this disclosure. This human-computer collaborative assembly system belongs to the same inventive concept as the digital twin and augmented reality-guided human-computer collaborative assembly methods in the above embodiments. For details not described in detail in the embodiments of the digital twin and augmented reality-guided human-computer collaborative assembly system, please refer to the embodiments of the digital twin and augmented reality-guided human-computer collaborative assembly methods described above.
[0088] Mobile terminals implementing various embodiments of the present application will now be described with reference to the accompanying drawings. In the following description, suffixes such as "module," "part," or "unit" used to denote elements are used only for the convenience of illustrating the embodiments of the present application and have no specific meaning in themselves. Therefore, "module" and "part" can be used interchangeably.
[0089] Figure 2 This is a schematic block diagram of the human-machine collaborative assembly system guided by digital twins and augmented reality in this embodiment. Figure 3 This is a schematic diagram illustrating the application scenario of the digital twin and augmented reality-guided human-machine collaborative assembly system in this embodiment, such as... Figures 2-3 As shown, the digital twin and augmented reality-guided human-machine collaborative assembly system includes: The visual perception module includes AR glasses 1 and multiple visual cameras 2 set around the workstation, wherein AR glasses 1 is used to acquire first-person image data of the assembly workbench, and visual cameras 2 are used to acquire multi-view image data in real time. The mechanical equipment module includes a collaborative robotic arm control submodule and a collaborative robotic arm 3. A sensing unit is deployed on the collaborative robotic arm 3. The collaborative robotic arm control submodule is used to receive the collaborative robotic arm's motion commands and execute the collaborative robotic arm's motion commands through the collaborative robotic arm 3. The sensing unit is used to acquire the collaborative robotic arm's sensing data. The scene assembly information generation module is used to extract features from multi-view image data and first-person image data, and then fuse the extracted image features with the sensor data of the collaborative robotic arm to generate scene assembly information. The digital twin module includes a mapping submodule and a data visualization interface 4. The mapping submodule is used to build a digital twin model of the assembly workstation based on the scene assembly information. The data visualization interface 4 is used to display the current assembly status, assembly accuracy detection data, force control feedback parameters and prediction simulation animation of the next assembly effect in real time. The identification and prediction module is used to configure the operation standards, standard assembly unit spatial states, and standard hand action templates for each assembly process stage. Based on the digital twin model, it identifies the actual assembly unit spatial state and calculates the matching degree with the standard assembly unit spatial state to determine the current assembly process stage; it identifies the operator's hand actions in the first-person image data and calculates the matching degree with the standard hand action template for the current assembly process stage to analyze the operation intention; and it combines the current assembly process stage and operation standards to predict the next assembly process, generating manual assembly action instructions, collaborative robotic arm action instructions, and assembly effect simulation data. AR interactive display module is used to overlay and display manual assembly action instructions in the field of view in the form of 3D models, arrow indicators and / or text prompts. Among them, the feature extraction and data fusion function of the scene assembly information generation module, the model dynamic update function of the mapping sub-module in the digital twin module, the object recognition and process prediction function of the recognition and prediction module, and the virtual and real space alignment function of the AR interactive display module are all supported by the computing host 5 deployed locally on the workstation.
[0090] The human-computer collaborative assembly system of this embodiment is used to realize a human-computer collaborative assembly method guided by digital twins and augmented reality, and the steps include: S1. Real-time acquisition of multi-view image data and collaborative robotic arm sensor data, and acquisition of first-person image data of the assembly workbench using AR glasses worn by the operator; S2. Extract features from multi-view image data and first-person image data, fuse the extracted image features with the sensor data of the collaborative robotic arm to generate scene assembly information, and build a digital twin model of the assembly workstation based on the scene assembly information. S3. Configure the standard assembly unit space status and standard hand gesture templates for each assembly process stage; Based on the digital twin model, the spatial state of the actual assembly unit is identified and the matching degree with the spatial state of the standard assembly unit is calculated to determine the current assembly process stage. Identify operator hand gestures in first-person image data and calculate the matching degree with standard hand gesture templates for the current assembly process stage to analyze the operator's intent; The spatial state of the assembly unit includes the position and orientation of the collaborative robotic arm, workpiece, and tools; S4. Based on the current assembly process stage and operating standards, predict the next assembly process, generate manual assembly action instructions, collaborative robotic arm action instructions, and assembly effect simulation data. The assembly effect simulation data includes robotic arm motion trajectory prediction, workpiece assembly three-dimensional morphology simulation, and prediction simulation animation of the next assembly effect. S5. Display the manual assembly action instructions in the AR glasses' field of view in the form of a 3D model, arrow indicators, and / or text prompts. At the same time, the collaborative robotic arm executes the collaborative robotic arm action instructions, and displays the current assembly status, assembly accuracy detection data, force control feedback parameters, and prediction simulation animation of the next assembly effect in real time through the visualization interface of the digital twin model.
[0091] In this embodiment, the operation flowchart of the human-machine collaborative assembly system is as follows: Figure 4 As shown.
[0092] This application also provides an electronic device for implementing various embodiments of this application, the electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor.
[0093] Those skilled in the art will understand that the electronic device structure involved in the embodiments of this application does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0094] Figure 5 A schematic diagram of the hardware structure of an electronic device for implementing the various embodiments of this application.
[0095] Electronic devices include, but are not limited to, components such as processors and memory. Those skilled in the art will understand that the electronic device structures described in the embodiments of this application do not constitute a limitation on the electronic device. An electronic device may include more or fewer components than illustrated, or may combine certain components, or have different component arrangements.
[0096] In embodiments of this application, electronic devices include, but are not limited to, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may also represent various forms of mobile devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of this application described and / or claimed herein.
[0097] In this application embodiment, the processor can be implemented using at least one of an Application-Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a processor, a controller, a microcontroller, a microprocessor, or an electronic unit designed to perform the functions described herein. In some cases, such implementations can be implemented within a controller. For software implementations, implementations such as processes or functions can be implemented with separate software modules that allow the performance of at least one function or operation. The software code can be implemented by a software application (or program) written in any suitable programming language, and the software code can be stored in memory and executed by the controller.
[0098] In addition, the electronic device includes some functional modules not shown, which will not be described in detail here.
[0099] Those skilled in the art will understand that the various aspects of the electronic device provided in this application can be implemented as a method, system, or program product. Therefore, the various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, which can be collectively referred to herein as a "circuit," "module," or "system."
[0100] This application also provides a storage medium storing a program product capable of implementing a human-computer collaborative assembly method guided by digital twins and augmented reality. In some possible implementations, various aspects of this disclosure can also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps described in the foregoing "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.
[0101] The storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0102] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A human-computer collaborative assembly method guided by digital twins and augmented reality, characterized in that, include: S1. Real-time acquisition of multi-view image data and collaborative robotic arm sensor data, and acquisition of first-person image data of the assembly workbench using AR glasses worn by the operator; S2. Extract features from multi-view image data and first-person image data, fuse the extracted image features with the sensor data of the collaborative robotic arm to generate scene assembly information, and build a digital twin model of the assembly workstation based on the scene assembly information. S3. Configure the operating standards, standard assembly unit space status, and standard hand gesture templates for each assembly process stage; Based on the digital twin model, the spatial state of the actual assembly unit is identified and the matching degree with the spatial state of the standard assembly unit is calculated to determine the current assembly process stage. Identify operator hand gestures in first-person image data and calculate the matching degree with standard hand gesture templates for the current assembly process stage to analyze the operator's intent; The spatial state of the assembly unit includes the position and orientation of the collaborative robotic arm, workpiece, and tools; S4. Based on the current assembly process stage and operating standards, predict the next assembly process, generate manual assembly action instructions, collaborative robotic arm action instructions, and assembly effect simulation data. The assembly effect simulation data includes robotic arm motion trajectory prediction, workpiece assembly three-dimensional morphology simulation, and prediction simulation animation of the next assembly effect. S5. Display the manual assembly action instructions in the AR glasses' field of view in the form of a 3D model, arrow indicators, and / or text prompts. At the same time, the collaborative robotic arm executes the collaborative robotic arm action instructions, and displays the current assembly status, assembly accuracy detection data, force control feedback parameters, and prediction simulation animation of the next assembly effect in real time through the visualization interface of the digital twin model.
2. The human-machine collaborative assembly method as described in claim 1, characterized in that, The multi-view image data includes images of the collaborative robotic arm, operator, and assembly workbench from different angles.
3. The human-machine collaborative assembly method as described in claim 1, characterized in that, In step S2, a deep convolutional neural network is used to extract features from multi-view image data, and a lightweight convolutional neural network is used to extract features from first-person image data.
4. The human-machine collaborative assembly method as described in claim 1, characterized in that, In step S2, the digital twin model includes a three-dimensional dynamic mapping of the collaborative robotic arm, workpiece, tools, and assembly workbench.
5. The human-machine collaborative assembly method as described in claim 1, characterized in that, In step S2, the scene assembly information includes the spatial coordinates of the end effector of the collaborative robotic arm, the three-dimensional pose matrix of the workpiece and the tool, and the assembly contact mechanical parameters. Data fusion employs state estimation algorithms to achieve spatiotemporal registration of multi-source data. The state estimation algorithms include Kalman filtering and particle filtering algorithms. The assembly contact mechanical parameters include at least one of pressure distribution and torque vector.
6. The human-machine collaborative assembly method as described in claim 5, characterized in that, In step S2, a digital twin model is constructed using a 3D modeling engine, including: Based on the spatial coordinates of the end effector of the collaborative robot arm, real-time pose data of each joint of the collaborative robot arm are generated through inverse kinematics calculation, and a kinematic simulation model of the collaborative robot arm is established. Based on the three-dimensional pose matrix of the workpiece and tool, the preset CAD model of the workpiece and tool is aligned with the real-time pose to generate a three-dimensional dynamic model of the assembly position of the workpiece and tool. By updating the contact state model between the tool and the workpiece using assembly contact mechanics parameters, a visual mapping of assembly contact forces can be achieved.
7. The human-machine collaborative assembly method as described in claim 1, characterized in that, In step S3, a geometric pose calculation algorithm is used to identify the position and posture of the collaborative robotic arm, workpiece and tool, and a hand key point detection algorithm is used to identify the operator's hand movements in the first-person image data. Among them, the geometric pose calculation algorithm includes the PnP algorithm or the ICP algorithm; The hand keypoint detection algorithm includes at least one of the following: a real-time detection model based on deep learning and a feature matching algorithm based on traditional image processing. The real-time detection model based on deep learning is selected from the MediaPipe hand detection framework, the OpenPose hand detection module, and the monocular 3D hand pose estimation model.
8. The human-machine collaborative assembly method as described in claim 1, characterized in that, In step S4, a pre-trained large model is used to predict the next assembly process. The large model is a sequence prediction model. Through the feature association mechanism, manual assembly action instructions and collaborative robotic arm action instructions are generated. The collaborative robotic arm action instructions include instruction data of posture constraints and trajectory planning parameters. The sequence prediction model is selected from the Transformer architecture or LSTM network, and the feature association mechanism includes attention mechanism or graph neural network.
9. The human-machine collaborative assembly method as described in claim 1, characterized in that, In step S4, the process of generating assembly effect simulation data includes: Based on the real-time pose data of the collaborative robotic arm in the digital twin model, the motion path of the end effector of the collaborative robotic arm in the next 5-10 seconds is predicted by the kinematic trajectory planning algorithm, and a trajectory coordinate sequence containing timestamps is generated. Based on the assembly tolerance requirements in the preset process standards, a physics engine is used to simulate the deformation of the assembled workpiece, generating a stress distribution cloud map of the workpiece contact surface and a three-dimensional model of the assembly gap. By fusing trajectory coordinate sequences with deformation simulation data, a lightweight rendering engine is used to generate real-time visualization simulation results of collaborative robotic arm motion trajectory prediction animation and workpiece assembly form.
10. A human-machine collaborative assembly system guided by digital twins and augmented reality, characterized in that, To implement the human-machine collaborative assembly method as described in any one of claims 1-9, comprising: The visual perception module includes AR glasses and multiple vision cameras set around the workstation, wherein the AR glasses are used to acquire first-person image data of the assembly workbench, and the vision cameras are used to acquire multi-view image data in real time. The mechanical equipment module includes a collaborative robotic arm control submodule and a collaborative robotic arm. The collaborative robotic arm is equipped with a sensing unit. The collaborative robotic arm control submodule is used to receive motion commands from the collaborative robotic arm and execute the motion commands through the collaborative robotic arm. The sensing unit is used to acquire sensing data from the collaborative robotic arm. The scene assembly information generation module is used to extract features from multi-view image data and first-person image data, and then fuse the extracted image features with the sensor data of the collaborative robotic arm to generate scene assembly information. The digital twin module includes a mapping submodule and a data visualization interface. The mapping submodule is used to build a digital twin model of the assembly workstation based on the scene assembly information. The data visualization interface is used to display the current assembly status, assembly accuracy detection data, force control feedback parameters and prediction simulation animation of the next assembly effect in real time. The identification and prediction module is used to configure the operation standards, standard assembly unit spatial states, and standard hand action templates for each assembly process stage. Based on the digital twin model, it identifies the actual assembly unit spatial state and calculates the matching degree with the standard assembly unit spatial state to determine the current assembly process stage; it identifies the operator's hand actions in the first-person image data and calculates the matching degree with the standard hand action template for the current assembly process stage to analyze the operation intention; and it combines the current assembly process stage and operation standards to predict the next assembly process, generating manual assembly action instructions, collaborative robotic arm action instructions, and assembly effect simulation data. The AR interactive display module is used to overlay manual assembly instructions in the field of view as 3D models, arrow indicators and / or text prompts.
Citation Information
Patent Citations
Man-machine collaborative assembly system based on augmented reality
CN116778119A
Method for remotely operating humanoid robot by identifying hand postures through virtual reality glasses
CN117746494A