AR assembly guidance generation method and system
By using multi-camera point cloud video acquisition and no-code interaction technology, combined with template point cloud registration and large language models, AR assembly guidance content is automatically generated, solving the problems of long update cycles and reliance on manual writing in traditional assembly guidance, and realizing efficient and automated assembly guidance generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-19
AI Technical Summary
Traditional assembly instructions rely on paper-based process manuals and two-dimensional drawings, which have long update cycles and limited information carrying capacity. They are difficult to meet the requirements of assembly efficiency and high reliability under rapidly changing production rhythms. Furthermore, the AR assembly instruction generation process has a high threshold, relies on manually written text and deep learning model training, and is difficult to adapt to multi-variety, small-batch production.
By acquiring point cloud video from multiple cameras, spatial teaching in a no-code interactive environment, and template point cloud registration, combined with hand trajectory analysis and large language models, AR assembly guidance content is automatically generated, reducing the need for manual measurement and programming, and enabling automatic determination of the spatial posture of the assembly object and automatic text generation.
It significantly reduces the barrier to entry and time cost of generating AR assembly instructions, improves the efficiency and consistency of assembly instruction generation, and supports rapidly changing production environments.
Smart Images

Figure CN122066902A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the fields of augmented reality and intelligent manufacturing technology, and in particular to an AR assembly guidance generation method and system. Background Technology
[0002] As the manufacturing industry moves towards flexibility, personalization, and high-variety, low-batch production, the assembly objects, assembly sequences, and process parameters on the production floor are undergoing frequent changes. Traditional assembly guidance mainly relies on paper-based worksheets and two-dimensional drawings, which have long update cycles, limited information capacity, and heavy reliance on operators' spatial imagination and experience. This makes it difficult to simultaneously meet the requirements of assembly efficiency and high reliability under rapidly changing production rhythms. Against this backdrop, augmented reality (AR) technology, which can overlay virtual 3D models, operational prompts, and process information onto real-world work scenarios, providing workers with intuitive, immediate, and visual assembly guidance, has become a crucial technological direction in the field of intelligent manufacturing for improving assembly training efficiency and operational consistency.
[0003] How to automatically generate content for AR assembly guide programs is a technical problem that urgently needs to be solved. Summary of the Invention
[0004] This application provides an AR assembly guide generation method and system, which can automatically generate the creation content of AR assembly guide programs.
[0005] This invention provides an AR assembly guidance generation method based on a no-code interactive environment, including: Obtain 3D data of the actual assembly process; Based on the obtained point cloud video sequence, spatial teaching is performed on the assembly object and spatial constraint data corresponding to the assembly steps are generated. Determine the spatial orientation of the assembly object based on spatial constraint data; Augmented reality (AR) assembly instructions are generated based on the determined spatial pose of the assembly object. The generated AR assembly guidance content is integrated with the spatial pose information of the assembly object to obtain AR assembly guidance data that can be loaded and executed by the augmented reality system.
[0006] In one exemplary instance, after determining the spatial pose of the assembly object and before generating the augmented reality (AR) assembly guidance content, the method further includes: The assembly direction of the assembly object is identified based on the hand movement trajectory during the assembly process.
[0007] In one exemplary instance, identifying the assembly direction of the assembly object includes: The key points of the operator's hand are detected in the color image corresponding to the point cloud video, and the detected key points of the hand are mapped to a unified three-dimensional coordinate system; The hand movement trajectory is calculated based on the changes in key hand points over time; Based on the generated spatial constraint data, effective trajectory segments related to the assembly action are extracted from the hand movement trajectory. Based on the spatial relationship between the effective trajectory segment and the assembly target position, the assembly direction of the assembly object is inferred.
[0008] In one exemplary instance, it also includes: When it is necessary to automatically determine whether the assembly steps are completed during the AR assembly guidance process, and to switch or prompt the steps according to the operation status, the assembly action is detected based on the spatial constraint data, and the progress of the assembly steps is controlled. The process of detecting assembly actions based on the spatial constraint data and controlling the progression of assembly steps includes: Based on the generated spatial constraint data, the spatial determination area of the current assembly action is determined; During the AR assembly guidance process, the spatial relationship between the operator's hand position or the position of the assembly object and the spatial judgment area is detected in real time to determine whether the current assembly action meets the preset completion conditions. When the assembly object is detected to have entered the spatial determination area and met the corresponding posture or position requirements, the current assembly step is determined to be completed, and the assembly guidance is triggered to enter the next assembly step, or the corresponding completion prompt or error correction prompt information is output to the operator.
[0009] In one exemplary instance, after generating the augmented reality (AR) assembly guidance content and before obtaining the AR assembly guidance data that can be loaded and executed by the augmented reality system, the method further includes: When it is necessary to reduce the amount of manually written AR assembly guidance text, or to maintain consistency in text expression across different assembly steps, AR assembly guidance text can be generated based on a post-trained large language model.
[0010] 6. The AR assembly guidance generation method according to claim 5, wherein the generation of AR assembly guidance text based on a post-trained large language model includes: During the generation of the AR assembly guidance content, the spatial pose information of the assembly object corresponding to the assembly step, the assembly step number, and the assembly direction information are obtained. The spatial attitude information, assembly step number, and assembly direction information are structured to form text generation input data that conforms to the preset input format. The text generation input data is input into a large language model trained with assembly domain data, and the post-trained large language model generates the AR assembly guidance text corresponding to the current assembly step. The generated AR assembly guidance text is associated with the corresponding assembly steps and encapsulated together with the AR assembly guidance content during the AR assembly guidance data output process for display to the operator in the augmented reality environment.
[0011] In one exemplary instance, the step of spatially teaching the assembly object and generating spatial constraint data corresponding to the assembly steps based on the obtained point cloud video sequence includes: In point cloud video, an initial 3D bounding box is created before the assembled object begins to move, and a final 3D bounding box is created after the assembly is completed and the object is stably in place, using AR / VR interaction. Extract bounding box parameters, including center position, spatial size, and spatial orientation, from the starting and ending 3D bounding boxes respectively; The bounding box parameters are associated with the corresponding assembly step numbers to form the spatial constraint data used to describe the range of spatial position changes of the assembled object.
[0012] In one exemplary instance, determining the spatial orientation of the assembly object based on spatial constraint data includes: Based on the CAD model of the assembly object, its surface is sampled to generate a template point cloud of the assembly object; Based on the generated spatial constraint data, real point cloud data within the spatial constraint range is extracted from the point cloud video and used as candidate point clouds for the assembly object; The captured real point cloud is preprocessed to remove noise points and background points; The preprocessed real point cloud is matched with the template point cloud, and the spatial transformation relationship between the two is calculated; Based on the spatial transformation relationship, the spatial attitude information of the assembly object in the unified world coordinate system is determined; wherein, the spatial attitude information includes the spatial position and spatial orientation of the assembly object, which is used to describe the actual placement state of the assembly object in three-dimensional space.
[0013] In one exemplary instance, generating AR assembly guidance content includes: Based on the determined spatial pose information of the assembly object, the display position and display pose of the assembly object in the augmented reality coordinate system are determined; The virtual 3D model of the assembly object is loaded into the augmented reality environment and placed according to the determined spatial posture; The spatial pose information of the assembly object is associated with the corresponding assembly steps to generate the AR assembly guidance content corresponding to the assembly steps.
[0014] In one exemplary instance, integrating the generated AR assembly guidance content with the spatial pose information of the assembly object includes: The spatial posture information of the assembly object, the assembly step number, and the corresponding AR assembly guidance content are integrated; The integrated information is encapsulated according to a predetermined data structure to obtain information in a unified data format. The packaged AR assembly guidance data is output to the augmented reality system for displaying the AR assembly guidance content in the augmented reality device.
[0015] This application also provides a computer-readable storage medium storing computer-executable instructions for performing the AR assembly guidance generation method described in any of the above claims.
[0016] This application embodiment further provides an AR assembly guidance generation system, including: a point cloud video acquisition module, a no-code space registration module, a template point cloud registration module, a first generation module, and a second generation module; wherein, The point cloud video acquisition module is configured to acquire 3D data of the actual assembly process; The no-code spatial registration module is configured to perform spatial teaching on the assembly object and generate spatial constraint data corresponding to the assembly steps based on the obtained point cloud video sequence. The template point cloud registration module is configured to determine the spatial pose of the assembly object based on spatial constraint data. The first generation module is set to generate AR assembly guidance content based on the determined spatial pose of the assembly object; The second generation module is configured to integrate the generated AR assembly guidance content with the spatial pose information of the assembly object to obtain AR assembly guidance data that can be loaded and executed by the augmented reality system.
[0017] In one exemplary instance, it also includes: a third generation module configured to generate AR assembly guidance text based on a post-trained large language model when it is necessary to reduce the need for manually writing AR assembly guidance text or to maintain consistency in textual expression across different assembly steps.
[0018] This application also provides an AR assembly guidance generation system, including: The expert assembly demonstration recording module is used to collect point cloud video data when experts perform the actual assembly process; The 3D space registration module for parts is used to spatially teach assembly objects in a no-code interactive environment and generate spatial constraint data corresponding to the assembly steps. The template point cloud registration module is used to register the template point cloud of the assembly object with the real point cloud based on spatial constraint data, so as to determine the spatial attitude information of the assembly object in a unified coordinate system. The no-code assembly instruction creation module is used to generate augmented reality assembly guidance content corresponding to the assembly steps based on the determined spatial posture information. The augmented assembly program generation module integrates the generated assembly guidance content, spatial attitude information, and assembly step information to generate an assembly guidance program that can be loaded and executed by the augmented reality system.
[0019] In one exemplary instance, an instruction text generation module is also included, for: During the generation of the AR assembly guidance content, the spatial pose information of the assembly object corresponding to the assembly step, the assembly step number, and the assembly direction information are obtained. The spatial attitude information, assembly step number, and assembly direction information are structured to form text generation input data that conforms to the preset input format. The text generation input data is input into a large language model trained with assembly domain data, and the post-trained large language model generates the AR assembly guidance text corresponding to the current assembly step. The generated AR assembly guidance text is associated with the corresponding assembly steps and encapsulated together with the AR assembly guidance content during the AR assembly guidance data output process for display to the operator in the augmented reality environment.
[0020] In one exemplary instance, an assembly direction recognition module is also included, which is used to: recognize the assembly direction of the assembly object based on the hand movement trajectory of the operator during the assembly process.
[0021] In one exemplary instance, an assembly process detection and step control module is also included, which is used to: detect whether the assembly action is completed based on the spatial constraint data during the execution of augmented reality assembly guidance, and control the automatic switching or prompt output of assembly steps.
[0022] The AR assembly guidance generation method provided in this application uses 3D data from a real assembly process as the basis for generating assembly guidance. It generates structured spatial constraint data through spatial teaching, transforming human experience into computable geometric information. Under spatial constraints, it automatically determines the spatial posture of the assembly object, avoiding manual measurement and calibration. It automatically generates assembly guidance content that can be used by augmented reality systems. This application significantly reduces the reliance on manual modeling and program development in the AR assembly guidance generation process, improving the efficiency and consistency of assembly guidance generation.
[0023] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description, claims, and drawings. Attached Figure Description
[0024] The accompanying drawings are used to provide a further understanding of the technical solutions of this application and constitute a part of the specification. They are used together with the embodiments of this application to explain the technical solutions of this application and do not constitute a limitation on the technical solutions of this application.
[0025] Figure 1 This is a flowchart illustrating the AR assembly guidance generation method in an embodiment of this application; Figure 2 This is an overview diagram of the AR assembly guidance generation system in the embodiments of this application; Figure 3 This is a schematic diagram of the overall architecture of the AR assembly guidance generation system and the registration of the engineering and user ends in this embodiment of the application; Figure 4 This is a schematic diagram of the overall architecture of the AR assembly guidance creation system in the embodiments of this application; Figure 5 This is a schematic diagram of point cloud acquisition and spatial calibration of registered objects using multiple RGBD cameras in an embodiment of this application. Figure 6 This is a schematic diagram illustrating the process of the spatial registration object annotation method based on no-code interaction in the embodiments of this application; Figure 7 This is a schematic diagram of the RANSAC+ICP registration process between the template point cloud and the real point cloud in the embodiments of this application. Figure 8 This is a schematic diagram of the assembly direction recognition based on hand trajectory in an embodiment of this application; Figure 9 This is a schematic diagram of the process of generating assembly guidance text with the assistance of a large language model in an embodiment of this application; Figure 10This is a schematic diagram illustrating the generation of an AR assembly guide based on a standardized data structure and its display in an AR headset, as described in this application embodiment. Figure 11 This is a schematic diagram of the composition structure of the AR assembly guidance generation system in the embodiments of this application. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be arbitrarily combined with each other.
[0027] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings, which illustrate embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of this application will be thorough and complete.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0029] It is understood that the terms "first" and "second" used in this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0030] It is understood that the term "connection" in the following embodiments should be understood as "electrical connection," "communication connection," etc., if the connected circuits, modules, units, etc., have electrical signal or data transmission with each other.
[0031] When used herein, the singular forms of “a,” “an,” and “the” may also include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising / including” or “having,” etc., specify the presence of the stated features, wholes, steps, operations, components, parts, or combinations thereof, but do not preclude the possibility of the presence or addition of one or more other features, wholes, steps, operations, components, parts, or combinations thereof. Meanwhile, the term “and / or” as used in this specification includes any and all combinations of the associated listed items.
[0032] The steps illustrated in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases the steps shown or described may be performed in a different order than that presented here.
[0033] AR assembly guidance systems typically assist assembly by pre-creating assembly content and displaying and interacting with it on AR devices. Content creation generally includes: importing the product's CAD model into an AR development environment (such as Unity), establishing assembly step logic and interaction flow for each step, completing spatial registration between the virtual model and the real scene (to determine the position and posture of the 3D model in the real environment), and configuring part assembly animation paths, text descriptions, and step switching logic, thereby providing step-by-step guidance on a head-mounted display or mobile terminal. To achieve spatial registration, common approaches include: one relies on manual measurement or alignment, where engineers calibrate the part installation position and posture parameters on-site to complete model placement; another uses vision / depth-based recognition or deep learning solutions, training the model to recognize parts or features and estimate their pose, thus achieving automatic or semi-automatic spatial alignment. Regarding the acquisition of teaching information, some solutions infer assembly paths and directions by recording teaching videos or operation process data, and generate insertion animations or guide arrows accordingly; for text descriptions, most still rely on engineers manually writing and maintaining the text for each step based on process documents and experience.
[0034] However, the above-mentioned implementation methods generally suffer from the following problems when facing assembly sites with multiple product types and high-frequency changes: First, AR content creation has a high barrier to entry, often requiring deep involvement from programmers. Engineers find it difficult to independently complete the configuration of assembly step logic, model binding, animation paths, and interaction logic, resulting in long development cycles, high costs, and difficulty in adapting to the rapid iteration of production sites. Second, the spatial registration process either relies on manual measurement and alignment, which is labor-intensive and prone to cumulative errors; or it relies on deep learning model training, which requires a large amount of labeled data and training processes, resulting in limited generalization ability, complex deployment, and high costs for migration across parts / workstations, hindering the rapid implementation of multi-product assembly scenarios. Furthermore, during the assembly process with manual teaching, parts are often obscured by the operator's hands, making it difficult for the system to directly extract the complete motion direction and installation path from the part trajectory. This makes it difficult to automatically obtain key information such as assembly direction and insertion animation, requiring manual supplementation and correction. Finally, assembly instruction texts usually rely on engineers to write and maintain them manually, which is not only time-consuming, but also makes it difficult to ensure the consistency of terminology and expression standards. It is also difficult to quickly iterate and update when faced with process changes, which in turn affects the reusability and large-scale promotion of AR instruction content.
[0035] To address at least one of the above problems, this application provides an AR assembly guidance generation method. Based on point cloud video collected by multiple cameras, it can complete spatial teaching of the assembly process through code-free interaction, determine the precise posture of the parts by combining a template point cloud registration algorithm, identify the assembly direction by using a hand trajectory analysis algorithm, and then automatically generate assembly text by a large language model. Finally, it generates an AR assembly guidance program that can be directly deployed, thereby significantly reducing the threshold and time cost of AR program production.
[0036] Figure 1 This is a flowchart illustrating the AR assembly guidance generation method in this application embodiment, based on a no-code interactive environment, such as... Figure 1 As shown, it may include: Step 100: Obtain 3D data of the actual assembly process.
[0037] In one exemplary instance, the acquisition of 3D data during the assembly process can be achieved using multi-camera acquisition to improve the integrity and spatial coverage of the point cloud data.
[0038] In one exemplary instance, multiple RGBD cameras can be used to simultaneously capture point cloud videos of the assembly process. For example, a unified coordinate system can be established using the spatial registration objects of AprilTags, and the point clouds from multiple cameras can be fused into a continuous three-dimensional point cloud video sequence to provide basic data for subsequent spatial analysis.
[0039] In one embodiment, step 100 may include: Multiple RGBD cameras (such as Azure Kinect) are deployed around the assembly area, and a unified world coordinate system is established through spatial calibration objects to achieve high-precision point cloud fusion; The depth data collected by each camera is converted into a 3D point cloud; The coordinate transformation and fusion of point clouds from different cameras are performed, and a continuous 3D point cloud video sequence is generated in chronological order.
[0040] Step 100 provides a unified 3D spatial data foundation for subsequent spatial teaching, point cloud registration, and motion analysis, ensuring that all spatial information in different steps is in the same coordinate system. Step 100 obtains complete, continuous, and computable 3D assembly process data, reducing occlusion and information loss problems caused by single-view acquisition.
[0041] Step 101: Based on the obtained point cloud video sequence, perform spatial teaching on the assembly object and generate spatial constraint data corresponding to the assembly steps.
[0042] In one exemplary instance, the starting and ending positions of a part can be marked in a point cloud video using a 3D bounding box through a VR / AR interface, achieving programming-free spatial teaching; the center position, scale, and rotation of the bounding box are recorded to construct spatial constraints for the assembly steps. In this embodiment, spatial constraints refer to geometric constraint information composed of the 3D spatial parameters corresponding to the assembly object in the assembly start state and assembly completion state. The geometric constraint information includes at least spatial position, spatial range, and spatial orientation, used to limit the spatial search range and target area in the subsequent assembly analysis and AR generation process.
[0043] In one embodiment, step 101 may include: In point cloud videos, engineers can create a starting 3D bounding box before the assembly object begins to move and a ending 3D bounding box after the assembly is completed and the object is stably in place, using AR or VR interaction. Extract bounding box parameters, including center position, spatial size, and spatial orientation, from the starting and ending 3D bounding boxes respectively; The bounding box parameters are associated with the corresponding assembly step numbers to form spatial constraint data that describes the range of spatial position changes of the assembled object.
[0044] The embodiments of this application enable spatial teaching without writing programs or manually inputting three-dimensional coordinates. This significantly reduces the manual threshold for generating assembly instructions and improves the stability and computational efficiency of subsequent algorithms.
[0045] In one exemplary instance, the spatial teaching process in step 101 can be completed through a no-code AR / VR interaction method, thereby lowering the barrier to entry for engineers when conducting spatial teaching.
[0046] In one embodiment, the no-code interactive environment is an AR or VR-based interactive environment. Operators can annotate and teach the spatial position and posture of the assembly object in point cloud video through graphical interactions such as dragging, selecting, box selection, and posture adjustment, without writing program code, to generate spatial constraint data corresponding to the assembly steps. In other words, the no-code interactive environment is used for creating assembly guidance content. It receives the user's definitions of assembly steps, the spatial range of the assembly object, and its posture through graphical operations, and automatically converts the user's interactive operations into structured instruction data for generating AR assembly guidance, without requiring the user to write program code.
[0047] Step 102: Determine the spatial orientation of the assembly object based on spatial constraint data.
[0048] In one exemplary instance, a corresponding template point cloud can be generated based on the CAD model of the assembly object, and the generated template point cloud can be matched with the real point cloud extracted from the point cloud video. The spatial pose of the assembly object in a unified coordinate system can be calculated by point cloud registration.
[0049] In one embodiment, step 102 may include: Based on the CAD model of the assembly object, its surface is sampled to generate a template point cloud of the assembly object; Based on the generated spatial constraint data, real point cloud data within the spatial constraint range is extracted from the point cloud video and used as candidate point clouds for the assembly objects; The captured real point cloud is preprocessed to remove noise points and background points; The preprocessed real point cloud is matched with the template point cloud, and the spatial transformation relationship between the two is calculated; Based on the spatial transformation relationship, the spatial attitude information of the assembly object in the unified world coordinate system is determined; wherein, the spatial attitude information includes at least the spatial position and spatial orientation of the assembly object, which is used to describe the actual placement state of the assembly object in three-dimensional space.
[0050] Step 102 relies on the spatial constraint data generated in step 101. By limiting the extraction range of the real point cloud through spatial constraints, the search space in the point cloud registration process is reduced, thereby improving the stability and computational efficiency of the pose determination process. Through step 102, without the need for manual measurement of the assembly object's pose, the spatial pose information of the assembly object in the real assembly scene is automatically obtained within the spatial range defined by the generated spatial constraint data. This provides an accurate and reliable data foundation for the accurate generation of subsequent AR assembly guidance content.
[0051] Step 103: Generate AR assembly guidance content based on the determined spatial pose of the assembly object.
[0052] This step generates assembly guidance content for augmented reality display based on the determined spatial pose information of the assembly object, realizing the transformation from the real assembly process to AR assembly guidance content.
[0053] In one exemplary instance, the corresponding virtual 3D model can be placed in the correct position in the augmented reality environment based on the spatial pose information of the assembly object, and assembly guidance information corresponding to the assembly steps can be generated.
[0054] In one embodiment, step 103 may include: Based on the determined spatial pose information of the assembly object, determine the display position and display pose of the assembly object in the augmented reality coordinate system; The virtual 3D model of the assembly object is loaded into the augmented reality environment and placed according to the determined spatial posture; The spatial pose information of the assembly object is associated with the corresponding assembly steps to generate AR assembly guidance content corresponding to the assembly steps; wherein, the AR assembly guidance content may include assembly guidance information for augmented reality display, which is used to indicate the spatial relationship and assembly sequence of the assembly object in the assembly process.
[0055] Step 103 comprehensively utilizes the obtained 3D assembly process data, generated spatial constraint data, and determined spatial posture information of the assembly object to achieve automatic generation of digital AR assembly guidance content from real assembly teaching. Step 103 avoids manually setting the position and posture of the assembly object in the AR development environment, thereby reducing the workload and technical threshold for creating AR assembly guidance content.
[0056] Step 104: Integrate the generated AR assembly guidance content with the spatial pose information of the assembly object to obtain AR assembly guidance data that can be loaded and executed by the augmented reality system.
[0057] In one embodiment, step 104 may include: Integrate the spatial posture information of the assembly object, the assembly step number, and the corresponding AR assembly guidance content; The integrated information is encapsulated according to a predetermined data structure to obtain information in a unified data format. The packaged AR assembly guidance data is output to the augmented reality system for displaying the AR assembly guidance content in the augmented reality device.
[0058] Step 104 completes the entire process from acquiring 3D data of the actual assembly process, spatial teaching, spatial attitude determination, to generating and outputting AR assembly guidance content, realizing the automatic generation and direct deployment of AR assembly guidance content.
[0059] The AR assembly guidance generation method provided in this application uses 3D data from a real assembly process as the basis for generating assembly guidance. It generates structured spatial constraint data through spatial teaching, transforming human experience into computable geometric information. Under spatial constraints, it automatically determines the spatial posture of the assembly object, avoiding manual measurement and calibration. It automatically generates AR assembly guidance content that can be used by augmented reality systems. This application significantly reduces the reliance on manual modeling and program development in the AR assembly guidance generation process, improving the efficiency and consistency of assembly guidance generation.
[0060] In one exemplary instance, after step 102 and before step 103, when the assembly object is obscured by the operator's hand during the teaching process, or when the motion direction cannot be reliably extracted directly from the point cloud of the assembly object, this embodiment of the application may further include: Step 1023: Identify the assembly direction of the assembly object based on the hand movement trajectory during the assembly process.
[0061] In one exemplary instance, step 1023 may include: Detect key points of the operator's hand in the color image corresponding to the point cloud video, and map the detected key points of the hand to a unified three-dimensional coordinate system; The hand movement trajectory is calculated based on the changes in key hand points over time. Combining the spatial constraint data generated in step 101, effective trajectory segments related to assembly actions are extracted from the hand movement trajectories. Based on the spatial relationship between the effective trajectory segment and the assembly target position, the assembly direction of the assembly object is inferred.
[0062] In this embodiment, the assembly direction information is independent of the spatial posture information of the assembly object and is used to describe the movement direction of the assembly object during the assembly process. In subsequent step 103, the assembly direction information serves as an optional supplementary input for generating assembly guidance content, enhancing the guidance effect of augmented reality assembly guidance.
[0063] This embodiment introduces assembly direction recognition based on hand trajectory. Even when the assembly object is obscured by a hand and the movement direction cannot be directly extracted from the point cloud of the assembly object, assembly direction information can still be reliably obtained, thereby improving the robustness and applicability of assembly direction recognition. Moreover, the obtained assembly direction information can be used as supplementary input for generating assembly guidance content to enhance the guidance effect of AR assembly guidance.
[0064] In one exemplary instance, after step 103 or step 104, when it is necessary to automatically determine whether the assembly steps are completed during the AR assembly guidance process, and to switch steps or provide prompts based on the operation status, the process may further include: Step 105: Detect assembly actions based on spatial constraint data and control the progress of assembly steps.
[0065] In one embodiment, step 105 may include: Based on the spatial constraint data generated in step 101, determine the spatial judgment area of the current assembly action; During the AR assembly guidance process, the spatial relationship between the operator's hand position or the position of the assembly object and the spatial judgment area is detected in real time to determine whether the current assembly action meets the preset completion conditions. When the assembly object is detected to have entered the spatial determination area and met the corresponding posture or position requirements, the current assembly step is determined to be completed, and the assembly guidance is triggered to enter the next assembly step, or the corresponding completion prompt or error correction prompt information is output to the operator.
[0066] By introducing an assembly action detection and step control implementation method, the automatic judgment and switching of assembly steps are realized without relying on manual confirmation, which improves the continuity and operational consistency of the AR assembly guidance process and reduces the impact of human operation errors on the assembly process.
[0067] In one exemplary instance, after step 103 and before step 104, when it is necessary to reduce the manual writing of AR assembly guidance text, or to maintain consistency in text expression between different assembly steps, the following may also be included: Step 1034: Generate AR assembly guidance text based on the post-trained large language model.
[0068] In one embodiment, step 1034 may include: During the generation of AR assembly guidance content, the spatial pose information of the assembly object corresponding to the assembly step, the assembly step number, and the assembly direction information are obtained. The spatial attitude information, assembly step number, and assembly direction information are structured to form text generation input data that conforms to the preset input format. The text generation input data is fed into a large language model trained with assembly domain data, and the post-trained large language model generates AR assembly guidance text corresponding to the current assembly step. The generated AR assembly guidance text is associated with the corresponding assembly steps and encapsulated together with the AR assembly guidance content during the subsequent AR assembly guidance data output process, for display to the operator in the augmented reality environment.
[0069] In this embodiment, by introducing a large language model trained with assembly domain data, automatic generation of assembly guidance text based on spatial pose information of assembly steps is achieved. Compared with manual writing or generating assembly guidance text based on fixed rules, this implementation significantly reduces the cost of manually writing AR assembly guidance text and ensures consistency in text description style and terminology usage between different assembly steps, thereby improving the overall professionalism, maintainability, and scalability of AR assembly guidance.
[0070] The AR assembly guidance generation method provided in this application automatically generates a directly deployable AR assembly guidance program by acquiring and analyzing multi-camera point cloud video of a real assembly process, combined with spatial teaching, 3D registration, assembly direction recognition, and assembly semantic generation. One embodiment can be found... Figure 2 , Figure 2 This application demonstrates that the AR assembly guidance generation scheme of this embodiment forms a closed-loop process from assembly demonstration data acquisition, spatial teaching, posture determination, assembly direction recognition, text generation to AR program encapsulation and output.
[0071] In some embodiments, the augmented reality assembly instruction creation method based on point cloud video analysis in this application solves the problems of long production cycles, high professional thresholds, difficulties in spatial registration, difficulty in automatically structuring teaching information, and complete reliance on manual labor for augmented reality program production. This application embodiment achieves automated conversion from real assembly teaching to AR assembly programs by integrating technologies such as multi-camera point cloud acquisition, code-free spatial registration, template point cloud registration, hand trajectory analysis, and large language model text generation.
[0072] The following is a detailed description with specific examples. For example... Figure 3 As shown, the AR assembly guidance generation system provided in this application includes an engineering-side assembly demonstration environment and a user-side assembly guidance environment. The engineering side uses a depth camera to capture the actual assembly process and completes spatial registration of the assembly demonstration environment; the user side uses an augmented reality device to complete registration of the assembly guidance environment, thereby achieving consistent presentation of assembly guidance content across different terminals. Figure 4 As shown, the AR assembly guidance creation system on the engineering side includes a point cloud video acquisition subsystem, a point cloud playback and timeline control subsystem, a no-code assembly step creation subsystem, a spatial constraint data management subsystem, and an assembly instruction data export subsystem, which are used to transform the expert assembly demonstration process into reusable assembly guidance data.
[0073] First, point cloud video capture of the assembly process.
[0074] Combination Figure 5 As shown, in this embodiment of the application, multiple RGBD cameras (such as Azure Kinect) are arranged around the assembly area, and a unified coordinate system is established by spatial registration objects such as AprilTag to achieve high-precision point cloud fusion.
[0075] 1) Multi-camera calibration and establishment of a unified coordinate system.
[0076] Several AprilTag markers are placed in the scene and observed by multiple RGBD cameras. For any camera... The system first detects the AprilTag plane it captures and calculates the transformation matrix from the camera coordinate system to the AprilTag coordinate system: ; Set the first camera coordinate system as the world coordinate system: Then the transformation matrix from all cameras to the world coordinate system is: .
[0077] 2) Multi-view point cloud fusion.
[0078] For cameras captured raw point cloud points Its transformation to the world coordinate system is as follows: ; Multiple camera point clouds are merged to obtain the complete point cloud for each frame: ; The final point cloud video sequence is formed: .
[0079] 3) The pseudocode for point cloud fusion is shown below: Next, register without code space.
[0080] Combination Figure 6 As shown, the code-free spatial registration in this embodiment means that engineers do not need to write any program code or manually input three-dimensional numerical coordinates. Instead, they can provide spatial prompts for the assembly parts in the form of a three-dimensional visual bounding box in the point cloud video. The system automatically records the spatial range and reference posture of the parts in the initial state and the completed installation state, providing constraints for subsequent template point cloud registration and AR assembly animation generation.
[0081] 1) Bounding box parameter structure. In the unified world coordinate system after multi-camera fusion, the point cloud in each frame of the point cloud video... All are in the same coordinate system Below. When engineers view point cloud videos in a VR / AR environment, they are essentially... The parts are selected by "box selection".
[0082] The system treats each 3D bounding box as an axis-aligned or rotated cuboid, whose geometry is uniquely determined by the following three factors: center position. ,size , attitude (quaternion) .
[0083] By defining a bounding box for the initial and final states of a part, the system provides the geometric constraint range for the part to complete one assembly operation in space. Subsequent algorithms only need to find the correspondence between the point cloud and the CAD model within this range, which greatly reduces the search space and improves the robustness and efficiency of registration.
[0084] 2) No code space annotation. Starting frame selection includes: The engineer loads a point cloud video in VR / AR, uses the timeline or playback controls to find the frame before a part is ready to begin moving, and uses this frame as the starting frame for the current assembly step of that part. Starting bounding box creation includes: The engineer uses a predefined gesture (such as raising a palm) to bring up the "Bounding Box Tool." The system generates a 3D bounding box of default size in the current view and displays it as a dashed box above the point cloud. Interactive editing of the bounding box includes: The engineer uses gestures such as pinching, dragging, and rotating to change the position, size, and orientation of the bounding box, making it as completely enclosing as possible the point cloud of the target part. Once the engineer confirms that the bounding box correctly covers the current part, they perform a "Confirm" gesture or press a virtual button. The system records the parameters of the bounding box, as well as the current frame number and step number. Terminating frame selection and termination bounding box creation include: The engineer continues playing the point cloud video, finds the frame where the part is fully assembled and stably positioned, uses this frame as the termination frame, repeats the above process of creating and editing the bounding box, adjusts the bounding box to just cover the position of the assembled part, and saves it. Multi-step repetitive teaching involves repeating the above operations for each part and each assembly step in the assembly process to build a complete "spatial teaching scheme." All bounding box information will serve as input for subsequent template point cloud registration and AR animation generation.
[0085] 3) The pseudocode without code space annotation is shown below: Then, the template point cloud registration (RANSAC + ICP) algorithm is used.
[0086] Combination Figure 7 As shown, the actual point cloud is automatically aligned with the template point cloud of the CAD model to obtain the part's pose matrix. The task of the template point cloud registration module is to calculate a rigid body transformation matrix based on the known point cloud of the CAD model and the actual acquired point cloud of the part, so that the two are optimally aligned in a unified coordinate system, thereby obtaining the part's spatial pose (position and planar orientation), which is then used for display and animation in augmented reality-assisted assembly programs. Due to the presence of noise and outliers (such as other parts or background points) in the actual point cloud, and the unknown initial alignment, a two-stage strategy can be adopted in this embodiment: In the first stage, a reliable initial transformation is obtained using the RANSAC algorithm with random sampling and robustness assessment. ; In the second phase, Based on this, the ICP (Iterative Closest Point) algorithm is used for local fine-tuning to obtain the final transformation. .
[0087] 1) Template Point Cloud Generation. Assuming the deformation of the part during assembly is negligible (i.e., the CAD model and the actual part are geometrically identical or have negligible errors), only rigid body transformations (rotation and translation) exist between them, without scaling or non-rigid deformation. Uniform sampling is performed on the CAD mesh model to obtain the template point cloud set: .
[0088] 2) Real point cloud extraction. Extract the real point cloud from the bounding box of a specified frame: Preprocessing of real point clouds includes: statistical filtering for noise reduction; RANSAC plane culling (removing desktops).
[0089] 3) Point cloud feature calculation. and Calculate the FPFH features separately: .
[0090] 4) Coarse registration (RANSAC algorithm). In this embodiment, the RANSAC coarse registration process is as follows: Constructing a candidate matching set... The initial transformation is solved through multiple rounds of RANSAC iterations: .
[0091] 5) Fine registration (ICP). Optimize the error function: Update rules: Until convergence: The final part orientation is obtained: .
[0092] The pseudocode for the template point cloud registration algorithm is shown below: Next, hand trajectory extraction and assembly direction recognition are performed. Combined with... Figure 8 As shown, this process addresses the core challenge of "parts being obscured by hands, making it impossible to directly extract the motion direction from the part's point cloud" during real-world assembly teaching. This invention infers the assembly direction based on the dynamic trajectory of key hand points, enabling the system to accurately identify the assembly direction even when the complete trajectory of the part is missing. It is applicable to various assembly actions such as insertion and sliding.
[0093] 1) Hand key point detection principle In this embodiment, the MediaPipe Hands hand keypoint detection algorithm is used to detect 21 three-dimensional keypoints of the hand in the RGB image corresponding to each frame of point cloud video, denoted as: ; among them, each key point All points are converted from camera depth maps to 3D points in the world coordinate system. Under a unified coordinate system, these key points can stably represent hand posture, grasping form, and positional changes.
[0094] 2) Hand center trajectory calculation. The hand center point is defined as the average position of 21 key points: The trajectory sequence is as follows: Due to jitter in actual data acquisition, a sliding window smoothing technique is applied to the trajectory to ensure the stability of the direction calculation: ,in, This is the window length.
[0095] 3) Valid trajectory segment selection. Since the hand moves before "picking up the part" and after "completing assembly", this embodiment utilizes the center point of the termination bounding box to avoid interference from irrelevant movements. The trajectory is filtered, including: calculating the distance to each smooth trajectory point. Select trajectory points that meet the following conditions as valid assembly paths: ,in, An empirical threshold (e.g., 10 to 15 cm) is used to ensure that the selected parts are the hand movements during the process of "transferring the parts to the target location".
[0096] 4) Assembly direction vector calculation principle. The first point of the effective trajectory segment is regarded as the assembly starting point: The target point is the center of the bounding box's termination position. The assembly direction vector is defined as: After normalization, the main installation direction is obtained: This direction is used for pointing AR guide arrows, calculating interpolation paths for part animations, and generating assembly description text.
[0097] After that, such as Figure 9 As shown, this is a data-driven text generation method based on a large language model (LLM). This step aims to address the problems of traditional AR assembly instruction texts, which rely on manual writing, are time-consuming, have inconsistent expressions, and are difficult to adapt to changing assembly processes. This invention proposes a data-driven text generation method that combines a large language model (LLM) with assembly semantic knowledge. This method can automatically generate text instructions that conform to industrial assembly standards based on information such as the assembly object sequence, part posture, and assembly direction obtained in the preceding steps.
[0098] 1) Overview of Method Principles. In this embodiment, the generation of assembly instruction text relies on a large language model that has undergone domain data augmentation and fine-tuning. The core idea is: taking structured data such as assembly object sequences, part motion relationships, bounding box spatial information, and assembly direction vectors as input; the fine-tuned large language model learns a large number of assembly action description templates; the large language model can convert structured information into natural language expressions; and finally, it outputs JSON assembly text with a fixed format, which is convenient for direct use by the AR system. This method automates the manual process, significantly reducing instruction compilation time.
[0099] 2) Dataset Construction for Assembly Tasks. To enable the model to understand assembly actions, this invention constructs a dataset specifically for generating assembly instructions. This dataset contains approximately 1000 organized assembly Q&A entries. Data sources include: mechanical assembly textbooks (standard expressions, terminology specifications); process guidance documents for specific assembly scenarios (such as process cards and work instructions); and task sequences derived from point cloud and CAD trajectory inversion (from the automatic analysis results of previous steps in this invention). In one embodiment, the dataset is organized in an Alpaca style and contains two core types of data: professional expression enhancement data: providing realistic assembly scenario descriptions, action targets, and precautions to improve the professionalism of the model-generated text; and assembly relationship parsing data: given "part chain relationships + motion direction + target constraints," training the model to correctly understand the logical relationship of "who assembles to whom."
[0100] 3) LLM Fine-tuning Process and Parameters. In this embodiment, the model is based on ChatGLM4-9B, which supports local deployment and has good Chinese semantic capabilities. This embodiment adopts the LoRA fine-tuning strategy to reduce training costs while maintaining the model's original expressive power.
[0101] Training configuration includes: Learning rate: Training rounds: 3; Maximum sequence length: 1024 tokens; Batch size: 2; Gradient accumulation: 8.
[0102] During the fine-tuning process, the system feeds the training set into the model in an "input → output" structure and optimizes the model parameters through a loss function, enabling the model to better describe the semantics of assembly actions.
[0103] 4) LLM Text Generation Process. During the AR assembly instruction generation phase, the system extracts the following from the generated assembly step data: assembly object sequence, part posture matrix (start / end), and assembly direction. The engineer provides the step attributes (step name, assembly type, tools involved, etc.) and fills them into a customized Prompt template. The model input is a structured query, and the output is a structured JSON. The system uses a text parser to extract the JSON from the model output and writes it to the AR instruction database.
[0104] Finally, a unified, standardized JSON database is constructed based on the spatial data, motion data, engineer input fields, and LLM output text collected in all the aforementioned steps, and an assembly guide that can be played on AR devices is automatically generated based on this database.
[0105] 1) Standardized file format and database structure. In this embodiment, a unified assembly instruction file format was designed, such as... Figure 10 As shown. The file structure encapsulates the following content in JSON format: Engineer input fields: step name, assembly type, tool type, start / end pose, etc.; Automatically generated fields: part pose matrix obtained from template point cloud registration, direction vector obtained from hand trajectory recognition, text instructions automatically generated by LLM, 3D model index of the target part, bounding box detection parameters (used for action judgment), and all steps are organized in array form, which the system can parse sequentially and play on the AR head-mounted display.
[0106] 2) The role of manually inputting data by engineers. To make AR instructions more closely resemble real-world processes, this embodiment allows engineers to input key metadata in the XR interactive environment, such as: assembly type (insertion, rotation, clamping, positioning, etc.), target part number, required tools (wrench, screwdriver, etc.), and start / end positions (already taught using bounding boxes). This information is written into JSON, providing semantic constraints for subsequent LLM text generation and AR animation generation.
[0107] 3) Automatically Generated Data and Processing Flow. The system automatically generates the following based on existing fields in the database: Part pose matrix, generated by the aforementioned template point cloud registration algorithm and directly written into JSON; Animation path and direction cues, generated from the direction vectors identified in the above steps. Combined with start and end poses, an interpolated animation is generated. Text descriptions are derived from the title and description fields automatically generated by LLM in the previous step. Real-time behavior detection logic (bounding box detection): the system automatically creates detection nodes in the AR program based on the bounding box range generated in the above steps.
[0108] 4) AR Program Generation Principle. The AR program generator reads the above JSON file and automatically completes: loading the part model and placing it in the corresponding pose, according to... and Generates SLERP / linear interpolation animations, displays text descriptions of LLM outputs, and builds step navigation interfaces (Next / Previous, etc.). This method generates complete AR guidance workflows without requiring engineers to write AR code.
[0109] The spatial pose information of the assembly object, the assembly step number, and the corresponding AR assembly guidance content are integrated. The integrated information is then encapsulated according to a predetermined data structure to obtain information in a unified data format. The encapsulated AR assembly guidance data is output to the augmented reality system for displaying the AR assembly guidance content in the augmented reality device. For example... Figure 10 As shown, AR assembly guidance data can be encapsulated using a standardized data structure and loaded by AR head-mounted display devices to achieve assembly step navigation, model overlay display, and text prompt output. For example... Figure 10 As shown, (A) shows the initial spatial position of the assembly object at the start of the assembly step, (B) shows the process of the assembly object moving towards the target assembly position during the assembly process, and (C) shows the state of the assembly object reaching the target spatial position and completing the current assembly step.
[0110] This application provides a no-code assembly guidance creation method. Engineers can complete spatial teaching by manually setting up 3D bounding boxes in a virtual reality or augmented reality environment without writing programs, which greatly reduces the production threshold of AR assembly guidance. It achieves high-precision 3D spatial registration without relying on deep learning models. Based on the RANSAC+ICP algorithm, the template point cloud registration method quickly estimates the posture matrix of the parts and achieves industrial-grade registration accuracy with an average translation error of about 0.0095m and a rotation error of about 5°. The system requires no large-scale training data, facilitating rapid adaptation to various assembly scenarios. It automatically identifies assembly direction and movement intent through hand trajectories, utilizing hand keypoint detection and trajectory analysis to infer assembly direction even when parts are obscured by hands, improving the accuracy and robustness of direction recognition. This makes it suitable for assembly actions such as insertion and sliding. The system automatically generates professional assembly text guidance through a large language model. Using a finely tuned large language model, it automatically generates standardized assembly guidance text based on spatial information, part relationships, and direction data for each step, significantly reducing manual writing costs and improving text consistency and professionalism. It automatically generates directly deployable AR assembly programs. The system writes part poses, animation paths, assembly text, detection logic, and other information into a unified data structure and automatically generates AR assembly guidance applications within the Unity3D physics engine, which can be directly deployed to AR headsets (such as HoloLens 2).
[0111] This application also provides a computer-readable storage medium storing computer-executable instructions for performing the AR assembly guidance generation method described in any of the preceding claims.
[0112] This application further provides a processing system, including a memory and a processor, wherein the memory stores the following instructions executable by the processor: for performing the steps of the AR assembly guidance generation method described in any of the preceding claims.
[0113] Figure 11 This is a schematic diagram of the composition structure of the AR assembly guidance generation system in the embodiments of this application, such as... Figure 11 As shown, it may include: a point cloud video acquisition module, a no-code space registration module, a template point cloud registration module, a first generation module, and a second generation module; wherein, The point cloud video acquisition module is configured to acquire 3D data of the actual assembly process. In one embodiment, the point cloud video acquisition module may include multiple RGBD cameras for synchronously acquiring point cloud videos of the assembly process. For example, a unified coordinate system can be established using the spatial registration objects of AprilTags, and the point clouds from multiple cameras can be fused into a continuous 3D point cloud video sequence to provide basic data for subsequent spatial analysis.
[0114] The no-code spatial registration module is configured to perform spatial teaching on the assembly object and generate spatial constraint data corresponding to the assembly steps based on the obtained point cloud video sequence. In one embodiment, the no-code spatial registration module can use a VR / AR interface to mark the start and end positions of the parts in the point cloud video using a 3D bounding box, thereby achieving programming-free spatial teaching; the center position, scale, and rotation of the bounding box are recorded to construct the spatial constraints of the assembly steps.
[0115] The template point cloud registration module is configured to determine the spatial pose of the assembly object based on spatial constraint data. In one embodiment, the template point cloud registration module can generate a corresponding template point cloud based on the CAD model of the assembly object, and match the generated template point cloud with the real point cloud extracted from the point cloud video. The spatial pose of the assembly object in a unified coordinate system is calculated through point cloud registration.
[0116] The first generation module is configured to generate AR assembly guidance content based on the determined spatial pose of the assembly object. In one embodiment, the first generation module can place the corresponding virtual 3D model in the correct position in the augmented reality environment based on the spatial pose information of the assembly object, and generate assembly guidance information corresponding to the assembly steps.
[0117] The second generation module is configured to integrate the generated assembly guidance content with the spatial pose information of the assembly object to obtain AR assembly guidance data that can be loaded and executed by the augmented reality system.
[0118] In one exemplary instance, the AR assembly guidance generation system provided in this application embodiment may further include: an assembly direction recognition module, configured to recognize the assembly direction of the assembly object based on the hand movement trajectory during the assembly process.
[0119] In one embodiment, the assembly orientation recognition module can be used for: Key points of the operator's hand are detected in the color image corresponding to the point cloud video, and the detected key points of the hand are mapped to a unified three-dimensional coordinate system. The hand movement trajectory is calculated based on the changes of the key points of the hand over time. Combined with the generated spatial constraint data, effective trajectory segments related to the assembly action are extracted from the hand movement trajectory. The assembly direction of the assembly object is inferred based on the spatial relationship between the effective trajectory segments and the assembly target position.
[0120] In one exemplary instance, the AR assembly guidance generation system provided in this application embodiment may further include: an assembly action detection module, configured to detect assembly actions based on spatial constraint data and control the advancement of assembly steps when it is necessary to automatically determine whether the assembly steps are completed during the AR assembly guidance process and to switch or prompt steps according to the operation status.
[0121] In one embodiment, the assembly action detection module can be used to: Based on the generated spatial constraint data, the spatial judgment area of the current assembly action is determined. During the AR assembly guidance process, the spatial relationship between the operator's hand position or the position of the assembly object and the spatial judgment area is detected in real time to determine whether the current assembly action meets the preset completion conditions. When the assembly object is detected to enter the spatial judgment area and meet the corresponding posture or position requirements, the current assembly step is determined to be completed, and the assembly guidance is triggered to enter the next assembly step, or the corresponding completion prompt or error correction prompt information is output to the operator.
[0122] In one exemplary instance, the AR assembly guidance generation system provided in this application embodiment may further include: a third generation module, configured to generate AR assembly guidance text based on a post-trained large language model when it is necessary to reduce the manual writing of AR assembly guidance text or to maintain consistency of text expression between different assembly steps.
[0123] In one embodiment, the third generation module can be used to: In the process of generating AR assembly guidance content, the spatial pose information, assembly step number, and assembly direction information of the assembly object corresponding to the assembly step are obtained; the spatial pose information, assembly step number, and assembly direction information are structured to form text generation input data that conforms to the preset input format; the text generation input data is input into a large language model trained with assembly domain data, and the post-trained large language model generates AR assembly guidance text corresponding to the current assembly step; the generated AR assembly guidance text is associated with the corresponding assembly step, and encapsulated together with the AR assembly guidance content in the subsequent AR assembly guidance data output process for display to the operator in the augmented reality environment.
[0124] The AR assembly guidance generation system provided in this application is an augmented reality assembly guidance creation system based on point cloud video analysis. It solves common problems in assembly guidance content creation, such as long production cycles, high professional barriers, reliance on manual spatial registration, difficulty in automatically structuring teaching information, and high dependence on manual development for augmented reality assembly programs. Targeting application scenarios where assembly objects and process steps frequently change in manufacturing, this application embodiment transforms the engineer's assembly teaching process into calculable and reusable digital assembly knowledge by acquiring and analyzing 3D point cloud video of the real assembly process. This achieves automated conversion from real assembly teaching to augmented reality assembly guidance programs, significantly reducing the production cost and technical barriers of AR assembly guidance content.
[0125] This application also provides an AR assembly guidance generation system, comprising at least: an expert assembly demonstration recording module, a part 3D space registration module, a template point cloud registration module, a no-code assembly instruction creation module, and an enhanced assembly program generation module; wherein... The expert assembly demonstration recording module is used to collect point cloud video data when experts perform the actual assembly process; The 3D space registration module for parts is used to spatially teach assembly objects in a no-code interactive environment and generate spatial constraint data corresponding to the assembly steps. The template point cloud registration module is used to register the template point cloud of the assembly object with the real point cloud based on spatial constraint data, so as to determine the spatial attitude information of the assembly object in a unified coordinate system. The no-code assembly instruction creation module is used to generate augmented reality assembly guidance content corresponding to the assembly steps based on the determined spatial posture information. The augmented assembly program generation module integrates the generated assembly guidance content, spatial attitude information, and assembly step information to generate an assembly guidance program that can be loaded and executed by the augmented reality system.
[0126] In one exemplary instance, an instruction text generation module is also included, for: During the generation of the AR assembly guidance content, the spatial pose information of the assembly object corresponding to the assembly step, the assembly step number, and the assembly direction information are obtained. The spatial attitude information, assembly step number, and assembly direction information are structured to form text generation input data that conforms to the preset input format. The text generation input data is input into a large language model trained with assembly domain data, and the post-trained large language model generates the AR assembly guidance text corresponding to the current assembly step. The generated AR assembly guidance text is associated with the corresponding assembly steps and encapsulated together with the AR assembly guidance content during the AR assembly guidance data output process for display to the operator in the augmented reality environment.
[0127] In one exemplary instance, an assembly direction recognition module is also included, which is used to: recognize the assembly direction of the assembly object based on the hand movement trajectory of the operator during the assembly process.
[0128] In one exemplary instance, an assembly process detection and step control module is also included, which is used to: detect whether the assembly action is completed based on the spatial constraint data during the execution of augmented reality assembly guidance, and control the automatic switching or prompt output of assembly steps.
[0129] In some embodiments, the AR assembly guidance generation system provided in this application comprehensively employs technologies such as multi-camera point cloud acquisition, no-code spatial teaching, template point cloud registration, hand trajectory analysis, and large language model text generation. Through the synergistic integration of multiple technologies, it achieves spatial perception, process understanding, and automatic generation of guidance content for the assembly process. Specifically, this embodiment acquires point cloud video data of the actual assembly process using multiple RGBD cameras and fuses it under a unified coordinate system, providing a complete and continuous three-dimensional data foundation for subsequent spatial analysis. Through no-code spatial teaching, engineers can complete the spatial annotation of assembly steps without writing programs. Based on this, the spatial posture of the assembly object is automatically determined using the template point cloud registration method, and the assembly direction is identified by combining hand trajectory analysis. Furthermore, AR assembly guidance text is generated through a large language model, ultimately automatically generating an augmented reality assembly guidance program that can be directly deployed.
[0130] The AR assembly guidance generation system provided in this embodiment has at least the following beneficial effects: The AR assembly guide generation system provided in this embodiment offers a no-code approach to creating assembly guides. Engineers do not need programming skills; they can complete spatial teaching of the assembly process simply by interactively setting up 3D bounding boxes in a virtual reality or augmented reality environment. This significantly reduces the barrier to entry and learning cost of creating AR assembly guide content.
[0131] The AR assembly guidance generation system provided in this embodiment achieves high-precision 3D spatial registration without relying on deep learning models. By using a template point cloud-based registration method, the spatial pose of the assembly object can be quickly estimated without a large amount of training data. This meets the accuracy and stability requirements of industrial assembly scenarios and is easily adapted to assembly scenarios with multiple varieties and small batches.
[0132] The AR assembly guidance generation system provided in this embodiment can automatically identify the assembly direction based on hand trajectory. Even when the assembly object is obscured by the operator's hand and the movement direction cannot be directly extracted from the part point cloud, it can still reliably infer the assembly direction and operation intention, thereby improving the robustness and applicability of assembly direction recognition. It is applicable to various assembly actions such as insertion and sliding.
[0133] Furthermore, the AR assembly guidance generation system provided in this embodiment introduces a large language model to automatically generate assembly guidance text. It can automatically generate standardized and professional assembly guidance text based on the spatial information of assembly steps, part relationships, and assembly direction data, which significantly reduces the workload of manually writing assembly instructions and improves the consistency and maintainability of assembly guidance text across different steps.
[0134] The AR assembly guidance generation system provided in this embodiment can automatically generate directly deployable augmented reality assembly guidance programs. The system uniformly writes information such as the spatial pose of the assembly object, the assembly animation path, the assembly guidance text, and the motion detection logic into a standardized data structure, and automatically generates the assembly guidance application in the augmented reality runtime environment. It can be directly deployed to augmented reality head-mounted display devices, thereby realizing a complete automated process from teaching to application of assembly guidance content.
[0135] Although the embodiments disclosed in this application are as described above, the content described is merely for the purpose of understanding this application and is not intended to limit this application. Any person skilled in the art to which this application pertains may make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in this application; however, the scope of patent protection of this application shall still be determined by the scope defined in the appended claims.
Claims
1. An AR assembly guidance generation method, characterized in that, Based on a no-code interactive environment, including: Obtain 3D data of the actual assembly process; Based on the obtained point cloud video sequence, spatial teaching is performed on the assembly object and spatial constraint data corresponding to the assembly steps are generated. Determine the spatial orientation of the assembly object based on spatial constraint data; Augmented reality (AR) assembly instructions are generated based on the determined spatial pose of the assembly object. The generated AR assembly guidance content is integrated with the spatial pose information of the assembly object to obtain AR assembly guidance data that can be loaded and executed by the augmented reality system.
2. The AR assembly guidance generation method according to claim 1, further comprising, after determining the spatial pose of the assembly object and before generating the augmented reality (AR) assembly guidance content: The assembly direction of the assembly object is identified based on the hand movement trajectory during the assembly process.
3. The AR assembly guidance generation method according to claim 2, wherein, The process of identifying the assembly direction of the assembly object includes: The key points of the operator's hand are detected in the color image corresponding to the point cloud video, and the detected key points of the hand are mapped to a unified three-dimensional coordinate system; The hand movement trajectory is calculated based on the changes in key hand points over time; Based on the generated spatial constraint data, effective trajectory segments related to the assembly action are extracted from the hand movement trajectory. Based on the spatial relationship between the effective trajectory segment and the assembly target position, the assembly direction of the assembly object is inferred.
4. The AR assembly guidance generation method according to claim 1 further includes: When it is necessary to automatically determine whether the assembly steps are completed during the AR assembly guidance process, and to switch or prompt the steps according to the operation status, the assembly action is detected based on the spatial constraint data, and the progress of the assembly steps is controlled. The process of detecting assembly actions based on the spatial constraint data and controlling the progression of assembly steps includes: Based on the generated spatial constraint data, the spatial determination area of the current assembly action is determined; During the AR assembly guidance process, the spatial relationship between the operator's hand position or the position of the assembly object and the spatial judgment area is detected in real time to determine whether the current assembly action meets the preset completion conditions. When the assembly object is detected to have entered the spatial determination area and met the corresponding posture or position requirements, the current assembly step is determined to be completed, and the assembly guidance is triggered to enter the next assembly step, or the corresponding completion prompt or error correction prompt information is output to the operator.
5. The AR assembly guide generation method according to claim 1, further comprising, after generating the augmented reality (AR) assembly guide content and before obtaining the AR assembly guide data that can be loaded and executed by the augmented reality system: When it is necessary to reduce the amount of manually written AR assembly guidance text, or to maintain consistency in text expression across different assembly steps, AR assembly guidance text can be generated based on a post-trained large language model.
6. The AR assembly guidance generation method according to claim 5, wherein, The AR assembly guidance text generated by the post-trained large language model includes: During the generation of the AR assembly guidance content, the spatial pose information of the assembly object corresponding to the assembly step, the assembly step number, and the assembly direction information are obtained. The spatial attitude information, assembly step number, and assembly direction information are structured to form text generation input data that conforms to the preset input format. The text generation input data is input into a large language model trained with assembly domain data, and the post-trained large language model generates the AR assembly guidance text corresponding to the current assembly step. The generated AR assembly guidance text is associated with the corresponding assembly steps and encapsulated together with the AR assembly guidance content during the AR assembly guidance data output process for display to the operator in the augmented reality environment.
7. The AR assembly guidance generation method according to claim 1, 2, 4 or 5, wherein, The step of spatially teaching the assembly object and generating spatial constraint data corresponding to the assembly steps based on the obtained point cloud video sequence includes: In point cloud video, an initial 3D bounding box is created before the assembled object begins to move, and a final 3D bounding box is created after the assembly is completed and the object is stably in place, using AR / VR interaction. Extract bounding box parameters, including center position, spatial size, and spatial orientation, from the starting and ending 3D bounding boxes respectively; The bounding box parameters are associated with the corresponding assembly step numbers to form the spatial constraint data used to describe the range of spatial position changes of the assembled object.
8. The AR assembly guidance generation method according to claim 1, 2, 4 or 5, wherein, Determining the spatial orientation of the assembly object based on spatial constraint data includes: Based on the CAD model of the assembly object, its surface is sampled to generate a template point cloud of the assembly object; Based on the generated spatial constraint data, real point cloud data within the spatial constraint range is extracted from the point cloud video and used as candidate point clouds for the assembly object; The captured real point cloud is preprocessed to remove noise points and background points; The preprocessed real point cloud is matched with the template point cloud, and the spatial transformation relationship between the two is calculated; Based on the spatial transformation relationship, the spatial attitude information of the assembly object in the unified world coordinate system is determined; wherein, the spatial attitude information includes the spatial position and spatial orientation of the assembly object, which is used to describe the actual placement state of the assembly object in three-dimensional space.
9. The AR assembly guidance generation method according to claim 1, 2, 4 or 5, wherein, The generated AR assembly guidance content includes: Based on the determined spatial pose information of the assembly object, the display position and display pose of the assembly object in the augmented reality coordinate system are determined; The virtual 3D model of the assembly object is loaded into the augmented reality environment and placed according to the determined spatial posture; The spatial pose information of the assembly object is associated with the corresponding assembly steps to generate the AR assembly guidance content corresponding to the assembly steps.
10. The AR assembly guidance generation method according to claim 1, 2, 4 or 5, wherein, The process of integrating the generated AR assembly guidance content with the spatial pose information of the assembly object includes: The spatial posture information of the assembly object, the assembly step number, and the corresponding AR assembly guidance content are integrated; The integrated information is encapsulated according to a predetermined data structure to obtain information in a unified data format. The packaged AR assembly guidance data is output to the augmented reality system for displaying the AR assembly guidance content in the augmented reality device.
11. An AR assembly guidance generation system, characterized in that, include: The expert assembly demonstration recording module is used to collect point cloud video data when experts perform the actual assembly process; The 3D space registration module for parts is used to spatially teach assembly objects in a no-code interactive environment and generate spatial constraint data corresponding to the assembly steps. The template point cloud registration module is used to register the template point cloud of the assembly object with the real point cloud based on spatial constraint data, so as to determine the spatial attitude information of the assembly object in a unified coordinate system. The no-code assembly instruction creation module is used to generate augmented reality assembly guidance content corresponding to the assembly steps based on the determined spatial posture information. The augmented assembly program generation module integrates the generated assembly guidance content, spatial attitude information, and assembly step information to generate an assembly guidance program that can be loaded and executed by the augmented reality system.
12. The AR assembly guidance generation system according to claim 11, further comprising an instruction text generation module, used for: During the generation of the AR assembly guidance content, the spatial pose information of the assembly object corresponding to the assembly step, the assembly step number, and the assembly direction information are obtained. The spatial attitude information, assembly step number, and assembly direction information are structured to form text generation input data that conforms to the preset input format. The text generation input data is input into a large language model trained with assembly domain data, and the post-trained large language model generates the AR assembly guidance text corresponding to the current assembly step. The generated AR assembly guidance text is associated with the corresponding assembly steps and encapsulated together with the AR assembly guidance content during the AR assembly guidance data output process for display to the operator in the augmented reality environment.
13. The AR assembly guidance generation system according to claim 11 further includes an assembly direction recognition module, used to: identify the assembly direction of the assembly object based on the hand movement trajectory of the operator during the assembly process.
14. The AR assembly guidance generation system according to claim 11 further includes an assembly process detection and step control module, used to: detect whether the assembly action is completed based on the spatial constraint data during the execution of the augmented reality assembly guidance, and control the automatic switching or prompt output of the assembly steps.