AR auxiliary assembly guide generation method and device, equipment, medium and product

By building an IDEF1X model and knowledge graph, combining a large visual language model, real-time identification of assembly status and generating dynamic guidance information, the problem of insufficient intelligence in AR-assisted assembly technology is solved, and assembly efficiency and accuracy are improved.

CN120429447AInactive Publication Date: 2025-08-05YANGTZE DEITA GRADUATE SCHOOI OF BEIJING INST OF TECH (JIAXING) +1

Patent Information

Application Number
CN202510933203.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-08-05
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing AR assisted assembly technology has limited intelligence level, lacks real-time perception and intelligent judgment capabilities, and cannot automatically update guidance information based on assembly status, resulting in inefficient assembly efficiency.

Method used

By building an IDEF1X model and knowledge graph, combining a large visual language model, we can identify assembled parts and states in real time, generate dynamic guidance information, and display and interact in real time through AR devices, including gestures, voice and eye interaction.

Benefits of technology

It has achieved the improvement of the intelligence level of the AR assisted assembly system, real-time update of guide information, improved assembly efficiency and accuracy, reduced manual operation, and enhanced the intelligence and real-timeness of the assembly process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429447A_ABST
    Figure CN120429447A_ABST
Patent Text Reader

Abstract

The invention discloses an AR auxiliary assembly guide generation method and device, equipment, a medium and a product, and relates to the field of AR auxiliary assembly, and the method comprises the steps: obtaining an assembly image at a current moment; inputting the assembly image at the current moment into the trained guide stage recognition model to obtain an assembly part and an assembly state at the current moment; the guiding stage recognition model is obtained by fusing an assembly information knowledge graph embedded vector into a visual big language model; the corresponding guide information is generated according to the assembly part and the assembly state at the current moment, real-time sensing can be conducted according to the assembly environment, the assembly state is determined in real time, and the guide information is automatically updated according to the assembly state determined in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of AR-assisted assembly, and in particular to an AR-assisted assembly guide generation method, device, equipment, medium and product. Background Art

[0002] AR-assisted assembly technology has been widely used in many fields in recent years, especially in the aerospace, automobile manufacturing, and electrical equipment manufacturing industries. Although AR-assisted assembly technology has made certain progress, some problems still exist.

[0003] 1. Limited level of intelligence: Most existing systems rely on preset assembly processes and guidance information, and lack the ability to perceive the assembly environment in real time and make intelligent judgments.

[0004] 2. Insufficient real-time performance of guidance information: The guidance information cannot be automatically updated according to real-time changes in assembly status, resulting in the need for assemblers to manually switch or adjust the guidance content. Summary of the Invention

[0005] The purpose of this application is to provide an AR-assisted assembly guidance generation method, device, equipment, medium and product, which can perform real-time perception based on the assembly environment, determine the assembly status in real time, and automatically update the guidance information based on the real-time determined assembly status.

[0006] To achieve the above objectives, this application provides the following solutions.

[0007] In a first aspect, the present application provides an AR-assisted assembly guide generation method, comprising: obtaining an assembly image at a current moment.

[0008] The assembly image at the current moment is input into the trained guidance stage recognition model to obtain the assembly parts and assembly status at the current moment; the guidance stage recognition model is obtained by fusing the assembly information knowledge graph embedding vector into the visual large language model.

[0009] Generate corresponding guidance information based on the current assembly parts and assembly status.

[0010] In one embodiment, the process of constructing the recognition model in the guidance phase specifically includes: extracting entity relationships from the IDEF1X model to obtain entities and relationships between entities; the IDEF1X model is obtained by processing each sample assembly image and part operation manual using the IDEF1X method.

[0011] Construct the first knowledge graph based on each entity and the relationship between entities.

[0012] Taking the assembly status of each part in the IDEF1X model and the image feature vectors of each sample assembly image as entities, a second knowledge graph is constructed based on the first knowledge graph.

[0013] The second knowledge graph is processed using a knowledge graph embedding algorithm to obtain an assembly information knowledge graph embedding vector.

[0014] The assembly information knowledge graph embedding vector is fused with the intermediate representation layer of the visual language model to obtain the guidance stage recognition model.

[0015] In one embodiment, after generating corresponding guidance information according to the current assembly parts and assembly status, the method further includes: processing the guidance information according to the assembler's gesture interaction signal, eye interaction signal, or voice interaction signal.

[0016] In one embodiment, the assembly state includes a preparation stage, an assembly stage, and a calibration stage.

[0017] The guidance information corresponding to the preparation stage includes an introduction to assembly parts information.

[0018] The guidance information corresponding to the assembly stage includes dynamic assembly animation.

[0019] The guidance information corresponding to the calibration stage includes prompt information and the next part to be assembled.

[0020] In one embodiment, after generating corresponding guidance information according to the assembly parts and assembly status at the current moment, the method further includes: fusing the guidance information with the real scene through a three-dimensional registration function.

[0021] In one embodiment, before obtaining the assembly image at the current moment, the method further includes: building guidance information in Unity software.

[0022] In a second aspect, the present application provides an AR-assisted assembly guide generation device, including: an AR device interaction system for obtaining an assembly image at the current moment.

[0023] The visual large language model recognition system is used to input the current assembly image into a trained guidance stage recognition model to obtain the assembly parts and assembly status at the current moment; the guidance stage recognition model is obtained by fusing the assembly information knowledge graph embedding vector into the visual large language model.

[0024] The guidance information generation system is used to generate corresponding guidance information according to the assembly parts and assembly status at the current moment.

[0025] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the above-described AR-assisted assembly guide generation methods.

[0026] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described AR-assisted assembly guide generation methods.

[0027] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned AR-assisted assembly guide generation methods.

[0028] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides an AR-assisted assembly guidance generation method, device, equipment, medium and product. This application proposes an AR-assisted assembly guidance method based on a visual large language model to identify the assembly status. By inputting the assembly image at the current moment into a trained guidance stage recognition model, the assembly parts and assembly status at the current moment are obtained, and corresponding guidance information is generated according to the assembly parts and assembly status at the current moment. The assembly status can be determined in real time, and the guidance information can be automatically updated according to the real-time determined assembly status, thereby improving the intelligence level of the AR-assisted assembly system, realizing real-time updating and recommendation of guidance information, and thus improving assembly efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0030] Figure 1 This is a schematic diagram of the IDEF1X model.

[0031] Figure 2 This is a diagram of the database system construction process.

[0032] Figure 3 Generate a flow chart for the guidance information.

[0033] Figure 4 A schematic diagram of the guidance information composition for each stage.

[0034] Figure 5 This is an architecture diagram of an AR-assisted assembly guide generation device provided in one embodiment of the present application.

[0035] Figure 6 A flowchart of a method for generating AR-assisted assembly guidance is provided in one embodiment of the present application.

[0036] Figure 7 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0037] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0038] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0039] In an exemplary embodiment, Figure 6 As shown, an AR-assisted assembly guide generation method is provided, including the following steps 201 to 203.

[0040] Step 201: Acquire the assembly image at the current moment.

[0041] Step 202: Input the current assembly image into the trained guidance phase recognition model to obtain the assembly parts and assembly status at the current moment; the guidance phase recognition model is obtained by fusing the assembly information knowledge graph embedding vector into the visual large language model.

[0042] Step 203: Generate corresponding guidance information according to the current assembly parts and assembly status.

[0043] In another exemplary embodiment of the present application, the process of constructing the recognition model in the guidance stage specifically includes: extracting entity relationships from the IDEF1X model to obtain entities and relationships between entities; the IDEF1X model is obtained by processing each sample assembly image and part operation manual using the IDEF1X method.

[0044] Construct the first knowledge graph based on each entity and the relationship between entities.

[0045] Taking the assembly status of each part in the IDEF1X model and the image feature vectors of each sample assembly image as entities, a second knowledge graph is constructed based on the first knowledge graph.

[0046] The second knowledge graph is processed using a knowledge graph embedding algorithm to obtain an assembly information knowledge graph embedding vector.

[0047] The assembly information knowledge graph embedding vector is fused with the intermediate representation layer of the visual language model to obtain the guidance stage recognition model.

[0048] In practical applications, IDEF1X is an entity-relationship modeling method used to describe entities, attributes, and relationships in a system. Specifically, it includes: entity definition: defining the core entities in the basic assembly process; relationship definition: describing the relationship between entities, such as "parts use tools", "process paths contain processes", etc. Attribute definition: defining attributes for each entity, such as the name, quantity, and appearance of the part, the model and location of the tool, etc. Semantic rules: ensuring the semantic consistency of entities and relationships in the model, so as to maintain data integrity and transparency in complex assembly processes. Through the IDEF1X method, complex assembly process information is structured to form a standardized information model, namely the IDEF1X model, such as Figure 1 This model can not only clearly describe the various elements and their relationships in the assembly process, but also provide basic data for the subsequent construction of the knowledge graph.

[0049] In practical applications, a knowledge graph is a graph structure used to represent knowledge, consisting of nodes (entities) and edges (relationships). Figure 1 The IDEF1X model shown in the figure is processed, and based on the structured information generated by the IDEF1X method, the assembly information knowledge graph, i.e., the second knowledge graph, is constructed. The steps are as follows.

[0050] 1. Extraction of entities and relationships: Extract entities (such as parts, tools, process paths, etc.) and relationships (such as "use", "include", etc.) from the IDEF1X model.

[0051] 2. Graph Construction: These entities and relationships are entered into a graph database to form an instance-relationship graph, the first knowledge graph. Semantic rules defined by the IDEF1X methodology ensure semantic consistency between entities and relationships in the knowledge graph, enabling information sharing and interaction between heterogeneous models.

[0052] 3. Based on the first knowledge graph, a second knowledge graph is constructed with the assembly status of each part in the IDEF1X model and the image feature vector of the sample assembly image as entities.

[0053] In another embodiment of the present application, the visual language model uses DeepSeek-VL2. DeepSeek-VL2 comprises a visual encoder, a language decoder, and an intermediate fusion module. The visual encoder processes input image data and extracts image feature vectors; the language decoder generates text descriptions or guidance information; and the intermediate fusion module is responsible for fusing visual features with language information.

[0054] In another embodiment of the present application, the knowledge graph embedding algorithm is the TransE algorithm, which uses the knowledge graph embedding algorithm to convert the entities and relationships in the second knowledge graph into vector form. These vectors can capture the semantic information of the entities and relationships, providing rich contextual support for the visual language model.

[0055] Assuming both entities and relations can be represented by vectors, the relation vector can be viewed as a translation from the head entity vector to the tail entity vector. The TransE algorithm minimizes the distance between the head entity vector plus the relation vector and the tail entity vector to obtain the assembly information knowledge graph embedding vector. This embedding vector is then embedded into the visual language model.

[0056] .

[0057] Among them: h represents the head entity vector, r represents the relationship vector, and t represents the tail entity vector.

[0058] In order to enhance the ability of the visual large language model in assembly state recognition and guidance information generation, structured knowledge is integrated into the visual large language model through knowledge graph embedding technology. In the hidden layer of the visual large language model, the assembly information knowledge graph embedding vector is fused with the intermediate representation layer of the model to enhance the model's understanding of the assembly scene. In another exemplary embodiment of the present application, the assembly information knowledge graph embedding vector is fused with the intermediate representation layer of the visual large language model to obtain a guidance stage recognition model, specifically: the sample assembly image is input into the visual encoder to obtain an image feature vector, which is then spliced or weighted summed with the assembly information knowledge graph embedding vector to form a fused feature representation. This fusion method enables the model to fully utilize the structured semantic knowledge in the knowledge graph while processing image information. For example, when identifying the assembly state of a part, the model not only makes judgments based on the visual features in the image, but also refers to the pre-constructed image vector information and assembly state information about the part in the knowledge graph, thereby improving the accuracy and robustness of recognition.

[0059] In another exemplary embodiment of the present application, the guidance stage recognition model is trained. The specific process is as follows: after the assembly information knowledge graph embedding vector is fused into the visual large language model, the entire model needs to be trained and optimized. Using an image dataset with assembly status annotations, the parameters of the guidance stage recognition model are adjusted through the back propagation algorithm so that the model can accurately identify the assembly state based on the fused feature representation and generate corresponding guidance information. In the process of the back propagation algorithm, the image feature vector is actually involved in the calculation. During the training process, the model will learn how to combine the semantic information in the knowledge graph with the image features to better complete the recognition of parts and assembly status.

[0060] In another exemplary embodiment of the present application, the mixed reality head-mounted device has high-resolution holographic projection, which can clearly display 3D assembly models and guidance information. Its field of view is wide, providing users with a wider field of view and reducing head movement when viewing information. It uses advanced sensors and processing chips to accurately track head posture and gestures to ensure the smoothness and accuracy of interaction. According to system requirements, the mixed reality head-mounted device is tested for adaptability to ensure that it meets the requirements of positioning accuracy, field of view, display effect, etc. in the assembly scene. At the same time, the hardware parameters of the device, such as resolution, refresh rate, sensor, etc., are configured to achieve the best interactive experience. This application uses a mixed reality head-mounted device as an AR device. When the assembler performs assembly operations, the head-mounted camera on the mixed reality head-mounted device he wears captures images and transmits the images to the visual large language model.

[0061] In another exemplary embodiment of the present application, after the assembly image at the current moment is input into the trained guidance phase recognition model, the processing steps within the trained guidance phase recognition model are specifically as follows.

[0062] (1): Image feature extraction.

[0063] 1. Dynamic tiling visual encoding strategy.

[0064] Image Preprocessing and Tiled Segmentation: The input image is first preprocessed and resized to a suitable size. DeepSeek-VL2 draws on the established slicing and tiling method to dynamically segment the high-resolution image into multiple local tiles. This approach avoids the limitations of traditional fixed-resolution encoders and excels at processing assembly images with varying aspect ratios and high resolution. For example, the model can segment a large assembly drawing or complex mechanical structure image into multiple, easily manageable parts, each containing key visual information.

[0065] 2. Visual feature extraction.

[0066] The pre-trained SigLIP-SO400M-384 visual encoder processes each tile and extracts local features. Operating at a base resolution of 384×384, it generates 27×27=729 1152-dimensional visual embeddings from each tile. These embeddings capture features such as the edges, shapes, and states of the parts within the tile, providing rich visual information for subsequent feature fusion.

[0067] 3. Question guidance mechanism and feature integration.

[0068] During the image feature extraction process, a question-guided mechanism provides contextual information to the model, helping it better understand the image content. For example, "Which part is being assembled? At what assembly step is it currently?" The generated question is converted into a vector and fused with the image feature vector. The fused feature vector not only contains the visual features of the parts in the image but also incorporates the semantic information provided by the question-guided mechanism. This provides richer information for accurately identifying assembled parts and their assembly states, improving recognition accuracy and reducing the likelihood of the model experiencing hallucinations.

[0069] .

[0070] .

[0071] Where: Q is the guided question, BERT() represents the text embedding generation process based on the BERT model, which converts text (such as the guided question) into vector form (embedding representation). These vectors can capture the semantic information of the text, allowing the trained guided phase recognition model to understand the meaning of the text. is the image feature vector, Embedding for the problem, is the fused feature vector, Concat ( , ) means and Connect them.

[0072] (2): After completing the above image feature extraction process, the resulting fused feature vector will be used for subsequent assembly state recognition. This fused feature vector will be used as input, and the deep learning model module in the visual language model will calculate the probability distribution of the assembly part category and state at the current moment to determine the assembly part and its assembly state at the current moment.

[0073] The identification formula for assembly parts is: .

[0074] in: represents the probability that the assembled part at the current moment is the kth part category, e represents the natural logarithm, is the weight vector of the i-th part category, Express Find the transpose, represents the weight vector of the k-th part category, Express Find the transpose, K represents the total number of part categories, is the image feature vector.

[0075] The identification formula for the assembly state is: .

[0076] in: It represents the probability that the current assembly state is the s-th assembly state under the condition that the assembly part at the current moment is the k-th part category, is the weight vector of the s-th assembly state, is the weight vector of the j-th assembly state, Express Find the transpose, Express Find the transpose, is the image feature vector, since the assembly status is divided into three categories, j=1, 2, 3.

[0077] The assembly parts and assembly status at the current moment are obtained by calculating the probability.

[0078] In another exemplary embodiment of the present application, the assembly state includes a preparation stage, an assembly stage, and a calibration stage. Preparation stage: The assembler turns his view to the assembly part, but has not yet started the actual assembly operation. At this time, the part may still be in the position to be assembled, not connected or fixed with other parts, and is in a relatively independent state, waiting for the assembler's next action. Assembly stage: The assembler has picked up the current part and is ready to assemble and connect it with other parts. At this time, the part may have partially approached or contacted the assembly position, but has not yet been fully installed in place. During the dynamic assembly process, its position and posture may continue to change until it is correctly installed in the specified position. Calibration stage: The part has been installed to the specified position, but there may be certain position deviations or loose connections. Further adjustments and calibrations are required to ensure the assembly accuracy and quality of the part. At this time, the part may have been preliminarily connected with other parts, but has not met the final assembly requirements. It needs to be calibrated to achieve the best assembly state.

[0079] like Figure 4As shown in the figure, the guidance information corresponding to the preparation stage includes an introduction to assembly parts information. Specifically, the assembly parts information is presented in text format, introducing the basic information of the current assembly part, including the part name, model, purpose, and basic parameters. This allows the assembler to have a comprehensive understanding of the part to be assembled. It also provides safety precautions and other important reminders related to the part to ensure a smooth assembly process.

[0080] Guidance information for the assembly phase includes dynamic assembly animations. These animations demonstrate the complete assembly process from the part's current position to the target assembly position, including key information such as the assembly path, direction, angle, and force. These animations intuitively guide the assembler through the correct operation, preventing assembly errors or part damage caused by improper operation. The animations can be played and paused in real time to ensure the accuracy and effectiveness of the guidance information, following the assembler's progress.

[0081] The guidance information during the calibration phase includes prompts and the next part to be assembled. These prompts are presented in text format and can include information such as the tools required for calibration, the steps involved, and the calibration standards. They also provide solutions to common problems and precautions to help the assembler complete the calibration smoothly. Furthermore, the assembler is informed of the name, location, and assembly method of the next part to be assembled, allowing them to prepare in advance and improve assembly efficiency.

[0082] When it is recognized that the assembly part is the last part and the assembly status is entering the calibration phase, the guidance information further includes text asking the assembler whether the assembly is completed.

[0083] In another exemplary embodiment of the present application, before acquiring the current assembly image, the process also includes: creating guidance information in Unity software. Based on the three assembly states corresponding to different parts, different guidance information scenarios are created in Unity using the content in the database and knowledge graph and saved to the scenario library for subsequent retrieval by the AR device interaction system.

[0084] The basic steps of scene construction are as follows.

[0085] 1. Model import: Import assembly parts and assemblies into Unity. These models can be created using SolidWorks 3D modeling software.

[0086] 2. Scene setup: Create a new scene in Unity and place the imported 3D model into the scene. Adjust the model's position, rotation, and scale according to the actual assembly environment to ensure that the model matches the real environment when displayed on the AR device. Add light sources (such as directional lights and point lights) to ensure that the model has good lighting effects in the AR environment.

[0087] 3. Create the guide information element.

[0088] Create a UI Canvas in Unity and create a Text element under the Canvas to display text information in the preparation stage, such as part name, model, and purpose.

[0089] Create an animation sequence for each part to show the complete assembly process from the current position to the target assembly position, define keyframes in the animation, and set the part's movement path, rotation angle, and scaling.

[0090] Create a UI Canvas in Unity and create a Text element within the Canvas to display textual information about the calibration phase, including information about the calibration tools, calibration steps, and calibration standards. Add buttons or interactive elements to the calibration phase to trigger the calibration operation or display assembly information for the next part.

[0091] In another exemplary embodiment of the present application, after generating corresponding guidance information based on the current assembly parts and assembly status, the process also includes: integrating the guidance information with the real-world scene through a 3D registration function. Specifically, this is based on Vuforia's 3D model registration function. 3D model registration is a key technology for augmented reality, used to precisely align 3D models with real-world objects. It identifies features of real-world objects and matches them with those of 3D models to determine the objects' position and orientation in space. This ensures that virtual content can be accurately placed on real-world objects and maintains alignment with them as the user's perspective or the object's position changes. This application utilizes Vuforia to complete 3D registration of assemblies. The basic steps are as follows: First, create a 3D model of the model to be assembled using software such as SolidWorks, UG, or Unity. The model should contain sufficient textures and feature points for Vuforia to recognize and align. The Vuforia engine is then imported into the Unity project, and the Vuforia database is configured. Finally, the virtual content is imported into the scene and tested in real-world conditions to complete the 3D registration of the model. This application completes the 3D registration stage of the assembly model based on Vuforia. Vuforia uses image recognition technology to extract the target's feature points, uses an efficient feature matching algorithm for fast and accurate matching, configures Vuforia's ObjectTarget in Unity, and uses the camera to scan the target in the real environment to achieve alignment between the virtual model and the real object.

[0092] In another exemplary embodiment of the present application, after generating corresponding guidance information according to the assembly parts and assembly status at the current moment, the method further includes: processing the guidance information according to the assembler's gesture interaction signal, eye interaction signal or voice interaction signal.

[0093] Mixed reality headsets support gesture interaction: Mixed reality headsets can recognize various gestures, such as air click and home gesture. Air click is a gesture where you raise your hand and click, often used to select an object, similar to a mouse click. The home gesture is to extend your hand, palm up, bring your fingers together, and then spread them apart to return to the home screen.

[0094] Assemblers can use air click gestures to select virtual models or buttons to trigger corresponding operations, such as starting assembly, playing assembly animations, etc. For example, during the assembly stage, assemblers can use air clicks to select the virtual model of a part to view its detailed information or start assembly operations. Using the home gesture to return to the main menu or switch between different assembly step menus allows assemblers to quickly switch operating modes, view different assembly information, or select other parts for assembly. Gestures are used to grab, move, rotate, and scale virtual objects, allowing assemblers to more clearly observe the structure and assembly position of the parts. For example, during the preparation stage, assemblers can grab and rotate the virtual model of a part to understand its appearance and features from different angles.

[0095] Mixed reality headsets support voice interaction: The voice interaction function of mixed reality headsets relies on the speech recognition system in the Mixed Reality Toolkit (MRTK), supports voice commands in multiple languages, and controls applications through voice commands.

[0096] Assemblers can control the assembly process through voice commands, such as "Start assembly," "Next," "Pause," or "Back." The system then executes the corresponding actions based on the commands, switching assembly states or displaying corresponding guidance information. For example, during the assembly process, the assembler can say "Next" to view the assembly instructions for the next part. Furthermore, assemblers can use voice commands to query assembly information, such as "Query part name," "Query assembly steps," or "Query precautions." The system will provide corresponding text information or voice prompts based on the commands. For example, during the preparation phase, the assembler can say "Query part usage" to understand the function and purpose of the current part.

[0097] Eye tracking technology on mixed reality headsets provides monocular gaze ray information about where the user is looking, enabling developers to design natural and intuitive input and interaction schemes.

[0098] By gazing at specific virtual buttons or areas, corresponding actions can be triggered, such as starting animations and playing voice prompts, achieving a more natural way of interaction. For example, if the assembler gazes at the virtual model of a part for a few seconds, the system will automatically play the assembly animation of the part to guide the assembler through the operation. During the assembly process, by gazing at different areas or objects, navigation and switching functions can be achieved, such as switching assembly steps and viewing detailed information on different parts. For example, when the assembler's eyes are fixed on the preview information of the next part, the system will automatically switch to the assembly guidance phase of that part.

[0099] This application provides a method for generating AR (augmented reality)-assisted assembly guidance, including: ① constructing an assisted assembly process database and knowledge graph based on professional information; ② building a guidance information scene in Unity; and ③ identifying assembly status and providing guidance information through a large visual language model. By integrating assembly knowledge, identifying assembly status in real time, and intelligently generating guidance information, this application improves the efficiency of AR-assisted assembly and enhances the intelligence level of the system. Different guidance information is provided according to different stages, and AR technology is used to provide assemblers with intuitive and convenient assembly guidance, thereby improving assembly efficiency.

[0100] During the knowledge graph construction stage, this application converts the collected data into a structured form through the IDEF1X entity relationship extraction step, constructs an information model of the basic assembly process, extracts entities and relationships from all knowledge documents (sample assembly images and parts operation manuals), and enters the obtained entities and the relationships between them into the Neo4j graph database to obtain an enhanced assembly knowledge graph. The knowledge graph is then embedded in a visual big language model to obtain a more specialized pre-trained visual big language model.

[0101] In the stage of identifying assembly status and providing guidance information, this application uses knowledge graph-based guidance and contextual prompts to allow the visual large language model to autonomously identify the current assembly parts and assembly status, and provide different guidance information to the assembler according to different assembly statuses.

[0102] After the AR-assisted assembly guidance generation method provided in this application is implemented in an AR device, the AR device will overlay virtual assembly guidance information with the real world based on the guidance information to guide the assembler through the assembly process. The specific steps are as follows: 1. Data Reception and Processing: The AR device receives guidance information packets sent from the server via a network module. These packets contain guidance content such as text and animation generated based on the current assembly status. The device's processor decodes the data and converts it into a format that can be displayed in the AR interface. 2. Information Fusion and Positioning: Through 3D registration, the AR device accurately integrates the guidance information with the real scene. Based on pre-identified assembly features, it determines the spatial position of the virtual guidance information, ensuring that text prompts, animations, etc. are accurately superimposed on the corresponding parts or assembly locations. 3. Guidance Information Display: Based on the current assembly status, the AR device displays appropriate guidance information at the corresponding location in the field of view. During the preparation phase, a text box containing basic information such as the name and model number is displayed above the part. During the assembly phase, a dynamic assembly animation is projected into the operating area, showing the assembly path and direction of the part in real time. During the calibration phase, prompt text is displayed near the calibration area and a preview of the next part is provided. 4. Interactive response: The assembler interacts with the guidance information through gestures, voice, or gaze. If the assembler uses gestures to click on a virtual button, the device's gesture vision large language model recognition system captures the action and triggers the corresponding operation, such as playing the next stage of guidance; if the assembler issues a voice command to inquire about assembly details, the voice vision large language model recognition system analyzes the command, requests additional information from the server, and displays it; gaze interaction uses eye tracking technology, and when the assembler looks at a specific virtual element, detailed information or operation options are automatically displayed. 5. Real-time feedback and adjustment: The device sensors continuously collect assembler operation data, such as gesture movements, assembly progress, etc., and upload it to the server in real time. The server evaluates the assembly accuracy and progress based on this data, and adjusts the guidance information when necessary, such as pushing correction prompts in the event of assembly errors, or loading the next stage of guidance content in advance according to the progress, to achieve accurate and real-time assembly assistance.

[0103] This application uses knowledge graph technology to accurately extract knowledge nodes such as component information, assembly procedures, and operating specifications involved in the assembly process, forming a clear assembly knowledge graph, making the relationship between assembly behavior and knowledge points clearer.

[0104] By leveraging the visual reasoning capabilities of the large visual language model, the assembly status can be judged in real time, avoiding the tedious operation of manually switching guidance information in traditional systems.

[0105] Knowledge graph technology is used to integrate structured and unstructured assembly-related data to form a local knowledge base encompassing assembly processes, component information, and operating specifications. This knowledge base provides rich semantic information support for the system. Leveraging the real-time interactive capabilities of AR devices, the system extracts corresponding guidance information from the local knowledge base based on the identified assembly state and displays it to the assembler in real time via the AR device. This dynamically updated guidance information reduces the time assemblers spend consulting manuals, improving overall assembly efficiency. Furthermore, intuitive visual guidance further enhances assembly accuracy and efficiency.

[0106] This application combines knowledge graphs with large visual language models, leveraging the structured knowledge and contextual information provided by the knowledge graph to effectively reduce the risk of hallucinations when the large visual language model generates assembly guidance information. The knowledge graph not only enhances the model's reasoning capabilities but also ensures the accuracy and reliability of the generated information by validating and constraining the model's output in real time.

[0107] During visual understanding, pre-defined prompts are provided based on the context of the assembly task, such as the current assembly state and the parts involved. These prompts help the model more accurately understand the image content. This context-based visual understanding technology can effectively improve the accuracy and relevance of visual reasoning.

[0108] The visual language model learns the visual features and context of the assembly scene to accurately identify the assembly progress. Guidance information can be dynamically updated based on the assembly progress, ensuring that assemblers always receive the latest and most accurate guidance.

[0109] Based on the same inventive concept, the present application also provides an AR-assisted assembly guide generation device for implementing the aforementioned AR-assisted assembly guide generation method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more AR-assisted assembly guide generation device embodiments provided below can be found in the above-mentioned limitations of the AR-assisted assembly guide generation method, and will not be repeated here.

[0110] In an exemplary embodiment, an AR-assisted assembly guide generation apparatus is provided, including: an AR device interaction system for acquiring an assembly image at a current moment.

[0111] The visual large language model recognition system is used to input the current assembly image into a trained guidance stage recognition model to obtain the assembly parts and assembly status at the current moment; the guidance stage recognition model is obtained by fusing the assembly information knowledge graph embedding vector into the visual large language model.

[0112] The guidance information generation system is used to generate corresponding guidance information according to the assembly parts and assembly status at the current moment.

[0113] In AR-assisted assembly guidance generation devices, accurately identifying the assembly state is key to providing precise guidance information to assemblers. In this application, the visual large language model analyzes and processes assembly images in real time, identifying the currently assembled parts and their assembly states (preparation, assembly, and calibration). This accurately transmits the current parts and their assembly states to the guidance information generation system, guiding it to accurately extract the corresponding guidance information from the guidance information generation library, ensuring the timeliness and accuracy of the guidance information.

[0114] In practical applications, the AR-assisted assembly guide generation device further includes: a database system for storing the second knowledge graph; a database system construction process such as Figure 2 As shown, the original data of the assembly components, namely the sample assembly images and parts operation manuals mentioned above, are obtained; the IDEF1X model is obtained based on the IDEF1X process modeling; the IDEF1X model is processed by Neo4j software to construct a Neo4j knowledge graph library; and the Neo4j knowledge graph library is stored to form a database system.

[0115] like Figure 5 As shown in the figure, a database system is pre-built and embedded into the visual language model to obtain a visual language model recognition system. The assembler wears an AR device interaction system and outputs the image to the visual language model recognition system to recognize parts and their status. The recognition results are then input into the guidance information generation system to generate specific guidance information, which is then transmitted to the AR device interaction system to guide the assembler in assembly. The visual language model is specifically: DeepSeek VL2, such as Figure 3 As shown, the image input is embedded into the DeepSeek VL2 database system, the parts and assembly status are identified, and then passed to the guidance information generation system.

[0116] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 7As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store AR-assisted assembly guidance generation data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an AR-assisted assembly guidance generation method is implemented.

[0117] Those skilled in the art will understand that Figure 7 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned method embodiments when executing the computer program.

[0118] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the above-mentioned method embodiments when executed by a processor.

[0119] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.

[0120] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0121] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0122] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0123] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0124] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for generating AR-assisted assembly guidance, characterized in that: The AR-assisted assembly guide generation method includes: Get the assembly image at the current moment; Input the current assembly image into the trained guidance phase recognition model to obtain the assembly parts and assembly status at the current moment; the guidance phase recognition model is obtained by fusing the assembly information knowledge graph embedding vector into the visual large language model; Generate corresponding guidance information based on the current assembly parts and assembly status.

2. The AR-assisted assembly guide generation method according to claim 1, characterized in that: The process of constructing the recognition model in the guidance phase specifically includes: Entity relationships are extracted from the IDEF1X model to obtain the relationships between entities. The IDEF1X model is obtained by processing sample assembly images and part operation manuals using the IDEF1X method. Construct a first knowledge graph based on each entity and the relationship between entities; Taking the assembly state of each part in the IDEF1X model and the image feature vectors of each sample assembly image as entities, a second knowledge graph is constructed based on the first knowledge graph; Process the second knowledge graph using a knowledge graph embedding algorithm to obtain an assembly information knowledge graph embedding vector; The assembly information knowledge graph embedding vector is fused with the intermediate representation layer of the visual language model to obtain the guidance stage recognition model.

3. The AR-assisted assembly guide generation method according to claim 1, characterized in that: After generating corresponding guidance information according to the current assembly parts and assembly status, the method further includes: processing the guidance information according to the assembler's gesture interaction signal, eye interaction signal or voice interaction signal.

4. The AR-assisted assembly guide generation method according to claim 1, characterized in that: The assembly state includes a preparation stage, an assembly stage and a calibration stage; The guidance information corresponding to the preparation stage includes the introduction of assembly parts information; The guidance information corresponding to the assembly stage includes dynamic assembly animation; The guidance information corresponding to the calibration stage includes prompt information and the next part to be assembled.

5. The AR-assisted assembly guide generation method according to claim 1, characterized in that: After generating corresponding guidance information according to the current assembly parts and assembly status, the following steps are also included: The guidance information and the real scene are integrated through the three-dimensional registration function.

6. The AR-assisted assembly guide generation method according to claim 1, characterized in that: Before obtaining the assembly image at the current moment, it also includes: building guidance information in Unity software.

7. An AR-assisted assembly guide generation device, characterized in that: The AR-assisted assembly guide generation device includes: AR device interaction system, used to obtain the assembly image at the current moment; The visual large language model recognition system is used to input the current assembly image into the trained guidance phase recognition model to obtain the current assembly parts and assembly status; the guidance phase recognition model is obtained by fusing the assembly information knowledge graph embedding vector into the visual large language model; The guidance information generation system is used to generate corresponding guidance information according to the assembly parts and assembly status at the current moment.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the AR-assisted assembly guide generation method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the AR-assisted assembly guide generation method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the AR-assisted assembly guide generation method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Aviation knowledge graph-based auxiliary assembly method and system

    CN114723817A

  • Ship pipeline assembly auxiliary method based on AI and knowledge graph

    CN119128175A

  • Embedded intelligent visual language large model knowledge base construction and application method, equipment, medium and product

    CN119476463A

Cited By

  • Equipment operation training auxiliary method and system

    CN121544440A

  • Assembly progress detection method and system based on few-sample vision and RAG

    CN122198576A

  • An assembly progress detection method and system based on few-shot vision and RAG

    CN122198576B