Intelligent assistant enhanced mixed reality remote collaborative assembly system and method

By introducing a mixed reality remote collaborative assembly system enhanced with an intelligent assistant, and utilizing a large language model to perceive and optimize virtual assets in real time, the system solves the problem of insufficient adaptive capability in MR remote collaborative assembly systems, and achieves efficient collaborative assembly of complex industrial products.

CN122492999APending Publication Date: 2026-07-31NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHWESTERN POLYTECHNICAL UNIV
Filing Date
2026-04-15
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing MR remote collaborative assembly systems lack adaptability, have fixed virtual asset representations, and cannot respond to dynamic situational needs, resulting in low information transmission efficiency, high cognitive load, and poor collaborative efficiency in collaborative communication.

Method used

A mixed reality remote collaborative assembly system with intelligent assistant enhancement is introduced. It uses a large language model to perceive multimodal interaction data in real time, generates virtual asset optimization and control strategies that match the current task context, dynamically adjusts the visual representation and spatial layout of virtual assets, and executes the optimization and control strategies through AR and VR clients to achieve adaptive collaborative assembly.

Benefits of technology

It significantly improves the collaborative efficiency of remote assembly of complex industrial products, reduces cognitive load, enables adaptive collaboration for complex assembly tasks, and enhances information transmission efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492999A_ABST
    Figure CN122492999A_ABST
Patent Text Reader

Abstract

This invention discloses a mixed reality remote collaborative assembly system and method enhanced with an intelligent assistant, belonging to the interdisciplinary field of mixed reality, large language model, and industrial assembly technology. The system includes an AR client, a VR client, a server, and an intelligent assistant. The intelligent assistant, driven by a large language model, perceives multimodal interaction data collected by the clients in real time to obtain the current task context. It then combines assembly domain knowledge with the broad domain knowledge of the large language model to perform reasoning, generating virtual asset control strategies and driving each client to execute them. This causes the virtual assets in each client's field of vision to dynamically change in accordance with the current task context. By introducing an intelligent assistant based on a large language model, this invention achieves dynamic adaptive optimization of virtual assets during mixed reality remote collaborative assembly, reducing user cognitive load and significantly improving the information interaction efficiency of remote collaborative assembly of complex products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of mixed reality, large language model and industrial assembly technology, and specifically relates to a mixed reality remote collaborative assembly system and method with intelligent assistant enhancement. Background Technology

[0002] In the industrial manufacturing sector, product assembly is one of the most critical links in the entire production cycle, accounting for a large proportion of time and cost, and directly determining the quality of the final product. Despite the rapid development of advanced manufacturing technologies, for high-precision industries such as aerospace and high-end equipment, the complex and precise products, due to their numerous models, small batches, complex structures, and high precision requirements, are still difficult to fully automate and require significant reliance on manual assembly. However, manual assembly is highly dependent on the skills and experience of the operators, and in actual operation, quality fluctuations often occur due to differences in personnel capabilities, or stagnation occurs due to unexpected problems. Given the current shortage of highly skilled workers, uneven distribution of personnel, and increasing demand for remote collaboration, on-site workers often cannot receive on-site professional assistance and need to seek guidance from remote experts.

[0003] Traditional remote collaboration methods such as voice calls and video conferencing suffer from limited information transmission and communication efficiency, failing to meet the spatial operation explanation and demonstration needs related to assembly tasks. Against this backdrop, Mixed Reality (MR) technology offers an efficient and feasible remote collaboration paradigm: adopting a classic layout of local Augmented Reality (AR) – remote Virtual Reality (VR), it adds 3D virtual assets to the real-world view of local users to enhance information, while providing remote experts with a consistent, immersive virtual space supporting free navigation and perception. This constructs a shared collaborative space that blends the virtual and real worlds, supporting multi-party collaboration across time and space. This space can integrate multimodal assets such as text, images, videos, virtual assets, and audio, supporting interactive methods such as voice commands and 3D virtual control, facilitating spatial information exchange and task advancement. Therefore, assembly operation assistance systems based on MR remote collaboration technology can not only present the required assembly process information to users in a more natural and intuitive way, but also support users to create virtual commands in real time to convey intentions and information. This can improve users' information perception and expression capabilities in the process of asking and answering questions about assembly problems, reduce cognitive load, and improve collaboration efficiency.

[0004] However, existing MR remote collaborative assembly systems largely rely on pre-development for specific task types, preparing the representation and layout of virtual assets such as interactive interfaces and virtual copies in the collaborative space, as well as the presentation of guidance information and instructions during the process. This content is static and singular, unable to adapt to dynamic user needs and task environment constraints as the task context changes. This results in low actual efficiency in the perception, expression, and understanding processes during collaborative information exchange, and may even lead to cognitive biases and additional workload, impacting the efficiency of remote collaborative assembly. In existing work assistance systems supported by Extended Reality (XR) technology, adaptive techniques are still insufficient and difficult to adapt to collaborative assembly tasks: In adaptive triggering mechanisms, they rely on single-dimensional input such as explicit physical actions by users, making it difficult to capture deep semantic information within the context; in decision-making mechanisms, pre-set rules, due to their inherent rigidity, cannot cope with unplanned ambiguity and uncertainty, and traditional deep learning models, with their poor generalization ability, are also unable to adapt to new situations; in terms of adjustment strategies, the types and levels of adjustable elements and executable actions are still in the early stages of exploration, and largely rely on limited pre-set solutions.

[0005] In recent years, the rise of Large Language Models (LLMs) has provided a revolutionary approach to overcoming this predicament. They possess technical characteristics such as handling multimodal input, deep natural language understanding, complex logical reasoning, structured decision generation, and multi-agent driving, enabling them to efficiently and conveniently transform fuzzy human needs in collaboration into executable instructions required for system adaptive optimization. However, current research has only explored the application of LLMs in the industrial field. How to adapt LLMs to remote collaborative assembly tasks using MR and leverage them to enhance the system's adaptive capabilities requires further practical exploration and research.

[0006] In summary, there is an urgent need to introduce a new adaptive mechanism with powerful interactive semantic triggering, dynamic contextual reasoning, and intelligent decision-making capabilities into the MR remote collaboration framework to solve the problems of insufficient adaptive capabilities, fixed virtual assets, and inability to respond to dynamic contextual requirements in existing MR remote collaborative assembly systems.

[0007] In view of this, the present invention is hereby proposed. Summary of the Invention

[0008] The purpose of this invention is to overcome the shortcomings of the prior art and provide a mixed reality remote collaborative assembly system and method with intelligent assistant enhancement. It is mainly used in the remote collaborative assembly process of complex and precision products such as aerospace and defense equipment, so as to solve the problems of insufficient adaptive capability, fixed virtual asset representation, and inability to respond to dynamic situational needs in existing mixed reality remote collaborative assembly systems, resulting in low information transmission efficiency, high cognitive load and poor collaborative efficiency in collaborative communication.

[0009] The objective of this invention is achieved through the following technical solution:

[0010] On one hand, the present invention provides a mixed reality remote collaborative assembly system with intelligent assistant enhancement, comprising:

[0011] The AR client is configured at the assembly site to collect on-site environmental data and user behavior data, and to overlay virtual information to the user.

[0012] The VR client, configured on a remote end, is used to recreate assembly scenes in an immersive virtual space for experts to browse and interact with.

[0013] The server communicates and connects with the AR client and VR client respectively, and stores and manages digital assets and knowledge base;

[0014] A smart assistant communicates with the AR client, VR client, and server. Driven by a large language model, the smart assistant is configured to: perceive multimodal interaction data collected by the AR and VR clients in real time to obtain the current task context; based on the current task context, and combining assembly domain knowledge and the general domain knowledge of the large language model, generate a virtual asset optimization and control strategy that matches the current task context. This optimization and control strategy includes adjusting the visual representation, spatial layout, or information push timing of the virtual assets; convert the optimization and control strategy into executable instructions and drive the AR and / or VR clients to execute them, causing the virtual assets in the client's field of vision to dynamically change in accordance with the current task context, thereby assisting information interaction during remote collaborative assembly.

[0015] Furthermore, the system is logically divided into a data layer, a service layer, and an application layer;

[0016] The data layer is deployed on the server and includes:

[0017] The virtual copy construction module retrieves the original CAD data of the parts from the product data management library based on the current assembly task, and constructs an interactive virtual copy with multi-level semantic features through format conversion, semantic feature classification and annotation.

[0018] The assembly process data processing module retrieves and parses the process documents of the current assembly task from the assembly process knowledge base, extracts structured processes, steps, and graphic information, prepares standardized visual assets that the system can load, and constructs an assembly domain knowledge base for the intelligent assistant to call.

[0019] The collaborative space initialization module retrieves the virtual copy and the standardized visual assets according to the current assembly task, instantiates the space in each client according to the preset layout rules, constructs a virtual-real integrated collaborative work space, and activates the intelligent assistant.

[0020] The service layer is deployed on the server and includes:

[0021] The context awareness module collects and integrates voice, gestures, gaze focus, interaction data and first-person perspective images from various clients in real time to form a standardized context status data package.

[0022] The prompt word generation module matches templates from a preset prompt word template library based on the context state data package and the current assembly task, and generates structured prompt words that drive the large language model to perform reasoning.

[0023] The reasoning and decision-making module, driven by the large language model, receives and parses the structured prompt words, combines them with the assembly domain knowledge base to perform deep reasoning, and generates optimized control strategies for virtual assets.

[0024] The instruction conversion module parses the optimization and control strategy and converts it into control instructions that can be executed by each module of the application layer.

[0025] The communication module is used to establish and maintain data transmission between the data layer, service layer, and application layer, as well as with the cloud-based large language model service;

[0026] The application layer is deployed on the AR client and VR client, including:

[0027] The virtual copy visual optimization module, in response to the control command, performs feature-level visual attribute adjustment on the virtual copy of the target component;

[0028] The information adaptive push module responds to the control command and dynamically adjusts the granularity, modality, and push timing of the guidance information according to the task context and user role.

[0029] The visual asset layout adjustment module responds to the control command, dynamically manages the spatial distribution of virtual information in the user's field of vision, and performs anti-occlusion adjustment and layout optimization.

[0030] The multimodal instruction response module responds to the control instructions and provides audiovisual feedback to the user's interaction requests.

[0031] Furthermore, the virtual copy visual optimization module incorporates an adaptive visual optimization strategy based on information supply and demand matching. This strategy is used to accurately filter the hierarchical information carriers of the virtual copy and dynamically reconstruct its visual presentation based on the dynamic information needs under different collaborative assembly task scenarios.

[0032] The feature enhancement operations performed by the virtual copy visual optimization module include one or more of the following: size enlargement, color highlighting, edge highlighting, center line rendering, centroid trajectory rendering, and model phantom rendering; the feature weakening operations include one or more of the following: feature deletion, transparent rendering, and wireframe rendering.

[0033] Furthermore, the adaptive information push module is configured as follows:

[0034] The information granularity is dynamically determined based on task complexity and user role. High-density system-level information is pushed to expert users, while procedural instructions that are strongly related to the current operation are pushed to worker users.

[0035] Based on cognitive load scenarios, push detailed text lists when cognitive load is low, and perform information simplification and enlarge UI display size when cognitive load is high or during urgent operations; and

[0036] Based on the timing of the push, a step-by-step node trigger strategy is used to push regular guidance information, knowledge supplement information is pushed when the user's gaze is stable or their hand is still, and immediate commands are given the highest push priority.

[0037] Furthermore, the visual asset layout adjustment module is configured as follows:

[0038] Based on the importance of the command, high-priority virtual assets are moved to the center of the user's field of vision, while low-priority assets are moved to the edge of the field of vision.

[0039] Real-time monitoring of object distribution within the user's field of view; using bounding box collision detection to determine if the UI panel obstructs the user's focus object, hand operation area, or copies of key components; and when occlusion is detected, using a blank area search algorithm to move the occluded assets to an unoccluded area; and

[0040] Automatically adjust text reading information to a comfortable viewing distance, and dynamically switch label information between field-following mode and world-anchoring mode based on object attributes.

[0041] Furthermore, the virtual copy construction module performs semantic-level feature classification and annotation on the converted virtual model according to preset feature classification rules. The feature classification rules include geometric features, appearance features and logical features, and automatically attach rendering attributes and physical interaction components.

[0042] Before rendering, the collaborative space initialization module performs a runtime integrity self-check on the loaded virtual assets to confirm that the necessary functional components of mixed reality are fully mounted.

[0043] Furthermore, the system prompt word architecture maintained by the prompt word generation module includes the following modules:

[0044] The task background module is used to build a large language model's global understanding of the task process, collaboration goals, and system architecture.

[0045] The role setting module is used to define the role and responsibilities of the large language model in collaborative tasks;

[0046] The standardized reasoning reference module is used to transform the standardized reasoning and decision-making ideas of domain experts into structured instructions;

[0047] The structured output specification module is used to define standardized data output formats;

[0048] The exception handling and precautions module is used to preset boundary conditions and response strategies;

[0049] The domain knowledge foundation module is used to integrate process documents, component attributes, and visual optimization strategies associated with the current task; and

[0050] The startup instruction module is used to trigger the large language model to enter the working state.

[0051] Furthermore, the system also includes a reliability enhancement mechanism for human-large language model collaboration, which is configured as follows:

[0052] Embedding standardized examples based on thought chains into system prompts guides the large language model to learn expert-level reasoning logic, and mandates that the large language model perform internal validation before outputting the final decision; and

[0053] When the large language model cannot parse the input, the inference result conflicts with the security rules, or there is great uncertainty, an abnormal circuit breaker is triggered, and a problem report is generated to request user intervention.

[0054] Furthermore, it supports users in actively monitoring and intervening in the output scheme of the intelligent assistant, responding to the user's active control commands to perform intent re-identification and reasoning correction, and realizing dynamic error correction and optimization of the human loop.

[0055] Furthermore, the system also includes an assembly quality inspection module, configured as follows:

[0056] In response to the user's trigger command, guide the AR client to collect multi-angle keyframe images of the assembly area;

[0057] The large language model is invoked to perform assembly state reasoning based on visual semantics, and the on-site images are compared with virtual copies of standard assembly states or knowledge base graphs to identify missing parts, incorrect poses, or foreign object residues.

[0058] Augmented reality annotations are generated through a multimodal command response module and combined with a voice broadcast correction scheme.

[0059] On the other hand, the present invention also provides a mixed reality remote collaborative assembly method enhanced by an intelligent assistant, comprising the following steps:

[0060] Step 1: Build an MR remote collaborative work environment

[0061] Step 1.1, Task Scene Recognition: Local workers wear augmented reality devices to enter their workstations and start the AR client. The collaborative space initialization module establishes a connection with the cloud-based large language model, submits scene recognition system prompts to activate the model's task analysis capabilities, responds to the user's interactive scene acquisition command, drives the context perception module to call the camera to capture key frame images on site, encapsulates them into multimodal single-round prompts through the prompt generation module, and sends them to the reasoning and decision-making module. The large language model performs visual understanding and reasoning on the images and outputs the type of the current assembly task and the list of parts involved.

[0062] Step 1.2, Task Data Asset Preparation: The assembly process data processing module receives the task identification results, retrieves the corresponding process documents and part information from the assembly process knowledge base, parses and prepares the graphic guidance information required for the task and the assembly domain knowledge base for the intelligent assistant to access; the virtual copy construction module retrieves the original CAD model from the product data management library based on the identified parts list, performs format conversion, semantic-level feature classification and annotation, and MR function additional processing to generate a structured virtual copy asset adapted to the MR system;

[0063] Step 1.3, Collaborative Space Instantiation and Intelligent Assistant Initialization: The collaborative space initialization module drives the AR client and VR client, loads virtual copies and graphic assets from the server according to predefined spatial layout rules, completes the anchoring and layout adjustment of virtual objects in the virtual and real spaces, and constructs a collaborative work space that integrates the virtual and real worlds; the prompt word generation module constructs the collaborative assembly intelligent assistant system prompt words for the current task stage, sends an initialization request through the communication module, and activates the intelligent assistant's analysis and decision-making functions;

[0064] Step 2: Collaboration between local workers and remote experts

[0065] Step 2.1, Local Operations and Collaboration Requests: When local workers encounter operational difficulties or decision-making bottlenecks while performing physical assembly tasks on the assembly site, they initiate collaboration requests to experts through the MR system. They describe the problem through voice and use gestures to control virtual replicas for auxiliary explanations. The context awareness module monitors and collects the worker's voice stream, gaze point, gestures, interactive object IDs, and first-person perspective images in real time, and uploads the multimodal data to the server.

[0066] Step 2.2, Immersive Remote Expert Guidance: Remote experts monitor the assembly site in real time via video stream on the VR terminal. When they detect worker errors or receive questions, they formulate guidance plans based on the status of the virtual copy synchronized with the system, worker gestures, and language. They demonstrate the correct assembly actions by manipulating the virtual copy and provide explanations of key elements with voice. The context awareness module synchronously collects and uploads the expert's voice commands, controller operation trajectory, and virtual interaction behavior data.

[0067] Step 3: Dynamic assistance from the intelligent assistant

[0068] Step 3.1 Real-time monitoring of user needs: When the context awareness module detects the communication behavior of the two collaborating parties, it triggers data upload. After receiving the data, the prompt word generation module performs the initial screening of the task stage and retrieves the corresponding prompt word template. It serializes the multimodal context data and fills it into the template slot. Combined with the retrieved knowledge fragments, it generates a single-round structured prompt word and sends it to the large language model interface.

[0069] Step 3.2, Contextual Understanding and Intent Reasoning: After receiving the prompt words, the reasoning and decision-making module performs data integrity and relevance checks in accordance with the system's preset intelligent assistant role responsibilities and task requirements. It uses the semantic understanding and visual analysis capabilities of the large language model to extract the collaborative sub-task type, user focus object and task goal from the contextual data, and analyzes the user's core collaborative intent and its implicit information type needs.

[0070] Step 3.3, Virtual Asset Optimization Scheme Decision: The reasoning and decision module supplements knowledge and expert strategies based on the assembly task in the system prompts, and conducts in-depth reasoning on virtual asset optimization schemes oriented towards task context and user needs, referring to the logic of the thinking chain and reasoning cases. The decision content covers the volume, carrier form and timing of the push of guiding information, the visual optimization representation of virtual copies, and the multimodal response form of user active commands, generating a structured decision scheme containing analysis results, execution objects, strategies and parameters.

[0071] Step 3.4, Adaptive Collaboration Scheme Execution: The instruction conversion module receives the structured decision scheme, parses it into control instructions executable by each client through the instruction mapping mechanism, and distributes them to the application layer modules for execution; the information adaptive push module selects suitable graphic and text asset carriers for guidance information based on the scheme and pushes them to the user's field of vision according to the decision timing; the virtual copy visual optimization module adjusts the visual attributes at the feature level of the virtual copy of the target component based on the optimization parameters in the scheme; the multimodal instruction response module generates a voice broadcast stream and coordinates with changes in visual assets to achieve adaptive matching of user collaborative communication needs and real-time response to proactive instructions;

[0072] Step 4: Efficient Collaboration Loop Supported by Intelligent Assistant

[0073] Step 4.1, Intelligent Assistant Proactive Response: Based on the dynamic assistance of the intelligent assistant, the visual appearance of the virtual copy in the collaborative space and the content of the guiding information adaptively match the user's needs. Through the audiovisual multimodal behavior of highlighting key features and weakening interfering information, the mapping and synchronization of communication information and presentation are realized, thereby achieving a proactive response to the user's implicit information expression and understanding needs.

[0074] Step 4.2, User-initiated interaction and intervention: When experts or workers find that the adaptation solution for virtual assets does not meet expectations or there are additional control requirements, they can issue an active intervention command through a combination of voice commands and gestures, or directly initiate a general knowledge Q&A session with the intelligent assistant; the multimodal command response module will instantly parse the active request and forcibly adjust the specified virtual assets or generate an audiovisual knowledge response.

[0075] Step 4.3, Implicit Control of Virtual Asset Layout: During the collaboration process, the context awareness module continuously monitors the user's visual attention area and operation behavior in the workspace. The visual asset layout adjustment module uses ray detection and bounding box collision algorithms to determine in real time whether the information panel obstructs the key operation area. When a conflict is detected, the interference items are automatically and smoothly moved to the free area outside the field of vision, and the core information is moved forward to the visual focus area.

[0076] Step 4.4, Intelligent Assembly Quality Inspection: After completing a specific process or overall task, in response to the user's trigger command, the system starts the quality inspection process. The context awareness module guides the user to collect keyframes of the assembly area from multiple angles through AR devices and upload the images. The reasoning and decision-making module calls the large language model to perform assembly state reasoning based on visual semantics, compares the features of the on-site real-time images with virtual copies of standard assembly states or knowledge base graphs, identifies missing parts, incorrect poses, or foreign object residues, and generates augmented reality annotations in the AR client's field of vision through the multimodal command response module, along with voice broadcasting correction schemes.

[0077] Compared with the prior art, the present invention has the following beneficial effects:

[0078] 1. This invention creatively introduces an intelligent assistant centered on a multilingual model, constructing an adaptive collaborative space that blends the virtual and real worlds. This space can dynamically adjust the visual representation, spatial layout, and information delivery of virtual assets according to the real-time task context, thereby greatly enhancing the expression and understanding efficiency of non-verbal cues in 3D assembly tasks and significantly reducing the cognitive load on both collaborating parties (local workers and remote experts). Through a complete end-to-end process encompassing "environment construction—collaborative operation—dynamic assistance—closed-loop feedback," this invention achieves adaptive collaboration across the entire remote assembly process of complex industrial products, significantly improving the efficiency and operational experience of cross-temporal and spatial collaboration, and providing a systematic solution for the assembly of high-precision equipment.

[0079] 2. This invention constructs an automated content generation and space initialization mechanism for new assembly tasks through the collaboration of a data layer (virtual copy construction module, assembly process data processing module, and collaborative space initialization module). This mechanism can parse heterogeneous process documents in real time, dynamically construct interactive virtual copies, and complete the instantiation and layout of the collaborative space according to preset semantic rules. This realizes a paradigm shift from "static prefabrication" to "dynamic generation," greatly reducing the threshold and deployment cycle of mixed reality content production, and ensuring the system's rapid migration capability for different products and tasks.

[0080] 3. This invention constructs an intelligent interaction optimization system at the application layer: the virtual copy visual optimization module establishes a mapping rule of "task - information demand - feature operation" based on the principle of information supply and demand adaptation, realizing precise control of the visual representation density of part features; the visual asset layout adjustment module constructs a dynamic anti-occlusion mechanism through real-time field-of-view analysis and collision detection algorithms. The above design effectively ensures the user's perceptual focus and interaction smoothness during long-term collaboration, solving the inherent problems of virtual information interference and cognitive overload in traditional MR systems.

[0081] 4. To address the potential decision-making uncertainties and unreliable output risks when general-purpose large language models are directly applied to MR remote collaborative assembly systems in professional scenarios, this invention innovatively designs a reliability enhancement mechanism for human-large language model collaboration. This mechanism constructs inherent stability constraints for the intelligent assistant oriented towards assembly tasks by embedding thought chain demonstrations, mandatory self-checks, and anomaly circuit breaking into the system's prompt word architecture. Simultaneously, a dynamic error correction mechanism based on the human-centric loop is introduced, integrating higher-order judgments from remote experts and local workers into the system's decision-making loop. This allows the system to proactively request user intervention when encountering uncertainties and allows users to instantly correct decision-making schemes through natural interaction methods such as voice and gestures. This design ensures the stability and effectiveness of the intelligent assistant's output decisions in complex industrial assembly scenarios, achieving an overall improvement in system-level reliability.

[0082] 5. This invention integrates a visual semantic-based intelligent assembly quality detection function, forming a complete quality closed loop from work instruction to result verification. Responding to user commands, this function uses AR devices to capture on-site assembly images and calls upon a large language model for visual semantic reasoning. It compares the actual images with standard assembly states to accurately identify defects such as missing parts and incorrect posture, and instantly generates augmented reality annotations and voice correction solutions. This closed-loop control mechanism provides reliable quality and safety assurance for remote collaborative assembly processes, ensuring the accuracy and reliability of assembly results. Attached Figure Description

[0083] The accompanying drawings are incorporated in and form part of this specification, and together with the description serve to explain the principles of the invention.

[0084] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0085] Figure 1 A schematic diagram of the overall framework of the intelligent assistant-enhanced mixed reality remote collaborative assembly system provided by the present invention;

[0086] Figure 2 Functional module architecture diagram of the intelligent assistant-enhanced mixed reality remote collaborative assembly system provided by the present invention;

[0087] Figure 3 A flowchart illustrating the workflow of the virtual copy construction module provided by this invention;

[0088] Figure 4 A flowchart of the assembly process data processing module provided by the present invention;

[0089] Figure 5 A flowchart illustrating the workflow of the collaborative space initialization module provided by this invention;

[0090] Figure 6 A flowchart illustrating the workflow of the context-aware module provided by this invention;

[0091] Figure 7 A flowchart illustrating the workflow of the prompt word generation module provided by this invention;

[0092] Figure 8 A flowchart of the reasoning and decision-making module provided by this invention;

[0093] Figure 9 A flowchart illustrating the workflow of the instruction conversion module provided by this invention;

[0094] Figure 10 A flowchart illustrating the operation of the communication module provided by this invention;

[0095] Figure 11 A flowchart illustrating the workflow of the virtual copy visual optimization module provided by this invention;

[0096] Figure 12 A flowchart illustrating the workflow of the visual asset layout adjustment module provided by this invention;

[0097] Figure 13 A flowchart illustrating the workflow of the adaptive information push module provided by this invention;

[0098] Figure 14 A flowchart illustrating the workflow of the multimodal command response module provided by this invention;

[0099] Figure 15 A schematic diagram of the system prompt word architecture for the intelligent assistant for MR collaborative assembly provided by the present invention;

[0100] Figure 16 A schematic diagram illustrating the reliability enhancement mechanism for the collaboration between the human and large language models provided by this invention;

[0101] Figure 17 This is a schematic diagram of a virtual replica adaptive visual optimization strategy based on the supply and demand matching concept provided by the present invention. Detailed Implementation

[0102] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples consistent with some aspects of the invention as detailed in the appended claims.

[0103] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0104] Example 1

[0105] Please see Figures 1 to 17This embodiment provides a mixed reality remote collaborative assembly system with intelligent assistant enhancement. The system mainly consists of four core parts: AR client (AR local client), VR client (VR remote expert client), server, and intelligent assistant. The system comprises several components: an AR client deployed at the assembly site to collect environmental and user behavior data, overlaying virtual information onto the user to assist local workers in completing assembly tasks; a VR client deployed remotely to recreate the assembly scene in an immersive virtual space, synchronizing the assembly status and providing remote experts with scene browsing, virtual operation demonstrations, and interactive guidance; a server communicating with both the AR and VR clients to establish data transmission and synchronization channels, while storing the digital assets (including virtual models of parts, process diagrams, etc.) and assembly domain knowledge base required by the management system, providing data support for system operation; and an intelligent assistant communicating with the AR, VR, and server, driven by a large language model, serving as the system's intelligent hub, coordinating the collaborative work of various components, enabling perception, understanding, reasoning, and decision-making regarding collaborative interaction content, driving the adaptive optimization of virtual assets (including 3D virtual models of parts overlaid on the real scene, assembly guidance diagrams, interactive operation interface panels, or combinations thereof), ensuring efficient and smooth execution of remote collaborative assembly.

[0106] Specifically, the AR client uses mixed reality devices, such as HoloLens 2 AR glasses, which mainly have the following functions: First, as a perception terminal, it uses integrated cameras, microphones, inertial measurement units (IMU), eye-tracking sensors, and gesture recognition systems to collect in real time the voice commands, gestures, gaze points, first-person perspective environmental images, and interaction events between workers and the virtual interface (UI) of local workers, forming multimodal interactive data describing the on-site situation; Second, as a virtual-real fusion presentation terminal, it accurately overlays, renders in real time, and dynamically adjusts virtual information such as 3D component virtual models, standardized assembly guidance graphics, and interactive operation panels in the local worker's real field of vision according to the control commands issued by the server, completing the enhanced display of on-site visual information.

[0107] The VR client uses a virtual reality system, such as the HTC Vive Pro 2 headset. Its core function is to build an immersive virtual assembly environment on the expert's end that strictly matches the layout of the on-site space. The expert gains a sense of immersion through the VR headset and uses the controllers for free navigation and precise interaction. The VR client receives synchronized on-site data from the server (such as the virtual scene status and worker operation video streams), allowing the expert to observe as if "on-site". At the same time, the VR client also has a behavioral data collection function, capturing and uploading in real time the expert's voice guidance content, controller spatial control trajectory, and refined interactive behavioral data such as grasping, rotating, aligning, and placing virtual models, providing raw data for backend intelligent analysis.

[0108] The server, serving as the system's data and communication hub, is deployed in the cloud or on a local area network. It runs a database management system and file services, centrally storing and managing all digital assets required for system operation, including: structured 3D virtual models generated by the data layer, standardized process drawings and documents, assembly domain knowledge bases, historical case data, etc. More importantly, the server, through its communication module based on network protocols such as TCP / IP, establishes and maintains data interaction, status synchronization, and reliable command distribution among the AR client, VR client, and intelligent assistant, ensuring unified timing and consistent data across multiple terminals.

[0109] The intelligent assistant is the core intelligent entity of this invention. It can be deployed on a dedicated local computing power node interconnected with the server, or remotely call the cloud API service of a compliant large language model. The intelligent assistant uses a large language model with multimodal semantic understanding, visual association analysis, and complex chain logic reasoning capabilities as its core computing power brain. It has a complete intelligent closed-loop process: First, it receives and integrates multimodal interaction data uploaded from both AR / VR terminals in real time, completes time sequence alignment and context fusion analysis, and deeply identifies the current assembly stage, operational bottlenecks, and the interaction intentions of both parties; Second, it retrieves the professional knowledge base and standard process specifications of the assembly field on the server side, and combines the general logic of the large language model itself with industry experience to carry out anthropomorphic deep chain reasoning; Finally, it outputs results that fit the real-time operation conditions. The virtual asset adaptive adjustment and optimization strategy includes: enhancing / weakening multi-level visual features of component virtual models, dynamically correcting the unobstructed spatial layout of UI information panels, and precisely controlling the granularity and timing of guide-type graphic and text information push. Finally, the abstract optimization strategy is standardized and parsed, and converted into rendering driving instructions and interactive control instructions that can be directly recognized and executed by AR and VR clients. This drives all virtual and real assets in the dual-end field of view to dynamically and adaptively adjust according to the collaborative process, comprehensively enhancing the efficiency of cross-end information transmission between local workers and remote experts, and realizing intelligent closed-loop collaborative assembly assistance.

[0110] To more clearly reveal how the system works collaboratively internally to achieve the aforementioned intelligence, Figure 2 The system's functional module architecture is demonstrated. Logically, this architecture is divided into a data layer, a service layer, and an application layer. These three layers work closely together to transform static data into dynamic intelligent services. Specifically, the data layer, serving as the system's foundational resource base, is deployed on the server and includes a virtual replica construction module, an assembly process data processing module, and a collaborative space initialization module. The service layer, as the system's core logical hub, is also deployed on the server and includes a context-aware module, a prompt word generation module, a reasoning and decision-making module, an instruction conversion module, and a communication module. The application layer, serving as the system's user interaction and execution terminal, is deployed on AR and VR clients and includes a virtual replica visual optimization module, an adaptive information push module, a visual asset layout adjustment module, and a multimodal instruction response module.

[0111] like Figure 3 As shown, the virtual copy construction module retrieves the original CAD data of the parts from the product data management database based on the current assembly task. After format conversion, semantic-level feature classification and annotation, it constructs an interactive virtual copy with multi-level semantic features. Specifically, based on the current assembly task identified by the reasoning and decision-making module, the module retrieves the original CAD data of the parts to be assembled corresponding to the task from the product data management database. Then, it calls an automated conversion script developed based on the Pixyz plugin to convert the original parametric model into a Unity real-time rendering compatible format, while preserving the geometric topology hierarchy of the model. Furthermore, it performs semantic-level feature classification and annotation on the model according to preset feature classification rules. The feature classification rules include geometric features (such as shape features, positioning features, assembly features), appearance features (such as color, scale, material, transparency), and logical features (such as implicit cognitive concepts derived from physical entities, used to represent phenomena, constraints, rules, or intentions in the assembly task). It also automatically attaches rendering attributes and physical interaction components (such as colliders and grab points), ultimately constructing a structured virtual copy that supports real-time rendering and multimodal interaction.

[0112] like Figure 4As shown, the assembly process data processing module retrieves and parses the process documents for the current assembly task from the assembly process knowledge base, extracts structured processes, steps, and graphic information, prepares standardized visual assets that the system can load, and constructs an assembly domain knowledge base for the intelligent assistant to access. Specifically, this module retrieves and calls up the original process documents (such as PDF or Word format) corresponding to the current assembly task from the enterprise assembly process knowledge base, as well as the associated component technical parameter tables and equipment operation manuals; subsequently, using natural language processing and document layout analysis technology, it performs automatic information extraction: identifying and segmenting processes, steps, and action levels, extracting key text information such as operation instructions, related part descriptions, equipment precautions, and safety warnings; separating the schematic diagrams or technical drawings corresponding to the steps from the documents, and maintaining their association with the text descriptions. Next, data assetization processing is performed: the extracted text and images are converted into standardized visual assets that can be directly loaded by the Unity client (such as JSON format process description text files and Texture2D format images), and stored in the server file system for real-time access; the above structured knowledge is stored in the server database as an external knowledge source for the subsequent intelligent assistant to obtain task background, core process flow and domain expertise by prompting the project.

[0113] like Figure 5 As shown, the collaborative space initialization module retrieves virtual copies and standardized visual assets based on the current assembly task, instantiates the space in each client according to preset layout rules, constructs a virtual-real integrated collaborative work space, and activates the intelligent assistant. Specifically, this module first initializes the intelligent assistant service instance, establishes a network communication connection with the large language model, and injects system prompts to activate the assistant's initialization phase role settings and responsibilities. Secondly, based on the identified assembly task context, it retrieves virtual copies of the parts to be assembled and process graphic assets prepared by the previous module from the server resource library, and instantiates the scene according to preset spatial layout rules; subsequently, it loads the process graphic data and fills it into the client's preset UI panel template to generate the initial state of the work instruction; finally, before rendering, this module performs a runtime integrity self-check on the loaded virtual assets, confirming that the necessary mixed reality functional components such as colliders, interaction scripts, and rendering materials are fully mounted, ensuring that the system as a whole is in a ready state.

[0114] like Figure 6As shown, the context-aware module collects and integrates voice, gestures, gaze focus, interaction data, and first-person perspective images from various clients in real time to form a standardized context state data package. Specifically, this module calls the sensor interfaces integrated into the user's mixed reality devices (including HoloLens 2 AR glasses and HTC Vive Pro 2 VR headset) to perform data collection tasks: First, it collects the voice dialogue stream during the collaboration process through a microphone array and converts it into text data in real time via Automatic Speech Recognition (ASR); simultaneously, it obtains the screen coordinates and dwell time of the user's gaze point through an eye-tracking sensor and identifies the virtual asset ID at the gaze focus point using a ray detection algorithm; second, it uses computer vision gesture tracking on the AR side and inertial sensor tracking technology on the VR side to obtain the user's hand movement state in real time and recognize gesture commands; in addition, it uses a virtual camera that follows the user's perspective to synchronously capture first-person field-of-view images when interactive behaviors occur, and uses an event listening mechanism to record the interaction trigger data between the user and the UI interface or virtual copy. Ultimately, based on the interaction rounds triggered by voice dialogue, the module timestamps and fuses the aforementioned modal context-aware data, encapsulating it into a standardized context state data package, which serves as the basic input for subsequent reasoning and decision-making.

[0115] like Figure 7 As shown, the prompt word generation module matches templates from a pre-set prompt word template library based on the context state data package and the current assembly task, generating structured prompt words to drive the large language model for reasoning. This module internally maintains a prompt word template library customized for mixed reality remote collaborative assembly tasks, which includes system prompts for defining model behavior benchmarks and single-turn prompts for driving specific interactions. System prompts include templates designed for different stages of the task, such as system initial startup, building the collaborative workspace, and starting the collaborative assembly task. The system will automatically trigger these prompts at different stages of startup, while local workers can specify their responsibilities through a dedicated UI interface. During runtime, upon receiving a contextual state data packet, the module first performs multimodal data serialization processing on the input speech text, sensor data, and field-of-view images, converting them into an encoding format compatible with the model interface. Subsequently, based on the current task stage and interaction intent characteristics, it performs dynamic template matching, selecting the most suitable single-round prompt template for the current context from the library, and accurately filling the serialized real-time dynamic contextual data into the context slots corresponding to the template. In addition, the module synchronously calls the knowledge base interface to retrieve process parameters, expert strategies, and historical cases related to the current assembly step, splicing them into the prompt word sequence as supplementary knowledge context, and finally generating structured prompt words containing complete multimodal contextual information.

[0116] like Figure 8As shown, the reasoning and decision-making module uses a large language model as its core driver. It receives and parses structured prompts, combines them with an assembly domain knowledge base for deep reasoning, and generates optimized control strategies for virtual assets. Specifically, this component uses a cloud-deployed multimodal large language model service (such as Gemini or DeepSeek) as its core reasoning engine, and relies on a communication module to connect the structured prompts generated by the prompt generation module with its API interface. Figure 15 As shown, to support the application of a general-purpose large language model to the professional field of mixed reality remote collaborative assembly, this invention proposes a system prompt word architecture for a smart assistant in mixed reality collaborative assembly. This architecture consists of seven modules: a task background module, used to establish the large language model's global understanding of the task process, collaborative goals, and overall system architecture; a role setting module, used to define the large language model's role and responsibilities in collaborative tasks; a standardized reasoning reference module, used to transform the standardized reasoning and decision-making ideas of domain experts into structured instructions understandable by the large language model; a structured output specification module, used to define standardized data output formats; an exception handling and precautions module, used to preset boundary conditions and response strategies; a domain knowledge foundation module, used to integrate process documents, component attributes, and visual optimization strategies associated with the current task; and a start command module, used to trigger the large language model to enter the working state.

[0117] like Figure 16As shown, this invention also designs a reliability enhancement mechanism for human-large language model collaboration to further improve the output stability and effectiveness of the large language model in the reasoning and decision-making process. This mechanism combines the internal canonical self-checking of the large language model with the high-order active supervision of human users to construct a rigorous output verification and error correction closed loop. On the one hand, the system prompt word architecture incorporates inherent reliability-oriented constraints: 1) Standardized demonstration based on chain-of-thought, introducing standard decision cases that include multimodal input, question instructions, step-by-step reasoning processes, and structured outputs, guiding the large language model to learn and imitate expert-level reasoning logic through contextual learning; 2) Mandatory self-checking based on role settings, requiring the large language model to perform internal verification before outputting the final decision solution, verifying the compliance of the output data format and the authenticity of the content; 3) Anomaly fallback and help-seeking mechanism to prevent hallucinations, clearly setting boundary conditions in "Anomaly Handling and Precautions," prohibiting the model from arbitrarily fabricating (fabricating) when the large language model encounters unparseable input, reasoning results that conflict with safety rules, or significant uncertainties, instead triggering "anomaly circuit breaker," directly generating a problem report on the mixed reality human-computer interaction interface, and actively requesting user intervention. On the other hand, a dynamic error correction and optimization mechanism based on Human-in-the-Loop (HITL) is constructed. Under this mechanism, humans are not only the service recipients of the intelligent assistant, but also its supervisors and collaborators: 1) The system passively seeks help and humans assist. When the large language model actively issues a request for assistance, users can correct, answer questions, or supplement missing contextual information through voice or gestures to help the system overcome reasoning obstacles; 2) Users actively intervene and the model iteratively corrects. During normal collaboration, users can monitor the output scheme of the intelligent assistant at any time. If deviations are found in visual optimization or information push, users can issue active control commands to intervene. After receiving the error correction command, the large language model immediately performs intent re-identification and reasoning correction, thereby achieving rapid iterative optimization of decision-making effects within a single task.

[0118] During the reasoning process, the large language model first utilizes its broad knowledge understanding capabilities across various domains, referring to the role settings and task requirements in the prompts, to perform in-depth analysis of the input multimodal data, accurately identifying the user's core collaborative intent (such as information query, operation confirmation, or visual assistance request) and related environmental situations (such as the object of gaze, the object of gesture operation, etc.). Subsequently, the model performs human-like deep reasoning based on the task background, expert decision rules, and pre-set thought chain logic. In this process, the model does not mechanically execute pre-set instructions, but rather, while maintaining the correctness of the expert guidance logic, it utilizes its emergent ability to consider the current dynamic situation and perform global adaptive optimization, thereby generating flexible feedback strategies that go beyond fixed rules. Finally, this component outputs a virtual asset optimization and layout scheme that meets the needs of the current situation, covering the load and modal form of guidance information, the visual attribute optimization parameters of the virtual copy, and the anti-occlusion layout coordinates of the visual assets in the field of view, and encapsulates the above decision scheme into a structured data format that is easy to parse later.

[0119] like Figure 9 As shown, the instruction conversion module parses and converts optimized control strategies into control instructions executable by various application layer modules. This module aims to solve the data protocol adaptation and format compatibility issues between heterogeneous modules in the system, achieving seamless connection between the inference decision-making end and the execution drive end. Specifically, for the prompt word transmission and model invocation process, this module performs a multimodal request encapsulation task: serializing structured prompt words containing multimodal data such as text and images into a data carrier conforming to the large language model API interface specification, and automatically configuring the authentication request header, inference parameters, and multimodal payload. For the decision-making process and client-driven operations, this module uses an automated instruction parsing script to decode response data and instantiate instructions. First, it receives and verifies the integrity of the structured response data output from the inference decision-making end, parses it into system-readable parameter objects, and extracts core decision fields. Then, based on a preset instruction mapping mechanism, it embeds the subject identifier, action semantics, and asset parameters from the core fields into a pre-built executable script in the VR / AR client. Simultaneously, it schedules visual assets such as images and videos to be transmitted from the server to the resource preloading area of ​​the client rendering system, converts voice feedback content into voice synthesis requests, and finally aggregates them to generate a system control instruction set that can directly drive the collaborative operation of the application layer.

[0120] like Figure 10As shown, the communication module is used to establish and maintain data transmission between the data layer, service layer, application layer, and the cloud-based large language model service. Specifically, based on the TCP / IP protocol, this module connects the VR expert client, AR worker client, and intelligent assistant with the server as the center. On the one hand, it is responsible for uploading context-aware data to the server and synchronizing direct interaction data such as voice and gestures, as well as the status of virtual assets (such as model pose and UI status) to the collaborating partner in real time, ensuring that remote experts and local workers have a consistent visual, auditory, and sensory experience in the virtual and real scenes and the collaborative communication experience of the real-time processes within them. On the other hand, it is responsible for data transmission between the system and the intelligent assistant, sending large language model API call requests to the cloud and receiving its processing results, and then distributing the parsed solution execution code to the corresponding client. In addition, this module is also responsible for data resource scheduling, supporting the client to request and load the required virtual asset files from the server on demand during the initialization phase and interaction process.

[0121] like Figure 11 As shown, the virtual copy visual optimization module responds to control commands and adjusts the visual attributes at the feature level of the virtual copy of the target component. For example... Figure 17As shown, this module has a built-in adaptive visual optimization strategy based on information supply and demand matching. It is used to accurately select the hierarchical information carrier of the virtual copy and dynamically reconstruct its visual presentation according to the dynamic information needs under different collaborative assembly task scenarios. Specifically, the system decouples the mixed reality collaborative assembly task into five typical sub-tasks (part inspection, part function explanation, assembly sequence formulation, assembly operation explanation, and assembly relationship confirmation), and summarizes the information needs associated with the task into four core information categories (part spatial location, part structural identification, assembly sequence, and assembly direction). Then, it constructs the mapping rules between sub-tasks and information types to quickly define the information demand set under specific task scenarios. Based on the above requirements, this module expands the concept of industrial part features into multi-level information carriers in three-dimensional space, specifically divided into three dimensions: (1) geometric features: including shape features, positioning features, assembly features, and processing features; (2) appearance features: covering visual attribute dimensions such as color, scale, material, and transparency; (3) logical features: implicit cognitive concepts derived from physical entities, used to characterize phenomena, constraints, rules, or intentions in assembly tasks. At the specific visual rendering control level, this module is configured with two types of optimization execution methods: one is feature enhancement, including size enlargement, color highlighting, edge highlighting, centerline rendering, centroid trajectory rendering, and model phantom rendering, used to highlight target requirement information; the other is feature weakening, including feature deletion, transparent rendering, and wireframe rendering, used to suppress redundant information interference. Ultimately, the above task mapping relationship, feature hierarchy definition, and optimization execution methods together constitute a set of built-in visual optimization guidelines (expert rule base) for the system. In actual collaborative processes, the intelligent assistant framework will strictly refer to these guidelines, reason about contextual requirements, and drive this module to automatically output corresponding visual optimization parameters and rendering instructions, thereby transforming the innovative adaptive visualization concept into an automated and executable system closed loop.

[0122] Specifically, this module responds to executable asset loading and rendering control commands issued by the instruction conversion, and modifies the visual attributes of target parts and their sub-feature objects in real time in the graphics rendering pipeline: for parts features marked as "cognitive focus", the module visually highlights them by switching the material to a highlight mode and overlaying contour lights; for features marked as "redundant interference", the rendering mode is dynamically switched to semi-transparent or only wireframe is displayed, so as to eliminate visual interference to the greatest extent while preserving the integrity of the overall spatial structure and reducing the user's cognitive load; at the same time, the module visually instantiates logical auxiliary objects (such as indicator arrows, center axes, and assembly trajectories) defined in the instructions that are not visible in physical entities but are crucial to the understanding of assembly, thereby achieving multi-dimensional visual enhancement covering "geometry-appearance-logic".

[0123] like Figure 12As shown, the visual asset layout adjustment module responds to control commands, dynamically manages the spatial distribution of virtual information in the user's field of vision, and performs anti-occlusion adjustments and layout optimizations. Specifically, based on frustum detection and ray detection technology, this module prioritizes high-priority virtual copies and explanatory graphics to the center of the user's field of vision during expert guidance or critical operation phases, while smoothly moving low-priority system UI panels and irrelevant assets to the edge of the field of vision. It also ensures the spatial stability of the currently manipulated model to avoid visual jumps. When the user understands the instructions or explores independently in the collaborative space, the module monitors the distribution of objects within the user's field of vision in real time. It uses a bounding box collision detection algorithm to determine whether floating UI panels obstruct the focus of the view, the hand operation area, or copies of key components. Once an obstruction is detected, the module uses a blank area search algorithm to calculate the optimal unobstructed coordinates, smoothly moving the obstructing assets to an empty area while maintaining the spatial anchoring position of assembled components or bases. Furthermore, this module optimizes the spatial anchoring of visual assets, including automatically adjusting text-based information to a comfortable viewing distance and dynamically switching label information between "field-following mode" and "world anchoring mode" based on object attributes, ensuring that key information is always accurately attached to the corresponding physical or virtual component surface.

[0124] like Figure 13 As shown, the adaptive information push module responds to control commands and dynamically adjusts the granularity, modality, and timing of information presentation based on the task context and user role. Specifically, based on the task complexity and user role attributes analyzed by the reasoning and decision-making component, the module dynamically determines the granularity, modality, and timing of information intervention: Regarding information granularity and modality selection, for expert users, the module pushes high-density, system-level schematic diagrams and complete technical parameters to support their global decision-making and knowledge retrieval; for worker users, it extracts only procedural instructions specified by the expert and strongly relevant to the current operation, shielding irrelevant background knowledge to focus on execution. In terms of cognitive load adaptation, in low cognitive load scenarios, the module pushes detailed text lists and auxiliary instructions; while in high cognitive load or emergency operation scenarios, the module performs information simplification, eliminating redundant text, retaining only core graphical symbols and key parameter values, and automatically enlarging the UI display size to achieve dynamic matching between information supply and cognitive surplus. Regarding the timing of push notifications, a step-by-step triggering strategy is adopted for regular guidance information, meaning that the next instruction is only pushed after the current physical action is detected to be completed, thus avoiding information accumulation. For knowledge supplement information, it is only presented when the user's gaze is stable or their hand is still. For immediate instructions such as safety warnings and error corrections issued by remote experts or generated by the system, the highest push priority is adopted to ensure that they are perceived as soon as possible.

[0125] like Figure 14As shown, the multimodal command response module responds to control commands and provides audiovisual feedback to user interaction requests. Specifically, this module has two core response mechanisms: First, it responds to user-initiated control commands: allowing users to proactively adjust the appearance attributes (such as color and transparency) and spatial layout of visual assets through voice commands or UI interaction, giving users autonomous control over the virtual environment. Second, it executes intelligent question answering: supporting responses to knowledge-based questions across a wide range of fields, including assembly process queries, tool usage method searches, and general information inquiries; after receiving a response command from the command conversion component, this module, on the one hand, calls the speech synthesis engine to convert the text content into a natural speech stream for playback; on the other hand, it simultaneously retrieves and renders auxiliary graphic descriptions or 3D models strongly related to the answer content, promoting users' rapid understanding of complex information through a multimodal information presentation method that synchronizes voice explanation and visual display.

[0126] In addition, this system also includes an assembly quality inspection module, the functions of which can be integrated into the application layer or used as an independent functional unit. Specifically, this module responds to user trigger commands, guiding the AR client to collect multi-angle keyframe images of the assembly area; through the inference decision module of the service layer, it calls a large language model to perform assembly state inference based on visual semantics, comparing the on-site images with virtual copies of standard assembly states or knowledge base graphs to identify missing parts, incorrect poses, or foreign object residues; finally, through the multimodal command response module of the application layer, it generates augmented reality annotations and coordinates with voice broadcast correction schemes.

[0127] Example 2

[0128] This embodiment provides a mixed reality remote collaborative assembly method based on the system described in Embodiment 1, which is enhanced by an intelligent assistant. This method can be applied to the remote collaborative assembly process of complex industrial products such as aerospace products. The specific steps are as follows:

[0129] Step 1: Build an MR remote collaborative work environment

[0130] Step 1.1, Task Scene Recognition: Local workers wearing augmented reality devices (such as HoloLens 2) enter their workstations and activate the AR client. The collaborative space initialization module establishes a connection with the cloud-based large language model, submits scene recognition system prompts to activate the model's task analysis capabilities, and responds to the user's interactive scene acquisition commands (voice, gestures, or UI buttons). This drives the context awareness module to call the camera to capture keyframe images of the scene. These images are then encapsulated into multimodal single-round prompts by the prompt generation module and sent to the inference and decision-making module via the instruction conversion module. The large language model performs visual understanding and inference on the images, outputting and parsing the type of the current assembly task and the list of involved parts.

[0131] Step 1.2, Task Data Asset Preparation: The assembly process data processing module receives the task identification results and retrieves the corresponding process documents, part data, and other data foundations from the enterprise assembly process knowledge base. It then parses and prepares the graphic guidance information required for the task and the domain knowledge base for use by the intelligent assistant. Simultaneously, the virtual copy construction module, based on the identified parts list, retrieves the original CAD model from the product data management system and performs processing steps such as format conversion, semantic-level feature grading and annotation, and MR function addition to generate a structured virtual copy asset adapted to the MR system.

[0132] Step 1.3, Collaborative Space Instantiation and Intelligent Assistant Initialization: The collaborative space initialization module drives the AR and VR clients to load the prepared virtual copies and graphic assets from the server according to predefined spatial layout rules, completing the anchoring and layout adjustment of virtual objects in the virtual and real spaces, and constructing a collaborative work space that integrates the virtual and real worlds. Then, the system-driven prompt generation module constructs a "collaborative assembly intelligent assistant system prompt" for the current task stage, sends an initialization request through the communication module, and officially activates the intelligent assistant's analysis and decision-making functions, thus completing the full construction of the collaborative work environment.

[0133] Step 2: Collaboration between local workers and remote experts

[0134] Step 2.1, Local Operations and Collaboration Requests: Local workers perform physical assembly tasks on the assembly site. When encountering operational difficulties or decision-making bottlenecks, they initiate collaboration requests to experts through the MR system. During this process, workers describe the problem via voice and use gestures to manipulate virtual replicas for supplementary explanations. The context-aware module runs automatically in the background, monitoring and collecting the worker's voice stream, gaze points, gestures, interaction object IDs, and first-person perspective images in real time, and uploading this multimodal data to the server.

[0135] Step 2.2, Immersive Remote Expert Guidance: Remote experts monitor the assembly site in real-time via video stream on the VR terminal. When a worker makes an error or a question is received, the expert formulates a guidance plan based on the synchronized virtual copy status, worker gestures, and verbal descriptions. Then, the expert demonstrates the correct assembly actions by manipulating the virtual copy, accompanied by voice explanations of key elements. During this process, the context-aware module also collects and uploads the expert's voice commands, controller operation trajectories, and the aforementioned virtual interactive behaviors and related data in real time.

[0136] Step 3: Dynamic assistance from the intelligent assistant

[0137] Step 3.1, Real-time Monitoring of User Needs: When the context-aware module detects communication behaviors (such as voice activation) between the collaborating parties, it triggers data upload. After receiving the data, the prompt word generation module performs the following processing: First, it performs initial screening of the task stage and retrieves the corresponding prompt word template; second, it serializes the multimodal context data and fills it into the template slots; finally, it combines the retrieved knowledge fragments to generate single-round structured prompt words and sends them to the large language model API interface, completing the structured definition of the current task context containing collaborative intent and user needs.

[0138] Step 3.2, Contextual Understanding and Intent Reasoning: After receiving the prompt words, the reasoning and decision-making module first performs data integrity and relevance checks based on the system's preset intelligent assistant role responsibilities and task requirements. Then, it uses the semantic understanding and visual analysis capabilities of the large language model to extract the collaborative sub-task type, user focus object, and task goal from the contextual data. Finally, it accurately analyzes the user's core collaborative intent (such as "query torque" or "request visual alignment") and its implicit information type requirements.

[0139] Step 3.3, Virtual Asset Optimization Scheme Decision: Based on the above analysis results, the reasoning and decision-making module, using the assembly task supplementary knowledge and expert strategies provided by the system prompts, and referring to the reasoning logic in the thought chain and the reasoning cases in the single-round prompts, conducts in-depth reasoning on virtual asset optimization schemes oriented towards the task context and user needs. The decision content covers: the volume, carrier form, and timing of guiding information pushes; the visual optimization representation of virtual copies (feature hierarchy and data appearance attributes); and the multimodal response forms of user-initiated commands. Finally, a structured decision scheme containing all analysis results, execution objects, strategies, and parameters is generated.

[0140] Step 3.4, Adaptive Collaboration Solution Execution: The instruction conversion module receives the structured decision-making scheme, parses it into control scripts that can be directly executed by each client through an instruction mapping mechanism, and distributes them to the application layer modules for execution; the information adaptive push module selects suitable graphic and textual asset carriers for guiding information based on the scheme, and pushes them to the user's field of vision according to the decision timing; the virtual copy visual optimization module performs customized rendering for VR / AR devices on the virtual copy of the target component based on the optimization scheme parameters in the scheme, such as highlighting, transparency, or logical instantiation; the multimodal instruction response module generates a voice broadcast stream and coordinates with changes in visual assets to achieve synchronized audiovisual feedback for user-initiated instructions. Ultimately, this achieves the system's adaptive matching of user collaborative communication needs and immediate response to proactive instructions.

[0141] Step 4: Efficient Collaboration Loop Supported by Intelligent Assistant

[0142] Step 4.1, Intelligent Assistant Proactive Response: Based on the dynamic assistance of the intelligent assistant, the visual appearance of the virtual copy in the collaborative space and the content of the guiding information adaptively match the user's needs. As a dynamic communication medium, it realizes the mapping and synchronization of communication information and presentation through audiovisual multimodal behaviors such as highlighting key features and weakening interfering information. This achieves proactive response to the user's implicit information expression and understanding needs, thereby significantly reducing communication costs and improving information exchange efficiency.

[0143] Step 4.2, User-Initiated Interaction and Intervention: When experts or workers find that the adaptation scheme for virtual assets does not perform as expected, or when there are additional control requirements, they can issue proactive intervention commands through a combination of voice commands and gestures (e.g., "Don't highlight this part"); or directly initiate general knowledge Q&A with the intelligent assistant (e.g., asking about tool usage guidelines). The multimodal command response module instantly parses such proactive requests, directly enforces control over the specified virtual asset, or generates audiovisual fusion knowledge responses, ensuring that system decisions can be taken over and corrected by the user at any time.

[0144] Step 4.3, Implicit Control of Virtual Asset Layout: During collaboration, the system's background context-aware module continuously monitors the user's visual focus area and operational behavior within the workspace; the visual asset layout adjustment module, based on ray detection and bounding box collision algorithms, determines in real time whether visual assets such as information panels obstruct key operational areas. Once a conflict is detected, the system automatically and smoothly moves the interfering items to an empty area outside the field of vision and moves the core information forward to the visual focus area, ensuring the transparency of the visual channel and the smoothness of interaction under the user's natural collaborative state.

[0145] Step 4.4, Intelligent Assembly Quality Inspection: Upon completion of a specific process or overall task, the system initiates the quality inspection process in response to user trigger commands (voice or UI trigger). The context awareness module guides the user to capture keyframes of the assembly area from multiple angles using AR devices and uploads the images. The reasoning and decision-making module calls a large language model to perform assembly state reasoning based on visual semantics: it compares the actual on-site images with virtual copies of standard assembly states or knowledge base graphs to identify whether there are missing parts, incorrect poses, or foreign object residues. Finally, the multimodal command response module generates augmented reality annotations in the AR field of view (such as red box warnings for error areas and ghost shadow guidance for correct positions), and, in conjunction with voice broadcast correction schemes, achieves closed-loop control of assembly quality.

[0146] Through the above steps, the method described in this embodiment achieves an efficient closed loop of "worker questioning - expert guidance - assistant assistance - system support - physical execution", which significantly improves the efficiency and user experience of remote collaborative assembly.

[0147] To further verify the effectiveness of the present invention, the above-mentioned intelligent assistant-enhanced mixed reality remote collaborative assembly system and method are applied to the remote collaborative assembly of a turbofan engine model. This model assembly has a complex structure, encompassing key functional components such as the air intake, compressor assembly, thrust reverser actuator assembly, and thrust reverser door. Its assembly process involves multiple stages, precision operations, and strict sequential constraints, fully reflecting the typical challenges faced in actual industrial assembly scenarios, such as complex spatial operations, high cognitive load, and intensive knowledge requirements. The specific implementation steps are as follows:

[0148] Step 1: Local workers wear HoloLens 2 AR glasses to activate the system, and under the guidance of the interface, collect and upload images of the assembly area. After recognizing the task scene, the system instantly overlays virtual models of the parts to be assembled, assembly process guidance information, and UI interaction interfaces onto the AR / VR field of view, completing the initial construction of the virtual-real fusion collaborative space for the turbofan engine model.

[0149] Step 2: Local workers perform physical assembly actions according to the assembly guidance information; remote experts wear HTC VivePro 2 VR headsets to monitor the correctness of the workers' operations through real-time video streams in an immersive virtual space.

[0150] Step 3: Local workers request assistance with the assembly challenges they face. Both parties work together to express the problems intuitively and provide preliminary solutions through voice descriptions, gestures, and interactive control of virtual replicas.

[0151] Step 4: The intelligent assistant collects the user's voice, gaze, and gesture data in real time during a single round of collaboration, autonomously analyzes and generates a virtual asset control plan, and drives the user's client to render and execute it through the instruction conversion module.

[0152] Step 5: Both parties will conduct a detailed explanation and demonstration of the operation, focusing on the visually optimized virtual replica and the automatically laid-out guide panel. At this point, the visual representation of the virtual asset and the expert's voice commands are dynamically synchronized, intuitively presenting the assembly logic and key operational points.

[0153] Step Six: For complex and difficult issues, experts issue proactive intervention commands in the form of "gesture pointing + voice commands", requiring the intelligent assistant to make specific visual attribute changes to the features of designated parts, and provide in-depth guidance with detailed voice explanations.

[0154] Step 7: After receiving expert guidance, local workers can directly ask the intelligent assistant questions about specific operating techniques or tool usage. The assistant will instantly search the knowledge base and return illustrated explanations and answers, along with voice prompts.

[0155] Step 8: Local workers perform actual assembly operations according to the instructions. During this process, the system monitors the workers' hand activity area in real time, automatically moving non-critical visual assets to the periphery of their field of vision to eliminate obstructions and interference to the workers' line of sight and ensure smooth interaction.

[0156] Step Nine: After local workers complete the physical assembly, they trigger the quality inspection function through interactive commands. Guided by the interface, workers take images of the areas to be inspected, and the intelligent assistant uses a visual model to compare and analyze the images, providing real-time feedback on the assembly results and correction suggestions.

[0157] This application example verifies the effectiveness and advancement of the system and method proposed in this invention. By introducing an intelligent assistant centered on a large language model, the system can dynamically adjust the representation and layout of virtual assets according to the task context, significantly enhancing the ability to express and understand non-verbal cues in 3D spatial tasks. The collaborative work of the virtual copy construction module, assembly process data processing module, and collaborative space initialization module enables rapid deployment for new assembly tasks. The virtual copy visual optimization module adaptively adjusts the visual performance of parts based on the principle of information supply and demand matching; the visual asset layout adjustment module introduces a dynamic anti-occlusion mechanism; the information adaptive push module dynamically adjusts the information granularity according to cognitive load; the reliability enhancement mechanism of human-large language model collaboration effectively suppresses model illusions; and the assembly quality inspection module realizes a complete closed loop from work guidance to result verification. This application example fully demonstrates that this invention can solve the technical problems of insufficient adaptive capability, fixed virtual assets, and inability to respond to dynamic situational needs in existing MR remote collaborative assembly systems, achieving the expected technical effects.

[0158] The above description is merely a specific embodiment of the present invention, intended to enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments without creative effort will be readily apparent to those skilled in the art; the basic principles defined herein can be applied to other embodiments without departing from the spirit and scope of the present invention.

[0159] It should be understood that this invention is not limited to the specific content described above, and various equivalent modifications or substitutions can be made without departing from its scope of protection. The scope of protection of this invention is defined only by the appended claims.

Claims

1. A mixed reality remote collaborative assembly system with intelligent assistant enhancement, characterized in that, include: The AR client is configured at the assembly site to collect on-site environmental data and user behavior data, and to overlay virtual information to the user. The VR client, configured on a remote end, is used to recreate assembly scenes in an immersive virtual space for experts to browse and interact with. The server communicates and connects with the AR client and VR client respectively, and stores and manages digital assets and knowledge base; The intelligent assistant communicates with the AR client, VR client, and server respectively. The intelligent assistant is driven by a large language model and is configured to: perceive multimodal interaction data collected by the AR client and VR client in real time to obtain the current task context. Based on the current task context, reasoning is performed by combining assembly domain knowledge and the general domain knowledge of the large language model to generate a virtual asset optimization and control strategy that matches the current task context. The optimization and control strategy includes adjusting the visual representation, spatial layout, or information push timing of the virtual assets. The optimized control strategy is converted into executable instructions and driven to execute by the AR client and / or VR client, so that the virtual assets in the client's field of vision dynamically change in accordance with the current task context, thereby assisting information interaction in the remote collaborative assembly process.

2. The intelligent assistant-enhanced mixed reality remote collaborative assembly system according to claim 1, characterized in that, The system is logically divided into a data layer, a service layer, and an application layer. The data layer is deployed on the server and includes: The virtual copy construction module retrieves the original CAD data of the parts from the product data management library based on the current assembly task, and constructs an interactive virtual copy with multi-level semantic features through format conversion, semantic feature classification and annotation. The assembly process data processing module retrieves and parses the process documents of the current assembly task from the assembly process knowledge base, extracts structured processes, steps, and graphic information, prepares standardized visual assets that the system can load, and constructs an assembly domain knowledge base for the intelligent assistant to call. The collaborative space initialization module retrieves the virtual copy and the standardized visual assets according to the current assembly task, instantiates the space in each client according to the preset layout rules, constructs a virtual-real integrated collaborative work space, and activates the intelligent assistant. The service layer is deployed on the server and includes: The context awareness module collects and integrates voice, gestures, gaze focus, interaction data and first-person perspective images from various clients in real time to form a standardized context status data package. The prompt word generation module matches templates from a preset prompt word template library based on the context state data package and the current assembly task, and generates structured prompt words that drive the large language model to perform reasoning. The reasoning and decision-making module, driven by the large language model, receives and parses the structured prompt words, combines them with the assembly domain knowledge base to perform deep reasoning, and generates optimized control strategies for virtual assets. The instruction conversion module parses the optimization and control strategy and converts it into control instructions that can be executed by each module of the application layer. The communication module is used to establish and maintain data transmission between the data layer, service layer, and application layer, as well as with the cloud-based large language model service; The application layer is deployed on the AR client and VR client, including: The virtual copy visual optimization module, in response to the control command, performs feature-level visual attribute adjustment on the virtual copy of the target component; The information adaptive push module responds to the control command and dynamically adjusts the granularity, modality, and push timing of the guidance information according to the task context and user role. The visual asset layout adjustment module responds to the control command, dynamically manages the spatial distribution of virtual information in the user's field of vision, and performs anti-occlusion adjustment and layout optimization. The multimodal instruction response module responds to the control instructions and provides audiovisual feedback to the user's interaction requests.

3. The intelligent assistant-enhanced mixed reality remote collaborative assembly system according to claim 2, characterized in that, The virtual copy visual optimization module has a built-in adaptive visual optimization strategy based on information supply and demand matching. The strategy is used to accurately filter the hierarchical information carriers of the virtual copy and dynamically reconstruct its visual presentation according to the dynamic information needs under different collaborative assembly task scenarios. The feature enhancement operations performed by the virtual copy visual optimization module include one or more of the following: size enlargement, color highlighting, edge highlighting, center line rendering, centroid trajectory rendering, and model phantom rendering; the feature weakening operations include one or more of the following: feature deletion, transparent rendering, and wireframe rendering.

4. The intelligent assistant-enhanced mixed reality remote collaborative assembly system according to claim 2, characterized in that, The adaptive information push module is configured as follows: The information granularity is dynamically determined based on task complexity and user role. High-density system-level information is pushed to expert users, while procedural instructions that are strongly related to the current operation are pushed to worker users. Based on the cognitive load scenario, push a detailed text list when the cognitive load is low, and perform information simplification and enlarge the UI display size when the cognitive load is high or an emergency operation is required; as well as Based on the timing of the push, a step-by-step node trigger strategy is used to push regular guidance information, knowledge supplement information is pushed when the user's gaze is stable or their hand is still, and immediate commands are given the highest push priority.

5. The intelligent assistant-enhanced mixed reality remote collaborative assembly system according to claim 2, characterized in that, The visual asset layout adjustment module is configured as follows: Based on the importance of the command, high-priority virtual assets are moved to the center of the user's field of vision, while low-priority assets are moved to the edge of the field of vision. Real-time monitoring of object distribution within the user's field of view; using bounding box collision detection to determine whether the UI panel obstructs the focus object, hand operation area, or key component copy; when obstruction is detected, calling the blank area search algorithm to move the obstructed asset to the non-obstructed area. as well as Automatically adjust text reading information to a comfortable viewing distance, and dynamically switch label information between field-following mode and world-anchoring mode based on object attributes.

6. The intelligent assistant-enhanced mixed reality remote collaborative assembly system according to claim 2, characterized in that, The virtual copy construction module performs semantic-level feature classification and annotation on the converted virtual model according to preset feature classification rules. The feature classification rules include geometric features, appearance features and logical features, and automatically attach rendering attributes and physical interaction components. Before rendering, the collaborative space initialization module performs a runtime integrity self-check on the loaded virtual assets to confirm that the necessary functional components of mixed reality are fully mounted.

7. The intelligent assistant-enhanced mixed reality remote collaborative assembly system according to claim 2, characterized in that, The system prompt word architecture maintained by the prompt word generation module includes the following modules: The task background module is used to build a large language model's global understanding of the task process, collaboration goals, and system architecture. The role setting module is used to define the role and responsibilities of the large language model in collaborative tasks; The standardized reasoning reference module is used to transform the standardized reasoning and decision-making ideas of domain experts into structured instructions; The structured output specification module is used to define standardized data output formats; The exception handling and precautions module is used to preset boundary conditions and response strategies; The domain knowledge foundation module is used to integrate process documents, component attributes, and visual optimization strategies associated with the current task. as well as The startup instruction module is used to trigger the large language model to enter the working state.

8. The intelligent assistant-enhanced mixed reality remote collaborative assembly system according to claim 1, characterized in that, The system also includes a reliability enhancement mechanism for human-large language model collaboration, which is configured as follows: Embedding standardized examples based on thought chains into system prompts guides large language models to learn expert-level reasoning logic and mandates that large language models perform internal verification before outputting the final decision solution. as well as When the large language model cannot parse the input, the inference result conflicts with the security rules, or there is great uncertainty, an abnormal circuit breaker is triggered, and a problem report is generated to request user intervention. Furthermore, it supports users in actively monitoring and intervening in the output scheme of the intelligent assistant, responding to the user's active control commands to perform intent re-identification and reasoning correction, and realizing dynamic error correction and optimization of the human loop.

9. The intelligent assistant-enhanced mixed reality remote collaborative assembly system according to claim 1, characterized in that, The system also includes an assembly quality inspection module, configured as follows: In response to the user's trigger command, guide the AR client to collect multi-angle keyframe images of the assembly area; The large language model is invoked to perform assembly state reasoning based on visual semantics, and the on-site images are compared with virtual copies of standard assembly states or knowledge base graphs to identify missing parts, incorrect poses, or foreign object residues. Augmented reality annotations are generated through a multimodal command response module and combined with a voice broadcast correction scheme.

10. A mixed reality remote collaborative assembly method enhanced with an intelligent assistant, characterized in that, Includes the following steps: Step 1: Build an MR remote collaborative work environment Step 1.1, Task Scene Recognition: Local workers wear augmented reality devices to enter their workstations and start the AR client. The collaborative space initialization module establishes a connection with the cloud-based large language model, submits scene recognition system prompts to activate the model's task analysis capabilities, responds to the user's interactive scene acquisition command, drives the context perception module to call the camera to capture key frame images on site, encapsulates them into multimodal single-round prompts through the prompt generation module, and sends them to the reasoning and decision-making module. The large language model performs visual understanding and reasoning on the images and outputs the type of the current assembly task and the list of parts involved. Step 1.2, Task Data Asset Preparation: The assembly process data processing module receives the task identification results, retrieves the corresponding process documents and part information from the assembly process knowledge base, parses and prepares the graphic guidance information required for the task and the assembly domain knowledge base for the intelligent assistant to access; the virtual copy construction module retrieves the original CAD model from the product data management library based on the identified parts list, performs format conversion, semantic-level feature classification and annotation, and MR function additional processing to generate a structured virtual copy asset adapted to the MR system; Step 1.3, Collaborative Space Instantiation and Intelligent Assistant Initialization: The collaborative space initialization module drives the AR client and VR client, loads virtual copies and graphic assets from the server according to predefined spatial layout rules, completes the anchoring and layout adjustment of virtual objects in the virtual and real spaces, and constructs a collaborative work space that integrates the virtual and real worlds; the prompt word generation module constructs the collaborative assembly intelligent assistant system prompt words for the current task stage, sends an initialization request through the communication module, and activates the intelligent assistant's analysis and decision-making functions; Step 2: Collaboration between local workers and remote experts Step 2.1, Local Operations and Collaboration Requests: When local workers encounter operational difficulties or decision-making bottlenecks while performing physical assembly tasks on the assembly site, they initiate collaboration requests to experts through the MR system. They describe the problem through voice and use gestures to control virtual replicas for auxiliary explanations. The context awareness module monitors and collects the worker's voice stream, gaze point, gestures, interactive object IDs, and first-person perspective images in real time, and uploads the multimodal data to the server. Step 2.2, Immersive guidance from remote experts: Remote experts monitor the assembly site in real time through video streams on the VR terminal. When they discover worker operation errors or receive questions, they formulate guidance plans by combining the status of the virtual copy synchronized with the system, worker gestures and language. They demonstrate the correct assembly actions by manipulating the virtual copy and explain key elements with voice. The context-aware module synchronously collects and uploads the expert's voice commands, controller operation trajectories, and virtual interaction behavior data; Step 3: Dynamic assistance from the intelligent assistant Step 3.1 Real-time monitoring of user needs: When the context awareness module detects the communication behavior of the two collaborating parties, it triggers data upload. After receiving the data, the prompt word generation module performs the initial screening of the task stage and retrieves the corresponding prompt word template. It serializes the multimodal context data and fills it into the template slot. Combined with the retrieved knowledge fragments, it generates a single-round structured prompt word and sends it to the large language model interface. Step 3.2, Contextual Understanding and Intent Reasoning: After receiving the prompt words, the reasoning and decision-making module performs data integrity and relevance checks in accordance with the system's preset intelligent assistant role responsibilities and task requirements. It uses the semantic understanding and visual analysis capabilities of the large language model to extract the collaborative sub-task type, user focus object and task goal from the contextual data, and analyzes the user's core collaborative intent and its implicit information type needs. Step 3.3, Virtual Asset Optimization Scheme Decision: The reasoning and decision module supplements knowledge and expert strategies based on the assembly task in the system prompts, and conducts in-depth reasoning on virtual asset optimization schemes oriented towards task context and user needs, referring to the logic of the thinking chain and reasoning cases. The decision content covers the volume, carrier form and timing of the push of guiding information, the visual optimization representation of virtual copies, and the multimodal response form of user active commands, generating a structured decision scheme containing analysis results, execution objects, strategies and parameters. Step 3.4, Adaptive Collaboration Scheme Execution: The instruction conversion module receives the structured decision scheme, parses it into control instructions executable by each client through the instruction mapping mechanism, and distributes them to the application layer modules for execution; the information adaptive push module selects suitable graphic and text asset carriers for guidance information based on the scheme and pushes them to the user's field of vision according to the decision timing; the virtual copy visual optimization module adjusts the visual attributes at the feature level of the virtual copy of the target component based on the optimization parameters in the scheme; the multimodal instruction response module generates a voice broadcast stream and coordinates with changes in visual assets to achieve adaptive matching of user collaborative communication needs and real-time response to proactive instructions; Step 4: Efficient Collaboration Loop Supported by Intelligent Assistant Step 4.1, Intelligent Assistant Proactive Response: Based on the dynamic assistance of the intelligent assistant, the visual appearance of the virtual copy in the collaborative space and the content of the guiding information adaptively match the user's needs. Through the audiovisual multimodal behavior of highlighting key features and weakening interfering information, the mapping and synchronization of communication information and presentation are realized, thereby achieving a proactive response to the user's implicit information expression and understanding needs. Step 4.2, User-initiated interaction and intervention: When experts or workers find that the adaptation solution for virtual assets does not meet expectations or there are additional control requirements, they can issue an active intervention command through a combination of voice commands and gestures, or directly initiate a general knowledge Q&A session with the intelligent assistant; the multimodal command response module will instantly parse the active request and forcibly adjust the specified virtual assets or generate an audiovisual knowledge response. Step 4.3, Implicit Control of Virtual Asset Layout: During the collaboration process, the context awareness module continuously monitors the user's visual attention area and operation behavior in the workspace. The visual asset layout adjustment module uses ray detection and bounding box collision algorithms to determine in real time whether the information panel obstructs the key operation area. When a conflict is detected, the interference items are automatically and smoothly moved to the free area outside the field of vision, and the core information is moved forward to the visual focus area. Step 4.4, Intelligent Assembly Quality Inspection: After completing a specific process or overall task, in response to the user's trigger command, the system starts the quality inspection process. The context awareness module guides the user to collect keyframes of the assembly area from multiple angles through AR devices and upload the images. The reasoning and decision-making module calls the large language model to perform assembly state reasoning based on visual semantics, compares the features of the on-site real-time images with virtual copies of standard assembly states or knowledge base graphs, identifies missing parts, incorrect poses, or foreign object residues, and generates augmented reality annotations in the AR client's field of vision through the multimodal command response module, along with voice broadcasting correction schemes.