Remote expert guidance intention visualization method, equipment and device for XR devices

By receiving multimodal instructions from remote experts and using visual language models to generate structured intent data packets, the problem in existing technologies that pure language instructions are difficult to accurately describe three-dimensional spatial relationships in remote collaboration is solved, and efficient and accurate remote guidance is achieved.

CN120431273BActive Publication Date: 2025-09-12HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510934248.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-12
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

Existing technologies rely on pure language instructions for guidance in remote collaboration, which makes it difficult to accurately describe three-dimensional spatial relationships, resulting in low guidance efficiency. This is especially prone to misoperation in complex scenes or when there are visual interferences.

Method used

By receiving multimodal instructions from remote experts, including voice, text, and visual markers, the system uses a visual language model to generate structured intent data packets, and based on this, generates dynamic 3D visual instructions that match the operation action intentions, and renders them in real time to the XR device screen of the on-site user.

Benefits of technology

It effectively solves the ambiguity problem of language description in three-dimensional space, improves the accuracy and expressiveness of guidance information, ensures the effective anchoring of guidance information in the three-dimensional environment, and significantly improves the efficiency and accuracy of remote guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431273B_ABST
    Figure CN120431273B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, and apparatus for visualizing the intention of remote expert guidance for XR devices. The method includes: receiving multimodal instructions issued by a remote expert in combination with a real-time scene preview image; inputting the multimodal instructions and the associated scene preview image into a visual language model to generate a structured intent data packet, and based on the structured intent data packet, generating dynamic 3D visual instructions that match the operation action intention in the structured intent data packet; spatially anchoring the 3D visual instructions to the real-time image of the XR device of the on-site user, and determining the rendering position of the 3D visual instructions based on the matching results, so as to render the 3D visual instructions in the real-time image of the XR device. Through the above method, experts can express complex operation intentions in a way that is consistent with human intuition, reducing communication ambiguity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of extended reality and artificial intelligence technology, and in particular to a method, device, and apparatus for visualizing remote expert guidance intentions for XR devices. Background Art

[0002] In the field of extended reality remote collaboration, existing technologies generally rely on experts to guide on-site users through voice or text instructions. However, purely verbal instructions have inherent flaws when describing three-dimensional spatial relationships. Experts find it difficult to accurately express the orientation, dynamic operation path, and relative spatial position of target objects in words. This forces on-site users to repeatedly confirm the operation object, significantly reducing the efficiency of guidance. Especially when there are visual distractions or complex spatial structures in the scene, the ambiguity of verbal descriptions easily leads to the risk of misoperation, which cannot meet the requirements of high-precision scenarios such as industrial maintenance and medical surgery.

[0003] To enhance the intuitiveness of guidance, some improvements incorporate manual annotation tools for experts, allowing them to draw two-dimensional markers on a shared video stream. However, these markers are merely static, planar overlays that lack the ability to adapt to dynamic changes in the scene: when the user shifts their perspective or the position of the target object changes, the markers cannot automatically correlate with the actual three-dimensional spatial properties of the object, resulting in misaligned visual guidance. More importantly, traditional marker systems rely entirely on manual expert drawing, unable to understand the operational intent behind the markers and struggling to generate dynamic visual feedback that matches the semantics of the actions.

[0004] Furthermore, existing 3D visual guidance solutions generally rely on precise 3D scene reconstruction and spatial anchoring technologies. These solutions require high-performance on-site equipment to construct environmental maps in real time and rely on predefined object recognition models to locate targets. In practice, 3D reconstruction of complex scenes is significantly time-consuming, making it difficult to meet the real-time requirements of remote guidance. Furthermore, predefined models cannot generalize to handle sudden and non-standardized operation objects, rendering visual guidance ineffective. Summary of the Invention

[0005] The embodiments of the present application provide a method, device, and apparatus for visualizing remote expert guidance intentions for XR devices to solve the above-mentioned technical problems.

[0006] On the one hand, embodiments of the present application provide a method for visualizing remote expert guidance intentions for XR devices, including:

[0007] Receiving multimodal instructions issued by a remote expert in combination with a real-time scene preview image; the multimodal instructions include at least one of the following: voice instructions, text instructions, or visual mark instructions;

[0008] Inputting the multimodal instruction and the associated scene preview image into a visual language model to generate a structured intent data packet, and based on the structured intent data packet, generating a dynamic 3D visual instruction that matches the operation action intent in the structured intent data packet;

[0009] The 3D visual instructions are spatially anchored with the real-time picture of the XR device of the on-site user, and the rendering position of the 3D visual instructions is determined according to the matching result, so as to render the 3D visual instructions in the real-time picture of the XR device.

[0010] In one implementation of the present application, receiving a multimodal instruction issued by a remote expert in combination with a real-time scene preview image specifically includes:

[0011] Send the remote expert a real-time preview video stream of the scene shared by the on-site user's XR device;

[0012] capturing a voice command or text command input by the remote expert in combination with the scene preview image frame in the scene preview video stream, and upon receiving the voice command, converting the voice command into a text format to determine a command text of the voice command or the text command;

[0013] Capturing a visual marking instruction drawn by the remote expert on the scene preview image frame in the scene preview video stream; the visual marking instruction includes the coordinates of the outline of the circled area or a sequence of path trajectory points;

[0014] The timestamps corresponding to the voice instruction or the text instruction and the visual mark instruction are collected to bind the coordinate data of the instruction text and the visual mark instruction to the corresponding scene preview image frames respectively as a time-series associated multimodal input set.

[0015] In one implementation of the present application, the multimodal instruction and the associated scene preview image are input into a visual language model to generate a structured intent data packet, specifically including:

[0016] Inputting the multimodal instruction and the associated scene preview image into a visual language model to locate spatial orientation descriptors of a target object in the instruction text using the visual language model, and parsing the spatial orientation descriptors to determine location information of the target object in the scene preview image; the location information includes at least one of the following: two-dimensional bounding box coordinates, pixel-level masks, or key points;

[0017] Extracting core operational verbs and related directional adverbs from the instruction text, and parsing the core operational verbs and the related directional adverbs into preset action intention codes;

[0018] The spatial relationship description word is identified, and an orientation offset vector of the target object relative to a scene reference object is generated.

[0019] In one implementation of the present application, parsing the spatial orientation descriptor to determine the location information of the target object in the scene preview image specifically includes:

[0020] Identifying the direction keyword and the reference object name in the spatial orientation description word, and determining the spatial position of the scene reference object in the scene preview image;

[0021] The vector relationship between the spatial position and the direction keyword is calculated, and the positioning coordinates of the target object are output; the positioning coordinates include the direction vector.

[0022] In one implementation of the present application, the dynamic 3D visual instruction includes a visual type, a spatial orientation parameter, and a dynamic effect template;

[0023] Based on the structured intent data packet, generating a dynamic 3D visual instruction that matches the operation action intention in the structured intent data packet, specifically including:

[0024] Selecting a visual type corresponding to the 3D visual instruction to be generated according to the action intention code in the structured intention data packet; the visual type includes direction indication, area focus, or operation demonstration;

[0025] Configuring spatial orientation parameters of the 3D vision instructions to be generated of the vision type based on the orientation offset vector in the structured intent data packet;

[0026] The dynamic effect template of the 3D visual instruction to be generated, which is associated with the core operation verb, is loaded; the dynamic effect template includes pulse intensity, motion trajectory or deformation sequence.

[0027] In one implementation of the present application, the present invention further includes:

[0028] When the visual type is direction indication, calculating the Euler angle rotation parameters and pulsation frequency parameters of the arrow in three-dimensional space according to the direction vector in the structured intent data packet;

[0029] Loading a predefined arrow base mesh, and applying the Euler angle rotation parameters and the pulsation frequency parameters to the arrow base mesh to generate a dynamic arrow instance;

[0030] In the case where the visual type is regional focus, generating a semi-transparent 3D highlight surface matching the outline of the target object based on the pixel-level mask of the target object in the structured intent data packet;

[0031] When the visual type is an operation demonstration, a mechanical motion animation sequence is generated for looping playback.

[0032] In one implementation of the present application, spatially anchoring the 3D visual instructions to the real-time image of the on-site user's XR device specifically includes:

[0033] Extracting, from the 3D vision instruction, a set of local feature descriptors of a target area in the scene preview image by a remote expert;

[0034] Searching for an area similar to the local feature descriptor in the real-time image of the on-site user, and when the matching similarity exceeds a preset matching threshold, taking the area with the highest matching similarity as the matching image area;

[0035] A central coordinate plane of the matching image area is determined, and a rendering origin of the 3D vision instruction is aligned to the central coordinate plane.

[0036] In one implementation of the present application, the present invention further includes:

[0037] Detecting the surface normal direction of the matching image area in the real-time image of the on-site user, and adjusting the rendering direction of the 3D visual instruction so that the 3D visual instruction is always perpendicular to the surface normal direction;

[0038] Dynamically update the perspective projection parameters of the 3D vision instruction based on the real-time posture data of the XR device of the on-site user.

[0039] On the other hand, an embodiment of the present application further provides a remote expert guidance intention visualization device for an XR device, the device comprising:

[0040] at least one processor;

[0041] and, a memory communicatively coupled to the at least one processor;

[0042] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the remote expert guidance intention visualization method for XR devices as described above.

[0043] On the other hand, an embodiment of the present application also provides a remote expert guidance intention visualization device for XR devices, the device including: a multimodal instruction receiving module for receiving multimodal instructions issued by a remote expert in combination with a real-time scene preview image; the multimodal instructions include at least one of the following: voice instructions, text instructions or visual mark instructions; a visual instruction generation module for inputting the multimodal instructions and the associated scene preview image into a visual language model to generate a structured intent data packet, and based on the structured intent data packet, generating a dynamic 3D visual instruction that matches the operation action intention in the structured intent data packet; a visual instruction rendering module for spatially anchoring the 3D visual instruction with the real-time screen of the XR device of the on-site user, and determining the rendering position of the 3D visual instruction based on the matching result, so as to render the 3D visual instruction in the real-time screen of the XR device.

[0044] The present application provides a method, device, and apparatus for visualizing remote expert guidance intentions for XR devices, which have at least the following beneficial effects:

[0045] By receiving and fusing multimodal instructions, the implicit three-dimensional spatial relationships in the expert's spoken descriptions are converted into computable data, resolving the ambiguity of pure language instructions in spatial positioning. This allows experts to express complex operational intentions in a way that is consistent with human intuition, significantly reducing communication ambiguity. Based on the joint analysis of instructions and scene images using a visual language model, the operation verbs and spatial attributes of the target objects are accurately parsed, and structured intent data packets are output to generate dynamic visual instructions that strictly match the action semantics. This avoids the disconnection between traditional static markings and operational semantics, and improves the expressiveness and accuracy of guidance information. By matching the target visual anchor point area in the real-time image and rendering 3D visual instructions based on the matching results, the difference in perspective or posture between the expert's preview image and the user's real-time scene is overcome, so that the generated dynamic arrows, highlighted areas and other visual elements always fit the actual spatial position of the object, ensuring the effective anchoring of the guidance information in the three-dimensional environment. Relying on feature matching to achieve spatial mapping rather than complex three-dimensional reconstruction greatly reduces computing power requirements and enables XR devices to respond to instruction updates in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0047] Figure 1 Schematic diagram of an application scenario of the method for visualizing remote expert guidance intentions for XR devices provided in an embodiment of the present application;

[0048] Figure 2A flowchart of a method for visualizing remote expert guidance intentions for XR devices provided in an embodiment of the present application;

[0049] Figure 3 A flowchart of a multimodal instruction receiving method for a remote expert provided in an embodiment of the present application;

[0050] Figure 4 A flow chart of a method for generating dynamic 3D visual instructions provided in an embodiment of the present application;

[0051] Figure 5 A schematic diagram of the internal structure of a remote expert guidance intention visualization device for XR devices provided in an embodiment of the present application;

[0052] Figure 6 Schematic diagram of the internal structure of the remote expert guidance intention visualization device for XR devices provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0054] The remote expert guidance intention visualization method for XR devices provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Figure 1 As shown, the application environment may include: on-site XR equipment 101, expert guidance terminal 102, communication network 103, intent processing server 104, and 3D vision engine 105.

[0055] On-site XR devices 101 (e.g., AR glasses or MR headsets) are equipped with environmental perception sensors and a real-time rendering module. These capture a video stream of the operational scene and transmit it to the expert via a communication network 103. They also receive and overlay dynamic 3D visual instructions. In industrial maintenance scenarios, these devices can provide enhanced guidance on the internal structure of the equipment through a transparent display.

[0056] The expert guidance terminal 102 (e.g., a tablet or workstation) integrates a multimodal input interface, including a voice acquisition microphone, a touch screen, and a preview display unit. Through this terminal, the expert observes the on-site scene preview stream and simultaneously inputs voice commands (e.g., "Tighten the top left bolt") and circles selections on the screen, forming a spatiotemporally aligned multimodal instruction set.

[0057] The communication network 103, serving as a low-latency data transmission channel, utilizes a 5G / UWB hybrid link architecture to provide two-way real-time video streaming between the on-site XR device 101 and the expert guidance terminal 102. It also establishes a high-bandwidth collaborative computing channel between the intent processing server 104 and the 3D vision engine 105. This network ensures that the latency of the entire instruction processing link is below the perceptual threshold.

[0058] The intent processing server 104 deploys the core algorithm of the visual language model (VLM), receives multimodal instructions and associated scene frames from the expert guidance terminal 102, and performs cross-modal joint reasoning: parses spatial orientation descriptors to generate target masks, extracts action verbs and maps them into intent codes, and outputs structured data packets to drive downstream generation modules.

[0059] The 3D vision engine 105 is based on the game engine architecture (such as Unity DOTS) and includes a parameterized generation pipeline: it activates the corresponding generator (arrow / highlight / animation) according to the visual type selector in the intent data packet, injects spatial orientation parameters and dynamic templates, and outputs a lightweight 3D instruction data stream, which is distributed to the on-site XR device 101 for real-time rendering via the communication network 103.

[0060] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0061] Figure 2 A flowchart of a method for visualizing remote expert guidance intentions for XR devices provided in an embodiment of the present application.

[0062] The analysis method involved in the embodiments of the present application can be implemented by a terminal device or a server, and the present application does not impose any special restrictions on this. For ease of understanding and description, the following embodiments are described in detail using a server as an example.

[0063] It should be noted that the server can be a single device or a system composed of multiple devices, that is, a distributed server, and this application does not make any specific restrictions on this.

[0064] like Figure 2 As shown, the embodiment of the present application provides a method for visualizing remote expert guidance intentions for XR devices, including:

[0065] Step 201: Receive a multimodal instruction issued by a remote expert in combination with a real-time scene preview image.

[0066] It should be noted that the multimodal instructions in the embodiments of the present application include at least one of the following: voice instructions, text instructions or visual mark instructions.

[0067] Extended Reality (XR) combines the real and virtual worlds through computers to create a virtual environment where humans and machines can interact. It is a general term for various technologies, including AR, VR, and MR. By integrating the visual interaction technologies of these three, users can experience an immersive experience with a seamless transition between the virtual and real worlds.

[0068] In one embodiment of the present application, receiving a multimodal instruction issued by a remote expert in combination with a real-time scene preview image specifically includes:

[0069] Figure 3 This is a flow chart of a multi-modal instruction receiving method for a remote expert provided in an embodiment of the present application. Figure 3 As shown, the embodiment of the present application provides a method for receiving multimodal instructions from a remote expert, which specifically includes the following steps:

[0070] Step 301: Send a scene preview video stream shared in real time by an on-site user's XR device to a remote expert;

[0071] Step 302: Capture the voice command or text command input by the remote expert in combination with the scene preview image frame in the scene preview video stream, and upon receiving the voice command, convert the voice command into a text format to determine the command text of the voice command or text command;

[0072] Step 303: capturing the visual marking instructions drawn by the remote expert on the scene preview image frame in the scene preview video stream; the visual marking instructions include the coordinates of the outline of the circled area or the sequence of path trajectory points;

[0073] Step 304 : Collect the timestamps corresponding to the voice command or text command and the visual markup command, so as to bind the coordinate data of the command text and the visual markup command to the corresponding scene preview image frames to form a time-series-associated multimodal input set.

[0074] In this embodiment, the on-site user's XR device captures the operating scene in real time through a camera, such as AR glasses, and encodes the operating scene into a low-latency video stream and transmits it to the client interface of the remote expert. Through the client interface, the expert observes the scene preview shared in real time by the on-site user's XR device. The scene preview can be an image or a video stream. The embodiment of this application uses a video stream as an example. Similarly, the scene preview image can also be directly transmitted. For example, in an industrial equipment maintenance scenario, the video stream can show a close-up of a local pipeline inside the industrial equipment.

[0075] Experts can enter text instructions directly into the input box of the client interface, or they can also speak the operation requirements into the microphone, such as tightening the pipe joint covered by oil or flipping the red switch on the screen upward. It should be noted that if the instructions received from the expert are voice commands, the voice recognition engine needs to convert the acquired voice commands into structured text in real time. This way, regardless of whether the expert enters voice commands or text commands, the corresponding command text will be obtained, making it easier to intuitively understand the specific content of the expert's instructions through text.

[0076] Experts trigger the annotation tool in the scene preview image and, through direct or indirect touch, draw a closed contour line, such as circling an oily area, or a path trajectory, such as drawing a line along the path of a pipe. The system records the normalized screen coordinate sequence or trajectory point set of the contour line.

[0077] It should be noted that the XR devices of on-site users can be mobile terminals, non-mobile terminals or virtual reality devices. Mobile terminals include mobile smart terminals such as mobile phones, tablets, laptops, etc., non-mobile terminals include smart large screens fixed in actual scenes, and virtual reality devices include AR glasses.

[0078] If the scene preview image is displayed on a mobile or non-mobile terminal, the expert can directly touch the terminal device, so drawing on the terminal device is done using direct touch. However, if the display device is a virtual reality device, the expert cannot directly touch the scene preview image displayed on the device. Drawing is achieved through sensing signals received within the sensing range of the virtual reality device, so in this case, drawing on the device is done using indirect touch.

[0079] It is understood that when capturing multimodal instructions, millisecond-level timestamps are added to the command text of voice or text instructions and the coordinate data of visual markup instructions, and these are bound to the corresponding video keyframes to form a time-correlated multimodal input set. Specifically, when the expert verbally describes the oil occlusion and simultaneously circles an area in the scene preview image, the system associates the oil keyword and the outline coordinates of the circled area as the same semantic unit, thus resolving the disconnect between the verbal description and the visual object in traditional solutions.

[0080] Step 202: Input the multimodal instructions and the associated scene preview images into the visual language model to generate a structured intent data packet, and based on the structured intent data packet, generate a dynamic 3D visual instruction that matches the operation action intention in the structured intent data packet.

[0081] The Vision-Language Model (VLM) is a pictographic feature encoding based on predicted representations of the real world. This allows AI to possess its own language, understand pixels, comprehend speech sequences, and perceive the world, truly possessing the core capabilities of embodied intelligence. Furthermore, this multimodal language encoding can be used for communication and interaction between embodied intelligences, building a ubiquitous machine intelligence society. The Vision-Language Model is the core of this multimodal general model.

[0082] In this embodiment, the VLM fuses spatial descriptors in the instruction text with the contour coordinates of the visual markers, performing joint inference in the scene preview frame. Specifically, the text encoder extracts embedding vectors for entity words such as pipe joints and oil stains. The image encoder calculates a heat map of the similarity between the visual features and the text vectors within the contour region, outputting a pixel-level mask of the target object, such as a binary segmentation map of the pipe joint behind the oil stain.

[0083] For example, the action parsing module extracts the core verb "tighten" and maps it to a preset action code. The spatial relationship parser generates an orientation offset vector of the target object relative to the scene reference object based on the occlusion of the orientation word, such as [0, 0, -0.2], which is used to indicate backward movement in the depth direction.

[0084] In one embodiment of the present application, based on the descriptive language in the expert's instructions, such as the red switch, the screw in the upper left corner, the place I circled, and the simple visual instructions that the expert may make on the scene preview image, the specific object, component or area pointed to by the expert is accurately located in the associated scene preview image or the corresponding video frame. Specifically, using the VLM-based Visual Grounding technology, the position detection of the target area can be achieved by inputting the scene preview image and text description. It should be noted that the output positioning information of the specific object, component or area pointed to by the expert can be a bounding box coordinate, a segmentation mask or a key point.

[0085] Extracting core action verbs and related directional adverbs from the instruction text helps understand the core actions the expert wants the user to perform, such as pulling upward, tightening, and observing, or the states they want the user to focus on. Furthermore, the expert analyzes the three-dimensional spatial relationships and action parameters implicit in the instruction text, such as "above A" and "along the edge of B," and action parameters such as "upward" and "slight."

[0086] In this embodiment, when parsing the spatial orientation descriptor to determine the location information of the target object in the scene preview image, the directional keywords and reference object names in the spatial orientation descriptor are identified, thereby determining the spatial position of the scene reference object in the scene preview image based on the directional keywords and reference object names. Then, the vector relationship between the spatial position and the directional keywords is calculated to obtain the location coordinates of the target object. It should be noted that the location coordinates in this embodiment of the application include the directional vector of the target object.

[0087] Artificial Intelligence Generated Content (AIGC) refers to content generated through artificial intelligence technology, primarily including text, images, audio, and video. The core concept of AIGC is to use artificial intelligence algorithms to generate content with a certain degree of creativity and quality, usually based on technologies such as Generative Adversarial Networks (GANs) and pre-trained models. These technologies can automatically generate relevant content by learning from large amounts of data, thereby greatly improving the speed and efficiency of content creation.

[0088] In one embodiment of the present application, generating a dynamic 3D visual instruction that matches the operation action intent in the structured intent data packet based on the structured intent data packet specifically includes:

[0089] Figure 4 This is a flow chart of a method for generating dynamic 3D visual instructions provided in an embodiment of the present application. Figure 4 As shown, the embodiment of the present application provides a method for generating dynamic 3D visual instructions, which specifically includes the following steps:

[0090] Step 401: Select a visual type corresponding to the 3D visual instruction to be generated based on the action intention code in the structured intention data packet; the visual type includes direction indication, area focus, or operation demonstration;

[0091] Step 402: Based on the orientation offset vector in the structured intent data packet, configure the spatial orientation parameters of the 3D vision instruction to be generated of the vision type;

[0092] Step 403: Load the dynamic effect template of the 3D visual instruction to be generated that is associated with the core operation verb; the dynamic effect template includes pulse intensity, motion trajectory or deformation sequence.

[0093] In this embodiment, a structured intent data packet is received from a visual language model, which includes key fields such as action intention encoding and orientation offset vector, and a 3D visual instruction that matches the operation action intention is generated accordingly. It should be noted that the construction of dynamic 3D visual instructions includes three core dimensions: visual type, spatial orientation parameters, and dynamic effect templates. Visual type refers to the classification of basic visual elements, including direction indication (such as arrows), area focus (such as highlighted outlines), and operation demonstrations (such as rotation animations). Spatial orientation parameters are used to determine the orientation, rotation angle, and scaling of visual elements in three-dimensional space. Dynamic effect templates are used to define the movement mode, deformation rules, or state change logic of visual elements.

[0094] The choice of visual type is directly related to the encoding of the action intent. Specifically, the action_intent field in the data packet is parsed, and a mapping table is queried to determine the visual type as an action demonstration. This design ensures that the visual feedback aligns with the semantics of the action, such as a rotating animation for tightening rather than a static arrow.

[0095] The spatial orientation parameters corresponding to the 3D vision instructions of the determined visual type are configured based on the orientation offset vector. The orientation offset vector is used to represent the relative position of the target object in three-dimensional space, such as [0.1, -0.3, -0.2]. The spatial orientation parameters include the origin coordinates, the Euler angle rotation, and the axial scaling factor. From the pre-built template library, load the dynamic effect template associated with the core operation verb and inject the spatial orientation parameters into the dynamic effect template.

[0096] In one embodiment of the present application, when the visual type is a direction indication, the direction vector in the structured data packet is parsed, such as [-0.5, 0, 0] which means horizontally to the left. The direction vector is converted into the Euler angle rotation parameter of the three-dimensional space, and according to the X, Y, and Z components in the direction vector, the rotation angle around the Y axis, such as -30°, and the tilt angle around the Z axis, such as 0°, are calculated respectively. For example, when the direction vector is an equipment maintenance instruction to move the switch to the left, the Euler angle (0, -30°, 0) is output. The pulsation frequency parameter is associated with the urgency of the operation. The action intention code MOVE_URGENT triggers high-frequency pulsation, such as 5Hz, and the action intention code MOVE_NORMAL enables low-frequency pulsation, such as 1Hz.

[0097] Load a predefined arrow base mesh containing vertex or patch data structures, apply Euler angle rotation parameters to apply a rotation transformation matrix to the mesh vertices, and inject a pulsation frequency parameter to drive the vertices to periodically displace along the axis, creating a stretching animation. Notably, this design gives the arrow dual semantics of pulsation and pointing in space.

[0098] In the case of regional focus, the VLM outputs a pixel-level mask, such as a binary segmentation map of a rusted area. Specifically, the mask boundary point set is extracted to generate a closed polygonal outline, which is then extruded along the surface normal to create a 3D surface. The thickness is adaptive to the scene scale, such as generating a 2mm thick surface for rust spots on a device surface.

[0099] Dynamic transparency control enhances visualization based on target properties. If the command contains severe rust, the surface will enable edge breathing, for example, with transparency periodically changing from 50% to 80%. If the command contains minor scratches, a constant translucency mode is used, for example, at 70% transparency. This ensures that the highlighted area precisely fits irregular surfaces, such as rust grooves on gears, avoiding the visual distraction of traditional rectangular frames.

[0100] In the case where the visual type is an operation demonstration, the basic motion unit is disassembled according to the action intention coding. For rotation operations such as tightening the valve, a wrench model is generated to rotate 0° to 360° along the valve stem axis, and the axial propulsion displacement is superimposed, such as moving forward 3cm during rotation to simulate the screw-in effect, and the loop playback parameters are set, such as a single cycle of 2 seconds. For pressing operations such as starting the red button, a spherical indicator is created to displace along the normal direction, such as a pressing depth of 50%, and the rebound physical curve is bound, as well as quick pressing + slow release rebound, and a 1-second pause is added to the loop interval. It is worth noting that the motion sequence is completely parameter-driven and there is no need to pre-store animation resources.

[0101] In this embodiment, if the intended action is simple pointing or directional indication, a parameterized 3D arrow or line generator can be invoked. The arrow's starting point, end point, direction, length, thickness, color, and dynamic effects are all controlled by parameters, with dynamic effects such as flashing, pulsing, and growing animations. If the intended action is to highlight a specific area, a 3D highlight volume or surface shading effect can be generated that roughly matches the shape of the target visual anchor point.

[0102] If the action intention involves a simple manipulation such as rotation or translation, a simplified 3D animation sequence can be generated, for example, an arrow rotating around a target, or a schematic marker moving along a specific path.

[0103] Step 203: spatially anchor the 3D visual instructions to the real-time image of the XR device of the on-site user, and determine the rendering position of the 3D visual instructions based on the matching result, so as to render the 3D visual instructions in the real-time image of the XR device.

[0104] In this example, a set of local feature descriptors, such as texture keypoints and oriented gradient histograms of an oil-contaminated pipe joint, is extracted from the target mask region of the scene preview image displayed by the remote expert client. Regions similar to these feature descriptors are then searched for within the current scene preview frame on the on-site user's XR device. It should be noted that when the highest similarity exceeds a preset threshold, the center coordinates of the matching region are used as the anchor reference plane.

[0105] In one embodiment of the present application, the surface normal direction of the matching area is detected using an XR device depth sensor or multi-view geometry calculation. The rendering engine dynamically rotates the 3D visual instructions to always be perpendicular to the normal direction, such as an arrow fitting the surface of a pipe joint.

[0106] It can be understood that the perspective projection matrix of the 3D visual instructions is dynamically calculated based on the real-time 6DoF pose data of the on-site user from the IMU / visual SLAM. For example, when the user moves around the device, the arrow depth value is adaptively adjusted to avoid visual distortion.

[0107] Finally, the dynamic arrow instance is rendered to the anchor position, and its spiral animation maintains a clockwise rotation effect in the 2D projection. This solution replaces SLAM reconstruction with feature matching, significantly reducing computing power requirements and supporting real-time operation on mobile XR devices.

[0108] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides a remote expert guidance intention visualization device for XR devices, whose structure is as follows: Figure 5 shown.

[0109] Figure 5 Schematic diagram of the internal structure of the remote expert guidance intention visualization device for XR devices provided in the embodiment of this application. Figure 5 As shown, the equipment includes:

[0110] at least one processor;

[0111] and, a memory communicatively coupled to the at least one processor;

[0112] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:

[0113] Receiving multimodal instructions from a remote expert in combination with a real-time scene preview image; the multimodal instructions include at least one of the following: voice instructions, text instructions, or visual markup instructions;

[0114] Input the multimodal instructions and the associated scene preview images into the visual language model to generate a structured intent data packet. Based on the structured intent data packet, generate dynamic 3D visual instructions that match the operation action intent in the structured intent data packet.

[0115] The 3D visual instructions are spatially anchored with the real-time image of the XR device of the on-site user, and the rendering position of the 3D visual instructions is determined based on the matching results to render the 3D visual instructions in the real-time image of the XR device.

[0116] like Figure 6 As shown, the embodiment of this specification also provides a remote expert guidance intention visualization device for XR devices. Figure 6 It can be seen that in one or more embodiments of this specification, the remote expert guidance intention visualization device for XR devices, the device 600 includes:

[0117] The multimodal instruction receiving module 601 is used to receive multimodal instructions issued by the remote expert in combination with the real-time scene preview image; the multimodal instructions include at least one of the following: voice instructions, text instructions, or visual mark instructions;

[0118] A visual instruction generation module 602 is configured to input the multimodal instruction and the associated scene preview image into a visual language model to generate a structured intent data packet, and based on the structured intent data packet, generate a dynamic 3D visual instruction that matches the operation action intent in the structured intent data packet;

[0119] The visual instruction rendering module 603 is used to spatially anchor the 3D visual instruction with the real-time image of the XR device of the on-site user, and determine the rendering position of the 3D visual instruction based on the matching result to render the 3D visual instruction in the real-time image of the XR device.

[0120] In some embodiments, receiving a multimodal instruction from a remote expert in combination with a real-time scene preview image specifically includes:

[0121] Send the remote expert a real-time preview video stream of the scene shared by the on-site user's XR device;

[0122] capturing a voice command or text command input by a remote expert in combination with a scene preview image frame in a scene preview video stream, and upon receiving a voice command, converting the voice command into a text format to determine a command text of the voice command or text command;

[0123] Capturing visual marking instructions drawn by a remote expert on a scene preview image frame in a scene preview video stream; the visual marking instructions include the coordinates of a circled area contour line or a path trajectory point sequence;

[0124] The timestamps corresponding to the voice instructions or text instructions and the visual mark instructions are collected to bind the coordinate data of the instruction text and the visual mark instructions to the corresponding scene preview image frames respectively as a time-series associated multimodal input set.

[0125] In some embodiments, a multimodal instruction and an associated scene preview image are input into a visual language model to generate a structured intent data packet, specifically including:

[0126] Inputting the multimodal instruction and the associated scene preview image into a visual language model, locating the spatial orientation descriptor of the target object in the instruction text through the visual language model, parsing the spatial orientation descriptor, and determining the location information of the target object in the scene preview image; the location information includes at least one of the following: two-dimensional bounding box coordinates, pixel-level mask, or key points;

[0127] Extract the core action verbs and related directional adverbs in the instruction text, and parse the core action verbs and related directional adverbs into preset action intention codes;

[0128] Identify spatial relationship descriptors and generate an orientation offset vector of the target object relative to the scene reference object.

[0129] In some embodiments, parsing the spatial orientation descriptor to determine the location information of the target object in the scene preview image specifically includes:

[0130] Recognize the direction keywords and reference object names in the spatial orientation description words, and determine the spatial position of the scene reference object in the scene preview image;

[0131] Calculate the vector relationship between spatial position and direction keywords and output the positioning coordinates of the target object; the positioning coordinates include the direction vector.

[0132] In some embodiments, the dynamic 3D visual instruction includes a visual type, a spatial orientation parameter, and a dynamic effect template;

[0133] Based on the structured intent data packet, dynamic 3D visual instructions are generated that match the operation action intent in the structured intent data packet, including:

[0134] According to the action intention code in the structured intent data packet, the visual type corresponding to the 3D visual instruction to be generated is selected; the visual type includes direction indication, area focus or operation demonstration;

[0135] Based on the orientation offset vector in the structured intent data packet, configure the spatial orientation parameters of the 3D vision instructions to be generated for the vision type;

[0136] A dynamic effect template of a 3D visual instruction to be generated and associated with a core operation verb is loaded; the dynamic effect template includes pulse intensity, motion trajectory or deformation sequence.

[0137] In some embodiments, further comprising:

[0138] When the visual type is direction indication, the Euler angle rotation parameters and pulsation frequency parameters of the arrow in three-dimensional space are calculated according to the direction vector in the structured intent data packet;

[0139] Load the predefined arrow base mesh and apply the Euler angle rotation parameters and pulsation frequency parameters to the arrow base mesh to generate a dynamic arrow instance;

[0140] When the vision type is regional focus, a semi-transparent 3D highlight surface matching the target object’s outline is generated based on the pixel-level mask of the target object in the structured intent data packet.

[0141] When the visual type is an operation demonstration, a looping mechanical motion animation sequence is generated.

[0142] In some embodiments, spatially anchoring the 3D visual instructions to the real-time image of the on-site user's XR device includes:

[0143] In 3D vision instructions, a set of local feature descriptors of the target area in the scene preview image of the remote expert is extracted;

[0144] Searching for an area similar to the local feature descriptor in the real-time image of the on-site user, and when the matching similarity exceeds a preset matching threshold, taking the area with the highest matching similarity as the matching image area;

[0145] Determine the central coordinate plane of the matching image region and align the rendering origin of the 3D vision instructions to the central coordinate plane.

[0146] In some embodiments, further comprising:

[0147] Detect the surface normal direction of the matching image area in the real-time image of the on-site user and adjust the rendering direction of the 3D vision instruction so that the 3D vision instruction is always perpendicular to the surface normal direction;

[0148] Dynamically update the perspective projection parameters of 3D vision instructions based on the real-time pose data of the on-site user's XR device.

[0149] The present application also provides a non-volatile computer storage medium storing computer-executable instructions. When the computer-executable instructions are executed, they can:

[0150] Receiving multimodal instructions from a remote expert in combination with a real-time scene preview image; the multimodal instructions include at least one of the following: voice instructions, text instructions, or visual markup instructions;

[0151] Input the multimodal instructions and the associated scene preview images into the visual language model to generate a structured intent data packet. Based on the structured intent data packet, generate dynamic 3D visual instructions that match the operation action intent in the structured intent data packet.

[0152] The 3D visual instructions are spatially anchored with the real-time image of the XR device of the on-site user, and the rendering position of the 3D visual instructions is determined based on the matching results to render the 3D visual instructions in the real-time image of the XR device.

[0153] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.

[0154] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0155] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0156] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0157] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0159] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0160] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0161] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0162] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0163] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A remote expert guidance intention visualization method for XR devices, characterized by: The method comprises: Receiving multimodal instructions issued by a remote expert in combination with a real-time scene preview image; the multimodal instructions include at least one of the following: voice instructions, text instructions, or visual mark instructions; Inputting the multimodal instruction and the associated scene preview image into a visual language model to generate a structured intent data packet, and based on the structured intent data packet, generating a dynamic 3D visual instruction that matches the operation action intent in the structured intent data packet; Spatially anchoring the 3D visual instruction with the real-time image of the on-site user's XR device, and determining a rendering position of the 3D visual instruction based on the matching result, so as to render the 3D visual instruction in the real-time image of the XR device; Input the multimodal instruction and the associated scene preview image into the visual language model to generate a structured intent data packet, specifically including: Inputting the multimodal instruction and the associated scene preview image into a visual language model to locate spatial orientation descriptors of a target object in the instruction text using the visual language model, and parsing the spatial orientation descriptors to determine location information of the target object in the scene preview image; the location information includes at least one of the following: two-dimensional bounding box coordinates, pixel-level masks, or key points; Extracting core operational verbs and related directional adverbs from the instruction text, and parsing the core operational verbs and the related directional adverbs into preset action intention codes; The spatial orientation descriptor is identified, and an orientation offset vector of the target object relative to a scene reference object is generated.

2. The method for visualizing remote expert guidance intention for XR devices according to claim 1, characterized in that: Receive multimodal instructions from remote experts combined with real-time scene preview images, including: Send the remote expert a real-time preview video stream of the scene shared by the on-site user's XR device; capturing a voice command or text command input by the remote expert in combination with the scene preview image frame in the scene preview video stream, and upon receiving the voice command, converting the voice command into a text format to determine a command text of the voice command or the text command; Capturing a visual marking instruction drawn by the remote expert on the scene preview image frame in the scene preview video stream; the visual marking instruction includes the coordinates of the outline of the circled area or a sequence of path trajectory points; The timestamps corresponding to the voice instruction or the text instruction and the visual mark instruction are collected to bind the coordinate data of the instruction text and the visual mark instruction to the corresponding scene preview image frames respectively as a time-series associated multimodal input set.

3. The method for visualizing remote expert guidance intention for XR devices according to claim 1, characterized in that: Parsing the spatial orientation descriptor to determine the location information of the target object in the scene preview image specifically includes: Identifying the direction keyword and the reference object name in the spatial orientation description word, and determining the spatial position of the scene reference object in the scene preview image; The vector relationship between the spatial position and the direction keyword is calculated, and the positioning coordinates of the target object are output; the positioning coordinates include the direction vector.

4. The method for visualizing remote expert guidance intention for XR devices according to claim 1, characterized in that: The dynamic 3D visual instruction includes a visual type, a spatial orientation parameter and a dynamic effect template; Based on the structured intent data packet, generating a dynamic 3D visual instruction that matches the operation action intention in the structured intent data packet, specifically including: Selecting a visual type corresponding to the 3D visual instruction to be generated according to the action intention code in the structured intention data packet; the visual type includes direction indication, area focus, or operation demonstration; Configuring spatial orientation parameters of the 3D vision instructions to be generated of the vision type based on the orientation offset vector in the structured intent data packet; The dynamic effect template of the 3D visual instruction to be generated, which is associated with the core operation verb, is loaded; the dynamic effect template includes pulse intensity, motion trajectory or deformation sequence.

5. The method for visualizing remote expert guidance intention for XR devices according to claim 4, characterized in that: The method further comprises: When the visual type is direction indication, calculating the Euler angle rotation parameters and pulsation frequency parameters of the arrow in three-dimensional space according to the direction vector in the structured intent data packet; Loading a predefined arrow base mesh, and applying the Euler angle rotation parameters and the pulsation frequency parameters to the arrow base mesh to generate a dynamic arrow instance; In the case where the visual type is regional focus, generating a semi-transparent 3D highlight surface matching the outline of the target object based on the pixel-level mask of the target object in the structured intent data packet; When the visual type is an operation demonstration, a mechanical motion animation sequence is generated for looping playback.

6. The method for visualizing remote expert guidance intention for XR devices according to claim 1, characterized in that: Spatially anchoring the 3D visual instructions to the real-time image of the on-site user's XR device specifically includes: Extracting, from the 3D vision instruction, a set of local feature descriptors of a target area in the scene preview image by a remote expert; Searching for an area similar to the local feature descriptor in the real-time image of the on-site user, and when the matching similarity exceeds a preset matching threshold, taking the area with the highest matching similarity as the matching image area; A central coordinate plane of the matching image area is determined, and a rendering origin of the 3D vision instruction is aligned to the central coordinate plane.

7. The method for visualizing remote expert guidance intention for XR devices according to claim 6, characterized in that: The method further comprises: Detecting the surface normal direction of the matching image area in the real-time image of the on-site user, and adjusting the rendering direction of the 3D visual instruction so that the 3D visual instruction is always perpendicular to the surface normal direction; Dynamically update the perspective projection parameters of the 3D vision instruction based on the real-time posture data of the XR device of the on-site user.

8. Remote expert guidance intention visualization device for XR devices, characterized by: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the remote expert guidance intention visualization method for XR devices as described in any one of claims 1-7.

9. A remote expert guidance intention visualization device for XR devices, characterized in that: The device comprises: A multimodal instruction receiving module is used to receive multimodal instructions issued by a remote expert in combination with a real-time scene preview image; the multimodal instructions include at least one of the following: voice instructions, text instructions, or visual mark instructions; A visual instruction generation module, configured to input the multimodal instruction and the associated scene preview image into a visual language model to generate a structured intent data packet, and based on the structured intent data packet, generate a dynamic 3D visual instruction that matches the operation action intent in the structured intent data packet; Input the multimodal instruction and the associated scene preview image into the visual language model to generate a structured intent data packet, specifically including: Inputting the multimodal instruction and the associated scene preview image into a visual language model to locate spatial orientation descriptors of a target object in the instruction text using the visual language model, and parsing the spatial orientation descriptors to determine location information of the target object in the scene preview image; the location information includes at least one of the following: two-dimensional bounding box coordinates, pixel-level masks, or key points; Extracting core operational verbs and related directional adverbs from the instruction text, and parsing the core operational verbs and the related directional adverbs into preset action intention codes; Identifying the spatial orientation descriptor and generating an orientation offset vector of the target object relative to a scene reference object; The visual instruction rendering module is used to spatially anchor the 3D visual instruction with the real-time picture of the XR device of the on-site user, and determine the rendering position of the 3D visual instruction based on the matching result, so as to render the 3D visual instruction in the real-time picture of the XR device.

Citation Information

Patent Citations

  • Mixed reality visual guidance method for remote cooperation between trainee and expert

    CN116540879A

  • Structured visual positioning method, system and equipment based on intention recognition

    CN119848277A