An immersive robotic teleoperation method, system, device, and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-08-11
AI Technical Summary
(1)操作者负担重,系统无法自主执行高层语义任务:现有方案多为低层次直接控制(手柄摇杆映射速度/位姿,或控制器位姿映射机械臂姿态),操作者需持续手动控制机器人的每一个运动指令
Smart Images

Figure CN122547239A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of embodied intelligence, and in particular to an immersive robot teleoperation method, system, device, and medium. Background Technology
[0002] Robot teleoperation technology has evolved from direct control to intelligent interaction. Early solutions involved direct control of robot joints or speed via handheld devices / keyboards; subsequently, supervised teleoperation based on 2D video observation emerged; in recent years, the combination of VR headsets and 3D reconstruction technology has provided immersive visualization interfaces, but operators still need to continuously manually control the underlying movements. 3D Gaussian Splash (3DGS) technology, with its advantages of explicit representation, real-time rendering, and strong editability, is applied to robot environmental perception and is more suitable for real-time interactive scenarios compared to implicit representations such as NeRF. Existing research has achieved immersive visualization by streaming 3DGS reconstructions to VR headsets, but it has not yet formed a closed loop with semantic understanding, intelligent interaction, and automatic command generation.
[0003] Specifically, existing robot teleoperation and related technologies can be broadly categorized into the following three types: I. Low-level direct control: This type of system directly maps robot motion parameters (speed, joint angles, robotic arm posture, etc.) to the pose of a joystick, keyboard, or controller. The operator must continuously manually manipulate the underlying motion; the system does not understand higher-level semantic intent. Some solutions combine with VR headsets to provide first-person 3D image feedback, but the viewpoint is fixed to the robot's current position, and the operator cannot freely switch viewing positions. This category includes traditional teleoperation systems and motion-mapping VR teleoperations.
[0004] II. 3D Reconstruction and Visualization: This approach utilizes 3D Gaussian Splash (3DGS) or Neural Radiation Field (NeRF) to reconstruct high-fidelity 3D scenes, which are then transmitted to a VR headset for free-viewpoint navigation. The operator can observe the scene from any position, gaining an immersive spatial perception. However, in this type of solution, 3D reconstruction is only used as a visualization tool. Gaussian points only contain geometric and appearance information, lacking semantic labels, supporting semantic object interaction, and not integrated with the robot control loop. When the environment changes, the virtual scene remains in the old state and cannot be updated synchronously.
[0005] III. Semantic Understanding Operations: This approach combines a Visual Language Model (VLM) with 3DGS to construct a semantic Gaussian field, supporting open-vocabulary object queries or language-guided robot operations. However, this type of solution is typically limited to single grasping tasks for desktop-level fixed robotic arms. The interaction method is text-based language command input, lacking VR immersive interaction; it only outputs a single grasping posture or simple action, without involving the generation of multi-step task sequences (such as navigation → localization → operation → confirmation); it does not explicitly construct the topological spatial relationships between objects, nor does it consider the impact of environmental constraints on operational feasibility; the 3D semantic field is statically reconstructed, remaining unchanged before and after execution, without an update mechanism.
[0006] The above three types of solutions have the following problems: (1) Heavy operator burden and inability of the system to autonomously execute high-level semantic tasks: Existing solutions are mostly low-level direct control (handle joystick mapping speed / pose, or controller pose mapping robotic arm posture), requiring operators to continuously and manually control every motion command of the robot. The system does not understand the operator's high-level intentions and lacks the semantic understanding of object categories, states, and spatial relationships in the scene. It cannot autonomously decompose semantic-level tasks such as "closing the valve" into sub-tasks such as navigation, positioning, and operation and execute them. Operators need to visually identify targets and determine operation methods, and cannot interact with the scene through natural means such as clicking objects or using voice commands. This leads to high concentration of attention during long-term operation, operator fatigue, and a high rate of error in complex tasks.
[0007] (2) Fixed perspective and limited spatial perception: Traditional solutions only provide a fixed first-person perspective of the robot (2D video stream or fixed-position 3D image). The operator's perspective is limited to the area that the robot's camera is currently pointing at, and it is impossible to examine the whole situation from any spatial position, making it difficult to obtain a truly immersive spatial perception. Even if some solutions provide 3D reconstruction visualization, the operator is still limited to the vicinity of the robot's current position and cannot freely switch the observation position and angle, making it difficult to detect blind spot hazards.
[0008] (3) Reconstruction and execution are independent of each other, and the virtual and real worlds are not synchronized after environmental changes: Explicit 3D reconstruction techniques such as 3D Gaussian splashing are only used for static visualization and do not form a closed loop with robot control. Semantic understanding operations construct a static semantic field, which remains unchanged before and after execution. After the robot performs actions that cause environmental changes (such as objects being moved or doors being opened), the virtual scene remains in the old state, and what the operator sees is inconsistent with the actual environment, which can easily lead to misjudgment and operational errors. Some solutions require data to be re-collected and the model to be retrained after environmental changes, and there is no real-time update mechanism.
[0009] (4) Limited application scenarios: Semantic understanding operation classes are mostly limited to single grasping tasks of fixed robotic arms, and do not involve full-body teleoperation of mobile robots or multi-step task execution in complex industrial scenarios.
[0010] (5) Lack of natural interaction capabilities: Low-level direct control classes and 3D reconstruction visualization classes do not support natural interaction with the scene through clicking semantic objects, voice commands, etc. Although semantic understanding operation classes support language commands, they are only text input and lack VR immersive interaction. Operators cannot examine the whole scene from any perspective and then issue commands through intuitive interaction.
[0011] Chinese patent CN120578299B discloses an immersive robot teleoperation method and system. By constructing a first-person mapping model, an interactive data acquisition model, a teleoperation processing model, an execution perception model, and an immersive perspective simulation model, it acquires and processes robot-perspective images, obtains the operator's posture control information in real time, converts it into control commands in the robot's coordinate system, and generates immersive 3D perspective feedback, thus achieving robot posture control and motion image acquisition. However, the VR scene constructed by this solution lacks high-level semantic information such as object categories, states, and spatial relationships. It can only perform low-level pose manipulation based on experience. Therefore, the operation method of this solution is that the operator must generate robot control commands in real time through continuous posture changes, and completing a task requires the operator to manually execute all process actions. This direct control mode not only demands extremely high levels of skill, attention, and physical exertion from the operator, but also struggles to express high-level object-oriented task intentions such as "closing valve A," resulting in low operational efficiency and inability to handle complex and delicate tasks requiring multiple consecutive operations.
[0012] Therefore, there is currently a lack of an intent-based robotic teleoperation solution that can free operators from continuous manual low-level control and achieve a closed-loop teleoperation process of intent issuance, automatic execution, and real-time feedback. Summary of the Invention
[0013] The purpose of this invention is to provide an immersive robot teleoperation method, system, device and medium to address the above-mentioned deficiencies in the prior art, reduce the operator's burden, realize the automatic execution of high-level semantic tasks, break the perspective limitation, realize immersive global spatial perception, and integrate reconstruction and execution to ensure virtual-real synchronization.
[0014] The objective of this invention can be achieved through the following technical solutions: An immersive robot teleoperation method includes the following steps: The robot collects perception data of the operating environment based on its multi-source perception unit, and reconstructs the initial three-dimensional Gaussian field based on the perception data. Semantic embeddings of multi-view environmental images are extracted from the perceptual data and mapped to the three-dimensional Gaussian points corresponding to the initial three-dimensional Gaussian field. Semantic labels are assigned to each three-dimensional Gaussian point to obtain the semantic Gaussian field. The semantic Gaussian field is transmitted to the VR terminal to support the operator to roam freely in the virtual scene rendered by the VR terminal and observe the scene from any position, and to collect the target semantic object selected by the operator based on VR ray interaction. The operator's interaction intent towards the target semantic object is obtained directly or inferred based on the pose, proximity relationship and spatial constraints of the target semantic object in the semantic Gaussian field, and the interaction intent is verified by querying a predefined semantic-action knowledge base; Map the validated interaction intent into a multi-step operation instruction sequence; The robot is controlled to execute the multi-step operation instruction sequence in sequence; During robot execution, environmental change information is collected in real time, and the semantic Gaussian field is updated based on the environmental change information, so that the virtual scene rendered by the VR terminal keeps synchronized with the real operating environment in real time.
[0015] The method for directly obtaining the operator's interaction intent regarding the target semantic object is as follows: The system acquires the operator's voice input information, performs voice recognition on the voice input information to obtain text information, or acquires the text information manually input by the operator through the interactive interface. Extract the interaction intent of the corresponding target semantic object from the text information.
[0016] The method for inferring the operator's interaction intent regarding the target semantic object based on the pose, proximity relationships, and spatial constraints of the target semantic object in a semantic Gaussian field is as follows: In response to the VR ray interaction hitting the target semantic object, multi-dimensional context information associated with the target semantic object is obtained. The multi-dimensional context information includes at least object identifier, semantic category, pose, size, current state, neighboring object information, robot current state, and operation history. The multi-dimensional contextual information is organized into structured data and input into a pre-trained semantic reasoning model, which then generates at least one candidate interaction intent based on the structured data. The at least one candidate interaction intent is fed back to the operator, and the operator's selection or regeneration instruction is received through VR ray interaction. In response to the selection instruction, the selected candidate interaction intent is determined as the interaction intent of the target semantic object; In response to the regeneration instruction, the semantic reasoning model is triggered to regenerate at least one new candidate interaction intent and feed it back to the operator.
[0017] The virtual scene rendered by the VR terminal is constructed based on a pure 3D Gaussian field rendering mode and a hybrid rendering mode. Initially, the pure 3D Gaussian field rendering mode is used. The hybrid rendering mode is triggered in response to the robot being in an execution state, the operator actively switching to the hybrid rendering mode, or the selected target semantic object being marked as being in a changing state. The current pose of the camera mounted on the robot and the real-time images it captures are obtained, and the virtual camera pose that is closest to the current pose of the camera is determined in the semantic Gaussian field. A three-dimensional Gaussian sputtering image is generated based on the virtual camera pose. The three-dimensional Gaussian sputtering image is dynamically mixed with the real-time image captured by the camera to generate a mixed rendering image. In the static observation mode, the proportion of the three-dimensional Gaussian sputtering image is higher than that of the real-time image. In the dynamic operation mode, the proportion of the real-time image is higher than that of the three-dimensional Gaussian sputtering image. The dynamic proportion is manually adjusted by the operator through the adjustment control of the VR terminal. During the generation of the hybrid rendering image, pixel-level spatial correspondence between images is achieved by aligning camera intrinsic and extrinsic parameters, and Alpha fusion is used at the hybrid boundary to eliminate stitching gaps. When the conditions for triggering the hybrid rendering mode are eliminated, switch to the pure 3D Gaussian field rendering mode to restore the six-degree-of-freedom roaming; wherein, the conditions for triggering the hybrid rendering mode are eliminated as follows: the robot execution state ends and the target semantic object state is stable and without abnormality, the operator actively exits the dynamic mode, or the marker of the selected target semantic object is restored to a non-changing state; if an abnormal interruption occurs during robot execution, the hybrid rendering mode or the switch to the pure 3D Gaussian field rendering mode is determined according to the degree of environmental change.
[0018] The method for mapping validated interactive intents to a multi-step operation instruction sequence is as follows: Transform the verified interaction intents into structured task objectives; The structured task objective is input into the task planning model, which combines the robot's current state with visual observations obtained from the semantic Gaussian field to decompose the structured task objective into a sequence of behavioral subtasks. Motion planning verification is performed on the behavior-level subtask sequence, including instantiating each subtask into an action instruction containing control parameters, and sequentially performing kinematic feasibility checks, collision detection, and reachability checks. If any of the checks fails, an alternative subtask is generated until a multi-step operation instruction sequence that passes the checks is obtained.
[0019] During the process of controlling the robot to execute the multi-step operation instruction sequence in sequence, the execution progress is transmitted back to the VR terminal in real time for visualization. At the same time, the robot's operating status is monitored in real time through multiple sensors. When an anomaly is detected, different responses are triggered according to the anomaly level. The multiple sensors include an encoder, a torque sensor, and a vision sensor. The encoder provides feedback on the joint position, the torque sensor detects the contact force, and the vision sensor provides feedback on the relative pose of the end effector and the target semantic object.
[0020] The method for collecting environmental change information in real time and updating the semantic Gaussian field based on the environmental change information is as follows: After determining that the robot has successfully performed the action and the environment has changed, the predicted image rendered by the semantic Gaussian field under the current virtual camera pose is obtained and compared with the actual image captured by the camera on the robot to determine the range of the changed three-dimensional region, i.e. the changed region. Local incremental reconstruction is performed on the changed region, updating the 3D Gaussian point parameters within the changed region and freezing the 3D Gaussian point parameters outside the changed region. For new regions within the changed region that were not previously observed, semantic labels are assigned according to a hierarchical semantic determination strategy: first, the semantic labels of neighboring regions with existing semantic labels are queried; if the color and depth features of the new region and its neighboring regions meet a preset similarity condition, the semantic label of the neighboring region is used; if the preset similarity condition is not met, the distribution of semantic labels in the surrounding neighboring regions is statistically analyzed, and the semantic label of the new region is determined according to the majority principle; if the degree of semantic label mixing in the surrounding neighboring regions exceeds a preset threshold or the majority principle still fails to determine the label, the visual language model is triggered to re-identify the new region and generate a semantic label. After the update is completed, the predicted pose of the semantic object in the virtual scene is compared with the measured pose of the corresponding object in the environment. If the deviation between the two exceeds the preset threshold, the area where the deviation occurs is locally corrected. If the local correction cannot meet the accuracy requirements, global realignment is triggered.
[0021] An immersive robotic teleoperation system includes: The three-dimensional Gaussian field reconstruction module is used to collect perception data of the operating environment based on the multi-source perception unit on the robot, and reconstruct the initial three-dimensional Gaussian field based on the perception data. The semantic mapping module is used to extract the semantic embedding of multi-view environmental images in the perceptual data and map it to the three-dimensional Gaussian points corresponding to the initial three-dimensional Gaussian field, and assign semantic labels to each three-dimensional Gaussian point to obtain the semantic Gaussian field. The VR interaction module is used to transmit the semantic Gaussian field to the VR terminal to support the operator to roam in the virtual scene rendered by the VR terminal with six degrees of freedom and observe the scene from any position, and to collect the target semantic object selected by the operator based on VR ray interaction. The intent acquisition and verification module is used to directly acquire or infer the operator's interaction intent for the target semantic object based on the pose, proximity relationship and spatial constraints of the target semantic object in the semantic Gaussian field, and to verify the interaction intent by querying a predefined semantic-action knowledge base; The instruction mapping module is used to map verified interactive intents into a multi-step operation instruction sequence; The control module is used to control the robot to execute the multi-step operation instruction sequence in sequence; The scene update module is used to collect environmental change information in real time during robot execution and update the semantic Gaussian field based on the environmental change information, so that the virtual scene rendered by the VR terminal keeps synchronized with the real operating environment in real time.
[0022] An electronic device includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the immersive robot teleoperation method.
[0023] A computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the immersive robot teleoperation method described above.
[0024] Compared with the prior art, the present invention has the following beneficial effects: (1) In view of the problem that the existing solutions have a heavy burden on operators and the system cannot autonomously perform high-level semantic tasks, the present invention constructs a semantic Gaussian field, which enables the system to have the semantic understanding of the object categories, states and spatial relationships in the scene. The operator only needs to click on the target semantic object through VR ray or express the intention through voice input. The system can automatically infer and verify the interaction intention, and then map it into a multi-step operation instruction sequence consisting of sub-tasks such as navigation, positioning and operation. There is no need to continuously manually control the underlying motion parameters, which transforms the operator from a low-level operator to a high-level decision-maker, thereby freeing him from continuous manual low-level operation, significantly reducing cognitive and physical burden, and improving the efficiency of complex task execution.
[0025] (2) In view of the problem that the existing solutions have fixed perspective and limited spatial perception, the present invention transmits the semantic Gaussian field to the VR terminal, which allows the operator to roam in the reconstructed virtual scene with six degrees of freedom and observe the scene from any position and angle (such as looking down from above the scene or observing from the side at close range). This breaks through the limitation of the traditional solution that only provides the robot with a fixed first-person perspective, enabling the operator to freely examine the global space, obtain true immersive spatial perception, and facilitate timely detection of blind spot hazards.
[0026] (3) In view of the problem that the reconstruction and execution of the existing scheme are independent and the virtual and real are not synchronized after the environment changes, the present invention collects environmental change information in real time during the robot execution, determines the changed area by comparing the predicted image with the actual image, and performs local incremental reconstruction and hierarchical semantic label update to complete the real-time synchronization of the virtual scene and the real environment, ensuring that the operator always observes the same environmental state as the actual one, and avoids misjudgment and operation error caused by virtual and real deviation.
[0027] (4) In view of the problem that the application scenarios of existing solutions are limited and lack natural interaction capabilities, this invention directly acts on the semantic objects in the semantic Gaussian field through natural interaction methods such as VR ray interaction and voice input, realizing the whole-body remote operation of the robot and the closed loop of multi-step complex tasks. The interaction method is intuitive and efficient, applicable to a wide range of industrial scenarios, and makes up for the shortcomings of existing semantic understanding operation classes that are limited to a single grasping task of a fixed robotic arm and lack VR immersive interaction. Attached Figure Description
[0028] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a flowchart of the semantic Gaussian field construction process of the present invention; Figure 3 This is a schematic diagram of the VR immersive arbitrary viewpoint interaction of the present invention; Figure 4 This is a schematic diagram of the virtual-real consistency closed-loop update of the present invention; Figure 5 This is a system architecture diagram of the present invention; Figure 6 This is a schematic diagram of the electronic device structure of the present invention. Detailed Implementation
[0029] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0030] It should be noted that, unless otherwise specified, the features in the following embodiments and implementation methods can be combined with each other.
[0031] Example 1
[0032] Figure 1 This is a flowchart illustrating an immersive robot teleoperation method according to an embodiment of the present invention. The robots mentioned in this embodiment include, but are not limited to, wheeled / tracked robots, quadrupedal robotic dogs, unmanned ground vehicles (UGVs), agricultural robots, unmanned aerial vehicles (UAVs), and fixed robotic arms. Figure 1 As shown, the immersive robot teleoperation method in this embodiment of the invention may include the following steps: S1: Based on the multi-source sensing unit mounted on the robot, the robot collects sensing data of the environment to be operated and reconstructs the initial three-dimensional Gaussian field based on the sensing data.
[0033] The robot is equipped with multi-view cameras and LiDAR to collect environmental data. The multi-view camera array is installed on the robot body in a surrounding or forward multi-view distribution to collect multi-view RGB images covering the operating area. The LiDAR is used to obtain dense 3D point clouds of the operating environment, providing accurate geometric priors and depth constraints for subsequent 3D reconstruction.
[0034] After entering the operating environment, the robot first executes an environmental scanning procedure. During the scanning process, the robot moves itself or rotates its gimbal to drive the multi-source sensing units to traverse the operating area, simultaneously acquiring multi-view RGB image sequences and LiDAR point cloud frames. The system utilizes simultaneous localization and mapping (SLAM) technology to estimate the camera pose of the robot at each acquisition moment in real time, and registers multiple frame point clouds to a unified world coordinate system, forming a global geometric prior for the operating environment.
[0035] In selecting the representation of the 3D scene, this embodiment preferably uses 3D Gaussian Splatting (3DGS) technology for scene reconstruction. 3DGS represents the scene as an explicit 3D Gaussian point cloud, where each Gaussian point contains parameters such as 3D position, covariance matrix, spherical harmonic coefficients, and opacity. Compared to implicit representation methods such as neural radiation fields, the explicit point cloud structure of 3DGS has the following advantages: First, it is highly editable; Gaussian points can be independently added, deleted, or modified, facilitating subsequent local incremental updates that only modify changed areas without affecting the global picture. Second, it has high real-time rendering efficiency; Gaussian points can be directly projected onto the image plane through a tile-based rasterization rendering pipeline, meeting the real-time requirements of high frame rate rendering in VR interactive scenes. Third, it is consistent with the point cloud data format, facilitating fusion and alignment with LiDAR point clouds. In another embodiment, 4D Gaussian Splatting can also be used instead of 3DGS. This technology is commonly used by those skilled in the art, and to avoid obscuring the purpose of this application, it will not be elaborated upon here.
[0036] Considering the computing power limitations of edge computing devices on the robot and the computing power advantages of high-performance GPU clusters in the cloud, this embodiment adopts a hierarchical processing strategy of edge-cloud collaboration. The specific division of labor is as follows: Sparse inference is performed at the edge: Embedded high-performance computing units (such as NVIDIA Jetson AGX Orin) deployed on the robot body receive the initial three-dimensional Gaussian field model parameters sent from the cloud, perform forward inference based on the robot's current pose, and generate a three-dimensional Gaussian sputtering rendering image from the current viewpoint, which is then used for subsequent mixed rendering with the real-time image from the camera.
[0037] Complete training is performed in the cloud: A high-performance GPU server in the cloud receives multi-view RGB image sequences, LiDAR point cloud frames, and corresponding camera poses uploaded by the robot, and performs complete 3D Gaussian field training. The training process first generates sparse point clouds from the multi-view images using a motion structure recovery method as the initial Gaussian point locations. Then, through differentiable rasterization rendering and photometric loss optimization, all parameters of each Gaussian point are iteratively updated until the difference between the rendered image and the real image converges to below a preset threshold. The LiDAR point cloud participates in training in two ways: first, by providing accurate point cloud location priors during the initialization phase to accelerate training convergence; and second, by serving as a depth constraint during optimization to improve the accuracy of reconstructed geometry. After training, an initial 3D Gaussian field is obtained, where each Gaussian point contains geometric and appearance information to support high-fidelity rendering for VR terminals.
[0038] That is, sparse inference is performed at the edge to ensure real-time performance, while complete training is performed in the cloud to ensure accuracy.
[0039] In one embodiment, the LiDAR can also be replaced by an RGB-D camera. An RGB-D camera can simultaneously acquire color and depth images. By aligning the depth map with the intrinsic parameters of the RGB image, it can provide 3D position information for each pixel. In this scheme, the system uses a multi-view RGB-D image sequence acquired by the RGB-D camera to backproject the depth map into a 3D point cloud, replacing the LiDAR point cloud as a geometric prior for initialization and depth constraints in 3D Gaussian field training. This alternative has the advantages of low cost and small size, making it suitable for robot platforms with strict cost or payload constraints; however, its depth measurement range and accuracy are limited by the technical parameters of the RGB-D camera, and its performance may be inferior to LiDAR in outdoor strong light or long-distance scenarios. Therefore, it can be flexibly selected according to the actual deployment environment.
[0040] S2, extract the semantic embedding of multi-view environmental images from the perceptual data, map it to the three-dimensional Gaussian points corresponding to the initial three-dimensional Gaussian field, assign semantic labels to each three-dimensional Gaussian point, and obtain the semantic Gaussian field.
[0041] Based on the initial three-dimensional Gaussian field obtained in step S1, the goal of step S2 is to assign a semantic label to each Gaussian point in the three-dimensional Gaussian field, so that it is upgraded from a simple geometric and appearance representation to a structured scene model with semantic information, namely a semantic Gaussian field.
[0042] Specifically, this embodiment utilizes a pre-trained visual language model to extract the semantic embedding of each pixel in a multi-view environmental image. The visual language model preferably employs a model with open-vocabulary recognition capabilities (such as CLIP, DINOv2, and other basic visual models). Its advantage lies in not being limited to a predefined set of closed categories, but being able to recognize arbitrary semantic concepts through text guidance, thereby supporting object queries and interactions using open-vocabulary concepts.
[0043] like Figure 2 As shown, the process specifically includes the following steps: First, the multi-view environmental images acquired in step S1 are input frame by frame into the image encoder of the visual language model to extract the feature vector of each pixel or image patch as a semantic embedding. This semantic embedding maps the visual appearance of a pixel to a high-dimensional feature space, in which pixels with similar semantics have similar embedding representations.
[0044] Subsequently, the semantic embedding of the 2D image needs to be mapped to 3D Gaussian points. Since the projection relationship between the 3D Gaussian points and the pixels of the multi-view image has been established in step S1 (i.e., each Gaussian point can be back-projected to the corresponding pixel position in multiple frames of the image), this embodiment achieves a reliable mapping of 2D semantic embedding to 3D semantic labels through a multi-view consistency voting mechanism. The specific process is as follows: Step 1, Single-frame projection: For each Gaussian point in the 3D Gaussian field, based on the camera poses of each frame known in step S1, the Gaussian point is projected onto the 2D image plane where the point is visible in each frame to obtain the corresponding pixel coordinates, and the feature vector of the pixel position is extracted from the semantic embedding map of the frame.
[0045] The second step is multi-view aggregation: Since a Gaussian point is usually observed by multiple frames of images, we collect all the semantic embedding vectors corresponding to the Gaussian point in each visible frame to form a semantic embedding set.
[0046] The third step, consensus voting, involves performing clustering analysis or similarity matching with predefined semantic category prototypes on the collected semantic embedding set. Cosine similarity is calculated by comparing the text embeddings of each candidate category in the open lexical semantic space, and the matching frequency of each candidate category across all observation frames is counted. If a semantic category is matched with the highest similarity in more than a preset threshold (e.g., 70%) of the observation frames, then that category is confirmed as the semantic label for that Gaussian point. If multiple categories have similar vote counts or the highest vote count does not reach the threshold, then the Gaussian point is marked as semantically undetermined, and can be further determined through spatial adjacency relationships or visual language model re-inference.
[0047] The fourth step is label assignment: the semantic label of the Gaussian point determined by consensus voting is written into the attribute field of the Gaussian point to form a structured semantic Gaussian point that simultaneously contains three-dimensional position, covariance matrix, spherical harmonic coefficient, opacity and semantic label.
[0048] Through the above steps, the initial 3D Gaussian field is upgraded to a semantic Gaussian field. This semantic Gaussian field simultaneously possesses the following three types of attributes: first, geometric attributes (the 3D position and covariance of Gaussian points, representing the spatial geometric structure); second, appearance attributes (spherical harmonic coefficients and opacity, supporting high-fidelity rendering from any viewpoint); and third, semantic attributes (semantic tags, supporting semantic-level object recognition and interaction).
[0049] Once the semantic Gaussian field is constructed, the operator can not only see a high-fidelity 3D scene rendering in the VR terminal, but also obtain the semantic label information of the corresponding Gaussian points in real time by hovering or clicking on any area of the scene using VR rays. For example, if the operator clicks on a pipe component in the scene, the system can immediately provide information such as the semantic label of the component as "valve A" and its current status. The open vocabulary query capability supported by the semantic Gaussian field means that the operator can perform semantic retrieval of the scene by inputting any natural language words (such as "find all red valves"). The system calculates the similarity between the text and the semantic embedding of each Gaussian point in the semantic embedding space, quickly locating and highlighting the target object.
[0050] S3, the semantic Gaussian field is transmitted to the VR terminal to support the operator to roam freely in the virtual scene rendered by the VR terminal and observe the scene from any position, and to collect the target semantic object selected by the operator based on VR ray interaction.
[0051] After the semantic Gaussian field is constructed in step S2, the cloud streams the semantic Gaussian field model parameters to the VR terminal via low-latency networks such as 5G / WiFi. Upon receiving the data, the VR terminal performs real-time rendering locally, presenting the operator with a 3D virtual scene that simultaneously possesses geometric, aesthetic, and semantic attributes. In this embodiment, the VR terminal can be a head-mounted VR display device, AR glasses supporting 3D scene display, or the AR mode of a tablet computer, to adapt to the deployment needs of different application scenarios.
[0052] like Figure 3 As shown, after entering the virtual scene wearing a VR terminal, the operator can freely roam within the space scanned and reconstructed by the robot, achieving six degrees of freedom. Six degrees of freedom refers to the operator's ability to freely change their observation position and perspective within the virtual scene along three spatial translation axes (forward / backward, left / right, up / down) and three rotation axes (pitch, yaw, roll). This means that the operator can not only observe the scene from a first-person perspective from the robot's current position, but also detach from the robot and move freely to the top of the scene for a global overview, to the side for close-up observation of specific equipment details, or to any corner to check blind spots, thus gaining an immersive spatial perception capability that transcends physical constraints.
[0053] In this step, a dynamic scene hybrid rendering mechanism is employed, which constructs a virtual scene based on a pure 3D Gaussian field rendering mode and a hybrid rendering mode. Specifically, When the robot has not yet begun executing commands in the real environment, the virtual scene rendered by the VR terminal defaults to a pure 3D Gaussian field rendering mode. In this mode, the scene observed by the operator is entirely generated by semantic Gaussian field rendering, possessing full six-degree-of-freedom roaming capabilities, allowing the operator to observe the entire operating environment from any perspective.
[0054] The hybrid rendering mode is automatically or manually triggered when any of the following preset trigger conditions are met: 1) The robot receives a multi-step operation instruction sequence and enters the execution state; 2) The operator actively switches to the hybrid rendering mode through the VR terminal interactive interface; 3) The target semantic object selected by the operator is marked as changing by the system.
[0055] The specific execution flow of the hybrid rendering mode is as follows: First, the current pose of the robot's camera and its real-time captured images are obtained. Then, the virtual camera pose closest to the current pose is determined within a semantic Gaussian field. Based on this virtual camera pose, a 3D Gaussian sputtering image is generated. This 3D Gaussian sputtering image is a predicted image rendered from the semantic Gaussian field according to the virtual camera's perspective using the 3DGS rasterization rendering pipeline. It reflects the system's prediction of the scene to be observed from the current perspective based on an existing scene model.
[0056] Subsequently, the 3D Gaussian sputtering image and the real-time image captured by the camera are dynamically alpha-blended to generate a mixed rendering image, which is then presented on the VR terminal. The blending ratio is set according to the following principles: In static observation mode, the operator primarily focuses on the spatial geometric consistency of the overall scene, with the 3D Gaussian sputtering image occupying a higher proportion (e.g., 70% to 90%) and the real-time image a lower proportion, ensuring the operator's perception of the continuity of the global space; in dynamic operation mode, the robot is performing fine maneuvers, and the operator needs to clearly observe the real-time details of the execution process, with the real-time image occupying a higher proportion (e.g., 60% to 80%) and the 3D Gaussian sputtering image a lower proportion, ensuring the operator can accurately judge the precision of the executed actions. In one embodiment, a ratio adjustment control (e.g., a virtual slider) can also be set in the VR interactive interface. The operator can manually drag the slider according to their observation needs to adjust the blending ratio of the 3D Gaussian sputtering image and the real-time image in real time to obtain the most suitable observation effect for the current task.
[0057] During the generation of the blended image, pixel-level spatial correspondence between the 3D Gaussian sputtering image and the real-time image is achieved through camera intrinsic and extrinsic parameter alignment. Intrinsic parameter alignment ensures that the two images have consistent focal length and principal point, while extrinsic parameter alignment ensures that the two images are rendered and acquired based on the same camera pose. Furthermore, for any seams that may arise at the boundary between the two images due to differences in rendering range or viewpoint, alpha fusion processing is employed. This involves a smooth, gradual transition of transparency in the boundary transition area based on pixel distance, eliminating visual seam artifacts and presenting a natural and unified blended image.
[0058] When the conditions for triggering the hybrid rendering mode are eliminated, the system automatically switches to the pure 3D Gaussian field rendering mode to restore the six-degree-of-freedom roaming. The conditions for triggering the hybrid rendering mode include: the robot execution state ends and the target semantic object state is stable and without abnormalities, the operator actively exits the dynamic mode, or the marker of the selected target semantic object is restored to a non-changing state.
[0059] If an abnormal interruption occurs during robot execution (such as collision alarm, action execution failure, communication interruption, etc.), the rendering mode is not switched immediately. Instead, the judgment is made based on the actual degree of environmental change: if the degree of environmental change is small and does not affect subsequent operation decisions, the pure 3D Gaussian field rendering mode can be switched so that the operator can troubleshoot the problem from a global perspective; if the degree of environmental change is large and the virtual-real deviation is significant, the hybrid rendering mode is maintained so that the operator can compare the virtual prediction screen with the real real screen at the same time, quickly identify the cause of the abnormality and take countermeasures.
[0060] like Figure 3 As shown, during the operator's observation of the virtual scene, the VR ray interaction mechanism collects the operator's intention to select a target semantic object. Specifically, the operator triggers a virtual ray through a controller button or gesture. This ray performs collision detection with a semantic Gaussian field in the virtual scene: when the ray intersects with a 3D Gaussian point contained in a semantic object in the semantic Gaussian field, a hit determination is triggered, returning the semantic object identifier of the hit object and its 3D pose in the virtual scene. At the same time, the object is highlighted in the VR interface (e.g., outline glow or color change), and its semantic label and brief operation prompt (e.g., "Valve A - Current Status: Open") are displayed near the object as floating labels. The operator can confirm the selection of the semantic object by releasing the ray or clicking again. This selection information will serve as the input trigger condition for intention inference in the subsequent step S4.
[0061] S4 directly obtains or infers the operator's interaction intent for the target semantic object based on the pose, proximity relationship and spatial constraints of the target semantic object in the semantic Gaussian field, and verifies the interaction intent by querying a predefined semantic-action knowledge base.
[0062] This embodiment provides two paths for obtaining interactive intent: a direct acquisition path and an inferred acquisition path, which can be used independently or as a complement to each other.
[0063] Path 1: Obtain directly.
[0064] Direct path acquisition allows operators to clearly express their operational intentions through voice or manual input, without requiring system inference.
[0065] When the operator uses voice input, the microphone array integrated into the VR terminal collects the operator's voice input in real time. The collected audio stream is processed by an automatic speech recognition engine deployed in the cloud, converting the speech signal into text information. The speech recognition engine preferably uses a pre-trained model with an end-to-end architecture and performs domain-adaptive fine-tuning for common command words in industrial operation scenarios (such as "close," "open," "rotate," "grab," "release," etc.) to improve recognition accuracy. This embodiment does not limit the specific model used; any existing speech recognition model can meet the requirements of this embodiment.
[0066] When the operator uses manual input, the VR terminal interface provides a virtual text input panel. The operator can directly obtain the text information entered by clicking the virtual keyboard with the controller or by using gesture recognition.
[0067] After obtaining the text information, the Natural Language Understanding (NLE) module extracts semantic elements from the text. Based on a pre-trained large language model, the NLE module accurately identifies the components of the operational intent from natural language expressions, including the target object name and the type of operation. Specifically, it links the text information with the semantic labels of each semantic object in a semantic Gaussian field, matching the object names mentioned in the text to determine the target semantic object of the operation; simultaneously, it extracts action words (such as "close," "tighten," "press," etc.) from the text, and after action normalization mapping, forms a structured interaction intent. For example, if the operator says "Turn off valve A," the interaction intent extracted after speech recognition and semantic understanding is: target object = "valve A," operation action = "close."
[0068] Path 2: Inference and Acquisition.
[0069] When the operator selects the target semantic object by clicking with a VR ray without providing a clear voice or text command, the inference acquisition path is initiated, and the operation intention is automatically inferred based on the multi-dimensional contextual information of the target semantic object.
[0070] In response to a VR raycasting interaction hitting a target semantic object, multi-dimensional contextual information associated with that target semantic object is retrieved from a semantic Gaussian field. This multi-dimensional contextual information is organized into a structured data record, including at least the following fields: 1. Object Identifier: A unique identifier for the object within the semantic Gaussian field; 2. Semantic Category: The semantic tag category of the object (such as "valve", "button", "switch", etc.); 3. Pose: The six-degree-of-freedom pose of the object in three-dimensional space; 4. Dimensions: The spatial bounding box dimensions of the object, used to assess the movement space required for operation; 5. Current status: The operable status attributes of the object (such as "on", "off", "locked", "faulted" etc.). This status information is provided by the Gaussian point status attribute field associated with the object in the semantic Gaussian field, or maintained in the previous operation history. 6. Neighboring object information: Identifiers, categories, and relative orientations of other semantic objects whose spatial distance is within a preset threshold range from this object; 7. Current robot status: current base position of the robot, current configuration of the robotic arm, type and status of the end effector, etc. 8. Operation History: A sequence record of recent operation steps, forming short-term task memory, used to identify the task context in continuous operations.
[0071] The structured data is input into a pre-trained semantic reasoning model, preferably a pre-trained large language model or a multimodal large model, which possesses strong common-sense reasoning capabilities and an understanding of industrial operation scenarios. The model performs semantic reasoning based on the multi-dimensional contextual information of the input, comprehensively analyzing factors such as object category, current state, surrounding environment, and operation history to generate at least one candidate interaction intent. For example, if the operator clicks to select "valve A," the structured data shows its current state as "open," and the semantic category as "valve." The model will infer that the most likely candidate interaction intent is "close valve A," and may simultaneously generate suboptimal candidate interaction intents such as "detect valve A" or "lock valve A," forming a list of candidate interaction intents.
[0072] In an alternative embodiment, when the semantic categories of the application scenario are limited and the operational logic is relatively fixed (e.g., the operating environment contains only a few categories of objects such as valves, buttons, and switches), a rule engine lookup table approach can be used instead of the semantic reasoning model. Specifically, an action mapping table is pre-built, using "semantic category + current state" as a composite index key to directly map to the corresponding preset operation action. For example, the mapping table stores rules such as: "valve + open → close", "valve + close → open", "button + standby → press to activate", etc. This method has fast reasoning speed and strong deterministic results, making it suitable for scenarios with a high degree of standardization in operational logic; however, its flexibility is limited, and it is difficult to handle complex multi-factor comprehensive reasoning scenarios.
[0073] The semantic reasoning model generates a list of candidate interaction intents, which are then fed back to the VR terminal and presented to the operator as selectable options on the VR interface. Each candidate intent is displayed as an independent virtual button or card in front of the operator's field of vision. The operator makes a selection via VR ray interaction: if an option in the candidate list matches their original interaction intent, the operator clicks that option to issue a selection command, and the system determines the selected candidate interaction intent as the final interaction intent. If none of the options in the candidate list meet the operator's expectations, the operator can click the regenerate button on the interface to issue a regeneration command. In response to the regeneration command, the semantic reasoning model is triggered to make internal adjustments and regenerate at least one new candidate interaction intent, which is then fed back to the operator for selection again, until the operator confirms or the preset maximum number of retries is reached.
[0074] Regardless of whether the interaction intent is obtained through direct path retrieval or inference, it must undergo compliance verification through a predefined semantic-action knowledge base before being used for subsequent instruction generation. This step matches the final selected interaction intent against the rules in the semantic-action knowledge base item by item: if the operation action in the interaction intent belongs to the allowed action set for the object category and current state, and all parameters are within the legal range and meet all safety constraints, then the verification is deemed successful, and the verified interaction intent is output for use in subsequent step S5. If the interaction intent conflicts with any rule, an alarm is triggered, the operator is notified of the conflict reason on the VR terminal interface, and the processing of this interaction intent is terminated. The operator can adjust the intent according to the alarm information and resubmit it, or select other candidate intents for verification.
[0075] S5 maps the verified interactive intent into a multi-step operation instruction sequence.
[0076] S51 transforms the verified interaction intent into a structured task objective.
[0077] First, the validated interaction intent is transformed into a structured task objective. This structured task objective is a machine-parseable data structure that contains at least the following fields: target object identifier (a unique identifier in a semantic Gaussian field), target object current state, and desired target state. For example, the structured task objective transformed from the interaction intent "Close valve A" is: {Target object: Valve A; Current state: Open; Target state: Closed}. This structured format provides a clear input interface for subsequent task planning models.
[0078] S52, the structured task objective is input into the task planning model, which combines the robot's current state and visual observations obtained from the semantic Gaussian field to decompose the structured task objective into a sequence of behavioral sub-tasks.
[0079] Subsequently, the structured task objective is input into the task planning model. In this embodiment, the task planning model preferably adopts a pre-trained visual language action model, i.e., a VLA model. A VLA model is an end-to-end model capable of simultaneously processing visual observations and language commands, and outputting robot action sequences. Typical examples include the GR00T N1 model and the OpenVLA model. These models, pre-trained with large-scale robot operation data, possess the generalization ability to decompose high-level language commands into multi-step action sequences.
[0080] During the inference phase, the VLA model receives two types of input: first, a structured task objective, provided in the form of text or structured prompts; and second, the robot's current state and visual observations from a semantic Gaussian field. The visual observations are not limited to 2D images, but include a multi-view 3D scene representation rendered from the semantic Gaussian field based on the robot's current pose. This representation contains rich spatial information such as the 3D pose of the manipulated object, the spatial relationships of neighboring objects, and traversable areas.
[0081] Based on the above input, the VLA model internally performs mapping reasoning from the task semantic space to the action space, decomposing the high-level task into a sequence of behavioral subtasks. The VLA model already has the ability to generate multi-step action sequences from visual and linguistic inputs. The specific task decomposition and control parameterization logic are completed internally by the VLA model. This embodiment directly utilizes this capability, and the specific reasoning process will not be described in detail.
[0082] This embodiment takes "closing valve A" as an example. The complete execution of this task requires crossing multiple behavioral stages. The VLA model decomposes it into the following typical seven-step behavioral sub-task sequence: Navigation to the working point: Based on the three-dimensional pose of valve A in the semantic Gaussian field, determine the path between the robot's current base position and the target operating position, generate chassis movement commands, and move the robot to the reachable operating range of the robotic arm; Robotic arm positioning: Move the end effector of the robotic arm to a pre-grabbing position at a preset distance above the handle of valve A. This position is obtained from the parameter template in the semantic-action knowledge base based on the semantic category and pose of valve A. Clamping handle: Controls the closing of the end gripper to grasp the valve A handle. The force control parameters are set according to the force / torque safety constraints of the valve category in the knowledge base. Rotary valve: Control the robotic arm wrist joint or end effector to perform rotational movements along the valve's rotation axis according to the rotation direction and angle specified in the task objective; Release: After rotation to the desired position, the control end gripper releases, releasing the valve handle; Retraction: Control the end effector of the robotic arm to move in the opposite direction of entry to a safe retraction position to avoid collision with valves and surrounding structures; Photo Confirmation: Control the robot's head or its onboard camera to take a picture of valve area A at a preset angle, and obtain a scene image after execution for subsequent change detection and status confirmation.
[0083] Each behavior-level subtask is further instantiated within the model as a motion instruction containing specific control parameters. For example, "robotic arm positioning" is instantiated as the target pose coordinates of the robotic arm end effector in Cartesian space, and "rotary valve" is instantiated as a parameterized motion instruction containing rotation axis vector, rotation angle, upper limit of angular velocity, and upper limit of torque.
[0084] S53, perform motion planning verification on the behavior-level subtask sequence, including instantiating each subtask into action instructions containing control parameters, and sequentially performing kinematic feasibility checks, collision detection, and reachability checks. If any of the checks fails, an alternative subtask is generated until a multi-step operation instruction sequence that passes the checks is obtained.
[0085] After the sequence of behavior-level subtasks is generated, motion planning verification is performed on each subtask sequentially to ensure its safe execution within the current environmental constraints and robot capabilities. The verification process includes three checks: kinematic feasibility check, collision detection, and reachability check. All three checks are performed using existing methods, and the execution process will not be elaborated upon in this embodiment. If all three checks pass, the subtask and its instantiated action instructions are confirmed as executable instructions and incorporated into the final multi-step operation instruction sequence. If any check fails, an alternative subtask generation mechanism is triggered: the reason for failure is input as negative feedback into the VLA model, which internally infers and generates alternative behavior strategies. For example, if the "robotic arm positioning" subtask fails because the target position is outside the current reachable range, the VLA model may generate an alternative solution, splitting the original "robotic arm positioning" into two subtasks: "chassis fine-tuning to a better position + robotic arm positioning." If the "rotary valve" fails due to collision detection, the VLA model may adjust the robotic arm entry angle or use a different end effector configuration. After the alternative subtask is generated, the motion planning verification loop is re-entered until a multi-step operation instruction sequence in which all subtasks have passed the check is obtained, or the preset maximum number of retries is reached, at which point an error is prompted to the operator.
[0086] It should be noted that existing VLA large-scale model solutions have the ability to generate multi-step action sequences from visual and verbal commands, but the input relies on verbal descriptions. Operators need to convert spatial intentions into text commands, and the 2D image-based input lacks immersive 3D spatial perception. The core difference in this approach lies in the input method and spatial perception—operators can freely roam in a 3D semantic field with six degrees of freedom through a VR headset, observe global spatial relationships from any perspective, and directly ray-click on semantic objects to issue intentions without language conversion. The system makes decisions based on object poses, proximity relationships, and spatial constraints in the 3D semantic field, rather than relying solely on 2D image features. Collision avoidance and operational space feasibility are considered during action sequence generation, transforming the operator from a low-level manipulator to a high-level decision-maker.
[0087] S6 controls the robot to execute a sequence of multi-step operation instructions sequentially.
[0088] After receiving a multi-step operation command sequence, the edge computing device on the robot executes it sequentially according to the subtask order. This command sequence is an ordered action queue, where each element contains a behavior-level subtask and its instantiated control parameters. The motion controller in the edge computing device retrieves the currently pending subtask from the queue and controls it according to the control parameters. After the previous subtask is completed and confirmed to be in place, the execution of the next subtask is automatically triggered. The system records the timestamp and execution status at the start and end times of each subtask, forming a complete operation log.
[0089] In this embodiment, during the process of controlling the robot to execute a sequence of multi-step operation instructions, the execution progress is transmitted back to the VR terminal in real time for visualization. Simultaneously, the robot's operating status is monitored in real time through multi-sensor fusion. When an anomaly is detected, different responses are triggered based on the anomaly level. In this embodiment, the multi-sensor includes encoders, torque sensors, and vision sensors. Encoders are installed at each joint of the robot to provide high-frequency feedback on the real-time angular position and angular velocity of each joint. Torque sensors are installed at each joint of the robotic arm or at the end effector to detect the contact forces and torques experienced by the robotic arm during movement and operation. The vision sensor includes a multi-view camera mounted on the robot, which continuously acquires images of the area where the end effector and the target semantic object are located during execution. Through visual servoing or feature matching technology, the relative pose deviation between the end effector and the target semantic object is calculated in real time. This deviation is used, on the one hand, for closed-loop correction of the end effector's motion trajectory to improve operational accuracy; on the other hand, it is used to monitor whether the execution is proceeding as expected. When the deviation between the actual position of the end effector and the target position observed by vision exceeds a preset threshold and cannot be corrected by servoing, it is determined to be a visual positioning anomaly.
[0090] When an anomaly is detected (such as collision, excessive force, or path deviation), different responses are triggered according to the anomaly level: minor anomaly triggers an alarm and continues execution; moderate anomaly pauses and awaits operator decision; severe anomaly triggers emergency braking and returns to a safe posture.
[0091] Simultaneously, throughout the entire process, the execution progress is transmitted back to the VR terminal in real time, presented to the operator in an intuitive and visual format. The VR terminal's interactive interface features a task execution dashboard, displaying the current overall task completion percentage as a progress bar and showcasing the sequence of subtasks in a list format. Currently executing subtasks are highlighted or flashing, completed subtasks are marked with a checkmark, and pending subtasks are displayed in gray. The operator can freely move their observation perspective within the virtual scene while clearly understanding the real-time progress of the task execution from this dashboard, without interrupting the immersive viewing experience.
[0092] S7 collects environmental change information in real time during robot execution and updates the semantic Gaussian field based on the environmental change information, so that the virtual scene rendered by the VR terminal keeps synchronized with the real operating environment in real time.
[0093] like Figure 4 As shown, the virtual-real consistency closed-loop update process described in step S7 includes: S71, after determining that the robot has successfully performed the action and the environment has changed, obtain the predicted image of the semantic Gaussian field rendered under the current virtual camera pose, and compare it with the actual image collected by the camera on the robot to determine the range of the changed three-dimensional region, i.e. the changed region.
[0094] This embodiment determines the change region by comparing pixels one by one. Specifically, it calculates the structural similarity index map or pixel-level photometric error map of the two images, and marks pixels with errors exceeding a preset threshold as difference pixels. Since the two images have achieved pixel-level spatial correspondence through camera intrinsic and extrinsic parameters, the difference pixels can be directly back-projected into the three-dimensional space of the semantic Gaussian field. Through the projection relationship between the three-dimensional Gaussian points and image pixels, the corresponding set of three-dimensional Gaussian points is determined, and the range of the three-dimensional region covered by this set is the change region.
[0095] S72 performs local incremental reconstruction on the changed region, updates the three-dimensional Gaussian point parameters within the changed region, and freezes the three-dimensional Gaussian point parameters outside the changed region.
[0096] In this embodiment, for new regions in the changing region that have not been previously observed, semantic labels are assigned according to a hierarchical semantic determination strategy: First, the semantic labels of neighboring regions with existing semantic labels are queried. If the color and depth features of the new region and the neighboring regions meet the preset similarity conditions, the semantic label of the neighboring region is used. If the preset similarity conditions are not met, the distribution of semantic labels of the surrounding neighboring regions is statistically analyzed, and the semantic label of the new region is determined according to the majority principle. If the degree of mixing of semantic labels of the surrounding neighboring regions exceeds the preset threshold or cannot be determined according to the majority principle, the visual language model is triggered to re-identify the new region and generate a semantic label.
[0097] Specifically, the strategy involves statistically analyzing the Gaussian points of a new region against their neighboring Gaussian points in terms of color features (distance in the RGB color space) and depth features (variance of depth values). If the color space distance is less than a color threshold and the depth variance is less than a depth threshold, the new region and its neighboring regions are considered to belong to the same semantic object, and the semantic label of that neighboring Gaussian point is directly adopted. This strategy is based on the prior assumption that the surface of the same object usually has continuous color and depth features, and is suitable for the common situation where the newly exposed region and the original object belong to the same material surface. If the Gaussian points of the new region do not meet the above conditions, the search radius is expanded, and the label distribution of existing semantically labeled Gaussian points within the spatial range surrounding the new region is statistically analyzed. The Gaussian point of the new region is assigned the semantic label with the highest frequency according to the majority principle. This strategy is suitable for situations where the newly exposed region and most surrounding objects belong to the same semantic category, such as a desktop area exposed after moving an object belonging to the same semantic category as the surrounding desktop areas. If the semantic label distribution of the surrounding neighboring regions is more mixed than a preset threshold (e.g., each label accounts for no more than 50%, with no obvious majority), or if it still cannot be determined after statistical voting, it indicates that the semantics of the new region are difficult to reliably infer from neighboring information, triggering the visual language model to re-identify the image region of the new region. The visual language model takes a cropped image of the new region from multiple perspectives as input, outputs the semantic label of the region, and assigns the label to the corresponding new Gaussian point.
[0098] S73, after the update is completed, compare the predicted pose of the semantic object in the virtual scene with the measured pose of the corresponding object in the environment. If the deviation between the two exceeds the preset threshold, the area where the deviation occurs will be locally corrected. If the local correction cannot meet the accuracy requirements, global realignment will be triggered.
[0099] In this embodiment, the deviation metric includes the Euclidean distance of the position deviation and the angular difference of the orientation deviation. When the deviation does not exceed a preset threshold (e.g., the position deviation is less than 2 cm and the orientation deviation is less than 5 degrees), the updated semantic Gaussian field geometric accuracy is determined to meet the requirements, and the update is completed.
[0100] When the deviation exceeds a preset threshold, priority is given to correcting the local area where the deviation occurred. The local correction method is as follows: using the measured pose as a constraint, and taking the pose deviation as one of the optimization objectives, the position parameters of the Gaussian points within the deviation area are fine-tuned to make the predicted pose of the Gaussian points after rendering approximate the measured pose, while keeping the appearance parameters of the Gaussian points unchanged. This local correction strategy can effectively eliminate the deviation in most cases, and the computational cost is far lower than that of global reconstruction.
[0101] When local correction fails to converge the deviation to within a threshold after multiple iterations, or when the deviation region involves a large area of inconsistent poses of multiple related objects, the local correction is deemed insufficient to meet the accuracy requirements, triggering global realignment. Global realignment performs a global bundle adjustment optimization on all Gaussian point positions in the entire semantic Gaussian field. The optimization objective is constructed using the visual feature constraints of all historical keyframes and the current LiDAR point cloud constraints, solving for a globally consistent scene structure and camera pose. After global realignment, the updated semantic Gaussian field is synchronously transmitted to the VR terminal, restoring precise synchronization between the virtual scene observed by the operator and the real environment.
[0102] Differences from existing technologies: Existing solutions use 3DGS for offline training data augmentation (re-collection and retraining are required after environmental changes) or static grasping pose generation (the semantic field remains unchanged); SLAM solutions, although updated online, are triggered by camera motion and do not involve semantic label preservation. This step is triggered by the robot's actions to update, forming a complete "execution-change-update-verification" closed loop through local incremental reconstruction, semantic consistency propagation, and pose deviation closed-loop verification, explicitly handling the inheritance and propagation of semantic labels. Differences from dynamic scene modeling solutions such as 4DGS: 4DGS models the inherent dynamism of the scene (such as a moving person), while this step deals with environmental changes caused by the robot's actions. The dynamism is introduced by external actions, and the update effect needs to be confirmed through closed-loop verification.
[0103] Example 2
[0104] This embodiment provides an immersive robot teleoperation system, including: The three-dimensional Gaussian field reconstruction module is used to collect perception data of the operating environment based on the multi-source perception unit on the robot, and reconstruct the initial three-dimensional Gaussian field based on the perception data. The semantic mapping module is used to extract the semantic embedding of multi-view environmental images in the perceptual data and map it to the three-dimensional Gaussian points corresponding to the initial three-dimensional Gaussian field, and assign semantic labels to each three-dimensional Gaussian point to obtain the semantic Gaussian field. The VR interaction module is used to transmit the semantic Gaussian field to the VR terminal to support the operator to roam in the virtual scene rendered by the VR terminal with six degrees of freedom and observe the scene from any position, and to collect the target semantic object selected by the operator based on VR ray interaction. The intent acquisition and verification module is used to directly acquire or infer the operator's interaction intent for the target semantic object based on the pose, proximity relationship and spatial constraints of the target semantic object in the semantic Gaussian field, and to verify the interaction intent by querying a predefined semantic-action knowledge base; The instruction mapping module is used to map verified interactive intents into a multi-step operation instruction sequence; The control module is used to control the robot to execute the multi-step operation instruction sequence in sequence; The scene update module is used to collect environmental change information in real time during robot execution and update the semantic Gaussian field based on the environmental change information, so that the virtual scene rendered by the VR terminal keeps synchronized with the real operating environment in real time.
[0105] like Figure 5 As shown, the system adopts a three-layer edge-cloud collaborative architecture, specifically including the robot end (edge end), cloud end and VR end. The three layers achieve low-latency data transmission through 5G / WiFi / wired network.
[0106] The robot is equipped with a multi-view camera array and a LiDAR (Light Detection and Ranging) system to collect environmental data, enabling it to gather perception data of the operating environment. The multi-view camera array is mounted on the robot body in a surround or forward-facing configuration, providing multi-view RGB / RGB-D images covering the operating area. The LiDAR acquires dense point clouds of the environment, complementing the multi-view image data and providing geometric prior constraints for 3D Gaussian field reconstruction. The robot's edge computing device receives multi-step operation command sequences from the cloud or VR terminal, converts them into robot joint control commands, and executes them in real time. Simultaneously, it collects the robot's operating status (joint angles, speed, current, end effector force, etc.) and feeds it back to the cloud and VR terminal.
[0107] The cloud platform utilizes high-performance GPU servers to perform full 3DGS training and complex VLM inference. Full 3DGS training includes global optimization of parameters such as the location, covariance, spherical harmonic coefficients, and opacity of Gaussian points. After environmental changes occur, the local incremental reconstruction results of the changed regions are further refined in the cloud to ensure global consistency. Complex VLM inference achieves complex inference for the following tasks by running a pre-trained Visual Language Model (VLM): Intent inference: Receive structured context information associated with the target semantic object selected by the operator through VR ray interaction, and generate candidate interaction intents; Semantic label re-inference: During the local incremental reconstruction process, for new observation areas where neither color nor depth features meet the similarity condition and the surrounding label distribution is mixed, VLM is triggered to re-identify the semantics of the area and output new semantic labels. Task planning: Using VLA (Visual Language Action Model) and other methods, the verified interaction intent is decomposed into a sequence of behavior-level sub-tasks, and kinematic feasibility checks, collision detection and reachability checks are performed to generate the final multi-step operation instruction sequence.
[0108] When entering a new environment for the first time, the edge device collects data and uploads it to the cloud. The cloud then trains the model and distributes it. During daily operation, the edge device performs local sparse inference and only uploads keyframes / abnormal data.
[0109] VR devices are immersive virtual reality equipment worn by the operator (such as VR headsets and matching controllers), and mainly include the following functions: 1. VR Headset Rendering: Semantic Gaussian field data is received from the cloud or robot edge devices via a low-latency network and rendered in real-time within the VR headset to generate a virtual scene. Furthermore, the virtual scene can be switched between a pure 3D Gaussian field rendering mode and a dynamic scene hybrid rendering mode based on the robot's execution status, operator-initiated switching commands, or the state of the target semantic object. In hybrid rendering mode, the 3D Gaussian sputtering image and the robot's real-time camera feed are dynamically alpha-fused and pixel-level spatial correspondence is achieved through camera intrinsic and extrinsic parameter alignment.
[0110] 2. Six degrees of freedom roaming: Operators can roam freely in the virtual scene with six degrees of freedom, and observe the global and local details of the environment to be operated from any position and angle.
[0111] 3. Multimodal interaction: This includes: VR ray interaction: emitting a virtual ray through the controller to select a target semantic object in the virtual scene; voice input: collecting the operator's voice commands through the microphone integrated into the VR headset; and control interaction: collecting the operator's operation commands for VR interactive interface controls such as the hybrid rendering ratio adjustment slider.
[0112] In one embodiment, the robot is a mobile robot or quadruped robot dog platform with multiple camera interfaces and a LiDAR interface, equipped with multiple RGB cameras, LiDAR and inertial measurement unit (IMU); the edge computing device is an NVIDIA Jetson Orin series or similar edge AI device; the VR headset is a Meta Quest series, Pico series or similar six-degrees-of-freedom VR device; the cloud server is equipped with a high-performance GPU.
[0113] In this embodiment, each module has mature open-source implementations or academic verification as its foundation: 3DGS real-time reconstruction has been verified by the LEGS framework and can achieve real-time inference on Jetson Orin; 3DGS semantic segmentation has been verified by academic works such as OpenSplat3D, ReferSplat, and Click-Gaussian; VLM semantic understanding has mature models such as CLIP, LLaVA, and Qwen-VL; VR real-time rendering has been verified by the Unity GaussianSplatting plugin and can run stably on Quest 3; and robot control has been supported by the Unitree SDK and general robot interface.
[0114] This embodiment innovatively combines and couples the aforementioned mature modules, especially deeply integrating VR immersive arbitrary viewpoint interaction with semantic object-driven instruction generation to form a complete technical solution with clear feasibility.
[0115] Example 3
[0116] This embodiment provides a specific implementation process of the method described in Embodiment 1.
[0117] Scenario description: The operator is located in the control center and needs to remotely control the robot to enter the industrial workshop to inspect and close specific valves.
[0118] Implementation steps: 1) The robot walks around the workshop, and multiple cameras and LiDAR collect environmental data, and the edge device builds a 3DGS basic scene in real time; 2) The cloud-based VLM performs semantic segmentation on the scene, identifying semantic objects such as "valve A", "valve B", "pipeline", "ground", and "control cabinet", forming a semantic Gaussian field; 3) The operator wears a VR headset and roams freely in the scanned space (the VR view position is independent of the robot's physical position). He looks down at the whole scene from above, teleports to the side of the valve to observe the details up close, and finds that valve A is in an abnormal open state. 4) When the operator ray-clicks on valve A, the VR interface displays the label "Valve A - Status: Open - Suggested Operation: Closed". The operator then clicks "Close", and the system automatically generates a sequence of instructions: navigate to the front of valve A, position the robotic arm on the valve handle, clamp, rotate, release, retreat, and take a photo for confirmation. 5) During the robot's closing action, VR automatically switches to hybrid rendering mode: the static workshop background reconstructed by 3DGS is superimposed with the valve area image transmitted back in real time by the robot's camera (the real-time image accounts for a larger proportion than the 3DGS background, mainly to see the details of the action, while retaining the geometric reference of the surrounding space), and the operator can clearly observe the real-time interaction process between the robotic arm and the valve. 6) After execution, the system automatically switches back to pure 3DGS mode, and the partial update displays the status of valve A as "closed", which is consistent with the actual status.
[0119] Traditional solutions require operators to manually control the robot's movement, adjust the robotic arm's posture, and manually write rotation commands, and can only observe from the robot's first-person perspective. In this embodiment, the robot first completes an environmental scan, allowing the operator to freely observe the entire environment from any perspective (top view, side view) via VR, and complete operations with a single click on semantic objects. Dynamic blending rendering during execution ensures that the operator can clearly see real-time actions while maintaining spatial reference. The operation steps are compressed from multiple steps to a minimalist interaction, significantly reducing the error rate and greatly improving spatial awareness.
[0120] The present invention also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 An immersive robot teleoperation method is provided.
[0121] The present invention also provides Figure 6 One of the corresponding Figure 1 A schematic diagram of the structure of an electronic device. (e.g.) Figure 6 At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for the business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 1 The method described herein. Of course, in addition to software implementation, this invention does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0122] Improvements in a technology can be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology can now be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement in methodology cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0123] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0124] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0125] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, in implementing this invention, the functions of each unit can be implemented in one or more software and / or hardware components.
[0126] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0127] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0128] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0129] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0130] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0131] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0132] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0133] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0134] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0135] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0136] The various embodiments in this invention are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0137] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for immersive robot teleoperation, characterized in that, Includes the following steps: The robot collects perception data of the operating environment based on its multi-source perception unit, and reconstructs the initial three-dimensional Gaussian field based on the perception data. Semantic embeddings of multi-view environmental images are extracted from the perceptual data and mapped to the three-dimensional Gaussian points corresponding to the initial three-dimensional Gaussian field. Semantic labels are assigned to each three-dimensional Gaussian point to obtain the semantic Gaussian field. The semantic Gaussian field is transmitted to the VR terminal to support the operator to roam freely in the virtual scene rendered by the VR terminal and observe the scene from any position, and to collect the target semantic object selected by the operator based on VR ray interaction. The operator's interaction intent towards the target semantic object is obtained directly or inferred based on the pose, proximity relationship and spatial constraints of the target semantic object in the semantic Gaussian field, and the interaction intent is verified by querying a predefined semantic-action knowledge base; Map the validated interaction intent into a multi-step operation instruction sequence; The robot is controlled to execute the multi-step operation instruction sequence in sequence; During robot execution, environmental change information is collected in real time, and the semantic Gaussian field is updated based on the environmental change information, so that the virtual scene rendered by the VR terminal keeps synchronized with the real operating environment in real time.
2. The immersive robot teleoperation method according to claim 1, characterized in that, The method for directly obtaining the operator's interaction intent regarding the target semantic object is as follows: The system acquires the operator's voice input information, performs voice recognition on the voice input information to obtain text information, or acquires the text information manually input by the operator through the interactive interface. Extract the interaction intent of the corresponding target semantic object from the text information.
3. The immersive robot teleoperation method according to claim 1, characterized in that, The method for inferring the operator's interaction intent regarding the target semantic object based on the pose, proximity relationships, and spatial constraints of the target semantic object in a semantic Gaussian field is as follows: In response to the VR ray interaction hitting the target semantic object, multi-dimensional context information associated with the target semantic object is obtained. The multi-dimensional context information includes at least object identifier, semantic category, pose, size, current state, neighboring object information, robot current state, and operation history. The multi-dimensional contextual information is organized into structured data and input into a pre-trained semantic reasoning model, which then generates at least one candidate interaction intent based on the structured data. The at least one candidate interaction intent is fed back to the operator, and the operator's selection or regeneration instruction is received through VR ray interaction. In response to the selection instruction, the selected candidate interaction intent is determined as the interaction intent of the target semantic object; In response to the regeneration instruction, the semantic reasoning model is triggered to regenerate at least one new candidate interaction intent and feed it back to the operator.
4. The immersive robot teleoperation method according to claim 1, characterized in that, The virtual scene rendered by the VR terminal is constructed based on a pure 3D Gaussian field rendering mode and a hybrid rendering mode. Initially, the pure 3D Gaussian field rendering mode is used. The hybrid rendering mode is triggered in response to the robot being in an execution state, the operator actively switching to the hybrid rendering mode, or the selected target semantic object being marked as being in a changing state. The current pose of the camera mounted on the robot and the real-time images it captures are obtained, and the virtual camera pose that is closest to the current pose of the camera is determined in the semantic Gaussian field. A three-dimensional Gaussian sputtering image is generated based on the virtual camera pose. The three-dimensional Gaussian sputtering image is dynamically mixed with the real-time image captured by the camera to generate a mixed rendering image. In the static observation mode, the proportion of the three-dimensional Gaussian sputtering image is higher than that of the real-time image. In the dynamic operation mode, the proportion of the real-time image is higher than that of the three-dimensional Gaussian sputtering image. The dynamic proportion is manually adjusted by the operator through the adjustment control of the VR terminal. During the generation of the hybrid rendering image, pixel-level spatial correspondence between images is achieved by aligning camera intrinsic and extrinsic parameters, and Alpha fusion is used at the hybrid boundary to eliminate stitching gaps. When the conditions for triggering the hybrid rendering mode are eliminated, switch to the pure 3D Gaussian field rendering mode to restore the six-degree-of-freedom roaming; wherein, the conditions for triggering the hybrid rendering mode are eliminated as follows: the robot execution state ends and the target semantic object state is stable and without abnormality, the operator actively exits the dynamic mode, or the marker of the selected target semantic object is restored to a non-changing state; if an abnormal interruption occurs during robot execution, the hybrid rendering mode or the switch to the pure 3D Gaussian field rendering mode is determined according to the degree of environmental change.
5. The immersive robot teleoperation method according to claim 1, characterized in that, The method for mapping validated interactive intents to a multi-step operation instruction sequence is as follows: Transform the verified interaction intents into structured task objectives; The structured task objective is input into the task planning model, which combines the robot's current state with visual observations obtained from the semantic Gaussian field to decompose the structured task objective into a sequence of behavioral subtasks. Motion planning verification is performed on the behavior-level subtask sequence, including instantiating each subtask into an action instruction containing control parameters, and sequentially performing kinematic feasibility checks, collision detection, and reachability checks. If any of the checks fails, an alternative subtask is generated until a multi-step operation instruction sequence that passes the checks is obtained.
6. The immersive robot teleoperation method according to claim 1, characterized in that, During the process of controlling the robot to execute the multi-step operation instruction sequence in sequence, the execution progress is transmitted back to the VR terminal in real time for visualization. At the same time, the robot's operating status is monitored in real time through multiple sensors. When an anomaly is detected, different responses are triggered according to the anomaly level. The multiple sensors include an encoder, a torque sensor, and a vision sensor. The encoder provides feedback on the joint position, the torque sensor detects the contact force, and the vision sensor provides feedback on the relative pose of the end effector and the target semantic object.
7. The immersive robot teleoperation method according to claim 1, characterized in that, The method for collecting environmental change information in real time and updating the semantic Gaussian field based on the environmental change information is as follows: After determining that the robot has successfully performed the action and the environment has changed, the predicted image rendered by the semantic Gaussian field under the current virtual camera pose is obtained and compared with the actual image captured by the camera on the robot to determine the range of the changed three-dimensional region, i.e. the changed region. Local incremental reconstruction is performed on the changed region, updating the 3D Gaussian point parameters within the changed region and freezing the 3D Gaussian point parameters outside the changed region. For new regions in the changed region that were not previously observed, semantic labels are assigned according to a hierarchical semantic determination strategy: first, the semantic labels of neighboring regions with existing semantic labels are queried; if the color and depth features of the new region and the neighboring regions meet a preset similarity condition, the semantic label of the neighboring region is used; if the preset similarity condition is not met, the distribution of semantic labels in the surrounding neighboring regions is statistically analyzed, and the semantic label of the new region is determined according to the majority principle; if the degree of semantic label mixing in the surrounding neighboring regions exceeds a preset threshold or the majority principle still cannot determine the label, the visual language model is triggered to re-identify the new region and generate a semantic label. After the update is completed, the predicted pose of the semantic object in the virtual scene is compared with the measured pose of the corresponding object in the environment. If the deviation between the two exceeds the preset threshold, the area where the deviation occurs is locally corrected. If the local correction cannot meet the accuracy requirements, global realignment is triggered.
8. An immersive robot teleoperation system, characterized in that, include: The three-dimensional Gaussian field reconstruction module is used to collect perception data of the operating environment based on the multi-source perception unit on the robot, and reconstruct the initial three-dimensional Gaussian field based on the perception data. The semantic mapping module is used to extract the semantic embedding of multi-view environmental images in the perceptual data and map it to the three-dimensional Gaussian points corresponding to the initial three-dimensional Gaussian field, and assign semantic labels to each three-dimensional Gaussian point to obtain the semantic Gaussian field. The VR interaction module is used to transmit the semantic Gaussian field to the VR terminal to support the operator to roam in the virtual scene rendered by the VR terminal with six degrees of freedom and observe the scene from any position, and to collect the target semantic object selected by the operator based on VR ray interaction. The intent acquisition and verification module is used to directly acquire or infer the operator's interaction intent for the target semantic object based on the pose, proximity relationship and spatial constraints of the target semantic object in the semantic Gaussian field, and to verify the interaction intent by querying a predefined semantic-action knowledge base; The instruction mapping module is used to map verified interactive intents into a multi-step operation instruction sequence; The control module is used to control the robot to execute the multi-step operation instruction sequence in sequence; The scene update module is used to collect environmental change information in real time during robot execution and update the semantic Gaussian field based on the environmental change information, so that the virtual scene rendered by the VR terminal keeps synchronized with the real operating environment in real time.
9. An electronic device, characterized in that, The device includes a memory and one or more processors, wherein the memory stores executable code, and the one or more processors execute the executable code to implement the immersive robot teleoperation method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores a program that, when executed by a processor, implements the immersive robot teleoperation method according to any one of claims 1-7.
Citation Information
Patent Citations
Immersive robot teleoperation method and system
CN120578299B