Remote telepresence sense control method, device and system for heterogeneous device cooperation

By using incremental quaternions, gravity compensation algorithms, and dynamic geofencing technology, the coordinate misalignment and high latency issues of heterogeneous devices are solved, enabling intuitive control and efficient interaction across devices. It also supports real-time interaction of virtual avatars in real space, improving the user experience and system response speed of remote control.

CN121887843APending Publication Date: 2026-04-17吴金河
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
吴金河
Filing Date
2026-01-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing remote control technologies face challenges in open social scenarios, such as coordinate misalignment of heterogeneous devices, single interaction medium, and significant perception of high dynamic latency, leading to user dizziness, difficulty in operation, and unsmooth interaction.

Method used

A heterogeneous attitude mapping algorithm based on incremental quaternions and gravity compensation is adopted, combined with dynamic geofencing and Kalman filter prediction technology, to achieve unified attitude mapping and low-latency interaction across devices, supporting real-time interaction of virtual avatars in real geographic space.

Benefits of technology

It solves the coordinate misalignment problem of heterogeneous devices, enables intuitive user operation, expands the flexibility of the interaction medium, reduces network latency perception, and provides a high-precision remote proxy interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887843A_ABST
    Figure CN121887843A_ABST
Patent Text Reader

Abstract

The invention discloses a remote presence sense control method, device and system for heterogeneous device cooperation. The method comprises the following steps: establishing a communication link between a first terminal and a controlled subject; acquiring a first terminal pose in real time, calculating a rotation increment quaternion, and mapping the rotation increment quaternion to a controlled main body coordinate system in combination with a gravity compensation matrix to generate a standardized instruction; and driving the controlled main body to execute the synchronous action. Through a heterogeneous attitude mapping algorithm, the problems of coordinate imbalance and control dizziness caused by different hardware structures of the control end and the controlled end are solved, and accurate control of a physical robot or a virtual image is supported. In addition, the invention further provides a social node matching mechanism based on the dynamic geofence, and temporary borrowing control over non-specific equipment is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision, human-computer interaction, augmented reality (AR) and artificial intelligence, and specifically to a remote presence control method, device, storage medium and system that supports collaboration of heterogeneous devices and the appearance of virtual avatars. Background Technology

[0002] With the widespread adoption of 5G communication technology and virtual reality / AR hardware, telepresence has become a crucial bridge connecting the physical and digital spaces. Existing remote control technologies are primarily used in industrial operations, telemedicine, and drone piloting. However, these technologies face significant bottlenecks when applied to open social and consumer applications:

[0003] 1. Coordinate Misalignment and Motion Sickness in Heterogeneous Devices: Existing control systems typically assume that the control and controlled devices are homogeneous (e.g., using a controller to control a drone). However, in open social scenarios, the control device might be a VR headset, while the controlled device could be a smartphone gimbal, a robotic dog, or a virtual avatar. Because different devices have different sensor gravity references and coordinate system definitions, direct mapping can lead to severe "non-linear motion deviations," causing motion sickness in users. Even when users control their own heterogeneous devices (e.g., using a VR controller to control a robotic dog at home), the lack of a unified coordinate mapping standard often requires users to spend a long time learning and adapting, making intuitive control difficult.

[0004] 2. Limited Interaction Platform: Traditional technologies often rely on pre-set, fixed robots or controlled nodes, lacking flexibility. In situations requiring instantaneous social interaction or assistance, they cannot quickly match the most suitable "on-site agent" from a massive pool of non-specific users. Furthermore, existing remote interaction methods primarily focus on controlling physical entities, lacking a universal architecture capable of using "virtual avatars" as controlled subjects and enabling them to interact with physical entities in multi-dimensional, real-time interactions at real geographical coordinates.

[0005] 3. Significant High Dynamic Latency Awareness: Remote control demands extremely high real-time command performance. In environments with fluctuating network speeds, command transmission delays can cause lag in the actions of the controlled device. Existing solutions lack the ability to predict the motion state of the controlled device in advance and provide visual feedback compensation, making it difficult for remote users to obtain a smooth, continuous interactive experience with a "first-person perspective." Summary of the Invention

[0006] 3.1 Purpose of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of existing remote immersive control technologies, such as coordinate misalignment of heterogeneous devices, single interaction medium, and significant latency perception in high dynamic environments, and to provide a remote immersive control method, device, and system for heterogeneous device collaboration. This invention enables unified posture mapping across physical and virtual entities, supporting users to perform high-precision, low-latency remote proxy interactions in real geographic space using virtual avatars or physical robots as media.

[0008] Implementation logic of devices, media and systems:

[0009] This invention also discloses a remote presence control device for heterogeneous device collaboration. The device includes a memory and a processor, the memory storing computer instructions for executing the aforementioned remote control method. Furthermore, this invention provides a computer-readable storage medium on which a set of instructions stored, when executed, can drive heterogeneous devices to extract and map pose increments.

[0010] Furthermore, the present invention discloses a remote immersive sensing control system for heterogeneous device collaboration, characterized by comprising a first terminal (control end), an audience terminal (execution end), and a cloud server. The server acts as a central hub, responsible for handling coordinate system compensation calculations between the first terminal and the audience terminal. When the controlled entity is a virtual carrier, the server further undertakes the real-time rendering and distribution of the virtual entity.

[0011] 3.2 Technical Solution

[0012] To achieve the above objectives, this invention provides a remote presence control method for heterogeneous device collaboration, the core technical solution of which includes the following three key dimensions of processing logic:

[0013] 1. Heterogeneous attitude mapping based on incremental quaternions and gravity compensation

[0014] To address the coordinate system misalignment problem caused by differences in hardware construction between the control end (such as a VR headset) and the controlled end (such as a mobile phone or virtual avatar), this invention proposes a standardized pose mapping algorithm:

[0015] Pose increment extraction: Real-time acquisition of pose data from the first terminal to calculate the current time step. Relative to the initial time Rotational Increment Quaternion This calculation process incorporates the gravity vector compensation matrix obtained from the gravity sensor. ,

[0016] The calculation formula is This eliminates the inherent differences in holding angle and physical horizontal plane between different devices.

[0017] Standardized instruction conversion: converting the calculated standardized pose increments... Map the coordinates to the local coordinate system of the controlled object. If the controlled object is a physical gimbal, convert the quaternion into Euler angle commands (Pitch, Yaw, Roll); if the controlled object is a virtual avatar, directly drive the rotation of its skeletal nodes to ensure that the user's observation intention and the motion response of the controlled object remain linearly consistent.

[0018] 2. Real-time matching of non-specific nodes based on dynamic geofencing

[0019] In one embodiment, addressing the limitations of existing technologies that rely on fixed-binding devices, this invention proposes a node matching mechanism based on spatiotemporal dynamic filtering. Specifically, it includes:

[0020] Dynamic fence construction: The server receives the access request from the first terminal (control terminal) for the target geographic coordinates (Lat, Lon), and constructs a dynamic geofence with the coordinates as the center and a preset distance (R) as the radius.

[0021] Heterogeneous node filtering: Within the fenced area, the system searches for active audience terminals (including physical robots, mobile gimbals, or handheld AR devices) in real time. By comparing the social tags (such as interest graphs and friend relationships) and permission levels of the first terminal and the audience terminals, the system filters out the optimal matching nodes from the non-specific population.

[0022] Link establishment: A low-latency, two-way real-time communication link is established between the first terminal and the selected audience terminals for transmitting control command streams and environmental perception feedback streams.

[0023] 3. Geographically Anchored Virtual Arrival and Closed Loop of Virtual-Real Perception

[0024] It allows users to intervene in real spaces without a physical presence, forming a two-way interactive closed loop:

[0025] Geographic coordinate anchoring: When the controlled subject is defined as a virtual avatar, the cloud engine instantiates the virtual model at the target geographic coordinates and aligns the origin of the virtual avatar's coordinates with the real physical environment (such as streets and squares) at the centimeter level using the Visual Positioning System (VPS) and GPS data.

[0026] Multi-terminal real-time distribution: The system encapsulates the real-time pose data of the virtual avatar into an augmented reality (AR) data packet, and pushes it to the AR terminals or mobile phone screens of all real users in the vicinity of the location in real time through dynamic geofencing distribution logic.

[0027] Perception feedback loop: The environmental images and interactive behaviors (such as waving and voice) captured by the AR terminal of the real user on site are transmitted back to the virtual reality interface of the first terminal in real time, thereby realizing the deep perception interaction between the remote virtual entity and the on-site physical entity in a unified spatiotemporal dimension.

[0028] 3.3 Beneficial Effects

[0029] Compared with existing technologies, the remote presence control method, device, and system based on social matching provided by this invention have the following significant advantages:

[0030] 1. Completely solves the coordinate misalignment and dizziness problems caused by heterogeneous device control. Existing remote control technologies usually require that the control end and the controlled end have isomorphic hardware; otherwise, it will lead to serious motion mapping deviations. This invention innovatively introduces a relative increment calculation and gravity vector compensation mechanism based on quaternions through a heterogeneous attitude mapping algorithm.

[0031] Technological advancements: Regardless of whether the user of the control end (such as a VR headset) is standing, lying flat, or on their side, and regardless of the initial physical orientation of the controlled end (such as a mobile phone gimbal or drone), the system can forcibly align the gravity planes of both and linearly map the user's "observation intention" into the "movement command" of the controlled end.

[0032] Practical benefits: It greatly reduces the cognitive load on users when remotely controlling devices across different devices, eliminates motion sickness caused by non-linear motion of the viewpoint, and enables ordinary users to accurately control professional industrial robots or remote cameras through consumer-grade VR devices.

[0033] 2. Breaking through the limitations of physical carriers, achieving "asymmetric virtual descent": Traditional immersive technologies rely on expensive physical robots as substitutes, limiting their application scope. The virtual descent interaction technology proposed in this invention supports the precise anchoring of the user's virtual avatar to real geographical coordinates through cloud rendering, even when no physical devices are available.

[0034] Technological advancements: Through dynamic geofencing and AR distribution technology, virtual avatars are projected in real time onto the mobile phones or AR glasses of real users on-site, constructing a perception loop between "physical entities" and "virtual entities" in a unified spatiotemporal dimension.

[0035] Practical benefits: It greatly expands the scenarios for remote social interaction. Users can "instantly" appear anywhere in the world (as long as there are active audience terminals there) without purchasing hardware, and can engage in in-depth interactions with people on-site, such as waving and talking, creating a new paradigm of virtual-real integrated social interaction.

[0036] 3. This invention achieves "on-demand" control based on the sharing economy model. Existing technologies typically involve fixed point-to-point equipment connections, resulting in poor flexibility. This invention introduces a dynamic geofencing and node matching mechanism.

[0037] Technological advancements: The system is able to filter the best access point in real time from a massive number of non-specific user devices (audience terminals) based on geolocation and social tags.

[0038] Practical benefits: Introducing the "sharing economy" concept into the field of visual control. In emergency situations (such as witnessing sudden events or providing remote assistance), users can borrow the perspective of a stranger's device without the need for pre-deployed hardware, significantly improving the system's response speed and coverage.

[0039] 4. Significantly reduces control latency perception in highly dynamic network environments. To address the pain point of extremely high real-time requirements for remote control, this invention integrates Kalman filter prediction and AR visual compensation technology.

[0040] Technological advancements: The system no longer passively waits for the screen to be transmitted back, but instead predicts the next frame action of the controlled subject based on the current instructions and overlays a HUD guidance indicator on the control end.

[0041] Practical benefits: It compensates for the perceived latency caused by physical network transmission from a "visual psychology" perspective. Even in environments with fluctuating network speeds, users can be assured that commands have been executed through HUD indicators, thus maintaining operational consistency and confidence, and keeping perceived latency within a comfortable millisecond range.

[0042] 5. A data closed loop from "real-time interaction" to "historical reconstruction" has been constructed. The data generated by this system has extremely high reuse value. The system persistently stores the command stream, motion trajectory, and environmental feedback during remote interaction by associating them with XYZT spatiotemporal tags.

[0043] Technological advancement: Through collaboration with related patents ("the applicant's patent application filed on the same day entitled 'A Data Processing Method, Apparatus, Medium and System Based on Spatiotemporal Index'"), the real-time interactive data generated by this invention can be directly converted into historical slices.

[0044] Practical benefits: Each remote interaction becomes more than just a momentary communication; it becomes an "asset accumulation" for the digital twin world. This data can be used to reconstruct historical events and train AI behavioral models, possessing profound commercial and scientific research value. Attached Figure Description

[0045] Figure 1 This is a schematic diagram of the overall architecture of the remote presence sensing control system provided in an embodiment of the present invention;

[0046] Figure 2This is a schematic diagram of the overall process of the heterogeneous device collaborative control method provided in an embodiment of the present invention;

[0047] Figure 3 This is a schematic diagram of the heterogeneous coordinate system pose mapping and quaternion calculation logic provided in an embodiment of the present invention;

[0048] Figure 4 The communication link adaptive control and delay prediction compensation logic diagram provided in the embodiments of the present invention;

[0049] Figure 5 A schematic diagram of the security and privacy circuit breaker mechanism and geofence boundary provided in this embodiment of the invention;

[0050] Figure 6 This is a logic diagram for multi-terminal collaborative perspective switching provided in an embodiment of the present invention;

[0051] Figure 7 This is an interactive architecture diagram of an asymmetric virtual presence carrier (virtual arrival) provided in an embodiment of the present invention. Detailed Implementation

[0052] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0053] 5.1 System Overall Architecture

[0054] Reference Appendix Figure 1 This system consists of a first terminal (control terminal), a cloud server, and an audience terminal (execution terminal).

[0055] The first terminal is usually a VR headset or a mobile device with a posture sensor, which is responsible for collecting the user's head / hand movement data.

[0056] Cloud server: The core processing unit, responsible for performing dynamic geofencing filtering, heterogeneous coordinate system transformation matrix calculation, and virtual avatar rendering.

[0057] Audience terminals include physical actuators (such as mobile phone gimbals and robots) or AR display devices that carry virtual avatars.

[0058] 5.1.1 Establishment of Communication Link

[0059] The communication link can be established directly based on the device ID (as in Example 1), or dynamically based on social matching (as in Example 2). The following focuses on describing the establishment process based on social matching:

[0060] Reference Appendix Figure 2 This system employs a dual screening mechanism based on both spatiotemporal and social dimensions. The specific steps are as follows:

[0061] 1. Spatiotemporal initial screening: The cloud server receives the access request initiated by the first terminal (including the target GPS coordinates and request type) and immediately builds a dynamic geofence centered on the coordinates (the radius R can be adaptively adjusted according to the density of active nodes, such as 50 meters to 500 meters).

[0062] 2. Social Tag Matching: The system retrieves all active audience terminals within the geofence area. It reads the user profile (interest graph, historical behavior, friend relationship chain) of the first terminal and compares it with the open tags of the audience terminals. For example, if the request type is "tourism," the system prioritizes matching audience terminals tagged "photography enthusiast" with sufficient battery power; if the request type is "emergency rescue," it prioritizes matching verified users with high credit ratings. It should be noted that the matching process is not simply a business-level screening; the system's underlying layer simultaneously performs communication link quality detection and spatial geometry calculation: the server uses ping testing and packet loss rate analysis to calculate the network topology stability between the first terminal and candidate audience terminals in real time; simultaneously, based on a dynamic geofence edge algorithm, terminal nodes located in signal multipath effect blind zones or experiencing excessive Doppler shift due to high-speed movement are eliminated, thus ensuring that the constructed control link possesses industrial-grade real-time performance and reliability.

[0063] 3. Non-specific node handshake: Once the optimal matching node is selected, the server sends a "proxy request" to the audience terminal. After the audience user confirms authorization (or based on a preset automatic order acceptance protocol), the system uses P2P penetration technology to establish a low-latency, two-way encrypted communication link between heterogeneous devices, completing the instant connection from "stranger" to "on-site substitute".

[0064] 5.2 Heterogeneous Pose Mapping Algorithm

[0065] Reference Appendix Figure 3 To address the issue of coordinate inconsistency between heterogeneous devices, this invention employs a "relative incremental mapping" strategy.

[0066] Assuming the control end (VR headset) is in The original quaternion of time is Initial calibration time The quaternion is .

[0067] The system first constructs a gravity compensation matrix using gravity sensor data. This matrix is ​​used to calibrate the non-horizontal holding state of the control end to the physical ground plane of the controlled end.

[0068] Rotational Increment Quaternion The calculation formula is as follows:

[0069]

[0070] in, To represent quaternion multiplication, Represents the conjugate (inverse) of a quaternion.

[0071] The rotation increment quaternion The calculations are not merely abstract mathematical operations; they have a clear physical purpose. Through calculations... The system achieves rigid body motion decoupling from the control end (first terminal) to the controlled end (audience terminal). Among these, The introduction of this technology aligns the gravity references of different heterogeneous devices, eliminating control deviations caused by differences in handheld angles. The resulting standardized control commands are then converted into underlying PWM drive pulse signals for the actuators of the target terminal (such as the gimbal motor of a mobile phone or the ESC of a drone), thereby transforming pose changes in virtual space into physical mechanical motion.

[0072] The gravity compensation matrix The generation process includes: real-time invocation of the raw gravity vector collected by the built-in inertial measurement unit (IMU) of the audience terminal (controlled end). This is used as the horizontal plane constraint of the reference coordinate system; simultaneously, the gravity offset of the first terminal (control end) is acquired. The two are aligned using a coordinate transformation matrix. This technique ensures... It only includes the user's pure intention rotation amount, completely decoupling the differences in physical placement angles of heterogeneous hardware.

[0073] Calculated It is then mapped to the controlled entity coordinate system:

[0074] The physical significance of the quaternion operation lies in the fact that, through incremental extraction at the mathematical level, it eliminates, at the physical level, the rigid body motion deviation caused by the different initial installation angles between the control device (such as a VR controller) and the controlled actuator (such as a three-axis motor). The system calculates the pure rotational increment. Converted into the underlying PWM motor drive signal or the rotation matrix of the virtual skeleton.

[0075] If the controlled entity is a physical gimbal, Convert to Euler angles instruction ;

[0076] If the controlled entity is a virtual avatar, It is applied directly to the head node of the virtual skeleton.

[0077] With this algorithm, the movement of the controlled subject remains linearly consistent with the user's subjective line of sight, regardless of whether the user is lying down or standing.

[0078] 5.3 Virtual Arrival and Interaction Between the Virtual and Real Worlds

[0079] Reference Appendix Figure 7 This invention supports the "asymmetric virtual presence carrier" mode.

[0080] When no physical devices are available in the target area, the system instantiates a virtual Avatar at the target's GPS coordinates.

[0081] 1. Coordinate Anchoring: Use VPS (Visual Positioning System) to accurately anchor the Avatar to the geographic coordinates of real streets.

[0082] 2. Two-way rendering:

[0083] Downlink: The viewpoint of the avatar in the virtual space is rendered and transmitted back to the remote user via the cloud.

[0084] Uplink distribution: The avatar's pose and movement are pushed in real time to AR glasses or mobile phones of users within the geofence.

[0085] 3. Interactive closed loop: Remote user waves -> Avatar waves at real street coordinates -> Passersby see the wave on the AR screen and respond -> Remote user hears the response through the Avatar's microphone.

[0086] 5.4 Delay Prediction and Network Adaptation

[0087] Reference Appendix Figure 4 To address network latency, the system introduces a Kalman filter prediction model.

[0088] The state equation is defined as:

[0089]

[0090] in, for The position and posture of the subject under constant control. This is the current control command.

[0091] The system uses this model to predict The system records the user's position and displays HUD guidance indicators (such as "pre-aiming star") on the control terminal's feedback screen, informing the user in advance of the execution point of the command, thus psychologically offsetting the visual lag. Simultaneously, the system monitors RTT in real time, automatically triggering a bit rate reduction mechanism when network congestion occurs, prioritizing the transmission of control command signaling.

[0092] 5.5 Safety Circuit Breaker Mechanism

[0093] Reference Appendix Figure 5 The system is equipped with dual security logic:

[0094] 1. Geofencing restricted areas: When the user's device enters a preset privacy-sensitive area (such as a private residence), the cloud automatically cuts off the video stream.

[0095] 2. Physical intervention priority: The end user has the highest authority. Any physical operation (such as touching the screen or blocking the lens) will instantly trigger the "circuit breaker" and forcibly take back control.

[0096] 5.6 Typical Application Examples

[0097] Example 1: User-controlled heterogeneous devices

[0098] The user wears a VR headset (the first terminal) and wants to control a robot dog that has been paired with them at home (the controlled entity).

[0099] Link establishment: Users can establish a communication link with the robot dog through direct ID connection or local area network discovery without going through geofencing screening.

[0100] Attitude mapping: The system acquires the Euler angles of the head-mounted display in real time and uses the formula described in claim 1 of this invention. Calculate the rotation increment.

[0101] Execution: The robot dog's head follows the user's head as it turns.

[0102] This embodiment demonstrates that the core of the present invention lies in the standardized mapping of heterogeneous postures, and social matching is only a specific way of establishing links.

[0103] Example 2: Non-specific viewpoint borrowing based on mobile phone gimbal

[0104] In this embodiment, user A wants to observe a sudden event remotely in real time. The system matches a real pedestrian B at the scene based on a dynamic geofencing algorithm and requests to borrow control of their mobile phone gimbal. User A rotates the VR headset, and the system calculates the quaternion increment of the headset's rotation and converts it into Pitch / Yaw commands for the mobile phone gimbal by combining it with a gravity compensation matrix. Pedestrian B only needs to hold the mobile phone, and the gimbal will automatically rotate according to user A's line of sight, realizing perspective transmission based on heterogeneous devices.

[0105] Example 3: Expert Collaborative Control of Remote Industrial Equipment

[0106] In this embodiment, the target terminal is a mobile robot located in the factory. A headquarters expert (the first terminal) controls the robot's robotic arm to perform precise operations through heterogeneous pose mapping. The system uses a Kalman filter algorithm to predict the expert's movement trend in the next second, and performs motion buffering on the robot side in advance when network latency fluctuates, ensuring smooth operation.

[0107] Example 4: Dynamic geofence inspection based on unmanned aerial vehicles (UAVs)

[0108] In this embodiment, the system locates a third-party drone over the target area based on a geofence. The user issues displacement commands via a motion-sensing glove. After gravity vector compensation, the commands eliminate yaw errors caused by the drone's own tilt, achieving high-precision remote flight guidance.

[0109] Example 5: Cross-device Augmented Reality (AR) Guided Interaction

[0110] In this embodiment, the first terminal marks the gaze point in real time on the returned image of the second terminal (the audience terminal). The HUD guidance markers seen by the second terminal holder through the AR screen are spatial coordinate points converted from the real-time gaze angle of the first terminal, thereby achieving the collaborative efficiency of "pointing and looking" and effectively compensating for the perceptual latency in interpersonal communication.

[0111] Referring to Figure 6, the system supports smooth viewpoint switching under multi-terminal collaboration. When the user issues a viewpoint switching command, the cloud engine first performs logical alignment of the two frames at the semantic level; if the target is a physical viewpoint, the system calls the physical terminal camera stream; if the target is a virtual viewpoint, the system calls the Avatar virtual camera stream; finally, a unified downlink video stream is synthesized and pushed to the control end, ensuring that there is no black screen or jump during the switching process.

[0112] Example 6: Real Street Interaction Based on Virtual Carrier

[0113] Scenario: User A wants to virtually tour a pedestrian street in a foreign country, but there are currently no available physical robots to control at that location. Virtual Descent Execution:

[0114] 1. Carrier Instantiation: When user A selects "Virtual Descent", the system renders a personalized 3D virtual avatar at the center point of the pedestrian street requested by user A based on VPS visual positioning technology, and anchors its coordinates in the real geographical environment (XYZ).

[0115] 2. Motion Mapping: User A walks around at home, and their motion capture data is mapped in real time onto a virtual avatar on the pedestrian street. At this time, real pedestrians within this dynamic geofence can see User A walking on the street like an "asymmetric virtual presence" through handheld AR devices.

[0116] 3. Two-way interaction: User A meets User B, who also appears in the virtual world, and the two engage in virtual social interaction on a real street; at the same time, User A greets a real pedestrian passing by, and after the pedestrian's device receives the social request, the two achieve cross-time and space voice and action communication.

[0117] 5.7 Cross-Patent Collaborative Logic The [Remote Presence Control Method, Device, and System for Heterogeneous Device Collaboration] described in this application is a component of a complete spatiotemporal information processing system. This system includes:

[0118] The distribution side (related patent / the applicant's patent application filed on the same day entitled "An Information Distribution Method, System, Device and Medium Based on Dynamic Geofencing") is responsible for the accurate push and persistent storage of discrete information based on dynamic geofencing and motion vectors;

[0119] Interactive side (this patent / a remote presence control method, device and system for heterogeneous device collaboration): responsible for enabling remote terminal to control the posture of physical space entities based on social distance and permissions;

[0120] The backtracking side (related patent / patent application filed by the applicant on the same day entitled "A Data Processing Method, Device, Medium and System Based on Spatiotemporal Index") is responsible for AI scene reconstruction and backtracking based on the discrete slices generated by the aforementioned patent using the XYZT index.

[0121] 5.8 Device, Medium and System Implementation

[0122] 5.8.1 Remote Presence Control Device Based on Heterogeneous Device Collaboration This embodiment of the invention provides a data processing device, the hardware structure of which includes:

[0123] Processor: This can be one or more central processing units (CPUs), graphics processing units (GPUs), or application-specific integrated circuits (ASICs). The processor is used to execute computer programs stored in memory to perform the node matching, heterogeneous pose mapping, quaternion calculation, and Kalman filter prediction steps described in Sections 5.1 to 5.5 above.

[0124] Memory: Used to store computer programs and temporary data (such as geofence lists and real-time pose queues). Memory may include high-speed random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0125] Communication interface: Used to establish a low-latency bidirectional communication link with cloud servers and other terminals, supporting wireless communication protocols such as 5G / WiFi, and responsible for sending and receiving video streams and control commands.

[0126] Bus: Used to connect the processor, memory and communication interface to enable high-speed transmission of internal data.

[0127] 5.8.2 Remote Presence Control System Referring to Figure 1, this embodiment of the invention also provides a remote presence control system, which includes:

[0128] The first terminal (control terminal): an electronic device (such as a VR headset, AR glasses, or smartphone) equipped with a posture sensor (such as an IMU) and a display module. It is used to collect the user's real-time head / hand movement data and display the visual images transmitted from the audience terminal and HUD guidance signs.

[0129] Audience terminal (execution end):

[0130] Physical type: Intelligent devices connected to physical actuators (such as mobile phones with gimbals, connected robots, and drones) are used to respond to control commands and perform mechanical actions.

[0131] Virtual type: Terminal devices with AR display capabilities are used to receive virtual avatar data distributed from the cloud and render and display it locally.

[0132] Cloud server: A computing cluster deployed with a dynamic geofencing algorithm engine and a heterogeneous pose mapping engine. It is responsible for handling multi-terminal route matching, coordinate system transformation matrix calculation, cloud rendering of virtual avatars, and security circuit breaker monitoring.

[0133] 5.8.3 Computer-readable storage medium This embodiment of the invention also provides a computer-readable storage medium (such as a USB flash drive, portable hard drive, ROM, optical disc, etc.) storing computer instructions thereon. When executed by a processor, these instructions can implement the steps of the methods described in Sections 5.1 to 5.5 of this specification and claims 1 to 7, including but not limited to: receiving access requests, filtering audience terminals, calculating rotation increment quaternions, generating a gravity compensation matrix, performing virtual arrival interaction, and performing time delay prediction compensation.

[0134] VI. Conclusion

[0135] This application complies with relevant data security regulations, and all interactions are based on explicit authorization.

[0136] It should be understood that this specification is only a preferred embodiment of the present invention. For those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention (for example, replacing quaternion operations with rotation matrix operations, or replacing Kalman filtering with particle filtering). These improvements and modifications should also be considered within the scope of protection of the present invention.

[0137] Furthermore, the terms "first terminal" and "audience terminal" mentioned in the embodiments of this invention are only used to distinguish different device objects and do not limit their specific hardware form. Those skilled in the art will understand that any computing device with attitude acquisition and communication capabilities can serve as a terminal node in this system. Content not described in detail in this specification belongs to prior art known to those skilled in the art.

Claims

1. A telepresence control method, characterized by, include: Communication link establishment steps: In response to a control request initiated by the first terminal for the controlled subject, a real-time communication link is established between the first terminal and the controlled subject; the controlled subject includes a physical actuator or a virtual avatar terminal; Heterogeneous attitude mapping steps: Real-time acquisition of pose data of the first terminal, calculation of the rotation increment quaternion of the first terminal relative to the initial state, and mapping of the rotation increment quaternion to the coordinate system of the controlled subject in combination with the gravity vector compensation matrix to generate standardized control commands; Remote agent execution steps: Send the standardized control command to the controlled subject side, drive the controlled subject to perform actions synchronized with the first terminal, and send real-time perception feedback back to the first terminal.

2. The method of claim 1, wherein, The communication link establishment steps specifically include: The server receives an access request from the first terminal for the target geographic coordinates; Based on location services (LBS) and dynamic geofencing algorithms, matching audience terminals within the neighborhood of the target geographic coordinates are selected as the controlled subject; The audience terminal is a non-specific active device that is within the dynamic geofence and meets the social tag matching conditions. There is no preset physical binding relationship between the first terminal and the audience terminal before the link is established.

3. The method according to claim 1, characterized in that, In the isometric pose mapping step, the rotation delta quaternion The calculation formula is: in, This is the real-time quaternion of the first terminal at the current moment. The calibration quaternion at the initial moment, This is the gravity compensation matrix calculated based on the gravitational acceleration vector. It represents quaternion multiplication operations; the gravity compensation matrix is ​​used to align the heterogeneous physical horizontal planes of the first terminal and the controlled subject to eliminate the control deviation caused by the difference in the holding angle of the device.

4. The method according to claim 1, characterized in that, When the controlled entity is a virtual avatar terminal, the remote agent execution steps specifically include: Based on the standardized control commands, the virtual avatar terminal is driven to perform synchronous displacement and skeletal posture changes in the virtual space corresponding to the target geographic coordinates; The real-time image of the virtual avatar terminal is transmitted to the first terminal, and the actions of the virtual avatar terminal are distributed in real time and superimposed on the display interface of other active terminals within the same dynamic geofence using augmented reality (AR) technology, thereby realizing the interaction between virtual entities and physical entities in the real geographic space.

5. The method according to claim 1, characterized in that, The method also includes time delay prediction and visual compensation steps: The Kalman filter algorithm is used to predict the motion state of the controlled subject in the next moment based on the control command at the current moment; In the returned image received by the first terminal, a head-up display (HUD) guide sign is superimposed based on the predicted motion state to compensate for the physical delay caused by network transmission at the visual and psychological level.

6. The method according to claim 1, characterized in that, The method also includes safety circuit breaking and privacy protection steps: Real-time monitoring of the round-trip time (RTT) of the communication link and the geographical location of the controlled entity; When the RTT exceeds a preset threshold, the bit rate of the returned image is automatically reduced and the transmission bandwidth of the control command is prioritized. When the controlled entity enters a preset geofence restricted area, or when physical intervention is detected on the controlled entity side, the control permissions of the first terminal are immediately interrupted.

7. The method according to claim 1, characterized in that, The instruction sequence issued by the first terminal, the motion trajectory of the controlled subject, and the feedback of perception are associated with spatiotemporal tags (XYZT) and persistently stored to support third-party asynchronous backtracking systems in performing scenario reconstruction and evolution simulation based on spatiotemporal indexes.

8. A remote presence control device, comprising a processor and a memory, the memory storing a computer program that, when executed by the processor, implements the method as described in any one of claims 1 to 7.

9. A remote presence control system, characterized in that, include: The first terminal is used to collect user pose data and display the transmitted screen. The audience terminal, as a controlled entity, is used to drive physical actuators or carry out virtual image rendering, and to feed back environmental perception data; The cloud server is equipped with a heterogeneous pose mapping engine, which is used to perform pose command transformations in heterogeneous coordinate systems and cloud rendering of virtual avatars.

10. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the method as claimed in any one of claims 1 to 7.