Multi-terminal scene synchronization control method and device, equipment and medium

The control terminal receives virtual scene information from the user terminal, generates a synchronous preview screen, disables operation input, and sets target viewing parameters, which solves the problem of insufficient viewing angle control in multi-terminal interaction, realizes dynamic guidance and picture control, and improves user experience and interaction efficiency.

CN120455639APending Publication Date: 2025-08-08CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510584068.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art cannot realize the control terminal's takeover and synchronous rendering control of the virtual scene position and viewing angle of the user terminal, resulting in the lack of effective viewing angle guidance and operation permission coordination mechanism in the multi-terminal interaction process, affecting user experience and interaction efficiency.

Method used

The control terminal receives the current virtual scene position and viewing angle information of the user terminal, generates a synchronous preview screen, and triggers the anti-control mode to disable the user terminal's operation input signal acquisition function, sets the target virtual scene position and viewing angle parameters through the interactive interface, and generates rendering execution instructions to control the rendering process of the user terminal.

Benefits of technology

It realizes dynamic guidance and picture control of the user terminal from the control terminal to avoid viewing angle conflicts, improves the consistency and controllability of multi-terminal interaction processes, and improves the synchronization and user experience of remote guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455639A_ABST
    Figure CN120455639A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of equipment operation and maintenance, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-terminal scene synchronization control method, which comprises the following steps: a control terminal receives current virtual scene position and current virtual scene view angle information sent when a user terminal starts a virtual scene, and generates a synchronization preview picture based on the information; the control terminal triggers an anti-control mode and sends an operation authority forbidding instruction to the user terminal, and the user terminal forbids an operation input signal acquisition function after receiving the instruction; and the control terminal sets a target virtual scene position and a target virtual scene view angle parameter of the user terminal through the interactive interface, generates a rendering execution instruction based on the target parameter, and controls the user terminal to load the instruction to complete virtual scene rendering. The target virtual scene position and the view angle parameter of the user terminal are set at the control end, dynamic guidance and picture control of the control end on the user terminal are achieved, and view angle conflicts are avoided in cooperation with an operation authority forbidding mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of equipment operation and maintenance technology, and in particular to a multi-terminal scene synchronization control method, device, equipment and storage medium. Background Art

[0002] As multi-terminal collaboration and immersive interactive scenarios become increasingly prevalent, remote visual guidance and state control of user terminals by control terminals are becoming increasingly essential across multiple industries. Currently, in virtual interactive environments, particularly those involving image perspective control and user state synchronization, a series of technical bottlenecks remain, severely impacting terminal usability and the user experience.

[0003] In scenarios like sales demonstrations and product showcases, terminal devices are often used to provide an immersive visual experience. However, due to user unfamiliarity with the operating procedures and uncontrollable feedback paths, the user's perspective can shift randomly or lag behind the explanation, affecting the pace of explanation and comprehension efficiency. Especially when users first encounter such devices, independent operation makes it difficult to accurately focus on the key points of the presentation. Sales staff lack the means to achieve real-time control of the terminal and visual guidance, resulting in a fragmented experience and erroneous information delivery.

[0004] In the fintech business, scenarios such as remote financial advisory, robo-advisory training, and complex product demonstrations feature complex financial products with high barriers to understanding. Users rely on visual guidance and explanations during the interaction process. However, existing solutions typically only provide passive content playback and cannot dynamically adjust the displayed content based on the user's perspective. When the end user is actively interacting, it is difficult for financial instructors to intervene or control the user's terminal's focus, resulting in delayed or out-of-focus presentation of key information.

[0005] In the healthcare sector, remote teaching, remote rehabilitation training, and surgical plan demonstrations are increasingly common. Participants in these applications often lack familiarity with interactive devices, particularly when it comes to controlling the position and perspective of visual scenes. Relying on users to perform these operations can result in key perspectives not being displayed in a timely manner, preventing instructors or supervising physicians from guiding users through effective observation or interaction at critical moments, impacting teaching or rehabilitation effectiveness.

[0006] Existing technologies typically only support user-initiated actions or fixed content delivery. The lack of an effective control path prevents the controller from proactively setting the user terminal's position and viewing angle. There's also no synchronous state feedback mechanism to generate real-time images for the controller. Furthermore, the lack of a means to take over user input permissions means that during processes requiring centralized guidance, user actions may conflict with controller instructions, leading to abrupt screen jumps, out-of-focus rendered content, and even discomfort or dizziness.

[0007] Therefore, the existing technology still has obvious shortcomings in multi-terminal screen status control, synchronous takeover of operation permissions, and generation and execution of control instructions based on perspective status. It is difficult to meet the actual needs of precise regulation of terminal user status and perspective consistency in multiple scenarios such as sales display, financial guidance and medical interaction. Summary of the Invention

[0008] The main purpose of the present invention is to provide a multi-terminal scene synchronization control method, device, equipment and storage medium, aiming to solve the technical problem that the existing technology cannot achieve the control end to take over and synchronously render the user terminal virtual scene position and perspective, resulting in a lack of effective perspective guidance and operation authority coordination mechanism during multi-terminal interaction.

[0009] To achieve the above objectives, the present invention provides a multi-terminal scene synchronization control method, comprising:

[0010] The control terminal receives the current virtual scene position and current virtual scene viewing angle information sent by the user terminal when the virtual scene is started;

[0011] The control terminal generates a synchronous preview image based on the received current virtual scene position and current virtual scene viewing angle information;

[0012] The control end triggers the reverse control mode and sends an operation permission disabling instruction to the user terminal, wherein the operation permission disabling instruction is used to disable the operation input signal collection function of the user terminal;

[0013] The control end sets the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through the interactive interface;

[0014] The control end generates a rendering execution instruction based on the target virtual scene position and the target virtual scene viewing angle parameters;

[0015] The control end controls the user terminal to load the rendering execution instruction to render the current virtual scene.

[0016] Furthermore, to achieve the above-mentioned purpose, the present invention provides a multi-terminal scene synchronization control device, comprising:

[0017] The virtual scene perception module is used for the control end to receive the current virtual scene position and current virtual scene viewing angle information sent by the user terminal when the virtual scene is started;

[0018] A synchronous preview rendering module is used for the control end to generate a synchronous preview image based on the received current virtual scene position and current virtual scene viewing angle information;

[0019] The authority control module is used to control the terminal to trigger the reverse control mode and send an operation authority disabling instruction to the user terminal, wherein the operation authority disabling instruction is used to disable the operation input signal collection function of the user terminal;

[0020] A target parameter setting module is used for the control end to set the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through an interactive interface;

[0021] An instruction generation module is used for the control end to generate a rendering execution instruction based on the target virtual scene position and the target virtual scene viewing angle parameters;

[0022] The remote rendering control module is used to control the user terminal to load the rendering execution instruction to render the current virtual scene.

[0023] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a multi-terminal scene synchronization control program stored in the memory and runnable on the processor. When the multi-terminal scene synchronization control program is executed by the processor, the steps of the multi-terminal scene synchronization control method described above are implemented.

[0024] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a multi-terminal scene synchronization control program is stored. When the multi-terminal scene synchronization control program is executed by a processor, the steps of the multi-terminal scene synchronization control method described above are implemented.

[0025] Beneficial effects: The present invention relates to the field of equipment operation and maintenance technology, and can be applied to business scenarios such as financial technology and medical health. A multi-terminal scene synchronization control method is disclosed, including: the control end receives the current virtual scene position and the current virtual scene perspective information sent by the user terminal when the virtual scene is started, and generates a synchronous preview screen based on the information; the control end triggers the reverse control mode and sends an operation permission disabling instruction to the user terminal, and the user terminal prohibits the operation input signal acquisition function after receiving the instruction; the control end sets the target virtual scene position and target virtual scene perspective parameters of the user terminal through the interactive interface, generates a rendering execution instruction based on the target parameters, and controls the user terminal to load the instruction to complete the virtual scene rendering. The present invention realizes dynamic guidance and screen control of the user terminal by the control end by setting the target virtual scene position and perspective parameters of the user terminal on the control end, and generates a rendering execution instruction based on the parameters, while cooperating with the operation permission disabling mechanism to avoid perspective conflicts, making the multi-terminal interaction process more consistent and controllable, and improving the screen synchronization and user experience of remote guidance. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0027] Figure 1 A schematic diagram of an application environment of a multi-terminal scene synchronization control method according to an embodiment of the present invention;

[0028] Figure 2 This is a flow chart of an embodiment of a multi-terminal scene synchronization control method of the present invention;

[0029] Figure 3 This is a functional module diagram of a preferred embodiment of the multi-terminal scene synchronization control device of the present invention;

[0030] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0031] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0032] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0033] The multi-terminal scene synchronization control method provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can receive the current virtual scene position and current virtual scene perspective information sent by the user terminal when the virtual scene is started through the control terminal of the user terminal, and generate a synchronous preview screen based on the information; the control terminal triggers the reverse control mode and sends an operation permission disabling instruction to the user terminal, and the user terminal prohibits the operation input signal acquisition function after receiving the instruction; the control terminal sets the target virtual scene position and target virtual scene perspective parameters of the user terminal through the interactive interface, generates a rendering execution instruction based on the target parameters, and controls the user terminal to load the instruction to complete the virtual scene rendering. The present invention realizes dynamic guidance and screen control of the user terminal by the control terminal by setting the target virtual scene position and perspective parameters of the user terminal on the control terminal, and generates a rendering execution instruction based on the parameters, while cooperating with the operation permission disabling mechanism to avoid perspective conflicts, making the multi-terminal interaction process more consistent and controllable, and improving the screen synchronization and user experience of remote guidance. Among them, the user terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0034] See also Figure 2 , Figure 2This is a flow chart of an embodiment of a multi-terminal scene synchronization control method provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than here.

[0035] like Figure 2 As shown, the multi-terminal scene synchronization control method proposed by the present invention includes the following steps:

[0036] S10, the control end receives the current virtual scene position and current virtual scene viewing angle information sent by the user terminal when starting the virtual scene;

[0037] In this embodiment, the control terminal can be an intelligent terminal device running on a desktop operating system or a mobile platform, such as a personal computer, a tablet computer, a smart phone, an interactive touch screen, a digital console, or a lightweight industrial control terminal that supports graphics rendering. The control terminal usually has a graphical interface output capability and a remote data interaction capability, and is suitable for scenarios that require real-time monitoring and management of user terminal behavior. In different application environments, the actual operating subject of the control terminal will also vary. For example, in a sales demonstration scenario, the control terminal can be operated by a salesperson, a lecturer, or a training instructor; in a medical teaching scenario, the control terminal can be controlled by a surgeon, a teaching lecturer, or a remote supervisor; in a financial product demonstration scenario, it can be managed by a financial advisor, an insurance agent, or a customer manager.

[0038] User terminals are usually device types with immersive display capabilities and spatial perception capabilities (such as VR devices), including but not limited to head-mounted displays, augmented reality glasses, panoramic projection terminals, or smart mobile devices equipped with spatial positioning modules. These terminals need to support graphics rendering, perspective tracking, and network communication capabilities to enable local loading and rendering of virtual scenes. In different application fields, the users of user terminals are also differentiated. For example, in retail experience scenarios, user terminals are operated by potential customers or product experiencers; in medical and health scenarios, user terminals can be used by trained medical staff, patients, or rehabilitation training participants; in financial education scenarios, they may be interacted with by financial product audiences, customer representatives, or compliance training participants.

[0039] The control end receives the current virtual scene position and the current virtual scene perspective information sent by the user terminal when the virtual scene is started. This process starts with the scene initialization operation of the user terminal. The current virtual scene position refers to the spatial coordinate position of the user terminal in the three-dimensional virtual environment. A three-dimensional vector (such as x, y, z) is usually used to represent its specific positioning in the virtual scene. The source of this position information is generally the physical space coordinate data generated by the spatial positioning sensor or inertial measurement unit. The user terminal then calls the preset spatial transformation rules or coordinate mapping matrix to convert it into three-dimensional coordinates in the virtual scene coordinate system. In virtual reality, human-computer interaction and smart wearable devices, this three-dimensional coordinate conversion usually uses Euler angle transformation, homogeneous transformation matrix, quaternion rotation and other forms to construct the mapping relationship between the virtual scene and the real space.

[0040] The current virtual scene perspective information represents the user's viewing direction and pitch angle within the virtual scene. This information is typically captured by the head posture tracking module, using a direction vector after the head posture changes and combined with the posture angle value to represent it. This perspective information can be obtained through the fusion of gyroscope and accelerometer data from the IMU module. Alternatively, it can be obtained through real-time tracking using a visual inertial navigation system (VINS) to generate combined directional angle values, such as horizontal yaw and vertical pitch. This perspective information not only reflects the user's current viewing direction but is also used in subsequent rendering logic to generate main view parameters, ensuring that the preview image is consistent with the user's viewpoint.

[0041] The user terminal needs to package the above-mentioned current virtual scene position and current virtual scene perspective information into a structured data object for transmission. During this process, the structured data packet often contains multiple fields, including the scene position field, the perspective angle field, the terminal identification field, the timestamp field, etc. The encapsulation method of this structured data packet can adopt JSON, Protocol Buffers or a custom binary protocol format to balance parsing efficiency and data integrity. To achieve secure data transmission across network links, the user terminal will call a communication module based on TLS or a custom encryption protocol after constructing the structured data packet to send the data to the intermediate backend service.

[0042] The backend service is responsible for forwarding and verification. After receiving a data packet from a user terminal, the backend service will perform a validity check based on the device identification field, including parsing the digital signature, key verification, or whitelist matching to ensure that the data is indeed generated by a valid user device. After the data validity verification is passed, the backend service will forward the parsed payload to the control terminal, which will then extract the virtual scene position and virtual scene perspective information contained therein and use it as input for subsequent synchronous preview screen generation and rendering instruction calculations.

[0043] After receiving this data, the control terminal's data parsing module first extracts location and viewing angle information from the structured data and converts this information into internal camera control parameters based on pre-set graphics engine standards. These parameters are used to construct the viewpoint and direction vector of the control terminal's virtual camera, accurately restoring the user's current viewing angle. This restoration process not only provides a real-time preview for sales personnel or control operators but also provides a precise input reference for subsequent guidance and control operations, ensuring consistent viewing angles and state synchronization across terminals.

[0044] In an edge computing environment, user terminals can process raw gesture data into structured data packets through local edge devices, which then forward and perform preliminary verification on their behalf, reducing transmission latency and cloud computing pressure. In a centralized architecture, all data transmissions can be distributed and synchronized through a unified messaging middleware system, such as message routing using topic subscriptions based on the MQTT protocol.

[0045] The design of the virtual scene coordinate transformation matrix can be adjusted based on the spatial perception accuracy of different terminals. For head-mounted displays equipped with high-precision SLAM modules, a real-time generated local map can be used as the basic reference frame for coordinate transformation. For lightweight devices that rely on IMU data, a static calibration matrix or external calibration results can be used for coordinate mapping. During information transmission, if the business scenario requires high data confidentiality, such as in medical education or financial product demonstrations, end-to-end encryption mechanisms can be used and signature verification can be implemented through national secret algorithms or elliptic curve cryptography mechanisms to ensure the security and reliability of the data link.

[0046] To accommodate different rendering platforms and terminal types, the field settings of structured data packets can also be dynamically adjusted. For example, in augmented reality devices with depth cameras, the scene position field can be accompanied by a depth map index value to assist with rendering calculations. In graphics workstations with high frame rate requirements, a frame synchronization flag field can be introduced to improve frame-level consistency between the image preview and the actual perspective of the user terminal.

[0047] Example: In the healthcare sector, during a remote surgical demonstration, a doctor wants to monitor the student's viewing angle while wearing the device in real time, so they can focus on specific organ structures during critical procedures. The doctor can obtain the student's scene position and viewing angle in real time and generate a corresponding image on their own to determine whether viewing angle correction is necessary.

[0048] In financial business training scenarios, instructors use user terminals to monitor trainees' observations in virtual investment simulations, assessing their understanding of the asset structure presented. If they deviate from key perspectives, they can provide subsequent instructions to guide their perspective. This mechanism makes high-value product explanations more targeted and pacing-controlled, improving the efficiency of remote marketing transactions.

[0049] In immersive retail scenarios, sales staff can check in real time whether the user's perspective remains on a specific product area, and generate interactive recommendation content based on this. At the same time, they can actively switch perspectives in conjunction with subsequent counter-control mechanisms to guide users through the key product introduction process, thereby greatly improving the interactive control and display effects in sales scenarios.

[0050] Example: In the medical and health business field, during remote surgical teaching, the control terminal can be deployed on the surgeon's tablet terminal or interactive control workstation to receive the current virtual scene position and current virtual scene perspective information sent by the user terminal. The user terminal worn by the student, such as a head-mounted visual device, synchronously encodes and uploads the three-dimensional coordinate position and head posture perspective in the virtual space. The doctor uses the control terminal to check in real time whether the student's perspective is focused on the surgical incision area or important anatomical structures. If the student's observation direction is found to deviate from the key area, the control terminal can further generate a corresponding preview screen for judgment, and prepare to enter the reverse control mode to guide the student's perspective, thereby ensuring the focus of the teaching process and the accuracy of the operation understanding.

[0051] In financial business training scenarios, the controller is typically operated by the instructor and can be a PC or a touch-enabled demonstration device. Trainees enter the virtual investment and financing environment through the user terminal, and their perspective data and scene coordinates are regularly pushed to the controller. The instructor can access the current virtual scene location and perspective information in real time on the controller to determine whether the user is accurately observing key visualizations such as asset distribution maps and return simulation charts. If the perspective wanders from the key training points, the controller can set the target virtual scene perspective parameters to guide the trainee back to the designated structure, achieving efficient, low-intrusion, and immersive teaching guidance.

[0052] In immersive retail scenarios, the control terminal is held by sales staff and is usually a mobile terminal or dedicated sales device that supports graphics rendering. User terminals such as VR glasses report the user's position in the current virtual store and the direction of product viewing through the spatial positioning module. After the control terminal receives and analyzes the current virtual scene position and the current virtual scene perspective information in real time, the sales staff can accurately determine whether the user is paying attention to the core product area or key booth. If the user stays in the marginal product area for a long time, the control terminal can further activate the anti-control mode and set the target virtual scene position and target virtual scene perspective parameters, triggering the scene adjustment process of the user terminal, guiding the user back to the key exhibit area, thereby improving display efficiency and sales conversion rate.

[0053] By encapsulating the user terminal's current position and viewing angle information as structured data, and parsing and restoring the image through the control end, the control end accurately perceives the user's current state. This process establishes a closed-loop link from user-side motion capture to control-end image feedback, providing a high-precision input foundation for viewing angle control and position guidance in subsequent interactions. This not only achieves information synchronization between the control end and the user terminal, but also improves the coherence and controllability of multi-terminal collaborative interactions, avoiding experience interruptions caused by information delays or viewing angle deviations.

[0054] S20, the control end generates a synchronous preview image based on the received current virtual scene position and current virtual scene viewing angle information;

[0055] In this embodiment, the core of the control end's generation of a synchronized preview screen lies in real-time mapping and restoring the current virtual scene position and perspective information uploaded by the user terminal to a local graphical display screen on the control end. The current virtual scene position typically refers to the spatial coordinates of the user terminal in three-dimensional virtual space, such as three-dimensional coordinate values based on the X, Y, and Z axes. The data source is generally the user terminal's spatial positioning sensor or inertial navigation unit. This position parameter is used to reconstruct the user's scene observation point on the control end.

[0056] The current virtual scene perspective information includes the orientation of the virtual camera in the user terminal, primarily composed of the horizontal yaw angle (yaw) and vertical pitch angle (pitch). This data typically comes from a head posture tracking module or gyroscope system. After receiving this data, the control terminal first parses the coordinate information and direction angle parameters to set a virtual camera node, ensuring that the camera node's position and orientation are consistent with the status in the user terminal.

[0057] The control end then loads 3D resource files consistent with the virtual scene model used by the user terminal. These resource files may contain information such as static models, dynamic textures, and ambient lighting parameters. The control end applies the reconstructed virtual camera parameters to the graphics rendering engine. Using the projection matrix and viewport configuration, the control end performs local real-time rendering of the user's current viewing position and direction. Physically Based Rendering (PBR) technology is typically used to achieve lighting and material consistency, and shadow mapping and anti-aliasing techniques can also be combined to improve image quality.

[0058] The final image generated is the synchronized preview image, which is essentially a mirror image generated in real time at the control end based on the current viewing angle of the user terminal. It is used to assist the operator at the control end in judging in real time whether the user's viewing content is focused on the key area.

[0059] In one implementation, the control end runs an independent rendering engine process specifically for parsing the data stream sent by the user terminal and constructing a virtual scene replica for display. The control end can use an open source engine (such as Unity or UnrealEngine) as the rendering backend, accepting structured data packets forwarded from the backend service through an open API interface, and refreshing camera parameters and scene status in real time. Another approach is to integrate the rendering service and scene data processing service on the server side. The control end only serves as a preview display terminal, and uses streaming rendering technology to achieve low-latency presentation of the generated image. This method is more suitable for mobile control end scenarios with limited rendering resources.

[0060] In the scenario of multiple users accessing concurrently, the control end can also maintain the viewing status of multiple user terminals at the same time, and display the viewing status of each user terminal in a multi-screen view on the same screen, thereby adapting to multi-point control scenarios such as medical teaching or corporate training.

[0061] Example: In the healthcare business, doctors use the control terminal to generate a real-time preview of the student's current perspective, allowing them to accurately determine whether the student is focusing on a certain organ area. If they find that the student is deviating from the critical surgical area, the doctor can decide whether to send a perspective adjustment instruction to improve the targeted teaching.

[0062] In financial business training, instructors can preview the trainees' perspectives through the control terminal to identify whether they have understood the key structures of the portfolio diagram or financial flow diagram. Once deviations or misinterpretation risks are found, prompts or interventions can be given to improve the efficiency and accuracy of the explanation.

[0063] In the retail experience scenario, sales staff can judge whether the customer has been staring at a certain product area for a long time based on the preview screen generated by the control end. If the customer is not focused on the core product, they can make product recommendations or auxiliary guidance based on the screen, thereby enhancing the interactivity of the sales behavior and the transaction conversion rate.

[0064] By generating a preview screen consistent with the user terminal in real time on the control end, the operator can understand the user's current position and viewing direction without wearing a VR terminal, thereby more efficiently identifying user attention deviation, observation blind spots or understanding deviations, providing basic data support for subsequent perspective adjustment and counter-control operations, and effectively improving the controllability and response efficiency during remote interaction.

[0065] S30, the control end triggers the reverse control mode and sends an operation permission disabling instruction to the user terminal, wherein the operation permission disabling instruction is used to disable the operation input signal collection function of the user terminal;

[0066] In this embodiment, triggering reverse control mode on the control end activates a state in which the control end directs the user terminal's behavior. In this state, the user terminal no longer responds to operation signals generated by local input devices and instead fully accepts control instructions from the control end. Reverse control mode is typically triggered by the control end operator selecting a corresponding control option in the interactive interface. This process generates a reverse control flag within the system and initiates the control permission switching process.

[0067] The Operation Permission Disable command is a control command packet sent from the control end to the user terminal. It contains fields such as an encrypted digital signature, a tamper-proof field, and a control-rejection status flag. This command is typically sent via a secure communication protocol like TLS to ensure data is not hijacked or modified by intermediate nodes during transmission. The core purpose of this command is to notify the user terminal to suspend or cancel its input device's ability to interact with the local environment.

[0068] After the user terminal receives the instruction, it first performs the legitimacy verification step of the digital signature to ensure that the instruction is indeed sent by the authorized control terminal and to prevent forged instructions from interfering with system behavior. Once the verification is passed, the user terminal enters the permission convergence process, which mainly includes two core actions: one is to call the device driver interface to close the signal acquisition channel with the input device (such as head-mounted interactive sensors, handles, motion capture modules, etc.) to prevent users from affecting scene behavior through local hardware; the other is to write the anti-control status flag bit in the system kernel space. This flag is used by the system process to determine whether the current device is in autonomous control state, thereby preventing other application modules from arbitrarily restoring user control permissions.

[0069] To improve system robustness, the user terminal activates a heartbeat detection thread, which periodically sends status confirmation requests to the control terminal and monitors the control terminal's responses. If multiple consecutive heartbeat responses are lost, the control terminal is considered disconnected. The system automatically clears the reverse control status flag and re-enables input signal acquisition to prevent prolonged user terminal inactivity due to control terminal anomalies.

[0070] In one specific implementation, the permission-disabling instruction is encapsulated in a JSON structure, containing encrypted fields, a device identifier, a reverse control state flag, and a timestamp. This instruction is pushed from the control end to the user terminal via a secure WebSocket channel. The user terminal integrates an input management module. Upon receiving the instruction, this module invokes the hardware access control interface provided by the operating system. For example, on Android, this module invokes the InputManagerService to change the input device state; on Linux, it invokes the evdev driver stack to disable the input device node.

[0071] In another specific implementation, the anti-control status flag is written to the user terminal's kernel-mode shared memory area, and a memory page locking mechanism prevents other threads from overwriting or modifying this state. The heartbeat mechanism uses a dual-channel timeout determination strategy: one channel is used to verify the control terminal status, and the other is used to detect activity on the data channel, ensuring accurate interruption detection even in complex network environments.

[0072] Furthermore, the control end can automatically enter counter-control mode without human intervention through pre-set triggering strategies, thereby managing the user terminal's operational permissions. The control end deploys a state perception and behavior judgment module to monitor the user terminal's behavioral parameters in the virtual scene in real time, including the rate of change of current position, frequency of viewpoint deviation, and dwell time distribution. When these parameters meet a pre-defined deviation threshold or abnormal behavior condition, the control end automatically generates a trigger signal, activating the counter-control mode initiation process. This initiation process includes automatically generating an operational permission disabling instruction, sending it to the user terminal via an encrypted channel, completing the instruction legitimacy verification, and disabling the input signal, without the need for human confirmation or approval. The control end can incorporate pre-set rule sets, such as when a user continuously moves within an invalid area for a specified period of time or when the viewpoint continuously deviates from the target area, as the basis for automatically determining whether counter-control mode is in effect.

[0073] To prevent false triggering or abuse, the system can also configure a cooldown mechanism. This means that after a certain period of time after the end of an automatic counter-control, the control terminal will not automatically initiate the next counter-control mode. Alternatively, a predictive prompt mechanism can be added to send an early warning message to the user terminal before the automatic counter-control is triggered, improving user predictability and acceptance. This automated mechanism improves management efficiency in multi-terminal collaboration scenarios and is particularly suitable for large-scale remote interactions where control resources are limited.

[0074] Example: In medical education, if a student frequently switches their viewing direction in a virtual scene, causing their observation to shift, the doctor on the control end can trigger a reverse control mode, causing the user terminal to temporarily ignore head movement input, preventing the student from straying from the view of key organs and enhancing the teaching focus. Alternatively, in remote medical training, if a student repeatedly shifts their view within a key explanation area and fails to focus on the key lesion model within a specified timeframe, the behavior judgment module on the control end will detect this behavior pattern and automatically trigger a reverse control mode, pausing the student terminal's input and forcing them to observe the organ position set by the doctor, thereby ensuring consistent teaching quality.

[0075] In financial simulation training scenarios, if instructors notice that trainees' frequent view switching affects their understanding of the content, they can disable their perspective control through a reverse control command, forcing them to focus on the currently displayed portfolio, thereby improving training efficiency. Alternatively, in financial education scenarios, if trainees repeatedly switch views within a complex asset structure chart and fail to maintain focus on the core data area, the control terminal will detect that the frequency of their perspective changes exceeds a threshold and automatically activate reverse control mode, redirecting their perspective to the target investment product page. This eliminates the need for manual intervention by the instructor and helps improve the stability of the overall training process.

[0076] In interactive retail displays, if sales staff determine that a user has skipped a key merchandise area, they can pause user input through the control terminal, keeping the user's view focused on high-value items and providing continuous visual guidance for subsequent recommendations. Alternatively, in a virtual retail experience, if a consumer lingers for an extended period in a non-key merchandise area, the control terminal system will recognize that their lingering behavior deviates from their purchase goals and, through strategic judgment, directly switch to a counter-control state, guiding the user's view to the promotional product page, thereby improving product display efficiency and achieving proactive marketing goals.

[0077] By precisely disabling the user terminal's input signal collection function, we can effectively prevent interference caused by accidental touches or misoperations during boot operations on the control end. Furthermore, the combination of a reverse control status flag and a heartbeat detection mechanism ensures the security and automatic recovery capabilities of control switching, improving the system's fault tolerance and user experience continuity in the event of abnormal disconnections.

[0078] S40, the control end sets the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through the interactive interface;

[0079] In this embodiment, the control end sets the target virtual scene position and target virtual scene viewing parameters of the user terminal through the interactive interface. The core is to build a visual configuration entrance for remotely setting the user's viewpoint and viewing direction. The interactive interface in the control end usually includes a visual display of the three-dimensional space coordinate axis, scene area annotation, and a positioning operation module for clicking and dragging. The target virtual scene position refers to the specific spatial coordinate position that the control end wants the user terminal to move to in the virtual scene, usually represented by a vector point in a three-dimensional coordinate system, such as (X, Y, Z). The setting of this position needs to refer to the current virtual scene space reference parameters, including the coordinate origin, axial direction, unit scale, etc., to ensure that the set position can be accurately restored in different terminals.

[0080] The target virtual scene's viewing angle parameters describe the target orientation of the user terminal's line of sight. These parameters primarily include the horizontal yaw angle (Yaw) and the vertical pitch angle (Pitch). These two angles form a complete three-dimensional viewing direction vector. The viewing angle parameters are typically calculated by combining the user terminal's current viewing angle information, as determined by spatial vector operations, to determine the desired viewing direction.

[0081] This interactive interface, connected to a 3D visualization engine, displays the user's current position and orientation within the virtual scene in real time on the control terminal. The operator sets the target position by clicking or dragging on the interface, and sets the target perspective using visual arrows, hotspot selection, or a standard perspective menu, completing the configuration of target parameters. The system then converts these settings into structured control parameters for subsequent generation of rendering execution instructions.

[0082] In one specific implementation, a 3D visualization component is embedded in the control interface, using an engine like WebGL or Unity to recreate and display the virtual scene. The operator can set the target coordinates of the user terminal by clicking on the 3D scene grid. The system automatically back-projects the clicked coordinates to obtain the world coordinates. The operator can also set the heading angle by dragging the virtual viewing angle indicator with the mouse. The system automatically calculates the yaw and pitch values as the target viewing angle parameters.

[0083] In another specific implementation, the control terminal can load preset viewpoint templates for key scene areas, such as a product display stand, a device control panel, or a data dashboard. Once the user selects a template, the system automatically populates the corresponding target virtual scene location and viewpoint parameters, eliminating the need for manual coordinate setting. This approach is suitable for batch control and standardized display processes.

[0084] You can also add auxiliary parameter setting modules to the control end, such as allowing the addition of current position offset, camera rotation inertia value or target position switching transition time, making the control action softer and the experience more natural.

[0085] Furthermore, in actual implementation, the control terminal's interactive interface can feature two-dimensional or three-dimensional visualization capabilities. Through this interface, control behaviors can be set manually or through an automated decision-making module that generates algorithm-driven parameters based on the current virtual scene state and user behavior data. This configuration approach is not limited to a single operating mode but supports strategies that combine multiple automated methods with manual intervention, including rule-driven, adaptive behavior prediction, and user behavior analysis.

[0086] By remotely setting the target virtual scene position and viewing angle parameters for the user terminal through the control terminal's interactive interface, users can be precisely guided to focus on the key areas or content the control terminal wants them to see, preventing them from getting lost in the virtual environment or straying from key interaction processes. This improves the controllability of user behavior paths during remote interactions and significantly enhances the consistency between scene guidance and content presentation, effectively resolving the issue of operational deviations caused by users' lack of spatial perception.

[0087] S50, the control end generates a rendering execution instruction based on the target virtual scene position and the target virtual scene viewing angle parameter;

[0088] In this embodiment, the control terminal generates rendering execution instructions based on the target virtual scene position and target virtual scene perspective parameters. This refers to processing existing three-dimensional spatial coordinates and direction angles into a rendering action sequence that can be executed sequentially on the user terminal to achieve remote guided content presentation. The target virtual scene position is typically represented as a three-dimensional coordinate vector, which can use floating-point numbers to represent the X, Y, and Z values of the spatial point. The target virtual scene perspective parameters are composed of horizontal deflection angles and vertical pitch angles, describing the direction of the viewing direction vector, and are often expressed as angle values or unit direction vectors.

[0089] The process of generating rendering execution instructions not only involves the acquisition of original parameters, but also involves key links such as parameter continuity reconstruction, execution rate adaptation, timing control identifier generation, and resource tag integration. To prevent the user terminal from feeling dizzy when the position or perspective suddenly changes, the control end needs to perform interpolation processing on the target virtual scene position, generate a continuous and smooth sequence of displacement vectors, and divide it into a set of position information that can be executed frame by frame at a fixed time step. Similarly, the target virtual scene perspective parameters also need to be converted into rotation steps executed on a frame-by-frame basis to achieve continuous changes in the viewing direction.

[0090] To adapt to the rendering performance of different user terminals, the control terminal typically needs to first obtain the current device's rendering performance parameters, including average frame rate, graphics processor load, resource usage, and other information. Based on this information, the control terminal dynamically sets the instruction sequence segmentation granularity and single-frame computational complexity. During the encapsulation process of rendering execution instructions, timestamp information for each frame sequence should be added to control the playback order. Priority identifiers are used to indicate the processing urgency or visual importance of a particular instruction segment, assisting the terminal in optimizing resource allocation when resources are limited.

[0091] The final rendering execution instruction is an encapsulated structure that contains fields such as timestamp, displacement vector, viewing angle step, resource priority, etc. Its purpose is to achieve guided evolution of spatial position and observation angle without relying on user operation, and to seamlessly connect with subsequent rendering processes.

[0092] In one specific implementation, the control end uses a spline interpolation algorithm to smoothly interpolate the target virtual scene position, generating a sequence of displacement vectors containing coordinate points at equal time intervals. This sequence is then linearly segmented into several frames of rotation step vectors based on the angular range between the starting and target viewing directions. These two vectors are then combined to form a pair of spatial and visual angle instructions that are executed on a per-frame basis. The system also automatically adjusts the interpolation density based on the terminal's rendering capabilities. For example, high-performance terminals can support higher frame rates and rendering step resolutions, while low-performance terminals can combine several frames into a single instruction segment to reduce processing complexity.

[0093] In another specific embodiment, the system uses a posture prediction model to dynamically adjust the viewing parameters of the target virtual scene, combines the user's current head movement trend to determine the possible natural turning path, and generates a step sequence with inertial smoothing characteristics in advance to improve the execution smoothness of instructions and reduce interference with the user experience.

[0094] The generated rendering execution instructions can also be stored as multiple hierarchical instruction packages, which can be called on demand based on task urgency, user behavior feedback, or system abnormality. For example, when the user deviates from the expected trajectory, a high-priority correction instruction package can be automatically switched to, and when the network fluctuates, the standard instruction package with low resource requirements can be downgraded to load, thereby achieving a highly dynamically adaptive guided execution strategy.

[0095] By generating rendering execution instructions on the control side and loading them into the user terminal for execution, the target's spatial position and viewing direction can be adjusted synchronously, improving the operational efficiency and visual comfort of remote guidance capabilities. The introduction of an interpolation mechanism for displacement vector sequences and viewing angle steps significantly reduces visual dizziness caused by sudden scene changes. Dynamic frame sequence partitioning and priority marking ensure that control actions remain executable and stable under varying terminal performance conditions.

[0096] S60: The control end controls the user terminal to load the rendering execution instruction to render the current virtual scene.

[0097] In this embodiment, the control end instructs the user terminal to load rendering execution instructions to render the current virtual scene. This means that the control end sends instruction data containing a sequence of continuous frame operations to the user terminal. The user terminal's graphics processing flow then parses and executes these instructions frame by frame, thereby presenting the corresponding virtual scene dynamics. The rendering execution instruction is essentially an executable data structure containing a sequence of displacement vectors and a viewpoint deflection step size, encapsulating the rendering task of each frame of continuous spatial movement and viewing direction adjustment.

[0098] During this process, the control end not only acts as a data generator but also assumes the responsibilities of a command dispatcher, including command timing control and communication scheduling. After receiving the rendering execution command sent by the control end, the user terminal first parses the executable frame sequence contained therein and determines the execution order based on the timestamp information attached to the command. To ensure smooth rendering and priority responsiveness, the user terminal also needs to extract the priority identifier from each frame sequence. This identifier is used to indicate the processing level of each rendering segment, such as critical actions, high-risk transitions, and static transitions.

[0099] Based on the priority classification results, the user terminal allocates different levels of rendering frame segments to the GPU resource pool. High-priority frame sequences are typically processed by the GPU's exclusive compute unit to ensure the stability and timeliness of critical rendering stages; low-priority frame sequences are scheduled for asynchronous execution in idle threads to avoid system lags or resource conflicts.

[0100] After the frame sequence is distributed, the user terminal loads each frame in timestamp order and inputs it into the graphics rendering pipeline. The graphics rendering pipeline updates the position and orientation of the virtual camera based on the displacement vector and viewing angle step size in each frame, presenting a continuously changing virtual scene frame by frame on the user terminal's display module, completing the entire remote rendering guidance process.

[0101] In one implementation, after generating a rendering execution instruction, the control end immediately calls a low-latency communication protocol (such as QUIC or TCP-optimized frame channel) to send it to the user terminal. Upon receiving it, the user terminal establishes a frame instruction buffer and loads the previous frames into the video memory in the form of a ring buffer to ensure the time continuity of data processing. Before each frame is executed, the system checks the local clock to determine whether its timestamp matches the current system time, thereby determining whether the frame enters the graphics rendering pipeline.

[0102] In another implementation, the user terminal's graphics processor is equipped with a dynamic resource scheduling module that dynamically adjusts the resource allocation ratio for each priority frame sequence based on the current GPU load. In the event of terminal resource constraints or background task conflicts, low-priority rendered frames will be downsampled or merged into keyframes, reducing rendering overhead without affecting the execution of the main view command.

[0103] A command packet status verification module can also be added to the rendering process. When the user terminal renders each frame sequence, it first verifies its data integrity and field legitimacy. When it detects a timestamp jump, data loss, or abnormal deflection value in the rendering execution instruction, it can automatically request the control end to resend the frame sequence and roll back the current rendering status to the previous valid key frame.

[0104] By centrally controlling the dispatch of rendering commands and frame-by-frame execution on user terminals, the control end effectively achieves spatial synchronization and perspective uniformity between remote terminals, ensuring the visual coherence and controllability of remote guidance content. Prioritizing rendering commands and combining them with timestamp-driven loading significantly improves the predictability and real-time nature of rendering behavior.

[0105] The present invention relates to the field of equipment operation and maintenance technology, and can be applied to business scenarios such as financial technology and medical health. A multi-terminal scene synchronization control method is disclosed, including: a control end receives the current virtual scene position and current virtual scene perspective information sent by the user terminal when the virtual scene is started, and generates a synchronous preview screen based on the information; the control end triggers a reverse control mode and sends an operation permission disabling instruction to the user terminal, and the user terminal prohibits the operation input signal acquisition function after receiving the instruction; the control end sets the target virtual scene position and target virtual scene perspective parameters of the user terminal through an interactive interface, generates a rendering execution instruction based on the target parameters, and controls the user terminal to load the instruction to complete the virtual scene rendering. The present invention realizes dynamic guidance and screen control of the user terminal by the control end by setting the target virtual scene position and perspective parameters of the user terminal at the control end, and generates a rendering execution instruction based on the parameters, while cooperating with the operation permission disabling mechanism to avoid perspective conflicts, making the multi-terminal interaction process more consistent and controllable, and improving the screen synchronization and user experience of remote guidance.

[0106] In one embodiment, the above step S10 includes:

[0107] S101, when the user terminal starts the virtual scene, the spatial positioning sensor collects three-dimensional coordinate data of the real physical space;

[0108] S102, the user terminal calls a preset virtual scene coordinate conversion matrix to map the three-dimensional coordinate data of the real physical space into three-dimensional coordinates in a virtual scene coordinate system;

[0109] S103, the user terminal generates a current virtual scene position according to the three-dimensional coordinates in the virtual scene coordinate system;

[0110] S104, the user terminal obtains view angle data through a head posture tracking module, and generates current virtual scene view angle information based on the view angle data;

[0111] S105, the user terminal encapsulates the current virtual scene position and the current virtual scene viewing angle information into a structured data packet, and adds a device identifier of the user terminal and an encapsulation timestamp to the structured data packet;

[0112] S106, the user terminal sends the structured data packet to the backend service through an encrypted communication protocol;

[0113] S107, after verifying the legitimacy of the device identifier, the backend service forwards the structured data packet to the control end;

[0114] S108: The control end parses the structured data packet and extracts the current virtual scene position and current virtual scene viewing angle information.

[0115] In this embodiment, when the user terminal starts the virtual scene, an initialization process is triggered. The process first calls the locally installed spatial positioning sensor system to obtain the three-dimensional coordinate data of the user in the real physical environment. This positioning data is usually achieved through an inertial measurement unit, a depth camera or a laser ranging device, and the collected spatial data is located in the real coordinate system defined locally by the user terminal. In order to enable the coordinate points in the real space to be mapped equivalently in the virtual scene, the user terminal will automatically call the preset virtual scene coordinate conversion matrix to convert the collected three-dimensional coordinates into the corresponding three-dimensional position points in the virtual coordinate system. The conversion matrix can be in the form of a homogeneous coordinate transformation, or it can be a composite matrix structure that combines transformation operations such as rotation, scaling and offset.

[0116] Once the coordinate mapping is complete, the system generates a virtual space position vector representing the user's current position, which is used to describe the user's terminal's positioning status in the virtual environment. Simultaneously, the user's terminal's head posture tracking module continuously collects information about the user's head orientation. This data is typically represented as Euler angles or quaternions. By parsing the current orientation angle data into specific deflection and pitch angles, virtual scene perspective information reflecting the viewing direction can be generated. The user's current viewing perspective, combined with their position in virtual space, forms a complete snapshot of the scene state.

[0117] To facilitate subsequent synchronization, the user terminal integrates the current virtual scene position and perspective information into a structured data packet, appending a device identifier for device verification and an encapsulation timestamp indicating the time the data was generated. This structured data packet is then sent to the backend service system via an encrypted communication protocol. The encryption protocol can be a standard TLS implementation or a lightweight security protocol specifically designed for end-to-end device secure communication. After receiving the structured data packet, the backend service first verifies the legitimacy of the device identifier to ensure the data source is trustworthy. Once verification is successful, the data packet is forwarded intact to the control terminal.

[0118] After receiving the data packet, the control end extracts the user's current virtual scene position and current virtual scene perspective information by parsing the content therein, providing a status basis for subsequent picture rendering or interactive control.

[0119] A positioning head-mounted display (HMD) with an integrated inertial navigation module can be installed on the user terminal, combining it with external fixed reference points or space base stations to obtain high-precision three-dimensional coordinates. The coordinate transformation matrix can be preset to a unit transformation structure consistent with the standard OpenGL or Unity coordinate system, or the developer can adjust it based on the modeling scale of the virtual scene. For head posture tracking, a composite solution that integrates visual SLAM and IMU information can be used to improve the stability of angular data. The device identifier can be a hardware-based unique ID or an authentication token issued by the system, and the timestamp is generated based on the local UTC time synchronization module.

[0120] In terms of communication, dual-channel TLS or QUIC protocol channels can be used to ensure encrypted data transmission with low latency. The backend service sets up an identity verification module to call the database interface to verify the legitimacy of the identifier after receiving it. If it passes the verification, the data is delivered to the control-side processing thread through the message queue system. The control-side parsing module writes the extracted data into the scene state cache table and triggers the corresponding rendering process.

[0121] Error detection logic can also be added to the user terminal. When the positioning data suddenly changes, is missing, or jumps, the correction algorithm will be triggered, and the nearest valid coordinates will be used for linear smooth interpolation in the virtual space to ensure the continuity and reliability of the transmitted data.

[0122] This embodiment performs three-dimensional position acquisition, coordinate system conversion, and perspective tracking on the user terminal, then structures and encapsulates the virtual scene state data before transmitting it to the control terminal. This not only ensures high-fidelity transmission of the scene state, but also lays the foundation for subsequent remote control, screen synchronization, and perspective guidance. In particular, without requiring active synchronization from the user on the control terminal, the system can automatically analyze the current user's position and perspective state, achieving accurate reproduction and perception of the user's current experience state. The introduction of encrypted communication and identifier authentication ensures the security and consistency of remote scene state data interaction.

[0123] In one embodiment, the above step S20 includes:

[0124] S201, the control end parses the three-dimensional coordinate data of the current virtual scene position, and determines the observation point position of the virtual camera in the virtual scene based on the three-dimensional coordinate data;

[0125] S202, analyzing the horizontal deflection angle and the vertical pitch angle of the current virtual scene viewing angle information, and setting the orientation parameters of the virtual camera in the virtual scene based on the horizontal deflection angle and the vertical pitch angle;

[0126] S203, loading a virtual scene 3D model resource file consistent with the user terminal into a graphics rendering pipeline;

[0127] S204, configuring the projection matrix and viewport parameters of the virtual camera according to the observation point position and orientation parameters of the virtual camera;

[0128] S205 , based on the projection matrix and viewport parameters of the virtual camera, calling the rasterization engine of the graphics rendering pipeline to generate a real-time preview image with lighting and shadow effects;

[0129] S206: Output the real-time preview image to the display interface of the control terminal for synchronous display.

[0130] In this embodiment, after obtaining the current virtual scene position and viewpoint information, the control end performs structured parsing and rendering parameter mapping on these two types of data to achieve high-fidelity synchronized scene rendering. First, the control end parses the 3D coordinate data, typically represented as floating-point vectors, as the source of information about the current virtual scene position. This data is used to determine the viewpoint position of the virtual camera in the virtual scene in 3D space. This viewpoint defines the camera's position and serves as the core reference for subsequent view matrix construction.

[0131] At the same time, the control end also extracts the horizontal yaw angle and vertical pitch angle from the current virtual scene perspective information. These two sets of angle data together constitute the camera's orientation parameters. The horizontal yaw angle determines the camera's rotation direction around the vertical axis, while the vertical pitch angle corresponds to the elevation control around the horizontal axis. These two angles are used to construct the camera's direction vector, which is typically used to construct the view matrix in graphics rendering systems.

[0132] To ensure that the preview content is identical to the virtual environment displayed on the user's terminal, the control terminal loads a 3D model resource file of the virtual scene with the same version and texture resources as the user terminal. These resources include static mesh data, texture mapping files, material parameter definition files, and even lighting configuration data for real-time rendering. After loading these resources, the graphics rendering system inputs them into the graphics rendering pipeline.

[0133] Next, the controller configures the virtual camera's projection matrix and viewport parameters based on the previously determined viewpoint position and orientation parameters. The projection matrix can be either perspective or orthographic. Perspective projection is suitable for general scene presentation and provides a realistic sense of depth. The viewport parameters determine the size, aspect ratio, and pixel accuracy of the rendered image output, matching the resolution and output capabilities of the controller's display.

[0134] The control end then invokes the rasterization engine in the graphics rendering pipeline to render the 3D scene data, after model loading and camera configuration, frame by frame. This rendering process combines a pre-set global illumination model and shading algorithm to output a real-time image that incorporates lighting, reflection, and shadow effects. This ensures that the image is not only precisely synchronized in terms of spatial structure, but also closely matches the rendering effect on the user's terminal in terms of light and shadow performance.

[0135] The final preview screen will be displayed in real time on the display interface of the control end through the output module of the control end, allowing the operator to synchronously view the position and observation direction of the user end in the virtual scene, and realize instant grasp and judgment of the user experience status.

[0136] The above rendering process can be implemented by deploying mainstream graphics rendering frameworks such as OpenGL, DirectX, or Vulkan on the control side. The viewpoint position and orientation parameters are used to construct the view matrix, which can be calculated using the LookAt function. The projection matrix is constructed based on the preset field of view angle and the distances to the near and far planes. For resource loading, common 3D model formats such as GLTF or FBX can be used, and a unified resource management module can be used to ensure consistency with the model version used by the user terminal.

[0137] A real-time synchronization module can also be introduced to monitor the model update version number of the user terminal. When a change in the model resource is detected, the 3D resource file on the control end is automatically refreshed to ensure model version synchronization.

[0138] In terms of display devices, the display interface can be configured according to the hardware capabilities of the control end. For example, ultra-high-resolution diagnostic displays can be used for medical teaching, while lightweight multi-window split-screen devices can be used in financial scenario explanations to facilitate parallel monitoring by multiple users.

[0139] Example: During remote surgical instruction in healthcare, the leading physician can use the control terminal to view the current viewing position and viewing direction of the student's device on a virtual mannequin. If the student misses a key surgical area, the physician can quickly identify the deviation through the synchronized preview generated by the control terminal and determine whether real-time corrective action is necessary, improving teaching focus and communication efficiency.

[0140] During financial business training, instructors use the real-time preview generated by the control terminal to determine whether trainees are focusing on the core layers of the financial product structure. If they find that their observation area lingers on the edge for a long time, they can subsequently initiate a counter-control instruction to guide them to quickly switch their perspective and focus on the key content, thereby improving the training rhythm.

[0141] In immersive shopping scenarios, shopping guides can use the control terminal to monitor changes in a customer's viewing angle in real time within the virtual store, and use preview images to determine their level of interest in a particular product. If the user is detected to be moving continuously and not lingering on a high-value item, the control terminal can proactively prepare rendering instructions to guide the customer's viewing angle, enabling precise user gaze management and recommended content push.

[0142] This embodiment, through precise analysis of virtual scene position and viewing angle parameters and complete configuration of the graphics rendering pipeline, enables the control end to synchronously generate the scene image currently seen by the user and display it in real time on the display interface. This synchronization process does not rely on the user end to push frame data, avoiding the experience deviation caused by video stream delay, while reserving more control logic and judgment capabilities on the control end. The synchronous preview mechanism provides a reliable data foundation and perception support for subsequent judgment, guidance, and rendering intervention, enhancing the stability and visual consistency of cross-end interactive control.

[0143] In one embodiment, the above step S30 includes:

[0144] S301, the control end generates an operation permission disabling instruction data packet including a digital signature and a reverse control status flag;

[0145] S302, the control end sends the operation permission disabling instruction data packet to the user terminal through a secure communication channel;

[0146] S303, the user terminal verifies the legitimacy of the digital signature in the operation permission disabling instruction data packet;

[0147] S304, after passing the legitimacy verification, the user terminal calls the device driver interface to disable the operation input signal acquisition circuit to disable the operation input signal acquisition function;

[0148] S305: The user terminal writes the reverse control status flag into the system kernel space and locks the write permission of the memory page where the reverse control status flag is located through the memory management unit;

[0149] S306, the user terminal starts a heartbeat detection thread to continuously monitor the communication connection status with the control terminal;

[0150] S307: If the heartbeat detection thread does not receive a response signal from the control terminal for a preset number of consecutive times, the user terminal clears the reverse control state flag and resumes the operation input signal collection function.

[0151] In this embodiment, when the control end triggers the reverse control mode, a complete control mechanism, encompassing identity authentication, secure communication, device interface calls, and system status locking, is required to ensure that the user terminal's control rights to the scene are promptly and reliably disabled while also avoiding security risks associated with unauthorized triggering. The control end first generates a data packet for permission deprivation. This data packet includes two core fields: an encrypted digital signature, used to identify the authenticity and integrity of the data source; and a reverse control status flag, which serves as the basis for subsequent user terminal status determination and kernel control.

[0152] The data packet is transmitted using a secure communication channel. The communication channel can use encryption protocols such as TLS and DTLS to ensure that it is not hijacked or tampered with by a third party during the transmission process. After the user terminal receives the data packet, it must immediately parse and verify the digital signature. The signature verification uses an asymmetric key system to ensure that only commands issued by the control terminal can be executed. After successful verification, the user terminal calls its underlying device driver interface and disables the input signal acquisition circuit corresponding to the handle, helmet, or other input device. This makes the user's physical operation actions no longer perceived by the system, creating a physical "operation failure" state.

[0153] To prevent the device from being unexpectedly restored or hijacked by malicious programs when in a disabled state, the user terminal will write an anti-control status flag in the system kernel space. The writing of this flag is performed in kernel mode and the memory management unit locks the permissions, setting the memory page where the flag is located to read-only attributes to prevent tampering in user mode or other threads.

[0154] To ensure the system remains in an unresponsive, anti-control mode for extended periods, the user terminal initiates a heartbeat detection thread, which continuously sends communication probe signals to the control terminal and waits for a response. Given that the control terminal may lose connection due to network disconnection, device failure, or other reasons, the user terminal automatically clears the anti-control status flag in the kernel and calls the device driver interface to resume all input signal acquisition, ensuring the system's self-healing and sustainable operation.

[0155] This embodiment constructs a multi-layered mechanism, encompassing data generation, transmission, and verification on the control side, as well as user terminal disabling and state maintenance. This allows the control side to precisely deprive the user of input permissions without interrupting the user's terminal operation. This mechanism, combined with kernel space flag writing and memory page locking, ensures that the system state cannot be tampered with. Furthermore, a heartbeat detection thread is introduced to ensure that the system can promptly restore input functionality when the control connection is interrupted, thereby improving stability and fault recovery capabilities. This overall mechanism implements reliable input permission deprivation control while avoiding issues such as user device deadlock or irreversible disabling.

[0156] In one embodiment, the above step S40 includes:

[0157] S401, obtaining virtual scene space reference parameters sent by the user terminal when the virtual scene is initialized;

[0158] S402, generating a three-dimensional coordinate system visualization grid aligned with the virtual scene space reference parameters on the interactive interface of the control terminal;

[0159] S403, generating a three-dimensional coordinate value of a target virtual scene position in response to a positioning operation of the three-dimensional coordinate system visualization grid;

[0160] S404: Based on the three-dimensional coordinate values of the target virtual scene position and the current virtual scene viewing angle information of the user terminal, a horizontal deflection angle value and a vertical pitch angle value of the target virtual scene viewing angle parameter are generated by a spatial vector analysis module.

[0161] In this embodiment, the control end sets the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through an interactive interface. The core is to synchronize the spatial structure information of the virtual scene in which the user is located with the visual control model of the control end. The virtual scene space reference parameters are used to characterize the spatial structure definition of the virtual scene loaded by the user terminal, including three core fields: the coordinate system origin definition indicates the position of the "zero point" in the virtual space, which serves as the reference starting point for all subsequent position calculations; the axial definition specifies the corresponding directions of the X, Y, and Z axes in three-dimensional space, for example, the X axis represents left and right, the Y axis represents up and down, and the Z axis represents front and back; the unit scale is used to quantify the mapping relationship between the coordinate unit and the actual model size, such as 1 unit corresponds to 1 meter or 10 centimeters.

[0162] After obtaining these parameters, the control terminal constructs a 3D coordinate grid consistent with the user terminal in its interactive interface and performs visual rendering, creating an operating environment with spatial logical consistency. Users can perform positioning operations within this 3D coordinate grid. These operations generate 3D coordinate values for the target virtual scene location by clicking the pointer, entering coordinates, or dragging the mouse. The generated position coordinates are aligned with the spatial reference parameters, ensuring that position calibration does not cause spatial drift due to coordinate system errors.

[0163] The control end also needs to synchronously generate the target virtual scene perspective parameters to specify the adjustment strategy of the user terminal's viewing direction. To this end, the spatial vector analysis module performs vector operations on the target position and the current perspective information. The current virtual scene perspective information refers to the unit direction vector corresponding to the current viewing direction of the user terminal, which is usually obtained in real time through head tracking or posture sensors. The spatial vector analysis module first constructs a direction vector based on the target virtual scene position and the current position of the user terminal, and then calculates the angle between the direction vector and the current perspective direction vector, and further decomposes it into deflection angles around the Y-axis (horizontal) and X-axis (vertical) dimensions, that is, generating horizontal deflection angle values and vertical pitch angle values. This process ensures that the user's perspective adjustment is directional and consistent, and can be used for subsequent smooth transitions or execution control.

[0164] The control terminal supports multiple coordinate systems for defining spatial reference parameters, such as right-handed and left-handed coordinate systems, to accommodate rendering standards of different 3D engines (such as Unity and Unreal). The coordinate system origin can be automatically retrieved from the scene configuration file using the scene identification information sent by the user terminal, or it can be uniformly set by the control terminal during initialization and then transmitted synchronously.

[0165] The visual grid can be rendered using graphical interfaces such as OpenGL or WebGL, supporting dynamic scaling, rotation, and grid spacing adjustment to meet the positioning accuracy requirements of complex models. The control terminal provides multiple positioning operation methods, including screen clicks to generate coordinates, mouse dragging to select the center of gravity coordinates of the area, and inputting coordinate values, enabling flexible interaction.

[0166] In the spatial vector analysis module, the direction vector is calculated using three-dimensional subtraction: the target position coordinates minus the current user terminal coordinates to obtain the heading vector. Angle calculation is performed using the vector dot product and the inverse cosine function, and is converted into Euler angles for output. The controller can set a threshold. If the angle is within a certain range, the view angle is not adjusted and only the position is moved, improving calculation efficiency.

[0167] Example description: In the remote assisted surgery training system in the medical and health business field, the doctor, as the teaching control terminal, obtains the virtual scene space reference parameters of the user terminal used by the student in his console interface. The parameters include the origin definition of the virtual operating room coordinate system (such as the center point of the operating table as the origin), the axial definition (X axis represents left and right, Y axis represents up and down, Z axis represents front and back) and the unit scale (such as each unit represents 1 cm). The interactive interface of the control end automatically generates a three-dimensional coordinate grid based on these parameters and spatially aligns it with the currently loaded virtual scene model. When demonstrating the key steps of vascular suturing, the doctor clicks on a coordinate point in the virtual patient's chest area in the control interface and sets it as the target virtual scene position. The generated coordinate value is (X=15, Y=12, Z=-20). At this time, the system automatically calls the spatial vector analysis module, subtracts the student's current coordinate position (for example, (X=5, Y=10, Z=-10)) from the target position to obtain the direction vector (10, 2, -10), and then performs a dot product calculation with the student's current virtual perspective direction vector (for example, normalized to (0.6, 0.1, -0.8)) and converts it into an angle. The inverse cosine function is used to obtain a horizontal deflection angle of approximately 28.4° and a vertical pitch angle of approximately 5.7°, thereby constructing the target virtual scene perspective parameters.

[0168] In the virtual asset structure explanation system in the fintech business field, the instructor, as the control-end user, discovered that a trainee was not focusing on the key shareholder path when explaining the virtual enterprise asset network diagram. The instructor directly selected the spatial coordinates of the shareholder node (such as (X=40, Y=0, Z=30)) in the interactive interface as the target position. After obtaining the trainee's current coordinates and observation direction, the system uses the spatial vector difference to generate the target deflection parameters, with a horizontal deflection angle of 17° and a vertical pitch angle of -3°. This process does not require manual input of angles, and is completely automatically derived from spatial geometric relationships, significantly improving the efficiency of intervention during the explanation process.

[0169] In immersive retail scenarios, the control end is driven by an AI sales engine, and its behavior is automatically executed based on the user's perspective behavior characteristics. If the user terminal does not focus on a promotional product area (such as the main booth 20 units forward on the Z axis) for 10 consecutive seconds, the system automatically calibrates the center point of the area (such as (X=0, Y=1.5, Z=20)) as the target virtual scene position, analyzes the directional difference between the user's current observation point and the target, and determines the required perspective deflection angle (such as Yaw=12°, Pitch=-2°). The perspective parameters are generated in real time and stored for direct use when the reverse control is triggered.

[0170] This embodiment makes the process of setting the remote target position and viewing angle more intuitive and accurate by constructing a three-dimensional spatial coordinate system consistent with the user terminal on the control end and visually rendering it in the interactive interface. A spatial vector analysis module analyzes the geometric relationship between the user's current position and the target position, automatically calculating the horizontal deflection angle and vertical pitch angle to ensure the continuity and rationality of viewing angle guidance. This process automatically derives the viewing angle adjustment parameters without requiring active user interaction, providing efficient support for subsequent automatic rendering and control end command generation.

[0171] In one embodiment, the above step S50 includes:

[0172] S501, performing interpolation processing on the three-dimensional coordinate values of the target virtual scene position to generate a displacement vector sequence;

[0173] S502, determining a viewing angle deflection step length between adjacent frames according to a horizontal deflection angle value and a vertical pitch angle value of the viewing angle of the target virtual scene;

[0174] S503, based on the current rendering performance parameters of the user terminal, dividing the displacement vector sequence and the viewing angle deflection step into multiple executable frame sequences;

[0175] S504, adding a timestamp and a priority identifier to each executable frame sequence;

[0176] S505: Encapsulate the executable frame sequence after adding the timestamp and priority identifier into a rendering execution instruction packet.

[0177] In this embodiment, the control end generates a rendering execution instruction for controlling the loading of the rendering process by the user terminal based on the target virtual scene position and the target virtual scene viewing angle parameters. First, the three-dimensional coordinate value of the target virtual scene position is generally expressed in the form of (X, Y, Z) in the Euclidean coordinate system. The three-dimensional coordinate value is not directly used for rendering, but requires an interpolation algorithm to generate a continuous displacement path. The interpolation process usually adopts linear interpolation, cubic spline interpolation or Bezier curve interpolation to generate a number of smooth intermediate position points between the target starting point and the end point. These points constitute a displacement vector sequence, and each vector represents the displacement between two consecutive frames.

[0178] The target virtual scene's perspective parameters include horizontal deflection and vertical pitch. The former represents the rotation angle around the vertical axis, while the latter represents the pitch angle around the horizontal axis. During rendering, to prevent motion sickness and image jitter, the overall perspective rotation process needs to be broken down into gradual changes across multiple frames. Therefore, the angle change step size between each frame, known as the perspective deflection step size, needs to be calculated. This calculation typically divides the total deflection angle by the number of frames to generate evenly or weightedly spaced angle increments for subsequent perspective gradients.

[0179] To adapt to the hardware performance of different user terminals, the control end needs to collect the current user terminal's rendering performance parameters, including frame rate cap, GPU utilization, memory remaining, etc. Based on this, the control end segments the generated displacement vector sequence and view deflection step length sequence into multiple executable frame sequences. The time span of each frame sequence matches the rendering complexity to avoid performance bottlenecks.

[0180] To ensure that multiple frame sequences are correctly identified and executed, each executable frame sequence is timestamped to indicate when it should begin execution. A priority identifier is also added to guide the GPU's allocation of computing resources to different types of frame sequences (such as those for interactive or background processing). Timestamps are integers that increment in milliseconds, and priorities can be expressed as multi-level identifiers (such as High, Medium, and Low) or numerical weights.

[0181] Finally, all frame sequences with timestamps and priority identifiers are packaged into a rendering execution instruction package, usually encapsulated in a structured format such as JSON, ProtocolBuffer, or binary protocol. This instruction package will be sent to the user terminal in the subsequent steps to drive its rendering process.

[0182] In one implementation, the control end uses a spline interpolation algorithm to generate a smooth sequence of displacement vectors. Based on the user terminal's current frame rate, the target frame count is set to 60, subdividing the entire position path into 60 consecutive coordinate points. For view deflection, the system captures target deflection angles of 30° horizontally and 10° vertically, automatically calculating deflection steps of 0.5° and 0.167° per frame.

[0183] The control end periodically collects GPU utilization from user terminals. When the load exceeds 85%, each frame sequence is automatically compressed to five frames, and the view deflection step size is dynamically halved to reduce the pressure of each frame rendering. In the resource allocation strategy, frames related to foreground interactions are marked as high priority, while background dynamic special effects frames are marked as low priority. The start timestamp of each frame sequence is assigned in ascending order based on the execution schedule.

[0184] In the instruction packet encapsulation, all frame sequences are merged into a transmission structure in a compressed format, and a check field is added for integrity verification.

[0185] An adaptive step size strategy can also be used, that is, the content volume of each frame sequence is dynamically adjusted after real-time detection of network bandwidth and delay on the control end to balance network transmission pressure and rendering continuity.

[0186] Example: During virtual lesion observation training in the healthcare field, the intern needs to slowly switch their perspective to a specific lesion structure. The control end sets the target virtual scene position to the coordinates of the lesion, and the deflection angle is 15° downward from the current perspective, then 20° to the left. The system uses linear interpolation to generate the displacement path and subdivides the total deflection angle into 45 frames, with a rotation step size of 0.33° per frame. Terminal performance analysis shows that the current device frame rate is 72Hz, and the GPU utilization is within an acceptable range. The system encapsulates the rendering execution instruction package according to the default step size.

[0187] In a multi-asset visualization teaching system for the financial industry, instructors guide users through a virtual investment structure comprised of different financial products. The controller uses a 3D graphical interface to set the target location and automatically generates a path. Under high system load, the controller proactively reduces the rendering granularity of each frame, reducing the number of frames per perspective change from 15 to 8, preventing device lag and maintaining the integrity of core product information.

[0188] In immersive exhibition hall applications for the general public, an AI recommendation system generates perspective optimization strategies based on user behavior trajectories. It automatically adjusts the target location to recommended exhibit areas when a user's viewing path stalls or their interest drifts. It also guides user perspective switching through optimized execution instruction packages, enhancing guidance and commercial conversion efficiency. Throughout this process, timestamps precisely control the rhythm of each switching segment, and a priority mechanism ensures exclusive computing resources for relevant exhibit views, thereby improving the overall performance of exhibition hall interactions.

[0189] This embodiment converts target position and viewing angle parameters into a high-resolution continuous frame sequence and performs frame-level optimization based on terminal performance parameters. This allows for efficient transmission and parsing of rendering commands while ensuring continuity and comfort. In particular, by implementing a priority and timestamp management mechanism, the control terminal can dynamically schedule different frame sequences, ensuring that critical interactive commands always receive priority while non-critical rendering tasks are deferred in the background, significantly improving the overall system's perceptual performance and user immersion.

[0190] In one embodiment, the above step S60 includes:

[0191] S601, the control end sends the rendering execution instruction to the user terminal;

[0192] S602, the user terminal receives the rendering execution instruction, and extracts an executable frame sequence from the rendering execution instruction;

[0193] S603: The user terminal sends a resource allocation instruction to a graphics processor according to the priority identifier in the executable frame sequence, and the graphics processor allocates graphics processor resources to executable frame sequences of different priorities according to the allocation instruction;

[0194] S604: The user terminal inputs the executable frame sequence into a graphics rendering pipeline according to the timestamp order of the executable frame sequence;

[0195] S605, the graphics rendering pipeline parses the displacement vector sequence and the perspective deflection step from the executable frame sequence, updates the position of the virtual camera according to the displacement vector sequence, adjusts the orientation of the virtual camera according to the perspective deflection step, completes the rendering of the current virtual scene, and outputs the rendering result of the current frame to the display module of the user terminal.

[0196] In this embodiment, the process of the control end controlling the user terminal to load the rendering execution instruction to render the current virtual scene first involves the control end sending the rendering execution instruction to the user terminal. The instruction contains a rendering data set consisting of multiple executable frame sequences. These frame sequences have been generated and packaged based on the target virtual scene position and viewpoint parameters in the previous step, and are accompanied by a timestamp and priority identifier.

[0197] After receiving the rendering execution command, the user terminal parses the encapsulated data packet and extracts the executable frame sequence. Each frame sequence represents a continuous rendering task, usually including the position and orientation change information of the target virtual camera.

[0198] After extraction is complete, the user terminal schedules the executable frame sequences based on their priority identifiers and issues resource allocation instructions to the graphics processing unit (GPU). The GPU then allocates computing resources accordingly. High-priority frame sequences receive exclusive GPU computing unit resources to ensure the stability and continuity of key scenes like interaction and perspective changes. Low-priority frame sequences, on the other hand, use the GPU's remaining idle resources through a sharing mechanism, ensuring core rendering tasks while maintaining overall system load balancing.

[0199] The user terminal then inputs each frame sequence into the graphics rendering pipeline in the order of the timestamps attached to it. Timestamps control the loading rhythm of each frame, avoiding abrupt perspective changes or lags. The graphics rendering pipeline, the main path for image processing, extracts a sequence of displacement vectors and perspective deflection steps from the frame sequence. The former is used to update the position of the virtual camera, while the latter is used to adjust the camera's orientation frame by frame. The displacement vectors and deflection steps are executed synchronously at the frame granularity to ensure that the virtual scene changes seen by the user are spatially continuous and motionally reasonable.

[0200] The rendered image is then output to the user's display module, typically a head-mounted display (HMD) or stereoscopic display. Within this module, the image undergoes processing such as parallax synthesis and distortion correction before being presented to the user, completing the entire process from rendering instructions to visual experience.

[0201] In one implementation, the control end transmits rendering execution instruction packets to the user terminal in real time using a TCP channel or the WebSocket protocol. Upon receiving the instruction packets, the user terminal parses the JSON format or binary structured instructions, extracts five executable frame sequences, each containing 30 frames, and identifies their priority identifiers as 3 (high), 2 (medium), and 1 (low).

[0202] The user terminal uses the CUDA interface to allocate resources to the GPU. Level 3 frames are bound to exclusive execution channels based on frame sequence priority, while level 2 and level 1 frames are loaded in batches through a cyclic shared thread queue. When GPU utilization exceeds a preset threshold (e.g., 85%), the loading frequency of low-priority frame sequences is dynamically reduced.

[0203] Before being input into the graphics rendering pipeline, frames are sorted with millisecond precision. The system uses a timestamp scheduling module to control the loading interval of each frame sequence to 16.6ms (corresponding to 60 frames per second). The graphics rendering pipeline uses OpenGL or Vulkan for frame-by-frame rendering. For each frame, the displacement vector is applied to the camera transformation matrix, and the view matrix is adjusted based on the deflection step size. The image calculation is completed in real time and transmitted to the display module.

[0204] It is also possible to pre-calculate the view deflection as a quaternion difference path in the enhanced rendering efficiency scenario, and use the GPU local cache mechanism to load part of the frame sequence in advance to reduce latency.

[0205] The eye tracking module can also be combined to optimize the rendering pipeline scheduling logic, allocating high-priority resources only to the area where the user is currently looking, and simplifying the rendering processing of other areas.

[0206] Example: In a remote surgery teaching platform, the instructing surgeon receives the current virtual scene position and viewpoint information from each student's user terminal on the control interface. Using the coordinate and posture analysis module, the control interface synchronizes and generates a virtual observation screen consistent with the student's viewpoint in real time, allowing the surgeon to understand their current focus. If the control interface detects that the student's viewpoint deviates from a critical anatomical structure (such as a heart valve or nerve bundle), it automatically triggers reverse control mode and sends a command to the corresponding user terminal to disable operation permissions. After legal verification, the user terminal disables the input signal acquisition function, blocks input sources such as head rotation and gestures, and writes a reverse control status flag locally. The control interface then sets the target virtual scene position to the "right atrium inner wall" and the target virtual scene viewpoint parameters to precisely align the student's viewpoint with the atrial valve. The spatial vector analysis module calculates the rotation angle based on the user's current position information and generates viewpoint deflection parameters. The control end generates a rendering execution instruction based on the target parameter, which includes a displacement vector sequence generated by continuous interpolation and a perspective deflection step size for frame-by-frame rotation. Combined with the GPU load and frame rate of the current user terminal, it is divided into high- and low-priority frame segments and timestamps are added. After the rendering execution instruction is sent to the user terminal, the user terminal inputs the executable frame sequence into the rendering pipeline. High-priority frames (such as perspective guide frames close to the target position) monopolize graphics processor resources, while low-priority frames (such as background rotation transitions) share idle threads. The displacement vector is used to update the position of the virtual camera, and the perspective deflection step size finely adjusts the camera direction. Finally, the image aligned with the heart valve area is output to the display module in real time, guiding students to efficiently complete the identification of key parts and learning of operation paths.

[0207] During a virtualized remote presentation on investment products for a high-net-worth client, the presenter, receiving information about the client's current virtual scene location and viewpoint from the client's user terminal, discovered that the client was not focusing on the core investment product segment within a complex asset sandbox. The presenter automatically or manually triggered anti-control mode and sent a permission-disable command to the client terminal. Upon successful verification, the client's head tracking and telemetry input were suspended, preventing the client from further deviating from the presentation logic. The presenter then selected the "Fixed Income + Equity-Linked Structured Products" segment within the presenter's interactive interface as the target virtual scene location. The system calculated the required deflection path based on the client's previous viewpoint information, generating new target virtual scene viewpoint parameters. The presenter converted this target scene and parameters into rendering instructions, which included an interpolated sequence of displacement vectors and a rotation path. The viewpoint deflection step size was dynamically adapted and segmented into three high-priority content frames to quickly direct the viewer's attention to the investment charts and backtest plots. The presenter then sent the complete rendering instructions to the client terminal, which parsed and allocated resources based on priority, with high-priority frames executed exclusively on the GPU. The virtual camera loads rendering instructions frame by frame on the user's terminal, automatically adjusting the viewing angle and quickly focusing the client's view on the target asset area, showcasing key yield points and liquidity curves, enhancing client understanding and investment decision-making confidence. The final rendered image is presented in real time via the user's head-mounted display, significantly enhancing the immersiveness and pacing of visual explanations of financial products.

[0208] This embodiment constructs a multi-level frame sequence priority mechanism and combines it with a timestamp-driven rendering loading order to ensure the rapid rendering of high-priority content while reasonably compressing the rendering cost of low-priority content, effectively avoiding competition for graphics processor resources. The synchronous execution of displacement vectors and viewing angle steps ensures the physical continuity of the scene transformation process. Combined with the optimized scheduling mechanism of the graphics rendering pipeline, it can significantly reduce rendering delays and the risk of freezes. The overall system achieves high adaptability and strong control over the rendering performance of user terminals while maintaining the realism of the scene, significantly enhancing the immersive interactive experience.

[0209] In one embodiment, a multi-terminal scene synchronization control device is provided, which corresponds one-to-one to the multi-terminal scene synchronization control method in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of the multi-terminal scene synchronization control device of the present invention. It includes a virtual scene perception module 10, a synchronous preview rendering module 20, a permission control module 30, a target parameter setting module 40, a command generation module 50, and a remote rendering control module 60. Each functional module is described in detail below:

[0210] The virtual scene perception module 10 is used for the control terminal to receive the current virtual scene position and current virtual scene viewing angle information sent by the user terminal when the virtual scene is started;

[0211] A synchronous preview rendering module 20 is used for the control terminal to generate a synchronous preview image based on the received current virtual scene position and current virtual scene viewing angle information;

[0212] The authority control module 30 is used to control the terminal to trigger the reverse control mode and send an operation authority disabling instruction to the user terminal, wherein the operation authority disabling instruction is used to disable the operation input signal collection function of the user terminal;

[0213] A target parameter setting module 40 is used for the control end to set the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through an interactive interface;

[0214] The instruction generation module 50 is used for the control terminal to generate a rendering execution instruction based on the target virtual scene position and the target virtual scene viewing angle parameters;

[0215] The remote rendering control module 60 is used to control the user terminal to load the rendering execution instruction to render the current virtual scene.

[0216] In one embodiment, the virtual scene perception module 10 is specifically configured to:

[0217] When the user terminal starts the virtual scene, the spatial positioning sensor collects the three-dimensional coordinate data of the real physical space;

[0218] The user terminal calls a preset virtual scene coordinate conversion matrix to map the three-dimensional coordinate data of the real physical space into three-dimensional coordinates in the virtual scene coordinate system;

[0219] The user terminal generates a current virtual scene position according to the three-dimensional coordinates in the virtual scene coordinate system;

[0220] The user terminal obtains viewing angle data through a head posture tracking module, and generates current virtual scene viewing angle information based on the viewing angle data;

[0221] The user terminal encapsulates the current virtual scene position and the current virtual scene viewing angle information into a structured data packet, and adds a device identifier of the user terminal and an encapsulation timestamp to the structured data packet;

[0222] The user terminal sends the structured data packet to the backend service via an encrypted communication protocol;

[0223] After verifying the legitimacy of the device identifier, the backend service forwards the structured data packet to the control end;

[0224] The control end parses the structured data packet and extracts the current virtual scene position and current virtual scene viewing angle information.

[0225] In one embodiment, the synchronous preview rendering module 20 is specifically configured to:

[0226] The control terminal analyzes the three-dimensional coordinate data of the current virtual scene position, and determines the observation point position of the virtual camera in the virtual scene based on the three-dimensional coordinate data;

[0227] Analyzing the horizontal deflection angle and the vertical pitch angle of the current virtual scene viewing angle information, and setting the orientation parameters of the virtual camera in the virtual scene based on the horizontal deflection angle and the vertical pitch angle;

[0228] Loading a virtual scene three-dimensional model resource file consistent with the user terminal into a graphics rendering pipeline;

[0229] Configuring the projection matrix and viewport parameters of the virtual camera according to the observation point position and orientation parameters of the virtual camera;

[0230] Based on the projection matrix and viewport parameters of the virtual camera, calling the rasterization engine of the graphics rendering pipeline to generate a real-time preview image with lighting and shadow effects;

[0231] The real-time preview image is output to the display interface of the control terminal for synchronous display.

[0232] In one embodiment, the authority control module 30 is specifically configured to:

[0233] The control end generates an operation permission disabling instruction data packet including a digital signature and an anti-control status flag;

[0234] The control terminal sends the operation permission disabling instruction data packet to the user terminal through a secure communication channel;

[0235] The user terminal verifies the legitimacy of the digital signature in the operation permission disabling instruction data packet;

[0236] After the legitimacy verification is passed, the user terminal calls the device driver interface to disable the operation input signal acquisition circuit to prohibit the operation input signal acquisition function;

[0237] The user terminal writes the reverse control status flag into the system kernel space, and locks the write permission of the memory page where the reverse control status flag is located through the memory management unit;

[0238] The user terminal starts a heartbeat detection thread to continuously monitor the communication connection status with the control terminal;

[0239] If the heartbeat detection thread does not receive a response signal from the control terminal for a preset number of consecutive times, the user terminal clears the reverse control state flag and resumes the operation input signal collection function.

[0240] In one embodiment, the target parameter setting module 40 is specifically configured to:

[0241] Obtaining virtual scene space reference parameters sent by the user terminal when the virtual scene is initialized;

[0242] Generating a three-dimensional coordinate system visualization grid aligned with the virtual scene space reference parameters on an interactive interface of the control terminal;

[0243] In response to the positioning operation of the three-dimensional coordinate system visualization grid, generating a three-dimensional coordinate value of a target virtual scene position;

[0244] Based on the three-dimensional coordinate values of the target virtual scene position and the current virtual scene viewing angle information of the user terminal, a horizontal deflection angle value and a vertical pitch angle value of the target virtual scene viewing angle parameter are generated by a spatial vector analysis module.

[0245] In one embodiment, the instruction generation module 50 is specifically configured to:

[0246] Performing interpolation processing on the three-dimensional coordinate values of the target virtual scene position to generate a displacement vector sequence;

[0247] Determining a viewing angle deflection step length between adjacent frames according to a horizontal deflection angle value and a vertical pitch angle value of the viewing angle of the target virtual scene;

[0248] Based on current rendering performance parameters of the user terminal, dividing the displacement vector sequence and the viewing angle deflection step into multiple executable frame sequences;

[0249] Adding a timestamp and a priority identifier to each executable frame sequence;

[0250] The executable frame sequence after adding the timestamp and priority identifier is encapsulated into a rendering execution instruction packet.

[0251] In one embodiment, the remote rendering control module 60 is specifically configured to:

[0252] The control end sends the rendering execution instruction to the user terminal;

[0253] The user terminal receives the rendering execution instruction and extracts an executable frame sequence from the rendering execution instruction;

[0254] The user terminal sends a resource allocation instruction to a graphics processor according to a priority identifier in the executable frame sequence, and the graphics processor allocates graphics processor resources to executable frame sequences of different priorities according to the allocation instruction;

[0255] The user terminal inputs the executable frame sequence into a graphics rendering pipeline according to a timestamp order of the executable frame sequence;

[0256] The graphics rendering pipeline parses the displacement vector sequence and the perspective deflection step from the executable frame sequence, updates the position of the virtual camera according to the displacement vector sequence, adjusts the orientation of the virtual camera according to the perspective deflection step, completes the rendering of the current virtual scene, and outputs the rendering result of the current frame to the display module of the user terminal.

[0257] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a multi-terminal scene synchronization control method.

[0258] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a multi-terminal scene synchronization control method.

[0259] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0260] The control terminal receives the current virtual scene position and current virtual scene viewing angle information sent by the user terminal when the virtual scene is started;

[0261] The control terminal generates a synchronous preview image based on the received current virtual scene position and current virtual scene viewing angle information;

[0262] The control end triggers the reverse control mode and sends an operation permission disabling instruction to the user terminal, wherein the operation permission disabling instruction is used to disable the operation input signal collection function of the user terminal;

[0263] The control end sets the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through the interactive interface;

[0264] The control end generates a rendering execution instruction based on the target virtual scene position and the target virtual scene viewing angle parameters;

[0265] The control end controls the user terminal to load the rendering execution instruction to render the current virtual scene.

[0266] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0267] The control terminal receives the current virtual scene position and current virtual scene viewing angle information sent by the user terminal when the virtual scene is started;

[0268] The control terminal generates a synchronous preview image based on the received current virtual scene position and current virtual scene viewing angle information;

[0269] The control end triggers the reverse control mode and sends an operation permission disabling instruction to the user terminal, wherein the operation permission disabling instruction is used to disable the operation input signal collection function of the user terminal;

[0270] The control end sets the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through the interactive interface;

[0271] The control end generates a rendering execution instruction based on the target virtual scene position and the target virtual scene viewing angle parameters;

[0272] The control end controls the user terminal to load the rendering execution instruction to render the current virtual scene.

[0273] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0274] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0275] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0276] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A multi-terminal scene synchronization control method, characterized in that: The following steps are involved: The control terminal receives the current virtual scene position and current virtual scene viewing angle information sent by the user terminal when the virtual scene is started; The control terminal generates a synchronous preview image based on the received current virtual scene position and current virtual scene viewing angle information; The control end triggers the reverse control mode and sends an operation permission disabling instruction to the user terminal, wherein the operation permission disabling instruction is used to disable the operation input signal collection function of the user terminal; The control end sets the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through the interactive interface; The control end generates a rendering execution instruction based on the target virtual scene position and the target virtual scene viewing angle parameters; The control end controls the user terminal to load the rendering execution instruction to render the current virtual scene.

2. The multi-terminal scene synchronization control method according to claim 1, characterized in that: The control terminal receives the current virtual scene position and current virtual scene viewing angle information sent by the user terminal when the virtual scene is started, including: When the user terminal starts the virtual scene, the spatial positioning sensor collects the three-dimensional coordinate data of the real physical space; The user terminal calls a preset virtual scene coordinate conversion matrix to map the three-dimensional coordinate data of the real physical space into three-dimensional coordinates in the virtual scene coordinate system; The user terminal generates a current virtual scene position according to the three-dimensional coordinates in the virtual scene coordinate system; The user terminal obtains viewing angle data through a head posture tracking module, and generates current virtual scene viewing angle information based on the viewing angle data; The user terminal encapsulates the current virtual scene position and the current virtual scene viewing angle information into a structured data packet, and adds a device identifier of the user terminal and an encapsulation timestamp to the structured data packet; The user terminal sends the structured data packet to the backend service via an encrypted communication protocol; After verifying the legitimacy of the device identifier, the backend service forwards the structured data packet to the control end; The control end parses the structured data packet and extracts the current virtual scene position and current virtual scene viewing angle information.

3. The multi-terminal scene synchronization control method according to claim 1, characterized in that: The control terminal generates a synchronous preview screen based on the received current virtual scene position and current virtual scene perspective information, including: The control terminal analyzes the three-dimensional coordinate data of the current virtual scene position, and determines the observation point position of the virtual camera in the virtual scene based on the three-dimensional coordinate data; Analyzing the horizontal deflection angle and the vertical pitch angle of the current virtual scene viewing angle information, and setting the orientation parameters of the virtual camera in the virtual scene based on the horizontal deflection angle and the vertical pitch angle; Loading a virtual scene three-dimensional model resource file consistent with the user terminal into a graphics rendering pipeline; Configuring the projection matrix and viewport parameters of the virtual camera according to the observation point position and orientation parameters of the virtual camera; Based on the projection matrix and viewport parameters of the virtual camera, calling the rasterization engine of the graphics rendering pipeline to generate a real-time preview image with lighting and shadow effects; The real-time preview image is output to the display interface of the control terminal for synchronous display.

4. The multi-terminal scene synchronization control method according to claim 1, characterized in that: The control end triggers the reverse control mode and sends an operation permission disabling instruction to the user terminal, where the operation permission disabling instruction is used to disable the operation input signal collection function of the user terminal, including: The control end generates an operation permission disabling instruction data packet including a digital signature and an anti-control status flag; The control terminal sends the operation permission disabling instruction data packet to the user terminal through a secure communication channel; The user terminal verifies the legitimacy of the digital signature in the operation permission disabling instruction data packet; After the legitimacy verification is passed, the user terminal calls the device driver interface to disable the operation input signal acquisition circuit to prohibit the operation input signal acquisition function; The user terminal writes the reverse control status flag into the system kernel space, and locks the write permission of the memory page where the reverse control status flag is located through the memory management unit; The user terminal starts a heartbeat detection thread to continuously monitor the communication connection status with the control terminal; If the heartbeat detection thread does not receive a response signal from the control terminal for a preset number of consecutive times, the user terminal clears the reverse control state flag and resumes the operation input signal collection function.

5. The multi-terminal scene synchronization control method according to claim 1, characterized in that: The control end sets the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through the interactive interface, including: Obtaining virtual scene space reference parameters sent by the user terminal when the virtual scene is initialized; Generating a three-dimensional coordinate system visualization grid aligned with the virtual scene space reference parameters on an interactive interface of the control terminal; In response to the positioning operation of the three-dimensional coordinate system visualization grid, generating a three-dimensional coordinate value of a target virtual scene position; Based on the three-dimensional coordinate values of the target virtual scene position and the current virtual scene viewing angle information of the user terminal, a horizontal deflection angle value and a vertical pitch angle value of the target virtual scene viewing angle parameter are generated by a spatial vector analysis module.

6. The multi-terminal scene synchronization control method according to claim 1, characterized in that: The control end generates a rendering execution instruction based on the target virtual scene position and the target virtual scene viewing angle parameters, including: Performing interpolation processing on the three-dimensional coordinate values of the target virtual scene position to generate a displacement vector sequence; Determining a viewing angle deflection step length between adjacent frames according to a horizontal deflection angle value and a vertical pitch angle value of the viewing angle of the target virtual scene; Based on current rendering performance parameters of the user terminal, dividing the displacement vector sequence and the viewing angle deflection step into multiple executable frame sequences; Adding a timestamp and a priority identifier to each executable frame sequence; The executable frame sequence after adding the timestamp and priority identifier is encapsulated into a rendering execution instruction packet.

7. The multi-terminal scene synchronization control method according to claim 1, characterized in that: The control end controls the user terminal to load the rendering execution instruction to render the current virtual scene, including: The control end sends the rendering execution instruction to the user terminal; The user terminal receives the rendering execution instruction and extracts an executable frame sequence from the rendering execution instruction; The user terminal sends a resource allocation instruction to a graphics processor according to a priority identifier in the executable frame sequence, and the graphics processor allocates graphics processor resources to executable frame sequences of different priorities according to the allocation instruction; The user terminal inputs the executable frame sequence into a graphics rendering pipeline according to a timestamp order of the executable frame sequence; The graphics rendering pipeline parses the displacement vector sequence and the perspective deflection step from the executable frame sequence, updates the position of the virtual camera according to the displacement vector sequence, adjusts the orientation of the virtual camera according to the perspective deflection step, completes the rendering of the current virtual scene, and outputs the rendering result of the current frame to the display module of the user terminal.

8. A multi-terminal scene synchronization control device, characterized in that: The multi-terminal scene synchronization control device includes: The virtual scene perception module is used for the control end to receive the current virtual scene position and current virtual scene viewing angle information sent by the user terminal when the virtual scene is started; A synchronous preview rendering module is used for the control end to generate a synchronous preview image based on the received current virtual scene position and current virtual scene viewing angle information; The authority control module is used to control the terminal to trigger the reverse control mode and send an operation authority disabling instruction to the user terminal, wherein the operation authority disabling instruction is used to disable the operation input signal collection function of the user terminal; A target parameter setting module is used for the control end to set the target virtual scene position and target virtual scene viewing angle parameters of the user terminal through an interactive interface; An instruction generation module is used for the control terminal to generate a rendering execution instruction based on the target virtual scene position and the target virtual scene viewing angle parameters; The remote rendering control module is used to control the user terminal to load the rendering execution instruction to render the current virtual scene.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a multi-terminal scene synchronization control program stored in the memory and capable of running on the processor. When the multi-terminal scene synchronization control program is executed by the processor, the steps of the multi-terminal scene synchronization control method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a multi-terminal scene synchronization control program, which, when executed by the processor, implements the steps of the multi-terminal scene synchronization control method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Road construction comprehensive management and control method and system based on Internet of Things

    CN120746497A

  • Auxiliary screen driving method and system based on virtual display architecture

    CN121277455A

  • Emergency equipment intelligent sensing system and method thereof

    CN121619342A

  • Toy remote interaction delay alignment method and system based on software development kit

    CN121645563A

  • Software development kit-based toy remote interaction delay alignment method and system

    CN121645563B