VR platform control method and system for synchronous interaction of multiple education terminals

By receiving user posture data and operation commands from educational terminals, and performing spatiotemporal consistency verification and dynamic weight allocation, the problem of asynchronous interaction in the synchronous control of multiple VR educational terminals is solved, improving the synchronicity and consistency of VR collaborative training, and enhancing immersion and teaching effectiveness.

CN122018701APending Publication Date: 2026-05-12BEIJING CHINESE EDUCATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING CHINESE EDUCATION TECH CO LTD
Filing Date
2026-04-10
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing VR multi-education terminal synchronous control methods are difficult to effectively coordinate the asynchronous interaction problems caused by network latency, differences in device performance, and inconsistent operation timing, which affects the sense of presence and teaching effectiveness of collaborative training.

Method used

By receiving user posture data and operation commands from the educational terminal, an initial set of interactive commands is generated, spatiotemporal consistency is checked, a synchronous interactive feature matrix is ​​constructed, and a multi-objective collaborative optimization algorithm is used for dynamic weight allocation and delay optimization to generate a collaborative control strategy and uniformly schedule the rendering content and interactive feedback timing of the terminal device.

Benefits of technology

It improves the interactivity and operational consistency of multiple educational terminals in VR collaborative training, enhancing the immersion and effectiveness of education and training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018701A_ABST
    Figure CN122018701A_ABST
Patent Text Reader

Abstract

The invention discloses a VR platform control method and system for synchronous interaction of multiple education terminals, and the method comprises the steps: receiving user posture data and operation instructions uploaded by multiple education terminal devices in real time, and generating an initial interaction instruction set of each education terminal based on the posture data and the operation instructions; performing space-time consistency verification on the initial interaction instruction set, constructing a synchronous interaction feature matrix, and decomposing the synchronous interaction feature matrix into a spatial position feature vector and a time sequence feature vector; performing dynamic weight distribution and delay optimization on the spatial position feature vector and the time sequence feature vector by using a multi-objective collaborative optimization algorithm to generate an optimized collaborative control strategy; and generating a multi-education-terminal synchronization instruction based on the cooperative control strategy, and uniformly scheduling the rendering content and the interaction feedback time sequence of each education terminal device. According to the embodiment of the invention, the interaction synchronism and operation consistency of multiple education terminals in VR collaborative training can be improved, and the immersion of education training and the actual teaching effect are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of VR interactive technology, and in particular to a VR platform control method and system for synchronous interaction of multiple educational terminals. Background Technology

[0002] Currently, virtual reality (VR) technology has been gradually applied in the education and training field, providing immersive and scenario-based learning experiences. In multi-person collaborative training scenarios, it is often necessary to connect multiple VR educational terminal devices to the same virtual environment to achieve synchronous interaction and unified teaching guidance. However, existing VR multi-education terminal synchronous control methods typically employ a centralized command distribution mechanism, which struggles to effectively coordinate the asynchronous interaction issues caused by network latency, differences in device performance, and inconsistent operation sequences among different educational terminal users. This easily leads to misaligned collaborative actions and delayed feedback responses in the virtual scene, severely impacting the sense of presence and teaching effectiveness of collaborative training. Especially in scenarios such as teaching demonstrations and team practical training, where high consistency of operation and real-time feedback are required, existing systems lack the ability to uniformly model and dynamically optimize the spatiotemporal interaction characteristics of multiple educational terminals, resulting in low collaborative efficiency and difficulty in supporting standardized training applications with high immersion and strong interaction. Summary of the Invention

[0003] The purpose of this invention is to provide a VR platform control method and system for synchronous interaction of multiple educational terminals, so as to overcome the shortcomings of the prior art, improve the interactive synchronization and operational consistency of multiple educational terminals in VR collaborative training, and enhance the immersion and teaching effectiveness of education and training.

[0004] One embodiment of this application provides a VR platform control method for synchronous interaction of multiple educational terminals, the method comprising: Receive user posture data and operation instructions uploaded in real time from multiple educational terminal devices, and generate an initial set of interaction instructions for each educational terminal based on the posture data and operation instructions; Spatiotemporal consistency verification is performed on the initial set of interactive instructions, a synchronous interactive feature matrix is ​​constructed, and the synchronous interactive feature matrix is ​​decomposed into spatial location feature vectors and time series feature vectors to capture the spatiotemporal characteristics of multi-education terminal interaction. A multi-objective collaborative optimization algorithm is used to dynamically assign weights and optimize delays to the spatial location feature vector and the time series feature vector to generate an optimized collaborative control strategy. Based on the aforementioned collaborative control strategy, multi-education terminal synchronization instructions are generated to uniformly schedule the rendering content and interactive feedback timing of each education terminal device.

[0005] Optionally, the step of receiving user posture data and operation commands uploaded in real time from multiple educational terminal devices, and generating an initial set of interaction commands for each educational terminal based on the posture data and operation commands, includes: The system receives user posture data and operation instructions uploaded by multiple educational terminal devices through a wireless communication module. The posture data includes head position, hand movements, and body orientation, while the operation instructions include button clicks and gesture recognition results, generating a raw data stream. The raw data stream is subjected to noise filtering and outlier removal, while the operation commands are standardized to ensure data consistency and integrity, generating a preprocessed attitude data stream and operation command stream. Based on the preprocessed posture data stream and operation command stream, key interaction features, including user intent labels, action trajectories and operation timestamps, are extracted to generate an interaction feature vector set. Based on the set of interactive feature vectors, the features are converted into specific interactive instructions through an instruction mapping algorithm, generating an initial set of interactive instructions for each educational terminal.

[0006] Optionally, the step of performing spatiotemporal consistency verification on the initial set of interactive instructions, constructing a synchronous interactive feature matrix, and decomposing the synchronous interactive feature matrix into spatial location feature vectors and time series feature vectors to capture the spatiotemporal characteristics of multi-education terminal interaction includes: The initial set of interactive instructions is checked for spatiotemporal consistency. The timestamp differences and spatial coordinate deviations of the instructions from each educational terminal are calculated. Inconsistent instructions are identified and marked, and a consistency check report is generated. Based on the consistency verification report, the initial set of interactive instructions is organized according to time sequence and spatial location to construct a synchronous interactive feature matrix. In the synchronous interactive feature matrix, the rows represent educational terminal devices, the columns represent time points, and the elements represent spatial coordinates and operation types. Singular value decomposition is performed on the synchronous interaction feature matrix to extract the main spatial patterns and temporal patterns, generating spatial location feature vectors and time series feature vectors. The spatial location feature vector and the time series feature vector are normalized to ensure consistent feature scale, thereby generating standardized spatial location feature vector and time series feature vector.

[0007] Optionally, the step of using a multi-objective collaborative optimization algorithm to dynamically assign weights and optimize delays in the spatial location feature vector and the time series feature vector to generate an optimized collaborative control strategy includes: Define the objective function of the multi-objective collaborative optimization algorithm, with objectives including minimizing interaction latency, maximizing spatial synchronization, and minimizing energy consumption, and generate an optimized objective configuration. Based on the optimized target configuration, dynamic weight allocation is performed on spatial location feature vectors and time series feature vectors. The weight coefficients are adjusted according to real-time network conditions and device performance to generate dynamic weight vectors. By using dynamic weight vectors, delay optimization is performed on time series feature vectors, and network latency is compensated by prediction algorithms to generate optimized time series feature vectors. By combining spatial location feature vectors and optimized time series feature vectors, a coordinated control strategy, including instruction scheduling order and resource allocation scheme, is output through a policy generator.

[0008] Optionally, the step of generating multi-education terminal synchronization instructions based on the collaborative control strategy, and uniformly scheduling the rendering content and interaction feedback timing of each education terminal device, includes: Analyze the collaborative control strategy, extract the instruction scheduling sequence and resource allocation scheme, and generate a multi-education terminal synchronous instruction template; Based on the multi-education terminal synchronization instruction template, specific rendering instructions and interactive feedback instructions are generated for each education terminal device, thus generating a specific instruction set for the education terminal. Based on the specific instruction set of the educational terminal, the rendering content and interactive feedback timing of each educational terminal device are uniformly scheduled through a timing coordinator to ensure that all educational terminals perform corresponding operations at the same time and generate a synchronous scheduling plan. Execute the synchronization scheduling plan, send synchronization instructions to each educational terminal device, monitor the execution status, and output the synchronization interaction results of multiple educational terminals.

[0009] Another embodiment of this application provides a VR platform control system for synchronous interaction of multiple educational terminals, the system comprising: The receiving module is used to receive user posture data and operation instructions uploaded in real time by multiple educational terminal devices, and generate an initial set of interaction instructions for each educational terminal based on the posture data and operation instructions. The module is used to perform spatiotemporal consistency verification on the initial set of interactive instructions, construct a synchronous interactive feature matrix, and decompose the synchronous interactive feature matrix into spatial location feature vectors and time series feature vectors to capture the spatiotemporal characteristics of multi-education terminal interaction. The optimization module is used to dynamically allocate weights and optimize delays for the spatial location feature vector and the time series feature vector using a multi-objective collaborative optimization algorithm, thereby generating an optimized collaborative control strategy. The generation module is used to generate synchronization instructions for multiple educational terminals based on the collaborative control strategy, and to uniformly schedule the rendering content and interactive feedback timing of each educational terminal device.

[0010] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.

[0011] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.

[0012] Compared with existing technologies, this invention provides a VR platform control method for synchronous interaction of multiple educational terminals. It receives user posture data and operation commands uploaded in real time from multiple educational terminal devices, generates an initial set of interaction commands for each educational terminal based on the posture data and operation commands, performs spatiotemporal consistency verification on the initial set of interaction commands, constructs a synchronous interaction feature matrix, and decomposes the synchronous interaction feature matrix into spatial position feature vectors and temporal sequence feature vectors. A multi-objective collaborative optimization algorithm is used to dynamically allocate weights and optimize delays for the spatial position feature vectors and temporal sequence feature vectors, generating an optimized collaborative control strategy. Based on the collaborative control strategy, synchronous commands for multiple educational terminals are generated, and the rendering content and interaction feedback timing of each educational terminal device are uniformly scheduled. This improves the interactive synchronization and operational consistency of multiple educational terminals in VR collaborative training, enhancing the immersion and teaching effectiveness of the training. Attached Figure Description

[0013] Figure 1 A hardware structure block diagram of a computer education terminal for a VR platform control method for synchronous interaction of multiple education terminals provided in an embodiment of the present invention; Figure 2 A flowchart illustrating a VR platform control method for synchronous interaction of multiple educational terminals provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a VR platform control system for synchronous interaction of multiple educational terminals provided in an embodiment of the present invention. Detailed Implementation

[0014] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0015] This invention first provides a VR platform control method for synchronous interaction of multiple educational terminals. This method can be applied to electronic devices, such as computer educational terminals, specifically ordinary computers.

[0016] The following is a detailed explanation using an example running on a computer education terminal. Figure 1This is a hardware structure block diagram of a computer education terminal for a VR platform control method that enables synchronous interaction among multiple education terminals, as provided in an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0017] The non-volatile storage medium can store the operating system and computer program. The computer program includes program instructions that, when executed, cause the processor to perform any VR platform control method for synchronous interaction across multiple educational terminals.

[0018] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0019] The internal memory provides an environment for the execution of computer programs in non-volatile storage media. When the computer program is executed by the processor, it enables the processor to execute any VR platform control method for synchronous interaction among multiple educational terminals.

[0020] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0021] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0022] See Figure 2 The present invention provides a VR platform control method for synchronous interaction of multiple educational terminals, which may include the following steps: S201, Receive user posture data and operation instructions uploaded in real time by multiple educational terminal devices, and generate an initial set of interaction instructions for each educational terminal based on the posture data and operation instructions; Specifically, the wireless communication module can receive user posture data and operation instructions uploaded by multiple educational terminal devices. The posture data includes head position, hand movements and body orientation, and the operation instructions include button clicks and gesture recognition results, generating a raw data stream. The system employs a wireless communication module adapted to the low latency requirements of multiple VR education terminals, specifically a Wi-Fi 6 (IEEE 802.11ax) module. It supports a 160MHz channel bandwidth, with a theoretical transmission rate of up to 9.6Gbps and end-to-end latency stably controlled at ≤20ms. It also supports concurrent access for up to 32 VR education terminals (such as MetaQuest3 and Pico4), meeting the needs of multi-person VR collaboration scenarios (such as VR team training and multi-person VR games).

[0023] Pose data needs to cover the core action dimensions of the user in VR space, and the specific definitions and formats are as follows: Head position: Described using three-dimensional Cartesian coordinates (x, y, z), in meters. Based on the VR scene's world coordinate system (origin set as the scene center, x-axis horizontal to the right, y-axis vertical upward, z-axis vertical outward from the screen). For example, a user's head position is (1.2, 1.8, 3.5), which means that the user's head is 1.2 meters from the origin in the x-direction, 1.8 meters in the y-direction (corresponding to the human standing height), and 3.5 meters in the z-direction. Hand movements: quantified by 26 joint angles (5 joints for each finger + 1 joint for the palm), in degrees. For example, the angle of the first joint of the right index finger is 45° (bent grasping position), the angle of the thumb joint is 0° (extended position), and the angle of the left wrist joint is -15° (turning to the left). Body orientation: Described using Euler angles (roll, pitch, yaw), in degrees. For example, a body yaw angle of 30° (30° to the right of the positive x-axis), a pitch angle of 5° (slight head tilt), and a roll angle of 0° (no left or right tilt).

[0024] The operation instructions cover the physical interaction and visual recognition results of the VR education terminal: Button clicks: including the trigger button, menu button, and grip button of the VR controller. The command status is divided into "PRESSED" (pressed) and "RELEASED" (released). For example, "trigger button pressed" corresponds to the command type "BTN_TRIGGER". Gesture recognition results: Recognized by the front-facing camera of the VR education terminal combined with AI algorithms, supporting common gestures such as "grab", "wave", and "point". Recognition results include types such as "GESTURE_GRAB" (grab successful) and "GESTURE_WAVE" (wave detection), with the coordinates of the target object (such as the target point of the pointing gesture (1.5, 1.8, 3.2)).

[0025] The raw data stream is encapsulated in JSON format. Each data entry contains "device_id" (a unique identifier for the educational terminal, such as "VR_DEV_001"), "timestamp" (data acquisition timestamp, accurate to milliseconds, such as "1699999999876"), "posture_data" (posture data object), and "operation_cmd" (operation command object). An example of the raw data is: {"device_id":"VR_DEV_001","timestamp":1699999999876,"posture_data":{"head_pos": {"x":1.2,"y":1.8,"z":3.5},"hand_joints":{"right_index_1":45,"right_thumb_0":0,"left_wrist":-15},"body_euler":{"roll":0 ,"pitch":5,"yaw":30}},"operation_cmd":{"type":"BTN_TRIGGER","state":"PRESSED","target_pos":{"x":1.5,"y":1.8,"z":3.2}}}.

[0026] The raw data stream is subjected to noise filtering and outlier removal, while the operation commands are standardized to ensure data consistency and integrity, generating a preprocessed attitude data stream and operation command stream. To address interference and format differences in the original data, the process involves three steps: 1. Noise filtering (attitude data): Attitude data is susceptible to sensor jitter (such as random fluctuations of ±2° in hand joint angles) and electromagnetic interference. A Kalman filter algorithm is used to smooth the data. This algorithm reduces noise through a "prediction-update" loop, with the core parameters set as follows: The process noise covariance Q = 0.01 (controls the confidence of the prediction model; the smaller the value, the more dependent it is on the prediction, avoiding excessive fluctuations). The measurement noise covariance R = 0.1 (controls the reliability of sensor data; the smaller the value, the more dependent it is on actual measurements, preserving the true trend of motion).

[0027] For example, the original sequence of the first joint angle of the index finger is [45,47,44,49,46], which is smoothed after filtering to [45,46,45,47,46], eliminating obvious jitter and restoring the trend of continuous grasping action.

[0028] 2. Outlier removal (attitude data): For abnormal values ​​that exceed the normal range of human movement (such as a sudden change in the head's y-axis position to 5.0 meters, far exceeding the reasonable range of 1.5-2.2 meters), the 3σ principle is used for judgment: Calculate the mean μ and standard deviation σ of a certain feature (such as the y-axis position of the head). For example, if μ = 1.8 meters and σ = 0.2 meters, the reasonable range is [μ-3σ, μ+3σ] = [1.2, 2.4]. If the data is outside the range (e.g., 5.0 meters), it is judged as an abnormal value and replaced with the normal data from the previous moment (e.g., 1.8 meters) to avoid misjudgment of subsequent interactive commands.

[0029] 3. Standardization of operating instruction formats: Eliminate naming differences in instructions across different educational terminals (e.g., some educational terminals refer to the "trigger button" as the "fire button"), and standardize instruction encoding and field formats: Command type encoding: "Trigger button pressed" is uniformly "CMD_BTN_TRIGGER_PRESSED", and "Grab gesture successful" is uniformly "CMD_GESTURE_GRAB_SUCCESS"; Additional required fields: Each instruction must include “device_id”, “timestamp”, “cmd_code” (encoding) and “cmd_param” (parameters, such as the target object ID). Missing fields (such as “timestamp”) should be filled with the data reception timestamp.

[0030] The example standardized instruction is: {"device_id":"VR_DEV_001","timestamp":1699999999876,"cmd_code":"CMD_BTN_TRIGGER_PRESSED","cmd_param":{"target_pos":{"x":1.5,"y":1.8,"z":3.2}}}.

[0031] The final generated preprocessed attitude data stream is a JSON array (containing filtered attitude parameters), and the operation command stream is a JSON array (containing standardized commands). The data integrity is 100%, and the noise error is ≤0.5° (joint angle) or 0.01m (position).

[0032] Based on the preprocessed posture data stream and operation command stream, key interaction features, including user intent labels, action trajectories and operation timestamps, are extracted to generate an interaction feature vector set. Extract features that reflect core user interaction needs from preprocessed data and construct structured feature vectors: 1. User intent tag extraction: By analyzing the correlation between posture data and operation commands, rule-based intent recognition is employed. Rule 1: If the operation command is "CMD_GESTURE_GRAB_SUCCESS", and the deviation between the end point of the hand trajectory and VR scene object A (coordinates (1.5,1.8,3.2)) is ≤0.1 meters, mark the intent label as "INTENT_GRAB_OBJ_A" (grabbing object A). Rule 2: If the operation command is “CMD_BTN_TRIGGER_PRESSED”, and the body yaw angle deviates from the direction of movement by ≤10°, mark the intention label as “INTENT_MOVE_FORWARD” (move forward).

[0033] The intent labels use one-hot encoding (e.g., “INTENT_GRAB_OBJ_A” corresponds to 1 in the 5th dimension and 0 in the rest), covering 10 common intents (grabbing, moving, rotating objects, etc.).

[0034] 2. Motion trajectory extraction: Capture the hand / head position sequence 1 second before and after the operation command is triggered, with a sampling frequency consistent with the VR education terminal (90Hz, 90 sampling points per second): Trajectory parameters include starting point (e.g., (1.3, 1.7, 3.3)), ending point (e.g., (1.5, 1.8, 3.2)), average speed (e.g., 0.2 m / s), and trajectory curvature (e.g., 0.1, reflecting the smoothness of the motion). Data compression: Take the first 30 key sampling points (covering the core action stage), each point contains x / y / z three-dimensional coordinates, for a total of 90 numerical features.

[0035] 3. Standardization of operation timestamps: Convert the command "timestamp" into a unified timeline for the VR scene (calculating relative time in seconds, with the scene start time as 0): Example: Original timestamp 1699999999876, scene start time 1699999998000, relative time = (1699999999876 - 1699999998000) / 1000 = 1.876 seconds; The time accuracy is retained to 3 decimal places to ensure that the time synchronization error of multiple educational terminals is ≤1ms.

[0036] 4. Construction of interactive feature vector set: The vector dimension is set to 128 dimensions, organized as "Intent Label (10 dimensions) + Motion Trajectory (90 dimensions) + Time and Device Attributes (28 dimensions)": 1-10 Dimensions: One-hot encoding of intent labels; 11-100 dimensional: x / y / z coordinates of 30 sampling points of the motion trajectory; 101-128 dimensions: 28 auxiliary features such as relative time, device type (e.g., “META_QUEST_3”), and action speed.

[0037] The example feature vector is [0,0,0,0,1,0,...,1.876,0.2,0.1,...] (only the key dimensions are non-zero), and the vector values ​​are all normalized to the range [0,1] for easy processing by subsequent algorithms.

[0038] Based on the set of interactive feature vectors, the features are converted into specific interactive instructions through an instruction mapping algorithm, generating an initial set of interactive instructions for each educational terminal.

[0039] The "feature-instruction" rule base mapping algorithm is used to convert abstract feature vectors into specific instructions that can be executed in the VR scene: 1. Construction of the mapping rule base: The rule base contains a mapping relationship between preset feature conditions and corresponding interaction commands, sorted by priority from high to low (to ensure command uniqueness): Rule 1 (High Priority): If the intent label is “INTENT_GRAB_OBJ_A”, the deviation between the trajectory termination point and object A is ≤0.1 meters, and the relative time is 1.5-2.0 seconds, it is mapped to “CMD_INTERACT_GRAB_OBJ_A”, with parameters including the object ID “OBJ_A” and a gripping force coefficient of 0.8 (0-1, the larger the value, the stronger the grip). Rule 2 (Medium Priority): If the intent label is “INTENT_MOVE_FORWARD”, and the body yaw angle deviates from the movement direction by ≤10°, it is mapped to “CMD_INTERACT_MOVE_FORWARD”, with parameters including a movement distance of 1.0 meter (configurable) and a movement speed of 0.5 meters / second; Rule 3 (low priority): When a high priority rule is not matched, it is mapped to "CMD_INTERACT_NO_OP" (no operation) to avoid invalid instructions.

[0040] 2. Generation of the initial set of interactive commands: Grouped by educational terminal device, each educational terminal's instruction set is a JSON array, and each instruction contains: "cmd_id": Unique command ID (e.g., "CMD_001"); "device_id": Educational terminal identifier; "scene_time": Relative scene time (e.g., 1.876 seconds); "interact_cmd": Interactive command name (e.g., "CMD_INTERACT_GRAB_OBJ_A"); "cmd_params": Execution parameters (e.g., {"obj_id":"OBJ_A","force":0.8}); "exec_status": Initial status "PENDING" (pending execution). The initial set of interaction commands for example VR_DEV_001 is: [{"cmd_id":"CMD_001","device_id":"VR_DEV_001","scene_time":1.876,"interact_cmd":"CMD_INTERACT_GRAB_OBJ_A","cmd_params":{"obj_id":"OBJ_A","force":0.8},"exec_status":"PENDING"},{"cmd_id":" The code `CMD_002","device_id":"VR_DEV_001","scene_time":2.150,"interact_cmd":"CMD_INTERACT_MOVE_FORWARD","cmd_params":{"distance":1.0,"speed":0.5},"exec_status":"PENDING"}` ensures that each instruction can be directly associated with the interaction logic of the VR scene (such as the physical collision and visual feedback of objects in the scene corresponding to grabbing object A).

[0041] S202, perform spatiotemporal consistency verification on the initial set of interactive instructions, construct a synchronous interactive feature matrix, and decompose the synchronous interactive feature matrix into spatial location feature vectors and time series feature vectors to capture the spatiotemporal characteristics of multi-education terminal interaction. Specifically, the initial set of interactive instructions can be checked for spatiotemporal consistency. The timestamp differences and spatial coordinate deviations of the instructions from each educational terminal can be calculated, inconsistent instructions can be identified and marked, and a consistency check report can be generated. Taking a multi-person VR team training scenario as an example, the initial set of interaction instructions includes 100 instructions from three educational terminals (VR_DEV_001, VR_DEV_002, VR_DEV_003). The core interaction target is "Device Button A" on the virtual control panel (the platform's preset spatial coordinates are (1.5, 1.8, 3.2) meters, and the interaction time reference is based on the system time T0 = 1699999999800ms when the platform receives the first instruction). Spatiotemporal consistency verification needs to eliminate "operation misalignment caused by time asynchrony" and "interaction misalignment caused by spatial deviation," specifically implemented as follows: 1. Timestamp Difference Calculation and Judgment: Difference calculation logic: The timestamp Ts of each instruction (the relative time of the scene when the education terminal uploads) needs to be converted into the platform's global time Tg=T0+Ts, and the timestamp difference ΔT=|Tg-T_ref|, where T_ref is the "baseline trigger time" of the interaction target - take the global time of the earliest education terminal instruction that initiated the interaction of the target (such as the instruction Tg1=1699999999876ms of VR_DEV_001, set as T_ref).

[0042] Consistency Threshold Setting: VR interaction is sensitive to latency. The timestamp difference threshold is set to 20ms (exceeding this value will result in a stuttering sensation where you "see someone else's action before triggering your own response"). For example, the instruction Tg2 of VR_DEV_002 is 1699999999900ms, ΔT=24ms>20ms, which is considered inconsistent in time; the instruction Tg3 of VR_DEV_003 is 16999999999886ms, ΔT=10ms≤20ms, which is considered consistent in time.

[0043] 2. Calculation and determination of spatial coordinate deviation: Deviation calculation logic: The "interactive target spatial coordinates (x, y, z)" in the instruction need to be compared with the platform's preset target coordinates (1.5, 1.8, 3.2). The spatial deviation ΔS = √[(x-1.5)]. 2 +(y-1.8) 2 +(z-3.2) 2 (Euclidean distance, reflecting the degree of misalignment between the target location perceived by the educational terminal and the platform baseline).

[0044] Consistency threshold setting: The spatial deviation threshold is set to 0.1 meters (exceeding this value will result in a misalignment where "your finger is clearly aligned with the button, but it cannot be triggered"). For example, the command coordinates of VR_DEV_002 are (1.58, 1.82, 3.25), ΔS = √[(0.08)]. 2 +(0.02) 2 +(0.03) 2The coordinates of VR_DEV_001 are approximately 0.088 meters and ≤0.1 meters, indicating spatial consistency. The command coordinates of VR_DEV_001 are (1.65, 1.8, 3.2), and ΔS = 0.15 meters > 0.1 meters, indicating spatial inconsistency.

[0045] 3. Generate a consistency verification report: The report uses a structured text format and includes four parts: "Basic Instruction Information," "Time Verification Results," "Spatial Verification Results," and "Inconsistency Markers." An example entry is: "Instruction ID: CMD_002; Educational Terminal ID: VR_DEV_002; Interaction Target: Device Button A; Global Time Tg: 1699999999900ms; Time Difference ΔT: 24ms (threshold 20ms, marked: time inconsistency); Interaction Coordinates: (1.58, 1.82, 3.25); Spatial Deviation ΔS: 0.088 meters (threshold 0.1 meters, marked: spatial consistency); Overall Judgment: Inconsistent (time synchronization needs optimization)." The report marks 12 inconsistent instructions, including 8 with time inconsistencies and 4 with spatial inconsistencies, providing a "valid instruction selection basis" for subsequent matrix construction.

[0046] Based on the consistency verification report, the initial set of interactive instructions is organized according to time sequence and spatial location to construct a synchronous interactive feature matrix. In the synchronous interactive feature matrix, the rows represent educational terminal devices, the columns represent time points, and the elements represent spatial coordinates and operation types. Based on the verification report, 88 valid instructions that are consistent in both time and space need to be selected. A matrix is ​​then constructed according to the three-dimensional relationship of "educational terminal-time-feature" to ensure that the matrix can intuitively reflect the interaction status of multiple educational terminals at different times. The specific implementation is as follows: 1. Definition of matrix dimension: Row (Education Terminal Dimension): Take 3 education terminals participating in the same interaction goal. Row index 0 corresponds to VR_DEV_001, 1 corresponds to VR_DEV_002, and 2 corresponds to VR_DEV_003. There are 3 rows in total, and each row represents the complete interaction sequence of an education terminal.

[0047] Columns (Time Dimension): Centered on the baseline trigger time T_ref=16999999999876ms, with a sampling interval of 10ms (the high-frequency sampling requirement for VR interaction to ensure continuous action capture), 49 time points are taken before and after, for a total of 100 columns (time points t0 to t99, corresponding to global time 16999999999827ms to 1699999999926ms), each column representing a time snapshot.

[0048] Element (feature dimension): Each matrix element is a 4-dimensional feature group (x,y,z,op_code), where (x,y,z) is the interaction target coordinate of the educational terminal at that point in time (extracted from valid instructions, and filled with the coordinates of the previous moment when there are no instructions), and op_code is the operation type code (1=button click, 2=gesture capture, 3=no operation, based on the instruction mapping result).

[0049] 2. Matrix construction example: Taking matrix A∈R^(3×100) as an example, some of the element values ​​are as follows: Row 0 (VR_DEV_001), Column 5 (t5=1699999999876+50=1699999999926ms): The element is (1.5,1.8,3.2,1), indicating that the educational terminal will perform a click operation 50ms after the reference time, aligning with the coordinates of button A. Row 1 (VR_DEV_002), Column 5: The element is (1.58, 1.82, 3.25, 1), indicating that at the same time point, the coordinates of the educational terminal are slightly off but within the threshold, and clicks are executed synchronously. Row 2 (VR_DEV_003), Column 3: The element is (1.48, 1.79, 3.19, 3), indicating that the educational terminal has not performed any operation 30ms after the reference time, and is marked as no operation.

[0050] After the matrix is ​​constructed, invalid time points in the columns that are “no educational terminal operation” (such as t90 to t99 without any educational terminal instructions) need to be removed. Finally, 85 columns of valid time points are retained, and the matrix dimension is adjusted to 3×85 to ensure the efficiency of subsequent decomposition.

[0051] Singular value decomposition is performed on the synchronous interaction feature matrix to extract the main spatial patterns and temporal patterns, generating spatial location feature vectors and time series feature vectors. Singular Value Decomposition (SVD) is a core algorithm for matrix dimensionality reduction and pattern extraction. It can separate the "spatial distribution of educational terminals" and "temporal operation patterns" from high-dimensional interaction matrices, adapting to the spatiotemporal characteristic capture requirements of VR multi-education terminals. The specific implementation is as follows: 1. SVD decomposition principle and parameter settings: The SVD decomposition formula for the synchronous interaction feature matrix A (3×85) is A=U×Σ×V^T, where: U∈R^(3×3): Left singular matrix, column vectors correspond to "educational terminal spatial pattern", reflecting the spatial positional relationship of different educational terminals in interaction (e.g., the spatial coordinates of educational terminals 1 and 2 are closer, and the correlation of the column vectors in U is higher). Σ∈R^(3×85): Singular value diagonal matrix, diagonal elements σ1≥σ2≥σ3≥0, the magnitude of the singular values ​​represents the "information contribution" of the corresponding pattern (the larger σ1 is, the more interactive information the pattern contains); V^T∈R^(85×85): Right singular matrix, row vectors correspond to "time series pattern", reflecting the changing pattern of multi-education terminal operation over time (e.g., operation frequency increases from t0 to t5, and decreases after t5).

[0052] 2. Main pattern extraction strategy: Singular value screening: Calculate the cumulative contribution rate of singular values ​​η = Σ(σi) / Σ(σall) (i = 1 to k), and set η ≥ 90% as the threshold (to ensure that more than 90% of the interaction information is retained and to avoid excessive dimensionality reduction). For example, after decomposition, σ1 = 12.5, σ2 = 3.2, σ3 = 0.8, Σσall = 16.5, and when k = 2, η = (12.5 + 3.2) / 16.5 ≈ 95.2% ≥ 90%, so the first two main patterns are extracted.

[0053] Spatial location feature vector generation: Take the first k columns (k=2) of the left singular matrix U to form a spatial feature matrix U_k∈R^(3×2), where each row corresponds to a 2D spatial location feature vector of an educational terminal. For example, the spatial vector of VR_DEV_001 is [0.85,0.12], VR_DEV_002 is [0.82,0.15], and VR_DEV_003 is [0.21,0.93]. The vector values ​​reflect the distribution of the educational terminal in the "main spatial dimension" (the first two columns correspond to the "x-axis deviation dimension" and "y-axis deviation dimension" respectively).

[0054] Time series feature vector generation: Take the first k rows (k=2) of the right singular matrix V^T to form the time feature matrix V_k∈R^(2×85), where each column corresponds to a 2D time series feature vector for a given time point. For example, the time vector for t5 (50ms after the baseline time) is [0.92, 0.08], and the vector for t10 is [0.75, 0.22]. The vector values ​​reflect the intensity of the time point in the "operation initiation dimension" and the "operation maintenance dimension".

[0055] 3. Interpretation of the physical significance of the model: Spatial Pattern 1 (Column 1 of U): This mainly reflects the x-axis deviation between the educational terminal and the interactive target. The larger the vector value (e.g., 0.85), the closer the educational terminal is to the target x-axis reference (1.5 meters), and the better the spatial synchronization. Time pattern 1 (the first line of V^T): mainly reflects the operation start timing. The peak value of the vector (0.92) corresponds to t5, indicating that multiple educational terminals reach the operation synchronization peak 50ms after the base time, which is consistent with the human behavior pattern of "delaying operation after seeing the target".

[0056] The spatial location feature vector and the time series feature vector are normalized to ensure consistent feature scale, thereby generating standardized spatial location feature vector and time series feature vector.

[0057] The original feature vectors have large numerical ranges (e.g., spatial vector values ​​range from -1 to 1, and time vector values ​​range from 0 to 10), which can cause subsequent optimization algorithms to favor features with large numerical ranges. Therefore, min-max scaling is needed to map them to the [0,1] interval to achieve scale uniformity. The specific implementation is as follows: 1. Determination of normalization formula and parameters: The minimum-maximum standardized formula is: x'=(x-x_min) / (x_max-x_min), where: x is the original feature value, x_min is the minimum value of the feature vector, and x_max is the maximum value; The normalization target interval is set to [0,1] (a common standardized range for VR interaction features, which makes it easier to intuitively understand the "importance ratio" when allocating weights later).

[0058] 2. Spatial location feature vector normalization: Taking the 2D spatial vectors of 3 educational terminals as an example: The first column of the spatial vector (x-axis deviation mode): the original values ​​are [0.85, 0.82, 0.21], x_min=0.21, x_max=0.85; The normalized value of VR_DEV_001 is: (0.85-0.21) / (0.85-0.21)=0.64 / 0.64=1.0; The normalized value of VR_DEV_002 is: (0.82-0.21) / 0.64≈0.61 / 0.64≈0.95; The normalized value of VR_DEV_003 is: (0.21-0.21) / 0.64=0.0; The second column of the spatial vector (y-axis deviation mode): the original values ​​are [0.12, 0.15, 0.93], x_min=0.12, x_max=0.93; The normalized value of VR_DEV_001 is: (0.12-0.12) / (0.93-0.12)=0.0; The normalized value of VR_DEV_002 is: (0.15-0.12) / 0.81≈0.03 / 0.81≈0.04; The normalized value of VR_DEV_003 is: (0.93-0.12) / 0.81=0.81 / 0.81=1.0; The standardized spatial location feature matrix is: [[1.0,0.0],[0.95,0.04],[0.0,1.0]], with values ​​all in [0,1]. This allows for a direct comparison of the performance of the educational terminal in different spatial modes (e.g., VR_DEV_001 has the best synchronization on the x-axis, while VR_DEV_003 has the best synchronization on the y-axis).

[0059] 3. Time series feature vector normalization: Taking a 2D time vector with 85 time points as an example: The first row of the time vector (operation start mode): original value range [0.05, 0.92], x_min=0.05, x_max=0.92; The normalized value of t5 is: (0.92-0.05) / (0.92-0.05)=0.87 / 0.87=1.0 (synchronization peak value); The normalized value of t0 is: (0.05-0.05) / 0.87=0.0 (initial stage of operation); The second row of the time vector (operation maintenance mode): original value range [0.02, 0.35], x_min=0.02, x_max=0.35; The normalized value of t10 is: (0.22-0.02) / (0.35-0.02)=0.20 / 0.33≈0.61; The standardized time series feature matrix is ​​a 2×85 matrix with all elements in [0,1], which eliminates the scale difference of "large operation start mode value and small maintenance mode value", and the weights can be fairly allocated in subsequent optimization.

[0060] After normalization, the numerical ranges of the standardized spatial vector and the time vector are completely unified, providing "unbiased feature input" for the multi-objective collaborative optimization in step four, and avoiding the optimization bias towards a certain type of feature due to scale issues.

[0061] S203, using a multi-objective collaborative optimization algorithm to dynamically assign weights and optimize delays for the spatial location feature vector and the time series feature vector, thereby generating an optimized collaborative control strategy; Specifically, the objective function of the multi-objective collaborative optimization algorithm can be set, with objectives including minimizing interaction latency, maximizing spatial synchronization, and minimizing energy consumption, thereby generating an optimized objective configuration; The multi-objective collaborative optimization algorithm uses the non-dominated sorting genetic algorithm (NSGA-II). This algorithm maintains population diversity through non-dominated sorting and crowding distance, and can efficiently balance multiple conflicting objectives (such as reducing latency may increase energy consumption). It is suitable for the complex requirements of VR multi-education terminal synchronization (taking a multi-person VR device operation training scenario as an example, there are 3 education terminals: VR_DEV_001 (main operation education terminal), VR_DEV_002 (auxiliary education terminal 1), and VR_DEV_003 (auxiliary education terminal 2)).

[0062] 1. Definition of objective function: Minimize interaction latency (f1): Interaction latency refers to the total time from when the platform generates a synchronization command to when the educational terminal receives and executes it (including transmission latency + educational terminal processing latency). The formula is: f1 = (T_recv - T_send), where T_send is the timestamp of the platform sending the command, and T_recv is the timestamp of the educational terminal's execution feedback. The target value should be ≤20ms (the critical value for smooth VR interaction). Example: For a certain command, T_send = 1699999999876ms, T_recv = 16999999999890ms, f1 = 14ms (meets the requirement).

[0063] Maximizing spatial synchronization (f2): Spatial synchronization measures the coordinate consistency of interactive targets across multiple educational terminals. It is expressed as the reciprocal of the average Euclidean distance between all educational terminals and the platform's reference coordinates. The formula is: f2=1 / [(1 / n)×Σ√((x_i-x_ref)] 2 +(y_i-y_ref) 2 +(z_i-z_ref) 2 [], where n is the number of educational terminals (n=3), (x_i,y_i,z_i) are the coordinates of the educational terminals, and (x_ref,y_ref,z_ref) are the platform reference coordinates (e.g., the coordinates of virtual button A are (1.5,1.8,3.2)). The larger the value of f2, the better the synchronization. Example: the coordinate deviations of the educational terminals are 0.08m, 0.09m, and 0.10m respectively, with an average deviation of 0.09m, f2=1 / 0.09≈11.11.

[0064] Minimize energy consumption (f3): Energy consumption encompasses both computing and data transmission energy consumption of the educational terminal. The formula is: f3 = (CPU_util × P_cpu) + (Data_size × P_trans), where CPU_util is the CPU utilization rate of the educational terminal (%), P_cpu is the power consumption per unit CPU utilization (e.g., 0.01W / %), Data_size is the instruction data volume (MB), and P_trans is the power consumption per unit data transmission (e.g., 0.05W / MB). The target value must be ≤5W (the energy consumption threshold for continuous operation of the VR educational terminal). Example: CPU_util = 60%, P_cpu = 0.01W / %, Data_size = 2MB, P_trans = 0.05W / MB, f3 = (60 × 0.01) + (2 × 0.05) = 0.6 + 0.1 = 0.7W (meets the requirement).

[0065] 2. Optimize target configuration generation: The configuration includes "algorithm parameters, target priority, and constraints": Algorithm parameters: Population size 50 (generating 50 optimization schemes per iteration), number of iterations 30 (ensuring convergence, objective function value fluctuation ≤5% after 30 iterations), crossover probability 0.8 (controlling scheme diversity), mutation probability 0.1 (avoiding local optima); Target priority: The initial priority is "interaction latency > spatial synchronization > energy consumption" (VR interaction latency is sensitive and should be prioritized), corresponding to initial weights w1=0.3, w2=0.4, w3=0.3; Constraints: Interaction delay ≤ 20ms, spatial synchronization f2 ≥ 10, energy consumption ≤ 5W. Schemes exceeding the constraints will be eliminated.

[0066] Based on the optimized target configuration, dynamic weight allocation is performed on spatial location feature vectors and time series feature vectors. The weight coefficients are adjusted according to real-time network conditions and device performance to generate dynamic weight vectors. The core of dynamic weight allocation is "adjustment on demand"—when the real-time state of a certain target deviates from the constraints, its weight is increased to prioritize optimization. The adjustment is based on "network status monitoring indicators" and "equipment performance monitoring indicators". These two types of indicators are collected through the real-time monitoring module of the VR platform (sampling frequency of 1Hz to ensure timely response to changes).

[0067] 1. Definition and thresholds of monitoring indicators: Real-time network status: including bandwidth (in Mbps, threshold ≥50Mbps), network latency (in ms, threshold ≤20ms), and packet loss rate (in %, threshold ≤1%). Device performance includes education terminal CPU utilization (in %, threshold ≤70%), battery power (in %, threshold ≥30%), and GPU rendering frame rate (in fps, threshold ≥90fps).

[0068] 2. Weighting adjustment rules and examples: Rule 1: When network latency exceeds the limit, increase the weight of "minimize interaction latency": If the real-time network latency increases from 15ms to 25ms (exceeding the threshold of 20ms), the interaction latency weight w1 is increased from 0.3 to 0.5, the spatial synchronization weight w2 is decreased from 0.4 to 0.2, and the energy consumption weight w3 remains at 0.3 (at this time, latency is reduced first, sacrificing some spatial synchronization tolerance). Rule 2: When the device battery is too low, increase the weight of "minimize energy consumption": If the battery power of VR_DEV_003 drops from 40% to 25% (below the threshold of 30%), the energy consumption weight w3 is increased from 0.3 to 0.4, the interaction latency weight w1 is decreased to 0.2, and the spatial synchronization weight w2 remains at 0.4 (at this time, the energy consumption of the educational terminal is reduced first to avoid shutdown). Rule 3: When spatial synchronization is not up to standard, increase the weight of “maximizing spatial synchronization”: If spatial synchronization f2 drops from 11.11 to 9.5 (below the threshold of 10), the spatial synchronization weight w2 is increased from 0.4 to 0.5, the interaction delay weight w1 is reduced to 0.2, and the energy consumption weight w3 is reduced to 0.3 (at this time, coordinate deviation is corrected first to ensure that the operation of multiple educational terminals is aligned).

[0069] 3. Dynamic weight vector generation: Taking the scenario of "network latency 25ms + VR_DEV_003 battery 25%" as an example, the adjusted weight vector is [w1=0.2, w2=0.4, w3=0.4], and the sum of each element in the vector is 1 (to ensure weight normalization). This vector will serve as an important basis for subsequent latency optimization - energy consumption and spatial synchronization have higher weights, and the relationship between the two and latency needs to be balanced in the optimization.

[0070] By using dynamic weight vectors, delay optimization is performed on time series feature vectors, and network latency is compensated by prediction algorithms to generate optimized time series feature vectors. The time series feature vector is a 2×85 normalized matrix (2 time series patterns × 85 time points). The core of latency optimization is to "predict future network latency and compensate in advance" to avoid the timing misalignment of instructions caused by latency. Long Short-Term Memory (LSTM) network is selected as the prediction algorithm (LSTM is good at capturing long-term dependencies of time series data and adapts to the fluctuation pattern of network latency).

[0071] 1. LSTM prediction model construction and training: Input and output definitions: The input is the network delay sequence of the past 10 sampling points (historical data of 100ms with a sampling interval of 10ms, such as [18,20,22,25,23,24,26,25,27,26]ms), and the output is the delay prediction value of the next sampling point (50ms later). Model structure: 1 input layer (10 dimensions), 2 LSTM hidden layers (32 neurons per layer, tanh activation function), 1 fully connected output layer (1 dimension). The training data is network latency logs from the past 24 hours (864,000 data points in total). The training loss function is mean squared error (MSE). The prediction error after training is ≤2ms. Example of prediction: Input historical delay [18,20,22,25,23,24,26,25,27,26]ms, the model outputs a predicted delay of 25ms (that is, the network delay is expected to be 25ms after 50ms).

[0072] 2. Delay compensation and time series vector optimization: Compensation logic: If the prediction delay is Δt_pred, advance the "instruction sending timestamp" of each time point in the time series feature vector by Δt_pred to ensure that the education terminal receives and executes the instruction at the expected time point (compensation formula: T_send_opt=T_send_org-Δt_pred, where T_send_org is the original sending timestamp and T_send_opt is the optimized timestamp). Optimization example: In the original time series feature vector, the instruction sending timestamp at t5 (corresponding to global time 1699999999926ms) is T_send_org=1699999999926ms, the prediction delay Δt_pred=25ms, and after optimization, T_send_opt=1699999999926-25=1699999999901ms; Optimized time series feature vector: Adjust the timestamps of all time points in the original matrix according to the above logic to generate a new 2×85 normalized matrix. For example, the original time series pattern value of t5 [1.0, 0.61] (operation start peak) becomes 1699999999901ms after optimization, while the pattern value remains unchanged (only the time dimension is adjusted, and the feature intensity is not changed).

[0073] 3. Optimize validity verification: After optimization, it is necessary to verify whether the interaction latency meets the standard: The actual measured transmission latency of the optimized command is 23ms (close to the predicted value of 25ms), which is 17.9% lower than the original 28ms, meeting the constraint of ≤20ms (due to the prediction error of 2ms, the actual latency is slightly higher than the threshold, which is acceptable).

[0074] By combining spatial location feature vectors and optimized time series feature vectors, a coordinated control strategy, including instruction scheduling order and resource allocation scheme, is output through a policy generator.

[0075] The strategy generator adopts a hybrid architecture of "rules + model" - first, it determines the priority of educational terminals based on spatial vectors, then determines the scheduling sequence by combining optimized time vectors, and finally allocates resources according to dynamic weights to ensure that the strategy simultaneously meets the requirements of "spatial alignment, timing synchronization, and controllable energy consumption".

[0076] 1. Instruction scheduling order is determined: Priority allocation: Based on the "spatial synchronization score" of spatial location feature vector (score = value of the first column of spatial vector × 0.6 + value of the second column × 0.4, with weights corresponding to the synchronization importance of the x-axis and y-axis), the higher the score, the higher the priority; Example: VR_DEV_001 spatial vector [1.0, 0.0], score = 1.0 × 0.6 + 0.0 × 0.4 = 0.6; VR_DEV_002 spatial vector [0.95, 0.04], score = 0.95 × 0.6 + 0.04 × 0.4 = 0.586; VR_DEV_003 spatial vector [0.0, 1.0], score = 0.0 × 0.6 + 1.0 × 0.4 = 0.4; Scheduling order: sorted from high to low priority, i.e., VR_DEV_001→VR_DEV_002→VR_DEV_003, with a scheduling interval of 5ms (to avoid network congestion caused by concurrent commands). For example, the command sending time for VR_DEV_001 is 1699999999901ms, for VR_DEV_002 it is 1699999999906ms, and for VR_DEV_003 it is 1699999999911ms.

[0077] 2. Resource allocation plan determined: Resource types include network bandwidth (in Mbps), educational terminal CPU computing power (in GHz), and GPU rendering channels (in units). The allocation is based on a dynamic weight vector [0.2, 0.4, 0.4] (energy consumption has the highest weight, and the allocation of high-energy-consuming resources needs to be controlled). Allocation example: Network bandwidth: Total bandwidth 100Mbps, allocated according to priority: VR_DEV_001 (high priority) 25Mbps, VR_DEV_002 (medium priority) 20Mbps, VR_DEV_003 (low priority) 15Mbps, and the remaining 40Mbps is reserved (to cope with sudden data surges). CPU computing power: The total CPU computing power of the education terminal is 6GHz, with VR_DEV_001 allocated 2GHz (to support complex rendering), VR_DEV_002 allocated 1.8GHz, and VR_DEV_003 allocated 1.2GHz (to reduce its energy consumption and keep the CPU utilization rate below 50%). GPU rendering channels: There are 10 channels in total, with 4 allocated to VR_DEV_001 (to ensure smooth operation feedback), 3 allocated to VR_DEV_002, and 3 allocated to VR_DEV_003.

[0078] The collaborative control strategy is output in JSON format, which includes three parts: "scheduling order", "resource allocation" and "constraints".

[0079] S204, Based on the aforementioned collaborative control strategy, generate multi-education terminal synchronization instructions to uniformly schedule the rendering content and interactive feedback timing of each education terminal device.

[0080] Specifically, it can analyze collaborative control strategies, extract instruction scheduling order and resource allocation schemes, and generate synchronous instruction templates for multiple educational terminals; The collaborative control strategy generated in the early stage is stored in JSON format (e.g., strategy ID: STR_001), containing three core parts: "scheduling_order", "resource_allocation", and "constraints". Parsing requires accurately extracting key parameters to provide a unified structural framework for subsequent instruction generation. The specific implementation is as follows: 1. Collaborative control strategy analysis process: The strategy file is processed using Python's built-in JSON parser (suitable for lightweight data parsing, with a parsing latency of ≤1ms) following a "hierarchical extraction + parameter validation" logic. The first layer extracts the "scheduling order": extract the device_id (such as VR_DEV_001, VR_DEV_002) and send_time (instruction sending timestamp, such as 1699999999901ms, 16999999999906ms) of each educational terminal from the scheduling_order field, and verify whether send_time conforms to the platform's global timeline (it must be after the scene start time T_start=1699999998000ms to avoid invalid time). The second layer extracts the "resource allocation scheme": extract the bandwidth (25Mbps, 20Mbps), CPU computing power (2.0GHz, 1.8GHz), and GPU rendering channels (4, 3) of each educational terminal from the resource_allocation field, and verify whether the resource parameters meet the hardware limit of the educational terminal (e.g., the maximum CPU computing power of VR_DEV_001 is 2.5GHz, and 2.0GHz is within a reasonable range). The third layer extracts "constraints": records constraints such as delay≤20ms and energy≤5W, which serve as verification standards for subsequent instruction execution.

[0081] 2. Generation of synchronized instruction templates for multiple educational terminals: The template must include a "fixed structure + placeholders" to ensure that different educational terminals can quickly fill in personalized parameters based on the template. The template field definitions are as follows: Fixed fields: template_id (unique template identifier, such as TPL_001), cmd_category (command category, divided into RENDER and FEEDBACK), protocol_version (communication protocol version, such as V1.0, to ensure compatibility with educational terminals); Placeholder fields: {device_id} (educational terminal identifier placeholder), {exec_time} (instruction execution time placeholder, calculated based on send_time + transmission delay), {resource_param} (resource parameter placeholder, such as bandwidth, computing power), {content_param} (content parameter placeholder, such as rendering object, feedback type).

[0082] Based on the multi-education terminal synchronization instruction template, specific rendering instructions and interactive feedback instructions are generated for each education terminal device, thus generating a specific instruction set for the education terminal. Rendering instructions define the display parameters of elements in the VR scene, ensuring visual consistency across multiple educational terminals. Interactive feedback instructions trigger hardware feedback from the educational terminals (such as controller vibration and sound effects) to enhance user immersion. Template placeholders need to be filled based on the educational terminal hardware parameters (such as screen resolution and controller vibration intensity range) and scene requirements (such as virtual objects). The specific implementation is as follows: 1. Rendering instruction generation (taking virtual button A interaction scene as an example): Scene element {scene_element}: filled with "virtual_button_A" (interaction target, platform default coordinates (1.5, 1.8, 3.2)); Resolution: Filled according to the screen parameters of the educational terminal. VR_DEV_001 (MetaQuest3) is "1920×1080" (single-eye resolution, combined with both eyes is 3840×2160, which meets the VR visual clarity requirements). VR_DEV_002 (Pico4) is the same (to ensure visual consistency across multiple educational terminals). Frame rate {frame_rate}: Fill with "90fps" (the critical frame rate for smooth VR scenes; below 90fps, it is easy to cause dizziness). Resource parameters {bandwidth}, "{cpu_power}", and "{gpu_channels}": extracted from the collaborative control strategy, VR_DEV_001 corresponds to 25Mbps, 2.0GHz, and 4 channels.

[0083] Based on the specific instruction set of the educational terminal, the rendering content and interactive feedback timing of each educational terminal device are uniformly scheduled through a timing coordinator to ensure that all educational terminals perform corresponding operations at the same time and generate a synchronous scheduling plan. The timing coordinator is the core module for synchronizing multiple educational terminals. It employs a "reference time anchoring + delay compensation" algorithm to eliminate timing deviations caused by transmission delays and device processing delays, ensuring that all educational terminals execute instructions at the same global time point. The specific implementation is as follows: 1. Core algorithm and parameters of the timing coordinator: Baseline time anchoring: Select the minimum value of send_time of all educational terminal instructions as the baseline time T_base. For example, in the collaborative control strategy, the minimum send_time is 1699999999901ms of VR_DEV_001. Set T_base=16999999999900ms (1ms in advance, reserve buffer). Delay Compensation: The "transmission delay ΔT_trans" (time from the platform to the educational terminal) and "processing delay ΔT_proc" (time from receiving the instruction to execution) of each educational terminal are obtained through real-time network monitoring. The total compensation delay ΔT_total = ΔT_trans + ΔT_proc. Example: VR_DEV_001 has ΔT_trans = 9ms, ΔT_proc = 2ms, and ΔT_total = 11ms; VR_DEV_002 has ΔT_trans = 10ms, ΔT_proc = 2ms, and ΔT_total = 12ms. Synchronous execution time calculation: The execution time of the instruction on the education terminal is T_exec = T_base + ΔT_total, ensuring that T_exec is consistent across all education terminals.

[0084] 2. Synchronization scheduling plan generation: The plan adopts a tabular structure (text format, not a table), which includes "plan_id, base_time, terminal_sync_info". Each record in terminal_sync_info corresponds to an educational terminal and includes device_id, cmd_ids (associated rendering + feedback command ID), T_exec, ΔT_trans, and ΔT_proc.

[0085] Execute the synchronization scheduling plan, send synchronization instructions to each educational terminal device, monitor the execution status, and output the synchronization interaction results of multiple educational terminals.

[0086] During the execution phase, it is necessary to ensure reliable command transmission and real-time status monitoring, and finally output a quantitative synchronization effect report. The specific implementation is as follows: 1. Synchronization command sending: Communication module: It adopts Wi-Fi 6 (IEEE 802.11ax) module, which supports low latency (≤10ms) and high reliability (packet loss rate ≤0.1%). The transmission protocol is UDP (User Datagram Protocol, connectionless overhead, suitable for real-time commands), supplemented by TCP retransmission mechanism (for critical commands such as rendering commands, if UDP packets are lost, TCP retransmission will be triggered). Sending timing: Based on the synchronization scheduling plan's T_base, the corresponding instructions are sent to each educational terminal at time T_base. Example: T_base = 1699999999898ms, at this time, REND_001 and FEED_001 instructions are sent to VR_DEV_001. The data packet contains an instruction checksum (CRC32 algorithm, to ensure data integrity; if the checksum fails, the educational terminal requests a retransmission).

[0087] 2. Execution status monitoring: Monitoring Mechanism: A dual mechanism of "heartbeat packet + status feedback" is adopted. The platform sends a heartbeat request to the education terminal every 10ms. After the instruction is executed (T_exec=1699999999910ms), the education terminal immediately returns status feedback, which includes "status_code (0x00=execution successful, 0x01=rendering failed, 0x02=feedback failure), actual_exec_time (actual execution timestamp), error_msg (reason for failure, such as "insufficient GPU channels")". Exception Handling: If the educational terminal does not return feedback within T_exec+5ms, the platform determines it as a timeout and triggers a retransmission (maximum of 2 retransmissions, 2ms interval). If it still fails, it is marked as "execution failed" and recorded in the results. Example: VR_DEV_003 timed out on the first transmission, returned 0x00 after retransmission, actual_exec_time=1699999999911ms (deviation of 1ms, acceptable).

[0088] 3. Synchronous interactive output of results from multiple educational terminals: The results are presented in a structured report format, including "result_report_id, sync_metrics, terminal_status_summary, and constraint_check". Synchronization metrics: Synchronization rate (number of successfully executed educational terminals / total number of educational terminals × 100%, e.g., 3 / 3 = 100%), average execution latency (average of |actual_exec_time - target_exec_time|, e.g., (1+1+0) / 3 ≈ 0.67ms), visual consistency score (by comparing screenshots of educational terminals, using the structural similarity index SSIM, e.g., 0.98, ≥0.95 is excellent); Summary of educational terminal status: VR_DEV_001 (Success, actual_exec_time=1699999999911ms), VR_DEV_002 (Success, 16999999999910ms), VR_DEV_003 (Success, 16999999999911ms). Constraint verification: Interaction delay (11ms≤20ms), energy consumption (VR_DEV_001=0.7W≤5W), and spatial synchronization (f2=11.2≥10) all meet the constraints.

[0089] Another embodiment of the present invention provides a VR platform control system for synchronous interaction of multiple educational terminals, see [link to relevant documentation]. Figure 3 The system may include: The receiving module 301 is used to receive user posture data and operation instructions uploaded in real time by multiple educational terminal devices, and generate an initial interaction instruction set for each educational terminal based on the posture data and operation instructions. The construction module 302 is used to perform spatiotemporal consistency verification on the initial set of interactive instructions, construct a synchronous interactive feature matrix, and decompose the synchronous interactive feature matrix into spatial location feature vectors and time series feature vectors to capture the spatiotemporal characteristics of multi-education terminal interaction. Optimization module 303 is used to perform dynamic weight allocation and delay optimization on the spatial location feature vector and time series feature vector using a multi-objective collaborative optimization algorithm to generate an optimized collaborative control strategy; The generation module 304 is used to generate multi-education terminal synchronization instructions based on the collaborative control strategy, and to uniformly schedule the rendering content and interactive feedback timing of each education terminal device.

[0090] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0091] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0092] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.

[0093] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.

Claims

1. A VR platform control method for synchronous interaction of multiple educational terminals, characterized in that, The method includes: Receive user posture data and operation instructions uploaded in real time from multiple educational terminal devices, and generate an initial set of interaction instructions for each educational terminal based on the posture data and operation instructions; Spatiotemporal consistency verification is performed on the initial set of interactive instructions, a synchronous interactive feature matrix is ​​constructed, and the synchronous interactive feature matrix is ​​decomposed into spatial location feature vectors and time series feature vectors to capture the spatiotemporal characteristics of multi-education terminal interaction. A multi-objective collaborative optimization algorithm is used to dynamically assign weights and optimize delays to the spatial location feature vector and the time series feature vector to generate an optimized collaborative control strategy. Based on the aforementioned collaborative control strategy, multi-education terminal synchronization instructions are generated to uniformly schedule the rendering content and interactive feedback timing of each education terminal device.

2. The method according to claim 1, characterized in that, The step involves receiving user posture data and operation commands uploaded in real time from multiple educational terminal devices, and generating an initial set of interaction commands for each educational terminal based on the posture data and operation commands, including: The system receives user posture data and operation instructions uploaded by multiple educational terminal devices through a wireless communication module. The posture data includes head position, hand movements, and body orientation, while the operation instructions include button clicks and gesture recognition results, generating a raw data stream. The raw data stream is subjected to noise filtering and outlier removal, while the operation commands are standardized to ensure data consistency and integrity, generating a preprocessed attitude data stream and operation command stream. Based on the preprocessed posture data stream and operation command stream, key interaction features, including user intent labels, action trajectories and operation timestamps, are extracted to generate an interaction feature vector set. Based on the set of interactive feature vectors, the features are converted into specific interactive instructions through an instruction mapping algorithm, generating an initial set of interactive instructions for each educational terminal.

3. The method according to claim 2, characterized in that, The process of performing spatiotemporal consistency verification on the initial set of interactive instructions, constructing a synchronous interactive feature matrix, and decomposing the synchronous interactive feature matrix into spatial location feature vectors and time series feature vectors to capture the spatiotemporal characteristics of multi-educational terminal interaction includes: The initial set of interactive instructions is checked for spatiotemporal consistency. The timestamp differences and spatial coordinate deviations of the instructions from each educational terminal are calculated. Inconsistent instructions are identified and marked, and a consistency check report is generated. Based on the consistency verification report, the initial set of interactive instructions is organized according to time sequence and spatial location to construct a synchronous interactive feature matrix. In the synchronous interactive feature matrix, the rows represent educational terminal devices, the columns represent time points, and the elements represent spatial coordinates and operation types. Singular value decomposition is performed on the synchronous interaction feature matrix to extract the main spatial patterns and temporal patterns, generating spatial location feature vectors and time series feature vectors. The spatial location feature vector and the time series feature vector are normalized to ensure consistent feature scale, thereby generating standardized spatial location feature vector and time series feature vector.

4. The method according to claim 3, characterized in that, The method of using a multi-objective collaborative optimization algorithm to dynamically assign weights and optimize delays in the spatial location feature vector and the time series feature vector to generate an optimized collaborative control strategy includes: Define the objective function of the multi-objective collaborative optimization algorithm, with objectives including minimizing interaction latency, maximizing spatial synchronization, and minimizing energy consumption, and generate an optimized objective configuration. Based on the optimized target configuration, dynamic weight allocation is performed on spatial location feature vectors and time series feature vectors. The weight coefficients are adjusted according to real-time network conditions and device performance to generate dynamic weight vectors. By using dynamic weight vectors, delay optimization is performed on time series feature vectors, and network latency is compensated by prediction algorithms to generate optimized time series feature vectors. By combining spatial location feature vectors and optimized time series feature vectors, a coordinated control strategy, including instruction scheduling order and resource allocation scheme, is output through a policy generator.

5. The method according to claim 4, characterized in that, The step of generating multi-education terminal synchronization instructions based on the collaborative control strategy, and uniformly scheduling the rendering content and interaction feedback timing of each education terminal device, includes: Analyze the collaborative control strategy, extract the instruction scheduling sequence and resource allocation scheme, and generate a multi-education terminal synchronous instruction template; Based on the multi-education terminal synchronization instruction template, specific rendering instructions and interactive feedback instructions are generated for each education terminal device, thus generating a specific instruction set for the education terminal. Based on the specific instruction set of the educational terminal, the rendering content and interactive feedback timing of each educational terminal device are uniformly scheduled through a timing coordinator to ensure that all educational terminals perform corresponding operations at the same time and generate a synchronous scheduling plan. Execute the synchronization scheduling plan, send synchronization instructions to each educational terminal device, monitor the execution status, and output the synchronization interaction results of multiple educational terminals.

6. A VR platform control system for synchronous interaction of multiple educational terminals, characterized in that, The system includes: The receiving module is used to receive user posture data and operation instructions uploaded in real time by multiple educational terminal devices, and generate an initial set of interaction instructions for each educational terminal based on the posture data and operation instructions. The module is used to perform spatiotemporal consistency verification on the initial set of interactive instructions, construct a synchronous interactive feature matrix, and decompose the synchronous interactive feature matrix into spatial location feature vectors and time series feature vectors to capture the spatiotemporal characteristics of multi-education terminal interaction. The optimization module is used to dynamically allocate weights and optimize delays for the spatial location feature vector and the time series feature vector using a multi-objective collaborative optimization algorithm, thereby generating an optimized collaborative control strategy. The generation module is used to generate synchronization instructions for multiple educational terminals based on the collaborative control strategy, and to uniformly schedule the rendering content and interactive feedback timing of each educational terminal device.

7. The system according to claim 6, characterized in that, The receiving module is specifically used for: The system receives user posture data and operation instructions uploaded by multiple educational terminal devices through a wireless communication module. The posture data includes head position, hand movements, and body orientation, while the operation instructions include button clicks and gesture recognition results, generating a raw data stream. The raw data stream is subjected to noise filtering and outlier removal, while the operation commands are standardized to ensure data consistency and integrity, generating a preprocessed attitude data stream and operation command stream. Based on the preprocessed posture data stream and operation command stream, key interaction features, including user intent labels, action trajectories and operation timestamps, are extracted to generate an interaction feature vector set. Based on the set of interactive feature vectors, the features are converted into specific interactive instructions through an instruction mapping algorithm, generating an initial set of interactive instructions for each educational terminal.

8. The system according to claim 7, characterized in that, The building module is specifically used for: The initial set of interactive instructions is checked for spatiotemporal consistency. The timestamp differences and spatial coordinate deviations of the instructions from each educational terminal are calculated. Inconsistent instructions are identified and marked, and a consistency check report is generated. Based on the consistency verification report, the initial set of interactive instructions is organized according to time sequence and spatial location to construct a synchronous interactive feature matrix. In the synchronous interactive feature matrix, the rows represent educational terminal devices, the columns represent time points, and the elements represent spatial coordinates and operation types. Singular value decomposition is performed on the synchronous interaction feature matrix to extract the main spatial patterns and temporal patterns, generating spatial location feature vectors and time series feature vectors. The spatial location feature vector and the time series feature vector are normalized to ensure consistent feature scale, thereby generating standardized spatial location feature vector and time series feature vector.

9. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-5 when it is run.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-5.