An asynchronous concurrent virtual reality synchronization method and system

CN122420581BActive Publication Date: 2026-09-01RESEARCH INSTITUTE OF TSINGHUA UNIVERSITY IN SHENZHEN
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610894461.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-01
Estimated Expiration
2046-06-22

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种异步并发的虚拟现实同步方法及系统,以解决上述背景技术中提到的现有的修改成本高、同步可靠性低、难以通过固定偏移量对时间轴进行精确矫正等问题

Benefits of technology

[0017]由上述技术方案可知,本发明与现有技术相比至少具备以下优点和积极效果:本发明将单一的线性时间轴拆分为多个异步并发的子内容,每个子内容拥有独立的时间轴矫正锚点;即使部分子内容因网络波动或识别错误导致矫正失败,其他子内容仍能正常同步;在相同网络丢包率下,整体画面出现明显失同步的概率从单次全场景失败降低为局部微小失同步,人眼难以察觉,从而大幅提高系统在实际复杂环境中的同步可靠性。采用碎片化内容实时合成渲染的方式,仅需替换或重新生成对应的一个或几个并发子内容,修改时间显著缩短,极大提升展览内容的迭代效率。通过引入基于空间位置的展项筛选机制,每个虚拟现实头盔仅同步距离阈值范围内的目标展项,避免全量同步带来的数据冗余;结合服务器线程池为每个头盔分配独立处理线程,有效降低单台设备的通信及数据处理压力,实现低延迟帧同步。针对视频展项、交互展项和机械装置展项分别设计差异化的同步策略(时间轴对齐、页面状态切换、传感器变换更新),并采用统一的数据结构进行封装,能够适应博物馆等复杂环境下多类型展项并存的同步需求;特别是对于不可控的物理机械装置,通过实时传感器数据或图像识别进行状态跟随,实现虚实场景的精确对齐。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122420581B_ABST
    Figure CN122420581B_ABST
Patent Text Reader

Abstract

This invention discloses an asynchronous concurrent virtual reality synchronization method and system, comprising: splitting the virtual reality content to be synchronized into multiple concurrently played sub-contents, and configuring an independent timeline correction anchor point for each sub-content; acquiring real-time status data of various types of external exhibits such as videos, interactions, and mechanical devices; filtering target exhibits within a distance threshold based on the real-time spatial position of the virtual reality display device; performing frame synchronization between the status data of the target exhibits and the timeline correction anchor points of the corresponding sub-contents to generate synchronization data; and rendering and synthesizing the multiple sub-contents in real-time based on the synchronization data, outputting a virtual reality image synchronized with the external exhibits. This invention, through its asynchronous concurrent structure and spatial filtering mechanism, improves the reliability and fault tolerance of synchronization, reduces content maintenance costs, and is suitable for virtual-real synchronization in complex scenarios such as museums.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality technology, specifically to an asynchronous concurrent virtual reality synchronization method and system. Background Technology

[0002] Currently, virtual reality systems are mainly divided into two categories: interactive virtual reality systems and linear playback virtual reality systems.

[0003] Interactive virtual reality systems typically enable user interaction through controllers, gestures, pupil tracking, or image recognition via helmet-mounted cameras, and are widely used in virtual reality games, simulation training, and other scenarios. While these systems can respond to real-time user actions, they cannot guarantee high-precision synchronization with linearly played content such as external large-screen videos. Even if the user triggers synchronization through specific gestures or communication, the fluctuating recognition rate results in poor repeatability with each trigger, making it difficult to accurately correct the timeline using a fixed offset.

[0004] Linear playback virtual reality systems, represented by egg chairs and virtual reality movies, consist of pre-rendered 3D films that users can only passively watch. These systems typically perform a timeline synchronization at the beginning of playback, but then do not correct it throughout the entire playback process. If the content or duration of the video on the external screen is adjusted in any way, the entire virtual reality content needs to be re-rendered and remade, resulting in extremely high modification costs. Furthermore, linear playback systems only perform correction at the beginning and end; any time drift occurring in the middle cannot be corrected, leading to a noticeable misalignment between the virtual content and the external visuals after prolonged playback.

[0005] Therefore, existing technologies cannot adequately address scenarios requiring high-precision synchronization of virtual reality content with externally played large-screen videos, dynamic interactive programs, or physical mechanical devices in real time. This is especially true in large exhibition environments such as museums, where exhibits are diverse (videos, interactive, mechanical) and numerous, and external content may be frequently updated. Traditional synchronization methods either have low reliability or are too costly to modify, making them impractical. Summary of the Invention

[0006] The purpose of this invention is to provide an asynchronous concurrent virtual reality synchronization method and system to solve the problems mentioned in the background art, such as high modification costs, low synchronization reliability, and difficulty in accurately correcting the time axis with a fixed offset.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: According to one aspect of the present invention, an asynchronous concurrent virtual reality synchronization method is provided, the method comprising: The virtual reality content to be synchronized is divided into multiple sub-contents that are played concurrently, and each sub-content is configured with an independent timeline correction anchor point; Real-time acquisition of status data of external multimedia exhibits, including video exhibits, interactive exhibits, and mechanical device exhibits; Based on the real-time spatial location of the virtual reality display device, target exhibits that are within a preset distance threshold for each exhibit are selected; The status data of the selected target exhibits is frame-synchronized with the timeline correction anchor points of the corresponding sub-contents to generate synchronization data. The virtual reality display device renders and synthesizes the multiple sub-contents in real time based on the synchronized data, and outputs virtual reality images synchronized with external multimedia exhibits.

[0008] Based on the aforementioned scheme, the step of splitting the virtual reality content to be synchronized into multiple concurrently played sub-contents includes: splitting according to the visual units, event triggering nodes, or independent presentation segments on the timeline that naturally appear in the virtual reality content. Each sub-content is an independent three-dimensional content unit with an independent spatial location, duration, and playback progress.

[0009] Based on the aforementioned scheme, configuring an independent timeline correction anchor point for each sub-content includes: assigning an independent playback controller to each sub-content, wherein the playback controller maintains a local time offset and independently receives correction parameters from the synchronization server, and adjusts the playback start time or playback progress of the sub-content according to the correction parameters, and the correction processes of different sub-contents do not block each other.

[0010] Based on the aforementioned solution, the real-time acquisition of status data of external multimedia exhibits includes: For the video item, obtain the current playback timestamp at the start of video playback or the start of each loop cycle. For the interactive display items, an event-triggered method is adopted to obtain the page number and event parameters when the user operation causes the page to switch or the component to be triggered; For the mechanical device exhibit, the position array and angle array are obtained in real time through built-in sensors, or the device posture is obtained through image recognition.

[0011] Based on the aforementioned scheme, the step of filtering out target exhibits that fall within a preset distance threshold for each exhibit includes: Pre-configure anchor point coordinates and independent distance thresholds for each exhibit in the global coordinate system; The spatial coordinates of the virtual reality display device are acquired in real time, and the distance between the spatial coordinates and the anchor point coordinates of each exhibit is calculated. If the distance is less than or equal to the distance threshold of the corresponding exhibit, then the corresponding exhibit is determined to be the target exhibit.

[0012] Based on the aforementioned scheme, the generation of synchronization data includes: When the target item is a video item, synchronous data for timeline alignment is generated, and the current timestamp of the video is sent to the correction anchor point of the corresponding sub-content to calibrate the playback progress of the sub-content. When the target exhibit is an interactive exhibit, synchronous data for state switching operation is generated, and the page number and event parameters are sent to the corresponding sub-content to switch the information or interface displayed by the sub-content. When the target exhibit is a mechanical device exhibit, synchronous data for transformation and update operations is generated, and the position array and angle array are sent to the corresponding sub-content to update the posture of the sub-content in virtual space.

[0013] Based on the aforementioned scheme, the synchronization data is encapsulated using a unified data structure. This data structure includes at least a timestamp, an item identifier, a playback control command, and a status payload. The content of the status payload varies depending on the item type and the playback control command. The data structure is serialized in JSON format and transmitted via the TCP protocol.

[0014] Based on the aforementioned scheme, when the virtual reality display device renders and synthesizes the multiple sub-contents in real time according to the synchronization data, if the synchronization data corresponding to a certain sub-content is lost or fails to arrive within a timeout period, the sub-content will continue to be rendered while maintaining the previous valid correction state, and the rendering of other sub-contents will not be affected; synchronization data lost during network transmission will not be retransmitted.

[0015] Based on the aforementioned scheme, the virtual reality display device renders and synthesizes the multiple sub-contents in real time according to the synchronized data, including: For video-synchronized sub-content, the playback controller adjusts the local playback progress according to the timestamp in the synchronization data, and maps the decoded video frames as texture maps onto the screen model or billboard model in the virtual scene. For interactive synchronization type sub-content, according to the page number and event parameters in the synchronization data, the corresponding 3D content is loaded from local storage or network cache, and the loaded 3D content is instantiated at a predetermined position in the scene; For the mechanical device synchronization type sub-content, the transformation values ​​of the corresponding bones or joints in the three-dimensional model are modified according to the position array and angle array in the synchronization data.

[0016] According to another aspect of the present invention, an asynchronous concurrent virtual reality synchronization system is provided, the system comprising: The content splitting module is used to split the virtual reality content to be synchronized into multiple sub-contents that are played concurrently, and to configure an independent timeline correction anchor point for each sub-content. The status acquisition module is used to acquire the status data of external multimedia exhibits in real time, including video exhibits, interactive exhibits, and mechanical device exhibits. The spatial filtering module is used to filter target exhibits that are within a preset distance threshold for each exhibit based on the real-time spatial location of the virtual reality display device. The frame synchronization module is used to synchronize the status data of the selected target items with the timeline correction anchor points of the corresponding sub-content to generate synchronization data. The rendering and compositing module, deployed in the virtual reality display device, is used to render and compose the multiple sub-contents in real time based on the synchronization data, and output virtual reality images synchronized with external multimedia exhibits.

[0017] As can be seen from the above technical solution, the present invention has at least the following advantages and positive effects compared with the prior art: The present invention splits a single linear timeline into multiple asynchronous concurrent sub-contents, each sub-content having an independent timeline correction anchor point; even if some sub-contents fail to correct due to network fluctuations or recognition errors, other sub-contents can still synchronize normally; under the same network packet loss rate, the probability of significant desynchronization of the overall picture is reduced from a single full-scene failure to a local minor desynchronization, which is difficult for the human eye to detect, thereby greatly improving the synchronization reliability of the system in real complex environments. By adopting the method of real-time synthesis and rendering of fragmented content, only one or a few corresponding concurrent sub-contents need to be replaced or regenerated, significantly shortening the modification time and greatly improving the iteration efficiency of exhibition content. By introducing an exhibit selection mechanism based on spatial location, each virtual reality headset only synchronizes target exhibits within a distance threshold range, avoiding data redundancy caused by full synchronization; combined with the server thread pool, each headset is allocated an independent processing thread, effectively reducing the communication and data processing pressure of a single device and achieving low-latency frame synchronization. Different synchronization strategies (timeline alignment, page state switching, and sensor change updates) are designed for video exhibits, interactive exhibits, and mechanical exhibits, and are encapsulated using a unified data structure to adapt to the synchronization needs of multiple types of exhibits coexisting in complex environments such as museums; especially for uncontrollable physical mechanical devices, the state is followed by real-time sensor data or image recognition to achieve precise alignment of virtual and real scenes.

[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 This is a schematic diagram of an asynchronous concurrent virtual reality synchronization method according to the present invention; Figure 2 This is a schematic diagram of the process for filtering target exhibits based on spatial location according to the present invention; Figure 3 This is a schematic diagram of the data flow of an asynchronous concurrent virtual reality synchronization system according to the present invention. Detailed Implementation

[0020] To more clearly illustrate the purpose, technical solutions, and advantages of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein. On the contrary, these embodiments are provided so that the present invention will be more comprehensive and complete, and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0021] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a full understanding of embodiments of the invention. However, those skilled in the art will recognize that the technical solutions of the invention can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the invention.

[0022] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0023] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0024] The present invention will now be described in detail with reference to specific embodiments.

[0025] Example 1

[0026] like Figure 1 , 2 As shown, this embodiment provides an asynchronous concurrent virtual reality synchronization method. In this embodiment, the virtual reality display device is a virtual reality headset, which has independent spatial positioning capabilities, real-time rendering capabilities, and network communication capabilities. It should be noted that the method described in this invention can also be applied to other types of virtual reality display devices (such as all-in-one VR glasses and external VR headsets). This embodiment uses a virtual reality headset as an example for description. The specific steps of the method are as follows: S1: Divide the virtual reality content to be synchronized into multiple sub-contents to be played concurrently, and configure an independent timeline correction anchor point for each sub-content.

[0027] In this embodiment, an asynchronous concurrent splitting operation is first performed on the original virtual reality content to achieve precise synchronization between the virtual reality content and external multimedia exhibits (such as large-screen videos, interactive programs, and mechanical devices). The splitting operation decomposes the complete virtual reality animation or scene logic into multiple independently playable and independently correctable sub-contents (hereinafter referred to as concurrent sub-contents) according to naturally occurring visual units, event triggering nodes, or independent presentation segments on the timeline. Each concurrent sub-content has a clear time start point, duration, playback progress, and spatial location attributes in the virtual reality space. Through asynchronous concurrent splitting and independent anchor point configuration, the overall synchronization accuracy increases with the number of sub-contents. Assuming the synchronization success probability of a single sub-content is p, the overall success probability of a traditional linear system is p, while the probability of all sub-contents failing simultaneously in this embodiment is p. (where n is the number of sub-contents), and the human eye has a high tolerance for local desynchronization in three-dimensional space, so visual synchronization is significantly better than that of a linear system. When the content of external exhibits is adjusted, only one or a few corresponding sub-contents need to be replaced or regenerated, without having to re-render all virtual reality content, thus greatly shortening the content iteration cycle.

[0028] Specifically, during the virtual reality content production stage, key change points in the original animation are identified through content editing tools or script configuration. These include: the switching moments between different 3D animation materials, user interaction response events, the start time of mechanical device movements, or specific frames appearing in the large-screen video that require synchronization. For example, a 3-minute virtual reality animation in a museum exhibition might consist of 20 independent 3D animation materials (such as rotating artifacts, emerging text, special effects particles, scene roaming, etc.) arranged chronologically and spatially. Traditional methods pre-render these 20 materials into a single video stream, allowing for global time synchronization only once at the start of playback. This embodiment abandons this pre-rendering method, instead treating each material as an independent concurrent sub-content, preserving its original data, animation parameters, and triggering conditions. To achieve this separation, an independent correction anchor point is generated for each concurrent sub-content. The correction anchor point is a timeline correction trigger, internally recording the expected absolute or relative timestamp of the sub-content's start, a synchronization reference signal corresponding to the external exhibit, and an independently operating local clock counter. During runtime, each concurrent sub-content does not depend on the global timeline. Instead, it receives correction instructions from the synchronization server based on its own correction anchor point, dynamically adjusting the playback progress to ensure that its start time is precisely aligned with the state of the corresponding external exhibit.

[0029] The implementation of the correction anchor point involves allocating an independent playback controller (e.g., a software thread or state machine) for each concurrent sub-content in the rendering pipeline of the virtual reality headset. This controller maintains a local time offset Δt_i. When a synchronization signal for that sub-content is received (e.g., external video plays to a specific frame, the interactive program switches to a specific page, or a mechanical sensor reaches a certain displacement value), the server sends a correction timestamp, the controller calculates Δt_i = T_server − T_local, and adjusts the subsequent playback progress according to this offset. Here, T_server is the correction timestamp issued by the synchronization server, representing the global reference time of the current playback moment of the external multimedia exhibit (such as a large-screen video). This timestamp is usually generated by the server based on the video frame sequence number, system clock, or external synchronization signal (such as NTP), and sent to the virtual reality headset via synchronization data packets. T_local is the local clock reading of the playback controller of this sub-content in the virtual reality headset at the instant it receives the synchronization signal. Δt_i is the calculated time offset. The actual playback time = local cumulative playback time + Δt_i. If Δt_i is positive, it means that the sub-content playback is lagging behind the external exhibit and needs to be fast-forwarded; if it is negative, it means that it is ahead and needs to be delayed or waited for. Through real-time adjustment of this offset, the timeline of the sub-content and the external exhibit are successively aligned. Since the correction of each sub-content is performed independently, the failure of correction of one sub-content due to network fluctuations or hardware latency will not affect the normal playback of other sub-contents.

[0030] In a preferred embodiment, all concurrent sub-contents are 3D animations. These 3D animations have different position, rotation, and scaling attributes in the 3D space rendered by the virtual reality headset. They are ultimately dynamically composited through the headset's real-time rendering pipeline to form a complete immersive scene. Since there is no traditional layer overlay concept, each sub-content is more independent, and the impact of local asynchrony on the overall visual experience is further reduced.

[0031] S2: Real-time acquisition of status data of external multimedia exhibits, including video exhibits, interactive exhibits, and mechanical device exhibits.

[0032] Real-time status data of various external exhibits is collected to achieve precise synchronization between virtual reality content and external multimedia exhibits. External multimedia exhibits include, but are not limited to: video playback exhibits, interactive touch exhibits, and mechanically driven exhibits. Different status data collection strategies are adopted based on the exhibit type, and the collection results are reported to the synchronization server using a unified data structure for subsequent frame synchronization.

[0033] For video playback exhibits, due to their linear and continuous timeline, status data only needs to be collected once at the start of video playback. The status data includes at least: a unique identifier for the exhibit, a timestamp of the current video playback (in milliseconds or frame number), and the video playback status (playing, paused, finished). For looping videos, status data is collected again at the beginning of each loop to eliminate accumulated clock drift over time. The status data for video exhibits is actively pushed to the synchronization server's status pool by the video player via a wired or wireless network interface, or obtained by the synchronization server through polling at a fixed frequency.

[0034] For interactive exhibits (such as touchscreens, motion-sensing interactive devices, and page-based multimedia navigation systems), their state changes are characterized by jumps and non-linear time. This embodiment uses a page numbering system as the synchronization benchmark, dividing the entire interactive program's interface structure hierarchically into pages, subpages, and interactive components within each page. Each identifiable interactive state corresponds to a unique page number and subpage number. State data is collected via event triggering. When a user performs actions such as clicking, swiping, or page navigation on an interactive exhibit, the interactive program immediately generates a state change event. The data structure of this event includes at least: a timestamp, program name, page number, trigger type (image, video, text description, etc.), and a type-specific number. To ensure the integrity and verifiability of data transmission, the above data is encapsulated in JSON format, and the receiver verifies data integrity through JSON parsing. The encapsulated state data is sent to the synchronization server in real-time via a TCP connection. After receiving the data, the server updates the latest state of the interactive exhibit to the state pool, overwriting the old state, thereby ensuring that the virtual reality client always obtains the currently valid interactive page information.

[0035] For dynamic mechanical devices driven by motors, cylinders, encoders, photoelectric sensors, etc., whose state changes are continuous and not entirely controlled by digital signals (e.g., random errors caused by physical inertia and load fluctuations), this embodiment employs two types of data acquisition methods: 1) Direct Sensor Data Acquisition Method. For mechanical devices with built-in encoders, angle sensors, displacement sensors, or distance sensors, the controller (such as a PLC or microcontroller) periodically reads the sensor values ​​and organizes the data into a predetermined format. Typical status data includes: a position array Position[] (used to describe multi-axis travel or spatial coordinates) and an angle array Rotation[] (used to describe multi-axis rotation). Each set of data is accompanied by a timestamp and device identifier, and is sent to a synchronization server via an industrial bus (such as Modbus or CAN) or Ethernet.

[0036] 2) Auxiliary Recognition Method. For mechanical devices lacking digital sensors, or in scenarios requiring precise positioning assistance, this invention utilizes a camera on a virtual reality headset or a separate image acquisition device to recognize QR codes or feature markers fixed on the mechanical device. The device number and spatial coordinates are obtained through QR code parsing, or the current posture of the device is determined through feature point matching. This recognition result is also converted into a unified state data structure and reported to a synchronization server. It should be noted that camera recognition serves as a supplement or backup to sensor data; in scenarios requiring high precision or low latency, built-in sensors are preferred.

[0037] All collected status data is converted into a general data structure, which includes at least: Time (a long integer timestamp, used to uniquely identify the data's time point and aid in sorting), ProgramName (a string used to distinguish different exhibits), and a payload field specific to the exhibit type. For video and interactive exhibits, the payload field contains the page number and trigger parameters; for mechanical devices, the payload field contains an array of positions and rotations. Data is serialized using JSON format before transmission and its integrity is verified at the receiving end. All status data is pushed to a state pool on the synchronization server for caching. The state pool maintains the latest status record for each exhibit according to its exhibit identifier and retains several recent historical states for synchronization correction.

[0038] S3: Based on the real-time spatial location of the virtual reality display device, filter out target exhibits that are within the preset distance threshold for each exhibit.

[0039] Because large venues such as museums contain numerous multimedia exhibits (video players, interactive programs, mechanical devices, etc.), sending the status data of all exhibits to every virtual reality headset would increase the communication load on the synchronization server and the data processing pressure on the headsets, leading to increased latency and decreased synchronization accuracy. Therefore, a location-based filtering mechanism is introduced, identifying only exhibits within a preset distance threshold range of the virtual reality headset as target exhibits and synchronizing their frames accordingly. This significantly reduces system resource consumption while ensuring a good user experience.

[0040] During the scene deployment phase, this embodiment pre-establishes a global spatial coordinate system. This coordinate system uses a Cartesian coordinate system (X, Y), with the scene ground or main activity plane as the reference. For multi-story building scenes, the floor height is automatically identified by the upper-level positioning system (such as an indoor positioning base station), eliminating the need to explicitly include the Z-axis coordinate in the exhibit location data. Each exhibit requiring synchronization is assigned a fixed anchor point coordinate in the coordinate system, corresponding to the exhibit's physical installation location or the optimal viewing position for the user. For example, the anchor point for a large-screen exhibit can be set at the ground directly in front of the center of the screen; interactive exhibits are set in the operating area in front of the touchscreen; and mechanical devices are set at the center point of their range of motion. The anchor point coordinates of all exhibits and their corresponding exhibit identifiers are pre-stored in the configuration database of the synchronization server.

[0041] The spatial position of the virtual reality headset is provided in real time by an external positioning system; this positioning system can be a laser positioning base station, an infrared optical positioning system, an ultra-wideband (UWB) positioning system, or a local position calculated by the headset itself through visual inertial odometry (VIO). This invention does not limit the specific positioning technology, but requires that the positioning system be able to output the headset's real-time coordinates (X, Y, θ) in the global coordinate system at a frequency of not less than 20Hz, where θ is the orientation angle (optional). Positioning data is transmitted from the positioning module on the headset to a synchronization server via a wired or low-latency wireless link, or broadcast from a separate positioning server to the synchronization server. Each position update is accompanied by a timestamp to indicate the freshness of the data.

[0042] Each exhibit has an independent trigger distance threshold R; this threshold is dynamically configured based on the exhibit's physical size, screen size, interaction range, and user attention. Specifically: for large screens or projection walls, where users can clearly view the content from a greater distance, the threshold R is set larger (e.g., 5-10 meters); for small LCD screens or close-range interactive devices, the threshold R is set smaller (e.g., 1-3 meters); for mechanical devices requiring precise operation or detailed displays, the threshold can be further reduced (e.g., 0.5-1.5 meters). The threshold parameters are stored in the exhibit configuration table on the synchronization server and can be adjusted online during operation via the management interface without requiring a system restart.

[0043] After receiving the latest position data from each virtual reality headset, the synchronization server performs the following filtering steps: 1) Iterate through all registered exhibits in the current scene and obtain the anchor point coordinates (x_e, y_e) and distance threshold R_e for each exhibit; 2) Calculate the Euclidean distance d=sqrt((x_h-x_e)^2+(y_h-y_e)^2) between the current position of the headset (x_h, y_h) and the anchor point of the exhibit; 3) If d≤R_e, the exhibit is determined to be the target exhibit of the current headset; otherwise, it is a non-target exhibit; 4) For the set of exhibits determined to be target exhibits, the synchronization server marks its state data (the latest value obtained from the state pool) as pending transmission and hands it over to the subsequent thread pool module for distribution to the corresponding headset.

[0044] The above filtering process is performed once for each helmet position update. Since the number of exhibits in a typical scenario is usually no more than 100, and the number of helmets is limited (e.g., a maximum of dozens per venue), the computational cost of traversal calculation is extremely low (less than 100 distance calculations per helmet per round), and no additional algorithm optimization is required. If the scene scale expands to hundreds of exhibits, spatial indexing (such as grid partitioning, quadtrees) can also be used to accelerate filtering.

[0045] This embodiment introduces a hysteresis mechanism to prevent frequent addition / removal of target exhibits (i.e., jitter) when the virtual reality headset moves slightly near the threshold boundary. Specifically, when the headset moves from outside the threshold to inside, the original threshold R_in = R is used; when the headset moves from inside the threshold to outside, a slightly larger exit threshold R_out = R + δ is used, where δ is the hysteresis margin (e.g., 0.2~0.5 meters). That is, when the virtual reality display device moves from a region where the distance from the exhibit's anchor point coordinates is less than or equal to the original distance threshold to a region where the distance is greater than the original distance threshold, an exit threshold greater than the original distance threshold is used as the exit judgment criterion to suppress jitter. An exhibit is only removed from the target set when the headset position satisfies d > R_out; otherwise, it remains a target. This hysteresis mechanism effectively avoids repeated switching caused by positioning noise or slight shaking, ensuring the continuity of synchronized data.

[0046] After the above screening process, each virtual reality headset receives a dynamically updated list of target exhibits. The synchronization server's thread pool allocates an independent thread for each headset. This thread retrieves the corresponding state data (video timestamps, interactive page numbers, mechanical device positions / angles, etc.) from the state pool only for exhibits in the list, packages it, and sends it to the headset. State data for exhibits not in the list is not sent, thus significantly saving network bandwidth and headset processing resources.

[0047] S4: Synchronize the status data of the selected target exhibits with the timeline correction anchor points of the corresponding sub-contents to generate synchronization data.

[0048] The target exhibit status data obtained after filtering in step S3 is further synchronized with the independent timeline correction anchor points of the corresponding concurrent sub-contents in the virtual reality headset to generate synchronization data that can be used for headset rendering. Frame synchronization adopts differentiated synchronization strategies based on exhibit type and sub-content attributes, converting information such as time, page, and position in the status data into playback control parameters for each concurrent sub-content.

[0049] The result of the frame synchronization operation is a structured data packet, which contains at least the following fields: target sub-content identifier, a unique ID of the concurrent sub-content corresponding to the exhibit status data; synchronization timestamp, the global time (in milliseconds) when the server generated the synchronization data; playback control instructions, including but not limited to "start," "jump to time point," "pause," "resume," and "update parameters"; and status payload, specific synchronization values ​​carried according to the exhibit type, such as video time points, page numbers, and arrays of mechanical device positions / angles. The synchronization data packet is encapsulated in JSON format and sent from the synchronization server to the corresponding virtual reality headset via a TCP connection.

[0050] For video playback exhibits, due to the linear and continuous nature of videos, and the fact that this embodiment has divided the virtual reality content into multiple concurrent sub-contents, each sub-content corresponds to a specific segment or keyframe in the video. The specific implementation of frame synchronization includes: a synchronization server receiving real-time status data of the video exhibit (current playback timestamp T_video and video identifier); the server determining the set of concurrent sub-contents corresponding to the video segment based on a preset mapping table, with each sub-content associated with a desired start time window [T_start, T_end]; when T_video falls into the start window of a sub-content, the server generates timeline-aligned synchronization data, where the playback control command is "align start," and the payload includes the T_video value; after receiving the synchronization data, the virtual reality headset locates the correction anchor point of the corresponding sub-content, forcibly setting the local playback clock of the sub-content to an offset synchronized with T_video; if the sub-content has not yet started, playback begins immediately; if it has already started playing, fast-forwarding or rewinding is adjusted based on the difference. For looped videos, the above process is repeated within each loop cycle to eliminate accumulated drift.

[0051] For interactive exhibits, state changes are triggered by user actions, exhibiting a non-continuous, event-driven characteristic. Frame synchronization does not require handling a continuous timeline; instead, it focuses on page state synchronization. Specifically, after receiving the state data of the interactive exhibit (program name, page number, trigger type, type-specific number, etc.), the synchronization server generates synchronization data for state switching operations. The playback control command is set to "switch page" or "trigger event," with the payload containing the complete page number and event parameters. Upon receiving this synchronization data, the virtual reality headset searches for the concurrent sub-content corresponding to the interactive program (usually a 3D UI panel, pop-up model, or explanatory animation). Based on the page number and event parameters, the headset loads the corresponding 3D content (such as video textures, image panels, or text description boxes) from its local resource library and attaches it to the preset display position in the virtual scene. If fine alignment of the model's spatial position is involved (such as combining QR code positioning), the transformation matrix of the sub-content is adjusted simultaneously. Since each interactive program corresponds to only one animation content, state changes only require switching the content resources of that animation, eliminating state conflicts between multiple interactive programs and thus requiring no arbitration logic.

[0052] For mechanical devices (such as rotating platforms, dynamic sculptures, and cylinder-driven models), their status data is reported in real time in the form of sensor values ​​(position array [], angle array []). The goal of frame synchronization is to ensure that the pose of the 3D model rendered in the virtual reality headset is consistent with that of the physical mechanical device. Specifically, after receiving the status data of the mechanical device, the synchronization server generates synchronization data for transformation update operations at a fixed frequency (e.g., 30Hz), and the playback control command is set to "update transformation". The load contains position arrays and angle arrays, with each array element corresponding to a degree of freedom value of a motion axis. After receiving the synchronization data, the virtual reality headset finds the concurrent sub-content bound to the mechanical device (i.e., the 3D model of the mechanical device and its animation controller). The headset directly assigns the position and angle values ​​from the synchronization data to the transformation components of the 3D model. If the motion of the mechanical device has a smooth interpolation requirement (such as avoiding jumps), the headset can use linear interpolation or cubic spline interpolation to generate a transition animation between two consecutive frames of synchronization data. For mechanical devices with built-in QR code positioning assistance, the synchronization data can also include spatial coordinate offsets to correct the absolute position of the model in the scene. This frame synchronization method does not rely on the prediction or modeling of the movement of mechanical devices, but completely follows the real state fed back by the sensors, thereby achieving real-time compensation for physical uncertainties (such as motor delay and load fluctuation).

[0053] Since a virtual reality headset may simultaneously receive synchronous data from multiple target exhibits, and each synchronous data corresponds to different concurrent sub-content, the headset needs to have multi-threading or asynchronous scheduling capabilities. Specifically, each concurrent sub-content has an independent playback controller and local clock. When synchronous data arrives, the data is dispatched to the corresponding controller based on the target sub-content identifier. Each controller performs correction operations independently without blocking each other. The rendering pipeline collects the latest state of all sub-content (position, angle, page resources, playback progress, etc.) in each frame and synthesizes the final image for output to the headset display.

[0054] In this embodiment, all synchronization data sent from the synchronization server to the virtual reality headset is encapsulated using a unified data structure to ensure consistency in parsing and processing of synchronization data corresponding to different types of exhibits. The data structure of the synchronization data includes at least the following fields: Timestamp (Time), a long integer representing the global time (in milliseconds) at which the synchronization data was generated, used for sorting, deduplication, and latency estimation at the receiving end; Exhibit Identifier (ProgramName), a string used to uniquely identify the exhibit (such as a video player, interactive program, or mechanical device) that generated the synchronization data; Playback Control Command (Command), a string or enumeration type indicating the type of operation to be performed by the receiving end. Depending on the exhibit type, this command may include, but is not limited to, timeline alignment operations, state switching operations, and transformation update operations; State Payload, a variable-structure data carrier whose specific content varies depending on the exhibit type and the playback control command. For example, for the timeline alignment operation of a video exhibit, the payload includes the current video timestamp; for the state switching operation of an interactive exhibit, the payload includes the page number, subpage number, trigger type, and number; for the transformation update operation of a mechanical device exhibit, the payload includes a position array and an angle array. To ensure the integrity and verifiability of data transmission, synchronized data is serialized in JSON format before transmission. JSON provides self-descriptiveness and cross-platform compatibility. Upon receiving the data, the receiving end (virtual reality headset) deserializes it using a JSON parser and can determine data corruption based on field integrity. If verification fails, the receiving end discards the packet without requesting retransmission.

[0055] The transport layer uses the TCP protocol to ensure that synchronized data arrives without loss or out-of-order. In the rare case of network interruption or server timeout, the headset does not request retransmission but discards the synchronized data, maintaining the previous valid correction state. The loss of a single synchronized data point only affects a short period of time in the corresponding sub-content or a single animation element, and will not cause a global misalignment of the entire virtual reality scene. Furthermore, since the human eye has limited ability to perceive localized, temporary desynchronization, the overall user experience remains at an acceptable level.

[0056] By using the frame synchronization method described above, the selected target exhibit status data is accurately mapped to the correction anchor points of each concurrent sub-content in the virtual reality headset, generating synchronization data that can be rendered in real time, thus completing a complete closed loop from external multimedia exhibit status acquisition to virtual content playback alignment.

[0057] S5: The virtual reality display device renders and synthesizes the multiple sub-contents in real time according to the synchronization data, and outputs a virtual reality screen synchronized with the external multimedia exhibits.

[0058] After receiving synchronization data from the synchronization server, the virtual reality headset performs real-time rendering and compositing operations based on the independent correction anchor points of each concurrent sub-content and its corresponding state payload, and finally outputs an immersive virtual reality screen that is synchronized with the external multimedia exhibits.

[0059] The virtual reality headset maintains an internal rendering manager that allocates independent rendering resources and a playback controller for each concurrent sub-content. Specifically, each concurrent sub-content is a complete 3D animation unit, including a 3D mesh model, skeletal animation, material textures, special effects particle systems, and transformation matrices (position, rotation, scaling). Unlike traditional 2D layers, these sub-contents have independent depth information and spatial positions in 3D space, and can intersect, occlude, or be side-by-side with each other. Each sub-content is bound to an independent playback controller; this controller maintains a local timeline offset Δt_i and adjusts Δt_i in real time or directly updates transformation parameters based on the synchronization data received in step S4 (such as "align start," "switch page," "update transformation," etc.). The controller's update frequency is decoupled from the helmet's rendering frame rate. Synchronization data may arrive at a lower frequency (e.g., 30Hz for mechanical devices, triggered by interactive events), while the rendering pipeline runs independently based on the helmet's refresh rate (e.g., 72Hz, 90Hz). Before each rendering frame begins, the controller calculates the playback progress of the current frame's sub-content based on the local clock and the most recent correction parameters (for video-type sub-content) or directly outputs the latest transformation matrix (for mechanical device-type sub-content).

[0060] Since all concurrent sub-contents are 3D animations and coexist in the same virtual scene, real-time compositing is performed using a standard 3D graphics rendering pipeline (such as those based on OpenGL, Vulkan, or Unreal Engine). The compositing process includes: the rendering manager organizing the scene graph according to the preset spatial positions and hierarchical relationships of the sub-contents (e.g., background objects in front, interactive panels behind), but without the fixed stacking order of traditional 2D layers; each sub-content independently contributes its geometry and materials. Based on the head pose of the virtual reality headset (provided by the positioning system or the headset's built-in IMU), the view matrix and projection matrix under the current viewpoint are calculated, transforming the 3D coordinates of all sub-contents to screen space. The graphics pipeline automatically handles the occlusion relationships between sub-contents using a depth buffer, eliminating the need for manual specification of the compositing order; if some pixels of a sub-content are occluded by other sub-contents, only the pixels closest to the viewpoint are displayed, thus ensuring realism. Each sub-content is drawn according to its independent animation progress and transformation matrix; for example, a rotating artifact model rotates in real-time according to the angle array in the synchronization data of the mechanical device; an interactively triggered pop-up model loads the corresponding texture and displays it according to the page number.

[0061] Because there is no concept of "layers" in a 3D scene, the rendering result of all sub-content is a natural blending at the pixel level, rather than a simple image overlay. This compositing method makes local desynchronization of individual sub-content (such as a 0.1-second delay in an animation at a distance) difficult for the human eye to notice, since visual attention is usually focused on the center of the scene or objects with significant movement.

[0062] Depending on the type of synchronized data, the headset executes targeted rendering logic, including: 1) Video Synchronization Sub-content. The playback controller adjusts the local playback progress based on the synchronization timestamp. During rendering, the decoded video frames are mapped as texture maps onto the screen model or billboard model in the virtual scene. If multiple video sub-contents are playing simultaneously, each updates its texture independently.

[0063] 2) Interactive Synchronous Sub-Content. The page number and event parameters in the synchronized data trigger resource loading. The headset retrieves the corresponding 3D content (such as a display board model with text descriptions, a video tutorial model, or a set of highlight effects) from local storage or network cache. After loading, the sub-content is instantiated at a predetermined location in the scene (usually within the user's visible range and without interfering with the main experience). Changes in the state of the interactive program do not cause scene reconstruction; only the corresponding sub-content resource is replaced or switched.

[0064] 3) Mechanical Device Synchronization Sub-content. After receiving the position and angle arrays, the rendering controller directly modifies the transformation values ​​of the corresponding bones or joints in the 3D model. To avoid model jitter caused by sensor data noise, low-pass filtering or moving average processing can be applied to continuously arriving synchronization data. For high-speed moving mechanical devices, predictive rendering can be used (e.g., estimating the intermediate pose of the current frame using data from the previous two frames).

[0065] Considering the potential for momentary network communication failures or loss of synchronization data, a degraded rendering mechanism is designed on the headset side. This includes: if a concurrent sub-content does not receive new synchronization data for an extended period (e.g., exceeding a set timeout threshold, such as 500 milliseconds), its playback controller pauses timeline correction and resumes playback based on the last valid correction parameters (for linear video sub-content) or maintains the last pose change (for mechanical device sub-content). When new synchronization data arrives, the controller performs smooth realignment based on the timestamp difference to avoid abrupt changes in the image. If multiple concurrent sub-contents fail simultaneously, the remaining normal sub-contents continue to render independently according to their respective anchor points, preventing a complete collapse of the overall scene. Because each sub-content occupies a limited area on the screen, users typically only perceive brief stuttering in localized areas without experiencing global dizziness.

[0066] After rendering and compositing all sub-contents, the graphics pipeline of the virtual reality headset outputs the generated left and right eye images to the headset's display screens respectively. The output screen includes synchronization features: alignment with the timeline of the external large screen video (achieved through correction anchor points of video-type sub-contents); consistency with the page state of the external interactive program (achieved through real-time switching of interactive sub-contents); and synchronization with the actual posture of the external mechanical device (achieved through real-time changes in sensor data).

[0067] When users watch using virtual reality headsets, the virtual content they see is highly consistent with the multimedia exhibits in the physical exhibition hall in terms of time, space, and behavior, thus achieving an immersive mixed reality experience.

[0068] Example 2

[0069] like Figure 3 As shown, this embodiment exemplifies an asynchronous concurrent virtual reality synchronization system, including a content splitting module, a state acquisition module, a spatial filtering module, a frame synchronization module, and a rendering and compositing module. In this embodiment, the virtual reality display device is a virtual reality headset, which has independent spatial positioning capabilities, real-time rendering capabilities, and network communication capabilities. It should be noted that the method described in this invention can also be applied to other types of virtual reality display devices (such as all-in-one VR glasses and external VR headsets). This embodiment uses a virtual reality headset as an example for description.

[0070] The content splitting module, deployed on the virtual reality content creation platform or content management server, is used to split the virtual reality content to be synchronized into multiple concurrently played sub-contents and configure independent timeline correction anchors for each sub-content. This module includes a content parser and an anchor point configurator. The content parser reads the description file of the original virtual reality scene (such as a 3D animation timeline or event list) and splits it according to naturally occurring visual units (such as independent animation clips), event trigger nodes (such as a model appearing when the user clicks), or independent presentation segments on the timeline (such as a sequence of actions every 5 seconds). Each split sub-content is saved as an independent 3D content unit, containing its 3D model, materials, animation data, and spatial location information. The anchor point configurator generates an independent correction anchor point record for each sub-content, which includes the sub-content identifier, the expected trigger time window, and the port or channel information required for communication with the synchronization server. At runtime, the anchor point configurator loads the correction anchors into the playback controller of the virtual reality headset.

[0071] The status acquisition module, deployed on the synchronization server or as a standalone data acquisition gateway, is used to acquire real-time status data of external multimedia exhibits, including video exhibits, interactive exhibits, and mechanical device exhibits. This module contains multiple adapters, each corresponding to a different type of exhibit.

[0072] The video capture adapter connects to the video player in the exhibition hall and subscribes to playback status events via the player's SDK or a common control protocol (such as RS-232 or TCP / IP). At the start of video playback or the beginning of each loop cycle, the adapter obtains the current playback timestamp (in milliseconds or frame number) and assembles it into a status data packet.

[0073] An interactive data capture adapter connects to interactive applications (such as touchscreen applications) and listens for callback interfaces triggered by page transitions or components. When user actions cause changes in page number, subpage number, or trigger type, the adapter captures event data, including the application name, page number, trigger type and number, and appends the current timestamp.

[0074] The mechanical data acquisition adapter connects to the controller of the mechanical device (such as a PLC or microcontroller) and periodically reads sensor data via industrial bus (Modbus, CAN) or Ethernet. The sensor data includes at least a position array (e.g., multi-axis travel, spatial coordinates) and an angle array (e.g., rotation angle). If the mechanical device lacks digital sensors, the adapter can trigger a camera on the helmet to perform QR code or feature recognition to determine the device's posture.

[0075] All collected status data is formatted uniformly and then written into the status pool of the synchronization server.

[0076] The spatial filtering module, deployed on the synchronization server, is used to filter target exhibits that are within a preset distance threshold for each exhibit based on the real-time spatial location of the virtual reality display device. This module includes a location receiver, a distance calculator, and a filtering decision-maker.

[0077] A position receiver receives real-time spatial coordinates (X, Y, θ) from a virtual reality headset positioning system via a low-latency network. The positioning system can be laser positioning, infrared optical positioning, UWB, or the headset's own visual inertial odometry; the position update frequency is no less than 20Hz.

[0078] The distance calculator reads the anchor point coordinates (pre-set in the global plane coordinate system) of all exhibits from the server configuration database, along with a distance threshold configured independently for each exhibit. For large screen exhibits, the threshold is set to 5-10 meters; for small interactive devices, the threshold is set to 1-3 meters. The calculator calculates the Euclidean distance between the helmet coordinates and the anchor point of each exhibit. The selection decision-maker determines whether each exhibit satisfies the condition that distance ≤ threshold; if so, it marks the exhibit as a target exhibit. To suppress helmet jitter near the threshold boundary, the selection decision-maker introduces a hysteresis mechanism: when the helmet moves from the inner area to the outer area, it must satisfy the condition that distance > (threshold + δ) before moving out of the target set, where δ is the hysteresis margin (e.g., 0.2-0.5 meters). The selection results are output to the frame synchronization module in the form of a target exhibit list.

[0079] The frame synchronization module, deployed in the synchronization server, is used to synchronize the status data of the selected target exhibits with the timeline correction anchor points of the corresponding sub-contents, generating synchronization data. This module includes a synchronization data generator and a data distributor.

[0080] The synchronization data generator generates synchronization data packets for each target exhibit, based on its exhibit type and corresponding operation type. Specifically: for video exhibits, it generates synchronization data for timeline alignment operations, with the payload containing the current video timestamp; for interactive exhibits, it generates synchronization data for state switching operations, with the payload containing the page number, trigger type, and number; for mechanical device exhibits, it generates synchronization data for transformation update operations, with the payload containing position and angle arrays. Each synchronization data packet uses a unified data structure, containing at least a timestamp (long integer), exhibit identifier (string), playback control instructions (enumeration or string), and state payload (variable structure). Data packets are serialized in JSON format before transmission to ensure verifiability.

[0081] The data distributor allocates an independent data sending thread to each connected virtual reality headset. Based on the target exhibit list output by the spatial filtering module, the distributor retrieves the latest state data of the corresponding exhibit from the state pool, calls the synchronization data generator to construct data packets, and sends them to the rendering and compositing module of the corresponding headset via the TCP protocol. If multiple target exhibits on the same headset are updated simultaneously, the distributor can merge multiple data packets or send them sequentially, without changing the independence of each sub-content.

[0082] The rendering and compositing module, deployed within the virtual reality display device, is used to render and composite the multiple sub-contents in real time based on synchronized data, outputting virtual reality visuals synchronized with external multimedia exhibits. This module includes multiple playback controllers, a resource manager, a 3D graphics rendering pipeline, and an output manager.

[0083] Each concurrent sub-content has its own independent playback controller. The controller maintains a local time offset Δt_i and listens for data packets from the frame synchronization module. Upon receiving synchronization data, the controller parses the operation type and payload: if it's a timeline alignment operation, it calculates Δt_i = T_server - T_local and adjusts the sub-content's local playback clock; if it's a state switching operation, it calls the resource manager to load the corresponding 3D content (such as textured panels, video models, or effects) and instantiates it in a preset position in the scene; if it's a transformation update operation, it directly modifies the position / angle values ​​in the skeleton or joint matrix of the sub-content's corresponding 3D model.

[0084] The resource manager is responsible for asynchronously loading 3D models, textures, videos, and other resources from the helmet's local storage or network cache, avoiding blocking the main rendering thread.

[0085] A 3D graphics rendering pipeline, implemented using OpenGL, Vulkan, or Unreal Engine. Before each rendering frame begins, it iterates through all playback controllers, obtaining the transformation matrix, animation progress, and texture data for each sub-content's current frame. For video texture sub-content, the pipeline maps video frames to the screen model in real time; for interactive sub-content, the pipeline places the loaded 3D model at a specified spatial location; for mechanical device sub-content, the pipeline directly applies the transformation matrix. All sub-content is seamlessly blended through depth testing, without any artificial layer overlay.

[0086] The output manager outputs the left and right eye images generated by the rendering pipeline to the helmet display, with the refresh rate synchronized with the helmet hardware. In the output screen, the spatiotemporal state of each sub-content is aligned with the external multimedia exhibits.

[0087] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims. It should be understood that the invention is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. An asynchronous concurrent virtual reality synchronization method, characterized in that, The method includes: The virtual reality content to be synchronized is divided into multiple sub-contents that are played concurrently, and each sub-content is configured with an independent timeline correction anchor point; Real-time acquisition of status data of external multimedia exhibits, including video exhibits, interactive exhibits, and mechanical device exhibits; Based on the real-time spatial location of the virtual reality display device, target exhibits that are within a preset distance threshold for each exhibit are selected; The status data of the selected target exhibits is frame-synchronized with the timeline correction anchor points of the corresponding sub-contents to generate synchronization data. The virtual reality display device renders and synthesizes the multiple sub-contents in real time based on the synchronization data, and outputs a virtual reality image synchronized with external multimedia exhibits; wherein, the virtual reality display device renders and synthesizes the multiple sub-contents in real time based on the synchronization data, including: For video-synchronized sub-content, the playback controller adjusts the local playback progress according to the timestamp in the synchronization data, and maps the decoded video frames as texture maps onto the screen model or billboard model in the virtual scene. For interactive synchronization type sub-content, according to the page number and event parameters in the synchronization data, the corresponding 3D content is loaded from local storage or network cache, and the loaded 3D content is instantiated at a predetermined position in the scene; For the mechanical device synchronization type sub-content, the transformation values ​​of the corresponding bones or joints in the three-dimensional model are modified according to the position array and angle array in the synchronization data.

2. The asynchronous concurrent virtual reality synchronization method according to claim 1, characterized in that, The process of splitting the virtual reality content to be synchronized into multiple concurrently played sub-contents includes: splitting the virtual reality content into visually independent three-dimensional content segments, event trigger nodes, or independently presented segments on the timeline. Each sub-content is an independent three-dimensional content unit with an independent spatial location, duration, and playback progress.

3. The asynchronous concurrent virtual reality synchronization method according to claim 1, characterized in that, The step of configuring an independent timeline correction anchor point for each sub-content includes: assigning an independent playback controller to each sub-content, wherein the playback controller maintains a local time offset and independently receives correction parameters from the synchronization server, and adjusts the playback start time or playback progress of the sub-content according to the correction parameters, and the correction processes of different sub-contents do not block each other.

4. The asynchronous concurrent virtual reality synchronization method according to claim 1, characterized in that, The real-time acquisition of status data for external multimedia exhibits includes: For the video item, obtain the current playback timestamp at the start of video playback or the start of each loop cycle. For the interactive display items, an event-triggered method is adopted to obtain the page number and event parameters when the user operation causes the page to switch or the component to be triggered; For the mechanical device exhibit, the position array and angle array are obtained in real time through built-in sensors, or the device posture is obtained through image recognition.

5. The asynchronous concurrent virtual reality synchronization method according to claim 1, characterized in that, The process of filtering out target exhibits that fall within a preset distance threshold for each exhibit includes: Pre-configure anchor point coordinates and independent distance thresholds for each exhibit in the global coordinate system; The spatial coordinates of the virtual reality display device are acquired in real time, and the distance between the spatial coordinates and the anchor point coordinates of each exhibit is calculated. If the distance is less than or equal to the distance threshold of the corresponding exhibit, then the corresponding exhibit is determined to be the target exhibit.

6. The asynchronous concurrent virtual reality synchronization method according to claim 1, characterized in that, The generation of synchronization data includes: When the target item is a video item, synchronous data for timeline alignment is generated, and the current timestamp of the video is sent to the correction anchor point of the corresponding sub-content to calibrate the playback progress of the sub-content. When the target exhibit is an interactive exhibit, synchronous data for state switching operation is generated, and the page number and event parameters are sent to the corresponding sub-content to switch the information or interface displayed by the sub-content. When the target exhibit is a mechanical device exhibit, synchronous data for transformation and update operations is generated, and the position array and angle array are sent to the corresponding sub-content to update the posture of the sub-content in virtual space.

7. The asynchronous concurrent virtual reality synchronization method according to claim 1, characterized in that, The synchronized data is encapsulated using a unified data structure, which includes at least a timestamp, an item identifier, a playback control command, and a status payload. The content of the status payload varies depending on the item type and the playback control command. The data structure is serialized in JSON format and transmitted via the TCP protocol.

8. The asynchronous concurrent virtual reality synchronization method according to claim 1, characterized in that, When the virtual reality display device renders and synthesizes the multiple sub-contents in real time based on the synchronization data, if the synchronization data corresponding to a certain sub-content is lost or fails to arrive within a timeout period, the sub-content will continue to be rendered while maintaining the previous valid correction state, and the rendering of other sub-contents will not be affected; synchronization data lost during network transmission will not be retransmitted.

9. An asynchronous concurrent virtual reality synchronization system for implementing the method as described in any one of claims 1-8, characterized in that, include: The content splitting module is used to split the virtual reality content to be synchronized into multiple sub-contents that are played concurrently, and to configure an independent timeline correction anchor point for each sub-content. The status acquisition module is used to acquire the status data of external multimedia exhibits in real time, including video exhibits, interactive exhibits, and mechanical device exhibits. The spatial filtering module is used to filter target exhibits that are within a preset distance threshold for each exhibit based on the real-time spatial location of the virtual reality display device. The frame synchronization module is used to synchronize the status data of the selected target items with the timeline correction anchor points of the corresponding sub-content to generate synchronization data. The rendering and compositing module, deployed in the virtual reality display device, is used to render and compose the multiple sub-contents in real time based on the synchronization data, and output virtual reality images synchronized with external multimedia exhibits.

Citation Information

Patent Citations

  • Immersive scene experience system and method based on mixed reality superposition projection

    CN121614031A

  • System for introducing physical experiences into virtual reality (VR) worlds

    US10362299B1