Virtual reality video interaction method and device, electronic equipment and storage medium
By establishing a video display sphere and an interaction trigger body associated with a head-mounted display device in a virtual reality scene, accurate generation of interactive feedback is achieved when the user's head moves, solving the interaction dislocation problem caused by a fixed coordinate system in the existing technology, and improving the immersion and naturalness of virtual reality applications.
Patent Information
- Application Number
- CN202511050147.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing technology, when the user's head moves and the viewing angle changes, the collision body of the fixed coordinate system cannot dynamically adjust to the video screen, resulting in interaction misalignment and inaccurate feedback.
In the virtual reality scene, a video display sphere associated with the spatial position of the head-mounted display device is established, an interaction trigger body is set, and the relative posture synchronization between the interaction trigger body and the video display sphere is ensured through real-time data transmission and transformation matrix, so as to achieve accurate generation of interactive feedback.
It solves the interaction misalignment problem caused by the fixed coordinate system, ensures that the interaction trigger body moves synchronously with the video display sphere, and provides accurate interaction feedback and a stable immersive experience.
Smart Images

Figure CN120711244A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of virtual reality technology, and in particular to a virtual reality video interaction method, device, electronic device, and storage medium. Background Art
[0002] Virtual Reality (VR) technology is widely used in gaming, education, simulation training, and other fields. VR applications based on panoramic video can provide a highly realistic immersive experience. However, traditional panoramic videos are typically generated using pre-rendering, resulting in static video streams that cannot support real-time user interaction with objects in the video. Therefore, how to enable dynamic user interaction with objects in pre-rendered videos without compromising video quality has become a pressing technical challenge in the VR field.
[0003] The existing technology mainly adopts the solution of superimposing video and real-time rendering objects, specifically: placing a transparent collision body at the position of the interactive object in the video screen to detect user operations, and binding the position of the collision body through a fixed coordinate system.
[0004] However, with this existing solution, when the user's head moves, causing the perspective to change, the collision body in the fixed coordinate system cannot dynamically adjust to the video image, causing the visual position of the collision body and the object in the video to gradually deviate, resulting in interaction misalignment and inaccurate feedback. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide a virtual reality video interaction method, device, electronic device, and storage medium to address the prior art problem that when the user's head moves and the viewing angle changes, the collision body in a fixed coordinate system cannot dynamically adjust to the video image. The specific technical solution is as follows:
[0006] In a first aspect, the present application provides a virtual reality video interaction method, comprising:
[0007] Establishing a video display sphere associated with a spatial position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a pre-rendered panoramic video and maintain a fixed image reference orientation;
[0008] An interaction trigger body corresponding to the interactive object in the panoramic video is set in the coordinate system corresponding to the video display sphere, and the coordinate system corresponding to the video display sphere is axially aligned with the virtual camera coordinate system when the panoramic video is rendered, and the relative posture of the interaction trigger body and the video display sphere is kept synchronized, so that the interaction trigger body maintains a fixed posture in the video display sphere coordinate system;
[0009] When an interaction event is detected between a user interaction operation and the interaction triggering body, interaction feedback is generated in an area corresponding to the visual position of the interaction object in the panoramic video screen.
[0010] In one possible implementation, establishing a video display sphere associated with a spatial position of a head-mounted display device in a virtual reality scene includes:
[0011] Establishing a real-time data transmission channel between the center point coordinates of the video display sphere and the spatial positioning coordinates of the head-mounted display device in the virtual reality scene, so that the center point coordinates of the video display sphere and the spatial positioning coordinates of the user's head-mounted display device are kept in real time synchronization through the real-time data transmission channel;
[0012] The screen reference orientation parameter of the video display sphere is set to the initial rendering viewing angle direction value of the panoramic video, so that the position of the video display sphere changes with the movement of the user's head and the screen orientation is fixed and does not change with the rotation of the user's head.
[0013] In a possible implementation, setting an interaction trigger body corresponding to an interactive object in the panoramic video in a coordinate system corresponding to the video display sphere includes:
[0014] Obtaining spatial transformation parameters of the interactive object in a virtual camera coordinate system from the rendering data of the panoramic video;
[0015] Based on the spatial transformation parameters, constructing a physical interactive body with a collision detection function in a coordinate system corresponding to the video display sphere;
[0016] The visualization rendering property of the physical interaction body is set to a closed state, so that the physical interaction body is invisible in the virtual reality scene, thereby completing the setting of the interaction trigger body.
[0017] In one possible implementation, the step of maintaining synchronization between the relative positions of the interaction triggering body and the video display sphere includes:
[0018] Calculating in real time the transformation matrix from the coordinate system corresponding to the video display sphere to the virtual reality world coordinate system;
[0019] The transformation matrix is applied to the pose calculation of the interaction triggering body, so that the spatial position of the interaction triggering body in the virtual reality scene and the visual position of the interactive object in the camera coordinate system continue to coincide.
[0020] In one possible implementation, the method further includes:
[0021] Monitoring the positional overlap deviation between the interaction triggering body and the interaction object;
[0022] When the position coincidence deviation exceeds a preset threshold, the posture synchronization process is paused, and the reference posture of the interaction triggering body is corrected according to the actual display position of the interactive object in the current video frame; and the posture synchronization process is restarted after the correction.
[0023] In one possible implementation, the method further includes:
[0024] Detecting the collision depth between the collision body of the interactive controller and the interactive trigger body in real time;
[0025] When it is detected that the collision depth exceeds a contact threshold, it is determined that an interaction event occurs between the user interaction operation and the interaction triggering body.
[0026] In one possible implementation, generating interactive feedback in an area corresponding to a visual position of an interactive object in the panoramic video includes:
[0027] Determining the interactive object type and collision force vector parameters corresponding to the interactive event;
[0028] Determining a visual special effects template corresponding to the interactive object type in a special effects resource library;
[0029] In the fragment shader stage of the video rendering pipeline, injecting the visual effect template into the texture space corresponding to the interaction event coordinates;
[0030] A three-dimensional sound effect corresponding to the collision force vector parameter is generated through an independent spatial audio processing thread.
[0031] In a second aspect, the present application provides a virtual reality video interaction device, comprising:
[0032] An establishment module, configured to establish a video display sphere associated with a spatial position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a pre-rendered panoramic video and maintain a fixed image reference orientation;
[0033] A setting module is used to set an interaction trigger body corresponding to the interactive object in the panoramic video in the coordinate system corresponding to the video display sphere, and the coordinate system corresponding to the video display sphere is axially aligned with the virtual camera coordinate system when the panoramic video is rendered, and to keep the relative posture of the interaction trigger body and the video display sphere synchronized, so that the interaction trigger body maintains a fixed posture in the video display sphere coordinate system;
[0034] The generating module is configured to generate interactive feedback in an area corresponding to a visual position of an interactive object in the panoramic video screen when an interactive event is detected between a user interactive operation and the interactive triggering body.
[0035] In one possible implementation, the establishment module is specifically configured to:
[0036] Establishing a real-time data transmission channel between the center point coordinates of the video display sphere and the spatial positioning coordinates of the head-mounted display device in the virtual reality scene, so that the center point coordinates of the video display sphere and the spatial positioning coordinates of the user's head-mounted display device are kept in real time synchronization through the real-time data transmission channel;
[0037] The screen reference orientation parameter of the video display sphere is set to the initial rendering viewing angle direction value of the panoramic video, so that the position of the video display sphere changes with the movement of the user's head and the screen orientation is fixed and does not change with the rotation of the user's head.
[0038] In one possible implementation, the setting module is specifically configured to:
[0039] Obtaining spatial transformation parameters of the interactive object in a virtual camera coordinate system from the rendering data of the panoramic video;
[0040] Based on the spatial transformation parameters, constructing a physical interactive body with a collision detection function in a coordinate system corresponding to the video display sphere;
[0041] The visualization rendering property of the physical interaction body is set to a closed state, so that the physical interaction body is invisible in the virtual reality scene, thereby completing the setting of the interaction trigger body.
[0042] In a possible implementation, the setting module is further configured to:
[0043] Calculating in real time the transformation matrix from the coordinate system corresponding to the video display sphere to the virtual reality world coordinate system;
[0044] The transformation matrix is applied to the pose calculation of the interaction triggering body, so that the spatial position of the interaction triggering body in the virtual reality scene and the visual position of the interactive object in the camera coordinate system continue to coincide.
[0045] In one possible implementation, the device further includes a correction module, configured to:
[0046] Monitoring the positional overlap deviation between the interaction triggering body and the interaction object;
[0047] When the position coincidence deviation exceeds a preset threshold, the posture synchronization process is paused, and the reference posture of the interaction triggering body is corrected according to the actual display position of the interactive object in the current video frame; and the posture synchronization process is restarted after the correction.
[0048] In one possible implementation, the apparatus further includes a determination module configured to:
[0049] Detecting the collision depth between the collision body of the interactive controller and the interactive trigger body in real time;
[0050] When it is detected that the collision depth exceeds a contact threshold, it is determined that an interaction event occurs between the user interaction operation and the interaction triggering body.
[0051] In one possible implementation, the generating module is specifically configured to:
[0052] Determining the interactive object type and collision force vector parameters corresponding to the interactive event;
[0053] Determining a visual special effects template corresponding to the interactive object type in a special effects resource library;
[0054] In the fragment shader stage of the video rendering pipeline, injecting the visual effect template into the texture space corresponding to the interaction event coordinates;
[0055] A three-dimensional sound effect corresponding to the collision force vector parameter is generated through an independent spatial audio processing thread.
[0056] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0057] Memory for storing computer programs;
[0058] The processor is configured to implement any of the method steps described in the first aspect when executing a program stored in the memory.
[0059] In a fourth aspect, a computer-readable storage medium is provided, characterized in that a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, any method step described in the first aspect is implemented.
[0060] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the above-described virtual reality video interaction methods.
[0061] Beneficial effects of the embodiments of the present application:
[0062] The embodiments of the present application provide a virtual reality video interaction method, device, electronic device and storage medium. In the present application, first, a spatial position association between a video display sphere and a head-mounted display device is established and the reference orientation of the screen is fixed, ensuring that the position of the video display sphere changes with the movement of the user's head but the orientation of the screen remains fixed; secondly, an interaction trigger body is set in the coordinate system corresponding to the video display sphere and its relative posture with the video display sphere is synchronized, so that the interaction trigger body always coincides with the visual position of the interactive object in the video; finally, feedback with precise position correspondence is generated when an interaction event is detected. This solution solves the interaction misalignment problem caused by the fixed coordinate system in the prior art by binding the interaction trigger body to the video display sphere coordinate system rather than the fixed world coordinate system and establishing a relative posture synchronization relationship.
[0063] Of course, it is not necessary to achieve all the advantages described above at the same time when implementing any product or method of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0065] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0066] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0067] Figure 1 A flowchart of a virtual reality video interaction method provided in an embodiment of the present application;
[0068] Figure 2 A flowchart of another virtual reality video interaction method provided in an embodiment of the present application;
[0069] Figure 3 A flowchart of another virtual reality video interaction method provided in an embodiment of the present application;
[0070] Figure 4 A flowchart of another virtual reality video interaction method provided in an embodiment of the present application;
[0071] Figure 5 A schematic diagram of the structure of a virtual reality video interaction device provided in an embodiment of the present application;
[0072] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0073] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0074] The disclosure below provides many different embodiments or examples for implementing different configurations of the present invention. To simplify the disclosure of the present invention, the components and configurations of specific examples are described below. Of course, these are merely examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or configurations discussed.
[0075] Figure 1 A flowchart of a virtual reality video interaction method provided in an embodiment of the present application. This method can be applied to one or more electronic devices such as smart phones, laptops, desktop computers, portable computers, servers, etc. In addition, the execution subject of this method can be hardware or software. When the above-mentioned execution subject is hardware, the execution subject can be one or more of the above-mentioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the above-mentioned execution subject is software, this method can be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.
[0076] like Figure 1 As shown, the method specifically includes:
[0077] S101. Establish a video display sphere associated with a spatial position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a pre-rendered panoramic video and maintain a fixed picture reference orientation.
[0078] Head-mounted display device: refers to a virtual reality headset with spatial positioning capabilities, such as the HTC Vive or Oculus Quest, which can track the position and rotation data of the user's head in real time. Video display sphere: A three-dimensional spherical model centered on the headset, with a panoramic video texture attached to its inner surface. The radius setting must ensure that it fully covers the user's field of view. Pre-rendered panoramic video: A 360-degree surround video sequence pre-generated by 3D rendering software, containing complete spherical texture mapping information. Fixed image reference orientation: refers to the default viewing direction determined during the initial rendering of the panoramic video.
[0079] In an embodiment of the present application, S101 may include the following steps: establishing a real-time data transmission channel between the center point coordinates of the video display sphere and the spatial positioning coordinates of the head-mounted display device in a virtual reality scene, and through the real-time data transmission channel, keeping the center point coordinates of the video display sphere and the spatial positioning coordinates of the user's head-mounted display device synchronized in real time; setting the picture reference orientation parameter of the video display sphere to the initial rendering viewing angle direction value of the panoramic video, so that the position of the video display sphere changes with the movement of the user's head, and the picture orientation is fixed and does not change with the rotation of the user's head.
[0080] Real-time data transmission channel: refers to the high-frequency data communication link established at the bottom layer of the system, which is used to transmit the position and orientation data of the head-mounted display device in real time; spatial positioning coordinates: contains complete positioning data of the three-dimensional spatial position and rotation direction, which is calculated by the built-in sensors of the device and the external positioning device; picture reference orientation parameters: records the reference value of the initial viewing angle direction of the video, using the spatial angle representation method; initial rendering viewing angle direction value: the starting shooting direction parameter of the virtual camera when producing panoramic video, which is saved in the video project file.
[0081] The solution first establishes a dedicated data transmission link within the virtual reality system, continuously transmitting the spatial positioning information of the head-mounted display device to the video display sphere, ensuring that the center point of the sphere always remains consistent with the device's position. Simultaneously, the initial shooting direction parameters are read from the video file and set as the fixed display direction of the video sphere. In this way, the video sphere changes position as the user's head moves, but does not change the screen orientation as the head rotates. This solution can accurately maintain the spatial correspondence between the video content and the user, effectively preventing screen misalignment caused by head movement. By fixing the screen orientation, viewing stability is guaranteed and an accurate spatial reference is provided for subsequent interactive functions. At the same time, this processing method can also reduce the system's computing burden and improve overall operating efficiency.
[0082] S102. An interaction trigger body corresponding to the interactive object in the panoramic video is set in the coordinate system corresponding to the video display sphere, and the coordinate system corresponding to the video display sphere is axially aligned with the virtual camera coordinate system during panoramic video rendering, and the relative posture of the interaction trigger body and the video display sphere is kept synchronized, so that the interaction trigger body maintains a fixed posture in the video display sphere coordinate system.
[0083] Interaction trigger volume: A transparent collision volume set in the virtual scene, whose shape and position precisely match the interactive objects in the video. Virtual camera coordinate system: The spatial reference system used in panoramic video production, whose axial definition aligns with the coordinate system of the video display sphere. Relative pose synchronization: The calculation process that ensures the position and orientation parameters of the interaction trigger volume remain dynamically consistent with the video display sphere.
[0084] In an embodiment of the present application, first, a corresponding interaction trigger body is created in the three-dimensional coordinate system corresponding to the video display sphere based on the object information that needs to be interacted with in the panoramic video. By axially aligning the video display sphere coordinate system with the virtual camera coordinate system when the original video is rendered, it is ensured that the two spatial reference systems have completely consistent coordinate definitions. On this basis, the interaction trigger body maintains a fixed spatial position and direction in the video display sphere coordinate system through real-time posture calculation. This solution achieves the alignment of the virtual trigger body and the video content through precise spatial coordinate mapping and posture synchronization algorithm, solving the technical problem that pre-rendered videos cannot interact directly.
[0085] S103: When an interaction event is detected between the user interaction operation and the interaction triggering body, an interaction feedback is generated in an area corresponding to the visual position of the interaction object in the panoramic video screen.
[0086] User interaction: refers to interactive commands input by users through VR controllers or gesture recognition. Interaction events: system responses triggered when a user's action effectively contacts an interaction trigger or enters a preset interaction range. Interaction feedback: the visual response effects and corresponding sound effects or vibrations generated by the system in real-time rendering.
[0087] In an embodiment of the present application, the contact state between the user input device and the interactive trigger body is first monitored in real time. When an effective interactive behavior that meets the preset conditions is detected, the system will generate corresponding feedback at the corresponding object position in the panoramic video screen according to the type of interactive object and the interaction intensity. The position calculation of this feedback is based on the coordinate mapping relationship between the interactive trigger body and the video display sphere, ensuring that the feedback effect can be accurately presented at the position overlapping with the interactive object in the video from the player's perspective. For example, there is a Great Wall in the video, and the interactive trigger body is aligned with its position and shape. When the player shoots an arrow and hits the interactive trigger body, a special effect of dust bursting, bullet marks, inserted arrows and other feedback effects will be displayed at the hit position. This breaks through the limitation that traditional panoramic videos can only be viewed passively, realizes rich interactive possibilities while maintaining video quality, and enhances the immersion and practicality of virtual reality applications.
[0088] Among them, whether an interaction event occurs can be determined by the following steps: real-time detection of the collision depth between the collision body of the interaction controller and the interaction trigger body; when it is detected that the collision depth exceeds the contact threshold, it is determined that an interaction event occurs between the user interaction operation and the interaction trigger body.
[0089] The collision volume of the interactive controller refers to the virtual collision detection area bound to the user's handheld controller, usually represented by a simplified geometric shape; collision depth is a quantitative value that describes the degree of mutual penetration between the controller collision volume and the interactive trigger volume; contact threshold is a pre-set critical value parameter used to determine whether the interactive behavior meets the effective triggering standard.
[0090] This solution accurately measures the degree of mutual penetration between the interactive controller collision body and the interactive trigger body by calculating the spatial position relationship between the two in real time. When the penetration depth is detected to exceed the preset critical value, the system determines that a valid interaction event has occurred. This judgment process runs continuously and can accurately capture various types of interactive actions, including touch, press and other interactive behaviors of varying intensities. In this way, the user's interactive intention can be accurately identified to avoid false triggering or missed triggering. The use of adjustable contact threshold parameters enables the system to adapt to the needs of different interactive scenarios, ensuring both the accuracy of interaction judgment and the response sensitivity of the system, providing a reliable trigger basis for subsequent interactive feedback.
[0091] In this application, first, a spatial position association is established between the video display sphere and the head-mounted display device, and the reference orientation of the screen is fixed, ensuring that the position of the video display sphere changes as the user's head moves but the orientation of the screen remains fixed; secondly, an interaction trigger body is set in the coordinate system corresponding to the video display sphere and its relative posture with the video display sphere is synchronized, so that the interaction trigger body always coincides with the visual position of the interactive object in the video; finally, feedback with precise position correspondence is generated when an interaction event is detected. This solution solves the interaction misalignment problem caused by the fixed coordinate system in the prior art by binding the interaction trigger body to the video display sphere coordinate system rather than the fixed world coordinate system and establishing a relative posture synchronization relationship.
[0092] See also Figure 2 , is a flow chart of another embodiment of a virtual reality video interaction method provided in an embodiment of the present application. Figure 2 The process shown in the above Figure 1 Based on the process shown in FIG, it is described how to set the interaction trigger body corresponding to the interactive object in the panoramic video in the coordinate system corresponding to the video display sphere. Figure 2 As shown, the process may include the following steps:
[0093] S201: Obtain spatial transformation parameters of the interactive object in a virtual camera coordinate system from rendering data of the panoramic video.
[0094] Rendering data: refers to the engineering file generated during the panoramic video production process that contains the spatial position information of scene objects; spatial transformation parameters: mathematical parameters that describe the position and rotation state of objects in three-dimensional space, usually including translation vectors and rotation matrices.
[0095] In this embodiment, the precise position and orientation data of the objects requiring interaction in the virtual camera coordinate system are extracted from the original engineering file of the panoramic video. This data records the position offset and rotation angle of the objects relative to the camera, providing the basic spatial information for the subsequent creation of the interaction trigger volume. By directly obtaining the object spatial information from the original rendering data, this solution ensures the accuracy of the interaction trigger volume position setting, laying the foundation for subsequent precise interaction.
[0096] Furthermore, when developers only have a panoramic video but lack the original rendering data, they can use a spatial pose estimation algorithm to obtain the pose parameters of the interactive object N. Specifically, this includes: 3D model reconstruction: creating a 3D model M in the virtual scene that matches the geometric features of the interactive object N; scene initialization: placing a video display sphere in the scene, loading the first frame of the video, and positioning the virtual camera A at the center of the sphere, with its rotation set to the initial rotation parameters of the panoramic video rendering camera (which can be obtained through video metadata parsing or manual specification); manual calibration: using the center of the sphere as the observation point (which can be rotated to any angle), the developer adjusts the pose of model M through the interactive device until, from the developer's perspective, the visual projection of M and the interactive object N in the video essentially overlap; pose calculation: recording the relative pose P of model M with respect to the virtual camera A as the baseline pose parameter of the interaction trigger body. This alternative solution effectively solves the problem of obtaining the pose of interactive objects when the original rendering project is lacking by combining 3D model reconstruction with manual calibration.
[0097] S202: Construct a physical interactive body with a collision detection function in a coordinate system corresponding to the video display sphere based on the space transformation parameters.
[0098] Physical interaction body: a collision detection area with physical characteristics created in the virtual scene; collision detection function: refers to the system's ability to detect contact and collision between virtual objects.
[0099] In this embodiment, a three-dimensional collision volume is created in the video display spherical coordinate system based on the acquired spatial transformation parameters, matching the shape of the interactive objects in the video. This collision volume is endowed with physical properties and responds to touch and collision with the controller. The physical interaction volume created by this solution accurately reproduces the spatial properties of the interactive objects in the video, providing a physical foundation for natural interaction.
[0100] S203: Setting the visualization rendering property of the physical interaction object to a closed state, so that the physical interaction object is invisible in the virtual reality scene, thereby completing the setting of the interaction triggering object.
[0101] Visual rendering properties: display parameters that control whether an object is visible in the scene; invisible state: refers to the state where an object exists in the scene but does not participate in the rendering of the picture.
[0102] In this embodiment, the visibility attribute of the created physical interactive object is set to invisible, so that it remains functional in the virtual scene but is not displayed, serving only as an invisible trigger area for interaction detection. By hiding the physical interactive object, this solution not only maintains the integrity of the video image, but also enables accurate interaction detection, enhancing the realism of the user experience. This complete technical solution achieves a seamless integration of video content and interactive functions.
[0103] Figure 2 The process shown here first ensures that the interaction triggering volume perfectly matches the position of the object in the video, achieving submillimeter alignment accuracy. Second, by setting up physical collision bodies, interaction detection can reflect customized effects such as physical or magical effects, enhancing the naturalness of interaction. Finally, by hiding the collision bodies, the integrity of the video image is maintained while enabling seamless interaction, resolving the technical challenge of traditional panoramic videos that prevent interaction. This solution provides users with a highly realistic interactive experience while ensuring system efficiency, offering reliable technical support for the development of VR applications based on pre-rendered video.
[0104] See also Figure 3 , is a flow chart of another embodiment of a virtual reality video interaction method provided in the embodiment of the present application. Figure 3 The process shown in the above Figure 1 Based on the process shown in the following, we describe how to achieve relative posture synchronization. Figure 3 As shown, the process may include the following steps:
[0105] S301, calculating in real time the transformation matrix from the coordinate system corresponding to the video display sphere to the virtual reality world coordinate system;
[0106] S302: Apply the transformation matrix to calculate the posture of the interaction triggering body, so that the spatial position of the interaction triggering body in the virtual reality scene and the visual position of the interactive object in the camera coordinate system continue to coincide.
[0107] For ease of understanding, steps S301-S302 are described in a unified manner below:
[0108] Virtual reality world coordinate system: the unified spatial reference system of the entire virtual scene, based on which the positions and orientations of all virtual objects are defined; camera coordinate system: the spatial reference system used when producing panoramic videos, which defines the original positional relationships of all objects in the video.
[0109] The solution first calculates the transformation relationship between the video display sphere coordinate system and the virtual reality world coordinate system through real-time calculation. This transformation relationship fully describes the relative position and orientation differences between the two coordinate systems. This transformation relationship is then applied to the position and orientation calculation of the interaction trigger volume, and the spatial state of the interaction trigger volume in the virtual scene is adjusted in real time through matrix operations. This process ensures that no matter how the user moves or rotates their head, the interaction trigger volume always accurately corresponds to the position of the interactive object in the video in the original camera coordinate system, achieving continuous visual overlap.
[0110] Figure 3The process shown here uses real-time calculation to transform the video display sphere coordinate system into the virtual reality world coordinate system. This transformation is then applied to the position and orientation of the interaction trigger volume. Matrix operations are then used to adjust the interaction trigger volume's spatial state within the virtual scene in real time. This process ensures that no matter how the user moves or rotates their head, the interaction trigger volume always accurately corresponds to the position of the interactive object in the video in the original camera coordinate system, achieving consistent visual overlap.
[0111] In addition, in another embodiment, the method may further include the following steps: monitoring the position overlap deviation between the interaction trigger body and the interactive object; when the position overlap deviation exceeds a preset threshold, pausing the posture synchronization process, and correcting the reference posture of the interaction trigger body according to the actual display position of the interactive object in the current video frame; and restarting the posture synchronization process after correction.
[0112] Position coincidence deviation: refers to the spatial position offset between the interactive trigger body and the corresponding interactive object in the video, calculated by the feature point matching algorithm; preset threshold: the maximum allowable offset value set according to the video resolution and interactive accuracy requirements, usually with pixel-level accuracy; posture correction: the process of dynamically adjusting the basic position and orientation parameters of the interactive trigger body.
[0113] In this embodiment of the present application, the spatial offset between the interaction trigger and the interactive object in the video frame is calculated by comparing their feature point positions in real time. If the offset exceeds a preset safety threshold, the system temporarily stops the pose synchronization process and recalculates and updates the baseline position and orientation parameters of the interaction trigger based on the actual display position of the interactive object in the current video frame. After the correction is completed, the system automatically resumes the pose synchronization process to ensure the accuracy of subsequent interactions.
[0114] This solution effectively addresses interaction position shifts caused by device errors or video distortion through a dynamic calibration mechanism, maintaining long-term stable interaction accuracy. This ensures a consistent interactive experience while improving the system's adaptability to diverse hardware devices and video content. Furthermore, the solution's adjustable threshold allows it to adapt to varying accuracy requirements.
[0115] See also Figure 4 , is a flow chart of another embodiment of a virtual reality video interaction method provided in the embodiment of the present application. Figure 4 The process shown in the above Figure 1 Based on the process shown in FIG, it is described how to generate interactive feedback in the area corresponding to the visual position of the interactive object in the panoramic video image. Figure 4 As shown, the process may include the following steps:
[0116] S401: Determine the interactive object type and collision force vector parameters corresponding to the interactive event.
[0117] Interactive object type: refers to a pre-defined object classification identifier used to distinguish interactive objects of different materials; collision force vector parameter: a three-dimensional physical parameter that describes the magnitude and direction of the interaction force.
[0118] In an embodiment of the present application, by analyzing the trigger position of the interaction event, the corresponding interactive object type identifier is identified, and the direction and magnitude of the force when the interaction controller collides with the triggering body are calculated, providing a parameter basis for subsequent feedback generation.
[0119] S402: Determine a visual special effects template corresponding to the interactive object type in a special effects resource library.
[0120] Special Effects Resource Library: A database that stores various interactive special effects templates; Visual Effects Templates: Pre-made visual effects resource packages, including material response animations and particle effects.
[0121] In an embodiment of the present application, according to the identified object type identifier, the corresponding visual effects template is matched from the resource library, including preset effects such as collision sparks and material deformation, to ensure that the feedback effect is consistent with the object characteristics.
[0122] S403: In the fragment shader stage of the video rendering pipeline, inject the visual effects template into the texture space corresponding to the interaction event coordinates.
[0123] Video rendering pipeline: the complete process of processing video images by the graphics processor; fragment shader: the rendering stage that determines the final color of the pixel.
[0124] In an embodiment of the present application, during the final rendering stage of the video image, the selected visual effects are precisely injected into the texture map of the interactive position to ensure seamless integration of the effects and the video content.
[0125] S404: Generate a three-dimensional sound effect corresponding to the collision force vector parameter through an independent spatial audio processing thread.
[0126] Spatial audio processing thread: an independently running 3D sound effect generation module; three-dimensional sound effect: stereo sound output with a sense of direction and distance.
[0127] In an embodiment of the present application, a dedicated audio thread is used to synthesize sound effects with spatial positioning characteristics in real time based on collision parameters, including impact sounds, friction sounds, etc., to enhance the realism of the interaction.
[0128] Figure 4The process shown achieves precise synchronous output of visual effects and three-dimensional sound effects through a multimodal feedback collaboration mechanism, provides differentiated interactive feedback based on the material properties of objects, and significantly enhances the realism and immersion of the user experience.
[0129] Based on the same technical concept, the embodiment of the present application also provides a virtual reality video interaction device, such as Figure 5 As shown, the device includes:
[0130] An establishing module 51 is configured to establish a video display sphere associated with a spatial position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a pre-rendered panoramic video and maintain a fixed reference orientation of the image;
[0131] A setting module 52 is used to set an interaction trigger body corresponding to the interactive object in the panoramic video in the coordinate system corresponding to the video display sphere, and the coordinate system corresponding to the video display sphere is axially aligned with the virtual camera coordinate system when the panoramic video is rendered, and to keep the relative posture of the interaction trigger body and the video display sphere synchronized, so that the interaction trigger body maintains a fixed posture in the video display sphere coordinate system;
[0132] The generating module 53 is configured to generate interactive feedback in an area corresponding to the visual position of the interactive object in the panoramic video image when an interactive event is detected between the user's interactive operation and the interactive triggering body.
[0133] In one possible implementation, the establishment module is specifically configured to:
[0134] Establishing a real-time data transmission channel between the center point coordinates of the video display sphere and the spatial positioning coordinates of the head-mounted display device in the virtual reality scene, so that the center point coordinates of the video display sphere and the spatial positioning coordinates of the user's head-mounted display device are kept in real time synchronization through the real-time data transmission channel;
[0135] The screen reference orientation parameter of the video display sphere is set to the initial rendering viewing angle direction value of the panoramic video, so that the position of the video display sphere changes with the movement of the user's head and the screen orientation is fixed and does not change with the rotation of the user's head.
[0136] In one possible implementation, the setting module is specifically configured to:
[0137] Obtaining spatial transformation parameters of the interactive object in a virtual camera coordinate system from the rendering data of the panoramic video;
[0138] Based on the spatial transformation parameters, constructing a physical interactive body with a collision detection function in a coordinate system corresponding to the video display sphere;
[0139] The visualization rendering property of the physical interaction body is set to a closed state, so that the physical interaction body is invisible in the virtual reality scene, thereby completing the setting of the interaction trigger body.
[0140] In a possible implementation, the setting module is further configured to:
[0141] Calculating in real time the transformation matrix from the coordinate system corresponding to the video display sphere to the virtual reality world coordinate system;
[0142] The transformation matrix is applied to the pose calculation of the interaction triggering body, so that the spatial position of the interaction triggering body in the virtual reality scene and the visual position of the interactive object in the camera coordinate system continue to coincide.
[0143] In one possible implementation, the device further includes a correction module, configured to:
[0144] Monitoring the positional overlap deviation between the interaction triggering body and the interaction object;
[0145] When the position coincidence deviation exceeds a preset threshold, the posture synchronization process is paused, and the reference posture of the interaction triggering body is corrected according to the actual display position of the interactive object in the current video frame; and the posture synchronization process is restarted after the correction.
[0146] In one possible implementation, the apparatus further includes a determination module configured to:
[0147] Detecting the collision depth between the collision body of the interactive controller and the interactive trigger body in real time;
[0148] When it is detected that the collision depth exceeds a contact threshold, it is determined that an interaction event occurs between the user interaction operation and the interaction triggering body.
[0149] In one possible implementation, the generating module is specifically configured to:
[0150] Determining the interactive object type and collision force vector parameters corresponding to the interactive event;
[0151] Determining a visual special effects template corresponding to the interactive object type in a special effects resource library;
[0152] In the fragment shader stage of the video rendering pipeline, injecting the visual effect template into the texture space corresponding to the interaction event coordinates;
[0153] A three-dimensional sound effect corresponding to the collision force vector parameter is generated through an independent spatial audio processing thread.
[0154] Based on the same technical concept, the embodiment of the present application also provides an electronic device, such as Figure 6As shown, it includes a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.
[0155] Memory 113, for storing computer programs;
[0156] The processor 111 is configured to execute the program stored in the memory 113 by performing the following steps:
[0157] Establishing a video display sphere associated with a spatial position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a pre-rendered panoramic video and maintain a fixed image reference orientation;
[0158] An interaction trigger body corresponding to the interactive object in the panoramic video is set in the coordinate system corresponding to the video display sphere, and the coordinate system corresponding to the video display sphere is axially aligned with the virtual camera coordinate system when the panoramic video is rendered, and the relative posture of the interaction trigger body and the video display sphere is kept synchronized, so that the interaction trigger body maintains a fixed posture in the video display sphere coordinate system;
[0159] When an interaction event is detected between a user interaction operation and the interaction triggering body, interaction feedback is generated in an area corresponding to the visual position of the interaction object in the panoramic video screen.
[0160] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0161] The communication interface is used for communication between the above electronic device and other devices.
[0162] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0163] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0164] In another embodiment provided by the present application, a computer-readable storage medium is further provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of any of the above-mentioned virtual reality video interaction methods are implemented.
[0165] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any virtual reality video interaction method in the above embodiments.
[0166] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.
[0168] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0169] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A virtual reality video interaction method, characterized in that: The method comprises: Establishing a video display sphere associated with a spatial position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a pre-rendered panoramic video and maintain a fixed image reference orientation; An interaction trigger body corresponding to the interactive object in the panoramic video is set in the coordinate system corresponding to the video display sphere, and the coordinate system corresponding to the video display sphere is axially aligned with the virtual camera coordinate system when the panoramic video is rendered, and the relative posture of the interaction trigger body and the video display sphere is kept synchronized, so that the interaction trigger body maintains a fixed posture in the video display sphere coordinate system; When an interaction event is detected between a user interaction operation and the interaction triggering body, interaction feedback is generated in an area corresponding to the visual position of the interaction object in the panoramic video screen.
2. The method according to claim 1, characterized in that The step of establishing a video display sphere associated with a spatial position of a head mounted display device in a virtual reality scene includes: Establishing a real-time data transmission channel between the center point coordinates of the video display sphere and the spatial positioning coordinates of the head-mounted display device in the virtual reality scene, so that the center point coordinates of the video display sphere and the spatial positioning coordinates of the user's head-mounted display device are kept in real time synchronization through the real-time data transmission channel; The screen reference orientation parameter of the video display sphere is set to the initial rendering viewing angle direction value of the panoramic video, so that the position of the video display sphere changes with the movement of the user's head and the screen orientation is fixed and does not change with the rotation of the user's head.
3. The method according to claim 1, characterized in that The step of setting an interaction trigger body corresponding to an interactive object in the panoramic video in a coordinate system corresponding to the video display sphere includes: Obtaining spatial transformation parameters of the interactive object in a virtual camera coordinate system from the rendering data of the panoramic video; Based on the spatial transformation parameters, constructing a physical interactive body with a collision detection function in a coordinate system corresponding to the video display sphere; The visualization rendering property of the physical interaction body is set to a closed state, so that the physical interaction body is invisible in the virtual reality scene, thereby completing the setting of the interaction trigger body.
4. The method according to claim 1, wherein The step of maintaining synchronization between the relative posture of the interaction triggering body and the video display sphere includes: Calculating in real time the transformation matrix from the coordinate system corresponding to the video display sphere to the virtual reality world coordinate system; The transformation matrix is applied to the pose calculation of the interaction triggering body, so that the spatial position of the interaction triggering body in the virtual reality scene and the visual position of the interactive object in the camera coordinate system continue to coincide.
5. The method according to claim 4, characterized in that The method further comprises: Monitoring the positional overlap deviation between the interaction triggering body and the interaction object; When the position coincidence deviation exceeds a preset threshold, the posture synchronization process is paused, and the reference posture of the interaction triggering body is corrected according to the actual display position of the interactive object in the current video frame; and the posture synchronization process is restarted after the correction.
6. The method according to claim 1, characterized in that The method further comprises: Detecting the collision depth between the collision body of the interactive controller and the interactive trigger body in real time; When it is detected that the collision depth exceeds a contact threshold, it is determined that an interaction event occurs between the user interaction operation and the interaction triggering body.
7. The method according to claim 1, characterized in that Generating interactive feedback in an area corresponding to a visual position of an interactive object in the panoramic video image includes: Determining the interactive object type and collision force vector parameters corresponding to the interactive event; Determining a visual special effects template corresponding to the interactive object type in a special effects resource library; In the fragment shader stage of the video rendering pipeline, injecting the visual effect template into the texture space corresponding to the interaction event coordinates; A three-dimensional sound effect corresponding to the collision force vector parameter is generated through an independent spatial audio processing thread.
8. A virtual reality video interaction device, characterized in that: The device comprises: An establishment module, configured to establish a video display sphere associated with a spatial position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a pre-rendered panoramic video and maintain a fixed image reference orientation; A setting module is used to set an interaction trigger body corresponding to the interactive object in the panoramic video in the coordinate system corresponding to the video display sphere, and the coordinate system corresponding to the video display sphere is axially aligned with the virtual camera coordinate system when the panoramic video is rendered, and to keep the relative posture of the interaction trigger body and the video display sphere synchronized, so that the interaction trigger body maintains a fixed posture in the video display sphere coordinate system; The generating module is configured to generate interactive feedback in an area corresponding to a visual position of an interactive object in the panoramic video screen when an interactive event is detected between a user interactive operation and the interactive triggering body.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; The processor is configured to implement the virtual reality video interaction method according to any one of claims 1 to 7 when executing a program stored in the memory.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the virtual reality video interaction method according to any one of claims 1 to 7 is implemented.