Virtual reality video interaction method and device, electronic equipment and storage medium
By establishing a video display sphere associated with the head-mounted display device in the VR scene and using pre-recorded posture data to calculate the posture of dynamic objects in real time, the problems of delay and resource consumption in the existing technology are solved, and real-time and precise synchronization of dynamic objects is achieved, thereby improving the accuracy and real-time nature of the interactive experience.
Patent Information
- Application Number
- CN202511050148.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-26
AI Technical Summary
Existing VR technology suffers from high system latency and large computing resource usage due to complex image processing calculations in dynamic object interactions, making it difficult to achieve real-time and accurate synchronization of the dynamic object postures, affecting the accuracy and real-time nature of the interactive experience.
By establishing a video display sphere associated with the position of the head-mounted display device in the virtual reality scene, using the pre-recorded pose data sequence to calculate the pose of the dynamic object in real time, and setting the dynamic interactive body in the video display sphere coordinate system to maintain the synchronization relationship, real-time and accurate matching of user operations and dynamic objects can be achieved.
It effectively avoids the delay and resource consumption problems caused by real-time image processing, realizes real-time and precise synchronization of the posture of dynamic objects, and improves the accuracy and real-time nature of the interactive experience.
Smart Images

Figure CN120711245A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of virtual reality technology, and in particular to a virtual reality video interaction method, device, electronic device, and storage medium. Background Art
[0002] With the rapid development of virtual reality (VR) technology, interactive applications based on panoramic video are becoming an important research direction. In typical application scenarios such as aviation simulation training and live sports events, users need to accurately interact with high-speed dynamic objects in the video (such as aircraft and athletes), which poses new challenges to VR interaction technology.
[0003] At present, the interaction of dynamic objects in VR videos mainly adopts real-time detection technology based on computer vision. The specific implementation method is: frame by frame analysis of video frames through target detection algorithms (such as YOLO, OpenCV, etc.), real-time identification of the spatial position of dynamic objects, and generation of corresponding interaction areas based on this.
[0004] However, since existing technical solutions require complex image processing calculations, the system latency is high and computing resources are consumed, making it difficult to achieve real-time and accurate synchronization of the posture of dynamic objects, which seriously affects the accuracy and real-time nature of the interactive experience. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide a virtual reality video interaction method, device, electronic device, and storage medium to address the problem that existing technical solutions require complex image processing calculations, resulting in high system latency, large computing resource usage, difficulty in achieving real-time and accurate synchronization of dynamic object positions, and seriously affecting the accuracy and real-time nature of the interactive experience. The specific technical solution is as follows:
[0006] In a first aspect, the present application provides a virtual reality video interaction method, comprising:
[0007] Establishing a video display sphere associated with the position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a panoramic video containing dynamic objects and maintain a fixed image reference orientation;
[0008] Determining the real-time posture of the dynamic object at the current video playback progress from a posture change data sequence of the dynamic object during video production;
[0009] Setting a dynamic interactive body in a coordinate system corresponding to the video display sphere according to the real-time posture, and maintaining a synchronous relationship between the dynamic interactive body and the video display sphere;
[0010] When an interaction event is detected between a user interaction operation and the dynamic interactive object, interaction feedback is generated in an area corresponding to the visual position of the dynamic object in the panoramic video picture.
[0011] In one possible implementation, establishing a video display sphere associated with a position of a head-mounted display device in a virtual reality scene includes:
[0012] Acquiring the three-dimensional spatial position of the head-mounted display device in the virtual reality scene in real time;
[0013] generating a spherical mesh model with a predetermined radius with the three-dimensional spatial position as the sphere center;
[0014] Dynamically mapping the panoramic video texture to the inner surface of the spherical mesh model;
[0015] The rotation properties of the spherical mesh model are set so that the spherical mesh model always maintains a preset screen reference orientation and does not change with the rotation of the head-mounted display device.
[0016] In one possible implementation, the posture change data sequence includes posture data and corresponding timestamps sampled at fixed intervals for the dynamic object during video production. Determining the real-time posture of the dynamic object at the current video playback progress from the posture change data sequence of the dynamic object during video production includes:
[0017] Locating adjacent first and second timestamps in the posture change data sequence according to the current video playback progress;
[0018] Perform spherical linear interpolation on the rotation components of the pose data corresponding to the two timestamps and output a real-time rotation quaternion;
[0019] Perform linear interpolation on the displacement components of the pose data corresponding to the two timestamps and output real-time displacement coordinates;
[0020] The real-time rotation quaternion and the real-time displacement coordinates are combined into the real-time posture at the current moment.
[0021] In one possible implementation, setting a dynamic interactive body in a coordinate system corresponding to the video display sphere according to the real-time posture includes:
[0022] Creating a transparent collision body that matches the three-dimensional outline of the dynamic object in a coordinate system corresponding to the video display sphere;
[0023] Writing the real-time posture into the transformation matrix of the transparent collision body so that the spatial posture of the transparent collision body is synchronized with the real-time posture of the dynamic object;
[0024] The parent node of the transparent collision body is set as the spatial anchor point of the video display sphere, and a binding relationship is established to form the dynamic interaction body.
[0025] In one possible implementation, maintaining the synchronization between the dynamic interactive object and the video display sphere includes:
[0026] Acquire the displacement change of the head mounted display device in the virtual reality scene in real time;
[0027] Converting the displacement change into a local coordinate system corresponding to the video display sphere to generate posture compensation parameters;
[0028] The transformation matrix of the dynamic interactive body is updated according to the posture compensation parameters so that the posture of the dynamic interactive body in the local coordinate system remains unchanged.
[0029] In one possible implementation, generating interactive feedback in an area corresponding to a visual position of a dynamic object in the panoramic video includes:
[0030] Obtaining the collision contact point coordinates and collision velocity values between the interactive controller and the dynamic interactive body;
[0031] Rendering visual feedback effects in the area corresponding to the coordinates of the collision contact point;
[0032] A tactile feedback module of the head mounted display device is driven to output vibration feedback of corresponding intensity based on the collision velocity value.
[0033] In one possible implementation, the method further includes:
[0034] Setting an optical reference mark at a spatial reference position of the video display sphere;
[0035] Identifying the actual display position of the optical reference mark in the video picture;
[0036] Calculating the offset between the actual display position and the preset theoretical position;
[0037] Generating a posture correction parameter according to the offset, and performing reverse adjustment on the current position and orientation of the dynamic interactive body according to the posture correction parameter.
[0038] In a second aspect, the present application provides a virtual reality video interaction device, comprising:
[0039] An establishment module, configured to establish a video display sphere associated with a head-mounted display device position in a virtual reality scene, wherein the video display sphere is used to play a panoramic video containing dynamic objects and maintain a fixed image reference orientation;
[0040] A determination module, configured to determine the real-time posture of the dynamic object at the current video playback progress from a posture change data sequence of the dynamic object during video production;
[0041] A setting module, configured to set a dynamic interactive body in a coordinate system corresponding to the video display sphere according to the real-time posture, and maintain a synchronous relationship between the dynamic interactive body and the video display sphere;
[0042] The generating module is configured to generate interactive feedback in an area corresponding to a visual position of the dynamic object in the panoramic video screen when an interactive event is detected between the user's interactive operation and the dynamic interactive object.
[0043] In one possible implementation, the establishment module is specifically configured to:
[0044] Acquiring the three-dimensional spatial position of the head-mounted display device in the virtual reality scene in real time;
[0045] generating a spherical mesh model with a predetermined radius with the three-dimensional spatial position as the sphere center;
[0046] Dynamically mapping the panoramic video texture to the inner surface of the spherical mesh model;
[0047] The rotation properties of the spherical mesh model are set so that the spherical mesh model always maintains a preset screen reference orientation and does not change with the rotation of the head-mounted display device.
[0048] In one possible implementation, the posture change data sequence includes posture data and corresponding timestamps sampled at fixed intervals for the dynamic object during video production, and the establishment module is specifically configured to:
[0049] Locating adjacent first and second timestamps in the posture change data sequence according to the current video playback progress;
[0050] Perform spherical linear interpolation on the rotation components of the pose data corresponding to the two timestamps and output a real-time rotation quaternion;
[0051] Perform linear interpolation on the displacement components of the pose data corresponding to the two timestamps and output real-time displacement coordinates;
[0052] The real-time rotation quaternion and the real-time displacement coordinates are combined into the real-time posture at the current moment.
[0053] In a possible implementation, the setting module is specifically configured to:
[0054] Creating a transparent collision body that matches the three-dimensional outline of the dynamic object in a coordinate system corresponding to the video display sphere;
[0055] Writing the real-time posture into the transformation matrix of the transparent collision body so that the spatial posture of the transparent collision body is synchronized with the real-time posture of the dynamic object;
[0056] The parent node of the transparent collision body is set as the spatial anchor point of the video display sphere, and a binding relationship is established to form the dynamic interaction body.
[0057] In a possible implementation, the setting module is further configured to:
[0058] Acquire the displacement change of the head mounted display device in the virtual reality scene in real time;
[0059] Converting the displacement change into a local coordinate system corresponding to the video display sphere to generate posture compensation parameters;
[0060] The transformation matrix of the dynamic interactive body is updated according to the posture compensation parameters so that the posture of the dynamic interactive body in the local coordinate system remains unchanged.
[0061] In one possible implementation, the generating module is specifically configured to:
[0062] Obtaining the collision contact point coordinates and collision velocity values between the interactive controller and the dynamic interactive body;
[0063] Rendering visual feedback effects in the area corresponding to the coordinates of the collision contact point;
[0064] A tactile feedback module of the head mounted display device is driven to output vibration feedback of corresponding intensity based on the collision velocity value.
[0065] In one possible implementation, the device further includes a correction module, configured to:
[0066] Setting an optical reference mark at a spatial reference position of the video display sphere;
[0067] Identifying the actual display position of the optical reference mark in the video picture;
[0068] Calculating the offset between the actual display position and the preset theoretical position;
[0069] Generating a posture correction parameter according to the offset, and performing reverse adjustment on the current position and orientation of the dynamic interactive body according to the posture correction parameter.
[0070] In a third aspect, an electronic device is provided, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0071] Memory for storing computer programs;
[0072] The processor is configured to implement any of the method steps described in the first aspect when executing a program stored in the memory.
[0073] In a fourth aspect, a computer-readable storage medium is provided, characterized in that a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, any method step described in the first aspect is implemented.
[0074] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the above-described virtual reality video interaction methods.
[0075] Beneficial effects of the embodiments of the present application:
[0076] The embodiments of the present application provide a virtual reality video interaction method, device, electronic device and storage medium. In the present application, first, by establishing a video display sphere associated with the position of the head-mounted display device, the panoramic video playback is dynamically bound to the VR scene coordinate system; then, based on the pre-recorded posture data sequence, the precise posture of the dynamic object is calculated in real time, avoiding the calculation delay caused by real-time visual detection in the existing technology; next, by setting a dynamic interactive body in the video display sphere coordinate system and maintaining a synchronous relationship, it is ensured that the interaction trigger area and the dynamic object in the video screen always maintain spatial consistency; finally, through a precise interactive feedback mechanism, real-time and precise matching of user operations and dynamic objects is achieved. This solution fundamentally avoids the delay and resource consumption problems caused by relying on real-time image processing in the existing technology through the design of pre-calculating posture data and dynamically associating the coordinate system, and realizes real-time and precise synchronization of the posture of dynamic objects, improving the accuracy and real-time nature of the interactive experience.
[0077] Of course, it is not necessary to achieve all the advantages described above at the same time when implementing any product or method of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0079] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0080] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0081] Figure 1 A flowchart of a virtual reality video interaction method provided in an embodiment of the present application;
[0082] Figure 2 A flowchart of another virtual reality video interaction method provided in an embodiment of the present application;
[0083] Figure 3 A schematic diagram of the structure of a virtual reality video interaction device provided in an embodiment of the present application;
[0084] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0085] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0086] The disclosure below provides many different embodiments or examples for implementing different configurations of the present invention. To simplify the disclosure of the present invention, the components and configurations of specific examples are described below. Of course, these are merely examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or configurations discussed.
[0087] Figure 1 A flowchart of a virtual reality video interaction method provided in an embodiment of the present application. This method can be applied to one or more electronic devices such as smart phones, laptops, desktop computers, portable computers, servers, etc. In addition, the execution subject of this method can be hardware or software. When the above-mentioned execution subject is hardware, the execution subject can be one or more of the above-mentioned electronic devices. For example, a single electronic device can execute this method, or multiple electronic devices can cooperate with each other to execute this method. When the above-mentioned execution subject is software, this method can be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.
[0088] like Figure 1 As shown, the method specifically includes:
[0089] S101. Establish a video display sphere associated with a position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a panoramic video containing dynamic objects and maintain a fixed picture reference orientation.
[0090] A VR (Virtual Reality) scene refers to an interactive digital environment with a three-dimensional sense of space generated by a computer; an HMD (Head-Mounted Display) refers to a virtual reality headset worn by a user; a video display sphere is a panoramic video playback carrier centered on the user's viewpoint, with a panoramic video texture mapped on its inner surface; a panoramic video is pre-produced in a 3D rendering software project, in which the position data P of dynamic objects relative to the rendering camera and the video timestamp t are recorded as a mapping array {(P1, t1), (P2, t2)...(Pn, tn)}); the image reference orientation refers to the pre-set initial spatial orientation of the panoramic video, for example, it can be aligned with the positive Z-axis of the video spherical coordinate system O1-X2Y2Z2, and axial fixation is achieved by locking the Yaw-axis rotation in the world coordinate system.
[0091] In an embodiment of the present application, S101 may specifically include the following steps: obtaining the three-dimensional spatial position of the head-mounted display device in the virtual reality scene in real time; generating a spherical mesh model with a predetermined radius with the three-dimensional spatial position as the center of the sphere; dynamically mapping the panoramic video texture to the inner surface of the spherical mesh model; setting the rotation properties of the spherical mesh model so that the spherical mesh model always maintains a preset screen reference orientation and does not change with the rotation of the head-mounted display device.
[0092] The three-dimensional spatial position refers to the X / Y / Z axis coordinate values of the HMD in the virtual reality scene coordinate system; the spherical mesh model is a spherical three-dimensional geometric body composed of a number of triangular facets; the panoramic video texture refers to the panoramic video image data processed by equidistant cylindrical projection; the predetermined radius can be set arbitrarily according to actual needs.
[0093] In this solution, the system includes three key coordinate systems: the virtual reality scene (such as the game world) coordinate system O-XYZ, the headset coordinate system O1-X1Y1Z1 (which changes with the position and orientation of the headset), and the video sphere coordinate system O1-X2Y2Z2 (the origin O1 coincides with the headset but the axis is fixed).
[0094] During specific implementation, the system obtains the three-dimensional coordinate values (X, Y, Z) of the HMD in the virtual reality scene in real time through the position tracking module. The coordinate values correspond to the real-time position of the HMD in the virtual reality scene coordinate system O-XYZ. Then, with this coordinate point as the center of the sphere, a spherical mesh model with a fixed radius is generated. The radius value is calculated and determined based on the optimal field of view of the human eye. Then, the pre-made panoramic video texture (generated in the 3D rendering project) is dynamically mapped to the inner surface of the sphere through a shader (such as a GPU shader or a CPU shader). Finally, by setting the rotation property lock matrix, the spherical model always maintains the preset screen reference orientation (i.e., the axial direction of the video spherical coordinate system O1-X2Y2Z2) and does not change with the rotation of the HMD. By establishing a video spherical coordinate system (O1-X2Y2Z2) dynamically bound to the HMD's position but independent of rotation, this solution not only achieves spatial synchronization between the video content and the user's viewpoint (O1 coincides with the HMD's position), but also ensures the stability of the image orientation (the X2Y2Y2 axes are fixed), providing the necessary spatial reference for the precise position synchronization of subsequent dynamic objects. This effectively solves the problem of image jitter caused by HMD rotation in traditional VR video interaction, while also providing a precise spatial positioning reference for dynamic interactive objects.
[0095] S102: Determine the real-time posture of the dynamic object at the current video playback progress from the posture change data sequence of the dynamic object during video production.
[0096] The pose change data sequence refers to the set of spatial position and posture data of dynamic objects relative to the rendering camera recorded during the video production stage. It contains the pose data and corresponding timestamps sampled at fixed intervals for dynamic objects during video production. It is specifically expressed as an array {(P1, t1), (P2, t2)...(Pn, tn)}, where P represents the pose (including three-dimensional coordinates and rotation) and t corresponds to the video timestamp. The real-time pose refers to the spatial state of the object dynamically calculated based on the current video progress.
[0097] In an embodiment of the present application, S102 may specifically include the following steps: according to the current video playback progress, locating the adjacent first timestamp and second timestamp in the posture change data sequence; performing spherical linear interpolation on the rotation components of the posture data corresponding to the two timestamps, and outputting a real-time rotation quaternion; performing linear interpolation on the displacement components of the posture data corresponding to the two timestamps, and outputting real-time displacement coordinates; and combining the real-time rotation quaternion and the real-time displacement coordinates into the real-time posture at the current moment.
[0098] The first and second timestamps refer to the two sampling time points adjacent to the current video playback time in the pose change data sequence {(P1, t1), (P2, t2)...(Pn, tn)}. Spherical linear interpolation (Slerp) is an interpolation algorithm for the rotation component of the quaternion. It can maintain unit length and constant speed during the interpolation process. In applications, the quaternion must be normalized to ensure the shortest arc characteristics of the interpolation path. Linear interpolation (Lerp) is a linear transition algorithm for the displacement component. The real-time rotation quaternion and real-time displacement coordinates represent the object's rotation state and position coordinates obtained after interpolation calculation, respectively.
[0099] This solution first locates two adjacent keyframes in the pose sequence based on the current playback time returned by the video decoder (i.e., the timestamp of the video playback progress). Then, the Slerp algorithm is used to calculate a smooth quaternion for the rotation component (avoiding the gimbal lock problem caused by Euler angle interpolation), and the Lerp algorithm is used to calculate the intermediate coordinates for the displacement component. Finally, the two are combined to form complete real-time pose data. By separating the rotation and displacement components and using the interpolation algorithm best suited for each type of data, this solution ensures the natural smoothness of dynamic object motion (no sudden changes in rotation interpolation) while achieving sub-millisecond pose calculation efficiency (low computational complexity for displacement interpolation), providing a precise pose reference for subsequent interactive triggering.
[0100] In addition, when it is impossible to obtain the pose change data sequence by obtaining the original rendering engineering data, the perspective projection inverse method can be used to estimate the pose: through the pixel coordinates of the dynamic object in the video frame and the known camera parameters (referring to the internal and external parameters of the virtual camera (or actual shooting equipment) when producing panoramic videos, these parameters are the basic data for pose estimation), the PnP algorithm is used to solve the approximate pose of the object in the virtual camera coordinate system as the real-time pose.
[0101] S103: Setting a dynamic interactive body in a coordinate system corresponding to the video display sphere according to the real-time posture, and maintaining a synchronous relationship between the dynamic interactive body and the video display sphere.
[0102] A dynamic interactive body refers to a transparent collision body generated in a virtual reality scene based on real-time posture data; the coordinate system corresponding to the video display sphere refers to the O1-X2Y2Z2 coordinate system with the headset position O1 as the origin and fixed axially.
[0103] In one embodiment, setting a dynamic interactive body in a coordinate system corresponding to a video display sphere according to a real-time posture may include the following steps: creating a transparent collision body that matches the three-dimensional outline of the dynamic object in the coordinate system corresponding to the video display sphere; writing the real-time posture into the transformation matrix of the transparent collision body so that the spatial posture of the transparent collision body is synchronized with the real-time posture of the dynamic object; setting the parent node of the transparent collision body as the spatial anchor point of the video display sphere, and establishing a binding relationship to form the dynamic interactive body.
[0104] A transparent collision volume refers to an interactive object that is rendered invisible by setting the alpha channel value of the renderer component to 0 (or disabling the rendering component) while retaining its physical collision properties. A transformation matrix is a 4×4 matrix containing position, rotation, and scale information. A spatial anchor point is a reference point for maintaining the spatial relationship of virtual objects. This solution first generates a collision volume that precisely matches the three-dimensional outline of the dynamic object in the video sphere coordinate system O1-X2Y2Z2. Then, the real-time pose P calculated by S102 is written into the collision volume's transformation matrix to achieve pose synchronization. Finally, a stable hierarchical relationship is established by setting the parent node to bind to the spatial anchor point of the video sphere. Through a hierarchical spatial binding mechanism and precise pose synchronization, this solution not only ensures submillimeter spatial consistency between the interaction trigger area and the dynamic object in the video image, but also maintains a stable spatial reference system through parent node binding, effectively avoiding interaction position drift caused by headset movement.
[0105] In one embodiment, maintaining the synchronization relationship between the dynamic interactive body and the video display sphere may include the following steps: obtaining the displacement change of the head-mounted display device in the virtual reality scene in real time; converting the displacement change to the local coordinate system corresponding to the video display sphere to generate posture compensation parameters; updating the transformation matrix of the dynamic interactive body according to the posture compensation parameters so that the posture of the dynamic interactive body in the local coordinate system remains unchanged.
[0106] The displacement change refers to the position offset vector ΔP (Δx, Δy, Δz) of the head HMD in the virtual reality scene world coordinate system O-XYZ; the posture compensation parameter is the result of converting the displacement change to the local coordinate system O1-X2Y2Z2 of the video display sphere. This solution first obtains its displacement change ΔP in real time through the HMD's positioning and tracking system; then, it converts this vector to the O1-X2Y2Z2 coordinate system (the conversion method includes but is not limited to the coordinate system rotation matrix multiplication); finally, it performs inverse compensation on the transformation matrix of the dynamic interactive body. This solution establishes a precise displacement compensation mechanism to ensure that the dynamic interactive body maintains a constant posture in the O1-X2Y2Z2 local coordinate system, effectively solving the problem of interactive position offset caused by the movement of the head display, and achieving sub-millimeter spatial synchronization accuracy between the interactive trigger area and the dynamic objects in the video screen. At the same time, the compensation calculation only involves simple matrix operations, and the computational efficiency is significantly better than the traditional visual relocalization method.
[0107] S104: When an interaction event is detected between the user interaction operation and the dynamic interactive object, an interaction feedback is generated in an area corresponding to the visual position of the dynamic object in the panoramic video image.
[0108] User interaction operations refer to input actions generated by VR controllers or gesture recognition; interaction events are system messages triggered when the collision detection result between the controller ray (in the application, it can also be a parabolic bullet or curve-tracing missile launched by the controller) and the dynamic interaction body (i.e., a transparent collision body) is true; interaction feedback includes but is not limited to visual effects, tactile vibration and other response forms.
[0109] In an embodiment of the present application, S104 may specifically include the following steps: obtaining the collision contact point coordinates and collision speed value between the interactive controller and the dynamic interactive body; rendering a visual feedback effect in the area corresponding to the collision contact point coordinates; and driving the tactile feedback module of the head-mounted display device to output vibration feedback of corresponding intensity based on the collision speed value.
[0110] The collision contact point coordinates are the coordinates corresponding to the collision position between the interactive controller and the dynamic interactive body; the collision velocity value refers to the relative motion velocity scalar when the conductor and the interactive body are in contact; the visual feedback effects include but are not limited to rendering effects such as particle systems and highlight shaders; the tactile feedback module refers to the vibration motor component integrated in the head-mounted display device. This solution first obtains the precise contact point coordinates and collision velocity through the collision detection system of the physics engine. Then, it renders the preset visual effects at the location of the collision contact point coordinates. At the same time, it queries the preset intensity mapping curve based on the collision velocity and drives the tactile module to output vibrations of corresponding intensity (the vibration frequency is positively correlated with the collision velocity). In the application, all interaction events are handled by the physical collision component of the trigger box (i.e., the dynamic interactive body). When a collision occurs, the trigger box passes the event parameters (including the collision point coordinates, collision force, etc.) to the central event processor, which then uniformly dispatches the visual / tactile / auditory feedback modules to ensure centralized management of the interaction logic.
[0111] Through precise coordinate transformation, this solution achieves pixel-level position matching between interactive feedback and dynamic objects in the video screen. At the same time, the dynamic adjustment of the intensity of multimodal feedback significantly improves the realism and operation accuracy of the interaction, solving the technical problem of position inaccuracy in traditional VR video interaction.
[0112] In this application, first, by establishing a video display sphere associated with the position of the head-mounted display device, the panoramic video playback is dynamically bound to the VR scene coordinate system; then, based on the pre-recorded posture data sequence, the precise posture of the dynamic object is calculated in real time, avoiding the calculation delay caused by real-time visual detection in the existing technology; next, by setting a dynamic interactive body in the video display sphere coordinate system and maintaining a synchronous relationship, it is ensured that the interaction trigger area and the dynamic object in the video screen always maintain spatial consistency; finally, through a precise interactive feedback mechanism, real-time and precise matching of user operations and dynamic objects is achieved. This solution fundamentally avoids the delay and resource consumption problems caused by relying on real-time image processing in the existing technology through the design of pre-calculated posture data and dynamic association of the coordinate system, and realizes real-time and precise synchronization of the posture of dynamic objects, improving the accuracy and real-time nature of the interactive experience.
[0113] See also Figure 2 , is a flow chart of another embodiment of a virtual reality video interaction method provided in the embodiment of the present application. Figure 2 As shown, the process may include the following steps:
[0114] S201 , setting an optical reference mark at a spatial reference position of the video display sphere.
[0115] Optical reference marks refer to visible identification points (such as high-contrast color blocks or special patterns) set at specific locations on the surface of the video display sphere. Spatial reference positions refer to geometric feature points selected on the surface of the video display sphere to establish a stable coordinate system, which may include the North Pole, South Pole and equatorial reference points.
[0116] In this embodiment, a spatial calibration benchmark is established by placing three optical markers at the north pole, south pole, and just in front of the equatorial plane of a spherical model. This optimizes the spatial distribution of the markers, providing high-precision reference points for subsequent pose calibration and resolving potential distortion errors in spherical video textures.
[0117] S202: Identify the actual display position of the optical reference mark in the video image.
[0118] The actual display position refers to the pixel coordinates of the optical marker in the rendered image, and the "recognition" process includes feature point detection and matching in the computer vision algorithm.
[0119] In this embodiment of the application, the GPU shader is used to capture the actual rendering position of the marker point in the video texture in real time. This solution uses image processing technology to achieve sub-pixel marker point positioning accuracy (error <0.5 pixel), providing a reliable data source for offset calculation.
[0120] S203: Calculate the offset between the actual display position and the preset theoretical position.
[0121] The offset includes the position deviation ΔP and rotation deviation ΔR of the marker point; the preset theoretical position is the UV coordinate of the marker point on the ideal spherical model.
[0122] In the embodiment of the present application, the overall offset matrix is calculated by comparing the actual / theoretical coordinate difference of each marker point using the least squares method. In this way, the spatial pose error can be reduced by jointly solving multiple marker points.
[0123] S204: Generate a posture correction parameter according to the offset, and reversely adjust the current position and orientation of the dynamic interactive body according to the posture correction parameter.
[0124] The pose correction parameter is a 4×4 transformation matrix containing displacement and rotation correction values; reverse adjustment refers to applying an inverse transformation to the dynamic interactive body.
[0125] In this embodiment, after decomposing the offset matrix, an inverse compensation correction is applied to the transformation matrix of the interactive object. This closed-loop calibration mechanism effectively eliminates the cumulative errors caused by factors such as device drift and video distortion, ensuring that the spatial matching accuracy between the interactive object and the video object remains at the sub-millimeter level for a long time.
[0126] Figure 2 The process shown in the figure sets optical reference marks at specific positions on the video display sphere and detects its actual rendering position in real time, accurately calculates the offset from the theoretical position, and then generates posture correction parameters to adjust the dynamic interactive body in reverse. This scheme effectively eliminates the spatial errors caused by device drift and video deformation by establishing a closed-loop calibration mechanism, so that the interaction trigger area and the dynamic objects in the video screen always maintain sub-millimeter precision matching. At the same time, through optimized marker layout and efficient offset calculation algorithm, the system computing overhead is significantly reduced while ensuring calibration accuracy, providing a stable and reliable spatial reference for VR video interaction.
[0127] In another embodiment of the present application, for a scene containing multiple dynamic objects, the method may further include the following steps: creating an independent pose data channel and interactive body instance for each dynamic object; synchronously updating all pose data channels through the video playback timestamp; and matching the interactive body and pose data according to the unique identifier of the dynamic object.
[0128] A pose data channel refers to an independent data structure that stores the pose sequence of a single dynamic object; an interactive body instance is an independent transparent collision body object generated for each dynamic object; a unique identifier is a digital ID or hash value assigned to each dynamic object.
[0129] This solution establishes a multi-object parallel processing architecture. First, a dedicated pose data container is assigned to each dynamic object in the scene. Then, a unified timestamp mechanism is used to synchronously update the pose status of all objects (ensuring the consistency of the timing of each object's motion). Finally, identifier matching is used to ensure the correct association between the interactive body and the pose data. In this way, through resource isolation and parallel processing mechanisms, it is possible to achieve independent and accurate tracking of multiple dynamic objects in complex scenes, and through unified time synchronization, the relative timing relationship of the motion between objects is guaranteed. At the same time, the identifier matching mechanism effectively prevents the misalignment of pose data between objects, improving the operating efficiency of the multi-object VR video interaction system.
[0130] Based on the same technical concept, the embodiment of the present application also provides a virtual reality video interaction device, such as Figure 3 As shown, the device includes:
[0131] An establishing module 31 is configured to establish a video display sphere associated with a head-mounted display device position in a virtual reality scene, wherein the video display sphere is used to play a panoramic video containing dynamic objects and maintain a fixed reference orientation of the image;
[0132] A determination module 32 is configured to determine the real-time posture of the dynamic object at the current video playback progress from a posture change data sequence of the dynamic object during video production;
[0133] A setting module 33, configured to set a dynamic interactive body in a coordinate system corresponding to the video display sphere according to the real-time posture, and maintain a synchronous relationship between the dynamic interactive body and the video display sphere;
[0134] The generating module 34 is configured to generate interactive feedback in an area corresponding to the visual position of the dynamic object in the panoramic video image when an interactive event is detected between the user's interactive operation and the dynamic interactive object.
[0135] In one possible implementation, the establishment module is specifically configured to:
[0136] Acquiring the three-dimensional spatial position of the head-mounted display device in the virtual reality scene in real time;
[0137] generating a spherical mesh model with a predetermined radius with the three-dimensional spatial position as the sphere center;
[0138] Dynamically mapping the panoramic video texture to the inner surface of the spherical mesh model;
[0139] The rotation properties of the spherical mesh model are set so that the spherical mesh model always maintains a preset screen reference orientation and does not change with the rotation of the head-mounted display device.
[0140] In one possible implementation, the posture change data sequence includes posture data and corresponding timestamps sampled at fixed intervals for the dynamic object during video production, and the establishment module is specifically configured to:
[0141] Locating adjacent first and second timestamps in the posture change data sequence according to the current video playback progress;
[0142] Perform spherical linear interpolation on the rotation components of the pose data corresponding to the two timestamps and output a real-time rotation quaternion;
[0143] Perform linear interpolation on the displacement components of the pose data corresponding to the two timestamps and output real-time displacement coordinates;
[0144] The real-time rotation quaternion and the real-time displacement coordinates are combined into the real-time posture at the current moment.
[0145] In a possible implementation, the setting module is specifically configured to:
[0146] Creating a transparent collision body that matches the three-dimensional outline of the dynamic object in a coordinate system corresponding to the video display sphere;
[0147] Writing the real-time posture into the transformation matrix of the transparent collision body so that the spatial posture of the transparent collision body is synchronized with the real-time posture of the dynamic object;
[0148] The parent node of the transparent collision body is set as the spatial anchor point of the video display sphere, and a binding relationship is established to form the dynamic interaction body.
[0149] In a possible implementation, the setting module is further configured to:
[0150] Acquire the displacement change of the head mounted display device in the virtual reality scene in real time;
[0151] Converting the displacement change into a local coordinate system corresponding to the video display sphere to generate posture compensation parameters;
[0152] The transformation matrix of the dynamic interactive body is updated according to the posture compensation parameters so that the posture of the dynamic interactive body in the local coordinate system remains unchanged.
[0153] In one possible implementation, the generating module is specifically configured to:
[0154] Obtaining the collision contact point coordinates and collision velocity values between the interactive controller and the dynamic interactive body;
[0155] Rendering visual feedback effects in the area corresponding to the coordinates of the collision contact point;
[0156] A tactile feedback module of the head mounted display device is driven to output vibration feedback of corresponding intensity based on the collision velocity value.
[0157] In one possible implementation, the device further includes a correction module, configured to:
[0158] Setting an optical reference mark at a spatial reference position of the video display sphere;
[0159] Identifying the actual display position of the optical reference mark in the video picture;
[0160] Calculating the offset between the actual display position and the preset theoretical position;
[0161] Generating a posture correction parameter according to the offset, and performing reverse adjustment on the current position and orientation of the dynamic interactive body according to the posture correction parameter.
[0162] Based on the same technical concept, the embodiment of the present application also provides an electronic device, such as Figure 4 As shown, it includes a processor 111, a communication interface 112, a memory 113 and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.
[0163] Memory 113, for storing computer programs;
[0164] The processor 111 is configured to execute the program stored in the memory 113 by performing the following steps:
[0165] Establishing a video display sphere associated with the position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a panoramic video containing dynamic objects and maintain a fixed image reference orientation;
[0166] Determining the real-time posture of the dynamic object at the current video playback progress from a posture change data sequence of the dynamic object during video production;
[0167] Setting a dynamic interactive body in a coordinate system corresponding to the video display sphere according to the real-time posture, and maintaining a synchronous relationship between the dynamic interactive body and the video display sphere;
[0168] When an interaction event is detected between a user interaction operation and the dynamic interactive object, interaction feedback is generated in an area corresponding to the visual position of the dynamic object in the panoramic video picture.
[0169] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0170] The communication interface is used for communication between the above electronic device and other devices.
[0171] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0172] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0173] In another embodiment provided by the present application, a computer-readable storage medium is further provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of any of the above-mentioned virtual reality video interaction methods are implemented.
[0174] In another embodiment provided by the present application, a computer program product including instructions is also provided, which, when executed on a computer, enables the computer to execute any virtual reality video interaction method in the above embodiments.
[0175] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0176] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, or of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiment.
[0177] It should be understood that the terms used herein are for the purpose of describing specific example embodiments only and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "comprise", "include", "contain" and "have" are inclusive and therefore specify the presence of stated features, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the specific order described or illustrated, unless the order of execution is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0178] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A virtual reality video interaction method, characterized in that: The method comprises: Establishing a video display sphere associated with the position of a head-mounted display device in a virtual reality scene, wherein the video display sphere is used to play a panoramic video containing dynamic objects and maintain a fixed image reference orientation; Determining the real-time posture of the dynamic object at the current video playback progress from a posture change data sequence of the dynamic object during video production; Setting a dynamic interactive body in a coordinate system corresponding to the video display sphere according to the real-time posture, and maintaining a synchronous relationship between the dynamic interactive body and the video display sphere; When an interaction event is detected between a user interaction operation and the dynamic interactive object, interaction feedback is generated in an area corresponding to the visual position of the dynamic object in the panoramic video picture.
2. The method according to claim 1, characterized in that The step of establishing a video display sphere associated with a position of a head mounted display device in a virtual reality scene includes: Acquiring the three-dimensional spatial position of the head-mounted display device in the virtual reality scene in real time; generating a spherical mesh model with a predetermined radius with the three-dimensional spatial position as the sphere center; Dynamically mapping the panoramic video texture to the inner surface of the spherical mesh model; The rotation properties of the spherical mesh model are set so that the spherical mesh model always maintains a preset screen reference orientation and does not change with the rotation of the head-mounted display device.
3. The method according to claim 1, characterized in that The posture change data sequence includes posture data and corresponding timestamps sampled at fixed intervals for the dynamic object during video production. Determining the real-time posture of the dynamic object at the current video playback progress from the posture change data sequence of the dynamic object during video production includes: Locating adjacent first and second timestamps in the posture change data sequence according to the current video playback progress; Perform spherical linear interpolation on the rotation components of the pose data corresponding to the two timestamps and output a real-time rotation quaternion; Perform linear interpolation on the displacement components of the pose data corresponding to the two timestamps and output real-time displacement coordinates; The real-time rotation quaternion and the real-time displacement coordinates are combined into the real-time posture at the current moment.
4. The method according to claim 1, wherein The step of setting a dynamic interactive body in a coordinate system corresponding to the video display sphere according to the real-time posture includes: Creating a transparent collision body that matches the three-dimensional outline of the dynamic object in a coordinate system corresponding to the video display sphere; Writing the real-time posture into the transformation matrix of the transparent collision body so that the spatial posture of the transparent collision body is synchronized with the real-time posture of the dynamic object; The parent node of the transparent collision body is set as the spatial anchor point of the video display sphere, and a binding relationship is established to form the dynamic interaction body.
5. The method according to claim 1, wherein Maintaining the synchronization relationship between the dynamic interactive object and the video display sphere includes: Acquire the displacement change of the head mounted display device in the virtual reality scene in real time; Converting the displacement change into a local coordinate system corresponding to the video display sphere to generate posture compensation parameters; The transformation matrix of the dynamic interactive body is updated according to the posture compensation parameters so that the posture of the dynamic interactive body in the local coordinate system remains unchanged.
6. The method according to claim 1, characterized in that Generating interactive feedback in an area corresponding to a visual position of a dynamic object in the panoramic video image includes: Obtaining the collision contact point coordinates and collision velocity values between the interactive controller and the dynamic interactive body; Rendering visual feedback effects in the area corresponding to the coordinates of the collision contact point; A tactile feedback module of the head mounted display device is driven to output vibration feedback of corresponding intensity based on the collision velocity value.
7. The method according to claim 1, characterized in that The method further comprises: Setting an optical reference mark at a spatial reference position of the video display sphere; Identifying the actual display position of the optical reference mark in the video picture; Calculating the offset between the actual display position and the preset theoretical position; Generating a posture correction parameter according to the offset, and performing reverse adjustment on the current position and orientation of the dynamic interactive body according to the posture correction parameter.
8. A virtual reality video interaction device, characterized in that: The device comprises: An establishment module, configured to establish a video display sphere associated with a head-mounted display device position in a virtual reality scene, wherein the video display sphere is used to play a panoramic video containing dynamic objects and maintain a fixed image reference orientation; A determination module, configured to determine the real-time posture of the dynamic object at the current video playback progress from a posture change data sequence of the dynamic object during video production; A setting module, configured to set a dynamic interactive body in a coordinate system corresponding to the video display sphere according to the real-time posture, and maintain a synchronous relationship between the dynamic interactive body and the video display sphere; The generating module is configured to generate interactive feedback in an area corresponding to a visual position of the dynamic object in the panoramic video screen when an interactive event is detected between the user's interactive operation and the dynamic interactive object.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; The processor is configured to implement the virtual reality video interaction method according to any one of claims 1 to 7 when executing a program stored in the memory.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the virtual reality video interaction method according to any one of claims 1 to 7 is implemented.