8KFOV content playing method and system capable of reducing swivel time delay

By obtaining the user's new perspective coordinates in the VR player, downloading and decoding new video chunks, and quickly rendering them to the texture, the problem of head-turning delay of existing VR players is solved, and the effect of users quickly seeing new high-definition video content is achieved, improving user experience and supporting high-reality applications.

CN120050468APending Publication Date: 2025-05-27SHANGHAI WONDERTEK SOFTWARE CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510123622.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing VR players have serious delay problems when users turn their heads, which causes users to wait for a long time to see new high-definition video content, affecting user experience and limiting the application of VR technology in some application scenarios with high real-time requirements.

Method used

By obtaining the user's new perspective coordinate position, downloading new visual area chunked HD video content, and using software decoding, quickly rendering new video images onto the texture, allowing users to see both the original and the new video content.

Benefits of technology

It significantly shortens the head-turning delay, and users can quickly see new high-definition video content, improve user experience, and support VR technology applications in application scenarios with high real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050468A_ABST
    Figure CN120050468A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multimedia, and discloses an 8KFOV content playing method and system capable of reducing turning time delay, and the method comprises the steps: S1, obtaining a new visual angle coordinate position after a user turns, and downloading a new visual region block high-definition video content to form a block high-definition file; s2, analyzing the new partitioned high-definition files through a new file analyzer, and independently decoding the partitioned high-definition files one by one; and S3, quickly acquiring the position information of the decoded blocks, and rendering the decoded video content to the new region for displaying the texture according to the position information, so that the original region high-definition video and the new region high-definition video are visible at the same time. Through the method, the 8KFOV content playing system capable of reducing the head turning time delay is realized, a user turns to watch contents of other areas, keeps an original playing area, downloads video contents of a new area, decodes the video contents by software, supplements new video pictures to textures, and can quickly see a new part of videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimedia technology, and in particular, to an 8K FOV content playback method and system for reducing head-turning latency. Background Art

[0002] In the current field of virtual reality (VR) technology, the VR player, as an important bridge connecting users with the virtual world, has attracted much attention for its performance and user experience. VR technology provides users with an immersive audio-visual experience, making users feel as if they are in a virtual environment and can interact naturally with virtual scenes. However, existing VR players have some significant problems in practical applications. Especially when users turn their heads to view content in different areas, their processing mechanisms have obvious limitations, seriously affecting the user experience.

[0003] Existing VR players usually adopt a relatively traditional video playback and download strategy. During normal playback, the VR player plays video content in sequence according to a pre-set time period. These video contents are often divided into multiple chunks for convenient storage and transmission. When the user turns their head, it means that the user hopes to view other areas in the virtual scene, which requires the VR player to obtain and play the corresponding area's video content according to the user's new viewing position. However, existing VR players have serious latency problems when dealing with the user's head-turning operation. Specifically, after the user turns their head, the player does not immediately respond to the user's viewing angle change. Instead, it continues to play the video content of the current time period until all the content of this time period is played, as shown in the accompanying drawings of the specification Figure 2 shown. Only when the video playback of the current time period ends and is about to enter the playback of the next time period, will the VR player start to perform corresponding operations according to the user's new position information.

[0004] This operation process includes multiple complex steps. First, the player needs to determine the visible area corresponding to the user's new viewing position, and then according to the coordinate information of this visible area, send a request to the server to download the new video chunk content, and then perform a series of operations such as parsing, decoding, and rendering, consuming a large amount of computing resources and time.

[0005] Due to the above series of complex operations that can only be carried out in sequence after the video playback of the current time period is completed, users need to wait for a relatively long time to see the high-definition video content of the new area they want to view. This long waiting time not only destroys the user's immersive experience, making users feel obvious lag and latency after turning their heads, but also limits the application of VR technology in some application scenarios with high real-time requirements (such as virtual reality games, real-time virtual training, etc.).

[0006] Therefore, how to optimize the video processing mechanism of the VR player when the user turns their head, reduce the user's waiting time, and improve the user experience have become important issues that need to be solved urgently in the current VR technology field. Summary of the Invention

[0007] The purpose of the present invention is to solve the above-mentioned disadvantages existing in the prior art, and provide an 8K FOV content playback method and system for reducing the head-turning time delay. When watching a high-definition video, when the user turns their head to view content in other areas, the original playback area is maintained, and at the same time, the video content of the new area is downloaded, decoded by software, and the new video frame is supplemented to the texture, so that the user can quickly see the new part of the video.

[0008] On the one hand, an 8K FOV content playback method for reducing the head-turning time delay is provided, including the following steps: S1: After the user turns their head, obtain the new perspective coordinate position, and download the high-definition video content of the new visible area block to form a block high-definition file; S2: Parse the new block high-definition file through a new file parser, and independently decode each of the block high-definition files; S3: Quickly obtain the position information of the decoded blocks, and render the decoded video content to a new area of the display texture according to the position information, so that the high-definition video of the original area and the high-definition video of the new area are visible at the same time.

[0009] Further, in step S1, the obtaining of the new perspective coordinate position includes: Use a sensor to track the rotation of the user's head to obtain the head rotation information, and the new perspective direction vector The calculation formula is as follows: Wherein, is the rotation matrix represented by the rotation information, is the user's previous perspective direction vector; According to the coordinate system of the scene and the position information of the user , combined with the new perspective direction vector , determine the boundary of the new visible area. Among them, the visible area is a cone with the user's position as the vertex, and the half vertex angle of the cone is , then a point in the visible area satisfies , and by traversing all the block positions in the scene, the new visible area blocks are determined.

[0010] Further, in step S1, the downloading of the high-definition video content of the new visible area block to form a block high-definition file further includes: The server stores the high-definition video chunk data of the entire scene. Each chunk includes a corresponding coordinate identifier. A download request is sent to the server according to the coordinates of the newly visible area chunks, and the server retrieves and returns the corresponding chunk data according to the request.

[0011] Further, in step S2, parsing the new chunk high-definition file by a new file parser includes file format parsing and data format parsing, including: The file parser reads the meta-information at the head of the high-definition file, including video coding format, resolution, and frame rate; The data inside the chunk high-definition file is organized according to a specific data structure for storing different parts of video frames. The file parser gradually parses the data of each video frame according to this data structure and stores it in a specific data structure in memory for subsequent decoding use.

[0012] Further, in step S2, independently decoding each of the chunk high-definition files includes: S21: Select a corresponding decoding algorithm according to the parsed video coding format for independent decoding of each block; S22: For the current block, convert the encoded symbols into quantization coefficients through an entropy decoding algorithm from the parsed video frame data; S23: Inverse-quantize the quantization coefficients obtained by entropy decoding to recover the transform coefficients, and then convert the transform coefficients back to pixel values in the spatial domain through an inverse transform; S24: Predict the pixel values of the current block according to the pixel values of the decoded pixels around the current block, and predict the pixel values of the current block using similar blocks in the reference frame; S25: Finally, add the predicted pixel values of the current block to the residuals obtained by the inverse transform to obtain the final decoded pixel values.

[0013] Further, in step S3, rendering to a new area of the display texture according to the new chunk position information includes: Convert the boundary coordinates of the new chunk in the scene to the screen coordinate space through a projection transformation to obtain the coordinates of the chunk in the scene , and obtain the boundary coordinates of the new chunk on the screen through calculation. The calculation formula is as follows: , , , where and are the coordinate values after the projection transformation, and the coordinates projected onto the two-dimensional plane are obtained after perspective division , is an intermediate value for perspective division, Represents a perspective projection matrix, and then maps the coordinates to the screen coordinate range and , and the calculation formula is as follows: wherein, is the resolution of the display screen, is the approximate coordinate position of the block on the screen finally; Determine the texture coordinate range of the block on the display texture according to the screen coordinates of the block, and complete the mapping conversion from screen coordinates to texture coordinates; Render the decoded video frame data as texture data to the corresponding area of the display texture according to the calculated texture coordinate range.

[0014] Preferably, step S3 further includes: During the rendering process, fuse the original texture area and the newly rendered texture area, set an appropriate transparency value through transparency blending to ensure that the new texture area and the new texture area are naturally fused and displayed, so that the user can see the high-definition video content of both the old and new areas at the same time.

[0015] More preferably, optimizing the playback process to implement playback control management further includes: Before obtaining the new coordinate position, first detect whether a head turn occurs in the downloaded thread. Once a head turn is detected, trigger the operation of obtaining the new coordinate position; Judge the end time of the period when the head turn occurs and the end time of the period when the currently playing content is about to end. If the time is too short, ignore the head turn behavior to avoid unnecessary new content download and processing caused by the head turn when the playback is about to end and the remaining time is too short, thereby reducing resource waste; Before downloading the new block high-definition file, judge whether the new block high-definition file already exists within the current playback period. If it does not exist and has not been downloaded, download the new block high-definition file.

[0016] On the other hand, a 8K FOV content playback system for reducing head turn latency is provided, including: A new perspective information acquisition module, which is used to obtain a new perspective coordinate position after the user turns their head, and download the new visible area block high-definition video content to form a block high-definition file; A parsing and decoding module, which is used to parse the new block high-definition file through a new file parser and decode each of the block high-definition files independently; The rendering management module is used to quickly obtain the position information of the decoded chunks, and render the decoded video content to a new area of the display texture according to the position information, so that the high-definition video in the original area and the high-definition video in the new area are visible at the same time.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: By downloading a part of the high-definition video content, the present invention separately obtains the video content of each video block by software decoding and supplements it to the new area, so as to realize the simultaneous display of the original video and the new video, enabling the user to quickly see the new high-definition video content after turning the head. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of an 8K FOV content playback method for reducing head-turning latency according to the present invention; Figure 2 is a schematic diagram of head-turning display in the prior art solution; Figure 3 is a schematic diagram of an 8K FOV content playback display for reducing head-turning latency according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0020] Quickly reduce the head-turning latency and improve the user viewing experience. Protection is adopted to use new sharded files to supplement the displayed video to achieve the effect of shortening the display of high-definition videos.

[0021] The following will illustrate the specific embodiments of the present invention in conjunction with the accompanying drawings and embodiments.

[0022] Embodiment 1 Please refer to Figure 1 , an 8K FOV content playback method for reducing head-turning latency provided in this embodiment, the technical solution includes the following steps: S1: After the user turns the head, obtain the new perspective coordinate position, and download the high-definition video content of the new visible area chunks to form a sharded high-definition file; S2: Parse the new chunked high-definition file through a new file parser and perform independent decoding on each of the chunked high-definition files; S3: Render to a new area of the display texture according to the new chunk position information, making the original area high-definition video and the new area high-definition video visible simultaneously.

[0023] Among them, in step S1, the obtaining of the new perspective coordinate position further includes: Use a sensor to track the rotation of the user's head to obtain head rotation information. In a virtual reality (VR) or augmented reality (AR) scenario, sensors (such as gyroscopes, accelerometers, etc.) are usually used to track the rotation of the user's head. The new perspective direction vector The calculation formula is expressed as follows: Among them, is the rotation matrix represented by the rotation information, is the user's previous perspective direction vector; According to the coordinate system of the scene and the user's position information , combined with the new perspective direction vector , determine the boundary of the new visible area. Among them, the visible area is a cone with the user's position as the vertex, and the half apex angle of the cone is , then a point within the visible area satisfies . By traversing all the chunk positions in the scene, the new visible area chunks are determined.

[0024] Then, downloading the new visible area chunk high-definition video content to form a chunked high-definition file further includes: The server stores the high-definition video chunk data of the entire scene. Each chunk includes a corresponding coordinate identifier. Send a download request to the server according to the coordinates of the new visible area chunk, and the server retrieves and returns the corresponding chunk data according to the request.

[0025] In this embodiment, Specifically, using the HTTP protocol, the requested URL can include the coordinate information of the chunk, and the server retrieves and returns the corresponding chunk data according to the request.

[0026] Then, the parsing of the new chunked high-definition file through the new file parser in step S2 includes file format parsing and data format parsing, and further includes: The file parser reads the meta information of the high-definition file header, including video coding format, resolution, and frame rate; The data inside the segmented high-definition file is organized according to a specific data structure, which is used to store different parts of video frames (such as I-frames, P-frames, B-frames, etc.). The file parser gradually parses the data of each video frame according to this data structure and stores it in a specific data structure in memory for subsequent decoding use.

[0027] In step S2, further independently decoding each of the segmented high-definition files includes: S21: Select a corresponding decoding algorithm according to the parsed video coding format (such as H.264, H.265, etc.) to perform independent decoding of each block. In this embodiment, taking H.264 as an example, its decoding process mainly includes steps such as entropy decoding, inverse quantization, inverse transformation, intra-frame prediction, and inter-frame prediction.

[0028] S22: For the current block, convert the encoded symbols into quantization coefficients through the entropy decoding algorithm from the parsed video frame data. In this embodiment, through the CAVLC algorithm, according to the predefined code table, map the received bitstream to the corresponding quantization coefficients; S23: Perform inverse quantization on the quantization coefficients obtained by entropy decoding to recover the transform coefficients, and then convert the transform coefficients back to pixel values in the spatial domain through inverse transformation. In this embodiment, we use the inverse discrete cosine transform DCT to convert the transform coefficients back to pixel values in the spatial domain. Assume the quantization coefficient is and the quantization step size is then the coefficient after inverse quantization is and then obtain the pixel value through inverse transformation; S24: For intra-frame prediction, predict the pixel values of the current block according to the decoded pixel values around the current block, and use the similar blocks in the reference frame to predict the pixel values of the current block; for inter-frame prediction, through motion estimation and motion compensation, use the similar blocks in the reference frame to predict the pixel values of the current block; S25: Finally, add the predicted pixel values of the current block to the residuals obtained by inverse transformation to obtain the final decoded pixel values.

[0029] Specifically, for all video frames in each new high-definition shard file, independent decoding is performed according to the above steps.

[0030] Next, in step S3, rendering to a new area of the display texture according to the new segmented position information further includes: Convert the boundary coordinates of the new segment in the scene to the screen coordinate space through projection transformation to obtain the coordinates of the segment in the scene , and obtain the boundary coordinates of the new segment on the screen through calculation. The calculation formula is as follows: , , , Among them, and are the coordinate values after projection transformation. After perspective division, the coordinates projected onto the two-dimensional plane are obtained , is an intermediate value for perspective division, represents the perspective projection matrix, and then the coordinates are mapped to the screen coordinate range and , and the calculation formula is as follows: Among them, is the resolution of the display screen, is the approximate coordinate position of this block on the screen finally; According to the screen coordinates of this block, determine its texture coordinate range on the display texture, and complete the mapping conversion from screen coordinates to texture coordinates. In this embodiment, it is assumed that the rectangular area occupied by the video block on the screen, the upper left screen coordinate is , and the lower right screen coordinate is , the width and height of the texture image are and , for any point on the screen within this block of the screen coordinates , its corresponding texture coordinates are calculated as follows: In this way, the mapping relationship from screen coordinates to texture coordinates is obtained. Through this method, we can accurately render the decoded video frame data (texture) to the corresponding area on the screen.

[0031] In addition, step S3 further includes: During the rendering process, fuse the original texture area and the newly rendered texture area. Through transparency blending, set an appropriate transparency value to ensure that the new texture area and the new texture area are naturally blended and displayed, so that the user can see the high-definition video content of both the old and new areas at the same time.

[0032] This method realizes play control management by optimizing the play process, and further includes: Before obtaining the new coordinate position, first detect whether there is a head turn in the downloaded thread. Once a head turn is detected, trigger the operation of obtaining the new coordinate position; When it is determined that during a head turn, the time to the end of the currently playing time period is too short, the head turn behavior is ignored to avoid unnecessary new content downloading and processing caused by the head turn when the playback is about to end and the remaining time is short, thereby reducing resource waste. Before downloading a new chunk of high-definition file, it is determined whether the new chunk of high-definition file already exists within the current playing time period. If it does not exist and has not been downloaded, then the new chunk of high-definition file is downloaded.

[0033] Through the above method, this embodiment provides a supplementary method for a new VR player to display high-definition videos, which quickly shortens the head turn latency and improves the user viewing experience. By using new sharded files to supplement the displayed video, the effect of shortening the high-definition video display is achieved. When watching a high-definition video, when the user turns their head to view content in other areas, the original playing area is maintained, and at the same time, the video content of the new area is downloaded, decoded by software, and the new video frame is supplemented to the texture, so that the user can quickly see the new part of the video. The waiting time is shortened from the original 1000 - 2000 milliseconds to 500 - 1000 milliseconds. As Figure 3 shown.

[0034] This embodiment also provides an 8K FOV content playback system for reducing head turn latency, including: A new perspective information acquisition module, used to acquire the new perspective coordinate position after the user turns their head, and download the high-definition video content of the new visible area to form a chunk of high-definition file; A parsing and decoding module, used to parse the new chunk of high-definition file through a new file parser and perform independent decoding on each of the chunk of high-definition files one by one; A rendering management module, used to quickly obtain the position information of the decoded chunks, and render the decoded video content to a new area of the display texture according to the position information, so that the high-definition video in the original area and the high-definition video in the new area are both visible.

[0035] Among them, the function implementation of each module and unit corresponds to each step in the above-mentioned embodiment of the 8K FOV content playback method for reducing head turn latency, and its function and implementation process will not be elaborated here one by one.

[0036] Finally, it should be noted that the above description is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention.

[0037] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

Claims

1. A method for playing 8K FOV content with reduced head-turning delay, characterized in that: The steps include: S1: After the user turns his head, the new view coordinate position is obtained, and the high-definition video content of the new visible area is downloaded to form a high-definition file; S2: parsing the new block-based high-definition file by a new file parser, and independently decoding the block-based high-definition files one by one; S3: Quickly obtain the position information of the decoded blocks, and render the decoded video content to a new area of ​​the display texture according to the position information, so that the original area high-definition video and the new area high-definition video are visible at the same time.

2. The 8K FOV content playback method with reduced head-turning delay according to claim 1, characterized in that: In step S1, obtaining a new viewing angle coordinate position further includes: Use sensors to track the user's head rotation to obtain head rotation information and the new viewing direction vector The calculation formula is as follows: in, is the rotation matrix represented by the rotation information, is the user's previous viewing direction vector; According to the scene's coordinate system and the user's location information , combined with the new viewing direction vector , determine the boundary of the new visible area, where the visible area is a cone with the user position as the vertex, and the semi-vertex angle of the cone is , then a point in the visible area satisfy , by traversing all the block positions in the scene to determine the new visible area block.

3. The 8K FOV content playback method with reduced head-turning delay according to claim 2, characterized in that: In step S1, downloading new visible area block high-definition video content to form block high-definition files further includes: The server stores high-definition video block data of the entire scene, each block includes a corresponding coordinate identifier, and sends a download request to the server according to the coordinates of the new visible area block. The server retrieves and returns the corresponding block data according to the request.

4. The 8K FOV content playback method with reduced head-turning delay according to claim 1, characterized in that: In step S2, parsing the new block high-definition file by a new file parser includes file format parsing and data format parsing, further comprising: The file parser reads the meta information in the header of the high-definition file, including the video encoding format, resolution and frame rate; The data inside the block HD file is organized according to a specific data structure for storing different parts of the video frame. The file parser gradually parses the data of each video frame according to this data structure and stores it in a specific data structure in the memory for subsequent decoding.

5. The 8K FOV content playback method with reduced head-turning delay according to claim 4, characterized in that: In step S2, independently decoding the block-by-block high-definition files one by one further comprises: S21: selecting a corresponding decoding algorithm according to the parsed video encoding format to perform independent decoding of each block; S22: for the current block, converting the encoded symbols into quantization coefficients from the parsed video frame data through an entropy decoding algorithm; S23: De-quantizing the quantized coefficients obtained by entropy decoding to restore transform coefficients, and then converting the transform coefficients back to pixel values ​​in the spatial domain through inverse transformation; S24: predicting the pixel value of the current block according to the decoded pixel values ​​around the current block, and predicting the pixel value of the current block using similar blocks in the reference frame; S25: Finally, the predicted pixel value of the current block is added to the residual obtained by inverse transformation to obtain the final decoded pixel value.

6. The 8K FOV content playback method with reduced head-turning delay according to claim 1, characterized in that: In step S3, rendering to a new area of ​​the display texture according to the new block position information further includes: The boundary coordinates of the new block in the scene are converted to the screen coordinate space through projection transformation, and the coordinates of the block in the scene are obtained. , the boundary coordinates of the new block on the screen are obtained by calculation, and the calculation formula is as follows: , , , in, and is the coordinate value after projection transformation, and after perspective division, the coordinate projected onto the two-dimensional plane is obtained , is an intermediate value used for perspective division, Represents the perspective projection matrix, and then maps the coordinates to the screen coordinate range and , the calculation formula is as follows: in, To display the screen resolution, The approximate coordinate position of the block on the screen; Determine the texture coordinate range of the block on the display texture according to the screen coordinates of the block, and complete the mapping conversion from the screen coordinates to the texture coordinates; The decoded video frame data is used as texture data and rendered onto the corresponding area of ​​the display texture according to the calculated texture coordinate range.

7. The 8K FOV content playback method with reduced head-turning delay according to claim 6, characterized in that: Step S3 further comprises: During the rendering process, the original texture area and the newly rendered texture area are merged, and through transparency blending, the appropriate transparency value is set to ensure that the new texture area and the new texture area are naturally merged and displayed, so that users can see the high-definition video content of the new and old areas at the same time.

8. The 8K FOV content playback method with reduced head-turning delay according to claim 3, characterized in that: Optimizing the playback process to achieve playback control management further includes: Before obtaining the new coordinate position, first detect whether the head turns in the download thread. Once the head turns are detected, the operation of obtaining the new coordinate position is triggered; Determine the end time of the playing time period when the head turn occurs. If the time is too short, ignore the head turn behavior to avoid unnecessary downloading and processing of new content caused by the head turn when the playback is about to end and the remaining time is too short, thereby reducing resource waste; Before downloading a new block HD file, it is determined whether the new block HD file already exists in the current playback time period. If it does not exist and has not been downloaded, the new block HD file is downloaded.

9. An 8K FOV content playback system with reduced head-turning delay, characterized in that: include: A new viewing angle information acquisition module is used to obtain the new viewing angle coordinate position after the user turns his head, and download the new visible area block HD video content to form a block HD file; A parsing and decoding module, used for parsing the new block-based high-definition files through a new file parser, and independently decoding the block-based high-definition files one by one; The rendering management module is used to quickly obtain the position information of the decoded blocks, and render the decoded video content to a new area of ​​the display texture according to the position information, so that the original area high-definition video and the new area high-definition video are visible at the same time.