Panoramic video playback method and apparatus, device, and storage medium

By obtaining the shooting device's posture information from VR videos to determine the target area and performing super-resolution processing, the problem of low-quality VR video playback experience is solved, achieving high-quality video playback while reducing performance consumption.

WO2025218295A1PCT designated stage Publication Date: 2025-10-23BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2025/072781
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-15
Filing Date
2025-01-16
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Low-quality VR videos result in lower clarity during playback, affecting the viewing experience. Furthermore, super-resolution processing consumes a lot of performance resources, which can easily lead to problems such as video stuttering, slow playback, and frame drops.

Method used

By acquiring the shooting device's pose information from the original video frames in the target panoramic video, the target area is determined, and the target texture is obtained using a mapping model. After super-resolution processing, the super-resolution texture and the original texture are fused to generate the target video frame for playback.

Benefits of technology

It improves video quality, reduces performance resource consumption, minimizes stuttering, slow startup, and frame drops, and enhances the playback experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025072781_23102025_PF_FP_ABST
    Figure CN2025072781_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a panoramic video playback method and apparatus, a device, and a storage medium. The panoramic video playback method comprises: acquiring photographing device orientation information corresponding to an original video frame in a target panoramic video; determining a target area of the original video frame on the basis of the photographing device orientation information, and on the basis of the target area, acquiring a target texture on a mapping model corresponding to the original video frame; and performing super-resolution processing on the target texture to obtain a super-resolution texture, fusing the super-resolution texture and an original texture to obtain a target video frame, and playing back the target video frame. Since a target area represents an audience focusing area, super-resolution processing is only performed on the target area, thereby reducing the performance resource overhead, reducing the problems of stuttering, slow playback starting, frame dropping, etc., and enabling audience to watch videos having high picture quality, thus improving the playback experience.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and storage medium for playing panoramic video

[0001] This application claims priority to Chinese Patent Application No. 202410452487.0, filed on April 15, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD

[0002] Embodiments of the present disclosure relate to a method, device, equipment and storage medium for playing panoramic video. BACKGROUND

[0003] Panoramic video, also known as Virtual Reality-video (VR video for short), is a video that can realize three-dimensional space display function. The VR video can display 360-degree panoramic image for the audience, improving the user viewing experience.

[0004] Currently, for low-quality VR video, there is a problem of low definition affecting the playback experience. In order to improve the playback experience, the low-quality VR video is usually processed by super-resolution, however, the virtual reality device consumes a large amount of performance resources, which is easy to cause video lag, slow start and frame loss and other problems. SUMMARY

[0005] The present disclosure provides a method, device, equipment and storage medium for playing panoramic video, which can improve the playback experience and reduce the performance resource consumption.

[0006] In a first aspect, the embodiments of the present disclosure provide a method for playing panoramic video, comprising:

[0007] obtaining shooting device pose information corresponding to an original video frame in a target panoramic video;

[0008] determining a target region of the original video frame according to the shooting device pose information, obtaining a target texture on a mapping model corresponding to the original video frame according to the target region, wherein a surface of the mapping model has a texture corresponding to the original video frame;

[0009] performing super-resolution processing on the target texture to obtain a super-resolution texture, fusing the super-resolution texture and an original texture to obtain a target video frame, and playing the target video frame, wherein the original texture represents a texture in the original video frame that does not belong to the target region.

[0010] In a second aspect, the embodiments of the present disclosure also provide a device for playing panoramic video, comprising:

[0011] an information obtaining module, configured to obtain shooting device posture information corresponding to an original video frame in a target panoramic video;

[0012] a texture obtaining module, configured to determine a target region of the original video frame according to the shooting device posture information, and obtain a target texture on a mapping model corresponding to the original video frame according to the target region, wherein a surface of the mapping model has a texture corresponding to the original video frame;

[0013] a texture fusion module, configured to perform super-resolution processing on the target texture to obtain a super-resolution texture, fuse the super-resolution texture and an original texture to obtain a target video frame, and play the target video frame, wherein the original texture represents a texture in the original video frame that does not belong to the target region.

[0014] In a third aspect, an electronic device is provided, and the electronic device includes:

[0015] one or more processors;

[0016] a storage device configured to store one or more programs,

[0017] when the one or more programs are executed by the one or more processors, the one or more processors implement a method for playing a panoramic video as described in any of the embodiments of the present disclosure.

[0018] In a fourth aspect, a storage medium containing computer executable instructions is provided, and the computer executable instructions, when executed by a computer processor, are used to perform a method for playing a panoramic video as described in any of the embodiments of the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0019] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.

[0020] FIG. 1 is a flow diagram of a method for playing a panoramic video according to an embodiment of the present disclosure;

[0021] FIG. 2 is a schematic diagram of a mapping model according to an embodiment of the present disclosure;

[0022] FIG. 3 is a flow diagram of another method for playing a panoramic video according to an embodiment of the present disclosure;

[0023] FIG. 4 is a schematic diagram of a device structure for playing a panoramic video according to an embodiment of the present disclosure; and

[0024] FIG. 5 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] Embodiments of the present disclosure will be described in more detail with reference to the drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art.

[0026] It should be understood that each step described in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0027] The term "comprising" and variations thereof as used herein are open-ended, that is, "comprising but not limited to." The term "based on" is "based, at least in part, on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related definitions are given below in the description of the application.

[0028] It should be noted that the terms "first", "second", and the like in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0029] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.

[0030] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0031] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0032] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will require obtaining and using personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware, such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solution of the present disclosure according to the prompt information.

[0033] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in a text manner. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.

[0034] It can be understood that the above notification and obtaining user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0035] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the obtaining or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0036] FIG. 1 is a flow diagram of a method for playing a panoramic video according to an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the case of playing a VR video, such as playing a VR video on a virtual reality device, such as a VR glasses or a VR helmet. Alternatively, the VR video is played on a smart phone, etc. The method can be performed by a device for playing a panoramic video. The device can be implemented in the form of software and / or hardware, and can be implemented by an electronic device, such as a smart phone or a virtual reality device, etc.

[0037] As shown in FIG. 1, the method comprises:

[0038] S110, obtaining shooting device pose information corresponding to an original video frame in a target panoramic video.

[0039] The target panoramic video represents a VR video to be played on an electronic device. Hereinafter, the VR video can refer to the target panoramic video. In the embodiment of the present disclosure, the VR video can be played by a client configured on the electronic device. The interactive page of the client can display a plurality of candidate panoramic videos for the user to select the target panoramic video therefrom. Alternatively, the interactive page of the client can display a search control for the user to search for the target panoramic video based on the search control.

[0040] In some embodiments, the VR video can include a panoramic video. The panoramic video can convert a static panoramic picture into a dynamic video image. The panoramic video can be viewed at any angle within 360 degrees up, down, left and right in the shooting angle to provide a sense of being there.

[0041] Optionally, the target panoramic video can also be a low-resolution gear corresponding to a high-resolution VR video selected by the user. For a high-resolution VR video, such as an 8K or 4K resolution VR video, the resolution is too large and can cause problems such as lag, slow start, and high memory consumption during playback. In addition, the code rate of the high-resolution VR video is also high, which increases the consumption of bandwidth resources. Therefore, after the user selects the high-resolution gear of the panoramic video, a lower resolution target panoramic video can be selected at the start.

[0042] Since the VR video can represent that the audience is wrapped in a sphere, the visual range of the audience can be considered as the front area of the current shooting device. The shooting device here is a virtual shooting device, and the shooting device pose information represents the pose of the virtual shooting device. Different shooting device poses correspond to different visual ranges. For example, the shooting device pose information includes a shooting device front direction vector. In the VR video playback process, the shooting device front direction vector is obtained in the target function of the target engine. For example, the shooting device front direction vector of the current shooting device pose information is obtained in the frame function update of the target engine. The target engine can be used to generate panoramic videos and perform super-resolution processing on panoramic videos, etc.

[0043] Illustratively, the shooting device pose information corresponding to the original video frame in the target panoramic video is obtained, including: analyzing a target function to obtain a shooting device front direction vector of the original video frame in the target panoramic video, wherein the target function is used to determine a shooting device pose corresponding to the target panoramic video.

[0044] Since the VR video provides a virtual reality scene (i.e., a VR scene) for the audience, if VR is enabled in Unity, all cameras in the VR scene can be directly rendered to a head-mounted display. In the VR scene, the human eye is only sensitive to the content in the front area of the shooting device, so the front direction vector of the current shooting device view angle can be obtained from the frame function update to determine the front area of the shooting device based on the shooting device front direction vector. It should be noted that the shooting device pose corresponding to the panoramic video is updated to the frame function update in real time.

[0045] S120, determining a target area of the original video frame according to the shooting device pose information, and obtaining a target texture on a mapping model corresponding to the original video frame according to the target area.

[0046] The surface of the mapping model has the texture corresponding to the original video frame. For example, the mapping model can include an equirectangular projection (ERP) mapping model and an equi-angular cubemap (EAC) mapping model.

[0047] In the embodiments of the present disclosure, the target region represents a region corresponding to the front of the shooting device in the mapping model of each original video frame of the target panoramic video. Since the human eye is usually sensitive to the content of the region in front of the shooting device, the target region can be determined based on the front direction of the shooting device at the current shooting device perspective. The mapping model can represent a model carrying the panoramic content corresponding to the original video frame. The texture corresponding to the original video frame can be projected onto the surface of the mapping model. Projection represents the process of unfolding the real scene of the full physical view to a 2D picture and restoring it to the VR device to achieve immersive viewing. For example, the projection mode can include cylindrical projection and octahedral projection. The cylindrical projection directly projects the spherical surface on the plane without the help of intermediate projection geometry. For example, the ERP mapping model is used to convert the spherical content to the plane. Alternatively, the cuboid projection format is a projection mode by unfolding each face of the cuboid model after projecting the spherical content on the cuboid model, and then splicing into a rectangle. The cuboid projection realizes the mapping from the spherical surface to the cuboid face through perspective. For example, the EAC mapping model first deforms the spherical film into a cuboid, then flattens the cuboid into a plane composed of six squares, and in the projection process, the angle corresponding to each section is adjusted to make each projected section have close pixel density.

[0048] The target texture represents the texture corresponding to the target region, and the target texture can be obtained by intercepting the texture corresponding to the target region on the mapping model.

[0049] Exemplarily, according to the shooting device posture information, the target region of the original video frame is determined, and the target texture on the mapping model corresponding to the original video frame is obtained according to the target region, including: determining the polar angle and azimuth angle of the intersection point of the front direction of the shooting device and the mapping model; determining the intersection point texture coordinates corresponding to the intersection point according to the polar angle and azimuth angle; determining the target region according to the size setting information of the target region and the intersection point texture coordinates, and obtaining the texture on the mapping model as the target texture according to the target region.

[0050] The video frames of the VR video are projected to a plane by a mapping model to be played by a virtual reality device. The shooting device front direction amount of each shooting device view angle intersects with the mapping model. If the mapping model is an ERP mapping model, for a certain shooting device view angle, the intersection point in the spherical coordinate system can be determined according to the three-dimensional coordinates of the intersection point of the shooting device front direction amount of the current shooting device view angle and the spherical model. The polar angle and the azimuth angle represent the latitude and the longitude of the intersection point. The intersection point texture coordinates (u, v) of the intersection point are calculated based on the polar angle and the azimuth angle by using a set formula. The user can set the size of the target region through an interactive page to obtain size setting information. For example, the user can set a rectangular region of w*h through the interactive page, and w and h are the size setting information of the target region. The target region is determined in combination with the size setting information of the target region and the intersection point texture coordinates. The corresponding texture region in the mapping model is searched based on the target region. In combination with the rendering to texture technology, the texture on the surface of the mapping model is cut based on the texture region to obtain the target texture. The rendering to texture technology allows the developer to create a texture map according to part of the image of the scene, which is saved for future deferred shading, multi-channel rendering, or more advanced rendering effects, etc.

[0051] Optionally, determining the target region according to the size setting information of the target region and the intersection point texture coordinates comprises: taking the intersection point as the center of the target region, determining the upper left corner coordinate of the target region according to the intersection point texture coordinates and the length setting information and the width setting information of the target region; and determining the target region according to the upper left corner coordinate, the length setting information and the width setting information.

[0052] For example, the intersection point is taken as the center of the target region, and the target region is formed by spreading around the intersection point. Specifically, the difference between the horizontal direction coordinate value of the intersection point texture coordinates and half of the length value corresponding to the length setting information is calculated as the horizontal direction coordinate value u' of the upper left corner vertex of the target region, that is, u' = u-w / 2. The difference between the vertical direction coordinate value of the intersection point texture coordinates and half of the width value corresponding to the width setting information is calculated as the vertical direction coordinate value v' of the upper left corner vertex of the target region, that is, v' = v-h / 2. The target region can be represented as (u', v', w, h). In combination with the rendering to texture technology, the texture on the surface of the mapping model is cut based on the target region (u', v', w, h) to obtain the target texture.

[0053] In an optional implementation, the target region of the original video frame is determined according to the shooting device posture information, and the target texture on the mapping model corresponding to the original video frame is obtained according to the target region, including: dividing each surface of the mapping model into a set number of sub-surfaces, and determining the texture region of each sub-surface; for each sub-surface of the mapping model, determining a target vector according to a set coordinate point on the sub-surface; determining a target sub-surface set according to the included angle between the front direction of the shooting device and the target vector corresponding to each sub-surface, and determining a target region according to the target sub-surface set; and obtaining the texture corresponding to the texture region of each target sub-surface in the target sub-surface set as the target texture.

[0054] If the mapping model is an EAC mapping model, each surface of the mapping model is divided into a set number of sub-surfaces. FIG. 2 is a schematic diagram of a mapping model provided by an embodiment of the present disclosure. As shown in FIG. 2, the projection mode of the EAC mapping model is to convert the texture of the surface of the spherical model 210 to the surface of the cubic model 220. Each face of the cubic model 220 is divided into four sub-surfaces, and an index is set. The texture region corresponding to each sub-surface is calculated, and the index and the texture region of the same sub-surface are stored in association. Optionally, the texture region can be represented by the texture coordinate of the upper left corner, the length of the texture region, and the width of the texture region. The length of the texture region can be determined based on the length of the sub-surface, and the width of the texture region can be determined based on the width of the sub-surface. The coordinates of the center point of each sub-surface are obtained. For each sub-surface of the cubic model, a target vector OA' is formed based on the center point A' of the current sub-surface and the origin O of the spherical model 210. Assuming that the front direction of the shooting device of the current shooting device view angle is OV, the included angle between the front direction of the shooting device OV and the target vector OA' is calculated. The included angle between the target vector corresponding to each sub-surface and the front direction of the shooting device OV can be calculated in the same way, and then the indexes of the sub-surfaces are arranged in ascending order according to the included angles, and a set number of indexes in the front are selected. If the indexes of the sub-surfaces are arranged in descending order according to the included angles, a set number of indexes in the rear are selected. The target sub-surface set is formed according to the sub-surfaces corresponding to the selected indexes. According to the indexes corresponding to each target sub-surface in the target sub-surface set and the association between the indexes and the texture regions, the texture region of each sub-surface is obtained. Based on the rendering-to-texture technology, the texture of the surface of the mapping model is cropped based on the texture region to obtain the target texture.

[0055] It should be noted that the number of sub-surfaces divided from each surface is not limited in the embodiments of the present disclosure, and each surface of the EAC mapping model can be divided into more sub-surfaces. In the embodiments of the present disclosure, some functions, components, models, etc. in the prior art can be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility in the implementation of the technical solutions of the present disclosure, but does not mean that the applicant has or will necessarily use them in the present disclosure.

[0056] S130, performing super-resolution processing on the target texture to obtain a super-resolution texture, fusing the super-resolution texture and the original texture to obtain a target video frame, and playing the target video frame.

[0057] The original texture represents a texture in the original video frame that does not belong to the target region.

[0058] Super-resolution processing refers to a process of restoring image details and other data information from known image information using optical and related optical knowledge. Super-resolution is used to increase the resolution of an image to prevent image quality degradation. In video playback, through super-resolution processing, the video quality can be significantly improved under the same resolution.

[0059] The super-resolution texture represents a new texture with improved resolution obtained by performing super-resolution processing on the target texture. For example, the pixels of the target region are super-resolved from 4K to 8K through super-resolution technology. The super-resolution result is represented by the target parameter of the target engine's on-screen function. If the super-resolution result is successful, the super-resolution texture and the original texture are fused and on-screen. If the super-resolution result fails, the original texture is used for on-screen. The on-screen function can represent a function for displaying a VR video. On-screen can be an operation for displaying a VR video.

[0060] Illustratively, fusing the super-resolution texture and the original texture to obtain a target video frame includes: obtaining a texture coordinate of a pixel to be rendered in the original video frame; if the texture coordinate belongs to the target region, obtaining a super-resolution texture corresponding to the texture coordinate as the texture of the pixel to be rendered; if the texture coordinate does not belong to the target region, obtaining an original texture corresponding to the texture coordinate as the texture of the pixel to be rendered; and determining a target video frame based on the texture of each pixel to be rendered in the original video frame.

[0061] For example, in the process of playing a VR video, a three-dimensional coordinate of a pixel to be rendered in an original video frame is converted into a texture coordinate by a vertex shader, and the texture coordinate is transmitted to a fragment shader. The fragment shader determines whether the texture coordinate belongs to the target region. If so, the super-resolution texture is sampled based on the texture coordinate, otherwise the original texture is sampled based on the texture coordinate to determine the texture of the pixel to be rendered. The pixels are shaded based on the texture of each pixel to be rendered in the original video frame to obtain a target video frame, which is on-screen to play the target video frame.

[0062] The technical scheme of the embodiments of the present disclosure determines the target region of the original video frame through the shooting device posture information, then obtains the target texture on the mapping model corresponding to the original video frame according to the target region, performs super-resolution processing on the target texture, in the playing process, fuses the super-resolution texture and the original texture corresponding to each original video frame to obtain the target video frame, and plays the target video frame. Since the target region represents the audience attention region, only the target region is subjected to super-resolution processing, the performance resource consumption is reduced, the problems such as lag, slow start and frame loss are reduced, and the audience can also watch a video with high quality, and the playing experience is improved.

[0063] FIG. 3 is a flowchart of another method for playing a panoramic video provided by an embodiment of the present disclosure. The embodiment of the present disclosure additionally limits the downshift super-resolution scheme on the basis of the above-mentioned embodiments. Downshift refers to reducing the resolution level of the VR video selected by the user.

[0064] As shown in FIG. 3, the method comprises:

[0065] In S310, if a panoramic video selection event is detected, the target resolution of the selected panoramic video is obtained.

[0066] For some video resources with multiple resolutions of VR videos, the multiple resolution levels can be displayed through the interactive page of the VR client when the user on-demand plays the VR video, so as to be selected by the user. The selection operation of the user on the VR video can trigger a VR video selection event. The VR client determines the target resolution of the selected VR video according to the selection operation.

[0067] In S320, it is judged whether the target resolution meets a preset super-resolution condition. If yes, S330 is executed, otherwise S340 is executed.

[0068] The preset super-resolution condition is used to determine whether the VR video can be subjected to a downshift super-resolution operation. The downshift super-resolution operation can refer to obtaining an original VR video with a resolution lower than the resolution of the VR video selected by the user, and performing super-resolution processing on a local region in the original VR video according to the shooting device posture information in the playing process of the original VR video, so as to make the local region of the original VR video reach a high-resolution quality.

[0069] In the VR scene, the high-resolution VR video is prone to problems such as lag, slow start and high memory consumption when playing. In addition, the code rate of the high-resolution VR video is also high, which increases the consumption of bandwidth resources. By using the downshift super-resolution, the low-resolution VR video can reach a high-resolution quality, so that a video stream with a lower code rate can be used on the basis of the same quality, so that the playing experience is smoother, and the bandwidth resource consumption and bandwidth cost are also reduced.

[0070] For example, if the VR video selected by the user has multiple resolution levels, and the user selects a higher resolution level, it can be determined that the preset super-resolution condition is met. Alternatively, whether the VR video supports downshift super-resolution is represented by a target identifier, and after the user selects the VR video, it is determined whether the preset super-resolution condition is met according to the corresponding target identifier. Alternatively, whether to enable the function of downshift super-resolution for the target panoramic video is controlled by a preset button.

[0071] S330, according to the identifier information of the selected panoramic video, the target panoramic video with a resolution lower than the target resolution is obtained.

[0072] For example, for the case where the preset super-resolution condition is met, the original VR video with the same low resolution as the selected VR video is obtained.

[0073] For example, the user selects a VR video with an 8K level to start playing, and the VR video has a 4K video resource. The 4K video resource can be obtained as the original VR video. The VR client plays the 4K VR video, and in the process of playing, local super-resolution is performed according to the shooting device posture information, so that the local area of each video frame in the VR video reaches high-resolution quality.

[0074] In some embodiments, a player for playing a VR video includes an on-demand SDK, a VR client, and a target engine. If super-resolution is enabled, the shooting device posture information is obtained in the frame function update of the target engine. The shooting device posture information includes a front direction amount of the shooting device. The intersection texture coordinates at the intersection of the front direction amount and the spherical model are calculated according to the front direction amount. The target area is determined based on the intersection texture coordinates. For example, the target area can be a rectangle, which is represented by the top-left corner texture coordinates (u, v) and the length w and the width h of the rectangle. Wherein, u, v, w, h ∈ (0, 1). The target texture in the target area is cropped by combining the rendering-to-texture technology. The target texture is super-resolution processed to obtain a super-resolution texture. The super-resolution result is stored in the on-screen function. If super-resolution is not enabled, the original texture is used for on-screen. The target engine returns the shooting device posture information to the VR client. The VR client transmits the shooting device posture information to the GPU and triggers the rendering texture event. The on-demand SDK responds to the rendering texture event by executing the code of the fragment shader through the rendering thread. The super-resolution result is obtained. If the super-resolution is successful, for the texture coordinates in the target area, the super-resolution texture is sampled, and for the texture coordinates not in the target area, the original texture is sampled, to realize the fusion on-screen of the super-resolution texture and the original texture. If the super-resolution fails, the original texture is used for on-screen.

[0075] S340, the selected panoramic video is taken as the target panoramic video.

[0076] S350, acquire shooting device posture information corresponding to the original video frame in the target panoramic video.

[0077] S360, determine a target region of the original video frame according to the shooting device posture information, and acquire target texture on a mapping model corresponding to the original video frame according to the target region.

[0078] S370, perform super-resolution processing on the target texture to obtain super-resolution texture, fuse the super-resolution texture and the original texture to obtain a target video frame, and play the target video frame.

[0079] The technical scheme of the embodiment of the present disclosure acquires a target VR video with a resolution lower than a target resolution corresponding to a selected gear of a VR video when the VR video is played. In the process of playing the VR video, a local region of an original video frame in the target VR video is determined according to shooting device posture information of the original video frame, a target texture in the local region is cropped for super-resolution processing to obtain super-resolution texture, and the super-resolution texture and the original texture are fused for on-screen display, thereby reducing the stuttering rate, the non-playing rate, the playing bandwidth cost, and the memory consumption, etc.

[0080] FIG. 4 is a schematic diagram of a device structure for playing panoramic video provided by an embodiment of the present disclosure. The device can be implemented in the form of software and / or hardware, and can be implemented by an electronic device such as a smart phone or a virtual reality device.

[0081] As shown in FIG. 4, the device includes an information acquisition module 410, a texture acquisition module 420, and a texture fusion module 430.

[0082] The information acquisition module 410 is configured to acquire shooting device posture information corresponding to an original video frame in a target panoramic video.

[0083] The texture acquisition module 420 is configured to determine a target region of the original video frame according to the shooting device posture information, and acquire target texture on a mapping model corresponding to the original video frame according to the target region, wherein a surface of the mapping model has texture corresponding to the original video frame.

[0084] The texture fusion module 430 is configured to perform super-resolution processing on the target texture to obtain super-resolution texture, fuse the super-resolution texture and the original texture to obtain a target video frame, and play the target video frame, wherein the original texture represents texture in the original video frame that does not belong to the target region.

[0085] Optionally, the information acquisition module 410 is specifically configured to:

[0086] The target function is analyzed to obtain a front direction amount of the shooting device of the original video frame in the target panoramic video, wherein the target function is used to determine the shooting device pose corresponding to the target panoramic video.

[0087] Optionally, the texture obtaining module 420 is specifically configured to:

[0088] determine a polar angle and an azimuth angle of the intersection point of the front direction amount of the shooting device and the mapping model;

[0089] determine an intersection point texture coordinate corresponding to the intersection point according to the polar angle and the azimuth angle;

[0090] determine the target region according to the size setting information and the intersection point texture coordinate of the target region, and obtain the texture on the mapping model as the target texture according to the target region.

[0091] Further, the determination of the target region according to the size setting information and the intersection point texture coordinate of the target region comprises:

[0092] take the intersection point as the center of the target region, and determine the top-left corner coordinate of the target region according to the intersection point texture coordinate and the length setting information and the width setting information of the target region;

[0093] determine the target region according to the top-left corner coordinate, the length setting information and the width setting information.

[0094] Optionally, the texture obtaining module 420 is further specifically configured to:

[0095] divide each surface of the mapping model into a set number of sub-surfaces, and determine the texture region of each sub-surface;

[0096] for each sub-surface of the mapping model, determine a target vector according to a set coordinate point on the sub-surface;

[0097] determine a target sub-surface set according to the included angle between the front direction amount of the shooting device and the target vector corresponding to each sub-surface, and determine a target region according to the target sub-surface set;

[0098] obtain the texture corresponding to the texture region of each target sub-surface in the target sub-surface set as the target texture.

[0099] Optionally, the texture fusion module 430 is specifically configured to:

[0100] obtain the texture coordinate of the pixel to be rendered in the original video frame;

[0101] if the texture coordinate belongs to the target region, obtain the super-resolution texture corresponding to the texture coordinate as the texture of the pixel to be rendered;

[0102] If the texture coordinate does not belong to the target region, an original texture corresponding to the texture coordinate is acquired as the texture of the pixel to be rendered.

[0103] The target video frame is determined based on the texture of each pixel to be rendered in the original video frame.

[0104] Optionally, the apparatus further includes:

[0105] The video acquisition module is configured to, if the panoramic video selection event is detected, acquire a target resolution of the selected panoramic video; and if the target resolution satisfies a preset super-resolution condition, acquire the target panoramic video with a resolution lower than the target resolution according to identification information of the selected panoramic video.

[0106] The apparatus for playing a panoramic video provided in the embodiments of the present disclosure can perform the method for playing a panoramic video provided in any of the embodiments of the present disclosure, and has the corresponding function modules and advantages of performing the method.

[0107] It should be noted that each unit and module included in the apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be implemented; in addition, the specific name of each functional unit is only for convenient mutual distinction, and does not serve to limit the protection scope of the embodiments of the present disclosure.

[0108] FIG. 5 is a structural schematic diagram of an electronic device provided in an embodiment of the present disclosure. Referring to FIG. 5, a structural schematic diagram of an electronic device (for example, a smart phone or a virtual reality device in FIG. 5) 500 suitable for implementing the embodiments of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as smart phones, notebook computers, VR glasses, VR headsets, and the like, and fixed terminals such as digital TVs, desktop computers, and the like. The electronic device shown in FIG. 5 is only an example, and should not bring any limitation to the functions and use range of the embodiments of the present disclosure.

[0109] As shown in FIG. 5, the electronic device 500 can include a processing apparatus (for example, a central processing unit, a graphics processing unit, and the like) 501, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or loaded into a random access memory (RAM) 503 from a storage apparatus 508. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing apparatus 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.

[0110] In general, the following devices can be connected to the I / O interface 505: input devices 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 508 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 509. The communication devices 509 can allow the electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG. 5 illustrates the electronic device 500 with various devices, it is understood that all of the illustrated devices are not required to be implemented or possessed. More or fewer devices can be alternatively implemented or possessed.

[0111] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 509, or installed from the storage devices 508, or installed from the ROM 502. When the computer program is executed by the processing devices 501, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.

[0112] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0113] The electronic device provided by the embodiments of the present disclosure and the method of playing a panoramic video provided by the above-described embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiments can be referred to the above-described embodiments, and the present embodiments have the same beneficial effects as the above-described embodiments.

[0114] The embodiments of the present disclosure provide a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method of playing a panoramic video provided by the above-described embodiments.

[0115] It should be noted that the computer readable medium in the above disclosure can be a computer readable signal medium or a computer readable storage medium or any combination of the above two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or component. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or component. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.

[0116] In some embodiments, the client, server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication (e.g., communication networks) of any form or medium, such as the Internet. Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0117] The above computer readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.

[0118] The above computer readable medium carries one or more programs, when the above one or more programs are executed by the electronic device, the electronic device:

[0119] Obtain the shooting device attitude information corresponding to the original video frame in the target panoramic video;

[0120] determine a target region of the original video frame according to the shooting device posture information, and acquire a target texture on a mapping model corresponding to the original video frame according to the target region, wherein a surface of the mapping model has a texture corresponding to the original video frame;

[0121] perform super-resolution processing on the target texture to obtain a super-resolution texture, fuse the super-resolution texture and an original texture to obtain a target video frame, and play the target video frame, wherein the original texture represents a texture in the original video frame that does not belong to the target region.

[0122] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0123] The flow and block diagrams in the drawings show architectural, functional, and operational representations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0124] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. In some cases, the names of the units do not constitute a limitation on the units themselves.

[0125] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0126] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0127] The above description is only preferred embodiments of the present disclosure and the explanation of the technical principles used. Those skilled in the art should understand that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0128] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0129] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A method for playing panoramic video, comprising: obtaining shooting device posture information corresponding to an original video frame in a target panoramic video; determining a target region of the original video frame according to the shooting device posture information, and obtaining target texture on a mapping model corresponding to the original video frame according to the target region, wherein a surface of the mapping model has texture corresponding to the original video frame; performing super-resolution processing on the target texture to obtain super-resolution texture, fusing the super-resolution texture and original texture to obtain a target video frame, and playing the target video frame, wherein the original texture represents texture in the original video frame that does not belong to the target region.

2. The method of claim 1, wherein, The obtaining of the shooting device posture information corresponding to the original video frame in the target panoramic video comprises: parsing a target function to obtain a shooting device front direction vector of the original video frame in the target panoramic video, wherein the target function is used to determine the shooting device posture corresponding to the target panoramic video.

3. The method of claim 2, wherein, The determining of the target region of the original video frame according to the shooting device posture information and the obtaining of the target texture on the mapping model corresponding to the original video frame according to the target region comprise: determining a polar angle and an azimuth angle of an intersection point of the shooting device front direction vector and the mapping model; determining intersection texture coordinates corresponding to the intersection point according to the polar angle and the azimuth angle; determining the target region according to size setting information of the target region and the intersection texture coordinates, and obtaining texture on the mapping model as the target texture according to the target region.

4. The method of claim 3, wherein, The determining of the target region according to the size setting information of the target region and the intersection texture coordinates comprises: taking the intersection point as a center of the target region, and determining a top-left corner coordinate of the target region according to the intersection texture coordinates and length setting information and width setting information of the target region; determining the target region according to the top-left corner coordinate, the length setting information and the width setting information.

5. The method of claim 2, wherein, The determining of the target region of the original video frame according to the shooting device posture information and the obtaining of the target texture on the mapping model corresponding to the original video frame according to the target region comprise: dividing each surface of the mapping model into a set number of sub-surfaces, and determining texture regions of the sub-surfaces; for each sub-surface of the mapping model, determining a target vector according to a set coordinate point on the sub-surface; determining a target sub-surface set according to an included angle between the shooting device front direction vector and the target vector corresponding to each sub-surface, and determining a target region according to the target sub-surface set; obtaining texture corresponding to the texture region of each target sub-surface in the target sub-surface set as the target texture.

6. The method according to any one of claims 1 to 5, wherein, The fusing of the super-resolution texture and the original texture to obtain the target video frame comprises: obtaining texture coordinates of a pixel to be rendered in the original video frame; if the texture coordinates belong to the target region, obtaining super-resolution texture corresponding to the texture coordinates as texture of the pixel to be rendered; if the texture coordinates do not belong to the target region, obtaining original texture corresponding to the texture coordinates as texture of the pixel to be rendered; determining a target video frame based on texture of each pixel to be rendered in the original video frame.

7. The method of any one of claims 1-6, further comprising: if a panoramic video selection event is detected, obtaining a target resolution of a selected panoramic video; if the target resolution satisfies a preset super-resolution condition, obtaining the target panoramic video with a resolution lower than the target resolution according to identification information of the selected panoramic video.

8. An apparatus for playing panoramic videos, comprising: an information obtaining module configured to obtain shooting device pose information corresponding to an original video frame in a target panoramic video; a texture obtaining module configured to determine a target region of the original video frame according to the shooting device pose information, and obtain a target texture on a mapping model corresponding to the original video frame according to the target region, wherein a surface of the mapping model has a texture corresponding to the original video frame; a texture fusion module configured to perform super-resolution processing on the target texture to obtain a super-resolution texture, fuse the super-resolution texture and an original texture to obtain a target video frame, and play the target video frame, wherein the original texture represents a texture in the original video frame that does not belong to the target region.

9. An electronic device, comprising: one or more processors; a storage device configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method for playing panoramic videos according to any one of claims 1-7.

10. A storage medium containing computer-executable instructions, wherein, the computer executable instructions, when executed by a computer processor, perform the method for playing panoramic videos according to any one of claims 1-7.

Citation Information

Patent Citations

  • 3D (three-dimensional) visualization method for coverage range based on quick estimation of attitude of camera

    CN103400409A

  • Video generating method, playing method, video generating device and playing device

    CN107197135A

  • Intelligent panoramic video playing method and device

    CN110324640A

  • Panoramic video rendering method and system

    CN112465939A

  • Method and device for playing panoramic video, equipment and storage medium

    CN118413745A

Cited By

  • Thermal imaging super-resolution method, device and system based on image reconstruction

    CN122415333A