A panoramic video evaluation method, device and electronic equipment

By preprocessing and analyzing panoramic videos, comparable evaluation results are generated, which solves the problem of evaluation difficulties caused by differences in panoramic video configuration parameters and improves the selection of panoramic video transmission strategies and user experience.

CN119342201BActive Publication Date: 2025-11-04BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310905124.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-21
Publication Date
2025-11-04
Estimated Expiration
2043-07-21

AI Technical Summary

Technical Problem

The differences in panoramic video configuration parameters provided by different service providers make it impossible to intuitively compare and evaluate the results, and it is difficult to uniformly evaluate panoramic videos with different configuration parameters.

Method used

By acquiring panoramic video and user perspective, video preprocessing is performed to generate preprocessed panoramic video. Based on the user perspective and preprocessed video, transmission information and bandwidth are determined, and evaluation results are generated. Evaluation is then conducted using perspective prediction and simulated transmission techniques.

Benefits of technology

It enables comparability evaluation of panoramic videos with different configuration parameters, helping to select the best transmission strategy and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119342201B_ABST
    Figure CN119342201B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of data processing, and particularly relates to a panoramic video evaluation method and device and electronic equipment, which are used to solve the problem of how to evaluate panoramic videos with different configuration parameters. The method comprises the following steps: acquiring a panoramic video and a historical user visual angle corresponding to each moment when a user watches the panoramic video; determining a theoretical image seen by the user at each historical user visual angle and a theoretical bandwidth occupied by the transmission of the theoretical image based on the historical user visual angle and the panoramic video; performing video preprocessing on the panoramic video to obtain a preprocessed panoramic video; determining transmission information when the preprocessed panoramic video is transmitted, an actual image seen by each target user visual angle, and an actual bandwidth occupied by the transmission of the actual image based on the preprocessed panoramic video and the historical user visual angle; and generating an evaluation result of the preprocessed panoramic video based on a target parameter.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a panoramic video evaluation method and device and electronic equipment. BACKGROUND

[0002] In recent years, with the popularization of virtual reality (VR) head-mounted hardware and the improvement of network infrastructure, more and more users watch panoramic videos through VR headsets to experience the immersive experience brought by panoramic videos.

[0003] In order to ensure the experience of panoramic videos, the panoramic videos need to be evaluated before being deployed to users for watching, and therefore each service provider providing panoramic videos has a set of evaluation methods suitable for itself. When comparing the evaluation results, since the configuration parameters of the panoramic videos provided by different service providers are different, and different service providers have their own evaluation methods, it is difficult to directly compare the evaluation results of different configuration parameters.

[0004] Therefore, how to evaluate panoramic videos with different configuration parameters has become a problem to be solved. SUMMARY

[0005] In view of this, the present disclosure provides a panoramic video evaluation method, device and electronic equipment to solve the problem of how to evaluate panoramic videos with different configuration parameters in the related art.

[0006] In order to achieve the above purpose, the present disclosure provides the technical solutions as follows.

[0007] In a first aspect, the present disclosure provides a panoramic video evaluation method, comprising: obtaining a panoramic video and a historical user perspective corresponding to each moment when a user watches the panoramic video; determining a theoretical image seen by the user at each historical user perspective and a theoretical bandwidth occupied by the transmission of the theoretical image based on the historical user perspective and the panoramic video; performing video preprocessing on the panoramic video to obtain a preprocessed panoramic video; wherein the configuration parameters of the preprocessed panoramic video are different from the configuration parameters of the panoramic video; determining transmission information when the preprocessed panoramic video is transmitted, an actual image seen by each target user perspective, and an actual bandwidth occupied by the transmission of the actual image based on the preprocessed panoramic video and the historical user perspective; wherein the target user perspective is calculated based on the historical user perspective; generating an evaluation result of the preprocessed panoramic video based on a target parameter; wherein the target parameter includes at least one of the transmission information, the theoretical image, the actual image, the theoretical bandwidth and the actual bandwidth.

[0008] As an optional implementation of the present disclosure, based on the pre-processed panoramic video and the historical user perspectives, determining the transmission information when transmitting the pre-processed panoramic video, the actual images seen by each target user perspective, and the actual bandwidth occupied by transmitting the actual images, comprises: performing perspective prediction based on the historical user perspectives to obtain predicted user perspectives, and taking the predicted user perspectives as target user perspectives; based on the target user perspectives and the pre-processed panoramic video, determining the transmission information when transmitting the pre-processed panoramic video, the actual images seen by each target user perspective, and the actual bandwidth occupied by transmitting the actual images.

[0009] As an optional implementation of the present disclosure, based on the pre-processed panoramic video and the historical user perspectives, determining the transmission information when transmitting the pre-processed panoramic video, the actual images seen by each target user perspective, and the actual bandwidth occupied by transmitting the actual images, comprises: based on the pre-processed panoramic video, determining the video frames transmitted in the next period; based on the video frames and the historical user perspectives in the next period, determining the actual images seen by the user in each target user perspective corresponding to the video frames, and the actual bandwidth occupied by transmitting the actual images; simulating transmission of the video frames to determine the total transmission data amount and the number of stalls in the next period; based on the total transmission data amount and the number of stalls, determining the transmission information when transmitting the pre-processed panoramic video.

[0010] As an optional implementation of the present disclosure, the target parameters include theoretical images, actual images, theoretical bandwidths, and actual bandwidths, and the theoretical images and the actual images correspond one-to-one; based on the target parameters, generating the evaluation result of the pre-processed panoramic video, comprises: based on the theoretical images and the actual images, determining a first total number of pixels contained in the theoretical images and a second total number of pixels contained in the actual images; based on the first total number, the second total number, the theoretical bandwidths, and the actual bandwidths, determining the evaluation result of the panoramic video.

[0011] As an optional implementation of the present disclosure, the target parameters include theoretical images and actual images; based on the target parameters, generating the evaluation result of the pre-processed panoramic video, comprises: based on the theoretical images and the actual images, determining the peak signal-to-noise ratio and the structural similarity between the theoretical images and the actual images; based on one or more of the peak signal-to-noise ratio and the structural similarity, obtaining the evaluation result of the panoramic video.

[0012] As an optional implementation of the present disclosure, the target parameters include transmission information, and the transmission information includes a total transmission data amount and a number of stalls; based on the target parameters, generating the evaluation result of the pre-processed panoramic video, comprises: based on the total transmission data amount and the number of stalls, determining the evaluation result of the panoramic video.

[0013] As an optional implementation of the present disclosure, the panoramic video is pre-processed to obtain a pre-processed panoramic video, including: performing format conversion on the panoramic video to generate an encoded panoramic video encoded according to a preset projection format; and encoding the encoded panoramic video according to preset encoding parameters to obtain the pre-processed panoramic video.

[0014] As an optional implementation of the present disclosure, the projection format includes any one of a rectangular spherical projection, a cubemap projection, an equiangular cubemap projection, a pyramid frustum projection, and a barrel distortion projection.

[0015] As an optional implementation of the present disclosure, the encoding parameters include one or more of a rectangular block size, a quantization parameter, a preset configuration parameter of a video encoding, a group of pictures size, an intra-frame encoded frame, a forward-backward predictive encoded frame, a bi-directional predictive interpolation encoded frame, and a video encoding standard.

[0016] In a second aspect, the present disclosure provides an evaluation device for a panoramic video, including: an acquisition unit configured to acquire the panoramic video and a historical user perspective corresponding to each time when a user watches the panoramic video; a processing unit configured to determine a theoretical image seen by the user at each historical user perspective and a theoretical bandwidth occupied by the theoretical image based on the historical user perspective acquired by the acquisition unit and the panoramic video acquired by the acquisition unit; the processing unit is further configured to pre-process the panoramic video acquired by the acquisition unit to obtain a pre-processed panoramic video; wherein the configuration parameters of the pre-processed panoramic video are different from the configuration parameters of the panoramic video; the processing unit is further configured to determine transmission information when the pre-processed panoramic video is transmitted, an actual image seen by each target user perspective, and an actual bandwidth occupied by the actual image based on the pre-processed panoramic video acquired by the acquisition unit and the historical user perspective acquired by the acquisition unit; wherein the target user perspective is calculated based on the historical user perspective; the processing unit is further configured to generate an evaluation result of the pre-processed panoramic video based on target parameters; wherein the target parameters include at least one of the transmission information, the theoretical image, the actual image, the theoretical bandwidth, and the actual bandwidth.

[0017] As an optional implementation of the present disclosure, the processing unit is specifically configured to perform perspective prediction based on the historical user perspective acquired by the acquisition unit to obtain a predicted user perspective, and take the predicted user perspective as the target user perspective; the processing unit is specifically configured to determine the transmission information when the pre-processed panoramic video is transmitted, the actual image seen by each target user perspective, and the actual bandwidth occupied by the actual image based on the target user perspective and the pre-processed panoramic video acquired by the acquisition unit.

[0018] As an optional implementation of the present disclosure, the processing unit is specifically configured to determine a video frame to be transmitted in a next period based on the preprocessed panoramic video; the processing unit is specifically configured to simulate transmission of the video frame to determine a total transmission data amount and a number of stalls of transmission in the next period; the processing unit is specifically configured to determine an actual image seen by the user at a target user perspective corresponding to each video frame and an actual bandwidth occupied by the actual image based on the video frame and a historical user perspective of the next period acquired by the acquisition unit; and the processing unit is specifically configured to determine transmission information when transmitting the preprocessed panoramic video based on the total transmission data amount and the number of stalls.

[0019] As an optional implementation of the present disclosure, the target parameters include theoretical images, actual images, theoretical bandwidths and actual bandwidths, and the theoretical images and the actual images correspond to each other in a one-to-one manner; the processing unit is specifically configured to determine a first total number of pixels contained in the theoretical images and a second total number of pixels contained in the actual images based on the theoretical images and the actual images; and the processing unit is specifically configured to determine an evaluation result of the panoramic video based on the first total number, the second total number, the theoretical bandwidths and the actual bandwidths.

[0020] As an optional implementation of the present disclosure, the target parameters include theoretical images and actual images; the processing unit is specifically configured to determine a peak signal-to-noise ratio and a structural similarity between the theoretical images and the actual images based on the theoretical images and the actual images; and the processing unit is specifically configured to obtain an evaluation result of the panoramic video based on one or more of the peak signal-to-noise ratio and the structural similarity.

[0021] As an optional implementation of the present disclosure, the target parameters include transmission information, and the transmission information includes a total transmission data amount and a number of stalls; and the processing unit is specifically configured to determine an evaluation result of the panoramic video based on the total transmission data amount and the number of stalls.

[0022] As an optional implementation of the present disclosure, the processing unit is specifically configured to perform format conversion on the panoramic video to generate an encoded panoramic video encoded in a preset projection format; and the processing unit is specifically configured to encode the encoded panoramic video according to preset encoding parameters to obtain the preprocessed panoramic video.

[0023] As an optional implementation of the present disclosure, the projection format includes any one of a rectangular spherical projection, a cubemap projection, an equiangular cubemap projection, a pyramid frustum projection and a barrel distortion projection.

[0024] As an optional implementation of the present disclosure, the encoding parameters include one or more of a rectangular block size, a quantization parameter, a preset configuration parameter of video encoding, a group of pictures size, an intra-frame encoded frame, a forward and backward prediction encoded frame, a bidirectional prediction interpolation encoded frame and a video encoding standard.

[0025] In a third aspect, the present disclosure provides an electronic device, comprising: a memory and a processor, the memory is configured to store a computer program; the processor is configured to, when executing the computer program, enable the electronic device to implement the panoramic video evaluation method according to the first aspect.

[0026] In a fourth aspect, the present disclosure provides a computer readable storage medium, the computer readable storage medium stores a computer program, when the computer program is executed by a computing device, the computing device is enabled to implement the panoramic video evaluation method according to the first aspect.

[0027] In a fifth aspect, the present disclosure provides a computer program product, when the computer program product is run on a computer, the computer is enabled to implement the panoramic video evaluation method according to the first aspect.

[0028] It should be noted that the above computer instructions can be stored on the first computer readable storage medium in whole or in part. The first computer readable storage medium can be packaged together with the processor of the panoramic video evaluation device, or can be packaged separately from the processor of the panoramic video evaluation device, and the present disclosure does not limit this.

[0029] The second aspect, the third aspect, the fourth aspect and the fifth aspect of the present disclosure can refer to the detailed description of the first aspect; and the beneficial effects of the second aspect, the third aspect, the fourth aspect and the fifth aspect can refer to the beneficial effect analysis of the first aspect, which will not be repeated here.

[0030] In the present disclosure, the name of the panoramic video evaluation device does not constitute a limitation on the device or functional module itself, and in actual implementation, these devices or functional modules can appear with other names. As long as the functions of each device or functional module are similar to the present disclosure, it belongs to the scope of the claims of the present disclosure and its equivalent technology.

[0031] These aspects or other aspects of the present disclosure will be more apparent in the following description.

[0032] The technical solutions provided by the present disclosure have the following advantages compared with the prior art:

[0033] The evaluation method of the panoramic video provided by the present disclosure can determine the transmission information when the preprocessed panoramic video is transmitted, the actual image seen by each target user perspective, and the actual bandwidth occupied by the transmission of the actual image based on the historical user perspective corresponding to each moment when the user watches the panoramic video and the preprocessed panoramic video. Then, based on the theoretical image seen by each historical user perspective when the user watches the preprocessed panoramic video, and at least one of the theoretical bandwidth occupied by the transmission of the theoretical image, the actual image and the actual bandwidth, the evaluation result of the preprocessed panoramic video is generated. Since the panoramic videos with different configuration parameters are preprocessed to obtain preprocessed panoramic videos with the same configuration parameters, the evaluation results of the generated preprocessed panoramic videos can be directly compared, thereby solving the problem of how to evaluate panoramic videos with different configuration parameters. BRIEF DESCRIPTION OF DRAWINGS

[0034] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure, together with the description.

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings required to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0036] Figure 1 A schematic diagram of a right-hand Cartesian coordinate system of the evaluation method of the panoramic video provided by the present disclosure;

[0037] Figure 2 One of the flowcharts of the evaluation method of the panoramic video provided by the present disclosure;

[0038] Figure 3 A schematic diagram of a video block of the evaluation method of the panoramic video provided by the present disclosure;

[0039] Figure 4 The second flowchart of the evaluation method of the panoramic video provided by the present disclosure;

[0040] Figure 5 The third flowchart of the evaluation method of the panoramic video provided by the present disclosure;

[0041] Figure 6 The fourth flowchart of the evaluation method of the panoramic video provided by the present disclosure;

[0042] Figure 7 The fifth flowchart of the evaluation method of the panoramic video provided by the present disclosure;

[0043] Figure 8 FIG. 6 is a flowchart of a method for evaluating a panoramic video according to an embodiment of the present disclosure;

[0044] Figure 9 FIG. 7 is a flowchart of a method for evaluating a panoramic video according to an embodiment of the present disclosure;

[0045] Figure 10 FIG. 8 is a structural diagram of an evaluation device for a panoramic video according to an embodiment of the present disclosure;

[0046] Figure 11 FIG. 9 is a structural diagram of an electronic device according to an embodiment of the present disclosure;

[0047] Figure 12 FIG. 10 is a structural diagram of a computer program product of a method for evaluating a panoramic video according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0048] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0049] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other different manners from those described herein; obviously, the embodiments described in the specification are only a part of the embodiments of the present disclosure, and not all the embodiments.

[0050] It should be noted that, in this document, relational terms such as “first” and “second”, and the like, are used solely to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms “comprises”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the phrase “comprising a……” does not exclude the existence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0051] H.264 in the embodiments of the present disclosure, which is also MPEG-4 Part 10, is a highly compressed digital video codec standard proposed by the Joint Video Team (JVT) composed of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG).

[0052] H.265 in the embodiments of the present disclosure is a new video coding standard formulated by ITU-T VCEG after H.264.

[0053] In the process of coding each video frame, in the embodiments of the present disclosure, in order to improve efficiency, a video frame is divided into multiple slices or tiles, each slice or tile of a video frame is independently coded, and the direct dependence between them is not high, so as to achieve the effect of quickly coding a video frame; the difference between slice and tile lies in the different division methods, the slice is divided in a strip shape, and the tile is in a rectangular shape.

[0054] The preset parameter in the embodiments of the present disclosure is used to specify the preset configuration parameter of video coding. The preset configuration is a set of predefined coding parameters, which can be conveniently applied to different video coding tasks.

[0055] JSON (JavaScript Object Notation) in the embodiments of the present disclosure is a lightweight data exchange format.

[0056] In the embodiments of the present disclosure, a right-handed Cartesian coordinate system of a three-dimensional space is established with the center point of the head of the user wearing the head-mounted display device as the center, as shown in Figure 1 The roll refers to rotation around the X axis, the pitch refers to rotation around the Y axis, and the yaw refers to rotation around the Z axis.

[0057] An example is taken as an example of a server 1 that executes the panoramic video evaluation method provided by the embodiments of the present disclosure to illustrate the panoramic video evaluation method provided by the embodiments of the present disclosure.

[0058] Figure 2 A flowchart of a panoramic video evaluation method according to an example embodiment is shown, as shown in Figure 2 The method comprises the following S11-S15.

[0059] S11, acquire a panoramic video and a historical user perspective corresponding to each moment when a user watches the panoramic video.

[0060] In some examples, when a user needs to watch a specified panoramic video, the user can access the server 2 storing the specified panoramic video through the head-mounted display device. Then, the server 2 transmits the video data of the specified panoramic video. Thus, the head-mounted display device decodes the video data received from the server 2 to obtain decoded data. Then, the head-mounted display device renders the decoded data, so that the user can watch the specified panoramic video through the head-mounted display device. In this process, the head-mounted display device can upload the user's view angle when the user watches the specified panoramic video to the server 2, so that the server 2 stores the user's view angle when the head-mounted display device plays the specified panoramic video.

[0061] In some examples, the server 1 can obtain the panoramic video and the historical user view angle corresponding to each moment when the user watches the panoramic video in the server 2. The panoramic video can be any specified panoramic video.

[0062] In some examples, the panoramic video and the historical user view angle corresponding to each moment when the user watches the panoramic video can be collected by the system.

[0063] S12, based on the historical user view angle and the panoramic video, determining a theoretical image seen by the user at each historical user view angle and a theoretical bandwidth occupied by the theoretical image.

[0064] In some examples, when the server 2 transmits the specified panoramic video to the head-mounted display device, the server 2 can record the theoretical bandwidth occupied by the transmission of different theoretical images, and record the correspondence between the theoretical bandwidth and the theoretical image. In this way, the server 1 can send query information to the server 2 to obtain the correspondence corresponding to the panoramic video. Then, in the correspondence, the theoretical image is queried to obtain the theoretical bandwidth occupied by the transmission of the theoretical image.

[0065] In some examples, the server 1 performs image cutting in the video frame corresponding to the historical user view angle based on the historical user view angle by aligning the historical user view angle with the panoramic video.

[0066] In some examples, the server 1 inputs the historical user view angle and the panoramic video into a view angle analysis model to obtain the theoretical image seen by the user at each historical user view angle. The training process of the view angle analysis model is as follows:

[0067] Obtain training sample data and labeled results of the training sample data; wherein the training sample data includes: a specified panoramic video and an actual user view angle corresponding to each moment when a user watches the specified panoramic video, and the labeled results include: an actual image seen by the user at each moment corresponding to the actual user view angle.

[0068] The training sample data is input into the neural network model for learning, to obtain a prediction result of the neural network model on the training sample data.

[0069] Based on the prediction result and the label result, the network parameters of the neural network model are adjusted until the neural network model converges, to obtain the view angle analysis model.

[0070] In some examples, the server 1 can determine a theoretical bandwidth occupied by the transmission of the theoretical image according to an image size of the theoretical image and a transmission duration.

[0071] S13, video pre-processing is performed on the panoramic video to obtain a pre-processed panoramic video. The configuration parameters of the pre-processed panoramic video are different from the configuration parameters of the panoramic video.

[0072] In some examples, the configuration parameters include at least one of a projection format and an encoding parameter; wherein the projection format includes any one of an equirectangular (ERP), a cubemap (CMP), an equi-angular cubemap (EAC), a truncated square pyramid (TSP), and a barrel, and the encoding parameter includes one or more of a tile size, a quantization parameter (QP), a preset configuration parameter (Preset parameter) of video encoding, a group of pictures (GoP) size, an intra-picture (I-frame), a predictive-picture (P-frame), a bi-directional interpolated prediction frame (B-frame), and a video encoding standard.

[0073] In some examples, the video encoding standard includes any one of H.264 and H.265.

[0074] In some examples, the video pre-processing includes one or more of format conversion and encoding. The format conversion of the panoramic video includes: performing format conversion on the panoramic video to generate an encoded panoramic video encoded in a preset projection format. In this way, panoramic videos in different projection formats can be converted into panoramic videos in the same projection format. The encoding of the panoramic video includes performing format conversion on the panoramic video to generate an encoded panoramic video encoded in a preset projection format. Then, the encoded panoramic video is encoded according to preset encoding parameters to obtain a pre-processed panoramic video. In this way, panoramic videos in the same projection format can be converted into panoramic videos encoded with the same encoding parameters, thereby facilitating the analysis of evaluation results of panoramic videos with different configuration parameters.

[0075] In some examples, the generated pre-processed panoramic video can be a JSON file.

[0076] S14, based on the pre-processed panoramic video and the historical user perspectives, determining transmission information when transmitting the pre-processed panoramic video, actual images seen by each target user perspective, and actual bandwidth occupied by the transmission of the actual images. The target user perspective is calculated based on the historical user perspective.

[0077] In some examples, the transmission information at least includes the number of stalls and the total transmission data amount. In order to obtain the transmission information when transmitting the pre-processed panoramic video, the actual images seen by each historical user perspective, and the actual bandwidth occupied by the transmission of the actual images, the server 1 can determine a video frame to be transmitted in the next period based on the pre-processed panoramic video. Based on the video frame and the historical user perspective in the next period, the actual images seen by the user in the historical user perspective corresponding to each video frame, and the actual bandwidth occupied by the transmission of the actual images are determined. The video frame is simulated to transmit to determine the total transmission data amount and the number of stalls of the transmission in the next period. Based on the total transmission data amount and the number of stalls, the transmission information when transmitting the pre-processed panoramic video is determined.

[0078] In some examples, the target user perspective is equal to the historical user perspective, and at this time the server 1 can determine the actual images seen by each target user perspective based on the historical user perspective. In this way, compared with the server 1 issuing panoramic images to the head-mounted display device, the server 1 can only issue the actual images seen by each target user perspective to the head-mounted display device, so that the amount of data sent by the server 1 to the head-mounted display device can be greatly reduced, and the occupancy rate of network resources can be reduced.

[0079] In some examples, the target user perspective is equal to the predicted user perspective, and the server 1 can perform perspective prediction based on the historical user perspective at the current time (e.g., motion trace of the historical user perspective such as roll, pitch, and yaw), and perform perspective prediction of the historical user perspective at the current time (e.g., motion trace of the historical user perspective such as roll, pitch, and yaw) using a regression algorithm (linear regression (LR), support vector regression (SVR), ridge regression (RR), or deep learning algorithm (long short-term memory (LSTM)) to obtain the predicted user perspective, such as the user perspective (pitch, yaw, roll) at the next period. In this way, the actual image actually seen by the user can be determined based on the user perspective at the next period and the preprocessed panoramic video, and the actual image is transmitted to the head-mounted display device. In this way, the amount of data transmitted by the server 1 to the head-mounted display device can be greatly reduced, and the occupancy rate of network resources can be reduced.

[0080] In some examples, the historical user perspective at the current time can be input into the perspective prediction model to obtain the predicted user perspective. The training process of the perspective prediction model is as follows:

[0081] Obtain training sample data and labeled results of the training sample data. The training sample data includes the actual user perspective corresponding to each time when the user watches a specified panoramic video, and the labeled results include the actual user perspective at the current time and the actual user perspective at the next time.

[0082] Input the training sample data into the neural network model for training to obtain the prediction results of the neural network model on the training sample data.

[0083] Based on the prediction results and the labeled results, adjust the network parameters of the neural network model until the neural network model converges to obtain the perspective prediction model.

[0084] It should be noted that the process of determining the actual image seen by the user at each video frame corresponding to the historical user perspective based on the video frame and the historical user perspective at the next period is similar to the process of determining the theoretical image seen by the user at each historical user perspective based on the historical user perspective and the panoramic video, which will not be described here.

[0085] It should be noted that the process of determining the actual bandwidth occupied by the transmission of the actual image is similar to the process of determining the theoretical bandwidth occupied by the transmission of the theoretical image, which will not be described here.

[0086] S15, generating an evaluation result of the preprocessed panoramic video based on the target parameter. The target parameter includes at least one of the transmission information, the theoretical image, the actual image, the theoretical bandwidth, and the actual bandwidth.

[0087] In some examples, the target parameter includes the theoretical image, the actual image, the theoretical bandwidth, and the actual bandwidth. The theoretical image and the actual image are in one-to-one correspondence. The server 1 determines a first total number of pixels contained in the theoretical image (i.e., a first total number of effective pixels (pixels contained in an image actually seen by the naked eye of a user)) and a second total number of pixels contained in the actual image based on the theoretical image and the actual image. Then, based on the first total number, the second total number, the theoretical bandwidth, and the actual bandwidth, an evaluation result of the panoramic video is determined, for example, we define as the bandwidth consumed by a unit of effective pixels. Wherein, B_others represents the actual bandwidth, and B_fov represents the theoretical bandwidth. PN_others represents the second total number, and PN_fov represents the first total number. Then, Norm(B) is taken as the evaluation result of the panoramic video.

[0088] In some examples, the first total number is equal to the resolution within a field of view (FoV) region (i.e., the historical user perspective of the theoretical image corresponding to the first total number).

[0089] In some examples, represents the average number of effective pixels of the panoramic video in the playback process, p h represents the pixel density of the video frame that is not down-sampled within the region of the target user perspective when the corresponding video frame of the preprocessed panoramic video is viewed using the target user perspective, p l represents the pixel density of the video frame that is down-sampled within the region of the target user perspective when the corresponding video frame of the preprocessed panoramic video is viewed using the target user perspective, w1 represents the area of the region that is not down-sampled within the region of the target user perspective, w2 represents the area of the region that is down-sampled within the region of the target user perspective, N represents the total number of video frames used, and w1+w2=1.

[0090] In some examples, the target parameter includes the theoretical image and the actual image. The server 1 determines the peak signal-to-noise ratio (PSNR) and the structural SIMilarity (SSIM) between the theoretical image and the actual image based on the theoretical image and the actual image. Based on one or more of the peak signal-to-noise ratio and the structural SIMilarity, an evaluation result of the panoramic video is obtained.

[0091] In some examples, the target parameter includes transmission information, and the transmission information includes a total transmission data amount and a number of stutters. The server 1 determines an evaluation result of the panoramic video based on the total transmission data amount and the number of stutters. For example, based on a ratio of the total transmission data amount to a total number of data frames corresponding to the total transmission data amount, an average data amount of each video frame corresponding to the total transmission data amount is determined; based on a ratio of a data amount of a current video frame to a total number of Tiles contained in the current video frame, an average data amount of each Tile in the current video frame is determined; based on the data amount of the current video frame, a data amount of each Tile is determined; based on a sum of differences of data amounts of adjacent Tiles, a first quality difference value of the adjacent Tiles in the current video frame is determined; based on a sum of differences of data amounts of Tiles at the same position in an nth frame image corresponding to the current video block and the next video block, respectively, a second quality difference value of the Tiles at the same position in the current video block and the next video block is determined (for example, the video block contains n video frames, and the video frames in the same video block are divided into Tiles in the same way, such as Figure 3 As shown in the same video block, the video frames are divided into four Tiles, namely Tile 1, Tile 2, Tile 3 and Tile 4. When calculating the second quality difference value, the first difference value of the data amount of Tile 1 in the first frame image of the current video block and the data amount of Tile 1 in the first frame image of the next video block, the second difference value of the data amount of Tile 2 in the first frame image of the current video block and the data amount of Tile 2 in the first frame image of the next video block, the third difference value of the data amount of Tile 3 in the first frame image of the current video block and the data amount of Tile 3 in the first frame image of the next video block, and the fourth difference value of the data amount of Tile 4 in the first frame image of the current video block and the data amount of Tile 4 in the first frame image of the next video block are calculated, respectively. Based on the sum of the first difference value, the second difference value, the third difference value and the fourth difference value, the first difference value of the current video block and the next video block is obtained. Similarly, the server 1 iteratively calculates the nth difference value of the nth frame image of the current video block and the nth frame image of the next video block (N represents the total number of video frames contained in the current video block, n ∈ [1, N], and n and N are both integers greater than or equal to 1); then, based on a ratio of the sum of the difference values (the sum of the first difference value, the second difference value, …, the Nth difference value) to N, the second quality difference value is obtained; based on the number of stutters, the average data amount, the first quality difference value and the second quality difference value, the evaluation result of the panoramic video is generated.

[0092] In some examples, the number of stutters, the average data amount, the first quality difference value and the second quality difference value can be brought into an evaluation formula to obtain an evaluation score; wherein the evaluation formula includes:

[0093] max{Q-w i (I1+I2)-ws XT stall}。

[0094] wherein, T stall represents the number of stalls, Q represents the average data volume, I1 represents the first quality difference value, I2 represents the second quality difference value, w i and w s are constants.

[0095] In this way, the number of stalls, the average quality value, the first quality difference value and the second quality difference value can be brought into the evaluation formula, and the sizes of w i and w s are set to make the value of Q-w i (I1+I2)-w s XT stall maximum, the maximum value of Q-w i (I1+I2)-w s XT stall is taken as the evaluation score, and the evaluation score is taken as the evaluation result of the panoramic video.

[0096] In some examples, the total transmission data volume and the number of stalls can be input into the evaluation model to obtain the evaluation result of the panoramic video. Wherein, the training process of the evaluation model is as follows:

[0097] Obtain training sample data and labeled results of the training sample data; wherein, the training sample data includes the total transmission data volume and the number of stalls of the panoramic video in each period in history, and the labeled results include the actual evaluation result of each period.

[0098] Input the training sample data into the neural network model for learning to obtain the prediction result of the neural network model on the training sample data.

[0099] Adjust the network parameters of the neural network model based on the prediction result and the labeled result until the neural network model converges to obtain the evaluation model.

[0100] In some examples, the target parameters include transmission information, a theoretical image, an actual image, a theoretical bandwidth, and an actual bandwidth. At this time, the server 1 can determine, based on the theoretical image and the actual image, a first total number of pixels contained in the theoretical image (i.e., a first total number of effective pixels (pixels contained in an image actually seen by the naked eye of a user)) and a second total number of pixels contained in the actual image. Then, based on the first total number, the second total number, the theoretical bandwidth, and the actual bandwidth, a first result is determined. At the same time, based on the theoretical image and the actual image, a peak signal-to-noise ratio and a structural similarity between the theoretical image and the actual image are determined. Then, based on the peak signal-to-noise ratio and the structural similarity, a second result is determined. At the same time, based on the total transmission data amount and the number of stutters, a third result is determined. Then, the first result, the second result, and the third result are summarized to obtain an evaluation result of the panoramic video.

[0101] It should be noted that the process of determining the first result based on the first total number, the second total number, the theoretical bandwidth, and the actual bandwidth is similar to the process of determining the evaluation result of the panoramic video based on the first total number, the second total number, the theoretical bandwidth, and the actual bandwidth, and will not be described here.

[0102] At the same time, the process of determining the second result based on the peak signal-to-noise ratio and the structural similarity is similar to the process of obtaining the evaluation result of the panoramic video based on one or more of the peak signal-to-noise ratio and the structural similarity, and will not be described here.

[0103] At the same time, the process of determining the third result based on the total transmission data amount and the number of stutters is similar to the process of determining the evaluation result of the panoramic video based on the total transmission data amount and the number of stutters, and will not be described here.

[0104] As can be seen from the above, the evaluation method of the panoramic video provided by the embodiments of the present disclosure can directly compare the evaluation results of the preprocessed panoramic videos with different configuration parameters, which is convenient for users to evaluate the panoramic videos with different transmission strategies (such as projection formats and encoding parameters) to find the best transmission strategy and ensure the user experience.

[0105] As an optional implementation of the present disclosure, in combination with Figure 2 As shown in Figure 4 S14 can be specifically implemented by S140 and S141.

[0106] S140, perspective prediction is performed based on historical user perspectives to obtain a predicted user perspective, and the predicted user perspective is taken as a target user perspective.

[0107] S141, determining, based on the target user perspective and the preprocessed panoramic video, transmission information when transmitting the preprocessed panoramic video, actual images seen by each target user perspective, and actual bandwidths occupied by the actual images.

[0108] As an optional implementation of the present disclosure, the target parameters include a theoretical image, an actual image, a theoretical bandwidth, and an actual bandwidth, the theoretical image and the actual image correspond to each other in a one-to-one manner; in combination with Figure 2 As shown in Figure 5 S14 can be specifically implemented by S142-S145.

[0109] S142, determining, based on the preprocessed panoramic video, video frames to be transmitted in a next period.

[0110] S143, determining, based on the video frames and historical user perspectives in the next period, actual images seen by users at target user perspectives corresponding to each video frame, and actual bandwidths occupied by the actual images.

[0111] S144, simulating transmission of the video frames to determine total transmission data amount and number of stalls in the transmission in the next period.

[0112] In some examples, to ensure the authenticity of the simulated transmission, a code rate adaptive algorithm can be used to simulate transmission of the video frames to determine the total transmission data amount and the number of stalls in the transmission in the next period.

[0113] S145, determining, based on the total transmission data amount and the number of stalls, the transmission information when transmitting the preprocessed panoramic video.

[0114] As an optional implementation of the present disclosure, the target parameters include a theoretical image, an actual image, a theoretical bandwidth, and an actual bandwidth, the theoretical image and the actual image correspond to each other in a one-to-one manner; in combination with Figure 2 As shown in Figure 6 S15 can be specifically implemented by S150 and S151.

[0115] S150, determining, based on the theoretical image and the actual image, a first total number of pixels contained in the theoretical image, and a second total number of pixels contained in the actual image.

[0116] S151, determining, based on the first total number, the second total number, the theoretical bandwidth, and the actual bandwidth, an evaluation result of the panoramic video.

[0117] As an optional implementation of the present disclosure, the target parameters include a theoretical image and an actual image; in combination with Figure 2 As shown in Figure 7 S15 can be specifically implemented by S152 and S153.

[0118] S152, determine the peak signal-to-noise ratio and structural similarity between the theoretical image and the actual image based on the theoretical image and the actual image.

[0119] S153, obtain the evaluation result of the panoramic video based on one or more of the peak signal-to-noise ratio and the structural similarity.

[0120] As an optional embodiment of the present disclosure, the target parameter includes transmission information, and the transmission information includes a total transmission data amount and a number of freezing times. Figure 2 As shown in Figure 8 S15 can be implemented by the following S154.

[0121] S154, determine the evaluation result of the panoramic video based on the total transmission data amount and the number of freezing times.

[0122] As an optional embodiment of the present disclosure, the target parameter includes transmission information, and the transmission information includes a total transmission data amount and a number of freezing times. Figure 2 As shown in Figure 9 S13 can be implemented by the following S130 and S131.

[0123] S130, format conversion is performed on the panoramic video to generate an encoded panoramic video encoded according to a preset projection format.

[0124] S131, encode the encoded panoramic video according to a preset encoding parameter to obtain a preprocessed panoramic video.

[0125] As an optional embodiment of the present disclosure, the projection format includes any one of a rectangular spherical projection, a cube map projection, an equiangular cube map projection, a pyramid frustum projection, and a barrel distortion projection.

[0126] As an optional embodiment of the present disclosure, the encoding parameter includes one or more of a rectangular block size, a quantization parameter, a preset configuration parameter of video encoding, a group of pictures size, an intra-frame encoding frame, a forward and backward prediction encoding frame, a bidirectional prediction interpolation encoding frame, and a video encoding standard.

[0127] The above mainly introduces the scheme provided by the embodiments of the present application from the perspective of method. In order to realize the above functions, it contains the hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of the examples described in the embodiments disclosed herein, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0128] The embodiments of the present application can divide the functional modules of the panoramic video evaluation device according to the above method examples. For example, each functional module can be divided according to each function, or two or more functions can be integrated in one processing module. The integrated module can be realized in the form of hardware or in the form of a software functional module. It should be noted that the division of the modules in the embodiments of the present application is illustrative, and is only a logical functional division. In actual implementation, there can be another division method.

[0129] As shown in Figure 10 The embodiments of the present application provide a structural schematic diagram of a panoramic video evaluation device 10. The panoramic video evaluation device 10 includes an acquisition unit 101 and a processing unit 102.

[0130] The acquisition unit 101 is configured to acquire a panoramic video and a historical user perspective corresponding to each moment when a user watches the panoramic video. The processing unit 102 is configured to determine a theoretical image seen by the user at each historical user perspective and a theoretical bandwidth occupied by the transmission of the theoretical image based on the historical user perspective acquired by the acquisition unit 101 and the panoramic video acquired by the acquisition unit 101. The processing unit 102 is further configured to perform video preprocessing on the panoramic video acquired by the acquisition unit 101 to obtain a preprocessed panoramic video. The configuration parameters of the preprocessed panoramic video are different from the configuration parameters of the panoramic video. The processing unit 102 is further configured to determine transmission information when the preprocessed panoramic video is transmitted, an actual image seen by each target user perspective, and an actual bandwidth occupied by the transmission of the actual image based on the preprocessed panoramic video acquired by the acquisition unit 101 and the historical user perspective acquired by the acquisition unit 101. The target user perspective is calculated based on the historical user perspective. The processing unit 102 is further configured to generate an evaluation result of the preprocessed panoramic video based on a target parameter. The target parameter includes at least one of the transmission information, the theoretical image, the actual image, the theoretical bandwidth, and the actual bandwidth.

[0131] As an optional embodiment of the present disclosure, the processing unit 102 is specifically configured to perform perspective prediction based on the historical user perspective acquired by the acquisition unit 101 to obtain a predicted user perspective, and take the predicted user perspective as the target user perspective. The processing unit 102 is specifically configured to determine the transmission information when the preprocessed panoramic video is transmitted, the actual image seen by each target user perspective, and the actual bandwidth occupied by the transmission of the actual image based on the target user perspective and the preprocessed panoramic video acquired by the acquisition unit 101.

[0132] As an optional implementation of the present disclosure, the processing unit 102 is specifically configured to determine the video frame transmitted in the next period based on the preprocessed panoramic video; the processing unit 102 is specifically configured to determine the total transmission data amount and the number of stalls of the transmission in the next period by simulating the transmission of the video frame; the processing unit 102 is specifically configured to determine the actual image seen by the user at the target user perspective corresponding to each video frame and the actual bandwidth occupied by the transmission of the actual image based on the video frame and the historical user perspective of the next period acquired by the acquisition unit 101; and the processing unit 102 is specifically configured to determine the transmission information when the preprocessed panoramic video is transmitted based on the total transmission data amount and the number of stalls.

[0133] As an optional implementation of the present disclosure, the target parameter includes the theoretical image, the actual image, the theoretical bandwidth and the actual bandwidth, and the theoretical image and the actual image correspond to each other; the processing unit 102 is specifically configured to determine the first total number of pixels contained in the theoretical image and the second total number of pixels contained in the actual image based on the theoretical image and the actual image; and the processing unit 102 is specifically configured to determine the evaluation result of the panoramic video based on the first total number, the second total number, the theoretical bandwidth and the actual bandwidth.

[0134] As an optional implementation of the present disclosure, the target parameter includes the theoretical image and the actual image; the processing unit 102 is specifically configured to determine the peak signal-to-noise ratio and the structural similarity between the theoretical image and the actual image based on the theoretical image and the actual image; and the processing unit 102 is specifically configured to obtain the evaluation result of the panoramic video based on one or more of the peak signal-to-noise ratio and the structural similarity.

[0135] As an optional implementation of the present disclosure, the target parameter includes the transmission information, and the transmission information includes the total transmission data amount and the number of stalls; and the processing unit 102 is specifically configured to determine the evaluation result of the panoramic video based on the total transmission data amount and the number of stalls.

[0136] As an optional implementation of the present disclosure, the processing unit 102 is specifically configured to perform format conversion on the panoramic video to generate an encoded panoramic video encoded according to a preset projection format; and the processing unit 102 is specifically configured to encode the encoded panoramic video according to a preset encoding parameter to obtain the preprocessed panoramic video.

[0137] As an optional implementation of the present disclosure, the projection format includes any one of a rectangular spherical projection, a cube map projection, an equiangular cube map projection, a pyramid frustum projection and a barrel distortion projection.

[0138] As an optional embodiment of the present disclosure, the encoding parameters include one or more of a rectangular block size, a quantization parameter, a preset configuration parameter of video encoding, a group of picture size, an intra-frame encoding frame, a forward-backward predictive encoding frame, a bidirectional predictive interpolation encoding frame, and a video encoding standard.

[0139] Wherein, all the related content of each step involved in the method embodiments can be cited to the function description of the corresponding function module, and the function is not described here.

[0140] Of course, the panoramic video evaluation device 10 provided by the embodiment of the present application includes but is not limited to the above-mentioned modules, for example, the panoramic video evaluation device 10 can also include a storage unit 103. The storage unit 103 can be used to store the program code of the panoramic video evaluation device 10, and can also be used to store the data generated by the panoramic video evaluation device 10 during operation, such as data in a write request.

[0141] Figure 11 A structural schematic diagram of an electronic device provided by the embodiment of the present application is shown in Figure 11 The electronic device can include at least one processor 51, a memory 52, a communication interface 53, and a communication bus 54.

[0142] The specific introduction of each component of the electronic device is as follows: Figure 11 The specific introduction of each component of the electronic device is as follows:

[0143] The processor 51 is the control center of the electronic device, which can be one processor or a plurality of processing elements. For example, the processor 51 is a central processing unit (CPU), which can also be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiment of the present application, such as one or more DSPs, or one or more field programmable gate arrays (FPGA).

[0144] In a specific implementation, as an embodiment, the processor 51 can include one or more CPUs, such as the CPU0 and CPU1 shown in Figure 11 As an embodiment, the electronic device can include a plurality of processors, such as the processor 51 and the processor 52 shown in Figure 11The processors 51 and 56 shown in FIG. 1 can each be a single-core processor (Single-CPU) or a multi-core processor (Multi-CPU). The processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0145] The memory 52 can be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disk storage, a magnetic disk storage or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and that can be accessed by a computer, but is not limited thereto. The memory 52 can exist independently, and is connected to the processor 51 through the communication bus 54. The memory 52 can also be integrated with the processor 51.

[0146] In a specific implementation, the memory 52 is configured to store data and software programs for implementing the present application. The processor 51 can execute various functions of the air conditioner by running or executing the software programs stored in the memory 52 and calling the data stored in the memory 52.

[0147] The communication interface 53 is configured to communicate with other devices or communication networks, such as a radio access network (RAN), a wireless local area network (WLAN), a terminal, a cloud, etc., using any transceiver-like device. The communication interface 53 can include an obtaining unit to implement an obtaining function.

[0148] The communication bus 54 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 11 Only one thick line is used to represent the bus in the figure, but it does not mean that there is only one bus or only one type of bus.

[0149] As an example, in conjunction with Figure 10 the function of the acquisition unit 101 in the panoramic video evaluation device 10 is the same as the function of the communication interface 53 in Figure 11 the function of the processing unit 102 in the panoramic video evaluation device 10 is the same as the function of the processor 51 in Figure 11 the function of the storage unit 103 in the panoramic video evaluation device 10 is the same as the function of the memory 52 in Figure 11 .

[0150] Another embodiment of the present application also provides a computer readable storage medium, and the computer program is stored on the computer readable storage medium, and when the computer program is executed by a computing device, the computing device implements the method shown in the above method embodiment.

[0151] In some embodiments, the disclosed method can be implemented as computer program instructions encoded in a machine-readable storage medium in a machine-readable format or encoded in other non-transitory media or articles.

[0152] Figure 12 The conceptual partial view of the computer program product provided by the embodiment of the present application is schematically shown, and the computer program product includes a computer program for executing a computer process on a computing device.

[0153] In one embodiment, the computer program product is provided using a signal bearing medium 410. The signal bearing medium 410 can include one or more program instructions, which when executed by one or more processors can provide the above-described functions or partial functions. Therefore, for example, with reference to the embodiment shown in Figure 2 , one or more features of S11-S15 can be assumed by one or more instructions associated with the signal bearing medium 410. In addition, Figure 2 the program instructions in Figure 12 also describe example instructions.

[0154] In some examples, the signal bearing medium 410 can comprise a computer- readable medium 411, such as, but not limited to, a hard disk drive, a compact disk (CD), a digital video disk (DVD), a memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electronically erasable programmable read-only memory (EEPROM), a floppy disk, a zip disk, a magnetic tape, other forms of magnetic, optical, or electrical storage media, or any other medium which can be used to carry or store desired program code elements in the form of computer-executable instructions or data structures and which can be accessed by a computer.

[0155] In some embodiments, the signal bearing medium 410 can comprise a computer- recordable medium 412, such as, but not limited to, a floppy disk, a zip disk, a magnetic tape, other forms of magnetic, optical, or electronic storage media, or any other medium which can be used to carry or store desired program code elements in the form of computer-executable instructions or data structures and which can be accessed by a computer.

[0156] In some embodiments, the signal bearing medium 410 can comprise a communication medium 413, such as, but not limited to, a digital and / or an analog communication medium (e.g., a fiber optic cable, a waveguide, a wired electrical cable, a wireless electrical cable, a wireless electromagnetic cable, etc.).

[0157] The signal bearing medium 410 can be communicated by a wireless form of the communication medium 413 (e.g., a wireless communication medium complying with the IEEE 802.41 standard or other transmission protocols). The one or more program instructions can be, for example, computer-executable instructions or logic-implemented instructions.

[0158] In some examples, such as for Figure 10 The evaluation apparatus 10 of the panoramic video described can be configured to provide various operations, functions, or actions in response to the one or more program instructions through the computer-readable medium 411, the computer-recordable medium 412, and / or the communication medium 413.

[0159] From the above description of embodiments, it is apparent that a person skilled in the art can clearly understand, for the convenience and brevity of description, only the above-mentioned division of functional modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above.

[0160] In several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented by other means. For example, the apparatus embodiments described above are only illustrative, for example, the division of the modules or units is only a logical function division, and in actual implementation, another division mode can be used, for example, a plurality of units or components can be combined or integrated into another apparatus, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed ones can be indirect coupling or communication connection through some interfaces, apparatuses or units, which can be electrical, mechanical or other forms.

[0161] The units described as separate components may or may not be physically separate, and the components displayed as units may be one physical unit or multiple physical units, that is, may be located in one place, or also may be distributed to multiple different places. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0162] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0163] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical scheme of the embodiments of the present application essentially or the part that contributes to the prior art or the whole or part of the technical scheme can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions to make a device (which can be a single-chip microcomputer, a chip, etc.) or a processor execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk and various program code storage media.

[0164] The above is only a specific embodiment of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of evaluating a panoramic video, characterized by, The method comprises: acquiring a panoramic video and a historical user perspective corresponding to each moment when a user watches the panoramic video; based on the historical user perspective and the panoramic video, determining a theoretical image seen by the user at each historical user perspective and a theoretical bandwidth occupied by the theoretical image; performing video preprocessing on the panoramic video to obtain a preprocessed panoramic video; wherein the configuration parameters of the preprocessed panoramic video are different from the configuration parameters of the panoramic video; based on the preprocessed panoramic video and the historical user perspective, determining transmission information when transmitting the preprocessed panoramic video, an actual image seen by each target user perspective, and an actual bandwidth occupied by the actual image; wherein the target user perspective is calculated based on the historical user perspective; based on target parameters, generating an evaluation result of the preprocessed panoramic video; wherein the target parameters include at least one of the transmission information, the theoretical image, the actual image, the theoretical bandwidth, and the actual bandwidth.

2. The method of claim 1, wherein, The method comprises: based on the historical user perspective, performing perspective prediction to obtain a predicted user perspective, and taking the predicted user perspective as a target user perspective; based on the target user perspective and the preprocessed panoramic video, determining transmission information when transmitting the preprocessed panoramic video, an actual image seen by each target user perspective, and an actual bandwidth occupied by the actual image.

3. The method of claim 1, wherein, The method comprises: based on the preprocessed panoramic video, determining a video frame to be transmitted in the next period; based on the video frame and the historical user perspective in the next period, determining an actual image seen by the user at each target user perspective corresponding to the video frame, and an actual bandwidth occupied by the actual image; performing simulation transmission on the video frame to determine the total transmission data amount and the number of stalls of transmission in the next period; based on the total transmission data amount and the number of stalls, determining transmission information when transmitting the preprocessed panoramic video.

4. The method of claim 1, wherein, The target parameters include the theoretical image, the actual image, the theoretical bandwidth, and the actual bandwidth, and the theoretical image and the actual image correspond one-to-one; The method comprises: based on the theoretical image and the actual image, determining a first total number of pixels contained in the theoretical image and a second total number of pixels contained in the actual image; based on the first total number, the second total number, the theoretical bandwidth, and the actual bandwidth, determining an evaluation result of the panoramic video.

5. The method of claim 1, wherein, The target parameters include the theoretical image and the actual image. The evaluation result of the preprocessed panoramic video is generated based on the target parameter, and the evaluation result of the preprocessed panoramic video is generated based on the target parameter. Based on the theoretical image and the actual image, the peak signal-to-noise ratio and the structural similarity between the theoretical image and the actual image are determined. Based on one or more of the peak signal-to-noise ratio and the structural similarity, the evaluation result of the panoramic video is obtained.

6. The method of claim 1, wherein, The target parameter includes transmission information, and the transmission information includes total transmission data volume and stuttering frequency. The evaluation result of the preprocessed panoramic video is generated based on the target parameter, and the evaluation result of the preprocessed panoramic video is generated based on the target parameter. Based on the total transmission data volume and the stuttering frequency, the evaluation result of the panoramic video is determined.

7. The method of claim 1, wherein, The video preprocessing of the panoramic video includes: The format conversion of the panoramic video is performed to generate an encoded panoramic video encoded in a preset projection format; The encoded panoramic video is encoded according to a preset encoding parameter to obtain a preprocessed panoramic video.

8. The method of evaluating a panoramic video according to claim 7, wherein, The projection format includes any one of rectangular spherical projection, cube map projection, equiangular cube map projection, pyramid frustum projection, and barrel distortion projection.

9. The method of claim 7, wherein, The encoding parameter includes one or more of rectangular block size, quantization parameter, preset configuration parameter of video encoding, group of pictures size, intra-frame encoding frame, forward and backward prediction encoding frame, bi-directional prediction interpolation encoding frame, and video encoding standard.

10. An apparatus for evaluating a panoramic video, characterized by comprising: It includes: An acquisition unit configured to acquire a panoramic video and a historical user perspective corresponding to each moment when a user views the panoramic video; A processing unit configured to determine a theoretical image seen by the user at each historical user perspective based on the historical user perspective acquired by the acquisition unit and the panoramic video acquired by the acquisition unit, and determine a theoretical bandwidth occupied by the transmission of the theoretical image; The processing unit is further configured to perform video preprocessing on the panoramic video acquired by the acquisition unit to obtain a preprocessed panoramic video; wherein the configuration parameter of the preprocessed panoramic video is different from the configuration parameter of the panoramic video; The processing unit is further configured to determine transmission information when the preprocessed panoramic video is transmitted, an actual image seen by each target user perspective, and an actual bandwidth occupied by the transmission of the actual image based on the preprocessed panoramic video acquired by the acquisition unit and the historical user perspective acquired by the acquisition unit; wherein the target user perspective is calculated based on the historical user perspective. The processing unit is further configured to generate an evaluation result of the preprocessed panoramic video based on a target parameter; wherein the target parameter includes at least one of the transmission information, the theoretical image, the actual image, the theoretical bandwidth, and the actual bandwidth.

11. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and when the computer program is executed by a computing device, the computing device implements the panoramic video evaluation method of any one of claims 1-9.

12. A computer program product, characterised in that, When the computer program product runs on a computer, the computer implements the panoramic video evaluation method as claimed in any one of claims 1-9.

Citation Information

Patent Citations

  • Panoramic video stream quality evaluation method and device

    CN111093069A

  • Video processing method and device, readable medium and electronic equipment

    CN113259601A