Panoramic video evaluation method, computer device and computer program product

By obtaining the deformed image features of panoramic video frames and calculating the video frame features and inter-frame correlation, the panoramic video quality is directly evaluated, and the problems of large and inefficient computing resources in the prior art are solved, and efficient and accurate panoramic video quality evaluation is achieved.

CN114972267BActive Publication Date: 2025-05-23TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210604868.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-05-23
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

The prior art has problems such as high consumption of computing resources and low efficiency in panoramic video quality evaluation, especially when training neural networks to generate reference videos.

Method used

By acquiring the deformed panoramic image and its image features of the panoramic video frame, calculating the video frame features based on these features, and calculating the video evaluation results of the panoramic video in combination with inter-frame correlation, the process of generating a reference video is avoided.

Benefits of technology

The calculation resources required for video quality evaluation are significantly reduced, the evaluation efficiency is improved, and the problem of insufficient panoramic video samples is compensated by deforming treatment, and the accuracy of evaluation results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972267B_ABST
    Figure CN114972267B_ABST
Patent Text Reader

Abstract

The present application relates to the field of video processing technology, and provides a panoramic video evaluation method, a computer device, and a computer program product, which can effectively improve the quality evaluation efficiency of panoramic videos. The method includes: obtaining multiple panoramic video frames of a panoramic video to be evaluated; obtaining a de-deformed panoramic image of each panoramic video frame and obtaining image features of the de-deformed panoramic image; obtaining video frame features of each panoramic video frame based on the image features of the de-deformed panoramic image of each panoramic video frame, and determining the inter-frame correlation of multiple panoramic video frames based on the video frame sequence of the multiple panoramic video frames and the video frame features of each panoramic video frame; obtaining a video evaluation result of the panoramic video according to the video frame features of each panoramic video frame and the inter-frame correlation of the multiple panoramic video frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video processing technology, and in particular to a panoramic video evaluation method, a computer device, and a computer program product. Background Art

[0002] With the development of computer technology, panoramic videos are increasingly widely used. To obtain better video display effects, panoramic videos can be evaluated for quality and related processing parameters can be adjusted based on the evaluation results. In traditional technologies, a neural network can be trained, and a reference video of the panoramic video can be generated through the trained neural network, and then the video quality evaluation result of the panoramic video can be obtained based on the reference video.

[0003] However, the above method requires a large amount of computing resources to train multiple parameters of the neural network, and there is a problem of low efficiency in panoramic video quality evaluation. Summary of the invention

[0004] Based on this, it is necessary to provide a panoramic video evaluation method, computer equipment and computer program product to address the above technical issues.

[0005] In a first aspect, the present application provides a panoramic video evaluation method. The method comprises:

[0006] Acquire multiple panoramic video frames of the panoramic video to be evaluated;

[0007] Acquire a de-deformed panoramic image of each panoramic video frame and acquire image features of the de-deformed panoramic image;

[0008] Based on the image features of the de-distorted panoramic image of each panoramic video frame, a video frame feature of each panoramic video frame is obtained, and based on the video frame sequence of the plurality of panoramic video frames and the video frame features of each of the panoramic video frames, an inter-frame correlation of the plurality of panoramic video frames is determined;

[0009] A video evaluation result of the panoramic video is obtained according to the video frame features of each of the panoramic video frames and the inter-frame correlation of the multiple panoramic video frames.

[0010] In a second aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0011] Acquire multiple panoramic video frames of the panoramic video to be evaluated;

[0012] Acquire a de-deformed panoramic image of each panoramic video frame and acquire image features of the de-deformed panoramic image;

[0013] Based on the image features of the de-distorted panoramic image of each panoramic video frame, a video frame feature of each panoramic video frame is obtained, and based on the video frame sequence of the plurality of panoramic video frames and the video frame features of each of the panoramic video frames, an inter-frame correlation of the plurality of panoramic video frames is determined;

[0014] A video evaluation result of the panoramic video is obtained according to the video frame features of each of the panoramic video frames and the inter-frame correlation of the multiple panoramic video frames.

[0015] In a third aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0016] Acquire multiple panoramic video frames of the panoramic video to be evaluated;

[0017] Acquire a de-deformed panoramic image of each panoramic video frame and acquire image features of the de-deformed panoramic image;

[0018] Based on the image features of the de-distorted panoramic image of each panoramic video frame, a video frame feature of each panoramic video frame is obtained, and based on the video frame sequence of the plurality of panoramic video frames and the video frame features of each of the panoramic video frames, an inter-frame correlation of the plurality of panoramic video frames is determined;

[0019] A video evaluation result of the panoramic video is obtained according to the video frame features of each of the panoramic video frames and the inter-frame correlation of the multiple panoramic video frames.

[0020] The above-mentioned panoramic video evaluation method, computer device and computer program product can obtain multiple panoramic video frames of the panoramic video to be evaluated, obtain the dedistorted panoramic image of each panoramic video frame through the pre-acquired dedistortion model, obtain the image features of the dedistorted panoramic image, and obtain the video frame features of each panoramic video frame based on the image features of the dedistorted panoramic image of each panoramic video frame, and then obtain the video evaluation result of the panoramic video according to the video frame features of each panoramic video frame and the frame correlation between the multiple panoramic video frames. The scheme of this embodiment can perform evaluation based on the video frame features of the panoramic video itself, without the need to train the network to generate a reference video, effectively alleviating the computing resources consumed in the reference video generation process, and evaluating the video quality based on the image features of the dedistorted panoramic image can effectively make up for the inaccurate evaluation results caused by insufficient panoramic video samples, thereby improving the accuracy of the evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a schematic flow chart of a panoramic video evaluation method in one embodiment;

[0022] Figure 2A schematic diagram of a flow chart of a step of obtaining a de-distorted panoramic image in one embodiment;

[0023] Figure 3 A schematic diagram of a process of obtaining a sampled panoramic image in an embodiment;

[0024] Figure 4 A schematic diagram of a process of obtaining video frame features in an embodiment;

[0025] Figure 5 A schematic diagram of a flow chart of a training generation model in an embodiment;

[0026] Figure 6 is a flow chart of another panoramic video evaluation method in one embodiment;

[0027] Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0029] In one embodiment, Figure 1 As shown, a panoramic video evaluation method is provided. This embodiment uses the method applied to a server as an example for illustration. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0030] S110: Acquire multiple panoramic video frames of a panoramic video to be evaluated.

[0031] Among them, panoramic video can be generated by using professional shooting equipment (such as VR panoramic camera) to capture the image information of the entire scene, or it can also be synthesized by video processing software. Compared with videos shot with traditional shooting equipment (such as cameras equipped with standard lenses), the field of view of panoramic video can cover the side, top, bottom and back that traditional shooting equipment cannot capture, but at the same time, due to the wide viewing angle, the video frames in the panoramic video will have imaging distortion.

[0032] In a specific implementation, multiple panoramic video frames in the panoramic video to be evaluated may be obtained.

[0033] Specifically, the panoramic video to be evaluated can be a video obtained in response to a specified trigger operation, such as responding to a video evaluation request sent by a terminal to obtain a panoramic video indicated by the video evaluation request, wherein the form of the trigger operation can be set according to actual conditions, such as a user clicking a preset video evaluation button or a video play button to trigger the evaluation of the panoramic video. Of course, the panoramic video to be evaluated can also be a panoramic video that is actively obtained, such as evaluating and reviewing multiple stored panoramic videos.

[0034] After the panoramic video to be evaluated is obtained, the panoramic video may be framed and multiple panoramic video frames in the panoramic video may be obtained. The multiple panoramic video frames may be all video frames of the panoramic video or partial video frames of the panoramic video.

[0035] S120: Obtain a de-deformed panoramic image and image features of the de-deformed panoramic image of each panoramic video frame.

[0036] As an example, the undeformed panoramic image may be an image in which deformation of a photographed object in a panoramic video frame is removed, and the undeformed panoramic image may correctly reflect the appearance of the photographed object. The image feature may be information reflecting the characteristics of the image content in the undeformed panoramic image. For example, the image feature may characterize the characteristics of at least one of the following information in the undeformed panoramic image: contour, edge, color, texture, shape, and object state in the undeformed panoramic image.

[0037] In practical applications, after acquiring multiple panoramic video frames, a dedistorted panoramic image of each panoramic video frame can be acquired to remove the image distortion in the panoramic video frame, and then feature extraction can be performed on the dedistorted panoramic image to obtain image features of the dedistorted panoramic image.

[0038] Optionally, when obtaining a dedistorted panoramic image of a panoramic video frame, existing panoramic image processing software (such as Pano2VR software) can be used to remove image deformation in the panoramic video frame to obtain a dedistorted panoramic image, or a trained model can be called to process the panoramic image to obtain a dedistorted panoramic image.

[0039] S130, obtaining video frame features of each panoramic video frame based on image features of the de-distorted panoramic image of each panoramic video frame, and determining inter-frame correlation of the plurality of panoramic video frames based on a video frame sequence of the plurality of panoramic video frames and the video frame features of each panoramic video frame.

[0040] The inter-frame correlation may refer to the temporal correlation between adjacent panoramic video frames or interval panoramic video frames. Compared with a still image, there is a motion relationship between the object contained in each panoramic video frame in the panoramic video and its previous and next frames.

[0041] In this step, after obtaining the image features of each dedistorted panoramic image, for each panoramic video frame, the video frame features of the panoramic video frame can be obtained based on the image features of the dedistorted panoramic image of the panoramic video frame. For example, for a panoramic video frame, the image features of the dedistorted panoramic image can be fused, and the fused image features can be used as the video frame features of the panoramic video frame. For another example, the image features of each dedistorted panoramic image can be screened, and the screened image features can be used as the video frame features.

[0042] Furthermore, after obtaining the video frame features of each panoramic video frame, the inter-frame correlation of the multiple panoramic video frames can be determined based on the video frame sequence of the multiple panoramic video frames and the video frame features of each panoramic video frame. For example, after obtaining the video frame features of each panoramic video frame, the video frame features of adjacent panoramic video frames can be compared according to the video frame sequence of the multiple panoramic video frames to determine the difference or change between the video frame features and obtain a comparison result, and then the changes that occur to the multiple panoramic video frames over time can be determined according to the multiple comparison results to obtain the inter-frame correlation of the multiple panoramic video frames.

[0043] S140, obtaining a video evaluation result of the panoramic video according to the video frame features of each panoramic video frame and the inter-frame correlation of the plurality of panoramic video frames.

[0044] In this step, after obtaining the video frame features corresponding to each panoramic video frame and the inter-frame correlation of multiple panoramic video frames, the video evaluation result of the panoramic video can be determined based on the multiple video frame features and the inter-frame correlation of multiple panoramic video frames. For example, based on the video frame features and inter-frame correlation of multiple panoramic video frames, evaluation information of one or more evaluation factors such as the quality of the panoramic video, the degree of jitter during video shooting, and whether the video is damaged during transmission can be determined, and then the video evaluation result of the panoramic video can be determined by combining multiple evaluation information.

[0045] By using the image features of the de-distorted panoramic image to determine the video frame features of each panoramic video frame, and based on this to evaluate the video quality of the panoramic video, the computing resources required in the video quality evaluation process can be significantly reduced, and reliable video evaluation results can be obtained. Specifically, on the one hand, the video evaluation results are determined by the video frame features of the panoramic video itself, without the need for reference video for evaluation, and a series of processing required to obtain the reference video (such as training a neural network model for generating a reference video and generating a reference video with a large amount of data through the trained neural network model) can be avoided; on the other hand, there is still a certain gap between the number of existing panoramic images and the number of ordinary images (images without distortion or images with a slight degree of distortion). When shooting videos or images, users often mainly shoot ordinary images, and the evaluation samples and evaluation results related to ordinary images or ordinary videos are also higher than those of panoramic images or panoramic videos. In this embodiment, by converting the panoramic image into a de-distorted panoramic image and obtaining the video frame features of the panoramic video frame based on the de-distorted panoramic image, the evaluation method of ordinary videos or ordinary images can be effectively transferred to the evaluation of panoramic videos, so that reliable and accurate evaluation results can still be obtained when there are fewer panoramic video samples.

[0046] In this embodiment, multiple panoramic video frames of the panoramic video to be evaluated can be obtained. After obtaining the de-deformed panoramic image of each panoramic video frame and the image features of the de-deformed panoramic image, the video frame features of each panoramic video frame can be obtained based on the image features of the de-deformed panoramic image of each panoramic video frame, and the inter-frame correlation of the multiple panoramic video frames can be determined based on the video frame sequence of the multiple panoramic video frames and the video frame features of each of the panoramic video frames. Then, the video evaluation result of the panoramic video can be obtained based on the video frame features of each panoramic video frame and the inter-frame correlation of the multiple panoramic video frames. The scheme of this embodiment can be evaluated based on the video frame features of the panoramic video itself, without the need to train the network to generate a reference video, effectively avoiding the use of a large amount of computing resources to obtain a reference video, and evaluating the video quality based on the image features of the de-deformed panoramic image can effectively make up for the inaccurate evaluation results caused by insufficient panoramic video samples, thereby improving the accuracy of the evaluation results.

[0047] In one embodiment, Figure 2 As shown, in S120, obtaining a de-distorted panoramic image of each panoramic video frame may include the following steps:

[0048] S210 , for each panoramic video frame, performing image sampling on the panoramic video frame to obtain a plurality of sampled panoramic images of the panoramic video frame.

[0049] In a specific implementation, panoramic videos have the characteristics of high resolution and large file size. That is, under the same conditions, the resolution and file size of panoramic videos shot with professional shooting equipment are larger than those shot with traditional shooting equipment. Correspondingly, the resolution and image size of panoramic video frames extracted from the panoramic video are also larger.

[0050] In this embodiment, after acquiring multiple panoramic video frames of a panoramic video, image sampling can be performed at multiple positions in the panoramic video frame to obtain multiple sampled panoramic images of the panoramic video frame. For example, multiple sampled panoramic images in the same sampled panoramic image can be different, that is, there is no identical image content in any two sampled panoramic images, so that the multiple sampled panoramic images can cover more image information in the panoramic image. Of course, in another example, the multiple sampled panoramic images can also have partially identical image content. For example, Figure 3 As shown, after obtaining the panoramic video frame, image sampling can be performed to obtain the following Figure 3 Five sample panoramic images are shown on the right.

[0051] S220 , performing a de-distortion process on each sampled panoramic image of each panoramic video frame to obtain each de-distorted panoramic image of each panoramic video frame.

[0052] Since image sampling is performed on the panoramic video frames, the original image distortion in the panoramic video frames will also be collected into the sampled panoramic images during sampling. After obtaining multiple sampled panoramic images of each panoramic video frame, the sampled panoramic images of each panoramic video frame can be converted into dedistorted panoramic images through the pre-acquired dedistortion model.

[0053] In this embodiment, for each panoramic video frame, image sampling can be performed on the panoramic video frame to obtain multiple sampled panoramic images of the panoramic video frame, and then each sampled panoramic image of each panoramic video frame can be dedistorted to obtain a dedistorted panoramic image. After image sampling, the dedistorted panoramic image can be obtained based on the sampled panoramic image, which can effectively reduce the amount of data processing in the video quality evaluation process and improve the evaluation efficiency of the panoramic video.

[0054] In one embodiment, S210 performs image sampling on the panoramic video frame to obtain multiple sampled panoramic images of the panoramic video frame, which may include: determining multiple image sampling positions; and performing image sampling on the panoramic video frame based on the multiple image sampling positions to obtain multiple sampled panoramic images.

[0055] Among them, the position distributions of multiple image sampling positions satisfy a preset distribution condition. Exemplarily, the preset distribution condition may be specified distribution positions. For example, the upper left, lower left, upper right, lower right, and middle positions in the panoramic video frame are used as the specified distribution positions. Of course, in another example, the specified distribution positions may also be determined according to information such as the image content, composition characteristics, or image blurring condition in the panoramic video frame. Or, the preset distribution condition may be a preset sampling position distribution density. When performing image sampling, multiple image sampling positions may be randomly determined in the panoramic video frame, but the multiple image sampling positions need to satisfy the preset distribution density, that is, to avoid the image sampling positions in a certain area of the panoramic video frame being too dense or too sparse.

[0056] In practical applications, although a region can be randomly cropped from the panoramic video frame as the sampled panoramic image, the sampled panoramic image obtained in this way often fails to comprehensively reflect various aspects of information in the panoramic video frame, which easily causes the reliability of the video frame features obtained based on the image features to decline. Therefore, in this embodiment, multiple image sampling positions whose position distributions satisfy the preset distribution condition can be determined.

[0057] After determining multiple image sampling positions, image sampling can be performed on the panoramic video frame based on the multiple image sampling positions to obtain multiple sampled panoramic images. For example, pictures with a size of 224 pixels * 224 pixels can be cropped at the upper left, lower left, upper right, lower right, and middle positions in the panoramic video frame as multiple sampled panoramic images.

[0058] In this embodiment, multiple image sampling positions that satisfy the preset distribution condition can be determined. Furthermore, image sampling can be performed on the panoramic video frame based on the multiple image sampling positions to obtain multiple sampled panoramic images. Without performing deformation removal processing on the entire panoramic video frame, it is ensured that multi-angle and all-round image features can be obtained based on the multiple sampled panoramic images that meet the distribution condition, and the information volume contained in the sampled panoramic video frame and the computing resources consumed during panoramic video preprocessing can be taken into account simultaneously.

[0059] In one embodiment, obtaining the image features of the de-deformed panoramic image in S120 may include: inputting each de-deformed panoramic image of the panoramic video frame into an image feature extraction network, and the image feature extraction network obtains the image features of each de-deformed panoramic image.

[0060] In practical applications, after obtaining multiple de-deformed panoramic images of each panoramic video frame, each de-deformed panoramic image can be input into an image feature extraction network, and the image feature extraction network performs feature extraction on each input de-deformed panoramic image to obtain the image features of each de-deformed panoramic image.

[0061] In one example, the image feature extraction network can be a residual network, such as a resnet50 residual network. The residual network is composed of a large number of convolution kernels, activation functions, etc., which can effectively obtain the image features of the input de-deformed panoramic image. The de-deformed panoramic image gradually becomes smaller in size and increases in dimension in the process of passing through the image feature extraction network, and the semantic information obtained by the image feature extraction network becomes richer and richer.

[0062] In S130, based on the image features of the dedistorted panoramic image of each panoramic video frame, the video frame features of each panoramic video frame are obtained, which may include: for each panoramic video frame, fusing the image features of the dedistorted panoramic images of the panoramic video frame, and using the fused image features as the video frame features of the panoramic video frame.

[0063] Specifically, for each panoramic video frame, after the image features of each de-distorted panoramic image of the panoramic video frame are acquired, the multiple image features may be fused, and the fused image features may be used as the video frame features of the panoramic video frame.

[0064] In one embodiment, the image features of each dedistorted panoramic image of a panoramic video frame can be fused by the following steps: obtaining feature weights of the image features of each dedistorted panoramic image; based on the feature weights, performing weighted summation on the image features of each dedistorted panoramic image of the panoramic video frame, and using the features obtained after the weighted summation as the fused image features.

[0065] Specifically, after determining the image features of each dedistorted panoramic image, the feature weight of each image feature can be obtained, and then based on the feature weights of each image feature, the image features of each dedistorted panoramic image of the panoramic video frame can be weighted summed, and the features obtained after the weighted summation can be used as the fused image features.

[0066] When obtaining the feature weights of each image feature, the same feature weights can be assigned to each image feature, that is, the mean of each image feature is obtained, and the mean is used as the fused image feature. For example, each image feature can be a feature vector with a dimension of X. When fusion of features is performed, the mean of the same dimension can be taken, and then a feature vector can be determined based on the mean of the X dimensions as the video frame feature. Figure 4 As shown, for the five sampled panoramic images cropped from the upper left, lower left, upper right, lower right and middle positions of the panoramic video frame, after obtaining the corresponding de-deformed panoramic images, feature extraction can be performed through the resnet50 residual network, and the feature vector output by the network for each de-deformed panoramic image is used as the image feature, and then the mean of the five feature vectors in the same dimension is taken to obtain the video frame feature of the panoramic video frame.

[0067] Alternatively, the feature weights of each image feature may be different. For example, the feature weights may be set according to the sampling position of the dedistorted panoramic image in the panoramic video frame or the image size of the dedistorted panoramic image. For the dedistorted panoramic image sampled at the center of the panoramic video frame, the dedistorted panoramic image containing the main photographed object in the panoramic video frame, or the dedistorted panoramic image with a larger image size, a larger feature weight may be set for its image features, while a smaller feature weight may be set for the dedistorted panoramic image sampled at the edge area of ​​the panoramic video frame, the dedistorted panoramic image not containing the main photographed object, or the dedistorted panoramic image with a smaller image size.

[0068] In this embodiment, the image features of each de-distorted panoramic image of the panoramic video frame can be fused, and the fused image features can be used as the video frame features of the panoramic video frame. Compared with some video quality evaluation methods that convert the panoramic video frame into a grayscale panoramic video frame with original image information, and extract features and evaluate the panoramic video quality based on the grayscale panoramic video frame, since the number of ordinary distortion-free videos in actual applications is much higher than the number of panoramic videos, this embodiment determines the video frame features based on the image features of multiple de-distorted panoramic images of the panoramic video frame, which can provide a basis for video quality evaluation data with the help of past distortion-free videos, avoid affecting the accuracy of the evaluation results due to insufficient panoramic video samples, enhance the versatility of this application, and effectively improve the evaluation efficiency of panoramic video quality.

[0069] In one embodiment, in S130, the video frame features of each of the multiple panoramic video frames are input into a coding network, and the coding network obtains the inter-frame correlation of the multiple panoramic video frames based on the multiple video frame features, which can include: obtaining the position code of each panoramic video frame; inputting the video frame features of each panoramic video frame and the position code of the panoramic video frame into the coding network, and the coding network determines the video frame order of the multiple panoramic video frames according to the position code and determines the inter-frame correlation of the multiple panoramic video frames based on the video frame order and the video frame features of each panoramic video frame.

[0070] In practical applications, the position codes of each panoramic video frame can be obtained. Specifically, the encoding can be performed based on the video frame sequence of each panoramic video frame to obtain each position code. After the position code is obtained, the video frame features of each panoramic video frame and its corresponding position code can be input into the encoding network for parallel processing. The encoding network can determine the video frame sequence of each video frame feature according to the position code, and determine the inter-frame correlation of multiple panoramic video frames in combination with the video frame features.

[0071] In this embodiment, by inputting the video frame features and position codes of each panoramic video frame into the encoding network, the encoding network can quickly determine the temporal order of appearance of each video frame feature, providing a basis for parallel processing of the encoding network.

[0072] In one embodiment, in S140, a video evaluation result of a panoramic video is obtained according to video frame features of each panoramic video frame and inter-frame correlation of multiple panoramic video frames, which may include the following steps: determining encoding features based on inter-frame correlation of multiple panoramic video frames and video frame features of each panoramic video frame; inputting the encoding features into a trained fully connected layer to obtain a video evaluation result of the panoramic video output by the fully connected layer.

[0073] In practical applications, the coding features of a panoramic video can be determined based on the inter-frame correlation and video frame features of the panoramic video frames. For example, after obtaining the video frame features of each panoramic video frame, each video frame feature can be input into a coding network, and the coding network obtains the inter-frame correlation of multiple panoramic video frames based on the input multiple video frame features. In one example, the coding network can be the coding part of a Transformer network, which is a coding and decoding model structure that can be used to process information with strong temporal correlation, such as in the field of natural language processing. After obtaining the inter-frame correlation of multiple panoramic video frames, the coding network can generate and output coding features based on the inter-frame correlation and the video frame features of each of the multiple panoramic videos.

[0074] After obtaining the coding features, the coding features can be input into the trained fully connected layer, and the fully connected layer performs regression operation based on the input coding features to evaluate the quality of the panoramic video, and obtains the video evaluation result of the panoramic video output by the fully connected layer. The video evaluation result output by the fully connected layer can be a vector with the same dimension as the video frame feature. For example, if the dimension of the video frame feature is 512, a 512*1 fully connected layer can be used for prediction.

[0075] Exemplarily, the fully connected layer can be trained together with the encoding network. During the training process, after the encoding network inputs the encoding features into the fully connected layer, the fully connected layer predicts the evaluation results of the panoramic video based on the input encoding features, and adjusts the parameters of the encoding network and the fully connected layer based on the predicted evaluation results and the evaluation result labels of the panoramic video.

[0076] In this embodiment, the coding features can be determined based on the inter-frame correlation of multiple panoramic video frames and the video frame features of each panoramic video frame, and the coding features can be input into the trained fully connected layer. The fully connected layer classifies and evaluates the video quality based on the coding features, and outputs accurate and reliable video evaluation results.

[0077] In one embodiment, obtaining the de-deformed panoramic image of each panoramic video frame may include: using a de-deformation model that has been pre-trained through adversarial training to obtain the de-deformed panoramic image of each panoramic video frame.

[0078] The de-deformation model may be a model obtained through adversarial training based on panoramic image samples and de-deformed panoramic image samples.

[0079] In a specific implementation, after obtaining a panoramic video frame of a panoramic video, the panoramic video frame can be input into a de-distortion model that has been pre-adversarially trained. The model removes image distortion in the panoramic video frame to obtain a de-distorted panoramic image of the panoramic video frame.

[0080] In one embodiment, the de-deformation model can be trained in the following manner: obtain panoramic image samples and de-deformed panoramic image samples of the panoramic image samples; perform adversarial training on the generative model and the discriminative model to be trained based on the panoramic image samples and the de-deformed panoramic image samples until the training end conditions are met; and use the trained generative model as the de-deformation model.

[0081] In a specific implementation, panoramic image samples for training the model and their corresponding dedistorted panoramic image samples can be obtained, wherein the image distortion in the panoramic image samples can be removed by image processing software to obtain the dedistorted panoramic image samples, or an undistorted image can be used as the dedistorted panoramic image sample, and the image content in the dedistorted panoramic image sample can be distorted to obtain the panoramic image sample.

[0082] After obtaining the de-deformed panoramic image samples corresponding to the panoramic image samples, adversarial training can be performed on the generative model and the discriminative model to be trained based on the panoramic image samples and the de-deformed panoramic image samples.

[0083] In practical applications, the generative model and the discriminative model are composed of neural networks. During the training process, Figure 5As shown, the parameters of the discriminant model can be fixed first, and the panoramic image samples can be input into the generative model to be trained. The generative model obtains the de-deformed panoramic image predicted by the panoramic image samples and inputs it into the discriminant model. At the same time, the de-deformed panoramic image samples can be input into the discriminant model, and then the discriminant model can obtain the image difference between the de-deformed panoramic image samples and the predicted de-deformed panoramic image, and determine the model loss of the generative model according to the difference, and adjust the model parameters of the generative model according to the loss. When the switching condition is met, the model parameters of the generative model are fixed, and the model parameters of the discriminant model are adjusted according to the image difference identified by the discriminant model. When the switching condition is met, it is switched to the training of the generative model again. This iteration is repeated many times until the training end condition is met, and a trained generative model and a trained discriminant model can be obtained, and then a de-deformed model trained by adversarial training can be obtained based on the generative model.

[0084] In this embodiment, a de-deformation model that has been pre-trained through adversarial training can be used to quickly obtain a de-deformed panoramic image of each panoramic video frame, thereby avoiding directly evaluating the video quality of the panoramic video based on the panoramic video frames, thereby improving the evaluation efficiency of the panoramic video.

[0085] In one embodiment, the following step may be further included before S110: in response to a panoramic video playback request from a terminal, determining a panoramic video to be evaluated based on a video identifier carried in the panoramic video playback request.

[0086] The panoramic video to be evaluated may be a video in a video playlist, or may be a panoramic video that is automatically loaded and played, such as a test video or a panoramic video in a game video stream.

[0087] In a specific implementation, the terminal can detect whether there is a playback trigger event for the panoramic video, where the playback trigger event may include at least one of the following: a trigger operation for a play button, the video playback progress reaches a preset progress (the preset progress may be the time point for playing the panoramic video), and enters a preset page (for example, entering a game page, an application introduction page, etc.).

[0088] When a playback trigger event for a panoramic video is detected, a panoramic video playback request may be generated, and a corresponding panoramic video may be determined based on the video identifier carried in the panoramic video playback request, and the panoramic video may be used as the panoramic video to be evaluated. For example, if the server receives a panoramic video playback request, the video identifier may be searched to find the corresponding panoramic video as the panoramic video to be evaluated.

[0089] After obtaining the video evaluation result of the panoramic video according to the video frame features of each panoramic video frame and the inter-frame correlation of multiple panoramic video frames in S140, the method may further include the following steps: determining the target playback parameters of the panoramic video according to the video evaluation results, and adjusting the playback parameters of the terminal used to play the panoramic video according to the target playback parameters.

[0090] In actual applications, panoramic videos with different video evaluation results may have differences in one or more aspects such as file size, distortion, jitter, color, etc. Based on this, after obtaining the video evaluation results of the panoramic video, the target playback parameters of the panoramic video can be determined according to the video evaluation results, and the current playback parameters of the terminal can be adjusted according to the target playback parameters, so that the terminal can display the panoramic video under the target playback parameters.

[0091] Among them, the target playback parameters can be video playback parameters that can optimize the panoramic video playback effect. For example, if the video evaluation results indicate that the panoramic video has low distortion, rich colors and clear picture quality, then high resolution (such as 1600*900 or 1280*720) can be used as one of the target playback parameters, and the original low resolution of the terminal can be adjusted to high resolution; or, the encoding method of the original panoramic video can be adjusted according to the video evaluation results, and the encoded panoramic video can be sent to the terminal to improve transmission efficiency.

[0092] In this embodiment, the target playback parameters of the panoramic video can be determined based on the video evaluation results, and the playback parameters of the terminal used to play the panoramic video can be adjusted based on the target playback parameters, so that the playback mode of the panoramic video can match its video evaluation results and the display effect of the panoramic video can be optimized.

[0093] In one embodiment, in S110, obtaining multiple panoramic video frames of the panoramic video to be evaluated may include: obtaining the sampling number of the video frames, and determining the extraction interval of the panoramic video frames according to the total number of video frames and the sampling number of the panoramic video to be evaluated; extracting multiple panoramic video frames from the panoramic video according to the extraction interval.

[0094] Specifically, the sampling number of video frames can be obtained. The sampling number of video frames can be a fixed sampling number, and can also be determined according to the file size of the panoramic video. For example, the mapping relationship between the file size interval and the sampling number of video frames can be set in advance. After the panoramic video to be evaluated is obtained, the sampling number of video frames is determined according to the file size of the video.

[0095] After determining the sampling number of the video frames, the panoramic video to be evaluated can be divided into frames to obtain the total number of video frames after division, and the extraction interval of the panoramic video frames can be determined based on the total number of video frames and the sampling number, and then multiple panoramic video frames can be extracted from the panoramic video frames after division according to the extraction interval.

[0096] In this embodiment, the extraction interval of the panoramic video frames can be determined according to the total number of video frames and the number of samples of the panoramic video to be evaluated, and multiple panoramic video frames can be extracted from the panoramic video according to the extraction interval. There is no need to process all the video frames of the panoramic video. The comprehensiveness of the evaluation can be ensured by using the multiple panoramic video frames extracted at intervals, while improving the efficiency of video quality evaluation.

[0097] In order to enable those skilled in the art to better understand the above steps, the embodiment of the present application is illustrated below by using an example, but it should be understood that the embodiment of the present application is not limited to this.

[0098] like Figure 6 As shown, after obtaining the panoramic video to be evaluated, the panoramic video can be preprocessed to obtain multiple panoramic video frames of the panoramic video, and sampled panoramic images of a preset size can be cropped at the upper left, lower left, upper right, lower right and middle positions of each panoramic video frame.

[0099] After obtaining the sampled panoramic image, the sampled panoramic image can be output to a pre-trained generation model to obtain a de-deformed panoramic image output by the generation model, and then input into the Resnt50 residual network for image feature extraction to obtain a feature vector for each de-deformed panoramic image. Furthermore, the feature vectors of the de-deformed panoramic images from the same panoramic video frame can be fused to obtain video frame features, and the video frame features of multiple panoramic video frames and their corresponding position codes are input into the encoding network to obtain the encoding features output by the encoding network, and then the encoding features are input into the fully connected network to obtain the video evaluation results output by the fully connected network.

[0100] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0101] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store panoramic videos to be evaluated. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a panoramic video evaluation method is implemented.

[0102] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0103] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0104] Acquire multiple panoramic video frames of the panoramic video to be evaluated;

[0105] Acquire a de-deformed panoramic image of each panoramic video frame and acquire image features of the de-deformed panoramic image;

[0106] Based on the image features of the de-distorted panoramic image of each panoramic video frame, a video frame feature of each panoramic video frame is obtained, and based on the video frame sequence of the plurality of panoramic video frames and the video frame features of each of the panoramic video frames, an inter-frame correlation of the plurality of panoramic video frames is determined;

[0107] A video evaluation result of the panoramic video is obtained according to the video frame features of each of the panoramic video frames and the inter-frame correlation of the multiple panoramic video frames.

[0108] In one embodiment, when the processor executes the computer program, the steps in the other embodiments described above are also implemented.

[0109] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:

[0110] Acquire multiple panoramic video frames of the panoramic video to be evaluated;

[0111] Acquire a de-deformed panoramic image of each panoramic video frame and acquire image features of the de-deformed panoramic image;

[0112] Based on the image features of the de-distorted panoramic image of each panoramic video frame, a video frame feature of each panoramic video frame is obtained, and based on the video frame sequence of the plurality of panoramic video frames and the video frame features of each of the panoramic video frames, an inter-frame correlation of the plurality of panoramic video frames is determined;

[0113] A video evaluation result of the panoramic video is obtained according to the video frame features of each of the panoramic video frames and the inter-frame correlation of the multiple panoramic video frames.

[0114] In one embodiment, when the computer program is executed by a processor, the steps in the other embodiments described above are also implemented.

[0115] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0116] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0117] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0118] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A panoramic video evaluation method, It is characterized in that The method comprises: Acquire multiple panoramic video frames of the panoramic video to be evaluated; Obtaining a de-deformed panoramic image of each panoramic video frame and obtaining image features of the de-deformed panoramic image; the image features include image features of a plurality of de-deformed sampled panoramic images in the de-deformed panoramic image; the sampled panoramic image is obtained by sampling and de-deforming the panoramic video frame according to a plurality of image sampling positions; the position distribution of the plurality of image sampling positions satisfies a preset distribution condition; the image features of the de-deformed sampled panoramic image correspond to different feature weights according to the distribution condition; Based on the image features of the dedistorted panoramic image of each panoramic video frame and the feature weight, the image features of each dedistorted panoramic image of the panoramic video frame are fused to obtain a video frame feature of each panoramic video frame, and based on the video frame sequence of the multiple panoramic video frames and the video frame features of each panoramic video frame, a change between the video frame features of adjacent panoramic video frames is determined to obtain an inter-frame correlation of the multiple panoramic video frames; The fully connected layer obtains the video evaluation result of the panoramic video according to the video frame features of each of the panoramic video frames and the coding features corresponding to the inter-frame correlation of the multiple panoramic video frames; the input of the fully connected layer is the coding features jointly determined by the video frame features and the inter-frame correlation.

2. The method according to claim 1, It is characterized in that The step of obtaining a de-distorted panoramic image of each panoramic video frame includes: For each panoramic video frame, performing image sampling on the panoramic video frame to obtain a plurality of sampled panoramic images of the panoramic video frame; De-deformation processing is performed on each sampled panoramic image of each panoramic video frame to obtain each de-deformed panoramic image of each panoramic video frame.

3. The method according to claim 2, It is characterized in that The performing image sampling on the panoramic video frame to obtain a plurality of sampled panoramic images of the panoramic video frame includes: A plurality of image sampling positions are determined, and image sampling is performed on the panoramic video frame based on the plurality of image sampling positions to obtain a plurality of sampled panoramic images; and a position distribution of the plurality of image sampling positions satisfies a preset distribution condition.

4. The method according to claim 1, It is characterized in that The step of obtaining the video frame features of each panoramic video frame based on the image features of the de-distorted panoramic image of each panoramic video frame includes: For each panoramic video frame, image features of the de-distorted panoramic images of the panoramic video frame are fused, and the fused image features are used as video frame features of the panoramic video frame.

5. The method according to claim 4, It is characterized in that The step of fusing the image features of the de-deformed panoramic images of each panoramic video frame and the feature weight based on the image features of the de-deformed panoramic images of each panoramic video frame to obtain the video frame features of each panoramic video frame includes: Based on the feature weights, weighted summation is performed on the image features of each de-distorted panoramic image of the panoramic video frame, and the features obtained after the weighted summation are used as fused image features.

6. The method according to claim 1, It is characterized in that The determining, based on the video frame sequence of the plurality of panoramic video frames and the video frame features of each of the panoramic video frames, the change between the video frame features of adjacent panoramic video frames to obtain the inter-frame correlation of the plurality of panoramic video frames includes: Obtaining position codes of each of the panoramic video frames; The video frame features of each of the panoramic video frames and the position codes of the panoramic video frames are input into a coding network, and the coding network determines the video frame order of the multiple panoramic video frames according to the position codes and determines the changes between the video frame features of adjacent panoramic video frames based on the video frame order and the video frame features of each of the panoramic video frames, so as to obtain the inter-frame correlation of the multiple panoramic video frames.

7. The method according to claim 1, It is characterized in that The step of obtaining the video evaluation result of the panoramic video by the fully connected layer according to the video frame features of each panoramic video frame and the coding features corresponding to the inter-frame correlation of the plurality of panoramic video frames includes: Determining encoding features based on inter-frame correlations of the plurality of panoramic video frames and video frame features of each of the panoramic video frames; The encoded features are input into a trained fully connected layer to obtain a video evaluation result of the panoramic video output by the fully connected layer.

8. The method according to claim 1, It is characterized in that The step of obtaining a de-distorted panoramic image of each panoramic video frame includes: Use the de-deformation model that has been pre-trained through adversarial training to obtain a de-deformed panoramic image for each panoramic video frame.

9. The method according to claim 1, It is characterized in that The step of obtaining a plurality of panoramic video frames of the panoramic video to be evaluated includes: Acquire the sampling number of the video frame, and determine the extraction interval of the panoramic video frame according to the total number of video frames of the panoramic video to be evaluated and the sampling number; A plurality of panoramic video frames are extracted from the panoramic video according to the extraction interval.

10. The method according to any one of claims 1 to 9, It is characterized in that Before obtaining a plurality of panoramic video frames of the panoramic video to be evaluated, the method further includes: In response to a panoramic video playback request from a terminal, determining a panoramic video to be evaluated based on a video identifier carried in the panoramic video playback request; After obtaining the video evaluation result of the panoramic video according to the video frame features of each of the panoramic video frames and the inter-frame correlation of the plurality of panoramic video frames, the method further includes: The target playback parameters of the panoramic video are determined according to the video evaluation result, and the playback parameters of the terminal are adjusted according to the target playback parameters.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program. It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.

12. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Method, device and device for repositioning using panoramic image

    CN109308678A

  • No-reference video quality evaluation method based on feature fusion and recurrent neural network

    CN110677639A

  • Virtual reality video quality evaluation method and system based on generative adversarial network

    CN112004078A

  • Virtual reality video quality evaluation method based on graph convolutional neural network

    CN113688686A