Video quality detection method and system, medium, equipment and program product

By rendering volumetric videos using a 3D Gaussian splash model, selecting key viewpoints, and generating short trajectory video sequences to obtain quality data, the efficiency and accuracy issues of volumetric video quality assessment are resolved, and the sensitivity and consistency of video quality detection are improved.

CN120897052APending Publication Date: 2025-11-04MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511257342.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies are insufficient to fully capture the three-dimensional spatial structure and real-time interactive requirements of volumetric videos. Furthermore, the three-dimensional Gaussian splashing technology may cause visual defects such as artifacts and color distortion during the rendering process, resulting in low efficiency and inaccurate results in quality assessment.

Method used

A 3D Gaussian splash model is used to render scenes with potential viewpoints, output uncertainty scores, filter candidate viewpoints, generate short camera trajectory video sequences, and obtain quality data under different video playback parameters. The video quality is quantified by calculating the quality data of the video sequences.

Benefits of technology

It enables accurate quality detection of volumetric videos, reduces computational and storage burden, improves the sensitivity and repeatability of evaluation, and ensures consistent image quality and immersive experience across different terminals and network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897052A_ABST
    Figure CN120897052A_ABST
Patent Text Reader

Abstract

The invention provides a video quality detection method and system, a medium, equipment and a program product, and relates to the technical field of video data processing, and the method comprises the steps: calling a three-dimensional Gaussian splash model to render a scene of each potential viewpoint in a to-be-detected video, and outputting an uncertainty score of each potential viewpoint; determining candidate viewpoints in the potential viewpoints according to the uncertainty score; for each candidate viewpoint, defining a short camera trajectory, and generating a video sequence of the short camera trajectory; acquiring video playing quality data of each video sequence under different video playing parameters; and calculating quality data of the to-be-detected video according to the video playing quality data. According to the method, the key observation viewpoints are intelligently selected by taking uncertainty as a driving criterion, and are preferentially focused on a potential artifact high-incidence area, so that the number of evaluation viewpoints is remarkably reduced, and time and computing resource consumption are effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video data processing, in particular to a video quality detection method, system, medium, device and program product. BACKGROUND

[0002] In recent years, virtual reality, augmented reality and immersive media technology have evolved rapidly, and volumetric video, as a key content form, has been widely used in games, film production, remote meetings and other scenarios.

[0003] However, the current mainstream video quality evaluation method still takes traditional two-dimensional video as the design benchmark, and it is difficult to fully capture the unique three-dimensional spatial structure and real-time interaction requirements of volumetric video. At the same time, three-dimensional Gaussian splatting (3DGS) technology has become the focus of the industry due to its efficient rendering capabilities, but it may induce visual defects such as splatting artifacts, color distortion and motion discontinuity during reconstruction and rendering. Therefore, it is urgent to establish a quality evaluation framework for the characteristics of volumetric video to accurately quantify and optimize user experience. SUMMARY

[0004] The purpose of the present application is to provide a video quality detection method, system, computer readable storage medium and electronic device, which can accurately detect video data quality.

[0005] To solve the above technical problems, the present application provides a video quality detection method, and the specific technical solutions are as follows:

[0006] A three-dimensional Gaussian splatting model is called to render the scenes of each potential viewpoint in the video to be detected, and an uncertainty score of each potential viewpoint is outputted;

[0007] According to the uncertainty score, a candidate viewpoint in the potential viewpoint is determined;

[0008] For each candidate viewpoint, a short camera track is defined, and a video sequence of the short camera track is generated;

[0009] Video playback quality data of each video sequence under different video playback parameters is obtained;

[0010] According to the video playback quality data, quality data of the video to be detected is calculated.

[0011] Optionally, calling a three-dimensional Gaussian splatting model to render the scenes of each potential viewpoint in the video to be detected, and outputting an uncertainty score of each potential viewpoint comprises:

[0012] Adjusting the Gaussian parameters to render the same video multiple times, and calculating the variance of pixel color between each rendering result;

[0013] calculating a Gaussian uncertainty metric contributed by each pixel based on a Gaussian covariance matrix;

[0014] training the three-dimensional Gaussian splatting model to predict pixel uncertainty using variational inference and the Gaussian uncertainty metric, and outputting a pixel uncertainty map and an average uncertainty score at different viewpoints.

[0015] Optionally, for each candidate viewpoint, a short camera trajectory is defined, and generating a video sequence of the short camera trajectory comprises:

[0016] for each candidate viewpoint, defining a short camera trajectory associated with the candidate viewpoint;

[0017] generating a video sequence corresponding to the short camera trajectory using a renderer; the video sequence has a sequence length greater than a set time length and includes different rendering parameter configurations.

[0018] Optionally, when determining the candidate viewpoint from the potential viewpoints according to the uncertainty score, the method further comprises:

[0019] representing the potential viewpoints as feature vectors; the feature vectors are used to reflect position and orientation information of the potential viewpoints in the scene;

[0020] filtering target feature vectors covering the scene using a clustering algorithm, and determining candidate viewpoints corresponding to the target feature vectors.

[0021] Optionally, obtaining video playback quality data of each video sequence under different video playback parameters comprises:

[0022] based on video playback parameters of each video sequence; the video playback parameters include geometric accuracy, color fidelity, motion smoothness, and video quality;

[0023] obtaining an average opinion score of each video sequence as video playback quality data.

[0024] Optionally, after obtaining the average opinion score of each video sequence as video playback quality data, the method further comprises:

[0025] calculating a correlation result between the average opinion score and model parameters of the three-dimensional Gaussian splatting model;

[0026] optimizing the model parameters of the three-dimensional Gaussian splatting model according to the correlation result.

[0027] The application also provides a video quality detection system, comprising:

[0028] a scene rendering module configured to call a three-dimensional Gaussian splatting model to render scenes of potential viewpoints in a video to be detected, and output uncertainty scores of the potential viewpoints.

[0029] a candidate view point module configured to determine a candidate view point from the potential view points according to the uncertainty score;

[0030] a video sequence generation module configured to define a short camera track for each of the candidate view points and generate a video sequence of the short camera track;

[0031] a quality data acquisition module configured to acquire video playback quality data of each of the video sequences under different video playback parameters;

[0032] a quality detection module configured to calculate quality data of the video to be detected according to the video playback quality data.

[0033] The application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method.

[0034] The application further provides an electronic device comprising a memory and a processor, wherein the memory has a computer program stored therein, and the processor, when invoking the computer program in the memory, implements the steps of the method.

[0035] The application further provides a computer program product comprising a computer program, wherein the computer program, when executed, implements the steps of the method.

[0036] The application provides a video quality detection method, comprising: rendering scenes of potential view points in a video to be detected by invoking a three-dimensional Gaussian splash model, and outputting uncertainty scores of the potential view points; determining a candidate view point from the potential view points according to the uncertainty scores; defining a short camera track for each of the candidate view points, and generating a video sequence of the short camera track; acquiring video playback quality data of each of the video sequences under different video playback parameters; and calculating quality data of the video to be detected according to the video playback quality data.

[0037] The application synchronously outputs uncertainty scores in the rendering stage by the three-dimensional Gaussian splash model, can find potential artifact high-risk areas in advance, and realizes active positioning of defects instead of passive detection; after screening candidate view points by taking uncertainty as a clue, only generates short track videos at key view points, significantly reduces the calculation and storage burden brought by full track traversal; plays back and collects objective quality data in the short track scale under multiple parameters, retains the interactive characteristics of volumetric videos, avoids interference of long time sequence contents on the evaluation model, and improves the repeatability and sensitivity of the results. By calculating the quality data of the video sequences, the quality indicators are quantified, thereby effectively improving the picture quality consistency and immersion of the volumetric video under different terminals and network conditions.

[0038] The application further provides a video quality detection system, a computer readable storage medium, an electronic device and a computer program product, which have the above beneficial effects, and details are not repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.

[0040] Figure 1 A flowchart of a video quality detection method provided by an embodiment of the present application;

[0041] Figure 2 A structural schematic diagram of a video quality detection system provided by an embodiment of the present application;

[0042] Figure 3 A structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0044] The existing volumetric video subjective quality evaluation method usually generates a 2D video sequence through fixed path rendering, does not consider the characteristics of 3D GS model, and may ignore the key quality problem area, resulting in low evaluation efficiency or inaccurate results.

[0045] To solve the above technical defects, see Figure 1 , Figure 1 A flowchart of a video quality detection method provided by an embodiment of the present application, the method comprising:

[0046] S101: calling a three-dimensional Gaussian splash model to render scenes of each potential viewpoint in a to-be-detected video, and outputting an uncertainty score of each potential viewpoint;

[0047] S102: determining a candidate viewpoint in the potential viewpoint according to the uncertainty score;

[0048] S103: defining a short camera trajectory for each of the candidate viewpoints, and generating a video sequence of the short camera trajectory;

[0049] S104: obtaining video playback quality data of each of the video sequences under different video playback parameters;

[0050] S105: calculating quality data of the video to be detected according to the video playback quality data.

[0051] Step S101 aims to analyze the rendering uncertainty under different viewpoints with three-dimensional Gaussian splatting model parameters (such as Gaussian density, covariance matrix), and identify the areas where artifacts are likely to occur. The three-dimensional Gaussian splatting model is composed of a group of 3D Gaussians, each of which is defined by parameters such as position, covariance matrix, color (usually using spherical harmonics to represent the viewing angle dependent effect), and opacity.

[0052] In a feasible implementation, the Gaussian parameters can be adjusted to render the same video multiple times, the variance of the pixel color between the rendering results of each time is calculated, the Gaussian uncertainty metric contributed by each pixel is calculated based on the Gaussian covariance matrix, and finally the three-dimensional Gaussian splatting model is trained to predict the pixel uncertainty using variational inference and the Gaussian uncertainty metric, and the pixel uncertainty map and the average uncertainty score under different viewpoints are output.

[0053] In this process, a pre-trained 3D GS model is input, which is usually generated through multi-view video or point cloud data.

[0054] For each potential viewpoint, the scene is rendered and the uncertainty of each pixel is estimated. Random rendering can be used to render the same viewpoint multiple times by slightly perturbing the Gaussian parameters (such as position, covariance), and the variance of the pixel color is calculated.

[0055] In analyzing the uncertainty, the uncertainty metric (such as the determinant of the covariance matrix) of the contribution Gaussian of each pixel is calculated based on the covariance matrix of the Gaussian. At the same time, the model is trained to predict the uncertainty of each pixel using variational inference (such as [Stochastic Gaussian Splatting]). Finally, the average uncertainty score of each viewpoint is output, reflecting the degree of possible impairment of the rendering quality.

[0056] In an example implementation, an additional learnable "confidence channel" can be attached to each 3D Gaussian sphere before training, which describes the degree of certainty of the Gaussian sphere on the final pixel color. To make the confidence channel statistically meaningful, it is treated as a random variable, whose distribution is predicted by another small neural network: the network takes the position, size, rotation and spherical harmonic coefficients of a Gaussian sphere as input, and outputs an average confidence and a confidence fluctuation range, which form an approximate posterior distribution. During rendering, along a pixel ray, all Gaussian spheres that the ray touches are blended in color according to the volume rendering formula, while the confidence of each sphere is also blended with the same weights during the blending process. To avoid repeated calculations, a Monte Carlo approach is used: different confidence values are sampled for the same ray multiple times, each time obtaining a slightly different pixel color, and the degree of dispersion of the color intuitively reflects the uncertainty of the pixel.

[0057] In step S102, candidate viewpoints in the potential viewpoints are determined according to the uncertainty scores. Before step S102 is performed, a candidate viewpoint set can be generated by uniform distribution or scene geometry-based sampling, and then the average uncertainty score of each candidate viewpoint is calculated. The first N viewpoints with the highest uncertainty scores are selected. Here, N is not specifically limited, and can be determined according to the evaluation requirements and resources (usually 5-10). At this time, the output is a set of key viewpoint coordinates and camera paths.

[0058] In a feasible implementation, the potential viewpoints can be represented as feature vectors, so that a clustering algorithm is used to screen target feature vectors covering the scene, and candidate viewpoints corresponding to the target feature vectors are determined. The feature vector is used to reflect the position and orientation information of the potential viewpoint in the scene.

[0059] When the potential viewpoints are processed, they are encoded as feature vectors with spatial position information and orientation attributes. Based on the feature vectors, clustering operations are carried out, which can automatically mine highly representative regions in both geometric structure and perspective characteristics, thereby effectively reducing the number of redundant viewpoints.

[0060] The clustering result has the advantage of intuitive visualization, which can clearly show the coverage of the scene, ensuring that the selected candidate viewpoints are evenly and reasonably distributed, and no key details are missed. In addition, the feature vector used has the characteristics of low dimensionality, and the calculation process is efficient and fast. This enables the system to complete the screening work in real time when facing a larger search space, without the need to render each point one by one, thereby greatly reducing the resource consumption in the preprocessing stage.

[0061] In this way, the candidate viewpoints obtained on the one hand significantly compress the sample quantity required in the subsequent quality evaluation link, and on the other hand fully retain diversified visual angle information. This optimization strategy makes the entire detection process achieve a good balance between running efficiency and the accuracy of the detection result.

[0062] In step S103, for each of the candidate viewpoints, a short camera trajectory is defined, and a video sequence of the short camera trajectory is generated. Specifically, a short camera trajectory related to the candidate viewpoint can be defined, such as a translation or a wrap-around path around the viewpoint. Then, a video sequence corresponding to the short camera trajectory is generated using a renderer. The sequence length of the video sequence is greater than a set time length to ensure that the video observer can perceive the quality, and different rendering parameter configurations, such as the number of Gaussians and transparency mixing modes, can be included.

[0063] After obtaining the video sequence in step S103, step S104 aims to obtain video playback quality data of each of the video sequences under different video playback parameters.

[0064] The video playback quality data can be divided into subjective data and objective data. The subjective data can be obtained by obtaining the video observer score. Specifically, the average opinion score of each of the video sequences can be obtained based on the video playback parameters of each of the video sequences as the video playback quality data. The video playback parameters include geometric accuracy, color fidelity, motion smoothness, and video quality. The objective data mainly includes the video quality of the video sequence.

[0065] For the subjective data, a VR headset (such as Meta Quest or Apple VisionPro) can be used in a laboratory environment to present the test sequence and provide a 6DoF interactive experience. The laboratory settings recommended by ITU-T P.910 are followed during the process to control environmental variables (such as lighting, distance, and display resolution). By recruiting a number of observers, it is ensured that there is no professional evaluation background to reduce bias. A 5-point MOS rating scale is used to evaluate the following dimensions:

[0066] Geometric accuracy: the realism of scene shapes and details.

[0067] Color fidelity: the accuracy of color and perspective-dependent effects.

[0068] Motion smoothness: the continuity of motion in dynamic scenes.

[0069] Overall quality: the comprehensive perceptual experience.

[0070] An interactive interface is provided to allow observers to freely explore and score in real time in the VR environment.

[0071] Thus, the observer scores of each video sequence are obtained as the subjective data.

[0072] For objective data, including but not limited to, peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), mean squared error (MSE), multi-scale SSIM (MS-SSIM), video quality metric model (VMAF), etc.

[0073] In step S105, the quality data of the to-be-detected video is calculated according to the video playback quality data.

[0074] Taking subjective data as an example, the MOS (Mean Opinion Score) of each test sequence can be calculated. The MOS of all sequences is aggregated using simple average or weighted average (based on uncertainty score) to obtain the overall quality score of the volumetric video. MOS is an index widely used to evaluate the subjective quality of multimedia content such as audio / video, image, and volumetric video.

[0075] The embodiment fully taps the inherent uncertainty estimation of the three-dimensional Gaussian scatter point (3DGS) model, intelligently selects key observation viewpoints with uncertainty as the driving criterion, and focuses on potential artifact-prone areas, thereby significantly reducing the number of evaluation viewpoints while effectively reducing the consumption of time and computing resources. For the splatter artifacts and color distortion specific to 3DGS, the embodiment can achieve more accurate and targeted quality evaluation. Further, six-degree-of-freedom interaction can be supported in a virtual reality environment, allowing observers to truly and comprehensively perceive the quality performance of volumetric video in an immersive space. In addition, the application has good scalability and can be smoothly migrated to dynamic scenes, virtual meetings, scene reconstruction and other 3DGS applications, continuously releasing its evaluation and optimization value.

[0076] As can be seen from the above process, the application synchronously outputs uncertainty scores in the rendering phase through the three-dimensional Gaussian splatter model, can discover potential artifact-prone areas in advance, and realizes active positioning of defects instead of passive detection. After screening candidate viewpoints based on uncertainty, only short-track videos are generated at key viewpoints, significantly reducing the computing and storage burden brought by full-track traversal. On the short-track scale, multiple parameters are played back and objective quality data is collected, which not only preserves the interactive characteristics of volumetric video, but also avoids the interference of long-time sequence content on the evaluation model, improving the repeatability and sensitivity of the results. By calculating the quality data of the video sequence, the quality indicators are quantified, thereby effectively improving the picture quality consistency and immersion of the volumetric video under different terminal and network conditions.

[0077] In an alternative embodiment, a correlation result between the mean opinion score and the model parameters of the three-dimensional Gaussian splash model can be calculated, so as to optimize the model parameters of the three-dimensional Gaussian splash model according to the correlation result.

[0078] In a specific application, the model parameter features can be extracted from the same set of rendering samples. The three-dimensional Gaussian splash model generally includes the mean, covariance, opacity, spherical harmonic coefficient, etc. of each Gaussian element. In order to reduce the dimension and avoid overfitting, the Gaussian elements can be clustered or divided according to a spatial grid first, and then the parameter statistics (mean, variance, extreme value, entropy, etc.) in each cluster or grid are counted, and finally an interpretable parameter feature vector is constructed for each rendering sample.

[0079] Subsequently, the MOS vector and the parameter feature vector are aligned in the time or spatial dimension, ensuring that they correspond one by one. Using Pearson correlation coefficient, mutual information, maximum information coefficient or feature importance based on random forest, the correlation measure between each parameter feature and MOS is calculated. The correlation analysis result will be presented in the form of a heat map or a sorted list, directly showing which parameters or statistics have the greatest impact on the perceived quality.

[0080] After obtaining the significantly correlated parameters, an optimization strategy is developed to enhance the parameters with positive correlation and low numerical value, and to suppress the parameters with negative correlation and high numerical value. For example, if the logarithm of the covariance determinant has a significant negative correlation with MOS, it means that the excessive volume of the Gaussian ellipsoid will reduce the visual quality, and at this time a regularization term can be added to the loss function to constrain the volume of the Gaussian element; if the mean of the opacity has a positive correlation with MOS, the network can be encouraged to increase the opacity of the visible Gaussian element during training.

[0081] The optimization stage adopts gradient-based backpropagation or reinforcement learning strategy. On the basis of the original reconstruction loss, a perception loss is introduced, so that the network directly maximizes the expected MOS while reducing the geometric error. After optimization, the subjective experiment is performed again to verify whether the MOS is improved; if the improvement is insufficient, the feature extraction, correlation analysis and parameter adjustment are iteratively performed until the MOS converges or reaches the target threshold. By quantifying the relationship between the subjective perception of the human eye and the model parameters, blind parameter tuning is avoided; the perception-driven regularization term significantly reduces artifacts, blurring and floating objects while maintaining geometric accuracy, making the rendering result more consistent with the human eye preference. Since the optimization target is directly related to the subjective score, the user satisfaction of the final model in the real application scenario is improved, and the subsequent manual tuning cost is reduced.

[0082] Reference Figure 2 , Figure 2 A video quality detection system structure schematic diagram provided by an embodiment of the present application, the system comprises:

[0083] a scene rendering module, configured to invoke a three-dimensional Gaussian splash model to render scenes of each potential viewpoint in a to-be-detected video, and output an uncertainty score of each potential viewpoint;

[0084] a candidate viewpoint module, configured to determine a candidate viewpoint from the potential viewpoints according to the uncertainty score;

[0085] a video sequence generation module, configured to define a short camera track for each candidate viewpoint, and generate a video sequence of the short camera track;

[0086] a quality data acquisition module, configured to acquire video playback quality data of each video sequence under different video playback parameters;

[0087] a quality detection module, configured to calculate quality data of the to-be-detected video according to the video playback quality data.

[0088] Based on the above embodiment, as a preferred embodiment, invoking a three-dimensional Gaussian splash model to render scenes of each potential viewpoint in a to-be-detected video, and outputting an uncertainty score of each potential viewpoint includes:

[0089] adjusting Gaussian parameters to render the same video multiple times, and calculating a variance of pixel colors between each rendering result;

[0090] calculating a Gaussian uncertainty measure contributed by each pixel based on a Gaussian covariance matrix;

[0091] training the three-dimensional Gaussian splash model to predict pixel uncertainty using variational inference and the Gaussian uncertainty measure, and outputting a pixel uncertainty map under different viewpoints and an average uncertainty score.

[0092] Based on the above embodiment, as a preferred embodiment, defining a short camera track for each candidate viewpoint, and generating a video sequence of the short camera track includes:

[0093] defining a short camera track related to the candidate viewpoint for each candidate viewpoint;

[0094] generating a video sequence corresponding to the short camera track using a renderer; the video sequence has a sequence length greater than a set time length and includes different rendering parameter configurations.

[0095] Based on the above embodiment, as a preferred embodiment, when determining a candidate viewpoint from the potential viewpoints according to the uncertainty score, further includes:

[0096] representing the potential viewpoint as a feature vector; the feature vector is used to reflect position and orientation information of the potential viewpoint in a scene;

[0097] The target feature vector covering the scene is screened by using a clustering algorithm, and a candidate viewpoint corresponding to the target feature vector is determined.

[0098] Based on the above-mentioned embodiments, as a preferred embodiment, obtaining the video playback quality data of each video sequence under different video playback parameters comprises:

[0099] Based on the video playback parameters of each video sequence, the video playback parameters include geometric accuracy, color fidelity, motion smoothness, and video quality.

[0100] Obtaining the mean opinion score of each video sequence as the video playback quality data.

[0101] Based on the above-mentioned embodiments, as a preferred embodiment, after obtaining the mean opinion score of each video sequence as the video playback quality data, the method further comprises:

[0102] Calculating the correlation result between the mean opinion score and the model parameters of the three-dimensional Gaussian splash model;

[0103] According to the correlation result, the model parameters of the three-dimensional Gaussian splash model are optimized.

[0104] The present application also provides an embodiment corresponding to a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the method described in the above method embodiment.

[0105] It can be understood that if the method in the above-mentioned embodiments is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and executes all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0106] The computer-readable storage medium provided in the present embodiment includes the above-mentioned method, and the effects are the same as above.

[0107] The present application also provides an electronic device, referring to Figure 3 , the structural diagram of an electronic device provided by the present application, as Figure 3 shown, can include a processor 1410 and a memory 1420.

[0108] The processor 1410 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 1410 can be implemented in at least one of a hardware form of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array). The processor 1410 can also include a main processor and a coprocessor, the main processor being a processor for processing data in an awake state, also known as a CPU (Central Processing Unit), and the coprocessor being a low-power processor for processing data in a standby state. In some embodiments, the processor 1410 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content required to be displayed by the display screen. In some embodiments, the processor 1410 can also include an AI (Artificial Intelligence) processor for processing computing operations related to machine learning.

[0109] The memory 1420 can include one or more computer-readable storage media that can be non-transitory. The memory 1420 can also include a high-speed random access memory, and a nonvolatile memory such as one or more disk storage devices, flash storage devices. In this embodiment, the memory 1420 is at least used to store the following computer program 1421, wherein the computer program is loaded and executed by the processor 1410, and can implement the related steps in the method executed by the electronic device side disclosed in any of the preceding embodiments. In addition, the resources stored by the memory 1420 can also include an operating system 1422 and data 1423, and the storage mode can be temporary storage or permanent storage. The operating system 1422 can include Windows, Linux, Android, and the like.

[0110] In some embodiments, the electronic device can also include a display screen 1430, an input / output interface 1440, a communication interface 1450, a sensor 1460, a power supply 1470, and a communication bus 1480.

[0111] Of course, Figure 3 The structure of the electronic device shown does not constitute a limitation on the electronic device in the embodiments of the present application, and in actual applications, the electronic device can include more or fewer components than those shown. Figure 3more or less parts, or combinations of parts, are shown.

[0112] The various embodiments described in the specification are presented for the purpose of illustration and description. Each of the embodiments highlights a different aspect of the application. The embodiments are not mutually exclusive, and parts of one embodiment can be combined with parts of another embodiment.

[0113] The principles and implementations of the present application are described in the specification with reference to specific examples. The above description of the embodiments is only intended to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.

[0114] It should also be noted that in the specification, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

Claims

1. A video quality detection method, characterized in that, include: The three-dimensional Gaussian splash model is called to render the scene of each potential viewpoint in the video to be detected, and the uncertainty score of each potential viewpoint is output. Candidate viewpoints among the potential viewpoints are determined based on the uncertainty score; For each candidate viewpoint, a short camera trajectory is defined, and a video sequence of the short camera trajectory is generated; Obtain video playback quality data for each of the video sequences under different video playback parameters; The quality data of the video to be tested is calculated based on the video playback quality data.

2. The video quality detection method according to claim 1, characterized in that, The three-dimensional Gaussian splash model is used to render the scene of each potential viewpoint in the video to be detected, and the uncertainty score of each potential viewpoint is output, including: Adjust the Gaussian parameters and render the same video multiple times, then calculate the variance of pixel color between each rendering result; Calculate the Gaussian uncertainty measure contributed by each pixel based on the Gaussian covariance matrix; The three-dimensional Gaussian splash model is trained using variational inference and the Gaussian uncertainty metric to predict pixel uncertainty, and pixel uncertainty maps and average uncertainty scores are output from different viewpoints.

3. The video quality detection method according to claim 1, characterized in that, For each candidate viewpoint, a short camera trajectory is defined, and a video sequence of the short camera trajectory is generated, including: For each candidate viewpoint, a short camera trajectory associated with the candidate viewpoint is defined; A video sequence corresponding to the short camera trajectory is generated using a renderer; the length of the video sequence is greater than a set duration and includes different rendering parameter configurations.

4. The video quality detection method according to claim 1, characterized in that, When determining candidate viewpoints among the potential viewpoints based on the uncertainty score, the method further includes: The potential viewpoint is represented as a feature vector; the feature vector is used to reflect the position and orientation information of the potential viewpoint in the scene; Clustering algorithms are used to filter target feature vectors covering the scene and to determine candidate viewpoints corresponding to the target feature vectors.

5. The video quality detection method according to claim 1, characterized in that, Obtaining video playback quality data for each of the video sequences under different video playback parameters includes: Video playback parameters based on each of the video sequences; the video playback parameters include geometric accuracy, color fidelity, motion smoothness, and video quality; The average opinion score of each video sequence is obtained as video playback quality data.

6. The video quality detection method according to claim 5, characterized in that, After obtaining the average opinion score of each video sequence as video playback quality data, the method further includes: Calculate the correlation results between the average opinion score and the corresponding model parameters of the three-dimensional Gaussian splash model; The model parameters of the three-dimensional Gaussian splash model are optimized based on the correlation results.

7. A video quality inspection system, characterized in that, include: The scene rendering module is used to call the three-dimensional Gaussian splash model to render the scene of each potential viewpoint in the video to be detected, and output the uncertainty score of each potential viewpoint. The candidate viewpoint module is used to determine candidate viewpoints among the potential viewpoints based on the uncertainty score. A video sequence generation module is used to define a short camera trajectory for each of the candidate viewpoints and generate a video sequence of the short camera trajectory; The quality data acquisition module is used to acquire video playback quality data of each video sequence under different video playback parameters; The quality detection module is used to calculate the quality data of the video to be detected based on the video playback quality data.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the method as claimed in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the steps of the method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, Includes a computer program, which, when executed, implements the steps of the method as described in any one of claims 1 to 6.