A stereoscopic video adaptive transmission method and system based on a video quality perception model
By adopting an adaptive transmission method based on a video quality perception model and a super-resolution model, the problem of insufficient stereoscopic video transmission quality under bandwidth constraints is solved, achieving high-quality video transmission and improved user experience under limited bandwidth.
Patent Information
- Application Number
- CN202410475591.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-19
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-04-19
AI Technical Summary
Existing technologies cannot guarantee the transmission quality of stereoscopic video under bandwidth constraints, resulting in a poor user experience.
An adaptive transmission method based on a video quality perception model is adopted. This method involves point cloud downsampling, training a super-resolution model, dividing the point cloud into blocks, and predicting video block transmission based on the user's viewpoint and bandwidth. Furthermore, a distortion prediction and rendering optimization network is used to improve video quality.
Under bandwidth-constrained conditions, it significantly improves the transmission quality and user experience of stereoscopic video, maximizing the perceived quality of video.
Smart Images

Figure CN118400507B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video transmission, in particular to a stereoscopic video adaptive transmission method and system based on a video quality perception model. BACKGROUND
[0002] With the continuous development of virtual reality (VR) and augmented reality (AR), stereoscopic video has wide application prospects in the fields of medical treatment, education and entertainment, etc. It can provide users with six degrees of freedom (6DoF) movement ability when watching videos, so that users can obtain highly immersive and interactive video watching experience. A stereoscopic video frame is usually represented by a point cloud or a polygon mesh. The point cloud contains rich information such as three-dimensional coordinates, color, classification value, intensity value, time, etc. and is the main research object at present. A single frame of stereoscopic video will contain more than 100 million points. If it is not compressed, at least 3.6 Gbps of bandwidth is required to ensure that the video can be played on the client at a frame rate of 30 frames per second. This is far beyond the current network bandwidth. Therefore, how to transmit the point cloud stream so that the limited bandwidth resources can meet the transmission demand of the video to improve the user experience has become a key problem.
[0003] The prior art discloses a stereoscopic video transmission method, which is divided into video depth map compression encoding, channel coding, network transmission and channel decoding. The channel coding includes the following steps: dividing the video sequence of the source into levels; virtually expanding each layer of data according to an expansion factor; dividing the obtained virtually expanded layer data into windows; using a tree algorithm to perform progressive analysis on the window data; performing LT coding and index replacement according to the robust solitary wave degree distribution; and obtaining the coded code word. Although this method can strengthen the protection of color video data, improve the picture quality of the terminal reconstructed virtual viewpoint, ensure the reliable transmission of stereoscopic video, realize the compatibility of stereoscopic display and planar display, and meet the needs of customers, it uses a traditional coding method to transmit the video, which requires a large bandwidth and cannot guarantee the video transmission quality under the condition of limited bandwidth. SUMMARY
[0004] The primary object of the present application is to overcome the problems existing in the prior art and provide a stereoscopic video adaptive transmission method based on a video quality perception model. The present application can guarantee the video transmission quality under the condition of limited bandwidth, thereby improving the user experience.
[0005] As another object of the present application, a system is provided which is adapted to the method described above.
[0006] In order to achieve the above primary object, the present application provides a stereoscopic video adaptive transmission method based on a video quality perception model, which comprises:
[0007] Step S1: obtaining a stereoscopic video to be transmitted;
[0008] Step S2: the server performs point cloud downsampling processing on the stereoscopic video to obtain a low-resolution point cloud;
[0009] Step S3: the server trains a super-resolution model using the low-resolution point cloud and the stereoscopic video to obtain a trained super-resolution model;
[0010] Step S4: the server divides the low-resolution point cloud to obtain a low-resolution point cloud block, and stores the low-resolution point cloud block and data information related to the low-resolution point cloud block;
[0011] Step S5: the client tracks a user's viewpoint and predicts network bandwidth to obtain a prediction result of the user's viewpoint and network bandwidth;
[0012] Step S6: the client sends request information to the server according to the viewpoint and the prediction result;
[0013] Step S7: the server transmits a stereoscopic video block and the trained super-resolution model to the client according to the request information using a video adaptive transmission algorithm based on a video quality perception model, the video quality perception model being determined by the following formula:
[0014]
[0015] wherein, denotes a downsampling rate multiplier, denotes a viewpoint movement speed multiplier, denotes a viewpoint rotation speed multiplier, denotes a user viewing distance, denotes a viewpoint movement speed, denotes a viewpoint rotation speed, denotes a downsampling rate; the video adaptive transmission algorithm aims to maximize the user's video perception quality, realizes video transmission by making quality level decisions on the video blocks to be transmitted, and the maximization of the user's video perception quality is determined by the following formula:
[0016]
[0017]
[0018]
[0019]
[0020] wherein, denotes a total bandwidth budget, Represents video block The corresponding video quality score, Represents video block Bandwidth consumed Indicates the index of the video block. An index representing the video quality level. Indicates the number of video blocks. Number of video quality levels Indicates whether the selected set of video blocks contains ,like =1 indicates that the selected video block contains , =0 indicates that the selected video block does not contain And if the selected video block contains ,express The selected video block does not contain the data being transmitted. , then it means It will not be transmitted;
[0021] Step S8: The client receives and assembles the stereoscopic video blocks to obtain a complete stereoscopic video, and processes the complete stereoscopic video using a trained super-resolution model to obtain a high-quality 2D rendering image. The player then plays the high-quality 2D rendering image to the user, completing the transmission.
[0022] Furthermore, a loss function is used in step S3. Optimize the training of the super-resolution model, the loss function Determined by the following formula:
[0023]
[0024] in, This refers to a rendered image. Represents a depth map. This refers to a high-quality rendered image generated by the network. and realistic high-quality renderings The mean square error between them Distortion indicator graph generated by the network And the actual depth noise map The mean square error between them This parameter is used to adjust the degree of impact of the two losses on the overall network.
[0025] Furthermore, the super-resolution model described in step S3 includes two sub-networks: a distortion prediction network and a rendering optimization network.
[0026] The distortion prediction network is configured to obtain a distortion indication map from the low-quality depth map, the distortion indication map being configured to indicate distortion of a low-quality rendering map;
[0027] The rendering optimization network is configured to obtain a high-quality rendering map from the distortion indication map and the low-quality rendering map.
[0028] Further, the low-quality depth map and the low-quality rendering map are obtained by pre-rendering processing of the low-resolution point cloud and user viewpoint information.
[0029] Further, the data information in step S4 includes spatial position, playing order, uniform resource locator, encoding format, resolution of the low-resolution point cloud block, and storage location of the super-resolution model.
[0030] Further, the request information in step S6 includes obtaining the data information, the stereoscopic video, and the trained super-resolution model.
[0031] Further, the maximum user video perceptual quality problem is solved by using a rate adaptation algorithm based on a pruned model control prediction architecture, the pruning operation including:
[0032] Different video blocks are assigned different code rates according to distance from the same user, and the closer the distance, the smaller the code rate;
[0033] If the video bandwidth consumption is greater than the bandwidth constraint when all video blocks are assigned the minimum code rate, the minimum code rate is assigned,
[0034] If the video bandwidth consumption is less than the bandwidth constraint when all video blocks are assigned the maximum code rate, the maximum code rate is assigned.
[0035] Within a single model control prediction window range, the code rate assignment decision is unchanged.
[0036] To achieve another object of the present application, the present application provides a stereoscopic video adaptive transmission system based on a video quality perception model, including a server and a client, the server being configured to obtain a stereoscopic video to be transmitted, perform point cloud downsampling processing on the stereoscopic video, train a super-resolution model using the low-resolution point cloud and the stereoscopic video, divide the low-resolution point cloud, store the low-resolution point cloud blocks and data information related to the low-resolution point cloud blocks, and transmit stereoscopic video blocks and the trained super-resolution model to the client according to request information using a video adaptive transmission algorithm based on a video quality perception model; the client being configured to track a user viewpoint and a prediction network bandwidth, send request information to the server according to the viewpoint and prediction results, receive and assemble the stereoscopic video blocks, obtain a complete stereoscopic video, and process the complete stereoscopic video using the trained super-resolution model to obtain a high-quality 2D rendering map.
[0037] Compared with the prior art, the present application has the beneficial effects that:
[0038] The present application transmits stereoscopic video blocks and trained super-resolution models to the client by using a video adaptive transmission algorithm based on a video quality perception model, wherein the video adaptive transmission algorithm takes bandwidth constraints, candidate code rates and user perspectives as inputs, sorts the video blocks to be transmitted according to the distances of the viewing distances when generating the candidate code rates, and assigns code rates to each block in compliance with the constraints in the pruning operation, the algorithm calculates the bandwidth consumption and the video quality score for each candidate code rate allocation scheme, and selects the code rate allocation scheme that maximizes the video quality under the condition of meeting the bandwidth limit; the super-resolution model includes two sub-networks, namely a distortion prediction network and a rendering optimization network, the distortion prediction network is used to obtain a distortion indication map according to a low-quality depth map, the distortion indication map is used to indicate the distortion of a low-quality rendering map, and the rendering optimization network is used to obtain a high-quality rendering map according to the distortion indication map and the low-quality rendering map, thereby improving the video transmission quality and improving the user experience. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 is a stereoscopic video adaptive transmission method flow chart based on a video quality perception model according to an embodiment of the present application;
[0040] Figure 2 is a stereoscopic video adaptive transmission system block diagram based on a video quality perception model according to an embodiment of the present application;
[0041] Figure 3 is a super-resolution model structure diagram of a stereoscopic video adaptive transmission method based on a video quality perception model according to an embodiment of the present application;
[0042] Figure 4 is a performance diagram of different code rate allocation algorithms of a stereoscopic video adaptive transmission method based on a video quality perception model in a Longdress data set transmission task according to an embodiment of the present application;
[0043] Figure 5 is a performance diagram of different code rate allocation algorithms of a stereoscopic video adaptive transmission method based on a video quality perception model in a Soldier data set transmission task according to an embodiment of the present application;
[0044] Figure 6 is a performance diagram of different model SR results of a stereoscopic video adaptive transmission method based on a video quality perception model on the VIF index according to an embodiment of the present application;
[0045] Figure 7This is a graph showing the performance of different models of a stereo video adaptive transmission method based on a video quality perception model in terms of the VIF index, according to an embodiment of the present invention. Detailed Implementation
[0046] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0047] Example 1
[0048] like Figure 1 and 3 As shown, a preferred embodiment of the present invention provides a stereoscopic video adaptive transmission method based on a video quality perception model, comprising:
[0049] Step S1: Acquire the stereoscopic video to be transmitted;
[0050] Step S2: The server performs point cloud downsampling processing on the stereoscopic video to obtain a low-resolution point cloud;
[0051] Step S3: The server uses low-resolution point cloud and stereo video to train a super-resolution model and obtain a trained super-resolution model.
[0052] In this embodiment, a loss function is used in step S3. Optimize the training of the super-resolution model, loss function Determined by the following formula:
[0053]
[0054] in, This refers to a rendered image. Represents a depth map. This refers to a high-quality rendered image generated by the network. and realistic high-quality renderings The mean square error between them Distortion indicator graph generated by the network And the actual depth noise map The mean square error between them denotes the parameter for adjusting the degree of influence of the two-part loss on the overall network. The loss function respectively quantifies the gap between the model's prediction results of the rendered image and the depth map and the actual expected results, and guides the optimization process of the model. In order to train the super-resolution network using a given stereoscopic video, the embodiment divides the video into groups of frames (GOF) of the same size, and each GOF contains 5 frames of pictures. During training, the first frame in each GOF is used to generate the corresponding rendered image and depth map according to different distances, viewpoints and down-sampling rates as a training set to train the network.
[0055] The super-resolution model in step S3 includes two sub-networks, a distortion prediction network and a rendering optimization network, the distortion prediction network is used to obtain a distortion indication map from a low-quality depth map, the distortion indication map is used to indicate the distortion of a low-quality rendered image; the rendering optimization network is used to obtain a high-quality rendered image from the distortion indication map and the low-quality rendered image.
[0056] Further, the low-quality depth map and the low-quality rendered image are obtained through pre-rendering processing according to a low-resolution point cloud and user viewpoint information. The structure of the super-resolution model is as shown in Figure 3 The model is expected to be able to hide the distortion of the point cloud rendered image caused by down-sampling, so we model the point cloud rendered image quality degradation model. In the 2D image rendered from the low-resolution point cloud, the pixels can be divided into two categories: real pixels and distorted pixels. Unlike common image degradation models, the degradation of the 2D image rendered from the point cloud caused by down-sampling can be mathematically represented as:
[0057]
[0058] wherein denotes the original rendered image (i.e. groundtruth), is the rendered image with distortion, and are the two degradations caused by point cloud down-sampling, denotes pixel-wise addition operation, and denotes pixel-wise subtraction operation. The introduction of is due to the occlusion relationship of the points in the point cloud. After removing the front points, the occluded points in the rear can naturally appear on the screen, while the distortion is caused by the lack of information of the corresponding pixels in the rendered image after removing the related points.
[0059] To better hide the distortion, the embodiment introduces a depth map into our network. The depth map of a point cloud is an image containing information about the distance of each point from a certain viewpoint. For an undistorted point cloud, its depth map is smooth, and the values are continuous. But when the point cloud is down-sampled, the depth map will become unsmooth due to the noise points. Based on this observation, the embodiment uses the depth map to better distinguish between real pixels and distortion, improving the effect of super-resolution of the rendered image. By Figure 3 It can be known that the super-resolution model based on three-dimensional point cloud rendering image proposed in the embodiment includes two sub-networks, i.e. a distortion prediction network and a rendering optimization network. For each low-resolution point cloud, the system will perform pre-rendering of a two-dimensional image based on the viewing angle of the current user through the rendering pipeline. Then, the pre-rendering generates a low-quality rendered image and a depth map which are transmitted to the subsequent network. Among them, the depth map will be sent to the distortion prediction network to estimate the distortion indication map. This map will be used to indicate the distortion of the low-quality rendered image, and the predicted distortion indication map and the low-quality rendered image Figure 1 are sent to the rendering optimization network, and then used to generate a prediction of a high-quality rendered image. For the two sub-networks, the embodiment adopts a U-Net architecture with symmetric skip connections and transpose convolution for design. The U-Net network architecture can be divided into two parts, an encoder (down-sampling) and a decoder (up-sampling). The encoder extracts features and reduces spatial dimensions through convolution and pooling layers, and the decoder restores spatial resolution and refines segmentation through deconvolution and concatenate skip layer connection. The size of the convolution kernel of all convolution layers in the network is 3x3, and ReLU is used as the activation function and placed after each layer except the output layer.
[0060] Step S4: The server divides the low-resolution point cloud to obtain a low-resolution point cloud block, and stores the low-resolution point cloud block and data information related to the low-resolution point cloud block.
[0061] In the embodiment, the data information in step S4 includes the spatial position, the playing order, the uniform resource locator, the encoding format, the resolution and the super-resolution model storage position of the low-resolution point cloud block.
[0062] Step S5: The client tracks the user's viewpoint and predicts the network bandwidth to obtain a prediction result of the user's viewpoint and the network bandwidth.
[0063] Step S6: The client sends request information to the server according to the viewpoint and the prediction result.
[0064] In the embodiment, the request information in step S6 includes obtaining the data information, the stereoscopic video and the trained super-resolution model.
[0065] Step S7: The server transmits the stereoscopic video block and the trained super-resolution model to the client using a video adaptive transmission algorithm based on the video quality perception model according to the request information;
[0066] In this embodiment, the code rate and quality level of the point cloud are controlled by controlling the down-sampling rate of the point cloud, and the conversion relationship between the video code rate and the down-sampling rate is as follows:
[0067]
[0068] wherein represents the size of a single frame file in the video, represents the video playback frame rate, represents the down-sampling rate. As can be seen from the above formula, the greater the down-sampling rate, the smaller the corresponding code rate, and therefore the poorer the visual quality of the point cloud frame. At the same time, since the user can freely change the viewing position and viewing angle when watching the stereoscopic video, the viewing distance of the video, the movement and rotation of the user's viewing angle will all affect the perceived quality of the video. Based on the above analysis, the video quality perception model is determined by the following formula:
[0069]
[0070] wherein, represents the down-sampling rate multiplier, represents the viewpoint movement speed multiplier, represents the viewpoint rotation speed multiplier, represents the user viewing distance, represents the viewpoint movement speed, represents the viewpoint rotation speed, represents the down-sampling rate.
[0071] Further, the video adaptive transmission algorithm in step S7 aims to maximize the user's video perceived quality, and realizes video transmission by making quality level decisions on the video blocks to be transmitted, and the maximization of the user's video perceived quality is determined by the following formula:
[0072]
[0073]
[0074]
[0075]
[0076] wherein, represents the total bandwidth budget, represents the video quality score corresponding to the video block represents the video block Bandwidth consumed Indicates the index of the video block. An index representing the video quality level. Indicates the number of video blocks. Number of video quality levels Indicates whether the selected set of video blocks contains ,like =1 indicates If selected, =0, then means It will not be transmitted.
[0077] Furthermore, a bitrate adaptive algorithm based on a pruned model-controlled prediction architecture is used to obtain the optimal solution for maximizing the user's perceived video quality. The pruning operation includes:
[0078] The bitrate is allocated based on the distance between different video blocks and the same user; the closer the distance, the lower the bitrate allocated.
[0079] If the video bandwidth consumption exceeds the bandwidth constraint when all video blocks are allocated the minimum bitrate, then the minimum bitrate will be allocated.
[0080] If the video bandwidth consumption is less than the bandwidth constraint when all video blocks are allocated the maximum bitrate, then the maximum bitrate is allocated.
[0081] Within a single model-controlled prediction window, the bitrate allocation decision remains unchanged.
[0082] The original model controls the prediction ( The algorithm requires making bitrate decisions for each video block individually, which is time-consuming. This embodiment greatly reduces the time required by pruning operations. The computational cost consumed by the algorithm. After pruning, the algorithm designed in this embodiment... The final computational complexity of the algorithm is ,in This represents the size of the decision space after pruning. The algorithm uses bandwidth constraints. Candidate code rate set The algorithm takes user-perspective features as input and outputs a download bitrate decision for video blocks. The candidate bitrate set is also included. Each element in the code represents a bitrate allocation method, for example... Marked in In the bitrate allocation method represented, this embodiment uses video blocks. , , The assigned quality levels are 3, 2 and 1 respectively. In generating the candidate rate allocation set, the video blocks to be transmitted are first sorted according to the distance of the viewing distance, and the code rate is assigned to each block in accordance with the constraints in the pruning operation. The code rate adaptive algorithm calculates the bandwidth consumption and video quality score for each candidate code rate allocation scheme, and selects the code rate allocation scheme that maximizes the video quality under the bandwidth limit.
[0083] Step S8: The client receives and assembles the stereoscopic video block, obtains a complete stereoscopic video, and processes the complete stereoscopic video using the trained super-resolution model to obtain a high-quality 2D rendering image. The player plays the high-quality 2D rendering image to the user, and the transmission is completed.
[0084] Embodiment two
[0085] As shown in Figures 4-7 The stereoscopic video adaptive transmission method based on a video quality perception model according to the embodiment of the application is compared with two other code rate allocation algorithms, namely Uniform and ViVo. In the Uniform algorithm, the total bandwidth available is used as a budget, and the bandwidth is evenly distributed to each video tile to be transmitted during video transmission. In the ViVo algorithm, a lookup table is established between the user viewing distance and the point cloud downsampling rate before video transmission, and the corresponding downsampling rate is obtained according to the distance of each tile during video transmission.
[0086] To verify the effectiveness of the super-resolution model, the embodiment compares it with two 2D image-based super-resolution algorithms and two point cloud-based super-resolution algorithms. The comparison algorithms are introduced as follows:
[0087] SwinIR: A Transformer with a top-level design that contains a sliding window operation is proposed, which uses a window-based attention module to more efficiently perform image super-resolution tasks;
[0088] Restormer: An encoder-decoder converter is included for multi-scale learning of high-resolution images, and parallelization is used to reduce inference time. The model can process high-resolution images for restoration tasks;
[0089] PU-GCN+: It is an extension of PU-GCN, and the optimization includes patch extraction based on octree, nearest neighbor interpolation on color attributes, and merging of SR input and output.
[0090] MPU+: It is an extension of MPU. Compared with PU-GCN+, MPU+ also uses a spherical kernel function (SKF) to speed up convolution when extracting features.
[0091] For each point cloud super-resolution model, four models with different up-sampling rates (e.g., x2, x5, x10, and x25) are trained respectively. For fair comparison, points outside the user's visible range are removed and will not be fed into the network. For the 2D image-based super-resolution model and RenDA-Net, only one model is trained to perform enhancement operation on distorted point cloud renderings under different down-sampling rates. Figure 4 With Figure 5 Different rate allocation algorithms are demonstrated 1 and 2 bandwidth conditions, and in the Longdress and Soldier two dataset transmission tasks. From the results in the figure, in terms of visual quality, the MPC-based algorithm proposed in this embodiment performs much better than the Uniform and ViVo algorithms under two different bandwidth limitations and on two datasets. By comparing two algorithms, it can be seen that: compared with the Uniform algorithm, the MPC algorithm can provide users with higher visual quality while occupying less bandwidth resources, because the MPC algorithm preferentially allocates bandwidth to tiles that bring higher visual quality gains, making the most of limited bandwidth resources; compared with the ViVo algorithm, the MPC algorithm not only considers the user's viewing distance when allocating code rate, but also takes into account the movement of the user's viewpoint. On the other hand, for the ViVo algorithm, when the initial allocated bandwidth exceeds the current allocatable bandwidth, the code rate level of all tiles will be uniformly reduced, and this rough adjustment results in a low utilization rate of bandwidth. In summary, the MPC-based code rate allocation algorithm in this embodiment can effectively utilize bandwidth resources while greatly improving user visual quality. Compared with the ViVo algorithm, the visual quality is improved by 25%.
[0092] Figure 6 With Figure 7The performances of the five models in the super-resolution task for the Longdress and Soldier datasets are demonstrated, specifically, the video quality of the model inference output is compared by describing the cumulative distribution function (CDF) of the super-resolution video frames on the PSNR and VIF indicators. First, by comparing the point cloud SR method (such as MPU+ and PU-GCN+) with the 2D image SR method (such as SwinIR, Restormer and RenDA-Net), it can be found that the method of enhancing the quality of the stereoscopic video based on 2D images is obviously better than the optimization method based on 3D model, because the current point cloud SR method has weak processing ability for the color attribute of the point cloud, and only uses the nearest neighbor difference method to infer the color attribute of the up-sampled point, so it cannot restore the complex texture. In addition, since the point cloud SR often introduces the deviation of the point position when up-sampling the geometric points, it causes the inconsistency between frames, which also seriously affects the visual quality of the stereoscopic video. Through the comparison among the 2D image SR methods, it can be found that the RenDA-Net designed in this paper is better than other SR methods in all four indicators, and in the stereoscopic video transmission task, it provides the best visual quality for the user. At the same time, by comparing the performances of the algorithms on the Longdress and Soldier datasets, it can be found that for the video with high texture complexity (such as Longdress), the effect gain of RenDA-Net will be more significant, and for the video with lower texture complexity (such as Soldier), the other two image SR algorithms can also perform well.
[0093] In order to better compare the decoding efficiency of different models, four kinds of down-sampling rates (such as 2, 5, 10, 25) are used to compress the Longdress dataset, and then five kinds of SR methods are used to decode it in turn and calculate the decoding time. The results are shown in the following table:
[0094]
[0095] Since different point cloud downsampling rates only affect the resolution of the point cloud itself and do not change the resolution of the corresponding rendered image, we can observe that as the downsampling rate increases (i.e., the bit rate decreases), the inference latency of point cloud SR models, represented by MPU+ and PU-GCN+, shows a significant decrease, while the SR algorithm based on 2D images does not show a significant change. Overall, point cloud SR methods perform poorly in terms of decoding efficiency, achieving only 3 FPS even at low bit rates (e.g., a downsampling rate of 25). For the SwinIR and Restormer algorithms, due to the complexity of their models, high inference latency is also inevitable, especially for the SwinIR model, where single-frame inference latency exceeds 10 seconds, far from meeting the requirements of stereoscopic video transmission systems. In contrast, RenDA-Net, with its simple network structure, exhibits excellent decoding speed, achieving 100-250 FPS on the test set. In summary, compared with point cloud SR methods, RenDA-Net can improve the perceptual quality of stereo video by 78% and increase the inference speed by 108 times; compared with 2D image SR methods, RenDA-Net can improve the perceptual quality of stereo video by 19% and increase the inference speed by at least 230 times.
[0096] Example 3
[0097] like Figure 2 As shown in the figure, an embodiment of the present invention provides a stereoscopic video adaptive transmission system based on a video quality perception model, comprising: a server and a client. The server is used to acquire stereoscopic video to be transmitted, perform point cloud downsampling processing on the stereoscopic video, train a super-resolution model using low-resolution point cloud and stereoscopic video, divide the low-resolution point cloud, store low-resolution point cloud blocks and related data information, and transmit stereoscopic video blocks and the trained super-resolution model to the client according to a request information using a video adaptive transmission algorithm based on a video quality perception model. The client is used to track the user's viewpoint and predict network bandwidth, send request information to the server according to the viewpoint and prediction results, receive and assemble stereoscopic video blocks to obtain a complete stereoscopic video, and process the complete stereoscopic video using the trained super-resolution model to obtain a high-quality 2D rendering image.
[0098] The embodiment transmits stereoscopic video blocks and a trained super-resolution model to a client by using a video adaptive transmission algorithm based on a video quality perception model, wherein the video adaptive transmission algorithm takes bandwidth constraints, candidate code rates and user perspectives as inputs, sorts the video blocks to be transmitted according to the distances of viewing distances when generating the candidate code rates, and assigns code rates to each block in accordance with the constraints in the pruning operation, the algorithm calculates bandwidth consumption and video quality scores for each candidate code rate allocation scheme, and selects the code rate allocation scheme that maximizes the video quality under the condition of meeting the bandwidth limit; the super-resolution model includes two sub-networks, namely a distortion prediction network and a rendering optimization network, the distortion prediction network is used to obtain a distortion indication map according to a low-quality depth map, the distortion indication map is used to indicate the distortion of a low-quality rendering map, and the rendering optimization network is used to obtain a high-quality rendering map according to the distortion indication map and the low-quality rendering map, so as to improve the video transmission quality and improve the user experience.
[0099] In summary, the embodiment of the present application provides a stereoscopic video adaptive transmission method and system based on a video quality perception model, which transmits stereoscopic video blocks and a trained super-resolution model to a client by using a video adaptive transmission algorithm based on a video quality perception model, wherein the video adaptive transmission algorithm takes bandwidth constraints, candidate code rates and user perspectives as inputs, sorts the video blocks to be transmitted according to the distances of viewing distances when generating the candidate code rates, and assigns code rates to each block in accordance with the constraints in the pruning operation, the algorithm calculates bandwidth consumption and video quality scores for each candidate code rate allocation scheme, and selects the code rate allocation scheme that maximizes the video quality under the condition of meeting the bandwidth limit; the super-resolution model includes two sub-networks, namely a distortion prediction network and a rendering optimization network, the distortion prediction network is used to obtain a distortion indication map according to a low-quality depth map, the distortion indication map is used to indicate the distortion of a low-quality rendering map, and the rendering optimization network is used to obtain a high-quality rendering map according to the distortion indication map and the low-quality rendering map, so as to improve the video transmission quality and improve the user experience.
[0100] The above only describes the preferred embodiments of the present application, and it should be noted that those skilled in the art can make several improvements and replacements without departing from the technical principles of the present application, and these improvements and replacements should also be considered as the protection scope of the present application.
Claims
1. A stereoscopic video adaptive transmission method based on a video quality perception model, characterized in that, The application relates to a stereoscopic video transmission method and device. The method comprises the following steps: S1: obtaining a stereoscopic video to be transmitted; S2: a server performs point cloud downsampling processing on the stereoscopic video to obtain a low-resolution point cloud; S3: the server trains an ultra-resolution model by using the low-resolution point cloud and the stereoscopic video to obtain a trained ultra-resolution model; S4: the server divides the low-resolution point cloud to obtain low-resolution point cloud blocks, and stores the low-resolution point cloud blocks and data information related to the low-resolution point cloud blocks; S5: a client tracks a user's viewpoint and predicts network bandwidth to obtain a prediction result of the user's viewpoint and network bandwidth; S6: the client sends request information to the server according to the viewpoint and the prediction result; wherein, represents a down-sampling rate multiplier, represents a viewpoint moving speed multiplier, represents a viewpoint rotating speed multiplier, represents a user viewing distance, represents a viewpoint moving speed, represents a viewpoint rotating speed, represents a down-sampling rate; the video adaptive transmission algorithm aims to maximize the video perceptual quality of the user, and realizes video transmission by making quality level decision on the video blocks to be transmitted, and the maximization of the video perceptual quality of the user is determined by the following formula: wherein, represents a total bandwidth budget, represents a video block corresponding video quality score, represents a video block bandwidth consumed, represents an index of a video block, represents an index of a video quality level, represents a number of video blocks, a number of video quality levels, represents whether a selected video block set contains , if = 1, represents that the selected video block contains , = 0, represents that the selected video block does not contain ; and if the selected video block contains , represents is transmitted, while the selected video block does not contain , represents will not be transmitted; S7: the server transmits stereoscopic video blocks and the trained ultra-resolution model to the client according to the request information by using a video adaptive transmission algorithm based on a video quality perception model, wherein the video quality perception model is determined by the following formula: 2.The stereoscopic video adaptive transmission method based on a video quality perception model of claim 1, wherein, The loss function used in step S3 optimizing the training of the super-resolution model, the loss function is determined by the following formula: wherein, denotes the rendered image, denotes the depth map, denotes the network-generated high-quality rendered image and the real high-quality rendered image the mean squared error between, denotes the network-generated distortion indication map and the actual depth noise map the mean squared error between, denotes a parameter for adjusting the degree of influence of the two-part loss on the overall network. 3.The stereoscopic video adaptive transmission method based on a video quality perception model of claim 1, wherein, S8: the client receives and assembles the stereoscopic video blocks to obtain a complete stereoscopic video, processes the complete stereoscopic video by using the trained ultra-resolution model to obtain a high-quality 2D rendering picture, and plays the high-quality 2D rendering picture to the user, thereby completing the transmission. The ultra-resolution model in step S3 comprises two sub-networks, namely a distortion prediction network and a rendering optimization network. The distortion prediction network is used for obtaining a distortion indication map according to a low-quality depth map, and the distortion indication map is used for indicating the distortion of a low-quality rendering picture.
4. The method of claim 3, wherein the method further comprises: The rendering optimization network is used for obtaining a high-quality rendering picture according to the distortion indication map and the low-quality rendering picture.
5. The method of claim 4, wherein the method further comprises: The low-quality depth map and the low-quality rendering picture are obtained by pre-rendering processing according to the low-resolution point cloud and user viewpoint information.
6. The method of claim 5, wherein the method further comprises: The data information in step S4 comprises the spatial position, playing order, uniform resource locator, encoding format, resolution of the low-resolution point cloud blocks and the storage position of the ultra-resolution model.
7. The method of claim 6, wherein the method further comprises: The request information in step S6 comprises the data information, the stereoscopic video and the trained ultra-resolution model. The optimal solution of the problem of maximizing the video perception quality of the user is obtained by using a code rate adaptive algorithm based on a pruned model control prediction architecture, and the pruning operation comprises the following steps: allocating code rates according to the distances of different video blocks to the same user, and the closer the distance, the smaller the allocated code rate; if the video bandwidth consumption is greater than the bandwidth constraint when all the video blocks are allocated the minimum code rate, then the minimum code rate is allocated, if the video bandwidth consumption is less than the bandwidth constraint when all the video blocks are allocated the maximum code rate, then the maximum code rate is allocated; the code rate allocation decision is unchanged within a single model control prediction window.
8. A stereoscopic video adaptive transmission system based on a video quality perception model, the system based on the method of any one of claims 1-7, comprising a server and a client, characterized in that, The server is used for obtaining a stereoscopic video to be transmitted, performing point cloud downsampling processing on the stereoscopic video, training a super-resolution model using the low-resolution point cloud and the stereoscopic video, dividing the low-resolution point cloud, storing the low-resolution point cloud blocks and data information related to the low-resolution point cloud blocks, transmitting stereoscopic video blocks and the trained super-resolution model to the client using a video adaptive transmission algorithm based on a video quality perception model according to the request information; the client is used for tracking a user viewpoint and predicting network bandwidth, sending request information to the server according to the viewpoint and prediction result, receiving and assembling the stereoscopic video blocks, obtaining a complete stereoscopic video, and processing the complete stereoscopic video using the trained super-resolution model to obtain a high-quality 2D rendering picture.
Citation Information
Patent Citations
Adaptive streaming media code rate selection system based on video perception quality
CN111490983A
Stereoscopic video transmission method based on asymmetric code rate allocation
CN111711810A