An Adaptive Video Stream Transmission System for Edge-Terminal Collaborative Super-Resolution

By using edge-end collaborative super-score technology and deep reinforcement learning algorithms in the video transmission system, video resolution and super-score decisions are made based on network conditions and equipment computing power, the shortcomings of video quality and service experience in the existing technology in low bandwidth environments are solved, and higher quality video transmission is achieved.

CN115633143BActive Publication Date: 2025-05-30TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211292686.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2025-05-30
Estimated Expiration
2042-10-21

AI Technical Summary

Technical Problem

The existing video transmission technology is difficult to ensure the quality and service experience of users watching videos when the network bandwidth is low, and it fails to effectively utilize the computing power of edge and mobile devices to perform video overscore.

Method used

An adaptive video streaming system with edge-end collaborative super-score is proposed. Decision-based video download resolution, super-score target resolution and super-score task ratio of mobile devices are used to optimize user service experience based on network conditions, computing power of edge devices and mobile devices, and playback buffer length.

Benefits of technology

Optimize user service experience in video quality, smoothness and lag, achieve better video services when network bandwidth is low, and effectively utilize the computing power of edge and mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115633143B_ABST
    Figure CN115633143B_ABST
Patent Text Reader

Abstract

The present invention discloses an adaptive video stream transmission system for edge-end collaborative super-resolution. The adaptive video stream transmission system includes a video server, an edge server, and a mobile device. It is characterized in that: the video server communicates with the edge server through a backhaul network, and the edge server communicates with the mobile device through a wireless access network. Among them: the video server is used to perform video preprocessing on the requested original video to obtain the PSNR gain of all video blocks; the edge server performs decision-making scheduling processing based on the PSNR gain of the video blocks to obtain super-resolution video blocks; the mobile device fuses the super-resolution video blocks to generate a super-resolution video partition and conveys it to the video playback buffer. The present invention solves the problem of super-resolution restoration of videos based on the computing power of the edge and the mobile terminal to provide a better service experience for users when the network bandwidth is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to the technical fields of edge computing, video super-resolution, and video transmission, and particularly to an adaptive video stream transmission system for edge-segment coordinated super-resolution. Background Art:

[0002] With the development of multimedia communication and network technologies and the popularization of mobile streaming media services, users have higher and higher requirements for video experience quality. Video picture quality and playback stuttering are key factors affecting user experience. In traditional adaptive bitrate video transmission technologies, the client will select the video with the most appropriate bitrate according to the current network bandwidth to ensure the user service experience. However, when the network bandwidth is low, the video picture quality will inevitably decline, and this method is difficult to ensure the user's video viewing quality and service experience when the network bandwidth is limited. With the enhancement of the computing power of mobile devices and the development of image super-resolution algorithms (SuperResolution, SR), researchers have proposed using image super-resolution neural networks to restore low-resolution videos on mobile devices, combining adaptive bitrate video transmission and super-resolution networks, and using the computing power of mobile devices to alleviate the limitation of network resources on user service experience. In a super-resolution neural network, the input low-resolution image will pass through a convolutional layer for feature extraction after passing through the upsampling layer to reconstruct a high-resolution image. The high-resolution image restored by the neural network model will contain more image details, texture details, etc. than the traditional linear interpolation method. However, the super-resolution neural network has a high demand for computational volume, and the computational volume is closely related to the size of the input and output resolutions. In order to reduce the latency of super-resolution inference on mobile devices, existing methods often deploy a lightweight super-resolution network model on the client, but this method sacrifices the improvement of super-resolution picture quality.

[0003] In the traditional service mode of video transmission provided by the cloud, the mobile terminal obtains and downloads video content from the remote server via the base station. With the rapid expansion of the user scale of video services and the improvement of users' requirements for video quality, the load and bandwidth overhead of the video server are extremely high. This service architecture that completely relies on the cloud is difficult to cope with the explosive growth of video service demands. Therefore, edge networks have emerged. Compared with mobile devices, edge servers have more sufficient computing power. Therefore, deploying a super-resolution network model at the edge can achieve better improvement in picture quality. Some researchers have proposed deploying a super-resolution model on edge devices to use the computing power of edge servers to reduce the impact on user service experience when the backhaul network bandwidth is low.

[0004] In the video transmission scenario at the cloud-edge-end, existing adaptive video transmission methods only consider using the computing power of edge servers and mobile devices separately for video super-resolution, without making full and effective joint use of edge-end computing power. Therefore, there is still much room for improvement in existing video transmission strategies. We allocate the super-resolution task to the edge and mobile devices. According to the characteristic that the computing power consumption of the super-resolution network is related to the input resolution, the video is sliced into different small regions to ensure that the mobile device has sufficient computing power to complete the super-resolution task. And we consider both the backhaul link bandwidth and the wireless access bandwidth for super-resolution decision-making to ensure the user's service experience.

[0005] The adaptive bitrate algorithm is the main tool used by content providers to optimize the user service experience. The adaptive bitrate algorithm runs on the mobile device and dynamically selects the bitrate for each video segment according to the current network state, the size of the playback buffer, or a combination of both. Although the user service experience of the adaptive bitrate algorithm has been significantly improved through deep reinforcement learning algorithms in recent years, it cannot provide high-quality video under limited network conditions. Inspired by the continuous enhancement of mobile device computing power and the latest progress in deep learning, some researchers have proposed using super-resolution to avoid the impact of network bandwidth on the user service experience. To use the super-resolution network on mobile devices, existing methods deploy lightweight super-resolution networks on mobile devices or only apply super-resolution to a few key frames of the video. The biggest limitation of deploying the super-resolution model on mobile devices is the limited computing power of mobile devices, resulting in limited improvement in video image quality.

[0006] By deploying computing and storage resources near mobile devices, edge computing has the potential to further improve the video quality in mobile video streaming and reduce the super-resolution latency. Compared with mobile devices, edge servers have greater computing power. Therefore, existing methods implement large-scale super-resolution networks on edge servers. If the wireless access network bandwidth is sufficient, the mobile device can receive the super-resolved high-definition video from the edge server. However, in practical applications, unstable access network bandwidth may still cause excessive transmission delay, greatly reducing the user service experience. In addition, the computing power of mobile devices is not utilized in edge-based video super-resolution transmission schemes.

[0007] From the above analysis, it can be seen that existing super-resolution enhanced video transmission schemes cannot effectively utilize the computing power of edge and mobile devices, and at the same time consider the backhaul link bandwidth and the wireless access bandwidth. In the actual scenario, both of these bandwidths will affect the user service experience. Although super-resolving the entire video frame on mobile devices will bring huge computing power consumption,

[0008] However, the computing power consumption of the super-resolution network is closely related to the input resolution, and the difficulty of super-resolution varies in different regions of the image. We can use a lightweight super-resolution network on mobile devices to perform super-resolution on regions with low super-resolution difficulty in the video. We need to utilize the computing resources of both the edge server and the mobile device, and consider the network bandwidth and edge computing resources to make decisions on the video transmission strategy. Summary of the Invention:

[0009] Aiming at the problems existing in the prior art, the present invention proposes an adaptive video stream transmission system for edge-side collaborative super-resolution. This system uses a deep reinforcement learning algorithm to solve for the video download resolution, reconstruction target resolution, and the proportion of the super-resolution task of the mobile device based on the network condition, the computing power of the edge device and the mobile device, and the length of the playback buffer, so as to optimize the user service experience in terms of video quality, smoothness, and stuttering. The main technical problems solved by the present invention are as follows:

[0010] 1) An adaptive video transmission method for edge-side collaborative super-resolution is proposed, which performs super-resolution restoration on the video based on the computing power of the edge and the mobile side to provide a better service experience for users when the network bandwidth is low.

[0011] 2) The image is divided into multiple different regions according to the super-resolution difficulty, which are respectively used for super-resolution on the edge and on the mobile side, and the video is super-resolved adaptively based on the video content to obtain a higher video quality improvement with reduced computational complexity.

[0012] 3) The video transmission process is modeled as a Markov decision process, and a deep reinforcement learning algorithm is used to find the optimal video download resolution, super-resolution resolution decision, and load allocation decision to maximize the user service experience.

[0013] The present invention solves its practical problems by adopting the following technical solutions:

[0014] An adaptive video stream transmission system for edge-side collaborative super-resolution, the adaptive video stream transmission system includes a video server, an edge server, and a mobile device; the video server communicates with the edge server through a backhaul network, and the edge server communicates with the mobile device through a wireless access network; wherein:

[0015] The video server is used to perform video preprocessing on the requested original video to obtain the PSNR gain of all video blocks;

[0016] The edge server makes decision scheduling processing based on the PSNR gain of the video blocks to obtain super-resolved video blocks;

[0017] The mobile device fuses the super-resolved video blocks to generate a super-resolved video partition and conveys it to the video playback buffer; wherein:

[0018] The edge server makes decisions on the download resolution and super-resolution target resolution of the video chunks sent by the video server based on the detected network bandwidth and the computing power video buffer data information of the edge device, and then sends a video request to the video server to download the video.

[0019] Further, the process of the video server obtaining the PSNR gain of all video chunks through video preprocessing of the requested original video: The video server encodes and partitions the requested original video content to obtain video chunks of different resolutions;

[0020] The original video is divided into N chunks at the video server side. The label of the video chunk is i, 1 ≤ i ≤ N, and the duration of each chunk is 1 second;

[0021] The video server records the PSNR gain of each video chunk in the video-related information file;

[0022] Each chunk is encoded into a version of the corresponding resolution according to the set of optional resolutions R. The encoded video is cut into H video chunks in the spatial dimension, and the video server records the super-resolution PSNR gain of each video chunk in the video-related information file;

[0023] The video server sends the corresponding video chunks in the video-related information file to the edge server according to the user request;

[0024] After the user requests a video, the video-related information file will be sent to the edge server, and the information therein will be provided to the edge server for scheduling after decision-making.

[0025] Further, the edge server includes a decision-making unit, a scheduler, and a first super-resolution network; the edge server performs decision-making scheduling on video chunks through a deep reinforcement learning method combined with an action branch network; the decision-making scheduling outputs download, super-resolution resolution decision-making, and edge-side load; the process of the edge server outputting the optimal decision-making scheduling for video chunks through a deep reinforcement learning method combined with an action branch network:

[0026] For video segment i, the decision affecting its transmission resolution is determined by the end-to-end delay of the video. The historical network bandwidth information, super-resolution time information, player buffer area information, etc. are used as the input information of the decision-making algorithm: The state composition of video segment i is as follows

[0027] S i ={c b ,c w ,l i ,x e ,x c ,Ω e ,Ω c}

[0028] c b : Backhaul link network bandwidth

[0029] c w : Wireless access network bandwidth

[0030] l i : Video playback buffer length

[0031] x e : The time that the previous video was super-resolved on the edge server

[0032] x c : The time it took to super-resolution videos on mobile devices in the past

[0033] Ω e :Computing power of edge servers

[0034] Ω c :Computing power of mobile devices

[0035] Action space: The decisions that need to be made during video transmission include two aspects: the first is the video transmission resolution decision r i and super-resolution target resolution decision r i ′, and the second is the decision of the proportion of super-resolution video blocks placed on mobile devices α i , so the action space of each video segment i is a i =(r i ,r i ′,α i );

[0036] The decision reward obtains the optimal video image through the following formula:

[0037] γ i =q(a i )-μ|q(a i )-q(a i-1 )|-λτ i

[0038] Among them: action decision a i Contains video download resolution i , super-resolution target resolution r′ i And the mobile device super-resolution ratio α i ;q(α i ) indicates that action decision a is adopted i The video after multi-method evaluation fusion; μ and λ represent the weights of image smoothness and freeze time respectively; τ i Indicates the freeze time of video playback;

[0039] In the Actor network, a branch network is used to output the transmission resolution decision of the video separately. i =(r i ,r′ i ) and super-resolution ratio decision α i ; The state of the environment S i Will be input into the Actor network and the Critic network respectively to obtain the action decision e i ,a i And the state value function V(s i ), after which the decision will continue to act on the environment and receive a reward γ i , a record (s i ,r i ,r′ i ,α i ,γ i ) is stored in the experience pool for gradient update of the A3C network; Θ and ω are the parameters of the Actor network and the Critic network respectively; the objective function of the Actor network learning is

[0040]

[0041] Where: H(Π θ (s i )) is the entropy value of the action, which is used to help the Actor network expand the action space during training;

[0042] Critic network predicts the state value function V(S i ), and its training objective function uses mean square error, that is,

[0043]

[0044] After calculating the updated gradients dΘ and dω through the records in the experience pool and the state value function, the parameters of the Actor network and the Critc network can be updated. The network parameters are updated until the algorithm converges to make a decision to obtain the best QoE.

[0045] Furthermore, the edge server will make decisions on the download resolution and super-resolution target resolution of the video block sent by the video server based on the detected network bandwidth and the computing power video buffer data information of the edge device, and then send a video request to the video server to download the video process:

[0046] The edge server will download the video with resolution r from the video server i The size of each video block is h(i,r). The time required to download from the video server to the edge is obtained by following the steps below.

[0047] When the video server transmits the i-th chunk video block, the decision-making unit will obtain the backhaul network bandwidth c based on bandwidth detection b , the wireless access bandwidth c w and information such as the buffer length of the player of the mobile device to output decisions: the video download resolution r i , the super-resolution target resolution r i ', and the ratio α of the super-resolved video blocks on the mobile device i

[0048] The edge server will download video blocks with a resolution of r i from the video server. According to the ratio α i , where the video blocks of (1 - α i )H will be super-resolved on the first super-resolution network, and the video blocks of α i H will be super-resolved on the second super-resolution network, and the target resolution of the super-resolution is r i ':

[0049] Since the bandwidth of the backhaul link is c b , the size of each video block is h(i, r), and the time taken for all video blocks to be transmitted from the video server to the edge is

[0050]

[0051] After all H video blocks are transmitted from the video server to the edge server, the scheduler will transmit α i H video blocks to the second super-resolution network for super-resolution, and the remaining video blocks will be super-resolved on the first super-resolution network; b e (r i , r' i ) and b c (r i , r' i ), respectively represent the time to super-resolve a video block at the edge and on the mobile device, and can be expressed as

[0052]

[0053] and

[0054]

[0055] where: Ω e and Ω c represent the computing capabilities of the edge server and the mobile device respectively, and represent the super-resolution networks on the edge server and the mobile device respectively to super-resolve a video block from r i to r' iComputing power consumption, t de and t en represent the decoding time and encoding time of a video block.

[0056] Furthermore, the mobile device includes a video playback buffer, a video fusion unit, and a second super-resolution network; the process of the mobile device fusing super-resolution video blocks to generate a super-resolution video partition and delivering it to the video playback buffer is as follows: Among them:

[0057] The mobile device scales the proportion α of the super-resolution video block i The process of sending the video block to the mobile end for super-resolution restoration: and represent the time when the video super-resolved at the edge is completed and transmitted and the time when the video super-resolved at the end is completed and restored, respectively, that is and are respectively

[0058]

[0059] and

[0060]

[0061] Therefore, the end-to-end delay t of transmitting the i-th video i is

[0062]

[0063] Among them: t fusion represents the time used for fusing the video on the mobile end. p i represents the length of the video playback buffer, then the playback stuttering time τ of the user's request for the i-th video can be obtained i is

[0064] τ i = max(t i - p i , 0).

[0065] Beneficial effects:

[0066] 1. An edge-end collaborative super-resolution adaptive video stream transmission technology, which proposes to use the computing power of the edge and mobile devices to collaboratively super-resolve and enhance the video image quality, achieving the goal of maximizing the user service experience.

[0067] 2. An edge-end collaborative super-resolution adaptive video stream transmission technology, which simultaneously considers the bandwidth of the backhaul link and the radio access network, and uses the deep reinforcement learning algorithm A3C to adaptively select the download and super-resolution target resolution of the video and the load of the mobile device, realizing the selection of the optimal video transmission decision under fluctuating network conditions.

[0068] 3. It is proposed to use an action branch network in the A3C algorithm to output the resolution decision and load decision of the video respectively, which solves the problem of slow algorithm convergence speed caused by too large action space in the present technology. Description of the Drawings:

[0069] Figure 1 It is a schematic structural diagram of an edge-cloud collaborative super-resolution adaptive video stream transmission system;

[0070] Figure 2 It is a training flow chart of the deep reinforcement learning algorithm involved in the present invention. Detailed Implementation Manner:

[0071] As Figure 1 shown, the present invention provides an edge-cloud collaborative super-resolution adaptive video stream transmission system. The adaptive video stream transmission system includes a video server, an edge server, and a mobile device; wherein: the video server communicates with the edge server through a backhaul network, and the edge server communicates with the mobile device through a wireless access network;

[0072] The video server is used to perform video preprocessing on the requested original video to obtain the PSNR gain of all video chunks;

[0073] The edge server performs decision-making scheduling processing based on the PSNR gain of the video chunks to obtain super-resolution video chunks;

[0074] The mobile device fuses the super-resolution video chunks to generate a super-resolution video partition and conveys it to the video play buffer. Among them:

[0075] The process of the video server performing video preprocessing on the requested original video to obtain the PSNR gain of all video chunks:

[0076] The video server encodes and partitions the requested original video content to obtain video chunks of different resolutions;

[0077] The original video is divided into N chunks at the video server side. The label of the video chunk is i, 1 ≤ i ≤ N, and the duration of each chunk is 1 second;

[0078] The video server records the PSNR gain of each video chunk in the video-related information file;

[0079] Each chunk is encoded into a corresponding resolution version according to the optional resolution set R. The encoded video is cut into H video chunks in the spatial dimension, and the video server records the super-resolution PSNR gain of each video chunk in the video-related information file.

[0080] The video server sends the corresponding video chunks in the video-related information file to the edge server according to the user request;

[0081] After the user requests a video, the video-related information file will be sent to the edge server, and the information therein will be provided to the scheduling module of the edge server for scheduling after decision-making.

[0082] The edge server uses a deep learning method to perform decision-making scheduling processing on the video chunks according to the PSNR gain of the video chunks to obtain super-resolution video chunks: the edge server includes a decision-making unit, a scheduler, and a first super-resolution network; the edge server performs decision-making scheduling on the video chunks through a deep reinforcement learning method combined with an action branch network; the decision-making scheduling outputs download, super-resolution resolution decision, and edge-side load. The edge server outputs the optimal decision-making scheduling process for the video chunks through a deep reinforcement learning method combined with an action branch network

[0083] For video segment i, the decision affecting its transmission resolution is determined by the end-to-end delay of the video, that is, it includes its transmission time and super-resolution reconstruction time. The historical network bandwidth information, super-resolution time information, player buffer area, etc. are used as input information for the decision-making algorithm. The state composition of specific video segment i is as follows

[0084] S i ={c b ,c w ,l i ,x e ,x c ,Ω e ,Ω c}

[0085] c b : Backhaul link network bandwidth

[0086] c w : Wireless access network bandwidth

[0087] l i : Video play buffer length

[0088] x e : Time taken for past videos to be super-resolved on the edge server

[0089] x c : Time taken for past videos to be super-resolved on the mobile device

[0090] Ω e : Computing power of the edge server

[0091] Ω c : Computing power of the mobile device

[0092] Action Space: The decisions to be made during video transmission involve two aspects. One is the decision on the video transmission resolution \(r\) i and the decision on the super-resolution target resolution \(r'\) i , and the other is the decision on the proportion \(\alpha\) of the super-resolution video blocks placed on the mobile device i . Therefore, the action space for each video segment \(i\) is \(a\) i = \((r\) i , \(r'\) i , \(\alpha\) i )

[0093] Decision Reward: The goal of this application is to maximize the user's QoE. This application uses the user's QoE of watching the video as the decision reward \(\gamma\) i , and its expression is

[0094] \(\gamma\) i = \(q(a\) i ) - \(\mu|q(a\) i ) - \(q(a\) i-1 )| - \(\lambda\tau\) i

[0095] After having the above MDP model, this problem can be solved using the deep reinforcement learning algorithm A3C. Since the action decision \(a\) i includes the video download resolution \(r\) i , the super-resolution target resolution \(r'\) i and the super-resolution proportion \(\alpha\) of the mobile device i , the combination of these decision variables in the discrete space results in an overly large action space, making it difficult for the training process to converge.

[0096] In the Actor network, a branch network is adopted to separately output the decision \(e\) on the video transmission resolution i = \((r\) i , \(r'\) i ) and the decision \(\alpha\) on the super-resolution proportion i . As Figure 2 shown, first, the state \(S\) of the environment i will be input into the Actor network and the Critic network respectively to obtain the action decision \(e\) i , \(a\) i and the state value function \(V(s\) i ). After that, the decision will continue to act on the environment to obtain the reward \(\gamma\) i . A record \((s\) i , \(r\) i , \(r'\) i , \(\alpha\) i , \(\gamma\) i ) will be stored in the experience pool for gradient update of the A3C network. \(\Theta\) and \(\omega\) are the parameters of the Actor network and the Critic network respectively. The objective function for the Actor network to learn is

[0097]

[0098] where: H(Π θ (s i )) is the entropy value of the action, which is used to help the Actor network expand the action space during training. The Critic network predicts the state value function V(S i ), and its training objective function uses the mean square error, that is

[0099]

[0100] After calculating the updated gradients dΘ and dω through the records in the experience pool and the state value function, the parameters of the Actor network and the Critc network can be updated, and the network parameters are updated until the algorithm can converge to make the best QoE decision.

[0101] The edge server will make decisions on the download resolution and super-resolution target resolution of the video chunks sent by the video server according to the detected network bandwidth and the computing power video buffer data information of the edge device, and then send a video request to the video server to download the video.

[0102] The edge server will download video chunks with a resolution of r i from the video server. The size of each video chunk is h(i, r), and the time required to download from the video server to the edge is obtained according to the following steps

[0103] When the video server transmits the i-th chunk video chunk, the decision-making unit will output decisions on the video download resolution r b , the wireless access bandwidth c w and the player buffer length of the mobile device, etc.: the video download resolution r i , the super-resolution target resolution r i ′ and the proportion α i of the super-resolution video chunks of the mobile device.

[0104] The edge server will download video chunks with a resolution of r i from the video server. According to the proportion α i , where (1 - α i )H of the video chunks will be super-resolved on the first super-resolution network, and α i H of the video chunks will be super-resolved on the second super-resolution network, and the super-resolution target resolution is r i ′.

[0105] Since the bandwidth of the backhaul link is c b, the size of each video block is h(i, r), and the time taken for all video blocks to be transferred from the video server to the edge is

[0106]

[0107] After all H video blocks are transferred from the video server to the edge server, the scheduler will transfer α i H video blocks to the second super-resolution network for super-resolution, and the remaining video blocks will be super-resolved in the first super-resolution network; b e (r i , r′ i ) and b c (r i , r′ i ), respectively represent the time for super-resolving a video block at the edge and on the mobile device, and can be expressed as

[0108]

[0109] and

[0110]

[0111] where: Ω e and Ω c represent the computing capabilities of the edge server and the mobile device, respectively, and t de and t en represent the decoding time and encoding time of a video block.

[0112] The mobile device fuses the super-resolved video blocks to generate a super-resolved video partition and delivers it to the video play buffer: The mobile device includes a video play buffer, a video fusion unit, and a second super-resolution network; where:

[0113] The mobile device sends the video blocks with the proportion α i of the super-resolved video blocks to the mobile side for the super-resolution restoration process: and represent the time when the super-resolved video at the edge is completed and transmitted and the time when the super-resolved video at the end is completed and restored, respectively, that is and are respectively

[0114]

[0115] and

[0116]

[0117] Therefore, the end-to-end delay t for transmitting the i-th videoi where

[0118]

[0119] t fusion represents the time taken for the video to be integrated on the mobile device. p i represents the length of the video playback buffer, then the playback stuttering time τ of the user's request for the i-th video can be obtained as i where

[0120] τ i = max(t i - p i , 0).

[0121] The mobile device obtains the optimal video image that maximizes the QoE for the user to watch the 1st to Nth chunks through the following formula:

[0122]

[0123] s.t. r i , r i ′ ∈ R, 0 ≤ α i ≤ 1,

[0124] 0 ≤ p i ≤ B max

[0125] where μ and λ represent the weights of video quality smoothness and stuttering time respectively, and B max is the maximum capacity of the buffer. The constraint is that the download resolution r i and the super-resolution target resolution r i ′ can only be selected from the set R, and the proportion of video chunks for super-resolution on the mobile device is a value between 0 and 1, and the video playback buffer length must be less than the maximum capacity B of the buffer max .

[0126] The present invention is not limited to the embodiments described above. The above description of the specific embodiments is intended to describe and illustrate the technical solutions of the present invention. The above specific embodiments are merely illustrative and not restrictive. Without departing from the spirit of the present invention and the scope protected by the claims, those of ordinary skill in the art can make many specific transformations in various forms under the inspiration of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. An adaptive video stream transmission system for edge-cloud collaborative super-resolution, the adaptive video stream transmission system comprising a video server, an edge server, and a mobile device; Characterized in that: The video server communicates with the edge server through a backhaul network, and the edge server communicates with the mobile device through a wireless access network; wherein: The video server is configured to perform video preprocessing on the requested original video to obtain the PSNR gain of all video blocks; The edge server performs decision-making scheduling processing based on the PSNR gain of the video blocks to obtain super-resolution video blocks; The mobile device fuses the super-resolution video blocks to generate a super-resolution video partition and conveys it to the video play buffer; wherein: The edge server makes decisions on the download resolution and the super-resolution target resolution of the video blocks sent by the video server according to the detected network bandwidth and the computing power video buffer data information of the edge device, and then sends a video request to the video server to download the video.

2. An adaptive video stream transmission system for edge-cloud collaborative super-resolution according to claim 1, Characterized in that: The process by which the video server performs video preprocessing on the requested original video to obtain the PSNR gain of all video blocks: The video server encodes and partitions the requested original video content to obtain video blocks of different resolutions; The original video is divided into N chunks at the video server side, the label of the video chunk is i, 1 ≤ i ≤ N, and the duration of each chunk is 1 second; The video server records the PSNR gain of each video block in a video-related information file; Each chunk is encoded into a version of the corresponding resolution according to the set of optional resolutions R, the encoded video is cut into H video blocks in the spatial dimension, and the video server records the super-resolution PSNR gain of each video block in the video-related information file; The video server sends the corresponding video blocks in the video-related information file to the edge server according to the user request; After the user requests a video, the video-related information file will be sent to the edge server, and the information therein will be provided to the edge server for scheduling after decision-making.

3. An adaptive video stream transmission system for edge-cloud collaborative super-resolution according to claim 1, Characterized in that: The edge server includes a decision-making unit, a scheduler, and a first super-resolution network; The edge server makes decision-making scheduling for the video blocks through a deep reinforcement learning method combined with an action branch network; the decision-making scheduling outputs download, super-resolution resolution decision, and edge-side load; The process by which the edge server outputs the optimal decision-making scheduling for the video blocks through a deep reinforcement learning method combined with an action branch network: For video segment i, the decision affecting its transmission resolution is determined by the end-to-end delay of the video. The historical network bandwidth information, super-resolution time information, player buffer information, etc. are used as input information for the decision-making algorithm: The state of video segment i is composed as follows S i = {c b , c w , l i , x e , x c , Ω e , Ω c} c b : Backhaul link network bandwidth c w : Wireless access network bandwidth l i : Video playback buffer length x e : Time taken for super-resolution of past videos on the edge server x c : Time taken for past videos to be used in mobile device super-resolution Ω e : Computing power of the edge server Ω c : Computing power of a mobile device Action space: The decisions to be made during the video transmission process include two aspects. One is the decision on the transmission resolution r of the video i and the decision on the super-resolution target resolution r i '. The other is the decision on the proportion α of the super-resolution video blocks placed on the mobile device i . Therefore, the action space for each video segment i is a i = (r i , r i ', α i ); The decision reward obtains the optimal video image through the following formula: γ i = q(a i ) - μ|q(a i ) - q(a i-1 )| - λτ i Among them: action decision a i includes the video download resolution r i , the super-resolution target resolution r i ', and the super-resolution ratio α of the mobile device i ; q(a i ) represents the multi-method evaluation and fusion of the video after adopting the action decision a i ; μ and λ respectively represent the weights of video quality smoothness and freezing time; τ i represents the freezing time of video playback; In the Actor network, a branch network is adopted to separately output the transmission resolution decision e of the video i =(r i , r i ′) and the super-resolution ratio decision α i ; The state S of the environment i will be separately input into the Actor network and the Critic network to obtain the action decision e i , a i and the state value function V(s i ). After that, the decision will continue to act on the environment to obtain the reward γ i . A record (s i , r i , r i ′, α i , γ i ) is stored in the experience pool for gradient update of the A3C network; Θ and ω are the parameters of the Actor network and the Critic network respectively; The objective function for the Actor network to learn is where: H(Π θ (s i )) is the entropy value of the action, which is used to help the Actor network expand the action space during training; The Critic network predicts the state value function V(S i ), and its training objective function uses the mean squared error, that is After calculating the updated gradients dΘ and dω through the records and state value functions in the experience pool, the parameters of the Actor network and the Critc network can be updated, and the network parameters are updated until the algorithm can converge to make the best QoE decision.

4. An edge-cloud collaborative super-resolution adaptive video stream transmission system according to claim 1, characterized in that: the edge server makes decisions on the download resolution and super-resolution target resolution of the video chunks sent by the video server according to the detected network bandwidth and the video buffer data information of the edge device, and then sends a video request to the video server to download the video process: The edge server downloads video chunks with a resolution of r from the video server i The size of each video chunk is h(i, r). The time required to download from the video server to the edge is obtained according to the following steps When the video server transmits the i-th chunk video block, the decision-making unit will obtain the backhaul network bandwidth c based on bandwidth detection b , the wireless access bandwidth c w and information such as the player buffer length of the mobile device to output decisions: the video download resolution r i , the super-resolution target resolution r i ′ and the ratio α of the super-resolution video block of the mobile device i The edge server downloads video chunks with a resolution of r i from the video server. According to the ratio α i , where the video chunks of (1 - α i )H are super-resolved on the first super-resolution network, and the video chunks of α i H are super-resolved on the second super-resolution network, and the target resolution for super-resolution is r i ': Since the bandwidth of the backhaul link is c b , the size of each video chunk is h(i,r), and the time taken for all video chunks to be transmitted from the video server to the edge is After all H video chunks of the video server are transmitted to the edge server, the scheduler will transfer α i H video chunks to the second super-resolution network for super-resolution, and the remaining video chunks will be super-resolved in the first super-resolution network; b e (r i ,r i ′) and b c (r i ,r i ′), respectively represent the time for a video block at the edge and on the mobile device for super-resolution, and can be expressed as and where: Ω e and Ω c represent the computing capabilities of the edge server and the mobile device respectively, and φ e (r i , r i ′) and φ c (r i , r i ′) represent respectively The computing power consumption of the super-resolution network on the edge server and mobile device for upscaling a video block from r i to r i , where t de and t en represent the decoding time and encoding time of a video block.

5. An edge-cloud collaborative super-resolution adaptive video stream transmission system according to claim 1, characterized in that: the mobile device includes a video playback buffer, a video fusion unit, and a second super-resolution network; the process of the mobile device fusing the super-resolved video chunks to generate a super-resolved video partition and delivering it to the video playback buffer: where: The ratio α of the super-resolution video block for the mobile device i The process of sending the video block to the mobile device for super-resolution restoration: and respectively represent the time when the super-resolution of the edge super-resolved video is completed and transmitted and the time when the super-resolution and restoration of the end super-resolved video are completed, that is and are respectively and Therefore, the end-to-end delay \(t\) for transmitting the \(i\)th video i is where: t fusion represents the time taken for the video to be integrated on the mobile device. p i represents the length of the video play buffer, and thus the playback stuttering time τ for the user to request the i-th video can be obtained as i is τ i = max(t i - p i , 0).

Citation Information

Patent Citations

  • Video cache updating method for adaptive code rate selection in mobile edge computing

    CN113114756A

  • Ultra-high-definition video transmission system and method based on queue learning and super-resolution

    CN115052182A