A virtual reality-oriented multicast video transmission QoE optimization method

By optimizing video segment selection through edge computing-assisted multicast network caching strategies and a dual sliding window method, the system addresses the differentiated data transmission requirements for user experience quality in virtual reality, achieving higher user QoE and lower latency, thus optimizing user experience quality.

CN119545113BActive Publication Date: 2025-12-26NANJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411710786.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-12-26
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

In virtual reality, 360° video transmission has different data transmission requirements due to different user viewing angles, which puts a lot of pressure on the network and may lead to a deterioration in network service quality and user experience.

Method used

We construct an optimization problem with QoE as the objective, adopt an edge computing-assisted multicast network caching strategy, and optimize video segment selection through a double sliding window method and an improved linear upper bound confidence algorithm. This problem is then transformed into a multi-armed slot machine problem for solution, thereby improving the quality of user experience.

Benefits of technology

It achieves higher user QoE and lower latency, thus optimizing the user experience quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119545113B_ABST
    Figure CN119545113B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video transmission, and discloses a multicast video transmission QoE optimization method for virtual reality, and a multicast network virtual video transmission system model based on edge computing assistance, the method comprising the following steps: constructing an optimization problem with a QoE value as an optimization target and edge computing servers and user caches, transmission time delays and calculation time delays as limiting conditions; determining a new video segment optimization selection method, so that users select to join a group with the most video cache resources through a calculation resource coverage rate, and the optimal cache strategy is obtained through a double sliding window method to realize optimized video segment selection; converting a non-convex and nonlinear QoE optimization problem into a multi-arm tiger machine problem, and applying an improved linear upper bound confidence algorithm to obtain an approximate optimal solution. The application constructs an optimization problem with a QoE value as a target, proposes an edge computing assisted multicast network cache strategy, and realizes the optimization and improvement of user experience quality.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video transmission, in particular to a multicast video transmission QoE optimization method for virtual reality. BACKGROUND

[0002] With the rapid development of the Internet and the Internet of Things technology, new services such as immersive extended reality in the new generation of wireless networks are emerging, and are widely penetrating into personal applications and other applications. The communication and sensing integration technology allocates network resources and computing resources according to services, forms a cross-layer sensing and computing integration design and unified arrangement management, and combines real-time state to reasonably allocate computing tasks, forming a mutual symbiosis architecture of data processing integration, resource allocation integration and service implementation integration, which is the main supporting technology for realizing high-speed, large-capacity and low-latency video transmission of extended reality and virtual reality services.

[0003] The 360° video transmission in virtual reality causes differentiated data transmission requirements due to different user viewing ports, which brings great transmission pressure to the network and may cause the quality of network service and user experience to deteriorate. SUMMARY

[0004] The present application provides a multicast video transmission QoE optimization method for virtual reality, which constructs an optimization problem with QoE value as the target, proposes an edge computing assisted multicast network caching strategy, and realizes the optimization and improvement of user experience quality.

[0005] The present application provides a multicast video transmission QoE optimization method for virtual reality, based on an edge computing assisted multicast network virtual video transmission system model, the multicast network virtual video transmission system model includes a cloud server, a wireless base station containing an edge computing server, and a D2D multicast network composed of mobile devices; the method comprises:

[0006] S1. According to the multicast network virtual video transmission system model, an optimization problem is constructed with QoE value as the optimization target, and edge computing server and user cache, transmission delay and computing delay as the constraint conditions;

[0007] S2. A new video segment optimization selection method is determined to enable users to select to join the group with the most video cache resources through computing resource coverage, and to obtain the optimal cache strategy through a double sliding window method to realize optimized video segment selection;

[0008] S3. The non-convex and non-linear QoE optimization problem is converted into a multi-arm bandit problem, and an improved linear upper bound confidence algorithm is applied to obtain an approximate optimal solution.

[0009] Further, in the multicast network virtual video transmission system model,

[0010] The edge computing server serves the users in the current base station coverage range who are watching the same video by providing video data and assisting user computing; the users will offload the tasks that need to consume a large amount of computing resources and are sensitive to delay to the edge computing server, the users closer to the edge computing server will have higher priority to obtain video or computing resources from the edge computing server, and the users outside the set range can obtain video resources from the intermediate users who are watching the same video or have finished watching the current video, and offload the computing tasks to the intermediate users, indirectly obtain video resources and computing resources through the intermediate users;

[0011] The users in the base station range who are watching the same video are connected with the edge computing server or other users through a D2D multicast network; all users are divided into a plurality of groups of different sizes, the groups are connected with each other through a D2D network, a user in each group acts as a group scheduling center, the user acting as the group scheduling center is in the middle of the other users in the group and the edge computing server, requests the video from the edge computing server, and forwards the video data to the other users in the same group.

[0012] Further, the step S1 comprises:

[0013] The video resolution is defined in the QoE model as:

[0014]

[0015] ω i,j ∈{0,1} represents whether the currently loaded tile is in the user's visual field, when ω i,j =1, it represents that the j th th block of the i th th video is in the user's visual field, otherwise, if the prediction is wrong, it is not in the user's visual field, then ω i,j =0.

[0016] The video segment quality change rate is defined in the QoE model as:

[0017]

[0018] The time for the user to wait for the video to be loaded is defined as:

[0019]

[0020] Wherein, T latency (i)+T finish time(i) -Δt means that the current video, i.e. the i-th video, should be downloaded and calculated in the required resolution within the playing time of the previous video, i.e. the (i-1)-th video, to play the current video, i.e. the i-th video, on time; if the downloading and calculation time is greater than the playing time of the previous video, it will cause the video to be stuck;

[0021] The QoE of a single user is defined as the weighted sum of the above three indicators:

[0022]

[0023] where λ, μ, ν are weight factors, representing the importance of the above three indicators, respectively;

[0024] If the user chooses to obtain the video segment from the edge computing server, the video will be transmitted between groups through the D2D network, and the group leader selects the resolution of the video according to the channel quality, and requests the video from the cloud server or the edge computing server, and then forwards the received video segment to other users in the group and other groups, at this time the formula representing the influence of the video playing delay time of the users in the group is:

[0025]

[0026] The QoE evaluation formula can be obtained:

[0027] QoE = 5.67 x I Q -6.72 x I R + 0.17 - 4.95 x I T - F T

[0028] Finally, the QoE target function is modeled as:

[0029]

[0030] where C1 and C2 limit the total size of the tile to be less than the cache capacity of the MEC server and the cache capacity of the user; C3 and C4 represent the total delay of a single user equipment processing task, including transmission delay and calculation delay and the total delay of transmitting data blocks between users; C5 represents that the encoding rate of the tile should fall within the given set.

[0031] Further, in the QoE model,

[0032] The size of a complete 360° video data can be represented by the following formula:

[0033]

[0034] A video library storing virtual reality videos is created on a cloud server. A 360° video is selected and divided into n consecutive segments. The duration Δt of each segment is fixed. Each video segment is further divided into m spatial tiles, each of which can be independently encoded and sent to the user. Let X = {X1, X2, ..., X...} N} is defined as the set of supported resolution versions, where X k It is k th The encoding bit rate at the resolution level is expressed as v(X). k ) indicates a resolution of X k The size of the tiles;

[0035] By r i,j ∈X represents fragment i th Chinese J th The resolution allocated on the block, i.e., the requested fragment i th The size will be The cache capacity of the edge computing server is D. MEC Its goal is to pre-cache the tiles that should be cached based on FoV predictions; i,j,k This represents the caching decision of the MEC server. If it decides to cache the i-th segment of the video with resolution k... th j th Tiles, then c i,j,k =1, otherwise c i,j,k =0;

[0036] When establishing the wireless connection model between users and base stations, large-scale path loss and shadowing effects are employed; when establishing the wireless connection model between users, free-space path loss and multipath effects are employed; the expected transmission delay for downloading tile i from an edge computing server or group is T. latency When downloading tiles from the edge computing server, the transmission latency is:

[0037]

[0038] When downloading tiles from the group, the transmission latency is:

[0039]

[0040] Among them, P i For transmission power, N is the channel gain, and N is the noise power. People For the number of users.

[0041] Furthermore, in step S3, the QoE objective function of the optimization objective consists of video clip image quality, video quality change rate, and video latency time, modifying the optimization problem in step S1 to obtain the revenue value of the multi-armed slot machine:

[0042] R t = 5.67 · (Q t + D t ) - 6.72 · F T - 0.17 · S t

[0043]

[0044] where Q t represents the video segment picture quality, including the picture quality of two sliding windows; S t is the switching frequency, used to measure the change of video quality when playing different segments of video; represents the coverage of two sliding windows.

[0045] Further, in the step S2, the video segment optimization selection method comprises:

[0046] S21, group selection algorithm: set a time slot with size L, at every Δt time, get the prediction of each video segment in the time slot (Δt, Δt+l) through FoV prediction; sort the adjacent groups: assuming that the number of groups is M, arrange in descending order according to the average transmission rate T latency of the group to the user, if there are multiple groups with the same speed, select the group with the fastest download speed; if there are multiple groups with the same speed, select these groups, compare the existing video tiles in these groups with the prediction results c i, j , k obtained by FoV prediction, and get the coverage C i of each group of video segment tiles; assuming that the current playing time is t, add the tile coverage of multiple video segments in the time period (Δt, Δt+l) of these groups, sort in descending order, and select one group with the highest coverage to join;

[0047] S22, video segment selection method: combine the sliding window and the internal buffer D USER , the current video is divided into n tiles, set two sliding windows with large and small sizes: a sliding window W with size w and a sliding window V with size v; set x as the currently played segment, as the video is played and the user pauses / fast forwards, the sliding window W will be dynamically updated according to the video playing, the starting positions of the two windows W and V always coincide; before transmitting the next part of the video segment, the user will check the tiles in the sliding window V range that constitute the video segment, that is, whether the tiles to be seen by the user in each video segment are all downloaded, use v​s i,j,k to represent the download status of a tile, v s i,j,k = 1, otherwise v s i,j,k = 0.

[0048] Further, the step S3 comprises:

[0049] In the improved linear upper bound confidence algorithm, the initialization and wherein, is a quadratic cumulative sum of the feature vector of the current action, when the action is selected, the feature vector x t is added to for updating the parameter estimation of the linear model; is a cumulative sum of the feature vector of the current action and the reward, when the action is selected, the feature vector x t is multiplied by the corresponding reward r t (i,j) and added to for updating the parameter estimation of the linear model; at the initialization, is a unit matrix, is a zero vector; at each time slot t, the video clip of the group with the highest expected reward is selected for caching to maximize the overall QoE; after each action is performed, the system receives a feedback reward r t (i,j) for updating A k (i,j) and b k (i,j):

[0050]

[0051] wherein, r t (i,j) is the feedback on the effect of the current action, x t is the feature vector containing the context information;

[0052] According to the group selection method and the video clip selection method, the context information is obtained to construct the feature vector x t For each tile, there are a 3D version that has been rendered and a 2D version that has not been rendered, and each selected tile to be played represents an arm, indicating the action of each user at the current step; represents the action selected at time t, k represents the kth tile, i=0 represents obtaining from the edge computing server, i=1 represents obtaining from the group, j=0 represents that the obtained is a 2D version and needs to be rendered, and j=1 represents that the obtained is a 3D version and does not need to be rendered; selecting an arm is equivalent to selecting the behavior of caching a tile from somewhere, and the expected reward of each arm is predicted and maximized by an algorithm, and the confidence upper bound corresponding to each arm That is:

[0053]

[0054] where R t is the reward, is a parameter vector, is the corresponding covariance matrix, and alpha is a parameter controlling exploration, M max is the maximum number of people that the current group can serve, is composed of two parts, which are predicting the expected reward and the uncertainty term; the former part is the estimation of R t by using a linear model parameter vector theta t and x t , and the greater this value is, the more likely this arm is to be selected; the latter is the uncertainty estimation of the action, which quantifies the confidence interval of the feature vector on the action; the action a t with the highest confidence upper bound is selected for execution:

[0055]

[0056] The action a t is executed, the current reward r t (i,j) is calculated, and the corresponding parameter vector and covariance matrix are updated. After each parameter update, the action a with the maximum confidence upper bound t is selected to maximize the cumulative reward R t .

[0057] The application also provides a multicast video transmission QoE optimization device for virtual reality, based on an edge computing assisted multicast network virtual video transmission system model, wherein the multicast network virtual video transmission system model comprises a cloud server, a wireless base station comprising an edge computing server, and a D2D multicast network composed of mobile devices; the device comprises:

[0058] A construction module is configured to construct, according to the multicast network virtual video transmission system model, an optimization problem with a QoE value as an optimization objective and an edge computing server and user cache, transmission delay and calculation delay as limit conditions.

[0059] The selection module is used for determining a new video segment optimization selection method, so that the user selects to join a group with the most video cache resources through a computing resource coverage rate, and an optimal cache strategy is obtained through a double sliding window method, and an optimized video segment selection is realized.

[0060] The computing module is used for converting the non-convex non-linear QoE optimization problem into a multi-arm bandit problem, and applying an improved linear upper bound confidence algorithm to obtain an approximate optimal solution.

[0061] The application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor realizes the steps of the above method when executing the computer program.

[0062] The application further provides a computer readable storage medium, which stores a computer program, and the computer program realizes the steps of the above method when executed by a processor.

[0063] The application has the following beneficial effects:

[0064] In the application, firstly, an edge computing assisted multicast network 360° video transmission system is considered, and an optimization problem is constructed with a QoE value as an optimization target and edge computing servers, user caches, transmission delays and computing delays as limitation conditions. Secondly, a new video segment optimization selection method is proposed, the user selects to join a group with the most video cache resources through a computing resource coverage rate, and an optimal cache strategy is obtained through a double sliding window method to realize optimized video segment selection. Finally, a non-convex non-linear QoE optimization problem is converted into a multi-arm bandit problem, and an improved linear upper bound confidence algorithm is applied to obtain an approximate optimal solution. Compared with the existing benchmark scheme, the linear upper bound confidence algorithm of the multi-arm bandit problem can realize higher user QoE and lower delay, and realizes optimization and improvement of user experience quality. BRIEF DESCRIPTION OF DRAWINGS

[0065] Figure 1 It is a flowchart of the multicast video transmission QoE optimization method for virtual reality of the application.

[0066] Figure 2 It is a structural diagram of the multicast network virtual video transmission system model in the application.

[0067] Figure 3 It is a first selection diagram of the double sliding window video segment selection method in the application.

[0068] Figure 4 It is a second selection diagram of the double sliding window video segment selection method in the application.

[0069] Figure 5 A schematic diagram of the relationship between the average learning reward value and the time slot in the application.

[0070] Figure 6 A schematic diagram of the relationship between the available bandwidth and the average QoE value of the user in the application.

[0071] Figure 7 A schematic diagram of the relationship between the available bandwidth and the cache hit rate in the application.

[0072] Figure 8 A schematic diagram of the relationship between the available bandwidth and the average delay of the user in the application.

[0073] Figure 9 A structural schematic diagram of the multicast video transmission QoE optimization device for virtual reality in the application.

[0074] Figure 10 A schematic diagram of the internal structure of the computer device in the application.

[0075] The implementation of the object of the application, functional features and advantages will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION

[0076] It should be understood that the specific embodiments described herein are only used to explain the application and do not limit the application.

[0077] The 360° video transmission in virtual reality causes differentiated data transmission requirements due to different user viewing viewports, which brings great transmission pressure to the network and may cause the network service quality and user experience quality to deteriorate. Compared with the network service quality model, the user experience quality model comprehensively considers the influencing factors of the service level, the user level and the environment level, is a subjective measurement of the overall quality of the network, and can better reflect the overall feeling and satisfaction of the user experience during video transmission. Therefore, the video transmission evaluation modeling mechanism and optimization based on the user experience quality provide a multicast video transmission QoE optimization method for virtual reality.

[0078] The application discloses a multicast video transmission QoE optimization method for virtual reality, constructs an optimization problem with a QoE value as a target, proposes an edge computing assisted double sliding window video segment selection method, and realizes optimization and improvement of user experience quality. First, considering an edge computing assisted multicast network 360-degree video transmission system, an optimization problem is constructed with a QoE value as an optimization target, and edge computing servers and user caches, transmission delay and computing delay are taken as constraint conditions. Secondly, a new video segment optimization selection method is proposed. Users select a group with the most video cache resources by calculating resource coverage, and the optimal cache strategy is obtained by a double sliding window method to realize optimized video segment selection. Finally, a non-convex non-linear QoE optimization problem is converted into a multi-armed bandit problem, and an improved linear upper bound confidence algorithm is applied to obtain an approximate optimal solution. Compared with the existing benchmark scheme, the linear upper bound confidence algorithm of the multi-armed bandit problem can realize higher user QoE and lower delay.

[0079] Specifically, as shown in the drawings, Figure 1 The application provides a multicast video transmission QoE optimization method for virtual reality, based on an edge computing assisted multicast network virtual video transmission system model, wherein the multicast network virtual video transmission system model comprises a cloud server, a wireless base station comprising an edge computing server, and a D2D multicast network composed of mobile devices.

[0080] The method comprises:

[0081] S1, according to the multicast network virtual video transmission system model, an optimization problem is constructed with a QoE value as an optimization target, and edge computing servers and user caches, transmission delay and computing delay are taken as constraint conditions.

[0082] (1) Multicast network virtual video transmission system model

[0083] As Figure 2As shown, the edge computing assisted multicast network virtual video transmission system model is composed of three parts: cloud server, wireless base station containing edge computing server, and D2D multicast network composed of mobile devices. The edge computing server located at the wireless base station provides video data and assists user computing to serve users watching the same video in the current base station coverage. Users can offload tasks that require a large amount of computing resources and are sensitive to delay to the edge computing server to achieve faster processing. Users closer to the edge computing server (base station) will have higher priority to obtain video or computing resources from the edge computing server, and users farther away can obtain video resources from intermediate users who are watching the same video or have finished watching the current video, and offload computing tasks to these intermediate users to indirectly obtain video resources and computing resources.

[0084] The edge computing server is deployed at the wireless base station, which enables users to obtain video resources more quickly. Users watching the same video within the base station coverage can connect to the edge computing server or other users through the D2D multicast network. Users are different in distance from each other, and in order to enable users farther away from the edge computing server to quickly obtain video resources and obtain as much computing resources as possible, users are divided into several groups of different sizes, which are connected to each other through the D2D network. Each group has one user acting as a group scheduling center, similar to a "group leader", which is in the middle of other users in the group and the edge computing server (base station). They request video from the edge computing server and forward video data to other users in the same group. The edge computing server can share computing tasks with users closer to it; and for users farther away from the edge computing server, video resources can be obtained from other devices in the group, and computing tasks of users in the group can be shared together.

[0085] When a user watches an extended reality video, he / she first sends a request to the edge server. The request is mainly to query whether the edge server caches the video data and to obtain the information of other users who watch the same video. After receiving the information returned by the edge server, the current user can choose to obtain the video data from the edge server (cloud server) or from another user, depending on the location of the user from the edge server. If the user chooses to obtain data from the edge server, there will be two cases: (1) the edge server has cached the video data, and the user can obtain it from the edge server; (2) the edge server does not cache the video data, and the user needs to access the cloud server through the core network to obtain the video data. If the user chooses to obtain data from the group, the user will join the nearest group through the D2D network, obtain the video segments from the group, and can offload the computing task to the edge server or other user devices in the group, reducing the computing load of the user.

[0086] (2) Virtual reality video transmission modeling

[0087] Optimization problem construction: when a user selects a video to start playing, he / she can choose to obtain video data from the edge server or the cloud server, or from other members in the same group, and can flexibly transmit the video. The size of a complete 360° video data can be represented by the following formula:

[0088]

[0089] Suppose that a video library of virtual reality videos is stored on the cloud server, and a 360° video is selected from the video library. The video is divided into n continuous segments, and the time length Δt of each segment is fixed. In addition, each video segment is divided into m spatial tiles, each of which can be independently encoded and sent to the user. Let X = {X1, X2,..., Xn} be the set of supported resolution versions, where Xk is the kth resolution level, and v(Xk) represents the size of the tile with resolution Xk. Finally, rikj represents the resolution allocated on the jth tile in the ith segment, i.e., the size of the requested segment i is N k th k k i,j th th th The cache capacity of the edge server is D, and the goal is to pre-cache the tiles that should be cached according to the FoV prediction. Let cikj represent the cache status of the jth tile in the ith segment, i.e., cikj = 1 if the tile is cached, and cikj = 0 otherwise. MEC i,j,k ​​​​​​​​​​This represents the caching decision of the MEC server, if it decides to cache the i-th segment of the video with resolution k. th j th Tiles, then c i,j,k =1, otherwise c i,j,k =0. This can be seen as a decimal of r. i,j For example, if c i,j,k =1, the user will see k th Video tile j with resolution.

[0090] When establishing a wireless connection model between a user and a base station, large-scale path loss and shadowing effects are employed; while when establishing a wireless connection model between users, free-space path loss and multipath effects can be used. The expected transmission delay for downloading tile i from an edge computing server or group is T. latency When downloading tiles from the edge computing server, the transmission latency is:

[0091]

[0092] When downloading tiles from the group, the transmission latency is:

[0093]

[0094] Among them, P i For transmission power, N is the channel gain, and N is the noise power. People For the number of users.

[0095] Use x i,j,k Let ∈{0,1} to represent the calculation of the tile, when x i,j,k =1 indicates that tiles are processed at the edge computing server; when x i,j,k =0 indicates that tiles are processed within the currently joined group. The computing power of the edge computing server and the group are represented by f. MEC and f c The time required to render all block projections of the i-th video segment can be expressed as:

[0096]

[0097] For virtual reality video users, QoE is mainly affected by three factors: (1) video clip image quality, (2) video clip quality change rate, and (3) video latency. The image quality of a video clip is determined by its resolution, and its impact on QoE follows the logarithmic law. In this paper's QoE model, video resolution is defined as:

[0098]

[0099] The FoV prediction method (ARMA model prediction method based on statistics) widely used at present still has some prediction bias, using ω i,j ∈{0,1} to represent whether the currently loaded tile is in the user's visual field, when ω i,j =1, it means that the j th th block of the i th th video is in the user's visual field, otherwise, if the prediction is wrong, it is not in the user's visual field, then ω i,j =0.

[0100] Another key indicator to measure the user's QoE is the video segment quality change rate, which is calculated by comparing the difference between the current resolution and the previous resolution, and usually taking the absolute value of the result to make it positive. This indicator helps to understand the fluctuation of user experience. The video segment quality change rate is defined as:

[0101]

[0102] Video delay time is taken as the last indicator. If the current video has been played and the next video has not been loaded, that is, the user's play buffer is empty, it will cause the video to be stuck. The time when the user waits for the video to be loaded is defined as:

[0103]

[0104] Where T latency (i)+T finish time (i)-Δt means that under the condition of selecting the required resolution, the download time and calculation time of the current video (i th video) should be less than the playing time of the previous video (i-1 th video), so that this segment of video (i th video) can be played on time. If the download time and calculation time are greater than the playing time of the previous video, it will cause the video to be stuck, affecting the user experience.

[0105] The QoE of a single user is defined as the weighted sum of the above three indicators:

[0106]

[0107] Where λ, μ, ν are weight factors, representing the importance of the three indicators respectively.

[0108] If the user chooses to obtain the video segment from the edge computing server, the video will be transmitted between groups through the D2D network, the group leader selects the resolution of the video according to the channel quality, and requests the video from the cloud server or the edge computing server, and then forwards the received video segment to other users in the group and other groups. In this case, the formula representing the influence of the user video playback delay time in the group is joined:

[0109]

[0110] The QoE evaluation formula can be obtained:

[0111] QoE = 5.67 x I Q -6.72 x I R +0.17-4.95 x I T -F T

[0112] Accordingly, the QoE target function is modeled as:

[0113] Max QoE = 5.67 x I Q -6.72 x I R +0.17-4.95 x I T -F T

[0114]

[0115] Wherein, C1 and C2 limit the total size of the tile cannot be greater than the cache capacity of the MEC server and the cache capacity of the user; C3 and C4 represent the total delay of a single user equipment processing task, including transmission delay and calculation delay and the total delay of transmitting data blocks between users; C5 represents that the encoding rate of the tile falls within the given set.

[0116] S2, determine a new video segment optimization selection method, so that users select to join the group with the most video cache resources through the calculation resource coverage rate, and obtain the optimal cache strategy through the double sliding window method, to realize the optimized video segment selection.

[0117] (1) Convert the optimization problem in step S1 into a multi-armed bandit problem

[0118] A complex video delivery environment requires choosing among multiple caching strategies to maximize group QoE. To solve this problem, the optimization problem in step S1 is transformed into a multi-armed bandit problem. Each caching strategy is considered as a selectable "action", and each time a caching strategy is chosen is equivalent to pulling a different lever in a bandit game to try to get the maximum reward. This analogy helps us simplify the complexity of the original problem into a series of discrete choices. According to the third section, the QoE objective function of the optimization objective is composed of video segment picture quality, video quality change rate, and video delay time. In order to make the upper bound improvement algorithm be able to solve this problem, the optimization problem is modified to obtain the reward value of the multi-armed bandit:

[0119] R t = 5.67 · (Q t + D t ) - 6.72 · F T - 0.17 · S t

[0120]

[0121] wherein Q t represents the video segment picture quality, including the picture quality in the two sliding windows; S t is the switching frequency, used to measure the change of video quality when different segments of video are played; indicates the coverage rate of the two sliding windows.

[0122] (2) Video segment selection method based on double sliding windows

[0123] In the designed system model, when multiple 360° video users request the same video content, because they are watching the same type of video, their viewports are basically similar, and they will produce about the same data transmission content. Through the multicast network, these users can share video content with each other, achieving lower delay and significant bandwidth savings. Therefore, a video segment selection method based on double sliding windows is proposed. This method has two steps, as shown in Figure 3 , 4 .

[0124] The first step is the group selection algorithm: set a time slot of size L, at every Δt time, get the prediction of each video segment in the time slot (Δt, Δt+l) through FoV prediction. Then sort the adjacent groups: assume that the number of groups is M, and according to the average transmission rate T latencyDescending order. If there are multiple groups with the same download speed, select these groups. The prediction result c i,j,k Compare with the existing video tiles in these groups, and get the coverage rate C of each group's video segment tile i . Assuming the current playback time is t, add the tile coverage rates of multiple video segments in the (Δt, Δt+l) time period of these groups , sort in descending order, and select the group with the highest coverage rate to join:

[0125]

[0126] The second step is the video segment selection method: combining sliding windows and internal buffer D USER . The current video is divided into n tiles, and two sliding windows are set: a sliding window W with a size of w and a sliding window V with a size of v (in video segments, w>v). Let x be the currently played segment. As the video is played and the user pauses / fast forwards, the sliding window W will be dynamically updated according to the video playback. The starting positions of the two windows W and V always coincide. Before starting to transmit the next part of the video segment, the user will check the tiles that make up the video segment within the sliding window V, i.e. whether all the tiles that the user's perspective will see within each video segment are completely downloaded. Use v s i,j,k to represent the download status of the tiles. If the download is complete, v s i,j,k =1, otherwise v s i,j,k =0.

[0127] S3, convert the non-convex non-linear QoE optimization problem into a multi-arm bandit problem, and apply an improved linear upper bound confidence algorithm to obtain an approximate optimal solution.

[0128] Use the improved linear upper bound confidence (Linear Upper Confidence Bound, LinUCB) to solve the problem. This algorithm can handle context information such as user features, item features, and cross features, thus finding a better balance between exploration and utilization. A key feature of this method is that it can update parameters in real time, allowing the algorithm to adapt to changing environments.

[0129] In the improved LinUCB, the initialization and where, is the quadratic cumulative sum of the feature vector of the current action, when the action is selected, the feature vector x t is added to to update the parameter estimation of the linear model; is the cumulative sum of the feature vector of the current action and the reward, when the action is selected, the feature vector x t is multiplied by the corresponding reward r t (i,j) and added to to update the parameter estimation of the linear model. At initialization, is the identity matrix, is the zero vector, meaning that each arm starts independent of each other, without prior correlation. As the iteration proceeds and the data collected increases, and are updated to reflect the correlation between features, better estimating the potential value of each decision. At each time slot t, the video segment of the group with the highest expected reward is selected for caching to maximize the overall QoE. After each execution of an action, the system receives a feedback reward r t (i,j), which is used to update A k (i,j) and b k (i,j):

[0130]

[0131] where r t (i,j) is the feedback of the effect of the current action, reflecting the actual performance of the action. Through the feedback of the reward, the algorithm can understand the performance of different actions in different contexts, and gradually learn to select the optimal action. x t is the feature vector, containing the context information.

[0132] By combining the group selection method and the video segment selection method with the LinUCB algorithm, the context features obtained by the group selection method and the video segment selection method are integrated into the feature vector of the LinUCB algorithm, so as to optimize the allocation of cache resources. According to the group selection method and the video segment selection method, the context information is obtained, and the feature vector x t is constructed. For each tile, it is assumed to have two versions, a rendered 3D version and a non-rendered 2D version. In the above model, each selected tile to be played represents an "arm", which represents the action of each user at the current step, i.e. whether the current video segment is obtained from the edge computing server or from the group, and what version is obtained. represents the action selected at time t, k represents the kth tile, i=0 represents obtaining from the edge computing server, i=1 represents obtaining from the group; j=0 represents that the obtained is a 2D version and needs to be rendered, j=1 represents that the obtained is a 3D version and does not need to be rendered. Selecting "arm" is equivalent to selecting the behavior of caching a tile from somewhere, and the expected reward (such as cache hit rate, bandwidth saving or QoE improvement, etc.) of each "arm" is predicted and maximized by an algorithm. The confidence upper bound corresponding to each "arm" is:

[0133]

[0134] where R t is the reward, is the parameter vector, is the corresponding covariance matrix, and a is a parameter that controls exploration, M max is the maximum number of people that the current group can serve. is composed of two parts, which are predicting the expected reward and the uncertainty term. The former part is the estimation of R t using the linear model parameter vector θ t and x t , and the greater this value is, the more likely this arm will be selected; the latter is the uncertainty estimate of the action, which quantifies the confidence interval of the feature vector on the action. The action a t with the highest confidence upper bound is selected for execution:

[0135]

[0136] The action a t is executed, the current reward r t (i,j) is calculated, and the corresponding parameter vector and covariance matrix are updated. After each parameter update, the action a with the maximum confidence upper bound t is selected to maximize the cumulative reward R t . In the video segment selection and group selection problem, the LinUCB algorithm can help determine the optimal caching strategy to improve the overall QoE.

[0137] The examples of the present application are specifically described as follows, the system simulation uses the jupyter notebook of Python, and the edge computing assisted multicast 360° video transmission network uses the parameters in the table to implement.

[0138]

[0139]

[0140] where the number of users randomly distributed in the base station range is 30-60, these users are unevenly distributed in 6 groups according to the video they watch, the computing power of the user end is 10 Mbps and the cache capacity is 2G. The number of videos is 5, each video is divided into 100 video segments, and each video segment is composed of 32 tiles. There are two versions of tiles, 2D and 3D, and 2D tiles have 4 resolutions to choose from, with corresponding bit rates X={8, 16, 24, 32} Mbps. The size of the compiled 3D tile is twice the size of the 2D tile. The popularity of the tile follows the Zipf distribution with a parameter of 0.7. For the edge computing server, the computing power is 500Mps and the cache capacity is 50G. The positioning performance of the proposed algorithm is compared with the following UCB algorithms:

[0141] 1) Combinatorial Upper Confidence Bound (CUCB) algorithm.

[0142] 2) Conversational Upper Confidence Bound (ConsUCB) algorithm.

[0143] And two common alternative algorithms are introduced:

[0144] 1) Least Recently Used (LRU) algorithm: LRU algorithm is a commonly used cache replacement algorithm, which replaces the data that is least recently used to make room for the latest data.

[0145] 2) Least Frequently Used (LFU) algorithm: LFU algorithm is another cache replacement algorithm, which replaces the data with the lowest access frequency to make room for the most frequently accessed data.

[0146] Figure 5 For the average learning reward value under different UCB algorithms, LinUCB algorithm can iterate to better performance faster than CUCB and ConsUCB algorithms, and has the highest average learning reward value, which shows that it learns the optimal strategy the fastest among the three algorithms, that is, LinUCB provides the best decision on the edge computing server cache video selection problem. ConsUCB and CUCB also show learning ability, but their performance is relatively poor.

[0147] Figure 6The Quality of Experience (QoE) score for users watching videos under different bandwidths, with the QoE score ranging from 1 to 5. The LinUCB algorithm is significantly ahead of the LRU algorithm, with the user QoE value of the LinUCB algorithm being nearly twice that of the LRU algorithm at the maximum difference, and basically maintaining a lead of 1.5-2. As for the second-performing CUCB algorithm, the user QoE value of the LinUCB algorithm is about 10% better. With the increase of bandwidth, the QoE scores of all algorithms also increase. Higher bandwidth generally means faster loading time, less buffering and higher video quality, resulting in better user experience. The LinUCB algorithm has the highest score in the entire bandwidth range, indicating that the algorithm is the most effective in cache decision-making, especially in low-bandwidth environments, the performance of the LinUCB algorithm is much better than that of other algorithms, which shows that the group selection and segment selection of the algorithm can better meet the user demand in the D2D network under resource constraints. As traditional cache strategies, LFU and LRU have lower QoE scores when the bandwidth is low, but their performance improves as the bandwidth increases. In the case of high bandwidth, their performance is close to CUCB, indicating that these traditional strategies can still provide reasonable user experience in the case of sufficient resources.

[0148] Figure 7 The Quality of Experience (QoE) score for users watching videos under different bandwidths, with the QoE score ranging from 1 to 5. The LinUCB algorithm is significantly ahead of the LRU algorithm, with the user QoE value of the LinUCB algorithm being nearly twice that of the LRU algorithm at the maximum difference, and basically maintaining a lead of 1.5-2. As for the second-performing CUCB algorithm, the user QoE value of the LinUCB algorithm is about 10% better. With the increase of bandwidth, the QoE scores of all algorithms also increase. Higher bandwidth generally means faster loading time, less buffering and higher video quality, resulting in better user experience. The LinUCB algorithm has the highest score in the entire bandwidth range, indicating that the algorithm is the most effective in cache decision-making, especially in low-bandwidth environments, the performance of the LinUCB algorithm is much better than that of other algorithms, which shows that the group selection and segment selection of the algorithm can better meet the user demand in the D2D network under resource constraints. As traditional cache strategies, LFU and LRU have lower QoE scores when the bandwidth is low, but their performance improves as the bandwidth increases. In the case of high bandwidth, their performance is close to CUCB, indicating that these traditional strategies can still provide reasonable user experience in the case of sufficient resources.

[0149] Figure 8The average latency of users watching videos under different bandwidth conditions. As the network bandwidth improves, video data can be transmitted faster, reducing waiting time and improving user viewing experience. The user average delay of the LinUCB algorithm and the CUCB algorithm is almost the same, but still less than the LRU algorithm. The average delay of the LinUCB algorithm is about 50% of the LRU algorithm. The LRU algorithm and the LFU algorithm show similar performance. They have high latency at low bandwidth but their performance improves significantly as the bandwidth increases. The performance of the CUCB algorithm is better than that of the LRU algorithm and the LFU algorithm, possibly because it is more effective in managing the cache in a way that reduces latency. The LinUCB algorithm, which is based on user context and predicts video tile requests to personalize the cache strategy, can more effectively utilize MEC server resources and D2D communication. It can better handle initial video content requests by effectively predicting which tiles are needed and caching them appropriately, and make better decisions to cache video segments that are more likely to be watched on edge computing servers, thereby reducing the latency of video download transmission. And according to the predicted user FoV, select the best group to join, ensure that the tiles are available within the user group, and minimize the use of remote edge computing servers. Through the above method, users can enjoy smoother video playback, less buffering and higher video quality.

[0150] As shown in Figure 9 The application also provides a virtual reality-oriented multicast video transmission QoE optimization device based on an edge computing-assisted multicast network virtual video transmission system model, which includes a cloud server, a wireless base station containing an edge computing server, and a D2D multicast network composed of mobile devices; the device includes:

[0151] A construction module 1 for constructing an optimization problem with QoE value as optimization target and edge computing server and user cache, transmission delay and computing delay as constraint conditions according to the multicast network virtual video transmission system model;

[0152] A selection module 2 for determining a new video segment optimization selection method to enable users to join groups with the most video cache resources by calculating resource coverage and obtain the optimal cache strategy by a double sliding window method to realize optimized video segment selection;

[0153] A calculation module 3 for converting the non-convex non-linear QoE optimization problem into a multi-armed bandit problem and applying an improved linear upper bound confidence algorithm to obtain an approximate optimal solution.

[0154] In one embodiment, the multicast network virtual video transmission system model includes a cloud server, a wireless base station containing an edge computing server, and a D2D multicast network composed of mobile devices;

[0155] The edge computing server serves the users in the current base station coverage range who are watching the same video by providing video data and assisting user computing; the users will offload the tasks that need to consume a large amount of computing resources and are sensitive to delay to the edge computing server, the users closer to the edge computing server will have higher priority to obtain video or computing resources from the edge computing server, and the users outside the set range can obtain video resources from the intermediate users who are watching the same video or have finished watching the current video, and offload the computing tasks to the intermediate users, indirectly obtain video resources and computing resources through the intermediate users;

[0156] The users in the base station range who are watching the same video are connected to the edge computing server or other users through a D2D multicast network; all users are divided into a plurality of groups of different sizes, the groups are connected to each other through a D2D network, a user in each group acts as a group scheduling center, the user acting as the group scheduling center is in the middle of the other users in the group and the edge computing server, requests the video from the edge computing server, and forwards the video data to the other users in the same group.

[0157] In one embodiment, the construction module 1 includes:

[0158] In the QoE model, the video resolution is defined as:

[0159]

[0160] The video resolution is defined as ω i,j ∈{0,1} represents whether the currently loaded tile is in the user's field of view, when ω i,j =1, it represents that the j th th block of the i th th video is in the user's field of view, otherwise, if the prediction is wrong, it is not in the user's field of view, then ω i,j =0.

[0161] In the QoE model, the video segment quality change rate is defined as:

[0162]

[0163] The time for which the user waits for the video to be loaded is defined as:

[0164]

[0165] Wherein, T latency (i)+T finish time(i) -Δt means that the current video segment, i.e. the i-th video segment, should be downloaded and calculated in the required resolution so that the playing time of the previous video segment, i.e. the (i-1)-th video segment, is less than the playing time of the current video segment, i.e. the i-th video segment, to play the current video segment on time; if the downloading time and the calculation time are greater than the playing time of the previous video segment, the video will be stuck;

[0166] The QoE of a single user is defined as the weighted sum of the above three indicators:

[0167]

[0168] where λ, μ, ν are weight factors, respectively representing the importance of the above three indicators;

[0169] If the user selects to obtain the video segment from the edge computing server, the video will be transmitted between groups through the D2D network, the group leader selects the resolution of the video according to the channel quality, and requests the video from the cloud server or the edge computing server, and then forwards the received video segment to other users in the group and other groups, at this time the formula representing the influence of the video playing delay time of the users in the group is:

[0170]

[0171] The QoE evaluation formula can be obtained:

[0172] QoE = 5.67 × I Q - 6.72 × I R + 0.17 - 4.95 × I T - F T

[0173] Finally, the QoE target function is modeled as:

[0174] Max QoE = 5.67 × I Q - 6.72 × I R + 0.17 - 4.95 × I T - F T

[0175]

[0176] where C1 and C2 limit the total size of the tile to be less than the cache capacity of the MEC server and the cache capacity of the user; C3 and C4 represent the total delay of a single user equipment processing task, including transmission delay and calculation delay and the total delay of transmitting data blocks between users; C5 represents that the coding rate of the tile should fall within the given set.

[0177] In an embodiment, in the QoE model,

[0178] The size of a complete 360° video data can be expressed by the following formula:

[0179]

[0180] The video library of virtual reality videos is stored on the cloud server, one of the 360° videos is selected, the video is divided into n continuous segments, the time length Δt of each segment is fixed, each video segment is divided into m spatial tiles, each spatial tile can be independently encoded and sent to the user; X = {X1, X2,..., Xn} is defined as the supported resolution version set, where Xn is the resolution of the nth segment, and v(Xn) is the size of the tile with resolution Xn. N} is defined as the supported resolution version set, where X k is the encoding bit rate of the k th th resolution level, and v(X k ) represents the size of the tile with resolution X k .

[0181] By expressing r i,j ∈X as the resolution allocated on the j th th block in the i th th segment, the size of the requested i th th segment will be The cache capacity of the edge computing server is D MEC , and the goal is to pre-cache the tiles that should be cached according to the FoV prediction; c i,j,k is expressed as the cache decision of the MEC server, if it is decided to cache the j th th tile with resolution k th in the i i,j,k th segment, then c i,j,k = 1, otherwise c latency = 0.

[0182] When establishing the wireless connection model between the user and the base station, a large-scale path loss and shadow effect are adopted; when establishing the wireless connection model between the users, a free space path loss and multipath effect are adopted; the expected transmission delay of downloading tile i from the edge computing server or group is T latency , when downloading the tile from the edge computing server, the transmission delay is:

[0183]

[0184] When downloading the tile from the group, the transmission delay is:

[0185]

[0186] Where P i is the transmit power, is the channel gain, N is the noise power, and N People is the number of users.

[0187] In one embodiment, in the selection module 2, the QoE objective function of the optimization objective is composed of the picture quality of the video segment, the video quality change rate, and the video delay time. The optimization problem of step S1 is modified to obtain the reward value of the multi-armed bandit:

[0188] R t = 5.67 · (Q t + D t ) - 6.72 · F T - 0.17 · S t

[0189]

[0190]

[0191] wherein Q t represents the picture quality of the video segment, including the picture quality in the two sliding windows; S t is the switching frequency, used to measure the change of the video quality when different segments are played; represents the coverage of the two sliding windows.

[0192] In one embodiment, in the selection module 2, the video segment optimization selection method comprises:

[0193] The group selection unit comprises a group selection algorithm: set a time slot with a size of L, at every Δt time, obtain the prediction of each video segment in the time slot (Δt, Δt+l) through FoV prediction; sort the adjacent groups: assuming that the number of groups is M, arrange the groups in descending order according to the average transmission rate T latency of the video tiles sent by the groups to the user, if there are multiple groups with the same speed, select the group with the fastest download speed; if there are multiple groups with the same speed, select these groups, compare the prediction result c i,j,k obtained by FoV prediction with the existing video tiles in these groups, and obtain the coverage C i of the video segment tiles of each group; assuming that the current playing time is t, add the tile coverages of multiple video segments in the time period (Δt, Δt+l) of these groups, sort them in descending order, and select one group with the highest coverage;

[0194] The video segment selection unit comprises a video segment selection method: combine the sliding window and the internal buffer D USER ​, the current video is divided into n tiles, and two sliding windows of large and small sizes are set: a sliding window W of size w and a sliding window V of size v; let x be the current played segment, as the video is played and the user pauses / fast forwards, the sliding window W is dynamically updated according to the video playing, and the starting positions of the two windows W and V always coincide; before starting to transmit the next part of the video segment, the user will check the tiles in the sliding window V range that make up the video segment, that is, whether the tiles that the user's perspective will see in each video segment are all downloaded, and v s i,j,k represents the download status of the tile, if the download is completed, v s i,j,k = 1, otherwise v s i,j,k = 0.

[0195] In an embodiment, the computing module 3 comprises:

[0196] In the improved linear upper bound confidence algorithm, the initializations and wherein, is a quadratic cumulative sum used to store the feature vector of the current action, when the action is selected, the feature vector x t is added to to update the parameter estimation of the linear model; is a cumulative sum used to store the feature vector of the current action and the reward, when the action is selected, the feature vector x t is multiplied by the corresponding reward r t (i,j) and added to to update the parameter estimation of the linear model; at initialization, is a unit matrix, is a zero vector; at each time slot t, the video segment of the group with the highest expected reward is selected for caching to maximize the overall QoE; after each action is performed, the system receives a feedback reward r t (i,j) for updating A k (i,j) and b k (i,j):

[0197]

[0198] wherein r t (i,j) is the feedback on the effect of the current action, x t is a feature vector containing context information;

[0199] According to the group selection method and the video segment selection method, the context information is obtained to construct the feature vector x tFor each tile, there is a rendered 3D version and an unrendered 2D version. Each tile selected for playback represents an arm, indicating the action of each user at this step. This represents the action selected at time t, where k represents the k-th tile, i=0 indicates it's retrieved from the edge computing server, and i=1 indicates it's retrieved from the group; j=0 indicates the 2D version is retrieved and still needs rendering, j=1 indicates the 3D version is retrieved and doesn't need rendering; selecting an arm is equivalent to choosing the action of caching a certain tile from a certain location, and the expected reward for each arm is predicted and maximized by an algorithm, with each "arm" having a corresponding upper confidence bound. That is:

[0200]

[0201] Among them, R t For profit, It is a parameter vector. This is the corresponding covariance matrix, α is the parameter controlling the exploration, and M... max This is the maximum number of people the current group can serve. It consists of two parts: predicting the expected reward and the uncertainty term; the first part uses the linear model parameter vector θ. t and x t For R t The first is an estimate; the larger this value, the more likely the arm is to be selected. The second is an estimate of the uncertainty of the action, quantifying the confidence interval of the feature vector on the action. The action a with the highest confidence upper bound is selected. t implement:

[0202]

[0203] Perform action a t Calculate the current reward r t (i,j), and update the corresponding parameter vector. Covariance Matrix After each parameter update, select the option with the largest confidence upper bound. action a t Maximize cumulative reward R t .

[0204] Each of the above modules and units is used to perform the respective steps in the above-described QoE optimization method for multicast video transmission in virtual reality. The specific implementation methods are as described in the above-described method embodiments, and will not be repeated here.

[0205] like Figure 3As shown, the present application also provides a computer device, which can be a server, and the internal structure thereof can be as shown in the figure. Figure 3 The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is configured to store all data required by the process of the multicast video transmission QoE optimization method for virtual reality. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the multicast video transmission QoE optimization method for virtual reality.

[0206] Those skilled in the art can understand that, Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied.

[0207] An embodiment of the present application also provides a computer readable storage medium having a computer program stored thereon, and the computer program is executed by the processor to implement any one of the multicast video transmission QoE optimization methods for virtual reality.

[0208] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiment methods can be included. Any reference to memory, storage, databases, or other media in this application and in examples provided herein, unless specifically stated otherwise, can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0209] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, device, article, or method that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, device, article, or method. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, device, article, or method that includes the element.

[0210] The above description is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, based on the content of the present application specification and drawings, is also included in the patent protection scope of the present application.

Claims

1. A method for optimizing QoE of multicast video transmission for virtual reality, characterized in that, The application discloses a multicast network virtual video transmission system model based on edge computing assistance, and the multicast network virtual video transmission system model comprises a cloud server, a wireless base station comprising an edge computing server, and a D2D multicast network composed of mobile devices; the method comprises the following steps: S1, according to the multicast network virtual video transmission system model, constructing an optimization problem with a QoE value as an optimization target and a cache capacity of the edge computing server and a user cache capacity, a transmission time delay and a calculation time delay as constraint conditions; S2, determining a video segment optimization selection method, so that a user selects to join a group with the most video cache resources through a calculation resource coverage rate, and an optimal cache strategy is obtained through a double sliding window method, and the optimized video segment selection is realized; S3, converting the non-convex non-linear QoE optimization problem into a multi-arm bandit problem, and applying an improved linear upper bound confidence algorithm to obtain an approximate optimal solution, and specifically: In the improved linear upper bound confidence algorithm, the initialization and where, is a quadratic cumulative sum of the feature vector of the current action, when the action is selected, the feature vector x t is added to to update the parameter estimation of the linear model; is a cumulative sum of the feature vector of the current action and the reward, when the action is selected, the feature vector x t is multiplied by the corresponding reward r t (i,j) and is added to to update the parameter estimation of the linear model; at the initialization, is a unit matrix, is a zero vector; at each time slot t, the video clip of the group with the highest expected reward is selected for caching to maximize the overall QoE; after each execution of the action, the system receives a feedback reward r t (i,j) for updating and : r t (i,j) = R t , where r t (i,j) is the feedback of the effect of the current action, x t is the feature vector, which contains the context information; According to the group selection method and the video segment selection method, context information is acquired, and a feature vector x is constructed t For each tile, there are a 3D version that has been rendered and a 2D version that has not been rendered, and each selected tile to be played represents an arm, indicating the action of each user at the current step; represents an action selected at time slot t, k represents the kth tile, i = 0 represents acquisition from an edge computing server, i = 1 represents acquisition from a group, j = 0 represents acquisition of a 2D version that still needs to be rendered, and j = 1 represents acquisition of a 3D version that does not need to be rendered; selecting an arm is equivalent to selecting a behavior of acquiring a tile from a certain cache, and an expected reward of each arm is predicted and maximized through an algorithm, and a confidence upper bound corresponding to each arm That is: where R t is the overall quality of experience QoE score at time t, used as a reward for the quantized arm; is the parameter vector transpose, is the inverse matrix of used to quantify the feature uncertainty of the action; a is a parameter that controls the exploration, M max is the maximum number of people that the current group can serve, is composed of two parts, which are the predicted expected reward and the uncertainty term; the former is the linear model parameter vector t combines R t with the predicted maximum QoE score (R t ) max computes the expected QoE reward of the action, the larger this value is, the more likely this arm will be selected; the latter is the uncertainty estimate of the action, which quantifies the confidence interval of the feature vector on the action; the action a t is executed: Perform action a t , compute the current reward r t (i,j), and update the corresponding and After each parameter update, maximize the cumulative reward R by selecting the action a t with the largest upper confidence bound t .

2. The VR-oriented multicast video transmission QoE optimization method of claim 1, wherein, In the multicast network virtual video transmission system model, The edge computing server serves users watching the same video in a current base station coverage range by providing video data and assisting user calculation; A user unloads a task requiring a large amount of calculation resources and being sensitive to a delay to the edge computing server, a user closer to the edge computing server has a higher priority to obtain video or calculation resources from the edge computing server, and a user outside a set range can obtain video resources from an intermediate user watching the same video or having watched the current video, and unloads a calculation task to the intermediate user to indirectly obtain video resources and calculation resources from the intermediate user; Users watching the same video in the base station range are connected with the edge computing server or other users through a D2D multicast network; all the users are divided into a plurality of groups of different sizes, the groups are connected with each other through the D2D network, a user in each group acts as a group scheduling center, the user acting as the group scheduling center is in the middle of other users in the group and the edge computing server, requests a video from the edge computing server, and forwards the video data to other users in the same group.

3. The VR-oriented multicast video transport QoE optimization method of claim 2, wherein, The step S1 comprises: In the QoE model, a video resolution is defined as: wherein n is the number of time segments of the video, m is the number of spatial region segments related to resolution, r i,j is the resolution parameter corresponding to the i,j segment, and ω i,j ∈{0,1} indicates whether the currently loaded tile is in the user's perspective field or not, when ω i,j =1, it indicates that the jth block of the ith video is in the user's perspective, otherwise, if the prediction is wrong, it is not in the user's perspective range, then ω i,j =0. In the QoE model, a video segment quality change rate is defined as: where v(r i,j ) the quality parameter of the ith time segment, jth sub-region, the time period that the user waits for the video to load is defined as: wherein T latency (i) is the delay time of the i-th instance, T finish time (i) is the completion time of the i-th instance, and Δt is a time threshold; T latency (i) + T finish time (i) - Δt means that, under the condition of selecting the required resolution, the download time and the calculation time of the current video, i.e., the i-th video, should be less than the playing time of the previous video, i.e., the (i-1)-th video, so that the current video, i.e., the i-th video, can be played on time; if the download time and the calculation time are greater than the playing time of the previous video, the video will be stuck; QoE of a single user is defined as a weighted sum of the above three indexes: Wherein, a, b and c are weight factors, and represent the importance of the above three indexes respectively; If a user selects to obtain a video segment from the edge computing server, the video is transmitted between groups through the D2D network, a scheduling center of a group selects a resolution of the video according to channel quality, requests the video from the cloud server or the edge computing server, and forwards the received video segment to other users in the group and other groups, at this time, a formula representing an influence of a video playing delay time of a user in the group is as follows: The QoE evaluation formula can be obtained as follows: QoE = 5.67 x I Q - 6.72 x I R + 0.17 - 4.95 x I T - F T Finally, the QoE target function is modeled as follows: Max QoE = 5.67 x I Q -6.72 x I R +0.17 - 4.95 x I T - F T s.t. where k is the resource block index, c i,j,k is the resource consumption coefficient, D MEC is the upper limit of the resource capacity of the edge server, ∈(t, t + l) is for all users in the time interval (t, t + l); v i,j,k s is the user-side resource limit coefficient, D USER is the upper limit of the resource capacity of the user equipment, ∈(t, t + l) is for all users in the time interval (t, t + l); T latency is the fixed delay term, T c (i) is the content-related delay of the user, T max is the upper limit of the total end-to-end delay that the user can tolerate; T latency (i,j) is the specific delay; r i,j is the quality parameter, X is the optional set of quality parameters; C1 and C2 limit the total size of the tile to be no larger than the cache capacity of the edge computing server and the cache capacity of the user; C3 and C4 represent the total time delay of a single user equipment processing a task, including transmission delay and computing delay, as well as the total delay of transmitting data blocks between users; C5 represents that the coding rate of the tile falls within the given set.

4. The VR-oriented multicast video transmission QoE optimization method of claim 3, wherein, In the QoE model, A size of a complete 360° video data can be represented by the following formula: where N is the total number of segments that the video is divided into, D segment (i) is the amount of resources required for the ith video segment, M is the total number of sub-blocks divided within each video segment, D tile (i,j) is the amount of resources required for the ith segment, jth sub-block, L is the number of resource dimensions divided, c i,j,k is the resource consumption coefficient for the ith segment, jth sub-block, kth resource dimension, r j,k is the video quality parameter corresponding to the jth sub-block, kth resource dimension, v(r j,k ) is the quality-resource mapping function; a video library of virtual reality videos is stored on a cloud server, one of the 360° videos is selected, the video is divided into n continuous segments, the time length Δt of each segment is fixed, each video segment is divided into m spatial tiles, each spatial tile can be independently encoded and sent to a user; X = {X1, X2,..., X N} is defined as the set of supported resolution versions, where X k is the kth resolution level encoding bit rate, v(X k ) represents the size of the tile with resolution X k ; By setting r i,j ∈X represents the resolution allocated on the jth tile of the ith segment, i.e. the size of the ith segment requested will be The cache capacity of the edge computing server is D MEC The goal is to predict in advance which tiles should be cached according to the FoV prediction; c i,j,k represents the cache decision of the edge computing server, if it decides to cache the tile on the jth tile of the ith segment of the video with the kth resolution, then c i,j,k = 1, otherwise c i,j,k = 0; In establishing the wireless connection model between the user and the base station, a large-scale path loss and shadow effect are adopted; in establishing the wireless connection model between the user and the user, a free space path loss and multipath effect are adopted; the expected transmission delay of downloading the tile from the edge computing server or group is T latency When downloading the tile from the edge computing server, the transmission delay is: wherein, is the delay of the e-th edge node when serving the ζ-th user, W is the transmission bandwidth; P e is the transmission power of the e-th edge node, is the channel coefficient from the e-th edge node to the ζ-th user; N ζ is the noise power at the location of the ζ-th user, N people is the set of users, denotes that the formula is applied to all ζ-th users belonging to the set; the transmission delay when downloading a tile from a group is: wherein, is the delay between the ζth user and the ηth communication entity, W is the transmission bandwidth, P ζ is the transmission power of the ζth user, is the channel coefficient between the ζth user and the ηth communication entity, N ζ is the noise power at the location of the ζth user, people is the set of users, denotes the restriction of the user side of the formula to all ζth users belonging to the set, N is the set of communication entities.

5. The VR-oriented multicast video transport QoE optimization method of claim 4, wherein, The QoE objective function of the optimization target in the step S2 is composed of video resolution, video segment quality change rate, user waiting time for video loading, and video playing delay time, and the optimization problem of the step S1 is modified to obtain the reward value of the multi-armed bandit: R t = 5.67 · (Q t + D t ) - 6.72 · F T - 0.17 · S t wherein R t is the overall quality of experience QoE score at time t, Q t is the quality contribution term, D t is the resource contribution term, representing the coverage of the two sliding windows, F t is the other negative interference term, S t is the quality fluctuation penalty term, used to quantify the penalty of the quality switching frequency when playing different segments of video; Q t represents the video segment picture quality, including the picture quality in the two sliding windows; V t is the effective resource set, W t is the total resource set, c i,j,k is the resource weight coefficient, v(r j,k ) is the quality mapping function, r j,k is the quality parameter; γ is the fluctuation penalty coefficient.

6. The VR-oriented multicast video transport QoE optimization method of claim 5, wherein, The video segment optimization selection method in the step S2 comprises: S21, group selection algorithm: set a time slot of size l, at every Δt moment, get the prediction of each video segment in time slot (Δt, Δt+l) through FoV prediction; sort adjacent groups: assume the number of groups is M, sort the groups in descending order according to the average transmission delay T of the group to the user latency , if there are multiple groups with the same speed, select the group with the fastest download speed; if there are multiple groups with the same speed, select these groups, compare the prediction results c i,j,k obtained by FoV prediction with the existing video tiles in these groups, and get the coverage rate C of the video segment tiles of each group i ; assume that the current playing time is t, add the tile coverage rates of multiple video segments in the (Δt, Δt+l) time period of these groups , sort in descending order, and select one group with the highest coverage rate to join S22, video segment selection method: combined with sliding window and internal buffer D USER , the current video is divided into n tiles, set a big and a small two sliding windows: a sliding window W of size w and a sliding window V of size v; set x as the current played segment, as the video is played and the user pauses and fast forwards, etc. operation, the sliding window W will be dynamically updated according to the video playing, the starting positions of the two windows W and V always coincide; before starting to transmit the next part of the video segment, the user will check the tiles within the sliding window V range that make up the video segment, that is, whether the tiles that the user's perspective will see in each video segment are all downloaded, use v s i,j,k to represent the download status of the tile corresponding to the i-th segment, j-th sub-block and k-th resource dimension, if it is downloaded, v s i,j,k =1, otherwise v s i,j,k =0.

7. A virtual reality oriented multicast video transmission QoE optimization apparatus, characterized in that, The edge computing auxiliary-based multicast network virtual video transmission system model comprises a cloud server, a wireless base station comprising an edge computing server, and a D2D multicast network composed of mobile devices; the device comprises: A construction module is configured to construct, according to the multicast network virtual video transmission system model, an optimization problem with a QoE value as an optimization target and an edge computing server and user cache, transmission delay, and calculation delay as limit conditions; A selection module is configured to determine a new video segment optimization selection method, so that a user selects to join a group with the most video cache resources through a calculation resource coverage rate and obtains an optimal cache strategy through a double sliding window method to realize optimized video segment selection; A calculation module is configured to convert a non-convex and nonlinear QoE optimization problem into a multi-armed bandit problem and solve the problem by using an improved linear upper bound confidence algorithm to obtain an approximate optimal solution, specifically: In the improved linear upper bound confidence algorithm, the initialization and where, is a quadratic cumulative sum of the feature vector of the current action, when the action is selected, the feature vector x t is added to to update the parameter estimation of the linear model; is a cumulative sum of the feature vector of the current action and the reward, when the action is selected, the feature vector x t is multiplied by the corresponding reward r t (i,j) and added to to update the parameter estimation of the linear model; at the initialization, is a unit matrix, is a zero vector; at each time slot t, the video clip of the group with the highest expected reward is selected for caching to maximize the overall QoE; after each execution of the action, the system receives a feedback reward r t (i,j) for updating and : where r t (i,j) is the feedback of the effect of the current action, x t is the feature vector, which contains the context information; According to the group selection method and the video segment selection method, context information is acquired, and a feature vector x is constructed t For each tile, there are a 3D version that has been rendered and a 2D version that has not been rendered, and each selected tile to be played represents an arm, indicating the action of each user at the current step; represents an action selected at a time slot t, k represents a kth tile, i = 0 represents acquisition from an edge computing server, i = 1 represents acquisition from a group, j = 0 represents acquisition of a 2D version that still needs to be rendered, and j = 1 represents acquisition of a 3D version that does not need to be rendered; selecting an arm is equivalent to selecting a behavior of acquiring a tile from a certain cache, and an expected reward of each arm is predicted and maximized through an algorithm, and a confidence upper bound corresponding to each arm That is: where R t is the overall quality of experience QoE score at time t, used as a reward for the quantized arm; is the parameter vector transpose, is the inverse matrix of the feature uncertainty of action a max is the maximum number of people that the current group can serve, consists of two parts, which are the predicted expected reward and the uncertainty term; the former is the linear model parameter vector θ t combines R t with the predicted maximum QoE score (R t ) max computes the expected QoE return of action a t perform: Perform action a t , compute the current reward r t (i,j), and update the corresponding and After each parameter update, maximize the cumulative reward R t by selecting the action a t with the largest upper confidence bound .

8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 6.

9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video cache updating method for adaptive code rate selection in mobile edge computing

    CN113114756A