User non-inductive resource efficient scheduling method and device for VR application and storage medium
By integrating the cross-attention mechanism and low-rank matrix fine-tuning model that integrates user trajectory and video content features, the problems of insufficient viewport prediction accuracy and insufficient dynamic resource allocation in VR applications are solved, and efficient resource scheduling in complex scenarios is achieved to ensure the quality of user experience.
Patent Information
- Application Number
- CN202510567476.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-08
AI Technical Summary
In existing VR applications, the viewport prediction accuracy is insufficient, the resource allocation dynamics are insufficient, and the lack of dynamic switching of video multi-stream data sources makes it difficult to guarantee business reliability.
Through the cross attention mechanism, a large model with fine-tuning of user trajectory timing characteristics and video content characteristics is built to construct a low-rank matrix fine-tuning large model for viewport prediction, dynamically select multi-stream data sources based on network status, set up viewport and network change trigger mechanisms, and optimize resource scheduling.
Improve the accuracy and generalization capabilities of viewport prediction, ensure the stability and reliability of resource allocation in complex user behavior and dynamic scenarios, optimize user viewing experience, and reduce bandwidth consumption.
Smart Images

Figure CN120281968A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of VR applications, and more specifically, relates to a method, device, and storage medium for efficient scheduling of user-insensitive resources in VR applications. Background Art
[0002] As a new network architecture for multi-modal service scenarios, the intelligent multi-modal network aims to provide unified and efficient bearer capabilities for various service modalities by integrating computing, storage, and network resources. Under this architecture, due to the significant differences in application characteristics, intention requirements, and execution strategies among different modal tasks, how to achieve transparent and efficient execution according to the characteristics of modal tasks has become a key issue in the current technological development. As a typical scenario in the intelligent multi-modal network, VR tasks have been widely applied in fields such as education, medical care, and tourism, and their core business is the real-time transmission of 360° video streams. Compared with traditional video services and other modal tasks, VR tasks exhibit characteristics in multiple aspects such as high bandwidth requirements, low latency requirements, complex dynamic interactions, and redundant transmissions.
[0003] Based on the characteristics of the human eye and the fixed field of view of the head-mounted display (HMD), when users watch 360° videos, they can only focus on 20% of the partial area, which is called the viewport. Existing VR applications for 360° video transmission mostly adopt a viewport adaptive transmission method based on tiles, which is divided into a viewport prediction module and a bitrate allocation module. First, the future viewport position of the user is predicted based on historical information, and then the 360° video is sliced into video chunks in the time domain and further divided into tiles in the spatial domain. Differentiated transmission is performed on the viewport and non-viewport areas based on the viewport prediction result to reduce network bandwidth consumption.
[0004] The existing viewport prediction methods can be classified into two categories from a technical principle perspective: content-independent and content-dependent. Content-independent prediction methods mainly predict the future viewport position through statistical models and deep learning methods based on the user's historical trajectory data. Typical technologies include linear regression models, clustering analysis, and the LSTM model based on time series. The above methods perform well under specific user behavior patterns, but show obvious limitations in long-term prediction windows and scenarios of sudden changes in user behavior. Content-dependent prediction methods improve the prediction accuracy to a certain extent by combining user trajectory features with video content information (such as salient regions or target object trajectories). However, the existing methods have relatively simple fusion strategies for trajectory features and video content features, which limits the robustness and generalization ability of the model in different scenarios. In addition, the existing resource allocation methods only adjust the tile bitrate based on the viewport prediction result and network status, lacking dynamic switching of multi-stream data sources for videos, resulting in insufficient dynamicity of resource allocation and difficulty in guaranteeing service reliability.
[0005] Therefore, the solution of the prior art has problems of insufficient viewport prediction accuracy, and only adjusting the tile bitrate according to the viewport prediction result and network status, lacking dynamic switching of video multi-stream data sources, resulting in insufficient dynamicity of resource allocation. Summary of the Invention
[0006] Aiming at the defects of the related art, the purpose of the present invention is to provide a method, device and storage medium for efficient resource scheduling without user perception in VR applications, aiming to solve the problems of insufficient viewport prediction accuracy, adjusting the tile bitrate according to the viewport prediction result and network status, lacking dynamic switching of video multi-stream data sources, resulting in insufficient dynamicity of resource allocation.
[0007] To achieve the above object, in the first aspect, the present invention provides a method for efficient resource scheduling without user perception in VR applications, including:
[0008] S100. Extracting trajectory time series features and video content features from the user trajectory data and image content data of the VR application raw data respectively;
[0009] S200. Calculating the attention weights of the trajectory time series features and video content features respectively through a cross-attention mechanism, adaptively and dynamically adjusting the attention weights of the trajectory time series features and video content features in different scenarios, and generating fusion feature information;
[0010] S300. Constructing a viewport prediction model by using a low-rank matrix to fine-tune a pre-trained large model; inputting the fusion feature information into the viewport prediction model, and outputting a viewport position prediction result within a future time window;
[0011] S400. Dynamically selecting a multi-stream data source for resource scheduling by using the viewport position prediction result and the current network status as resource scheduling conditions; for the high-priority area that the user is concerned about in the viewport position prediction result, allocating a first bitstream channel and network resources meeting the set QOS standard, and using a second bitstream channel for the low-priority area; when the user's viewport changes and the network condition deteriorates, pulling a downscaled version of the pre-cached high-priority area from the local CDN.
[0012] Optionally, it further includes:
[0013] Setting a viewport change trigger mechanism and a network change trigger mechanism;
[0014] When the offset of the viewport area corresponding to the viewport position prediction result within an adjacent time window is greater than a preset value, triggering the viewport change trigger mechanism to reposition the high-priority area of the viewport area;
[0015] If any indicator of the current network status drops to the first judgment condition, the viewport change trigger mechanism is triggered, the overall regional bit rate setting is dynamically adjusted according to the current network status, and the data source is switched; wherein the first judgment condition threshold includes: bandwidth B t The bandwidth is less than or equal to the preset bandwidth or the jitter time is greater than or equal to the preset jitter time or the delay time is greater than or equal to the preset delay time.
[0016] Optionally, after the viewport change trigger mechanism is triggered, a key data block of the new viewport area is requested from a local CDN node. If the local CDN node does not have the key data block cached, the key data block is pulled from an edge node, and a background pre-cache is triggered. Meanwhile, the bitrate of the low-priority area is reduced to a preset ratio of the baseline value.
[0017] After the network change trigger mechanism is triggered, if all indicators of the current network status meet the second judgment condition, the original high-quality data stream is pulled from the central cloud server and the edge node; wherein the second judgment condition includes: bandwidth B t Greater than the preset bandwidth or the jitter time is less than the preset jitter time or the delay time is less than the preset delay time;
[0018] If any indicator of the current network status drops to a first threshold, the pre-cached reduced-bitrate version is pulled from the local CDN. If it misses, the real-time transcoded low-quality data is pulled from the edge server.
[0019] Optionally, step S100 specifically includes:
[0020] S101, collecting trajectory data of the user when watching the video through sensors, obtaining image content data by extracting frames from the video, and obtaining VR application raw data; performing data cleaning on the VR application raw data, eliminating abnormal values in the data collection process, and using linear interpolation to fill in the missing parts;
[0021] S102, using the LSTM model to extract the trajectory time series feature code P′=LSTM(P)∈R from the user trajectory data of the VR application raw data d , where P represents the input trajectory time series data, R d Indicates that the time series feature result is a d-dimensional vector, where d is a positive integer;
[0022] S103, using the VIT model to extract content features V′=VIT(image)∈R from the image content data of the VR application raw data D , where image represents the input image content data, R D It means that the corresponding image content feature is a D-dimensional vector, where D is a positive integer.
[0023] Optionally, step S200 specifically includes:
[0024] S201. Use the trajectory time series feature as the query matrix Q p , and use the video content feature as the key matrix K v and V v . Calculate the correlation between the two through the cross-attention mechanism, and perform weighted summation on the video content feature V′ through the attention score weight matrix to generate the preliminary fusion feature Z′;
[0025]
[0026] Z′ = Attention(Q p , K v , V v ) * V′;
[0027] S202. For different mode scenarios, further integrate the trajectory time series feature P′ and the preliminary fusion feature Z′ through adaptive weighting, and dynamically adjust the influence weights α and β of the two features; in the scenario where the user behavior pattern is stable, adaptively increase the weight β of the trajectory time series feature; in the scenario where there is a region in the video frame that significantly attracts the user's attention, adaptively increase the weight α of the video content feature; to enhance the representativeness of the fusion feature and form the final fusion feature matrix Z t :
[0028] Z t = α * Z′ + β * P′
[0029] α + β = 1.
[0030] Optionally, step S300 specifically includes:
[0031] During the training process of constructing the viewport prediction model,
[0032] S301. The training process of the viewport prediction model includes: adding a trainable low-rank matrix, using the low-rank matrix to freeze the weights of the pre-trained large model, and adjusting the large model structure to adapt to the viewport prediction task;
[0033] S302. In the inference stage of the viewport prediction model, design a prompt template to describe the viewport prediction task type, embed relevant data information, and uniformly guide the large model to generate a prediction result; gradually analyze the fusion feature information in the input prompt through a multi-layer attention mechanism; in each iteration, predict the viewport position at the future time step, and gradually generate a complete prediction in the form of a time series, and post-process to obtain the region number to be transmitted first and the corresponding time step to form the final viewport position prediction result F t .
[0034] Optionally, step S400 specifically includes:
[0035] S401. Obtain and store the network quality data within a preset length of time interval to form the current network status information matrix N t =[B t , J t , L t , where B t represents bandwidth, J t represents jitter, and L t represents latency;
[0036] S402. Dynamically integrate the viewport position prediction result F t and the current network status N t to form the decision information input D for the dynamic execution policy of network resources by analyzing the changes in the viewport position prediction result and the current network status information t =[N t , F t ;
[0037] S403. Expand the video corresponding to the viewport position prediction result into M*N tiles according to the equal rectangle mapping ERP, and calculate the priority value P of each tile;
[0038]
[0039] Among them, each tile is marked as T ij , and the priority value P is determined by the normalized distance d from the viewport prediction center; d is the normalized distance from the tile center to the viewport prediction center, with a range of [0,1]; R is the normalized radius of the viewport area, which is fixed at a preset size according to the field of view size;
[0040] S404. Divide the tiles with a priority value greater than or equal to 0.7 into high priority, and divide the tiles with a priority value less than 0.7 into low priority; divide the bitstream channels according to the priority value P of the tiles, allocate the first bitstream channel and network resources meeting the set QOS standard for the high-priority tiles, and allocate the second bitstream channel for the low-priority tiles;
[0041] S405. Dynamically adjust the resource allocation policy according to the decision information input D t . Among them, the resource allocation policy includes: pulling high-quality tile data from the central cloud, caching the top preset proportion of tile data with the most access times in the past preset time at the edge node, and caching the high-priority tiles predicted in real time in the local CDN.
[0042] Optionally, it further includes:
[0043] Adopting a reinforcement learning algorithm to optimize the resource scheduling policy;
[0044] After the transmission of the multi-stream data source is completed, the multi-stream data source is synchronously integrated and hierarchically rendered on the client side, while dynamically adapting to modal resources.
[0045] In a second aspect, the present invention also provides a device for efficiently scheduling user-insensitive resources in a VR application, including:
[0046] A feature extraction module, configured to extract trajectory time-series features and video content features from the user trajectory data and image content data of the VR application raw data respectively;
[0047] A feature fusion module, configured to calculate the attention weights of the trajectory time-series features and video content features respectively through a cross-attention mechanism, adaptively and dynamically adjust the attention weights of the trajectory time-series features and video content features in different scenarios, and generate fused feature information;
[0048] A viewport prediction module, configured to construct a viewport prediction model by fine-tuning a pre-trained large model with a low-rank matrix; input the fused feature information into the viewport prediction model, and output a viewport position prediction result within a future time window;
[0049] A resource scheduling module, configured to use the viewport position prediction result and the current network state as resource scheduling conditions to dynamically select a multi-stream data source for resource scheduling; for the high-priority area that the user is concerned about in the viewport position prediction result, allocate a first bitstream channel and network resources that meet the set QOS standard, and use a second bitstream channel for the low-priority area; when the user's viewport changes and the network condition deteriorates, pull the downscaled version of the pre-cached high-priority area from the local CDN.
[0050] In a third aspect, the present invention also provides a computer-readable storage medium, where the computer-readable storage medium includes a stored computer program, and when the computer program is run by a processor, it controls the device where the storage medium is located to execute the method provided in any item of the first aspect.
[0051] Through the above technical solutions conceived by the present invention, compared with the prior art, the following beneficial effects can be achieved:
[0052] 1. The present invention provides a method for efficient resource scheduling without user perception in VR applications, which extracts features from trajectory time-series data and video image data, adjusts feature weights according to different scenario modes, dynamically fuses user trajectory features and video content features, and further captures complex dependencies between features by using a fine-tuned large model to improve the accuracy and generalization ability of viewport prediction; based on the viewport prediction results, a resource dynamic scheduling strategy is further designed to dynamically select multi-stream data sources under different conditions. In complex user behavior patterns and dynamic scenarios, this solution exhibits excellent prediction stability and reliability. Especially in long-term prediction tasks, it can effectively cope with the dynamic changes of user behavior and scenario switching, ensure the continuous stability of prediction performance, and guarantee the transmission of high-priority area content with high resolution through efficient resource dynamic scheduling without user perception in VR applications, ensuring that the viewing effect of users in the core area is not affected.
[0053] 2. The present invention provides a method for efficient resource scheduling without user perception in VR applications, which sets a viewport change trigger mechanism and a network change trigger mechanism. Under the condition of sufficient bandwidth, the system is set to pull the original high-quality data stream from the central cloud server and edge nodes, and appropriately improve the transmission quality of the background area to further optimize the user viewing experience; under the condition of limited bandwidth, pull the pre-cached version with reduced bit rate from the local CDN to ensure the content transmission of the user's concerned area and ensure that the viewing effect of the core area is not affected.
[0054] 3. The present invention provides a method for efficient resource scheduling without user perception in VR applications, which dynamically adjusts the priority and transmission quality of 360° video chunks based on an optimization algorithm. The content of high-priority areas is preferentially guaranteed to be transmitted with high resolution, while the content of low-priority areas is processed with reduced bit rate or delayed transmission according to the network status, further ensuring the quality of experience of users in different situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a schematic flow chart of a method for efficient resource scheduling without user perception in VR applications provided by an embodiment of the present invention;
[0056] Figure 2 is an application schematic diagram of a method for efficient resource scheduling without user perception in VR applications provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0058] The following describes the content involved in the above embodiments in conjunction with a preferred embodiment.
[0059] Embodiment 1
[0060] As Figure 1 shown, a method for efficient scheduling of user-insensitive resources in a VR application includes:
[0061] S100. Extracting trajectory time-series features and video content features from the user trajectory data and image content data of the VR application raw data respectively;
[0062] S200. Calculating the attention weights of the trajectory time-series features and the video content features respectively through a cross-attention mechanism, adaptively and dynamically adjusting the attention weights of the trajectory time-series features and the video content features in different scenarios, and generating fused feature information;
[0063] S300. Constructing a viewport prediction model by using a low-rank matrix to fine-tune a pre-trained large model; inputting the fused feature information into the viewport prediction model, and outputting a viewport position prediction result within a future time window;
[0064] S400. Dynamically selecting multi-stream data sources for resource scheduling by using the viewport position prediction result and the current network state as resource scheduling conditions; for the high-priority areas concerned by the user in the viewport position prediction result, allocating a first bitstream channel and network resources meeting the set QoS standard, using a second bitstream channel for the low-priority areas, and dynamically selecting multi-stream data sources according to different resource scheduling conditions; when the user's viewport changes and the network condition deteriorates, pulling a downscaled version of the pre-cached high-priority area from the local CDN.
[0065] In the context of existing VR applications, the resource allocation method for viewport prediction is relatively simple in terms of feature fusion and fails to fully explore the deep dependencies between user trajectories and video content features. This limitation makes it difficult for the model to accurately model complex behavior patterns or content scene changes, significantly affecting the accuracy and robustness of prediction results. Secondly, in long-term prediction tasks, the existing technology has insufficient modeling ability for dynamic user behavior changes. Especially when user behavior suddenly changes or the scene switches, the prediction accuracy of the model drops significantly, making it difficult to stably capture the changing rules of the user's perspective. In addition, the existing methods are highly dependent on the distribution of training data, resulting in weak generalization ability of the model and difficulty in adapting to diverse actual application scenarios. The technical solution of this application proposes a systematic method for efficient resource scheduling in VR applications. By deeply exploring the multi-modal correlation between user behavior trajectories and video content features, a dynamic fusion and adaptive prediction model is constructed to achieve high-precision prediction of viewport positions and optimized resource scheduling in complex scenarios. The technical solution first extracts trajectory time-series features from time-series data such as the user's head direction and line-of-sight trajectory using an LSTM model, and at the same time uses a ViT model to analyze the video content features of video frames to form a two-channel feature input. On this basis, the two types of features are dynamically fused through a cross-attention mechanism and an adaptive weighting module, and the feature weights are adjusted in real time according to the scene mode (such as user behavior stability, content salience) to generate a more robust fused feature representation. Further, based on the low-rank matrix fine-tuning technology, a pre-trained large model is adapted to construct a viewport prediction model. The fused features are analyzed through a hierarchical attention mechanism to output the predicted results of the viewport positions within the future time window, significantly improving the stability of long-term prediction and the response ability in mutation scenarios. Finally, combined with the real-time network status (bandwidth, delay, jitter) and the viewport prediction results, multi-source data streams (central cloud, edge nodes, local CDN) are dynamically scheduled to preferentially ensure high-quality transmission in high-priority areas and seamlessly switch to a lower bitrate version during network fluctuations to achieve efficient resource allocation without user awareness.
[0066] This solution solves the key technical problems in current viewport prediction and resource dynamic scheduling strategies with a highly systematic process, optimizes the resource dynamic scheduling strategy under different user behavior patterns and network states, not only significantly reduces the bandwidth consumption of VR application 360° video transmission, but also enables users to obtain a good viewing experience without awareness, has good generalization ability and practical application prospects, and provides a new technical solution for the transparent and efficient execution of VR applications in multi-modal networks.
[0067] Optionally, step S100 specifically includes:
[0068] S101. Collect the trajectory data of the user when watching the video through sensors, obtain the image content data by extracting frames from the video, and get the original VR application data; perform data cleaning on the original VR application data, remove the outliers in the data collection process, and use the linear interpolation method to complete the missing part;
[0069] S102. Use the LSTM model to extract the trajectory time series feature encoding P′ = LSTM(P) ∈ R d , where P represents the input trajectory time series data, and R d represents that the time series feature result is a d-dimensional vector, and d is a positive integer;
[0070] S103. Use the VIT model to extract the content feature V′ = VIT(image) ∈ R D , where image represents the input image content data, and R D represents that the corresponding image content feature is a D-dimensional vector, and D is a positive integer.
[0071] Obtain the user trajectory data and perform preprocessing. The user trajectory data is collected through the sensors of the Head-Mounted Display (HMD), and the coordinate sequences of the user's head direction H and line-of-sight direction E are recorded and transmitted. These data are standardized and converted into a time series P at a fixed time interval:
[0072] P = {H, E}, H = {h1, h2,..., h t}, E = {e1, e2,..., e t}
[0073]
[0074] e i = {(x i , y i )|0 ≤ x i ≤ 1, 0 ≤ y i ≤ 1}
[0075] where h i is the historical data of the head direction at the i-th moment, and φi, respectively represent the yaw angle (Yaw) and pitch angle (Pitch) of the user's head at the i-th moment; e i is the historical data of the line-of-sight direction at the i-th moment, and (x i , y i ) represents the two-dimensional coordinate position of the user's line-of-sight direction in the projection image of the video frame plane.
[0076] To ensure the continuity and integrity of the data, the system eliminates outliers during the acquisition process and uses linear interpolation to complete the missing parts. For example, 60 groups of direction data sampled per second can be uniformly reduced to 30 groups of standardized time series per second to ensure data consistency and processing efficiency.
[0077] By encoding the temporal sequence features of the trajectory, the finally generated trajectory embedding can accurately reflect the user's historical trajectory pattern, providing strong support for subsequent feature fusion.
[0078] Obtain video frames and extract image content features, specifically including: expanding the video frames of spherical projection into two-dimensional planar images image through the Equirectangular Projection (ERP) method. Each frame of the image is divided into M*N tiles, and a pre-trained Vision Transformer (ViT) model is used to extract features from each frame of the image to form image content features V'. In this embodiment, when using the VIT-Base model, D = 768; the image content features are used to describe the semantic information of the image, covering the priority of the areas that the user may be interested in.
[0079] Optionally, step S200 specifically includes:
[0080] S201. Take the trajectory temporal sequence features as the query matrix Q p , and the video content features as the key-value matrix K v and V v . Calculate the correlation between the two through the cross-attention mechanism, and perform weighted summation on the video content features V' through the attention score weight matrix to generate the preliminary fusion feature Z';
[0081]
[0082] Z' = Attention(Q p , K v , V v ) * V';
[0083] S202. For different mode scenarios, further integrate the trajectory temporal sequence features P' and the preliminary fusion feature Z' through adaptive weighting, and dynamically adjust the influence weights α and β of the two features; in the scenario where the user's behavior pattern is stable, adaptively increase the weight β of the trajectory temporal sequence features; in the scenario where there are areas in the video frame that significantly attract the user's attention, adaptively increase the weight α of the video content features; to enhance the representativeness of the fusion features and form the final fusion feature matrix Z t :
[0084] Z t = α * Z' + β * P'
[0085] α + β = 1.
[0086] Calculate the attention weight between the trajectory timing feature and the video content feature. The attention score weight matrix performs a weighted sum on the video content feature V′ to generate a preliminary fusion feature Z′. Further, in combination with a specific VR application scenario, the feature weights are changed to dynamically fuse the features.
[0087] In a scenario where the user behavior pattern is stable (such as staring at a static blackboard for a long time in a VR education scenario), the system enhances the trajectory timing feature weight (β ≥ 0.8), and uses the long-term modeling ability of the LSTM model for the head movement trend to effectively suppress the interference of the local dynamics of the video content (such as the gradual display of the blackboard writing) on the prediction, ensuring the continuity of the viewport prediction; while in a scenario where the content saliency is prominent (such as the appearance of a dynamic target in a VR game), by strengthening the video content feature weight (α ≥ 0.8), combined with the high-weight semantic regions (Aij ≥ 0.8) extracted by the ViT model, quickly capture the sudden change of the user's line-of-sight focus and correct the lag of the trajectory prediction. For example, when the user's head rotates rapidly (yaw angular velocity ≥ 30° / s) and the line of sight focuses on a dynamic target, the dynamic weighting mechanism balances the trajectory mutation trend and the content semantic priority, reducing the prediction error rate by 35%.
[0088] Optionally, step S300 specifically includes:
[0089] S301. The training process of the viewport prediction model includes: adding a trainable low-rank matrix, using the low-rank matrix to freeze the weights of the pre-trained large model, and adjusting the large model structure to adapt to the viewport prediction task;
[0090] S302. In the inference stage of the viewport prediction model, design a prompt template to describe the viewport prediction task type, embed relevant data information, and uniformly guide the large model to generate a prediction result; gradually parse the fusion feature information in the input prompt through a multi-layer attention mechanism; in each iteration, predict the viewport position at the future time step, and gradually generate a complete prediction in the form of a time series, and post-process to obtain the region number and corresponding time step to be preferentially transmitted, forming the final viewport position prediction result F t .
[0091] Fine-tune the pre-trained large model through a low-rank matrix and construct a viewport prediction model. Adding a small number of trainable low-rank matrices can reduce the computational overhead and storage requirements of the model. By adjusting the number of samples and time range in the prompt template, flexibly control the window size of the model prediction to meet different actual needs.
[0092] The specific content of the prompt template is as follows:
[0093] {Task: Predict future viewports by adaptively fusing trajectory and video features.
[0094] Given historical data features for the past {time_window} frames:......
[0095] Predict the next {prediction_window} viewports in the format (pitch, yaw, roll).}
[0096] Iteratively generate the viewport prediction results, and the generated results are parsed through post-processing into the region numbers to be preferentially transmitted and the corresponding time steps, forming the final structured viewport prediction results.
[0097] Optionally, step S400 specifically includes:
[0098] S401. Obtain and store the network quality data within a preset-length time interval to form the current network status information matrix N t = [B t , J t , L t , where B t represents bandwidth, J t represents jitter, and L t represents latency;
[0099] S402. Dynamically integrate the viewport position prediction result F t and the current network status N t to form the decision information input D t = [N t , F t by analyzing the changes in the viewport position prediction result and the current network status information;
[0100] S403. Expand the video corresponding to the viewport position prediction result into M * N tiles according to the equirectangular projection (ERP), and calculate the priority value P of each tile;
[0101]
[0102] Among them, each tile is marked as T ij, the priority value P is determined by the normalized distance d from the center of the viewport prediction; d is the normalized distance from the center of the tile to the center of the viewport prediction, with a range of [0, 1]; R is the normalized radius of the viewport area, which is fixed at a preset size according to the size of the field of view;
[0103] S404. Divide the tiles with a priority value greater than or equal to 0.7 into high-priority tiles, and divide the tiles with a priority value less than 0.7 into low-priority tiles; divide the bitstream channels according to the priority value P of the tiles, allocate the first bitstream channel and network resources meeting the set QOS standard for the high-priority tiles, and allocate the second bitstream channel for the low-priority tiles;
[0104] S405. According to the decision information input D t , dynamically adjust the resource allocation strategy; among them, the resource allocation strategy includes: pulling high-quality tile data from the central cloud, caching the top preset proportion of tile data with the most access times in the past preset time at the edge node, and locally caching high-priority tiles predicted in real time by the CDN.
[0105] Calculate the priority of the tiles according to the viewport prediction result; before calculating the priority P of each tile, expand the 360° video into M*N tiles by ERP projection, and each tile is marked as T ij , in this embodiment, M = 6 and N = 4. Among them, the preset size with a fixed field of view size is 0.2.
[0106] In this embodiment, the high-quality bitstream channel has a bitrate ≥ 20Mbps, supports a resolution ≥ 3840×2160 (4K), a frame rate ≥ 60fps, and complies with the H.265 / HEVC coding standard; the low-quality bitstream is 720p / 1080p@30fps, H.264 coding, and a bitrate of 7.5Mbps. The network resources meeting the set QOS standard are: bandwidth ≥ 20Mbps, latency ≤ 100ms, and jitter ≤ 30ms. Divide high-quality data and low-quality data according to the priority, corresponding to different bitrate channels, so as to save network resource consumption.
[0107] Decision information input D t , the system dynamically adjusts the resource allocation strategy according to the viewport prediction result and the network status. Combining the changes in user behavior and network status, set the viewport change trigger mechanism and network change trigger mechanism to dynamically change the multi-stream data source. Pull high-quality tile data from the central cloud, cache the top preset proportion of tile data with the most access times in the past preset time at the edge node, and locally cache high-priority tiles predicted in real time by the CDN. Among them, in this embodiment, the preset time is 1 minute and the preset proportion is the top 40%. Change the source of multi-stream data according to the changes in the viewport and the network to ensure the user experience.
[0108] Based on the above embodiments, it further includes:
[0109] Set a viewport change trigger mechanism and a network change trigger mechanism;
[0110] When the offset of the viewport area corresponding to the viewport position prediction result within adjacent time windows is greater than a preset value, trigger the viewport change trigger mechanism to reposition the high-priority area of the viewport area;
[0111] If any index of the current network state drops to reach the first judgment condition, trigger the viewport change trigger mechanism, dynamically adjust the overall area bitrate setting according to the current network state, and switch the data source; wherein, the first judgment threshold includes: bandwidth B t Less than or equal to the preset bandwidth or the jitter time is greater than or equal to the preset jitter time or the delay time is greater than or equal to the preset delay time; specifically in this embodiment: bandwidth B t ≤20Mbps, jitter time J t ≥30ms, delay time L t ≥100ms.
[0112] Optionally, after triggering the viewport change trigger mechanism, request the key data blocks of the new viewport area from the local CDN node. If the local CDN does not cache them, pull them from the edge node and trigger background pre-caching. At the same time, reduce the bitrate of the low-priority area to a preset ratio of the reference value;
[0113] After triggering the network change trigger mechanism, if all indicators of the current network state meet the second judgment condition, pull the original high-quality data stream from the central cloud server and the edge node; wherein, the second judgment condition includes: bandwidth B t Greater than the preset bandwidth or the jitter time is less than the preset jitter time or the delay time is less than the preset delay time; specifically in this embodiment: bandwidth B t >20Mbps, jitter J t <30ms, delay L t <100ms;
[0114] If any index of the current network state drops to reach the first judgment condition, pull the pre-cached bitrate-reduced version from the local CDN. If it misses, pull the low-quality data of real-time transcoding from the edge server.
[0115] Among them, setting the viewport change trigger mechanism includes: defining the pitch / yaw angle position change threshold Δθ th =15°, when the offset Δθ of the predicted viewport within adjacent time windows (0.5s) ≥Δθ thWhen (i.e., when the viewport movement speed v = Δθ / Δt ≥ 30° / s), the relocation of the high-priority area is triggered. If the local CDN does not cache, while reducing the bitrate of the low-priority area, the bandwidth ratio of the high-priority area is increased to ensure the user experience. The above dynamic adjustment of the overall area bitrate setting specifically includes: if the network state drops less, ensure that the quality of the viewport area remains basically unchanged, and the quality of the non-viewport area drops first; if the network state drops significantly, appropriately reduce the quality of the viewport area. Among them, the preset ratio is 50%.
[0116] Optionally, it further includes:
[0117] Adopt a reinforcement learning algorithm to optimize the resource scheduling strategy;
[0118] After the multi-stream data source transmission is completed, the multi-stream data source is synchronously integrated and hierarchically rendered on the client side, and at the same time, the modal resources are dynamically adapted.
[0119] According to the set conditions, the system continuously optimizes the resource scheduling strategy. In this embodiment, a reinforcement learning algorithm is used to optimize the resource scheduling strategy.
[0120] Optimize the resource scheduling strategy by real-time learning of the network state and user feedback information. For example, set the policy reward score as:
[0121]
[0122] Among them, MSE is used to measure the user's video viewing quality, C b is the cost of the streaming media transmission bandwidth, C s is the cost of video data storage, C c is the computational cost of real-time processing or transcoding, and α, ω1, ω2, ω3 are the corresponding weight values, finally realizing an efficient dynamic resource execution strategy.
[0123] Furthermore, a derivative-free optimization algorithm or a heuristic algorithm can also be used to optimize the resource scheduling strategy. By setting the optimization objective function to achieve resource allocation optimization with the goal of improving resource utilization and user experience quality, or a transmission priority configuration can be quickly generated based on heuristic rules.
[0124] After the multi-stream data source transmission is completed, it is necessary to perform synchronous integration and hierarchical rendering on the client side, and at the same time, dynamically adapt relevant modal resources such as GPU, CPU, and network bandwidth to optimize the user experience. Specifically, by adopting a chunk-based progressive transmission method; when the viewport changes, transmit the key frames of the new viewport area, and then complete the relevant predicted frames; when the network bandwidth is insufficient, use the SVC hierarchical coding method to transmit the low-quality base layer and then gradually enhance the high-quality predicted layer. Allocate a low-latency channel for the video synchronization information and the control flow of the task scheduling, and the control flow transmission interval Δt c= 10 ms, synchronize the timing information of multi-stream data to ensure seamless rendering on the client side. The client player synchronizes and integrates the received high and low video stream data, and restores the content in different priority regions to a complete video frame through the synchronization information and chunk numbers of the control stream, ensuring the smoothness and consistency of the user's viewing.
[0125] Compared with the traditional viewport prediction tile-based video stream transmission scheme based on content association, this scheme has achieved a certain performance improvement in typical VR scenarios. In scenarios where the user's behavior changes dynamically (such as rapid view rotation in a VR game scenario), the prediction accuracy of the viewport position for the next 2 s is improved by about 9%; in scenarios where the network bandwidth fluctuates (such as the bandwidth drops from 20 Mbps to 10 Mbps), through the priority-driven chunk bitrate allocation and multi-source data switching strategy, the overall bandwidth consumption is reduced by about 11%, and the peak signal-to-noise ratio (PSNR) in the core viewport area remains at the set requirement, and the quality of the user experience does not decrease significantly.
[0126] In the embodiment of the present invention, by extracting the features of the trajectory timing data and video image data, adjusting the feature weights according to different scene modes, dynamically fusing the user trajectory features and video content features, and using the fine-tuned large model to further capture the complex dependencies between features, the accuracy and generalization ability of viewport prediction are improved; based on the viewport prediction results, a resource dynamic scheduling strategy is further designed to dynamically select multi-stream data sources under different conditions. It solves the technical problems of insufficient viewport prediction accuracy and lack of dynamic switching of video multi-stream data sources when adjusting the tile bitrate according to the viewport prediction results and network status, resulting in insufficient dynamic resource allocation, and realizes excellent prediction stability and reliability in complex user behavior patterns and dynamic scenarios, realizes user-invisible and efficient resource dynamic scheduling, and ensures that the viewing effect of the user in the core area is not affected.
[0127] Embodiment 2
[0128] The present invention also provides a user-invisible resource efficient scheduling device for VR applications, including:
[0129] A feature extraction module for respectively extracting trajectory timing features and video content features from the user trajectory data and image content data of the VR application original data;
[0130] A feature fusion module for respectively calculating the attention weights of the trajectory timing features and video content features through a cross-attention mechanism, adaptively and dynamically adjusting the attention weights of the trajectory timing features and video content features in different scenarios, and generating fused feature information;
[0131] A viewport prediction module is used to fine-tune the pre-trained large model using a low-rank matrix to construct a viewport prediction model; input the fused feature information into the viewport prediction model, and output a viewport position prediction result within a future time window;
[0132] The resource scheduling module is used to dynamically select a multi-stream data source for resource scheduling by taking the viewport position prediction result and the current network status as resource scheduling conditions; allocate the first code stream channel and network resources that meet the set QOS standard to the high-priority area that the user is concerned about in the viewport position prediction result, and use the first code stream channel for the low-priority area; when the user's viewport changes and the network condition decreases, pull the pre-cached reduced-bitrate version of the high-priority area from the local CDN.
[0133] A user-imperceptible resource efficient scheduling device for VR applications provided in an embodiment of the present invention is used to execute a user-imperceptible resource efficient scheduling method for VR applications provided in any embodiment of the present invention, and has corresponding beneficial effects.
[0134] Embodiment 3
[0135] The present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is executed by a processor, the device where the storage medium is located is controlled to execute a method as provided in any one of the embodiments.
[0136] It will be easily understood by those skilled in the art that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for efficient scheduling of user - insensitive resources in a VR application, characterized in that, Including: S100: Extracting trajectory time-series features and video content features from the user trajectory data and image content data of the VR application original data respectively; S200: Calculating the attention weights of the trajectory time-series features and the video content features respectively through a cross-attention mechanism, adaptively and dynamically adjusting the attention weights of the trajectory time-series features and the video content features in different scenarios, and generating fused feature information; S300: Using a low-rank matrix to fine-tune a pre-trained large model to construct a viewport prediction model; inputting the fused feature information into the viewport prediction model, and outputting the viewport position prediction result within a future time window; S400: Dynamically selecting multi-stream data sources for resource scheduling with the viewport position prediction result and the current network state as resource scheduling conditions; for the high-priority areas that the user is concerned about in the viewport position prediction result, allocating a first bitstream channel and network resources that meet the set QOS standard, and using a second bitstream channel for the low-priority areas; when the user's viewport changes and the network condition deteriorates, pulling the downscaled version of the pre-cached high-priority area from the local CDN.
2. The method according to claim 1, wherein Also including: Setting a viewport change trigger mechanism and a network change trigger mechanism; When the offset of the viewport area corresponding to the viewport position prediction result within adjacent time windows is greater than a preset value, triggering the viewport change trigger mechanism to re-locate the high-priority areas of the viewport area; When any index of the current network state drops and reaches the first judgment condition, the viewport change trigger mechanism is triggered, the overall region bitrate setting is dynamically adjusted according to the current network state, and the data source is switched; wherein, the first judgment condition threshold includes: bandwidth B t Less than or equal to the preset bandwidth or the jitter time is greater than or equal to the preset jitter time or the delay time is greater than or equal to the preset delay time.
3. The method according to claim 2, wherein After triggering the viewport change trigger mechanism, requesting the key data blocks of the new viewport area from the local CDN node. If the local CDN does not cache them, pulling them from the edge node and triggering background pre-caching. At the same time, reducing the bitrate of the low-priority area to a preset proportion of the benchmark value; After triggering the network change triggering mechanism, when all indicators of the current network state meet the second judgment condition, pull the original high-quality data stream from the central cloud server and the edge node; wherein, the second judgment condition includes: the bandwidth B t is greater than the preset bandwidth or the jitter time is less than the preset jitter time or the delay time is less than the preset delay time; If any index of the current network state drops to a first threshold, pulling the pre-cached downscaled version from the local CDN. If not hit, pulling the low-quality data of real-time transcoding from the edge server.
4. The method according to claim 1, characterized in that, Step S100 specifically includes: S101: Collecting trajectory data when the user watches a video through a sensor, obtaining image content data by frame extraction of the video, and getting the VR application original data; cleaning the VR application original data, removing outliers in the data collection process, and using linear interpolation to fill in the missing parts; S102. Extract the trajectory time series feature encoding P' = LSTM(P) ∈ R from the user trajectory data of the VR application original data, where P represents the input trajectory time series data, and R d represents that the time series feature result is a d-dimensional vector, and d is a positive integer; d S103. Extract the content feature V' = VIT(image) ∈ R from the image content data of the VR application's original data using the VIT model D , where image represents the input image content data, and R D indicates that the corresponding image content feature is a D-dimensional vector, and D is a positive integer.
5. The method according to claim 1, characterized in that, Step S200 specifically includes: S201. Use the trajectory time series feature as the query matrix Q p , use the video content feature as the key matrix K v and V v . Calculate the correlation between the two through the cross-attention mechanism, and perform weighted summation on the video content feature V ′ using the attention score weight matrix to generate the preliminary fusion feature Z ′ ; Z ′ = Attention(Q p , K v , V v ) * V'; S202. For different mode scenarios, adaptively weight the trajectory time-series feature P ′ and the preliminary fusion feature Z ′ for further integration, and dynamically adjust the influence weights α and β of the two features; in scenarios where the user behavior pattern is stable, adaptively increase the weight β of the trajectory time-series feature; in scenarios where there are regions in the video frame that significantly attract the user's attention, adaptively increase the weight α of the video content feature; to enhance the representativeness of the fusion feature and form the final fusion feature matrix Z t : Z t = α * Z'+ β * P' α+β=1。 6. The method according to claim 1, characterized in that Step S300 specifically includes: S301: The training process of the viewport prediction model includes: adding a trainable low-rank matrix, freezing the weights of the pre-trained large model using the low-rank matrix, and adjusting the large model structure to adapt to the viewport prediction task; S302. During the inference stage of the viewport prediction model, design a prompt template to describe the viewport prediction task type, embed relevant data information, and uniformly guide the large model to generate prediction results; gradually analyze the fused feature information in the input prompt through a multi-layer attention mechanism; in each iteration step, predict the viewport position at future time steps, and gradually generate a complete prediction in the form of a time series. After post-processing, obtain the region numbers to be preferentially transmitted and the corresponding time steps to form the final viewport position prediction result F t .
7. The method according to claim 1, characterized in that Step S400 specifically includes: S401. Obtain and store network quality data within a preset length time interval to form the current network status information matrix N t =[B t , J t , L t , where B t represents bandwidth, J t represents jitter, and L t represents latency; S402. Dynamically integrate the viewport position prediction result F t and the current network status N t to form the decision information input D for the dynamic execution policy of network resources by analyzing the changes in the viewport position prediction result and the current network status information t = [N t , F t ; S403: Unfolding the video corresponding to the viewport position prediction result into M*N tiles according to the equirectangular projection (ERP), and calculating the priority value P of each tile; Among them, each tile is marked as T ij , and the priority value P is determined by the normalized distance d from the center of the viewport prediction; d is the normalized distance from the tile center to the viewport prediction center, with a range of [0, 1]; R is the normalized radius of the viewport area, which is fixed at a preset size according to the size of the field of view; S404: Dividing the tiles with a priority value greater than or equal to 0.7 into high-priority, and dividing the tiles with a priority value less than 0.7 into low-priority; dividing the bitstream channels according to the priority value P of the tiles, allocating a first bitstream channel and network resources that meet the set QOS standard for the high-priority tiles, and allocating a second bitstream channel for the low-priority tiles; S405. Dynamically adjust the resource allocation policy according to the decision information input D t , where the resource allocation policy includes: pulling high-quality tile data from the central cloud, caching the top preset percentage of tile data with the most accesses in the past preset time at the edge node, and caching high-priority tiles predicted in real time by the local CDN.
8. The method according to claim 1, characterized in that Also including: Optimizing the resource scheduling strategy using a reinforcement learning algorithm; After the transmission of the multi-stream data source is completed, the multi-stream data source is synchronously integrated and hierarchically rendered on the client side, while dynamically adapting to modal resources.
9. An efficient scheduling device for user - insensitive resources of a VR application, characterized in that, Including: A feature extraction module, configured to extract trajectory time-series features and video content features from the user trajectory data and image content data of the VR application raw data respectively; A feature fusion module, configured to calculate the attention weights of the trajectory time-series features and the video content features respectively through a cross-attention mechanism, adaptively and dynamically adjust the attention weights of the trajectory time-series features and the video content features in different scenarios, and generate fused feature information; A viewport prediction module, configured to construct a viewport prediction model by using a large pre-trained model fine-tuned with a low-rank matrix; input the fused feature information into the viewport prediction model, and output the viewport position prediction result within a future time window; A resource scheduling module, configured to use the viewport position prediction result and the current network state as resource scheduling conditions to dynamically select a multi-stream data source for resource scheduling; for the high-priority area that the user is concerned about in the viewport position prediction result, allocate a first bitstream channel and network resources that meet the set QOS standard, and use a second bitstream channel for the low-priority area; when the user's viewport changes and the network condition deteriorates, pull the downscaled version of the pre-cached high-priority area from the local CDN.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein when the computer program is run by a processor, it controls the device where the storage medium is located to execute the method provided in any one of claims 1-8.
Citation Information
Cited By
Intelligent switching method for remote video signals
CN121284179A
Intelligent switching method of remote video signals
CN121284179B
Video stream lagging intelligent optimization method and device based on AI prediction
CN121603691A
Resource downloading bandwidth allocation method in same application interface and processing terminal
CN121644369A