Data transmission method and apparatus, communication device, readable storage medium, and program product

WO2026144044A9PCT designated stage Publication Date: 2026-09-03CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/104345
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-31
Filing Date
2025-06-27
Publication Date
2026-09-03

Smart Images

  • Figure CN2025104345_03092026_PF_FP_ABST
    Figure CN2025104345_03092026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to a data transmission method and apparatus, a communication device, a readable storage medium, and a program product. The method comprises: receiving historical saliency information from a terminal, the historical saliency information comprising region saliency information of a VR image frame at a first time, and the region saliency information representing the importance of each image block in the VR image frame; predicting future saliency information on the basis of the historical saliency information, the future saliency information comprising a saliency prediction value of each image block at a second time, and the first time being earlier than the second time; and transmitting the VR image frame to the terminal on the basis of a transmission strategy determined by the future saliency information.
Need to check novelty before this filing date? Find Prior Art

Description

Data transmission methods, apparatus, communication equipment, readable storage media and program products

[0001] Related applications

[0002] This application claims priority to Chinese patent application No. 2024119958707, filed on December 31, 2024, entitled "Data Transmission Method, Apparatus, Communication Equipment, Readable Storage Medium and Program Product", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of network technology, and in particular to a data transmission method, apparatus, communication device, readable storage medium, and program product. Background Technology

[0004] Cloud VR (Cloud Virtual Reality) is a new type of multimedia service that stores VR (Virtual Reality) data in the cloud (such as edge cloud), thereby reducing the storage burden on the terminal and providing more users with high-quality VR content. This data is encoded in the cloud, transmitted over the network to the terminal, and finally decoded and displayed on the terminal's display device. However, limited network bandwidth resources can affect the quality of cloud VR services. Summary of the Invention

[0005] This application provides a data transmission method, apparatus, communication device, readable storage medium, and program product, which can improve the quality of cloud VR services.

[0006] Firstly, this application provides a data transmission method applied to a cloud server, the method comprising:

[0007] Receive historical saliency information from the terminal; the historical saliency information includes the regional saliency information of the VR image frame at the first moment; the regional saliency information indicates the importance of each image block in the VR image frame;

[0008] Future saliency information is predicted based on historical saliency information; future saliency information includes the predicted saliency values ​​of each image patch at the second time point; the first time point is earlier than the second time point;

[0009] Based on the transmission strategy determined by future saliency information, VR image frames are transmitted to the terminal.

[0010] In some embodiments, the first time includes a preset time period prior to the current time; the second time includes at least one of the following: at least one future time and at least one future time period.

[0011] In some embodiments, the regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; the regional saliency values ​​are obtained based on the user saliency value and the image saliency value of the image blocks.

[0012] In some embodiments, the user saliency value is determined by the pixels in the user's visual field region in different image blocks; the user's visual field region represents the visual boundary of the VR user in the VR image; the image saliency value is obtained from the quantized heatmap of the image block.

[0013] In some embodiments, predicting future saliency information based on historical saliency information includes:

[0014] Historical significance information is input into the prediction model, and the prediction model outputs future significance information.

[0015] In some embodiments, the cloud server transmits image blocks of different resolutions to the terminal through different transmission links; wherein the resolution of each image block is determined based on a segmentation threshold; the segmentation threshold is determined based on the saliency prediction value and the transmission capacity of different transmission links.

[0016] In some embodiments, transmitting image blocks of different resolutions to the terminal via different transmission links includes:

[0017] When an image patch is identified as a highly saliency region based on a segmentation threshold, it is transmitted via a high-bandwidth, high-reliability transmission link.

[0018] When an image patch is determined to be a moderately salient region or a lowly salient region based on a segmentation threshold, it is transmitted via a normal bandwidth transmission link.

[0019] Secondly, this application also provides a data transmission method applied to a terminal, the method comprising:

[0020] Send historical saliency information to the cloud server; the historical saliency information includes the regional saliency information of the VR image frame at the first moment; the regional saliency information indicates the importance of each image block in the VR image frame;

[0021] Historical saliency information is used to instruct the cloud server to predict future saliency information and, based on the future saliency information, to determine the transmission strategy for transmitting VR image frames with the terminal; future saliency information includes the saliency prediction values ​​of each image block at the second time; the first time is earlier than the second time.

[0022] In some embodiments, the first time includes a preset time period prior to the current time; the second time includes at least one of the following: at least one future time and at least one future time period.

[0023] In some embodiments, the regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; the regional saliency values ​​are obtained based on the user saliency value and the image saliency value of the image blocks.

[0024] In some embodiments, the user saliency value is determined by the pixels in the user's visual field region in different image blocks; the user's visual field region represents the visual boundary of the VR user in the VR image; the image saliency value is obtained from the quantized heatmap of the image block.

[0025] In some embodiments, before sending historical salience information to the cloud server, the method further includes:

[0026] Sample VR service data within a preset time period to obtain each VR service frame and the corresponding user field of view data within the time series;

[0027] Equidistant cylindrical projection is used to convert VR service frames into VR image frames, and user field-view area data is converted into user field-view area in equidistant cylindrical projection format.

[0028] The VR image frames are divided into blocks to obtain individual image blocks.

[0029] In some embodiments, the terminal receives image blocks of different resolutions transmitted by the cloud server through different transmission links; wherein the resolution of each image block is determined based on a segmentation threshold; the segmentation threshold is determined based on a saliency prediction value and the transmission capacity of different transmission links.

[0030] In some embodiments, the method further includes:

[0031] The received data from each transmission link is aligned to obtain encoded service data;

[0032] The encoded business data is decoded and regions are stitched together to obtain VR image frames.

[0033] In some embodiments, the method further includes:

[0034] Super-resolution processing is performed on the edges of image blocks of different resolutions in VR image frames.

[0035] Thirdly, this application provides a data transmission apparatus for use on a cloud server, the apparatus comprising:

[0036] The information receiving module is used to receive historical saliency information from the terminal; the historical saliency information includes the regional saliency information of the VR image frame at the first moment; the regional saliency information indicates the importance of each image block in the VR image frame;

[0037] The information prediction module is used to predict future saliency information based on historical saliency information; the future saliency information includes the predicted saliency values ​​of each image patch at the second time point; the first time point is earlier than the second time point;

[0038] The image frame transmission module is used to transmit VR image frames with the terminal based on a transmission strategy determined by future saliency information.

[0039] Fourthly, this application also provides a data transmission apparatus for use in a terminal, the apparatus comprising:

[0040] The information sending module is used to send historical saliency information to the cloud server; the historical saliency information includes the regional saliency information of the VR image frame at the first moment; the regional saliency information indicates the importance of each image block in the VR image frame;

[0041] Historical saliency information is used to instruct the cloud server to predict future saliency information and, based on the future saliency information, to determine the transmission strategy for transmitting VR image frames with the terminal; future saliency information includes the saliency prediction values ​​of each image block at the second time; the first time is earlier than the second time.

[0042] Fifthly, this application provides a communication device, including: a transmitter, a processor, and a receiver;

[0043] The receiver is used to receive historical saliency information from the terminal; the historical saliency information includes the regional saliency information of the VR image frame at the first moment; the regional saliency information indicates the importance of each image block in the VR image frame;

[0044] The processor is used to predict future saliency information based on historical saliency information; the future saliency information includes the predicted saliency values ​​of each image patch at the second time point; the first time point is earlier than the second time point;

[0045] The processor is also used to control the transmission of VR image frames between the transmitter and the terminal based on a transmission strategy determined by future saliency information.

[0046] Sixthly, this application also provides a communication device, including a processor and a transmitter;

[0047] The processor controls the transmitter to send historical saliency information to the cloud server; the historical saliency information includes the regional saliency information of the VR image frame at the first moment; the regional saliency information indicates the importance of each image block in the VR image frame;

[0048] Historical saliency information is used to instruct the cloud server to predict future saliency information and, based on the future saliency information, to determine the transmission strategy for transmitting VR image frames with the terminal; future saliency information includes the saliency prediction values ​​of each image block at the second time; the first time is earlier than the second time.

[0049] In a seventh aspect, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the methods described above.

[0050] Eighthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the methods described above.

[0051] The aforementioned data transmission method, apparatus, communication equipment, readable storage medium, and program product allow a cloud server to receive historical saliency information transmitted by a terminal. This historical saliency information includes regional saliency information of VR image frames at a first time, representing the importance of each image block within the VR image frame. The cloud server then predicts future saliency information based on this historical saliency information. This future saliency information includes the predicted saliency values ​​of each image block at a second time, with the first time being earlier than the second time. Furthermore, the server can transmit VR image frames to the terminal based on a transmission strategy determined by the future saliency information. This application embodiment enables dynamic adjustment of cloud VR service resource transmission strategies, achieving an efficient balance between network load and user subjective experience, and improving the service quality of cloud VR services. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the disclosed drawings without creative effort.

[0053] Figure 1 is an application environment diagram of a data transmission method in one embodiment;

[0054] Figure 2 is a schematic diagram of a VR image frame based on equidistant cylindrical projection (ERP) encoding;

[0055] Figure 3 is a flowchart illustrating a data transmission method in one embodiment;

[0056] Figure 4 is a schematic diagram of the user's field of view and image blocks in one embodiment;

[0057] Figure 5 is a schematic diagram of the regional saliency values ​​of an image block in one embodiment;

[0058] Figure 6 is a schematic diagram of constructing training samples by shifting a time window in one embodiment;

[0059] Figure 7 is a schematic diagram of the structure of a two-layer LSTM model in one embodiment;

[0060] Figure 8 is a structural block diagram of a data transmission device in one embodiment;

[0061] Figure 9 is an internal structure diagram of a communication device in one embodiment;

[0062] Figure 10 is an internal structural diagram of a communication device in another embodiment. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0064] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0065] In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0066] Figure 1 is a schematic diagram of an application scenario of a data transmission method provided in an embodiment of this application. As shown in Figure 1, in this scenario, terminal 102 communicates with server 104 via a network. The data storage system can store the data that server 104 needs to process. The data storage system can be integrated on server 104 or placed on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc.

[0067] Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, such as edge cloud, cloud platform servers, cloud servers, and cloud resource pools. It should be noted that servers configured with VR content and servers controlling the distribution of VR resources both belong to the cloud platform in the functional framework.

[0068] Virtual reality is one of the most important technological branches in the rapidly developing metaverse world. A typical VR business scenario involves users wearing VR glasses or headsets, where the device delivers panoramic visual information and immersive auditory information to give users an immersive experience. VR devices can use sensors integrated into the device or placed externally in the room to determine the user's current head position and field of vision, and display the corresponding visible area of ​​the VR service (such as panoramic images or panoramic videos) on the device screen.

[0069] For cloud VR, the commonly used encoding method for VR services is equirectangular projection (ERP). This method can store VR image frames at the tens of millions of pixels level (4K / 8K resolution), which leads to VR service data consuming a large amount of network bandwidth and device memory during transmission and decoding. With limited network bandwidth resources, serious problems such as high latency and high packet loss may occur in cloud VR services, resulting in a poor user experience and potentially causing dizziness in VR users.

[0070] Cloud VR services are highly user-dependent: the cloud VR data ultimately displayed on the terminal is closely related to the user's field of vision and attention information. The user's future field of vision is a comprehensive decision based on the user's current viewpoint, historical viewpoint trajectories, and attention influenced by user and content information. Transmitting too much data beyond the user's field of vision will consume unnecessary network and computing resources. Furthermore, traditional cloud VR services only support global dynamic resolution transmission.

[0071] Based on the aforementioned traditional technologies, and under conditions of limited network resources, this application proposes a multi-link transmission scheme for cloud VR service data based on regional saliency prediction. This scheme enables dynamic adjustment of cloud VR service resource transmission strategies, achieving an efficient balance between network load and user experience, and improving the service quality of cloud VR services. Furthermore, this application can realize dynamic VR service data transmission based on regional saliency prediction, dynamically providing high-definition, sub-high-definition, and low-definition content, optimizing the performance of the cloud VR system, and improving user experience and data transmission efficiency.

[0072] In addressing the issue of network link congestion and high terminal decoding pressure caused by excessive data volume during traditional VR service data transmission, this application embodiment divides VR image frames based on ERP encoding into blocks, calculates regional saliency scores based on the user's field of view and image saliency, and predicts the user's field of view saliency of each image block in future image frames using a two-layer LSTM (Long Short-Term Memory) model. High-bandwidth, high-reliability transmission links are used to transmit video data in the central region of the user's field of view with high saliency, while ordinary bandwidth transmission links are used to transmit video data in the edge and outer regions of the field of view. This achieves a balance between network resource utilization and user service experience by transmitting as much service data as possible with minimal network resources.

[0073] It is understood that this application can be applied to multi-link intelligent transmission scenarios for cloud VR service data. For example, this application can be applied to intelligent routing scenarios for dynamic transmission of service data streams in large-volume multimedia services such as cloud VR and 8K video. The embodiments of this application can provide dynamic and intelligent service data transmission solutions for large-volume multimedia services / operator-operated IPTV (Internet Protocol TV) ultra-high-definition video, VR, and other services for industry customers such as multimedia application vendors, which is beneficial for improving operational efficiency, optimizing resource utilization, and reducing business operating costs.

[0074] It should be noted that the beneficial effects or technical problems solved by the embodiments of this application are not limited to this one, but may also be other implicit or related problems. For details, please refer to the description of the embodiments below.

[0075] Before introducing the specific embodiments of this application, the technical terms involved in this application will be explained:

[0076] Cloud VR (Cloud Virtual Reality): Cloud-based virtual reality is a new type of multimedia service that provides users with an immersive experience based on cloud-edge-device collaboration, effectively reducing the storage and operational burden on terminals.

[0077] ERP (Equirectangular Projection): Equidistant cylindrical projection, the most commonly used frame encoding format for VR services. Its advantage is that, like traditional video frame formats, the image area is rectangular. Its disadvantage is that the image distortion is obvious at the four edges (especially the top and bottom) (as shown in Figure 2).

[0078] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0079] In an exemplary embodiment, as shown in FIG3, a data transmission method is provided. Taking the application of this method to the cloud server in FIG1 as an example, the method includes the following steps 202 to 206.

[0080] Step 202: Receive historical saliency information from the terminal; the historical saliency information includes regional saliency information of the VR image frame at the first time; the regional saliency information indicates the importance of each image block in the VR image frame.

[0081] The terminal can collect and upload historical saliency information to the cloud server. In this embodiment, the historical saliency information may include the regional saliency information of the VR image frame at the first time, which is used to represent the importance of each image block in the VR image frame.

[0082] In some embodiments, VR image frames can be segmented to obtain image blocks. For example, each image block can be understood as an image block region (referred to as a block region) obtained by dividing the VR image frame into regions. For each image block in the VR image frame, the regional saliency information can effectively and quantitatively measure the importance of each region, thereby providing supervision for subsequent regional saliency prediction.

[0083] For example, the regional saliency information can be determined based on image saliency and visual saliency. It is understood that the aforementioned regional saliency information can also take other forms, not limited to those already mentioned in the above embodiments, as long as it can achieve the function of quantitatively measuring the importance of each region.

[0084] The first time can refer to a period of time in the past (or a period of time before the current moment), i.e., historical time. The terminal can periodically upload historical saliency information. In some embodiments, the terminal collects the user's field of view information over a period of time, combines it with image saliency to quantitatively describe the saliency of each image region, and periodically uploads this information (i.e., historical saliency information) to the edge cloud in order to dynamically adjust the data transmission method of VR services.

[0085] Step 204: Predict future saliency information based on historical saliency information; future saliency information includes the predicted saliency values ​​of each image patch at the second time point; the first time point is earlier than the second time point.

[0086] The cloud server receives historical saliency information transmitted from the terminal and can predict future saliency information based on this information. The future saliency information includes the predicted saliency values ​​of each image patch in the VR image frame at the second time point. It can be understood that the second time point refers to a future time, that is, a time after the current moment.

[0087] In this application embodiment, the first time is earlier than the second time; for example, the future saliency information can be the saliency prediction value of each image block in the future VR image frame obtained based on the historical saliency information. The saliency prediction value can represent the predicted value of the importance of each image block in the VR image frame, that is, this application estimates the importance of each image block in the future based on the historical saliency information.

[0088] Step 206: Based on the transmission strategy determined by the future saliency information, transmit VR image frames with the terminal.

[0089] Cloud servers can determine transmission strategies based on future saliency information, and then transmit VR image frames to the terminal according to the transmission strategy. The transmission strategy can be used to dynamically adjust the data transmission method for VR services.

[0090] This application adjusts the transmission method of each area by predicting the importance of each image block in the future, which can ensure the reliable transmission of important business data and achieve an effective balance between improving the user's business experience and saving network resources.

[0091] In the aforementioned data transmission method, based on the regional saliency information representing the importance of each image block in a VR image frame, future saliency information is predicted through historical saliency information. Based on the transmission strategy determined by the future saliency information, VR image frames are transmitted with the terminal, thereby supporting dynamic transmission by region (e.g., dynamic resolution transmission). Thus, under the condition of limited network resources, dynamic adjustment of cloud VR service resource transmission strategy can be achieved, realizing an efficient balance between network load and user experience, and improving the service quality of cloud VR services.

[0092] In one embodiment, the first time period includes a preset time period preceding the current time; the second time period includes at least one of the following: at least one future time period and at least one future time period. A future time period refers to a time period following the current time, and a future time period refers to a time period following the current time.

[0093] The first time period can be a past period of time, specifically a preset time period before the current time. In some embodiments, the end point of the preset time period is the current time, that is, a period of time up to the current time. For example, the value of the preset time period can be 30 seconds to 10 minutes.

[0094] The second time period may include at least one of the following: at least one future moment, and at least one future time period. Future saliency information includes the predicted regional saliency values ​​of each image patch within the future moment / future time period, i.e., the saliency prediction values.

[0095] In some embodiments, the regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; the regional saliency values ​​are obtained based on the user saliency value and the image saliency value of the image blocks.

[0096] Regional saliency information may include the regional saliency values ​​of all image blocks (i.e., all image block regions) in each VR image frame within the time series. In some embodiments, the length of the time series is determined according to a preset time period.

[0097] Taking head-mounted devices as an example, when users experience VR services using head-mounted devices or other terminals, their field of vision will move as their attention changes. In response, the terminal can collect user field of vision information within a preset time period, combine it with image saliency to quantitatively describe the saliency of each image block, and upload historical saliency information to the cloud server in order to dynamically adjust the VR service data transmission method.

[0098] In some embodiments, VR image frames can be image frames in an equidistant cylindrical projection format.

[0099] In some embodiments, the VR image frame in this application can be an EPR-encoded VR service frame. The terminal can sample VR service data within a preset time period to obtain each VR service frame in the time series and the corresponding user field of view area data; then, the terminal can use equidistant cylindrical projection to convert the VR service frame into a VR image frame, and convert the user field of view area data into a user field of view area in equidistant cylindrical projection format; and perform block processing on the VR image frame to obtain each image block.

[0100] The embodiments of this application, based on VR image frame region segmentation and the calculation of region saliency scores based on the number of pixels in the field of view within each region and autoregressive image saliency information, can effectively and quantitatively measure the importance of each region, providing supervision for subsequent region saliency prediction.

[0101] In one possible implementation, the user saliency value is determined by the pixels in the user's visual field region within different image patches; the user's visual field region represents the visual boundary of the VR user in the VR image; and the image saliency value is obtained from the quantized heatmap of the image patch.

[0102] The user saliency value can be determined by the pixels in the user's field of vision area in different image blocks. For example, the user saliency value can represent the percentage of pixels in the image block that fall within the user's field of vision area; where the user's field of vision area represents the visual boundary of the VR user in the VR image.

[0103] The image saliency value is obtained from the quantized heatmap of the image patch; for example, the image saliency value can be obtained by quantizing the heatmap corresponding to the image patch.

[0104] As shown in Figure 4, the terminal can use sensors and other devices to calculate a period of time up to the current moment (i.e., a preset time period, the length of which is T, and the frame rate of the service is N) based on the collected user pose data. FPS Total T*N FPS The user field of view data (the user field of view area corresponding to the service frame) is used to represent the visual boundary of the VR user in the VR screen. For example, the user field of view area can be the area that the user can see in the service data provided by the VR device, such as 120 degrees in the horizontal direction and 90 degrees in the vertical direction.

[0105] It is understandable that user pose data can include information such as the user's head rotation angle and position, such as 6DOF (Degree of Freedom), which means that user pose data includes the spatial coordinates (xyz) of the head and the angle of the line of sight (three angular coordinates).

[0106] If the number of service frames is too large, the original data can be sampled, with 1 frame sampled every S frames, resulting in a time series of length T. SAMPLE =T*N FPS / S is the sampled service frame sequence and the corresponding user field of view area data.

[0107] For each frame of VR service data in the above time series, it can be presented as a rectangular image frame using ERP rectangular encoding, with its width and height set to x pixels and y pixels respectively; the user's field of view area data is also converted into the form of ERP rectangular encoding; as shown in Figure 4, the rectangular image frame is evenly divided into m groups and n groups in the width and height directions, thereby evenly dividing the image into m*n rectangular image blocks (i.e., image block regions, or simply block regions), with the size of each image block being (x / m) pixels * (y / n) pixels.

[0108] Taking an image patch as an example, in this embodiment, the regional saliency value is obtained based on the user saliency value and the image saliency value of the patch. The user saliency value represents the percentage of pixels in the user's field of view that fall within the patch, and the image saliency value is obtained by quantizing the heatmap corresponding to the patch. For example, as shown in Figure 5, the regional saliency value of the patch can be calculated based on the user's field of view and the image saliency. The regional saliency can be measured by combining the following two methods to comprehensively derive the regional saliency value:

[0109] The user saliency value, which represents user saliency, can be obtained by counting the number of pixels contained in the user's field of view within each image block region and dividing it by the total number of pixels in that block region. This ratio is a value within the range of [0,1], and it is the user saliency value of that block region in the VR image frame.

[0110] For the image saliency value representing the saliency of an image, a self-clustering algorithm (such as SLIC (Simple Linear Iterative Clustering)) can be used to predict the importance of image information in an unsupervised manner based on information such as pixel color differences and pixel positions, resulting in a heatmap of the importance of image information (as shown in Figure 5, the high image saliency region, i.e., the high heat region). This heatmap is quantized to the range of [0,1], and the average heat value of all pixels in each block region is calculated to obtain the image saliency value of that block region in the VR image frame. Among them, the heatmap is quantized to the range of [0,1], which can be obtained by using RGB grayscale values ​​of 0-255, dividing each channel by 255.

[0111] The region saliency value = k * user saliency value + (1-k) * image saliency value, where k ranges from [0,1] and is determined manually or by an adaptive algorithm; this application does not limit this value. For T... SAMPLE The above operations are performed on all business image frames encoded by ERP, and finally a set of data with dimensions (mn,T) is obtained, which contains the regional saliency values ​​of all image blocks in each VR image frame within the time series.

[0112] The terminal can periodically send messages carrying the regional saliency values ​​of each block region within the time series to the cloud server (e.g., edge cloud).

[0113] The aforementioned data transmission method supports the quantitative identification of regional salience based on user field-of-view variation patterns and image content salience information, and predicts future regional salience, thereby achieving more accurate regional data clarity management. Furthermore, it can select network links with different bandwidths and reliability for transmission based on the future salience values ​​of each image patch, ensuring reliable transmission of critical business data and achieving an effective balance between improving the user's subjective business experience and conserving network resources.

[0114] In practical applications, predictive models can be used to predict the saliency information of each image patch in future image frames, thereby effectively improving the accuracy of region saliency prediction. In some embodiments, predicting future saliency information based on historical saliency information includes:

[0115] Historical significance information is input into the prediction model, and the prediction model outputs future significance information.

[0116] Cloud servers can input historical saliency information into a prediction model to obtain future saliency information. For example, the prediction model can be a model that outputs predicted saliency by constructing training samples. The prediction model includes, but is not limited to, a pre-trained neural network model.

[0117] Taking edge cloud servers as an example, based on their computing power, neural network models can be trained to predict the salience of image patches in the future. For instance, before prediction, the neural network model can be trained to learn the changing patterns of regional salience. This application's embodiments use deep learning methods to predict future regional salience, thereby achieving more accurate regional data clarity management.

[0118] Taking a pre-trained neural network model as an example, in some embodiments, the pre-trained neural network model is obtained through the following steps:

[0119] Obtain training samples, which include training data and test data constructed by translating regional saliency information;

[0120] The initial neural network model is trained based on the training samples to obtain a trained neural network model. The initial neural network model includes a first-level LSTM and a second-level LSTM. The input of the first-level LSTM is the training data, and the input of the second-level LSTM is the output of the first-level LSTM.

[0121] Specifically, the cloud server can obtain training samples. For example, it can use the regional saliency information included in the historical saliency information as the original training data to construct training samples. The training samples can include training data and test data constructed by translating the regional saliency information.

[0122] Furthermore, the cloud server can train the initial neural network model based on the training samples to obtain a trained neural network model; wherein, the initial neural network model includes a first-level LSTM and a second-level LSTM, the input of the first-level LSTM is the training data, and the input of the second-level LSTM is the output of the first-level LSTM.

[0123] This application uses a two-layer LSTM model as the initial neural network model. By training the user's visual field prediction model based on the two-layer LSTM recurrent neural network structure, it can simultaneously learn the temporal variation pattern of the user's visual field and the implicit spatial correlation between each block at the same time, thereby effectively improving the accuracy of regional saliency prediction.

[0124] Taking image patches as the segmented regions as an example, the training of the future segmented region saliency prediction model is shown in Figure 6. Training samples can be constructed by shifting the time window, and a model can be trained by inputting the previous T... train Using time-to-time data to predict the Tth time train+1 The model for the saliency of each block region at time T can be constructed using training samples in a corresponding format. Each training sample contains the following two parts: ① Training data, i.e., the data from the first T time step. train ① The saliency of each block region at time T; ② Test data, i.e., the data at time T. train+1 The saliency of each block region at time step. This method of constructing training samples through translation allows us to obtain samples with a total length of T. SAMPLE Construct T from the data SAMPLE -T train One sample.

[0125] It should be noted that the labels in Figure 6 represent the predicted results, i.e., the expected model output. X1, X2…X 13It can represent significant data (e.g., regional significance values) at different timestamps. The value of the time window can be set according to requirements; for example, the value of this parameter is determined based on the effect in practical applications, and this application does not limit this. Regarding the method of constructing training samples by translation, it is possible to start from a total length of T. SAMPLE Construct T from the data SAMPLE -T train There are 10 samples, which can be understood as a total of T samples. SAMPLE Data with timestamps, each sample selected from consecutive T... train A sample consists of 10 data points (e.g., sample 1 consists of data points from t1 to T). train Sample 2 is from t2 to T train+1 ...), and thus, through this method of constructing samples, samples of a total length of T can be obtained. SAMPLE Construct T from the data SAMPLE -T train One sample.

[0126] For example, taking the two-layer LSTM model structure shown in Figure 7 as an example, the training process of the two-layer LSTM model may include: the two-layer LSTM model consists of two LSTM stages, wherein the input of the first-stage LSTM is all the time-length T constructed in the above steps. train The training samples have dimensions mn, and the region saliency data at a certain time point is input into each layer according to the time dimension. The output is Y, which is obtained by processing the hidden state H learned by each LSTM layer in the time dimension through the activation function, and contains N neurons. hidden The total amount of output data is T. train ×N hidden .

[0127] Furthermore, the input to the second-level LSTM is the output of the previous-level LSTM, which is sequentially input into each layer of the second-level LSTM model in the order of time. The output is also the result of processing the learned hidden states H of each layer through an activation function. The advantage of a two-layer LSTM lies in its greater model depth, providing better representation capabilities for high-dimensional and complex data. Based on this, each convolutional layer of the first-level LSTM can update its own coefficients, allowing the output data to better enable the training of the second-level LSTM to achieve better results. In addition, the training of the first and second-level LSTMs can be performed simultaneously, enabling the first-level LSTM to learn the implicit relationships between different image regions within the same image frame and construct feature dimensions reflected in the output data of the first-level LSTM, which is then input into the second-level LSTM so that it can better learn the most representative feature dimensions.

[0128] The output of the second LSTM layer can be followed up with a fully connected layer Dense and all the data can be aggregated into a set of data with dimensions [mn, 1]. This data can then be used as the final result for predicting the saliency of each block region of the VR image frame at future moments.

[0129] In some embodiments, during the training of a two-layer LSTM model, multiple sets of training samples can be input into the model for prediction, and the calculated loss function can be continuously passed to the two-layer LSTM model and the fully connected layer through backpropagation to optimize its parameters until the model training converges. Different models can be trained as needed for prediction times at different second time intervals. The loss function can be understood as the difference between the saliency prediction values ​​and the true values ​​of mn region blocks, and the loss function can be obtained by superimposing L1 / L2 values ​​(i.e., direct addition or sum of squares).

[0130] Furthermore, it is possible to predict the saliency of future regions. After training the two-layer LSTM model, the time series obtained by sampling a period of time before the current moment has a length of T. train The saliency of each block region is used as input data, and the Tth region... train+1 The saliency (i.e., the predicted saliency value) of each block region at a given time (i.e., the future time to be predicted) can be obtained by feeding the input data into a trained two-layer LSTM model.

[0131] Furthermore, regarding the setup of the two-layer LSTM model, for example, the training data can be input into a dataset consisting of two layers, each with N hidden neurons. hidden The model is trained using a two-layer LSTM with a resolution of 128. The Return_Sequences parameter is set to True to allow the two layers to pass parameters. At the same time, an Early Dropping strategy is used during the training process to avoid severe overfitting during model training.

[0132] Based on the predicted saliency values ​​of each image patch within a future VR image frame, this application enables multi-link service data transmission. In some embodiments, the cloud server transmits image patches of different resolutions to the terminal via different transmission links;

[0133] The sharpness of each image block is determined based on a segmentation threshold, which is determined according to the saliency prediction value and the transmission capacity of different transmission links.

[0134] Specifically, the transmission strategy in this application embodiment may include the cloud server transmitting each image block to the terminal through different transmission links according to the clarity of each image block in the VR image frame. Based on this transmission strategy, the cloud server can transmit image blocks of different clarity to the terminal through different transmission links.

[0135] The transmission capacity of different transmission links may include, but is not limited to, the number of available links, the current carrying capacity of the links, and the transmission speed. The segmentation threshold can be determined based on the predicted saliency values ​​of each image patch within a future VR image frame, as well as the transmission speed and number of different transmission links. It should be noted that the segmentation threshold can be customized or generated by an adaptive algorithm.

[0136] For example, different transmission links may include a high-speed transmission link (also called a first transmission link) and a low-speed transmission link (also called a second transmission link), wherein the bandwidth of the first transmission link is higher than that of the second transmission link. For example, a link with a bandwidth of 400G or higher is a high-speed transmission link, and a link with a bandwidth of 100G or lower is a low-speed transmission link.

[0137] Taking image blocks as segmented regions as an example, segmentation thresholds can be used to divide each segmented region according to its sharpness. For example, regions with a saliency prediction value higher than a first segmentation threshold (i.e., image blocks) are defined as high-definition regions (corresponding to high-definition content); image blocks with a saliency prediction value between the first and second segmentation thresholds are defined as medium-definition regions (corresponding to medium-definition content); and image blocks with a saliency prediction value lower than the second segmentation threshold are defined as low-definition regions (corresponding to low-definition content). It can be understood that high-definition regions can represent highly saliency regions, such as the highly saliency region in the center of the user's field of vision; medium-definition regions can represent moderately saliency regions, such as the edge of the field of vision; and low-definition regions can represent low-saliency regions, such as the region outside the field of vision. It should be noted that the center of the user's field of vision is not necessarily a highly saliency region. This embodiment of the application evaluates the sharpness of segmented regions through a comprehensive assessment of image saliency.

[0138] For example, clarity can represent the resolution value of a segmented region, with high-definition, sub-high-definition, and low-definition regions having different resolution values.

[0139] Furthermore, the transmission strategy can include transmitting each segment of the VR image frame to the terminal via different transmission links based on the resolution of each segment region. This is known as multi-link service data transmission. Cloud servers (such as edge cloud) can transmit the corresponding VR service data to the terminal via different transmission links based on the resolution of each segment region. This achieves a balance between network resource utilization and user experience by transmitting as much service data as possible with minimal network resources.

[0140] In some embodiments, the cloud server transmits image blocks of different resolutions to the terminal via different transmission links, including:

[0141] When an image patch is identified as a highly saliency region based on a segmentation threshold, it is transmitted via a high-bandwidth, high-reliability transmission link.

[0142] When an image patch is determined to be a moderately salient region or a lowly salient region based on a segmentation threshold, it is transmitted via a normal bandwidth transmission link.

[0143] This can also be understood as the cloud server transmitting image patches with different saliency prediction values ​​to the terminal through different transmission links based on a segmentation threshold, including:

[0144] When the saliency prediction value of an image patch is higher than the segmentation threshold, the image patch is transmitted through the first transmission link.

[0145] When the saliency prediction value of an image patch is lower than the segmentation threshold, the image patch is transmitted through the second transmission link, wherein the bandwidth of the first transmission link is higher than the first bandwidth threshold, and the bandwidth of the second transmission link is lower than the second bandwidth threshold.

[0146] For example, the saliency prediction value of an image patch ranges from [0,1], and the segmentation threshold can be set to 0.7. Image patches with a saliency prediction value greater than or equal to 0.7 are identified as high-saliency regions, and these image patches can be transmitted via a high-bandwidth, high-reliability transmission link (i.e., the first transmission link). Image patches with a saliency prediction value less than 0.7 are identified as medium-saliency or low-saliency regions, and these image patches can be transmitted via a normal-bandwidth transmission link. High-saliency regions can be transmitted via a high-bandwidth, high-reliability transmission link. High-saliency regions may include, but are not limited to, the central region of the user's field of vision, which can characterize the visual boundary of the VR user in the VR image.

[0147] For moderately salient or lowly salient regions, transmission can be achieved through ordinary bandwidth transmission links. Moderately salient regions may include, but are not limited to, the edge of the user's field of vision, while lowly salient regions may include, but are not limited to, regions outside the user's field of vision. For example, high-bandwidth, high-reliability transmission links have higher bandwidth than ordinary bandwidth transmission links; for instance, high-bandwidth, high-reliability transmission links have a bandwidth of 400G or higher, while ordinary bandwidth transmission links have a bandwidth of 100G or lower.

[0148] In some embodiments, during multi-link service data transmission, low-latency high-reliability transmission links are used to transmit high-definition content in highly salient areas (e.g., the central area within the user's field of view) to ensure high-quality image presentation. Low-latency high-reliability transmission links are used to transmit sub-high-definition content in moderately salient areas (e.g., the edge areas within the user's field of view), balancing bandwidth and quality requirements. High-latency low-reliability transmission links are used to transmit low-definition content in less salient areas (e.g., areas outside the field of view) to assist in display and ensure the integrity of content display within the screen area.

[0149] For example, a low-latency, high-reliability transmission link can be equivalent to a high-bandwidth, high-reliability transmission link, and a high-latency, low-reliability transmission link can be equivalent to a normal-bandwidth transmission link. In some embodiments, a high-bandwidth, high-reliability transmission link can be used to transmit video data in the highly salient central area of ​​the user's field of vision, while a normal-bandwidth transmission link can be used to transmit video data in the edge and outer areas of the field of vision. This achieves a balance between network resource utilization and user experience by transmitting as much service data as possible with minimal network resources.

[0150] Taking edge cloud servers as an example, edge cloud can store high-definition, sub-high-definition, and low-definition content versions for each VR service. Edge cloud can possess the training and prediction capabilities of LSTM models, as well as the ability to dynamically adjust data transmission links (i.e., control the routing of router devices on the link from the cloud to the terminal), which can be accomplished through controllers and other devices. Furthermore, in multi-link service data transmission, the link scheduling algorithm executed by edge cloud can consider factors such as bandwidth utilization, user experience, and system load to achieve a balance between network capacity and user services. Edge cloud can also establish a robust monitoring mechanism to evaluate transmission performance in real time and make adjustments, such as through cloud platform capacity, network capacity, terminal capacity, and user service quality evaluation.

[0151] The embodiments described above enable dynamic VR service data transmission based on regional saliency prediction, dynamically providing high-definition, sub-high-definition, and low-definition content, optimizing the performance of the cloud VR system, and improving user experience and data transmission efficiency. Since the user's field of view movement patterns are constantly changing, the above steps can be repeated periodically to continuously update the LSTM prediction model in the cloud, thereby obtaining accurate prediction results.

[0152] In an exemplary embodiment, a data transmission method is provided, which is illustrated using the method applied to the terminal in Figure 1 as an example, and includes the following steps:

[0153] Send historical saliency information to the cloud server; the historical saliency information includes the regional saliency information of the VR image frame at the first moment; the regional saliency information indicates the importance of each image block in the VR image frame;

[0154] Historical saliency information is used to instruct the cloud server to predict future saliency information and, based on the future saliency information, to determine the transmission strategy for transmitting VR image frames with the terminal; future saliency information includes the saliency prediction values ​​of each image block at the second time; the first time is earlier than the second time.

[0155] Specifically, the terminal can send historical saliency information to the cloud server, thereby instructing the cloud server to dynamically adjust the cloud VR service resource transmission strategy, achieve an efficient balance between network load and user subjective experience, and improve the service quality of cloud VR services.

[0156] It is understood that the solution provided by the data transmission method implemented from the terminal perspective is similar to the solution described in the data transmission method implemented from the cloud server perspective. Therefore, the specific limitations of the following one or more data transmission method embodiments implemented from the terminal perspective provided in this application can be found in the limitations of the data transmission method implemented from the cloud server perspective above, and will not be repeated here.

[0157] In some embodiments, the first time period includes a preset time period preceding the current time; the second time period includes at least one of the following: at least one future time period and at least one future time period. A future time period and a future time period refer to a time period or time period following the current time.

[0158] In some embodiments, the end point of the preset time period is the current moment.

[0159] In some embodiments, the regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; the regional saliency values ​​are obtained based on the user saliency value and the image saliency value of the image blocks.

[0160] In some embodiments, the user saliency value is determined by the pixels in the user's visual field region in different image blocks; the user's visual field region represents the visual boundary of the VR user in the VR image; the image saliency value is obtained from the quantized heatmap of the image block.

[0161] In some embodiments, before sending historical salience information to the cloud server, the method further includes:

[0162] Sample VR service data within a preset time period to obtain each VR service frame and the corresponding user field of view data within the time series;

[0163] Equidistant cylindrical projection is used to convert VR service frames into VR image frames, and user field-view area data is converted into user field-view area in equidistant cylindrical projection format.

[0164] The VR image frames are divided into blocks to obtain individual image blocks.

[0165] In some embodiments, the terminal receives image blocks of different resolutions transmitted by the cloud server through different transmission links; wherein the resolution of each image block is determined based on a segmentation threshold; the segmentation threshold is determined based on a saliency prediction value and the transmission capacity of different transmission links.

[0166] In some embodiments, the method further includes:

[0167] The received data from each transmission link is aligned to obtain encoded service data;

[0168] The encoded business data is decoded and regions are stitched together to obtain VR image frames.

[0169] Aligning received data from various transmission links can refer to aligning multiple data streams based on clock synchronization.

[0170] In this process, the server transmits the corresponding VR service data to the terminal through different transmission links. After receiving the data, the terminal can synchronize multiple data streams based on clock synchronization, decode the encoded service data, and perform region stitching to obtain the VR service image frame. Taking an edge cloud server as the server and image blocks as segmented regions as an example, the edge cloud server can transmit the corresponding VR service data to the terminal through different transmission links according to the specified resolution of each segmented region. After receiving the data, the terminal synchronizes multiple data streams based on clock synchronization, decodes the encoded service data, and performs region stitching to reconstruct the VR service image frame.

[0171] In some embodiments, the method may further include:

[0172] Super-resolution processing is performed on the edges of image blocks of different resolutions in VR image frames.

[0173] Specifically, taking image blocks as segmented regions as an example, the terminal can perform super-resolution processing on the edges of segmented regions of different resolutions in a VR image frame, thereby ensuring a smooth transition between segmented regions of different resolutions. For example, for the boundary between high-definition / sub-high-definition / low-definition regions, the terminal can use bicubic sampling to perform super-resolution on the lower-resolution part, ensuring a smooth transition between adjacent regions.

[0174] To further illustrate the solution of this application, a specific example is provided below. Taking the cloud server as the edge cloud server and the image block as the segmented area as an example, the embodiment of this application proposes a VR service transmission strategy. In this strategy, a low-latency, high-reliability 400G color light link is used to transmit high-definition content in high-salience areas (generally the central area within the user's field of view), a low-latency, high-reliability 100G gray light link is used to transmit sub-high-definition content in medium-salience areas (generally the edge area within the user's field of view), and a high-latency, low-reliability 100G gray light link is used to transmit low-definition content in low-salience areas (generally the area outside the field of view).

[0175] Furthermore, the embodiments of this application can realize dynamic strategy adjustment. The terminal collects the user's field of view information and the image saliency information of the service frame (e.g., calculates the saliency information through some unsupervised methods), calculates the saliency score (i.e., the region saliency value) of each block region, and uploads the saliency score sequence information within this 1 second to the edge cloud server every 1 second. The edge cloud server inputs the sequence data into a two-layer LSTM model to perform saliency prediction and obtains the saliency prediction value of each block at future time. Based on the pre-set saliency thresholds of high-definition region / second-high-definition region / low-definition region, the image frame regions of the three resolution types are encoded based on ERP and transmitted through different transmission links in the next 1 second. The image information is restored on the terminal side and the edges of different resolution regions are super-resolution restored using Bicubic encoding to smooth the connection areas between image blocks.

[0176] In one possible implementation, the edge cloud server can receive the saliency score sequence information uploaded by the terminal every 1 second and predict the saliency of the image region in the next 1 second. It can adjust the clarity of the transmitted content and the link selection in real time to ensure a stable and smooth user experience and efficient and economical network transmission.

[0177] The method described in this application, which divides VR image frames into regions and calculates region saliency scores based on the number of pixels in the visual field within each region and autoregressive image saliency information, can effectively and quantitatively measure the importance of each region, providing supervision for subsequent region saliency prediction. Simultaneously, the user visual field region prediction model trained using a two-layer LSTM recurrent neural network structure can simultaneously learn the temporal variation patterns of the user's visual field region and the implicit spatial relationships between different blocks at the same time, thereby effectively improving the accuracy of region saliency prediction. Furthermore, this application selects network links with different bandwidths and reliability for transmission based on the region saliency values ​​of each image block within a future time period / future time interval, ensuring reliable transmission of the most critical business data and achieving an effective balance between improving the user's subjective business experience and conserving network resources.

[0178] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0179] Based on the same inventive concept, this application also provides a data transmission apparatus for implementing the data transmission method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, specific limitations in one or more data transmission apparatus embodiments provided below can be found in the limitations of the data transmission method described above, and will not be repeated here.

[0180] In an exemplary embodiment, as shown in FIG8, a data transmission apparatus is provided, applied to a cloud server, the apparatus comprising:

[0181] The information receiving module 801 is used to receive historical saliency information from the terminal; the historical saliency information includes the regional saliency information of the VR image frame at the first moment; the regional saliency information indicates the importance of each image block in the VR image frame;

[0182] The information prediction module 802 is used to predict future saliency information based on historical saliency information; the future saliency information includes the predicted saliency values ​​of each image patch at the second time point; the first time point is earlier than the second time point;

[0183] The image frame transmission module 803 is used to transmit VR image frames with the terminal based on a transmission strategy determined by future saliency information.

[0184] In some embodiments, the first time includes a preset time period prior to the current time; the second time includes at least one of the following: at least one future time and at least one future time period.

[0185] In some embodiments, the regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; the regional saliency values ​​are obtained based on the user saliency value and the image saliency value of the image blocks.

[0186] In some embodiments, the user saliency value is determined by the pixels in the user's visual field region in different image blocks; the user's visual field region represents the visual boundary of the VR user in the VR image; the image saliency value is obtained from the quantized heatmap of the image block.

[0187] In some embodiments, the information prediction module 802 is used to input historical significance information into the prediction model, and the prediction model outputs future significance information.

[0188] In some embodiments, the cloud server transmits image blocks of different resolutions to the terminal through different transmission links; wherein the resolution of each image block is determined based on a segmentation threshold; the segmentation threshold is determined based on the saliency prediction value and the transmission capacity of different transmission links.

[0189] In some embodiments, the image frame transmission module 803 is configured to transmit the image block via a high-bandwidth, high-reliability transmission link when the image block is determined to be a high-saliency region based on a segmentation threshold, and via a normal-bandwidth transmission link when the image block is determined to be a medium-saliency region or a low-saliency region based on a segmentation threshold.

[0190] In one exemplary embodiment, a data transmission apparatus is provided for use in a terminal, the apparatus comprising:

[0191] The information sending module is used to send historical saliency information to the cloud server; the historical saliency information includes the regional saliency information of the VR image frame at the first moment; the regional saliency information indicates the importance of each image block in the VR image frame;

[0192] Historical saliency information is used to instruct the cloud server to predict future saliency information and, based on the future saliency information, to determine the transmission strategy for transmitting VR image frames with the terminal; future saliency information includes the saliency prediction values ​​of each image block at the second time; the first time is earlier than the second time.

[0193] In some embodiments, the first time period includes a preset time period preceding the current time; the second time period includes at least one of the following: at least one future time period and at least one future time period. A future time period refers to a time period following the current time, and a future time period refers to a time period following the current time.

[0194] In some embodiments, the regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; the regional saliency values ​​are obtained based on the user saliency value and the image saliency value of the image blocks.

[0195] In some embodiments, the user saliency value is determined by the pixels in the user's visual field region in different image blocks; the user's visual field region represents the visual boundary of the VR user in the VR image; the image saliency value is obtained from the quantized heatmap of the image block.

[0196] In some embodiments, the apparatus further includes:

[0197] The sampling module is used to sample VR service data within a preset time period to obtain VR service frames and corresponding user field of view data within the time series.

[0198] The format conversion module is used to convert VR service frames into VR image frames using equidistant cylindrical projection, and to convert user field-view area data into user field-view area in equidistant cylindrical projection format.

[0199] The image segmentation module is used to segment VR image frames into blocks to obtain individual image blocks.

[0200] In some embodiments, the terminal receives image blocks of different resolutions transmitted by the cloud server through different transmission links; wherein the resolution of each image block is determined based on a segmentation threshold; the segmentation threshold is determined based on a saliency prediction value and the transmission capacity of different transmission links.

[0201] In some embodiments, the apparatus further includes:

[0202] The alignment module is used to align the received transmission data from each transmission link to obtain encoded service data.

[0203] The splicing module is used to decode the encoded business data and splice the regions to obtain VR image frames.

[0204] In some embodiments, the apparatus further includes:

[0205] The super-resolution processing module is used to perform super-resolution processing on the edges of image blocks of different resolutions in VR image frames.

[0206] Each module in the aforementioned data transmission device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0207] Figure 9 is a schematic diagram of the communication device provided in an embodiment of this application. Taking the communication device as a server as an example, the communication device may include a receiver 31, a memory 32, a processor 33, at least one communication bus 34, and a transmitter 35. The communication bus 34 is used to realize communication connections between components. The memory 32 may include a high-speed RAM memory, and may also include non-volatile memory (NVM), such as at least one disk storage. The memory 32 can store various programs for performing various processing functions and implementing the method steps of this embodiment. In this embodiment, the transmitter 35 can be a radio frequency processing module or a baseband processing module in the communication device, and the receiver 31 can also be a radio frequency processing module or a baseband processing module in the communication device. The transmitter 35 and the receiver 31 can be integrated together to form a transceiver. Both the transmitter 35 and the receiver 31 can be coupled to the processor 33, and can perform receiving or transmitting actions under the instruction or control of the processor 33.

[0208] In this embodiment, receiver 31 is used to receive historical saliency information from the terminal; the historical saliency information includes regional saliency information of the VR image frame at a first time; the regional saliency information indicates the importance of each image block in the VR image frame.

[0209] The processor 33 is configured to predict future saliency information based on historical saliency information; the future saliency information includes the predicted saliency values ​​of each image block at a second time; the first time is earlier than the second time; and is also configured to control the transmitter 35 to transmit VR image frames with the terminal based on the transmission strategy determined by the future saliency information.

[0210] In one embodiment, the first time period includes a preset time period prior to the current time period; the second time period includes at least one of the following: at least one future time period and at least one future time period.

[0211] In one embodiment, the regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; the regional saliency values ​​are obtained based on the user saliency value and the image saliency value of the image blocks.

[0212] In one embodiment, the user saliency value is determined by the pixels in the user's visual field region within different image patches; the user's visual field region represents the visual boundary of the VR user in the VR image; and the image saliency value is obtained from the quantized heatmap of the image patch.

[0213] In one embodiment, the processor 33 is also configured to input historical significance information into a prediction model, the prediction model outputting future significance information.

[0214] In one embodiment, the processor 33 is configured to control the transmitter 35 to transmit image blocks of different resolutions to the terminal via different transmission links; wherein the resolution of each image block is determined based on a segmentation threshold; the segmentation threshold is determined based on a saliency prediction value and the transmission capacity of different transmission links.

[0215] In one embodiment, the processor 33 is configured to control the transmitter 35 to transmit via a high-bandwidth, high-reliability transmission link when an image patch is determined to be a high-saliency region based on a segmentation threshold; the high-saliency region includes the central region of the user's field of vision, which refers to the region within a preset distance from the center point of the user's field of vision, and the preset distance can be determined according to actual needs; the user's field of vision represents the visual boundary of the VR user in the VR image; when an image patch is determined to be a medium-saliency region or a low-saliency region based on a segmentation threshold, the processor 33 controls the transmitter 35 to transmit via a normal-bandwidth transmission link; the medium-saliency region includes the edge region of the user's field of vision, which refers to the region outside the central region of the user's field of vision; the low-saliency region includes the region outside the user's field of vision; wherein, the bandwidth of the high-bandwidth, high-reliability transmission link is higher than a first bandwidth threshold, and the bandwidth of the normal-bandwidth transmission link is lower than a second bandwidth threshold.

[0216] In one embodiment, a communication device is provided, as shown in FIG10. FIG10 is a schematic diagram of the structure of a terminal provided in an embodiment of the present application. The terminal 700 shown in FIG10 includes: at least one processor 701, a memory 702, at least one network interface 704, and a user interface 703. The various components in the terminal 700 are coupled together through a bus system 705. It is understood that the bus system 705 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 705 also includes a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 705 in FIG10. In addition, the embodiments of the present application also include a transceiver 706, which may be multiple elements, including a transmitter and a receiver, providing a unit for communicating with various other devices over a transmission medium.

[0217] The user interface 703 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).

[0218] It is understood that the memory 702 in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Sync Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 702 of the systems and methods described in this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0219] In some implementations, memory 702 stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof: operating system 7021 and application program 7022.

[0220] The operating system 7021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 7022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this application embodiment can be included in the application program 7022.

[0221] In this embodiment, by calling the program or instructions stored in memory 702, specifically the program or instructions stored in application program 7022, the processor controls the transmitter to send historical saliency information to the cloud server. The historical saliency information includes regional saliency information of the VR image frame at a first time. The regional saliency information indicates the importance of each image block in the VR image frame. The historical saliency information is used to instruct the cloud server to predict future saliency information and, based on the transmission strategy determined by the future saliency information, to transmit VR image frames with the terminal. The future saliency information includes the saliency prediction value of each image block at a second time. The first time is earlier than the second time.

[0222] The methods disclosed in some or all of the above embodiments of this application can also be applied to processor 701, or implemented by processor 701, or implemented by processor 701 in conjunction with other components (e.g., transceivers). Processor 701 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 701 or by instructions in the form of software. The processor 701 mentioned above may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 702, and processor 701 reads the information from memory 702 and, in conjunction with its hardware, completes the steps of the above method.

[0223] It is understood that the embodiments described in this application can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in this application, or combinations thereof.

[0224] For software implementation, the technology described in the embodiments of this application can be implemented by modules (e.g., procedures, functions, etc.) that perform the functions described in the embodiments of this application. The software code can be stored in memory and executed by processor 701. The memory can be implemented in processor 701 or external to processor 701.

[0225] In one embodiment, the first time period includes a preset time period prior to the current time period; the second time period includes at least one of the following: at least one future time period and at least one future time period.

[0226] In one embodiment, the regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; the regional saliency values ​​are obtained based on the user saliency value and the image saliency value of the image blocks.

[0227] In one embodiment, the user saliency value is determined by the pixels in the user's visual field region within different image patches; the user's visual field region represents the visual boundary of the VR user in the VR image; and the image saliency value is obtained from the quantized heatmap of the image patch.

[0228] In one embodiment, the processor is further configured to sample VR service data within a preset time period to obtain each VR service frame and corresponding user field of view data within the time series; convert the VR service frames into VR image frames using equidistant cylindrical projection, and convert the user field of view data into a user field of view in equidistant cylindrical projection format; and perform block processing on the VR image frames to obtain each image block.

[0229] In one embodiment, the processor is further configured to control the receiver to receive image blocks of different resolutions transmitted by the cloud server through different transmission links; wherein the resolution of each image block is determined based on a segmentation threshold; the segmentation threshold is determined based on a saliency prediction value and the transmission capacity of different transmission links.

[0230] In one embodiment, the processor is further configured to perform alignment processing on the received transmission data from each transmission link to obtain encoded service data; and to decode and splice the encoded service data to obtain VR image frames.

[0231] In one embodiment, the processor is also configured to perform super-resolution processing on the edges of image blocks of different resolutions in a VR image frame.

[0232] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0233] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0234] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0235] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0236] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0237] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A data transmission method applied to a cloud server, the method comprising: Receive historical saliency information from the terminal; the historical saliency information includes regional saliency information of VR image frames at a first time. The region saliency information indicates the importance of each image block in a VR image frame; Future saliency information is predicted based on the historical saliency information; the future saliency information includes the predicted saliency values ​​of each image patch at the second time point; The first time is earlier than the second time; Based on the transmission strategy determined by the future saliency information, VR image frames are transmitted with the terminal.

2. The method according to claim 1, wherein, The first time includes a preset time period prior to the current time; the second time includes at least one of the following: at least one future time and at least one future time period.

3. The method according to claim 2, wherein, The regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; The region saliency value is obtained based on the image saliency value of the image block and the user saliency value.

4. The method according to claim 3, wherein, The user saliency value is determined by the pixels in the user's visual field area in different image blocks; the user's visual field area represents the visual boundary of the VR user in the VR screen; the image saliency value is obtained from the quantized heatmap of the image block.

5. The method according to any one of claims 1 to 4, wherein, The step of predicting future significance information based on the historical significance information includes: The historical significance information is input into the prediction model, and the prediction model outputs the future significance information.

6. The method according to any one of claims 1 to 4, wherein, The cloud server transmits image blocks of different resolutions to the terminal through different transmission links; The sharpness of each image block is determined based on a segmentation threshold, which is determined according to the saliency prediction value and the transmission capacity of different transmission links.

7. The method according to claim 6, wherein, The step of transmitting image blocks of different resolutions to the terminal via different transmission links includes: When an image patch is determined to be a highly saliency region based on the segmentation threshold, it is transmitted via a high-bandwidth, high-reliability transmission link. When an image patch is determined to be a moderately salient region or a lowly salient region based on the segmentation threshold, it is transmitted via a normal bandwidth transmission link.

8. A data transmission method applied to a terminal, the method comprising: Send historical saliency information to the cloud server; the historical saliency information includes the regional saliency information of the VR image frame at the first time. The region saliency information indicates the importance of each image block in a VR image frame; The historical saliency information is used to instruct the cloud server to predict future saliency information and, based on the transmission strategy determined by the future saliency information, to transmit VR image frames with the terminal; the future saliency information includes the saliency prediction values ​​of each image block at a second time. The first time is earlier than the second time.

9. The method according to claim 8, wherein, The first time includes a preset time period prior to the current time; the second time includes at least one of the following: at least one future time and at least one future time period.

10. The method according to claim 9, wherein, The regional saliency information includes the regional saliency values ​​of all image blocks in each VR image frame within the time series; The region saliency value is obtained based on the user saliency value of the image patch and the image saliency value.

11. The method according to claim 10, wherein, The user saliency value is determined by the pixels in the user's visual field area in different image blocks; the user's visual field area represents the visual boundary of the VR user in the VR screen; the image saliency value is obtained from the quantized heatmap of the image block.

12. The method according to claim 10, wherein, Before sending historical saliency information to the cloud server, the following is also included: The VR service data within the preset time period is sampled to obtain each VR service frame and the corresponding user field of view data within the time series. The VR service frame is converted into the VR image frame using equidistant cylindrical projection, and the user field of view area data is converted into the user field of view area in equidistant cylindrical projection format. The VR image frames are divided into blocks to obtain individual image blocks.

13. The method according to any one of claims 8 to 12, wherein, The terminal receives image blocks of different resolutions transmitted by the cloud server through different transmission links; The sharpness of each image block is determined based on a segmentation threshold, which is determined according to the saliency prediction value and the transmission capacity of different transmission links.

14. The method according to claim 13, wherein, The method further includes: The received transmission data from each of the transmission links is aligned to obtain encoded service data; The encoded service data is decoded and regions are stitched together to obtain the VR image frame.

15. The method according to claim 14, wherein, The method further includes: Super-resolution processing is performed on the edges of image blocks of different resolutions in the VR image frame.

16. A data transmission apparatus, applied to a cloud server, the apparatus comprising: The information receiving module is used to receive historical salience information from the terminal; The historical saliency information includes the regional saliency information of the VR image frame at the first time. The region saliency information indicates the importance of each image block in a VR image frame; An information prediction module is used to predict future saliency information based on the historical saliency information; the future saliency information includes the saliency prediction values ​​of each image patch at the second time point; The first time is earlier than the second time; The image frame transmission module is used to transmit VR image frames with the terminal based on the transmission strategy determined by the future saliency information.

17. A data transmission apparatus applied to a terminal, the apparatus comprising: The information sending module is used to send historical saliency information to the cloud server; the historical saliency information includes the regional saliency information of the VR image frame at the first time. The region saliency information indicates the importance of each image block in a VR image frame; The historical saliency information is used to instruct the cloud server to predict future saliency information and, based on the transmission strategy determined by the future saliency information, to transmit VR image frames with the terminal; the future saliency information includes the saliency prediction values ​​of each image block at a second time. The first time is earlier than the second time.

18. A communication device, comprising: Transmitter, processor, and receiver; The receiver is used to receive historical saliency information from the terminal; The historical saliency information includes the regional saliency information of the VR image frame at the first time; the regional saliency information indicates the importance of each image block in the VR image frame; The processor is configured to predict future saliency information based on the historical saliency information; the future saliency information includes the predicted saliency values ​​of each image patch at a second time point; the first time point is earlier than the second time point; The processor is also configured to control the transmitter and the terminal to transmit VR image frames based on the transmission strategy determined by the future saliency information.

19. A communication device, comprising a processor and a transmitter; The processor is used to control the transmitter to send historical saliency information to the cloud server; the historical saliency information includes regional saliency information of the VR image frame at a first time; the regional saliency information indicates the importance of each image block in the VR image frame; The historical saliency information is used to instruct the cloud server to predict future saliency information and, based on the transmission strategy determined by the future saliency information, to transmit VR image frames with the terminal; the future saliency information includes the saliency prediction values ​​of each image block at a second time; the first time is earlier than the second time.

20. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15.

21. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 15.