Video data processing method and related device

By hierarchically encode video data in cloud desktop service and sending different image frames according to the terminal network status, the problem of large bandwidth occupancy of video data transmission in cloud desktop service is solved, and effective bandwidth reduction and high frame rate transmission of video data are achieved.

WO2025130048A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/109144
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-07
Filing Date
2024-08-01
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

When transmitting video data, cloud desktop services do not support frame extraction operations due to strong dependencies between video frames to be encoded, resulting in increased bandwidth usage.

Method used

By layering the video data on the cloud desktop server, the basic layer image frame and M group enhancement layer image frame are generated, and different image frames are sent according to the different network status of the terminal to realize the frame extraction function and reduce bandwidth consumption.

Benefits of technology

It effectively reduces bandwidth consumption, especially when the terminal is in a weak network state, the frame extraction function is realized by sending basic layer image frames, reducing the bandwidth usage of video data transmission; at the same time, when the terminal is in a non-weak network state, the basic layer and enhancement layer image frames are sent to ensure the high frame rate and fluency of video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024109144_26062025_PF_FP_ABST
    Figure CN2024109144_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a video data processing method and a related device, which are used for saving bandwidth. The method applicable to a cloud desktop server comprises: on the basis of attribute information of an image frame comprised in video data, processing the image frame, so as to obtain video data to be encoded; encoding said video data in a layered manner to obtain a base layer image frame and M groups of enhancement layer image frames, the base layer image frame being independently decodable and being obtained by performing frame extraction on layered encoding results, and decoding of the M groups of enhancement layer image frames depending on the base layer image frame; sending the base layer image frame and first attribute information of the base layer image frame to a first terminal in a weak network state, the first attribute information indicating attributes of image blocks comprised in the base layer image frame; and sending to a second terminal in a non-weak network state the base layer image frame, the first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames, the second attribute information indicating attributes of image blocks comprised in the N groups of enhancement layer image frames, M and N both being positive integers, and N being less than or equal to M.
Need to check novelty before this filing date? Find Prior Art

Description

Video data processing method and related equipment

[0001] This application claims priority to Chinese patent application No. 202311791366.0, filed with the State Intellectual Property Office on December 22, 2023, entitled “A Method and Apparatus for Deduplicating Remote Desktop Image Data,” and Chinese patent application No. 202410263180.6, filed with the State Intellectual Property Office on March 7, 2024, entitled “Video Data Processing Method and Related Device.” The entire contents of each of these patent applications are incorporated herein by reference. Technical Field

[0002] The present application relates to the field of cloud computing, and in particular to a method for processing video data and related equipment. Background Art

[0003] With the development of computer technology, cloud desktop services have gained widespread application. Cloud desktop services are cloud computing-based desktop services. Users can log in to the purchased cloud desktop service through a terminal to achieve desktop virtualization. Cloud desktop services also enable screen sharing between multiple terminals, displaying the same desktop content on different terminals.

[0004] In the related technical solutions, for the transmission of video data, there is a strong dependency between the video frames to be encoded processed by the cloud desktop server, and frame extraction operations are not supported, which makes the number of frames per second (fps) of the video data transmitted by the cloud desktop server to the client high, resulting in increased bandwidth occupancy.

[0005] Summary of the Invention

[0006] The present application provides a method for processing video data and related equipment for reducing bandwidth.

[0007] In the first aspect, the present application provides a method for processing video data, which is applied to a cloud desktop server. The cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the screen presented by the cloud desktop server, that is, multiple terminal screens are shared. In other words, the video processing method provided by the present application is applied in the screen sharing scenario of the cloud desktop service. The video processing method includes:

[0008] The cloud desktop server processes the image frames based on the attribute information of the image frames included in the video data to obtain the video data to be encoded. The attribute information of the image frames is used to indicate whether the image blocks included in the image frames are identical to the historical image blocks. The image blocks being identical to the historical image blocks mean that the information included in the image blocks is identical to the information included in the historical image blocks, that is, the images displayed by the image blocks are identical to the images displayed by the historical image blocks. The historical image blocks are image blocks obtained by the cloud desktop server from processing the historical video data. Layered encoding of the video data to be encoded is performed to obtain a base layer image frame and M groups of enhancement layer image frames, where M is a positive integer. Specifically, the cloud desktop server may perform temporal layered encoding of the video data to be encoded based on the temporal scalability included in the scalability of video coding. The base layer image frames retain the minimum frame rate information and can be independently decoded. The decoding of the enhancement layer image frames depends on the base layer image frames. The decoded enhancement layer data combined with the decoded base layer data has a higher frame rate. In other words, the M groups of enhancement layer image frames of the base layer image frames are all extracted from the image frames obtained after layered encoding of the video data. That is, both the base layer image frame and the enhanced layer image frame are partial image frames in the video data to be encoded. The cloud desktop server sends different image frames to different terminals according to the different network states of the multiple terminals. For example, the base layer image frame and the first attribute information of the base layer image frame are sent to the first terminal in a weak network state among the multiple terminals, and the first attribute information indicates the image block attributes included in the base layer image frame. The base layer image frame, the first attribute information, N groups of enhanced layer image frames and the second attribute information of N groups of enhanced layer image frames are sent to the second terminal in a non-weak network state among the multiple terminals, and the second attribute information indicates the image block attributes included in the N groups of enhanced layer image frames, where N is a positive integer less than or equal to M.

[0009] In the present application, after the cloud desktop server obtains the base layer image frames and M groups of enhanced layer image frames by layered encoding of the video data to be encoded, different image frames are sent to terminals with different network states. For the first terminal in a weak network state, the cloud desktop server sends the base layer image frames, which are obtained by extracting the frames of the video data to be encoded, thereby realizing the frame extraction function and reducing bandwidth consumption. In addition, for the second terminal in a non-weak network state, the cloud desktop server sends the base layer image frames and the enhanced layer image frames, which will not burden the transmission bandwidth and can also ensure that the second terminal obtains video data with a higher frame rate and displays smoother video.

[0010] In some optional implementations of the first aspect, the cloud desktop server processes the image frames based on attribute information of the image frames included in the video data, including processing image blocks whose attribute information in the image frames is a base layer hit attribute or an enhancement layer hit attribute as a background color. This is because both the base layer hit attribute and the enhancement layer hit attribute indicate that the corresponding image blocks are the same as historical image blocks, that is, these image blocks are previously transmitted and do not need to be encoded again.

[0011] In this application, based on the attribute information of the image frame included in the video data, the cloud desktop server can process the image block in the image frame that is the same as the historical image block as the background color, reducing the image content for video encoding and further reducing the video encoding bandwidth.

[0012] In some optional implementations of the first aspect, the cloud desktop server includes a first cache pool and a second cache pool, the first cache pool being used to store image block identifiers whose attribute information is a base layer hit attribute or a base layer newly added attribute, and the second cache pool being used to store image block identifiers whose attribute information is an enhancement layer hit attribute or an enhancement layer newly added attribute, wherein both the base layer newly added attribute and the enhancement layer newly added attribute indicate that the corresponding image block is different from the historical image block. Therefore, when preprocessing the image frames included in the video data, the desktop server can compare the attribute information of the image frames included in the video data with the first cache pool and / or the second cache pool to determine whether the image blocks in the image frames included in the video data are the same as the historical image blocks.

[0013] In this application, image block identifiers of different attributes are stored in the cache pool of the cloud desktop server, which provides a basis for preprocessing of image frames included in the video data and provides technical support for the implementation of the technical solution of this application.

[0014] In some optional implementations of the first aspect, when the attribute information of the image frame of the video data indicates that the attribute of the target image block in the image frame is a new attribute of the base layer, the identifier of the target image block is stored in the first cache pool.

[0015] In the present application, the cloud desktop server can also update the first cache pool to provide a basis for preprocessing subsequent image data, which is conducive to identifying historical image blocks in subsequent image data, thereby reducing the content that needs to be encoded.

[0016] In some optional implementations of the first aspect, after processing the image frames based on attribute information of the image frames included in the video data, the cloud desktop server may further process the second buffer pool. Specifically, the cloud desktop server deletes image block identifiers in the M groups of enhancement layer image frames in the second buffer pool whose attribute information is enhancement layer hit attributes or enhancement layer newly added attributes.

[0017] In the present application, since the encoding of the enhanced layer image data does not affect the encoding of the base layer image data, after preprocessing the image frames included in the video data, the cloud desktop server deletes the image block identifiers of the enhanced layer attributes in the image frames of the video data in the second cache pool, which does not affect other image data and can also reduce the storage resources used by the cloud desktop server.

[0018] In the second aspect, the present application provides a method for processing video data, which is applied to a client, and the client runs on a terminal. The terminal is connected to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the pictures presented by the cloud desktop server. The cloud desktop server is used to perform layered encoding on the video data to be encoded to obtain a base layer image frame and M groups of enhanced layer image frames. The base layer image frame is independently decoded and is obtained by extracting the image frame obtained by layered encoding. The decoding of the enhanced layer image frame depends on the base layer image frame. The video data to be encoded is obtained by the cloud desktop server processing the image frame based on the attribute information of the image frame included in the video data. The attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block, and M is a positive integer.

[0019] The client receives a base layer image frame and first attribute information of the base layer image frame from a cloud desktop server. The first attribute information indicates attributes of image blocks included in the base layer image frame. The base layer image frame retains information about a minimum frame rate and can be independently decoded. The client manages a third buffer pool, which is used to store attribute information including image block identifiers and image content information of base layer hit attributes. Based on the third buffer pool and the first attribute information, the client processes first image data obtained by decoding the base layer image frame.

[0020] Alternatively, the client receives a base layer image frame, first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames from a cloud desktop server, where the second attribute information indicates the image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M. The decoding of the enhancement layer image frame depends on the base layer image frame, and the decoded enhancement layer data plus the decoded base layer data have a higher frame rate. The client manages a fourth cache pool, which is used to store attribute information including image block identifiers and image content information of enhancement layer hit attributes. Based on the third cache pool, the fourth cache pool, the first attribute information, and the second attribute information, the client processes the first image data obtained by decoding the base layer image frame and the N groups of second image data obtained by decoding the N groups of enhancement layer image frames.

[0021] In this application, the client may receive different types of image frames sent by the cloud desktop server. The client manages the relevant information used to store image blocks with different attributes. Therefore, different types of image frames can be processed accordingly, thereby improving the practicality of the technical solution of this application.

[0022] In some optional implementations of the second aspect, regardless of whether the terminal is in a weak or non-weak network state, the client can receive the base layer image frame and the first attribute information. The client processes the first image data based on the third buffer pool and the first attribute information, including: if the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, the client determines from the third buffer pool a second image block with the same identifier as the first image block, and replaces the image content information of the first image block with the image content information of the second image block. The base layer hit attribute indicates that the corresponding image block is the same as the historical image block. That is, for an image block in the base layer image frame that is the same as the historical image block, when the client decodes the image block, the image block is actually the background color, reducing decoding bandwidth. After decoding, the image content information stored in the third buffer pool is used to restore the image data based on the first attribute information. It should be noted that if the terminal device is in a weak network state and only receives the base layer image frame but not the enhancement layer image frame, the client obtains the video data to be displayed after processing the first image data, which can be displayed on the terminal's display device.

[0023] In this application, after decoding the base layer image frame, the client can fill the historical image block according to the first attribute information to obtain a complete image frame. In other words, during the decoding process, the historical image block is actually the background color, which reduces the decoding bandwidth.

[0024] In some optional implementations of the second aspect, the process of the client processing the first image data based on the third cache pool and the first attribute information is similar to the above and will not be repeated here. The client also processes N groups of second image data based on the fourth cache pool and the second attribute information, or processes N groups of second image data based on the third cache pool, the fourth cache pool and the second attribute information. This depends on the content of the second attribute information. If the attribute of the image block indicated by the second attribute information includes the enhanced layer hit attribute, the client compares the second image data with the fourth cache pool. If the attribute of the image block indicated by the second attribute information includes the base layer hit attribute, the client compares the second image data with the third cache pool. That is, when the second attribute information indicates that the attribute of the third image block in the N groups of second image data is the enhanced layer hit attribute, the client determines a fourth image block with the same identifier as the third image block from the fourth cache pool, and fills the image content information of the third image block with the image content information of the fourth image block. And / or, when the second attribute information indicates that the attribute of the fifth image block in the N groups of second image data is a base layer hit attribute, the client determines a sixth image block with the same identifier as the fifth image block from the third cache pool, and fills the image content information of the fifth image block with the image content information of the sixth image block. Among them, both the enhancement layer hit attribute and the base layer hit attribute indicate that the corresponding image block is the same as the historical image block. That is to say, for the image block in the enhancement layer image frame that is the same as the historical image block, when the client decodes, the image block is actually the background color, which reduces the decoding bandwidth. After decoding, according to the second attribute information, the image content information stored in the third cache pool and / or the fourth cache pool is filled to restore. It should also be noted that if the terminal device is in a non-weak network state and receives the base layer image frame and N groups of enhancement layer image frames, then after processing the first image data and N groups of second image data, the client obtains the video data to be displayed, which can be displayed through the terminal's display device.

[0025] In this application, for the enhanced layer image frame, after decoding, the client fills the historical image block according to the second attribute information to obtain a complete image frame. In other words, during the decoding process, the historical image block is actually the background color, which reduces the decoding bandwidth.

[0026] In some optional implementations of the second aspect, the client may process the third cache pool. Specifically, if the first attribute information indicates that the attribute of the seventh image block in the first image data is a new attribute added to the base layer, then the terminal stores the identifier and image content information of the seventh image block in the third cache pool. And / or, if the second attribute information indicates that the attribute of the eighth image block in the N groups of second image data is a new attribute added to the base layer, then the terminal stores the identifier and image content information of the eighth image block in the third cache pool. Among them, the image block indicated by the enhanced attribute of the base layer is different from the historical image block. In general, the client stores the identifier and image content information of the image block with the attribute information being the new attribute added to the base layer in the third cache pool.

[0027] In this application, the client can also update the third cache pool to provide a basis for the processing of subsequent image data, which is conducive to identifying historical image blocks in subsequent image data and facilitating the splicing of decoded images into a complete image, thereby improving the practicality of the technical solution of this application.

[0028] In some optional implementations of the second aspect, the client may process the fourth buffer pool. Specifically, if the second attribute information indicates that the attribute of the ninth image block in the N sets of second image data is a newly added attribute of the enhancement layer, the client stores the identifier and image content information of the ninth image block in the fourth buffer pool.

[0029] In some optional implementations of the second aspect, after processing the N sets of second image data obtained by decoding the enhancement layer image frame, that is, after obtaining the complete enhancement layer image data, the client may update the fourth buffer pool. This includes deleting, from the fourth buffer pool, identifiers and image memory information of image blocks with enhancement layer hit attributes or enhancement layer newly added attributes indicated by the second attribute information.

[0030] In the present application, since the processing of the enhanced layer image data does not affect the processing of the base layer image data, after the N groups of second image data obtained by decoding the enhanced layer image frame are processed, the client deletes the identifier and image memory information of the image block with the enhanced layer-related attributes corresponding to the second attribute information in the fourth cache pool, which will not affect other image data and can also reduce the storage resources used by the terminal.

[0031] In a third aspect, the present application provides a cloud desktop server, which is connected to multiple terminals, and the screens of the multiple terminals display the images presented by the cloud desktop server; the cloud desktop server includes:

[0032] The processing unit is configured to process the image frames based on attribute information of the image frames included in the video data to obtain video data to be encoded, wherein the attribute information of the image frames is used to indicate whether an image block included in the image frames is the same as a historical image block. The video data to be encoded is layered encoded to obtain a base layer image frame and M groups of enhancement layer image frames, wherein the base layer image frames are independently decoded and are obtained by extracting the image frames obtained by the layered encoding, and the decoding of the enhancement layer image frames depends on the base layer image frames, and M is a positive integer.

[0033] The transceiver unit is configured to send a base layer image frame and first attribute information of the base layer image frame to a first terminal in a weak network state among multiple terminals, where the first attribute information indicates attributes of image blocks included in the base layer image frame. The transceiver unit is configured to send the base layer image frame, the first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames to a second terminal in a non-weak network state among the multiple terminals, where N is a positive integer less than or equal to M.

[0034] The cloud desktop server is used to implement the method shown in the aforementioned first aspect or any possible implementation of the first aspect. Its beneficial effects are similar to those of the aforementioned first aspect or any possible implementation of the first aspect, and will not be repeated here.

[0035] In a fourth aspect, the present application provides a client, which runs on a terminal, the terminal is connected to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the pictures presented by the cloud desktop server; the cloud desktop server is used to perform layered encoding on the video data to be encoded to obtain a base layer image frame and M groups of enhancement layer image frames, the base layer image frames are independently decoded and are obtained by extracting the image frames obtained by layered encoding, the decoding of the enhancement layer image frames depends on the base layer image frames, the video data to be encoded is obtained by the cloud desktop server processing the image frames based on the attribute information of the image frames included in the video data, the attribute information of the image frames is used to indicate whether the image blocks included in the image frames are the same as the historical image blocks, and M is a positive integer.

[0036] The client includes: a transceiver unit configured to receive a base layer image frame and first attribute information of the base layer image frame from a cloud desktop server, wherein the first attribute information indicates attributes of image blocks included in the base layer image frame; a processing unit configured to process first image data obtained by decoding the base layer image frame based on a third buffer pool and the first attribute information, wherein the third buffer pool is managed by the client and is configured to store attribute information including image block identifiers and image content information of base layer hit attributes;

[0037] Alternatively, the client includes a transceiver unit configured to receive a base layer image frame, first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames from a cloud desktop server, where the second attribute information indicates attributes of image blocks included in the N groups of enhancement layer image frames, where N is a positive integer less than or equal to M. A processing unit configured to process first image data obtained by decoding the base layer image frame and N groups of second image data obtained by decoding the N groups of enhancement layer image frames based on a third buffer pool, a fourth buffer pool, the first attribute information, and the second attribute information, wherein the fourth buffer pool is managed by the client, and the fourth buffer pool is configured to store attribute information including image block identifiers and image content information of enhancement layer hit attributes.

[0038] The client is used to implement the method shown in the aforementioned second aspect or any possible implementation of the second aspect, and its beneficial effects are similar to those of the aforementioned second aspect or any possible implementation of the second aspect, and will not be repeated here.

[0039] In a fifth aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster implements the method disclosed in the first aspect or any possible implementation of the first aspect. The beneficial effects thereof are similar to those of the first aspect or any possible implementation of the first aspect, and are not further described here.

[0040] In a sixth aspect, the present application provides a terminal comprising a processor and a memory, wherein the processor stores instructions. When the instructions stored in the memory are executed on the processor, the terminal implements the method described in the second aspect or any possible implementation of the second aspect. The beneficial effects thereof are similar to those of the second aspect or any possible implementation of the second aspect, and are not further described here.

[0041] In a seventh aspect, the present application provides a computer program product comprising instructions. When the instructions are executed by a computer device cluster, the computer device cluster implements the method disclosed in the first aspect and any possible implementation of the first aspect. Alternatively, when the instructions are executed by a terminal, the method disclosed in the second aspect or any possible implementation of the second aspect is implemented.

[0042] In an eighth aspect, the present application provides a computer-readable storage medium storing computer program instructions. When the computer program instructions are executed by a computing device cluster, the computer program instructions implement the method described in the first aspect or any possible implementation of the first aspect. Alternatively, when the computer program instructions are executed by a terminal or client, the computer program instructions implement the method described in the second aspect or any possible implementation of the second aspect.

[0043] The beneficial effects shown in any of the seventh and eighth aspects are similar to those of the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] FIG1 is a schematic diagram of a system architecture provided by an embodiment of the present application;

[0045] FIG2 is a schematic diagram of an application scenario provided by an embodiment of the present application;

[0046] FIG3 is a schematic diagram of another application scenario provided by an embodiment of the present application;

[0047] FIG4 is a flow chart of a method for processing video data provided by an embodiment of the present application;

[0048] FIG5 is another flowchart of a method for processing video data provided by an embodiment of the present application;

[0049] FIG6 is another flowchart of a method for processing video data provided by an embodiment of the present application;

[0050] FIG7 is another flowchart of a method for processing video data provided by an embodiment of the present application;

[0051] FIG8 is another flowchart of a method for processing video data provided by an embodiment of the present application;

[0052] FIG9 is another flowchart of a method for processing video data provided by an embodiment of the present application;

[0053] FIG10 is a schematic diagram of the structure of a client provided in an embodiment of the present application;

[0054] FIG11 is another schematic structural diagram of a terminal provided in an embodiment of the present application;

[0055] FIG12 is a schematic diagram of the structure of a cloud desktop server provided in an embodiment of the present application;

[0056] FIG13 is a schematic diagram of a structure of a computing device provided in an embodiment of the present application;

[0057] FIG14 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0058] FIG15 is another structural diagram of a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] An embodiment of the present application provides a method for processing video data to reduce bandwidth.

[0060] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0061] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way are interchangeable when appropriate, and this is merely a way of distinguishing objects of the same attributes when describing the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or device comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or devices. In addition, "at least one" refers to one or more, and "a plurality" refers to two or more. "and / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: the situation where A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0062] First, please refer to FIG1 , which is a schematic diagram of the system architecture provided in an embodiment of the present application.

[0063] As shown in Figure 1, a tenant logs into cloud platform 30 via client 10 over the internet 20 using the account and password registered on cloud platform 30. Cloud platform 30 manages infrastructure, which includes multiple data centers located in different regions. Each region has at least one cloud data center. For example, Region 1 shown in Figure 1 includes Cloud Data Center 1 and Cloud Data Center 2, while Region 2 includes Cloud Data Center 3 and Cloud Data Center 4. Each cloud data center is equipped with multiple servers, each of which runs business instances (including at least one of virtual machines, containers, and dedicated hosts).

[0064] The cloud platform 30 may provide interfaces related to cloud computing services, such as configuration pages (i.e., interfaces) or APIs for tenants to access cloud services. The cloud service used in this application is a cloud desktop service. If a tenant uses this service, the desktop displayed on the client 10 is the cloud desktop provided by the cloud desktop service. In addition, multiple terminals can use the same cloud service at the same time to achieve screen sharing.

[0065] Below, the application scenario of the video data processing method provided in the embodiment of the present application is briefly described. Please refer to Figures 2 and 3. Figures 2 and 3 are both schematic diagrams of the application scenario provided in the embodiment of the present application.

[0066] The scenario shown in Figure 2 is a "cloud desktop + cloud conferencing" scenario, with conferencing software running within a service instance of a cloud desktop server. In traditional technical solutions, for multi-party collaboration (i.e., screen sharing), the cloud desktop server encodes the video data, and the conferencing software then performs a secondary encoding. Consequently, the client also needs to perform multiple decoding operations before displaying the data.

[0067] In the traditional "point-to-multipoint" or "1-to-N" scenario shown in Figure 3, a service instance running on a cloud desktop server encodes video data and broadcasts it to both the master and collaborating clients. This means that each client receives the same data, potentially exceeding the bandwidth of a terminal with a weak network connection, limiting video data transmission.

[0068] When the video data processing method provided by the embodiment of the present application is applied to the aforementioned scenario, the server can perform layered encoding based on the attribute information of the image frames included in the video data when encoding the video data, thereby obtaining multiple code streams to adapt to terminals in different network states, thereby achieving the effect of reducing bandwidth. The following is a detailed description with reference to the schematic diagram:

[0069] Please refer to FIG4 , which is a flowchart of a method for processing video data provided in an embodiment of the present application, including:

[0070] 401. The cloud desktop server processes the image frame based on the attribute information of the image frame included in the video data to obtain the video data to be encoded, where the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block.

[0071] In an embodiment of the present application, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the screen presented by the cloud desktop server. In other words, the screens of the multiple terminals connected to the cloud desktop server are shared, and the multiple terminals are in a collaborative state.

[0072] Before encoding the image frames included in the video data, the cloud desktop server pre-processes the image frames based on the attribute information of each image frame. The attribute information of the image frame indicates whether the image block in the image frame is the same as the historical image block. The same here means that the information included in the image block is the same as the information included in the historical image block, that is, the picture displayed by the image block is the same as the picture displayed by the historical image block. In addition, the historical image block refers to the image block obtained by the cloud desktop server from processing the historical video data. When the image block is the same as the historical image block, it means that the information included in the image block has been encoded and transmitted. Then, when the cloud desktop server processes the video data to be encoded, it can not encode the image block, thereby reducing the encoding bandwidth.

[0073] In the field of video coding, video data can include multiple types of image frames, some of which can be decoded independently, while the decoding of some image frames depends on other image frames. For example, in a two-layered video data, a group of pictures (GOP) can typically include intra-frames (I frames), small predictive-coded picture frames, and large P frames. The attributes of the image blocks included in the I frames and large P frames are base layer-related attributes, while the attributes of the image blocks included in the small P frames are enhancement layer-related attributes. Image blocks with base layer-related attributes can be decoded independently, while the decoding of image blocks with enhancement layer-related attributes depends on image blocks with base layer attributes.

[0074] In summary, based on the attribute information of the image frame included in the video data, the image frame is processed, including: processing the image block whose attribute information in the image frame is the basic layer hit attribute or the enhanced layer hit attribute as the background color, and both the basic layer hit attribute and the enhanced layer hit attribute indicate that the corresponding image block is the same as the historical image block.

[0075] The following is a detailed description of the process of processing image frames based on attribute information with reference to Figures 5 to 7, which are flowcharts of the method for processing video data provided by embodiments of the present application.

[0076] The embodiments shown in Figures 5 to 7 all use the preprocessing of a single image frame as an example. Based-add represents the added attributes of the base layer, Based-hit represents the hit attributes of the base layer, SVC-add represents the added attributes of the enhancement layer, and SVC-hit represents the hit attributes of the base layer.

[0077] In the embodiment shown in Figure 5, a single frame image is divided into six image blocks. The attribute information for blocks identified as ID4, ID5, and ID6 is the same as the base layer hit attribute, meaning these blocks are identical to the historical image blocks. The cloud desktop server processes these blocks as background colors, resulting in the preprocessed image frame shown in Figure 5, where the gray portion represents the background color.

[0078] In the embodiment shown in Figure 6, a single frame is divided into six image blocks. The attribute information for blocks identified as ID4, ID5, and ID6 is enhancement layer hit attributes, meaning these blocks are identical to the historical image blocks. The cloud desktop server processes these blocks as background colors, resulting in the preprocessed image frame shown in Figure 6, where the gray portion represents the background color.

[0079] Regardless of the embodiment shown in Figure 5 or Figure 6, the attributes of each image block in a single image frame are related attributes of the same layer. In actual applications, the same image frame can also include related attributes of different layers. For example, as shown in Figure 7, a single frame image is divided into 6 image blocks, wherein the attribute information of the image blocks with image block identifiers ID1, ID2 and ID3 are related attributes of the basic layer, and the attribute information of the image blocks with image block identifiers ID4, ID5 and ID6 are related attributes of the enhanced layer. In addition, the attribute information of the image block with image block identifier ID3 is the basic layer hit attribute, and the attribute information of the image blocks with image block identifiers ID5 and ID6 is the enhanced layer hit attribute, indicating that these image blocks are the same as the historical image blocks. The cloud desktop server processes these image blocks as background colors to obtain the preprocessed image frame shown in Figure 6, where the gray part in the figure represents the background color.

[0080] It should be noted that the embodiments shown in Figures 5 to 7 take the example of dividing a single frame image into 6 image blocks. In actual applications, each image frame can also be divided into more or fewer image blocks, and the sizes of the image blocks can be completely or partially the same, or can be completely different, which is not limited here.

[0081] Based on the above description of Figures 5 to 7, it can be seen that in an embodiment of the present application, based on the attribute information of the image frame, the historical image blocks in the image frame can be processed as background colors, which reduces the image content for video encoding and reduces the bandwidth of video encoding.

[0082] 402. The cloud desktop server performs layered encoding on the encoded video data to obtain a base layer image frame and M groups of enhanced layer image frames. The base layer image frames are independently decoded and are obtained by extracting the image frames obtained by layered encoding. The decoding of the enhanced layer image frames depends on the base layer image frames, and M is a positive integer.

[0083] After the cloud desktop server obtains the video data to be encoded, it can perform time-domain layered encoding on the video data to be encoded according to the time-domain scalability characteristics in the video coding scalability to obtain a base layer image frame and M groups of enhanced layer image frames. It can be understood that time-domain layered coding includes frame extraction operations. Then the base layer image frame and the M groups of enhanced layer image frames are both partial image frames in the video data to be encoded, or in other words, the base layer image frame and the M groups of enhanced layer image frames are both image frames extracted from the image frames obtained by layered encoding of the video data to be encoded. Among them, the base layer image frame retains the information of the minimum frame rate and can be decoded independently. The decoding of the enhanced layer image frame depends on the base layer image frame, and the decoded enhanced layer data plus the decoded base layer data have a higher frame rate.

[0084] In addition, the attribute information of the image blocks included in the base layer image frame is all base layer related attributes, and the attribute information of the image blocks included in the enhancement layer image frame includes enhancement layer related attributes, or enhancement layer related attributes and base layer related attributes. This can be determined based on the dependency relationship between the image blocks in the enhancement layer image frame and the image blocks in other frames.

[0085] 403. The cloud desktop server sends a base layer image frame and first attribute information of the base layer image frame to a first terminal in a weak network state, where the first attribute information indicates attributes of image blocks included in the base layer image frame.

[0086] The cloud desktop server can determine whether a terminal connected to the cloud desktop server is in a weak network state based on parameters such as round-trip time (RTT), network jitter, or packet loss rate. The type, number, and thresholds of the parameters used to determine whether a terminal is in a weak network state or not can be defined based on actual business needs and network conditions, and are not specifically defined here.

[0087] Optionally, a weak network state can be defined based on multiple types of parameters. For example, a weak network state can be defined as one with an RTT greater than 100ms, a network jitter range exceeding ±10ms, and a packet loss rate greater than 15%. A non-weak network state can be defined as one that is not within this range.

[0088] Optionally, a weak network state can be defined based on a single parameter type. For example, a network state with an RTT greater than 80ms is defined as a weak network state. A network state outside this range is defined as a non-weak network state.

[0089] The first terminal, which is in a weak network state and has limited bandwidth resources, sends the base layer image frame and first attribute information of the base layer image frame to the first terminal. Because the base layer image frame retains the minimum frame rate information and can be independently decoded, the first terminal can still decode and obtain smooth video data within the limited bandwidth resources. In other words, the base layer image frame sent by the cloud desktop server to the first terminal is actually a frame extraction, or the result of frame loss, of the complete encoded video data.

[0090] 404. The cloud desktop server sends the base layer image frame, the first attribute information, N groups of enhancement layer image frames and the second attribute information of the N groups of enhancement layer image frames to the second terminal in a non-weak network state, where the second attribute information indicates the image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M.

[0091] For a second terminal in a non-weak network state, the cloud desktop server not only sends the base layer image frame and the first attribute information of the base layer image frame, but also sends N groups of enhancement layer image frames and the second attribute information of the N groups of enhancement layer image frames. Because the decoding of the enhancement layer image frame depends on the base layer image frame, and the decoded enhancement layer data plus the decoded base layer data have a higher frame rate, the video decoded by the second terminal is smoother than the video decoded by the first terminal.

[0092] In addition, among the second terminals in a non-weak network state, the network status of each terminal may be different, and the number of enhanced layer image frames sent by the cloud desktop server may also be different. For example, assuming that the network status of terminal A among the second terminals is better than the network status of terminal B among the second terminals, the number of groups of enhanced layer image frames sent by the cloud desktop server to terminal A is greater than the number of groups of enhanced layer image frames sent by the cloud desktop server to terminal B.

[0093] For example, assume that a weak network state is defined as one with an RTT greater than 100ms, a network jitter range exceeding ±10ms, and a packet loss rate greater than 15%, and a non-weak network state is defined as one outside this range. Furthermore, the cloud desktop server performs layered encoding on the video data to be encoded, generating a base layer image frame and three sets of enhancement layer image frames, i.e., M = 3.

[0094] For example, if Terminal 1 has an RTT of 150ms, a jitter range of ±15ms, and a packet loss rate of 16%, Terminal 2 has an RTT of 80ms, a jitter range of ±8ms, and a packet loss rate of 10%, and Terminal 3 has an RTT of 50ms, a jitter range of ±5ms, and a packet loss rate of 5%, then Terminal 1 is in a weak network state, while Terminals 2 and 3 are both in a strong network state, and Terminal 3's network state is better than Terminal 2's.

[0095] Then, the cloud desktop server can send the base layer image frame and the first attribute information to terminal 1; send the base layer image frame, the first attribute information, 1 group of enhancement layer image frames and the second attribute information corresponding to the group of enhancement layer image frames to terminal 2; and not only send the base layer image frame and the first attribute information to terminal 3, but also send 2 or 3 groups of enhancement layer image frames and the corresponding second attribute information.

[0096] It should be noted that no matter how many sets of enhancement layer image frames the cloud desktop server sends to the second terminal, the attribute information of each enhancement layer image frame will be sent to the second terminal, so that the second terminal processes the enhancement layer image frame according to the second attribute information.

[0097] Based on the above description, it can be seen that in this application, after the cloud desktop server obtains the base layer image frame and M groups of enhanced layer image frames by layered encoding of the video data to be encoded, different image frames are sent to terminals with different network states. For the first terminal in a weak network state, the cloud desktop server sends the base layer image frame, which is obtained by extracting the frame of the video data to be encoded, realizing the frame extraction function and reducing the bandwidth. In addition, for the second terminal in a non-weak network state, the cloud desktop server sends the base layer image frame and the enhanced layer image frame, which will not burden the transmission bandwidth, but also ensure that the second terminal obtains video data with a higher frame rate and displays a smoother video.

[0098] In some optional embodiments, the cloud desktop server adopts a hierarchical cache pool, and cache pools at different levels are used to store image block identifiers of different layer attributes. Specifically, the cloud desktop server includes a first cache pool and a second cache pool, the first cache pool is used to store image block identifiers whose attribute information is a base layer hit attribute or a base layer newly added attribute, and the second cache pool is used to store image block identifiers whose attribute information is an enhancement layer hit attribute or an enhancement layer newly added attribute. Among them, the base layer hit attribute and the enhancement layer hit attribute both indicate that the corresponding image block is the same as the historical image block, and the base layer newly added attribute and the enhancement layer newly added attribute both indicate that the corresponding image block is different from the historical image block.

[0099] In this application, image block identifiers of different attributes are stored in the cache pool of the cloud desktop server, which provides a basis for preprocessing of image frames included in the video data and provides technical support for the implementation of the technical solution of this application.

[0100] When processing the video data, the cloud desktop server may also update the cache pool, which will be described with reference to the examples in Figures 5 to 7 .

[0101] As shown in Figure 5, the attribute information for image blocks identified as ID1, ID2, and ID3 is newly added to the base layer, indicating that these image blocks are different from historical image blocks and have never been transmitted before. Furthermore, because these image blocks are in the base layer, the decoding of other image blocks depends on them. Therefore, the cloud desktop server adds the identifiers of these image blocks to the first cache pool, completing the update to the first cache pool required for this preprocessing.

[0102] As shown in Figure 6, the attributes of the image blocks identified by ID1, ID2, and ID3 are newly added to the enhancement layer, indicating that these image blocks are different from the historical image blocks and have never been transmitted before. The cloud desktop server adds the identifiers of these image blocks to the second cache pool, completing the update required for this preprocessing.

[0103] As shown in Figure 7, the attribute information of the image blocks identified as ID1 and ID2 is a new attribute added to the base layer, indicating that these image blocks are different from historical image blocks and are image blocks that have never been transmitted before. In addition, since these image blocks are in the base layer, the decoding of other image blocks depends on these image blocks. Therefore, the cloud desktop server adds the identifiers of these image blocks to the first cache pool, completing the update of the first cache pool required for this preprocessing. The attribute information of the image block identified as ID4 is a new attribute added to the enhancement layer, indicating that the image block is different from historical image blocks and is an image block that has never been transmitted before. The cloud desktop server adds the identifier of the image block to the second cache pool, completing the update of the second cache pool required for this preprocessing.

[0104] In the present application, the cloud desktop server can also update the first cache pool to provide a basis for preprocessing subsequent image data, which is conducive to identifying historical image blocks in subsequent image data, thereby reducing the content that needs to be encoded.

[0105] In some optional embodiments, after processing the image frames based on the attribute information of the image frames included in the video data, the method further includes: deleting, from the second buffer pool, image block identifiers of the M groups of enhancement layer image frames whose attribute information is enhancement layer hit attributes or enhancement layer newly added attributes. In other words, after preprocessing is completed, the cloud desktop server deletes the image block identifiers related to the current preprocessing from the second buffer pool.

[0106] For example, in the embodiments shown in Figures 6 and 7, since the image blocks in the second cache pool are enhancement layers, the decoding of other image blocks does not depend on these image blocks. Therefore, after the preprocessing is completed, the cloud desktop server can also delete the image block identifiers (including ID1-6) related to the preprocessing here in the second cache pool.

[0107] In the present application, since the encoding of the enhanced layer image data does not affect the encoding of the base layer image data, after preprocessing the image frames included in the video data, the cloud desktop server deletes the image block identifiers of the enhanced layer attributes in the image frames of the video data in the second cache pool, which does not affect other image data and can also reduce the storage resources used by the cloud desktop server.

[0108] The preceding text describes the operations performed by the cloud desktop server in the video data processing method provided in the embodiments of this application. In the embodiments of this application, a client runs on the terminal. The cloud desktop server sends encoded data to the terminal. From a software perspective, this means that the client running on the terminal obtains the encoded data. The client's operations vary depending on the encoded data obtained, and are described separately below.

[0109] 1. A client running on a terminal in a weak network state (hereinafter referred to as the first terminal).

[0110] The client manages a third cache pool, which is used to store attribute information including image block identifiers and image content information of base layer hit attributes. In other words, the first terminal includes a third cache pool, which is managed by the client. In this application, the client's management of the cache pool includes adding, deleting, changing the mapping relationship of the identifiers and / or image content information of the image blocks stored in the cache pool, and also includes operations such as deleting the cache pool that can change the cache pool itself or change the content stored in the cache pool.

[0111] After obtaining the base layer image frame and the first attribute information, the client running on the first terminal first decodes the base layer image frame to obtain decoded first image data. Then, the client processes the first image data according to the third buffer pool and the first attribute information to obtain video data to be displayed.

[0112] The third cache pool processes the first image data based on the first attribute information, including: if the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block having the same identifier as the first image block from the third cache pool, and filling the image content information of the first image block with the image content information of the second image block. This process can also be understood as an image stitching process, wherein the base layer hit attribute indicates that the corresponding image block is the same as the historical image block.

[0113] For example, reference is made to Figure 8 , which is a flowchart illustrating a method for processing video data provided in an embodiment of the present application. The embodiment illustrated in Figure 8 uses the preprocessing of a single base layer image frame as an example. Based-add indicates newly added attributes to the base layer, and Based-hit indicates hit attributes to the base layer.

[0114] In the embodiment shown in Figure 8, a single-frame base layer image is divided into six image blocks, of which the attribute information of the image blocks identified as ID4, ID5, and ID6 is the base layer hit attribute, which means that these image blocks are the same as the historical image blocks. When the cloud desktop server encodes the base layer image frame, these image blocks are the background color. After the client running on the first terminal decodes the base layer image frame, these image blocks still have the background color. The image in the upper left corner of the figure represents the decoded image, that is, the first image data.

[0115] Based on the identifiers of the image blocks, the client running on the first terminal identifies image blocks in the third buffer pool that have the same identifiers as the image blocks, as well as the image content information corresponding to the image blocks. Furthermore, the third buffer pool also stores position information for the image blocks, indicating their positions within the image frames, thereby ensuring accurate positioning of the image frames during splicing.

[0116] Based on this, the client daunt running on the first terminal can obtain the position information and image content information of the image blocks identified as ID4, ID5, and ID6 from the third buffer pool, fill them into the corresponding image blocks, and obtain the filled image frame. Similar operations are performed on each base layer image frame to obtain the image to be displayed.

[0117] In this application, after decoding the base layer image frame, the client can fill the historical image block according to the first attribute information to obtain a complete image frame. In other words, during the decoding process, the historical image block is actually the background color, which reduces the decoding bandwidth.

[0118] 2. A client running on a terminal in a non-weak network state (hereinafter referred to as the second terminal).

[0119] The client manages a third cache pool and a fourth cache pool. The third cache pool is used to store attribute information including image block identifiers and image content information of base layer hit attributes, and the fourth cache pool is used to store attribute information including image block identifiers and image content information of enhancement layer hit attributes. In other words, the second terminal includes the third cache pool and the fourth cache pool.

[0120] After obtaining the base layer image frame and the first attribute information, the client running on the second terminal first decodes the base layer image frame to obtain decoded first image data. The first image data is then processed based on the third buffer pool and the first attribute information. The specific process is similar to the process of processing the base layer image frame by the client running on the first terminal, and is not further described here.

[0121] The client running on the second terminal further processes N sets of enhancement layer image frames, including: first decoding the N sets of enhancement layer image frames to obtain N sets of second image data. Then, processing the N sets of second image data based on the fourth buffer pool and the second attribute information. Alternatively, processing the N sets of second image data based on the third buffer pool, the fourth buffer pool, and the second attribute information.

[0122] It can be understood that for the client running on the second terminal, after receiving the base layer image frames and N sets of enhancement layer image frames, decoding these image frames to obtain first image data and M sets of second image data, and then processing the first image data and M sets of second image data based on the corresponding attribute information to obtain the video data to be displayed. Compared with the video data to be displayed on the first terminal, the video data to be displayed on the second terminal has a higher frame rate, and the video playback is smoother.

[0123] The process of processing N sets of second image data includes: if the second attribute information indicates that the attribute of the third image block in the N sets of second image data is an enhancement layer hit attribute, determining a fourth image block from the fourth buffer pool having the same identifier as the third image block, and filling the image content information of the third image block with the image content information of the fourth image block. And / or, if the second attribute information indicates that the attribute of the fifth image block in the N sets of second image data is a base layer hit attribute, determining a sixth image block from the third buffer pool having the same identifier as the fifth image block, and filling the image content information of the fifth image block with the image content information of the sixth image block. Both the enhancement layer hit attribute and the base layer hit attribute indicate that the corresponding image block is the same as the historical image block. In summary, for the historical data block in the decoded image data, based on the attribute information, the image content information of the image block is obtained from the corresponding buffer pool and filled in.

[0124] For example, FIG9 is provided for illustration. FIG9 is a flowchart illustrating a method for processing video data according to an embodiment of the present application. The embodiment shown in FIG9 uses the preprocessing of a single enhancement layer image frame as an example. Based-add indicates a newly added attribute of the base layer, Based-hit indicates a hit attribute of the base layer, SVC-add indicates a newly added attribute of the enhancement layer, and SVC-hit indicates a hit attribute of the base layer.

[0125] In the embodiment shown in Figure 9, a single-frame enhancement layer image is divided into six image blocks. The attribute information of the image block identified as ID3 is the base layer hit attribute, while the attribute information of the image blocks identified as ID5 and ID6 is the enhancement layer hit attribute, which means that these image blocks are the same as the historical image blocks. When the cloud desktop server encodes the enhancement layer image frame, these image blocks are the background color. After the client running on the second terminal decodes the image frame, these image blocks still have the background color. The image in the upper left corner of the figure represents the decoded image, i.e., the second image data.

[0126] Based on the base layer image block's ID3, the client running on the second terminal identifies an image block in the third buffer pool that matches the ID3 of the image block, along with the image content information corresponding to the image block. Furthermore, the third buffer pool also stores the image block's position information, which indicates the image block's location within the image frame, thereby ensuring accurate positioning of the image frame splicing.

[0127] Based on the IDs ID5 and ID6 of the enhancement layer image blocks, the client running on the second terminal identifies an image block in the fourth buffer pool that has the same ID as the two image blocks, as well as the image content information corresponding to each of the two image blocks. Furthermore, the fourth buffer pool also stores positional information for the image blocks, indicating their location within the image frame, thereby ensuring accurate positioning of the image frame splicing.

[0128] Based on this, the client running on the second terminal can obtain the image block location information and image content information for the image block identified by ID3 from the third buffer pool; and obtain the image block location information and image content information for the image blocks identified by ID5 and ID6 from the fourth buffer pool. The image content information from the buffer pool is then applied to the corresponding image blocks, resulting in a populated image frame. Similar operations are performed on each enhancement layer image frame, and the client running on the second terminal obtains M sets of reprocessed second image data. Combined with the reprocessed first image data corresponding to the base layer image frame, the video data to be displayed is obtained.

[0129] It should be noted that, in the embodiment shown in FIG9 , an example is taken in which the image blocks of a single enhancement layer image frame include base layer related attributes and enhancement layer related attributes. In actual applications, the image blocks of a single enhancement layer image frame may only include enhancement layer related attributes, which is not specifically limited here.

[0130] In this application, for the enhanced layer image frame, after decoding, the client fills the historical image block according to the second attribute information to obtain a complete image frame. In other words, during the decoding process, the historical image block is actually the background color, which reduces the decoding bandwidth.

[0131] Based on the previous description of the decoding process, the client / terminal can employ a hierarchical buffer pool system to isolate base layer attributes from enhancement layer attributes. Furthermore, base layer frames are decoded and reprocessed solely using a third buffer pool storing base layer attributes. This ensures that even if enhancement layer frames are discarded, the remaining bitstream can still be decoded properly, providing a frame extraction feature and robustness against weak network conditions, provided the base layer frame stream is intact.

[0132] In some optional implementations, after processing the decoded image data based on the attribute information of the image frame, the client may further process the buffer pool, as illustrated in conjunction with the examples of FIG8 and FIG9 .

[0133] Optionally, when the first attribute information indicates that the attribute of the seventh image block in the first image data is a new attribute of the base layer, it indicates that the seventh image block is an image block that has never been transmitted before. In addition, because the image block is in the base layer, the decoding of other image blocks depends on the seventh image block. Therefore, the client stores the identifier and image content information of the seventh image block in the third cache pool. Optionally, the client may also store the position information of the seventh image block in the third cache pool, where the position information indicates the position of the seventh image block in the first image data.

[0134] For example, as shown in FIG8 , the attribute information of the image blocks identified as ID1, ID2, and ID3 are newly added attributes to the base layer, and the client adds the identification, image content information, and location information of these image blocks to the third cache pool to complete the update of the third cache pool required for this reprocessing.

[0135] Optionally, when the second attribute information indicates that the attribute of the eighth image block in the N groups of second image data is a new attribute of the base layer, it means that the eighth image block is different from the historical image block and is an image block that has never been transmitted before. In addition, since the image block is in the base layer, the decoding of other image blocks depends on the eighth image block. Therefore, the client stores the identifier and image content information of the eighth image block in the third cache pool. Optionally, the client may also store the position information of the eighth image block in the third cache pool, where the position information indicates the position of the eighth image block in the second image data.

[0136] For example, as shown in FIG9 , the attribute information of the image blocks identified as ID1 and ID2 are newly added attributes to the base layer, and the client adds the identification, image content information and location information of these image blocks to the third cache pool to complete the update of the third cache pool required for this reprocessing.

[0137] In this application, the client can also update the third cache pool to provide a basis for the processing of subsequent image data, which is conducive to identifying historical image blocks in subsequent image data and facilitating the splicing of decoded images into a complete image, thereby improving the practicality of the technical solution of this application.

[0138] Optionally, when the second attribute information indicates that the attribute of the ninth image block in the N sets of second image data is a newly added attribute of the enhancement layer, this indicates that the ninth image block is an image block that has never been transmitted before. The client stores the identifier and image content information of the ninth image block in the fourth cache pool. Optionally, the client may also store position information of the ninth image block in the fourth cache pool, where the position information indicates the position of the ninth image block in the second image data.

[0139] For example, as shown in FIG9 , the attribute information of the image block identified as ID4 is a new attribute of the enhancement layer. The client adds the identification, image content information and location information of the image block to the fourth cache pool to complete the update of the fourth cache pool required for this reprocessing.

[0140] In some optional implementations, after reprocessing the second image data, the client may further process the fourth buffer pool. In summary, the client may delete the identifiers, image memory information, and location information of image blocks with enhanced layer hit attributes or newly added attributes indicated by the second attribute information from the fourth buffer pool. This is because image blocks with enhanced layer attributes do not affect the processing of image blocks with attributes in other layers during decoding and reprocessing. Deleting the attribute information does not affect subsequent decoding and reprocessing.

[0141] For example, in the embodiment shown in FIG9 , since the image blocks in the fourth cache pool are enhancement layers, the decoding of other image blocks does not depend on these image blocks. Therefore, after the reprocessing (i.e., filling) is completed, the client can also delete the image block identifiers (including ID4-6) in the fourth cache pool that are related to the reprocessing here.

[0142] In some optional embodiments, considering that image blocks with enhancement layer attributes do not affect the processing of image blocks with attributes in other layers during decoding and reprocessing, the identifier and image content information of the image blocks corresponding to attribute information indicating newly added enhancement layer attributes may not be added to the fourth buffer pool, further reducing memory usage.

[0143] In the present application, since the processing of the enhanced layer image data does not affect the processing of the base layer image data, after the N groups of second image data obtained by decoding the enhanced layer image frame are processed, the client deletes the identifier and image memory information of the image block with the enhanced layer-related attributes corresponding to the second attribute information in the fourth cache pool, which will not affect other image data and can also reduce the storage resources used by the client.

[0144] Based on the foregoing description, it can be seen that in the video data processing method provided in the embodiment of the present application, the cloud desktop server can encode the video data to obtain multiple code streams, and distribute different code streams for different network conditions.

[0145] In the "cloud desktop + cloud conferencing" application scenario shown in Figure 2, the cloud desktop server encodes video data to obtain base layer image frames and enhancement layer image frames. If the network status of the cloud conferencing server and each client is not weak, then the cloud desktop server can send the base layer image frames and enhancement layer image frames. If the cloud conferencing server is in a weak network state, the cloud desktop server can send the base layer image frames to implement the frame extraction function and reduce bandwidth.

[0146] In the collaborative desktop application scenario shown in Figure 3, the cloud desktop server encodes video data to obtain base layer image frames and enhancement layer image frames, and sends the base layer image frames to the client in a weak network state to implement frame extraction and reduce bandwidth. The base layer image frames and enhancement layer image frames are then sent to the client in a normal network state.

[0147] In addition, in the related art, there is also a relationship of mutual reference between local macroblocks of the previous and next frames, which is called Ref (the current frame macroblock fully references a local image of the previous frame) and RefAdd (the current frame macroblock references a local image of the previous frame, and there is also newly added image content). In the technical solution of the present application, by defining the attribute information of the image frame in the video data to be encoded, this relationship is changed from "previous frame" to "previous base layer frame", that is, the enhanced layer frame image cannot participate in the calculation of Ref / RefAdd to prevent the occurrence of scenarios where subsequent frames depend on the enhanced frame. At the same time, in the image frame, the identifier and image content information of the image block related to the base layer attributes can be added to the corresponding cache pool (the first cache pool, the third cache pool); the identifier and image content information related to the enhanced layer attributes can be added to the corresponding cache pool (the second cache pool, the fourth cache pool). In addition, from the perspective of reducing the amount of calculation, the content indicated by the relevant attributes used for encoding or decoding of this frame may not be added to the cache pool.

[0148] Please refer to Figure 10, which is a structural diagram of the client provided by the embodiment of the present application. The client provided by the embodiment of the present application runs on a terminal, the terminal is connected to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the pictures presented by the cloud desktop server; the cloud desktop server is used to perform layered encoding on the video data to be encoded, and obtain a base layer image frame and M groups of enhancement layer image frames. The base layer image frame is independently decoded and is obtained by extracting the image frame obtained by layered encoding. The decoding of the enhancement layer image frame depends on the base layer image frame. The video data to be encoded is obtained by the cloud desktop server processing the image frame based on the attribute information of the image frame included in the video data. The attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block, and M is a positive integer.

[0149] In some optional embodiments, client 1000 includes a transceiver unit 1001 for receiving a base layer image frame and first attribute information of the base layer image frame from a cloud desktop server, wherein the first attribute information indicates attributes of image blocks included in the base layer image frame. A processing unit 1002 is configured to process first image data obtained by decoding the base layer image frame based on a third buffer pool and the first attribute information, wherein the third buffer pool is managed by the client and is configured to store attribute information including image block identifiers and image content information of base layer hit attributes.

[0150] In some optional embodiments, the transceiver unit 1001 is configured to receive a base layer image frame, first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames from a cloud desktop server, where the second attribute information indicates attributes of image blocks included in the N groups of enhancement layer image frames, where N is a positive integer less than or equal to M. The processing unit 1002 is configured to process first image data obtained by decoding the base layer image frame and N groups of second image data obtained by decoding the N groups of enhancement layer image frames based on the third buffer pool, the fourth buffer pool, the first attribute information, and the second attribute information, wherein the fourth buffer pool is managed by the client, and the fourth buffer pool is configured to store attribute information including image block identifiers and image content information of enhancement layer hit attributes.

[0151] In some optional embodiments, the processing unit 1002 is specifically configured to: if the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determine from the third cache pool a second image block having the same identifier as the first image block, and the base layer hit attribute indicates that the corresponding image block is the same as the historical image block; and fill the image content information of the first image block with the image content information of the second image block.

[0152] In some optional embodiments, processing unit 1002 is specifically configured to: if the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determine from the third cache pool a second image block having the same identifier as the first image block, where the base layer hit attribute indicates that the corresponding image block is the same as the historical image block; and fill the image content information of the first image block with the image content information of the second image block. If the second attribute information indicates that the attribute of the third image block in the N sets of second image data is an enhancement layer hit attribute, determine from the fourth cache pool a fourth image block having the same identifier as the third image block, where the enhancement layer hit attribute indicates that the corresponding image block is the same as the historical image block; and fill the image content information of the third image block with the image content information of the fourth image block. And / or, if the second attribute information indicates that the attribute of the fifth image block in the N sets of second image data is a base layer hit attribute, determine from the third cache pool a sixth image block having the same identifier as the fifth image block; and fill the image content information of the fifth image block with the image content information of the sixth image block.

[0153] In some optional embodiments, the processing unit 1002 is further configured to: if the first attribute information indicates that the attribute of the seventh image block in the first image data is a new attribute added to the base layer, then store the identifier and image content information of the seventh image block in the third cache pool. And / or, if the second attribute information indicates that the attribute of the eighth image block in the N sets of second image data is a new attribute added to the base layer, then store the identifier and image content information of the eighth image block in the third cache pool. Wherein, the image block indicated by the base layer enhanced attribute is different from the historical image block

[0154] In some optional embodiments, the processing unit 1002 is further used to: if the second attribute information indicates that the attribute of the ninth image block in the N groups of second image data is a new attribute of the enhancement layer, then the identification and image content information of the ninth image block are stored in the fourth cache pool.

[0155] In some optional embodiments, the processing unit 1002 is further configured to: delete, from the fourth buffer pool, identifiers and image memory information of image blocks with enhanced layer hit attributes or enhanced layer newly added attributes indicated by the second attribute information, wherein the image blocks indicated by the enhanced layer newly added attributes are different from the historical image blocks.

[0156] The client 1000 is used to implement the operations performed by the master client and / or the collaborative client in the embodiments shown in Figures 2 and 3, and the client running on the first terminal and / or the client running on the second terminal in the embodiments shown in Figures 4 to 9, which will not be repeated here.

[0157] Please refer to Figure 11, which is a structural diagram of the terminal provided in an embodiment of the present application. The terminal 1100 includes a processor 1101, a memory 1102, a communication interface 1103 and a bus 1104. Among them, the processor 1101, the memory 1102, the communication interface 1103 communicate through the bus 1104, and can also achieve communication through other means such as wireless transmission. The memory 1102 stores program code, and the processor 1101 can call the program code stored in the memory 1102 to execute the operations performed by the master client and / or the collaborative client in the embodiments shown in Figures 2 and 3, and the client running on the first terminal and / or the client running on the second terminal in the embodiments shown in Figures 4 to 9, which will not be repeated here.

[0158] It should be understood that in the embodiment of the present application, the processor 1101 may be a CPU, or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0159] The memory 1102 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1101. The memory 1102 may also include a non-volatile random access memory. For example, the memory 1102 may also store device type information.

[0160] The memory 1102 may be a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memories. The nonvolatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0161] In addition to the data bus, bus 1104 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus 1104 in the figure. Bus 1140 may be a Peripheral Component Interconnect Express (PCIe) bus, an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), or the like. Bus 1140 may be divided into an address bus, a data bus, a control bus, or the like.

[0162] The terminal 1100 may also include one or more communication interfaces, one or more operating systems, such as Windows Server 2003, ... and Windows Server 2003R. TM , Mac OS X TM, Unix TM ,Linux TM , FreeBSD TM wait.

[0163] Please refer to Figure 12, which is a schematic diagram of the structure of the cloud desktop server provided in an embodiment of the present application. In which, the cloud desktop server 1200 is connected to multiple terminals, and the screens of the multiple terminals display the screens presented by the cloud desktop server 1200.

[0164] In some optional implementations, the cloud desktop server 1200 includes a processing unit 1201 and a transceiver unit 1202 .

[0165] Processing unit 1201 is configured to process image frames included in video data based on their attribute information to obtain video data to be encoded. The attribute information of the image frames indicates whether image blocks included in the image frames are identical to historical image blocks. The video data to be encoded is layered encoded to obtain base layer image frames and M sets of enhancement layer image frames. The base layer image frames are independently decoded and are extracted from the layered encoded image frames. The decoding of the enhancement layer image frames depends on the base layer image frames. M is a positive integer.

[0166] Transceiver unit 1202 is configured to send a base layer image frame and first attribute information of the base layer image frame to a first terminal in a weak network state among multiple terminals, where the first attribute information indicates attributes of image blocks included in the base layer image frame. Send the base layer image frame, the first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames to a second terminal in a non-weak network state among the multiple terminals, where N is a positive integer less than or equal to M.

[0167] In some optional embodiments, the processing unit 1201 is specifically used to: process the image block whose attribute information in the image frame is the basic layer hit attribute or the enhanced layer hit attribute as the background color, and both the basic layer hit attribute and the enhanced layer hit attribute indicate that the corresponding image block is the same as the historical image block.

[0168] In some optional embodiments, the cloud desktop server 1200 includes a first cache pool and a second cache pool, the first cache pool is used to store image block identifiers whose attribute information is a base layer hit attribute or a base layer newly added attribute, and the second cache pool is used to store image block identifiers whose attribute information is an enhancement layer hit attribute or an enhancement layer newly added attribute, and both the base layer newly added attribute and the enhancement layer newly added attribute indicate that the corresponding image block is different from the historical image block.

[0169] In some optional implementations, the processing unit 1201 is further configured to: delete, from the second buffer pool, image block identifiers whose attribute information is enhancement layer hit attributes or enhancement layer newly added attributes in the M groups of enhancement layer image frames.

[0170] The processing unit 1201 and the transceiver unit 1202 can be implemented in software or hardware. For example, the implementation of the processing unit 1201 will be described below using the processing unit 1201 as an example. Similarly, the implementation of the transceiver unit 1202 can refer to the implementation of the processing unit 1201.

[0171] As an example of a software functional unit, the processing unit 1201 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the processing unit 1201 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.

[0172] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.

[0173] As an example of a hardware functional unit, the processing unit 1201 may include at least one computing device, such as a server. Alternatively, the processing unit 1201 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0174] The multiple computing devices included in processing unit 1201 can be distributed in the same region or in different regions. The multiple computing devices included in processing unit 1201 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in processing unit 1201 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.

[0175] It should be noted that the processing unit 1201 and the transceiver unit 1202 respectively implement different steps in the data processing method to realize all functions of the cloud desktop server 1200. The cloud desktop server 1200 is used to perform the operations performed by the cloud desktop server in the embodiments shown in Figures 2 to 9 above to implement the video data processing method applied to the cloud desktop server provided in the embodiments of the present application, and will not be described in detail here.

[0176] Please refer to Figure 13, which is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. The computing device 1300 includes a processor 1301, a communication interface 1302, a bus 1303, and a memory 1304. The processor 1301, the communication interface 1302, and the memory 1304 communicate with each other via the bus 1303. In practical applications, communication can also be achieved through other means such as wireless transmission, which is not limited here.

[0177] The computing device 1300 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1300.

[0178] The processor 1301 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0179] The communication interface 1302 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1300 and other devices or a communication network.

[0180] Bus 1303 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG13 shows a single line, but this does not imply a single bus or type of bus. Bus 1303 may include a path for transmitting information between various components of computing device 1300 (e.g., memory 1304, processor 1301, and communication interface 1302).

[0181] The memory 1304 may include volatile memory, such as random access memory (RAM). The memory 1304 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0182] Memory 1304 stores executable program code. Processor 1301 executes the executable program code to implement the functions of processing unit 1201 and transceiver unit 1202, thereby implementing the video data processing method for a cloud desktop server. In other words, memory 1304 stores instructions for executing the video data processing method for a cloud desktop server.

[0183] Alternatively, the memory 1304 stores executable code, and the processor 1301 executes the executable code to implement the functions of the aforementioned processing unit 1201 and the transceiver unit 1202, thereby implementing the video data processing method applied to the cloud desktop server. In other words, the memory 1304 stores instructions for executing the video data processing method applied to the cloud desktop server.

[0184] The embodiment of the present application further provides a computing device cluster, which includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center.

[0185] Please refer to Figures 14 and 15, which are both structural diagrams of the computing device cluster provided in embodiments of the present application.

[0186] As shown in Figure 14, the computing device cluster includes at least one computing device 1300. The memory 1304 in one or more computing devices 1300 in the computing device cluster may store the same instructions for executing the video data processing method applied to the cloud desktop server provided in the embodiment of the present application.

[0187] In some possible implementations, the memory 1304 of one or more computing devices 1300 in the computing device cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 1304 can jointly execute instructions for executing the video data processing method applied to the cloud desktop server.

[0188] It should be noted that the memory 1304 in different computing devices 1300 in the computing device cluster can store different instructions, each for executing a portion of the functions of the data processing apparatus. In other words, the instructions stored in the memory 1304 in different computing devices 1300 can implement the functions of one or more units of the processing unit 1201 and the transceiver unit 1202.

[0189] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), among others. FIG. 15 illustrates a possible implementation. As shown in FIG. 15 , two computing devices 1300A and 1300B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1304 in the computing device 1300A stores instructions for executing the functions of the transceiver unit 1202. Simultaneously, the memory 1304 in the computing device 1300B stores instructions for executing the functions of the processing unit 1201.

[0190] The connection method between the computing device clusters shown in Figure 15 can be based on the consideration that in the video data processing method applied to the cloud desktop server provided in this application, processing operations and operations other than processing operations are performed separately, that is, the function of the processing unit 1201 is considered to be executed by the computing device 1300B, and the function of the transceiver unit 1202 is considered to be executed by the computing device 1300A.

[0191] It should be understood that the functionality of the computing device 1300A shown in FIG15 may also be implemented by multiple computing devices 1300. Similarly, the functionality of the computing device 1300B may also be implemented by multiple computing devices 1300.

[0192] The present application also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to the connection method of the computing device cluster described in Figures 14 and 15, which will not be repeated here.

[0193] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the aforementioned method for processing video data.

[0194] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned video data processing method.

[0195] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for processing video data, characterized in that: The method is applied to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals are used to display the pictures presented by the cloud desktop server; the method includes: Based on attribute information of an image frame included in the video data, the image frame is processed to obtain video data to be encoded, wherein the attribute information of the image frame is used to indicate whether an image block included in the image frame is the same as a historical image block; Performing hierarchical coding on the video data to be coded, obtaining a base layer image frame and M groups of enhancement layer image frames, wherein the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by hierarchical coding, the decoding of the enhancement layer image frame depends on the base layer image frame, and M is a positive integer; Sending the base layer image frame and first attribute information of the base layer image frame to a first terminal in a weak network state among the multiple terminals, where the first attribute information indicates attributes of image blocks included in the base layer image frame; The base layer image frame, the first attribute information, N groups of enhancement layer image frames and the second attribute information of the N groups of enhancement layer image frames are sent to the second terminal in a non-weak network state among the multiple terminals, wherein the second attribute information indicates the image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M.

2. The method according to claim 1, characterized in that The step of processing the image frame based on the attribute information of the image frame included in the video data comprises: An image block whose attribute information in the image frame is a base layer hit attribute or an enhanced layer hit attribute is processed as a background color, and both the base layer hit attribute and the enhanced layer hit attribute indicate that the corresponding image block is the same as the historical image block.

3. The method according to claim 1 or 2, characterized in that: The cloud desktop server includes a first cache pool and a second cache pool, the first cache pool is used to store image block identifiers whose attribute information is a base layer hit attribute or a base layer newly added attribute, and the second cache pool is used to store image block identifiers whose attribute information is an enhancement layer hit attribute or an enhancement layer newly added attribute, and the base layer newly added attribute and the enhancement layer newly added attribute both indicate that the corresponding image block is different from the historical image block.

4. The method according to claim 3, characterized in that After processing the image frame based on the attribute information of the image frame included in the video data, the method further includes: In the second buffer pool, image block identifiers whose attribute information is enhancement layer hit attributes or enhancement layer newly added attributes in the M groups of enhancement layer image frames are deleted.

5. A method for processing video data, characterized in that: The method is applied to a client, the client runs on a terminal, the terminal is connected to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the images presented by the cloud desktop server; The cloud desktop server is used to perform hierarchical encoding on the video data to be encoded, and obtain a base layer image frame and M groups of enhanced layer image frames, wherein the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by hierarchical encoding, and the decoding of the enhanced layer image frame depends on the base layer image frame, and the video data to be encoded is obtained by the cloud desktop server processing the image frame based on the attribute information of the image frame included in the video data, and the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block, and M is a positive integer; The method comprises: Receiving the base layer image frame and first attribute information of the base layer image frame from the cloud desktop server, wherein the first attribute information indicates attributes of image blocks included in the base layer image frame; Based on a third buffer pool and the first attribute information, processing the first image data obtained by decoding the base layer image frame, wherein the third buffer pool is managed by the client, and the third buffer pool is used to store attribute information including an image block identifier and image content information of a base layer hit attribute; or, Receiving the base layer image frame, the first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames from the cloud desktop server, wherein the second attribute information indicates image block attributes included in the N groups of enhancement layer image frames, where N is a positive integer less than or equal to M; Based on the third cache pool, the fourth cache pool, the first attribute information and the second attribute information, the first image data obtained by decoding the base layer image frame and the N groups of second image data obtained by decoding the N groups of enhancement layer image frames are processed, wherein the fourth cache pool is managed by the client, and the fourth cache pool is used to store attribute information including image block identifiers and image content information of enhancement layer hit attributes.

6. The method according to claim 5, characterized in that The processing, based on the third buffer pool and the first attribute information, of first image data obtained by decoding the base layer image frame includes: If the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block having the same identifier as the first image block from the third cache pool, the base layer hit attribute indicating a historical image block; The image content information of the first image block is filled as the image content information of the second image block.

7. The method according to claim 5, characterized in that The processing, based on the third buffer pool, the fourth buffer pool, the first attribute information, and the second attribute information, of first image data obtained by decoding the base layer image frame and second image data obtained by decoding the N groups of enhancement layer image frames includes: If the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block having the same identifier as the first image block from the third cache pool, the base layer hit attribute indicating that the corresponding image block is the same as the historical image block; Filling the image content information of the first image block with the image content information of the second image block; If the second attribute information indicates that the attribute of the third image block in the N groups of second image data is an enhancement layer hit attribute, determining a fourth image block having the same identifier as the third image block from the fourth cache pool, and the image block corresponding to the enhancement layer hit attribute indication is the same as the historical image block; Filling the image content information of the third image block as the image content information of the fourth image block; and / or, If the second attribute information indicates that the attribute of the fifth image block in the N groups of second image data is a base layer hit attribute, determining a sixth image block having the same identifier as the fifth image block from the third cache pool; The image content information of the fifth image block is filled as the image content information of the sixth image block.

8. The method according to any one of claims 5 to 7, characterized in that The method further comprises: If the first attribute information indicates that the attribute of the seventh image block in the first image data is a newly added attribute of the base layer, the identifier and image content information of the seventh image block are stored in the third cache pool; and / or, If the second attribute information indicates that the attribute of the eighth image block in the N groups of second image data is a new attribute added to the base layer, the identifier and image content information of the eighth image block are stored in the third cache pool; wherein the image block indicated by the base layer enhanced attribute is different from the historical image block.

9. The method according to claim 7, characterized in that: The method further comprises: If the second attribute information indicates that the attribute of the ninth image block in the N groups of second image data is a new attribute of the enhancement layer, the identifier and image content information of the ninth image block are stored in the fourth cache pool, wherein the image block indicated by the new attribute of the enhancement layer is different from the historical image block.

10. The method according to any one of claims 5 to 9, characterized in that The method further comprises: In the fourth buffer pool, the identifier and image memory information of the image block with the enhancement layer hit attribute or the enhancement layer newly added attribute indicated by the second attribute information are deleted.

11. A cloud desktop server, characterized in that: The cloud desktop server is connected to a plurality of terminals, and the screens of the plurality of terminals are used to display the images presented by the cloud desktop server; The cloud desktop server comprises: a processing unit, configured to process the image frame based on attribute information of the image frame included in the video data to obtain the video data to be encoded, wherein the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block; The processing unit is further used to perform hierarchical encoding on the video data to be encoded, to obtain a base layer image frame and M groups of enhancement layer image frames, wherein the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by hierarchical encoding, the decoding of the enhancement layer image frame depends on the base layer image frame, and M is a positive integer; A transceiver unit, configured to send the base layer image frame and first attribute information of the base layer image frame to a first terminal in a weak network state among the multiple terminals, wherein the first attribute information indicates an attribute of an image block included in the base layer image frame; The transceiver unit is also used to send the base layer image frame, the first attribute information, N groups of enhancement layer image frames and second attribute information of the N groups of enhancement layer image frames to a second terminal in a non-weak network state among the multiple terminals, wherein the second attribute information indicates image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M.

12. The cloud desktop server according to claim 11, characterized in that: The processing unit is specifically used for: An image block whose attribute information in the image frame is a base layer hit attribute or an enhanced layer hit attribute is processed as a background color, and both the base layer hit attribute and the enhanced layer hit attribute indicate that the corresponding image block is the same as the historical image block.

13. The cloud desktop server according to claim 11 or 12, characterized in that: The cloud desktop server includes a first cache pool and a second cache pool, the first cache pool is used to store image block identifiers whose attribute information is a base layer hit attribute or a base layer newly added attribute, and the second cache pool is used to store image block identifiers whose attribute information is an enhancement layer hit attribute or an enhancement layer newly added attribute, and the base layer newly added attribute and the enhancement layer newly added attribute both indicate that the corresponding image block is different from the historical image block.

14. The cloud desktop server according to claim 13, characterized in that: The processing unit is further used for: In the second buffer pool, image block identifiers whose attribute information is enhancement layer hit attributes or enhancement layer newly added attributes in the M groups of enhancement layer image frames are deleted.

15. A client, characterized in that: The client runs on a terminal, the terminal is connected to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the images presented by the cloud desktop server; The cloud desktop server is used to perform hierarchical encoding on the video data to be encoded, and obtain a base layer image frame and M groups of enhanced layer image frames, wherein the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by hierarchical encoding, and the decoding of the enhanced layer image frame depends on the base layer image frame, and the video data to be encoded is obtained by the cloud desktop server processing the image frame based on the attribute information of the image frame included in the video data, and the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block, and M is a positive integer; The client comprises: A transceiver unit, configured to receive the base layer image frame and first attribute information of the base layer image frame from the cloud desktop server, wherein the first attribute information indicates an attribute of an image block included in the base layer image frame; a processing unit, configured to process first image data obtained by decoding the base layer image frame based on a third buffer pool and the first attribute information, wherein the third buffer pool is managed by the client, and the third buffer pool is used to store attribute information including an image block identifier and image content information of a base layer hit attribute; Alternatively, the client comprises: a transceiver unit, configured to receive the base layer image frame, the first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames from the cloud desktop server, wherein the second attribute information indicates image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M; A processing unit is used to process the first image data obtained by decoding the base layer image frame and the N groups of second image data obtained by decoding the N groups of enhancement layer image frames based on the third cache pool, the fourth cache pool, the first attribute information and the second attribute information, wherein the fourth cache pool is managed by the client, and the fourth cache pool is used to store attribute information including image block identifiers and image content information of enhancement layer hit attributes.

16. The client according to claim 15, characterized in that: The processing unit is specifically used for: If the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block having the same identifier as the first image block from the third cache pool, the base layer hit attribute indicating a historical image block; The image content information of the first image block is filled as the image content information of the second image block.

17. The client according to claim 15, wherein the processing unit is specifically configured to: If the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block having the same identifier as the first image block from the third cache pool, the base layer hit attribute indicating that the corresponding image block is the same as the historical image block; Filling the image content information of the first image block with the image content information of the second image block; If the second attribute information indicates that the attribute of the third image block in the N groups of second image data is an enhancement layer hit attribute, determining a fourth image block having the same identifier as the third image block from the fourth cache pool, and the image block corresponding to the enhancement layer hit attribute indication is the same as the historical image block; Filling the image content information of the third image block as the image content information of the fourth image block; and / or, If the second attribute information indicates that the attribute of the fifth image block in the N groups of second image data is a base layer hit attribute, determining a sixth image block having the same identifier as the fifth image block from the third cache pool; The image content information of the fifth image block is filled as the image content information of the sixth image block.

18. The client according to any one of claims 15 to 17, characterized in that: The processing unit is also used to: If the first attribute information indicates that the attribute of the seventh image block in the first image data is a newly added attribute of the base layer, storing the identifier and image content information of the seventh image block in the third cache pool; and / or, If the second attribute information indicates that the attribute of the eighth image block in the N groups of second image data is a new attribute added to the base layer, the identifier and image content information of the eighth image block are stored in the third cache pool; wherein the image block indicated by the base layer enhanced attribute is different from the historical image block.

19. The client according to claim 17, characterized in that: The processing unit is further used for: If the second attribute information indicates that the attribute of the ninth image block in the N groups of second image data is a newly added attribute of the enhancement layer, the identifier and image content information of the ninth image block are stored in the fourth cache pool.

20. The client according to any one of claims 15 to 19, characterized in that: The processing unit is also used to: In the fourth cache pool, the identifier and image memory information of the image block with the enhanced layer hit attribute or the enhanced layer newly added attribute indicated by the second attribute information are deleted, wherein the image block indicated by the enhanced layer newly added attribute is different from the historical image block.

21. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 4.

22. A terminal, characterized in that: comprising a processor coupled to a memory; Instructions are stored in the memory, and when the instructions are executed on the processor, the method according to any one of claims 5 to 10 is implemented.

23. A computer program product comprising instructions, characterized in that When the instruction is executed by a computing device cluster, the computing device cluster executes the method as described in any one of claims 1 to 4; or, when the instruction is executed by a terminal, the terminal executes the method as described in any one of claims 5 to 10.

24. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes computer program instructions, and when the computer program instructions are executed by a computing device cluster, the method according to any one of claims 1 to 4 is implemented; or when the computer program instructions are executed by a terminal or a client, the method according to any one of claims 5 to 10 is implemented.

Citation Information

Patent Citations

  • Image transmission method and apparatus for virtual desktop

    CN107145340A

  • Image data sending method and device and related parts

    CN111831366A

  • Content caching method based on multi-network channel transmission

    CN116828268A

  • Predictive bit-plane coding for progressive fine-granularity scalable (PFGS) video coding

    EP1511324A1

  • Virtual desktop infrastructure server, computer implemented video streaming method, and non-transitory computer readable storage medium thereof

    US20150149593A1