Video data processing method and related equipment
By layering the video data in the cloud desktop service and adjusting the transmission of image frames according to the terminal network status, the problem of high bandwidth occupation of video data in the cloud desktop service is solved, and efficient video data transmission in different network environments is achieved.
Patent Information
- Application Number
- CN202410263180.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
In cloud desktop services, the high frame rate and high bandwidth usage of video data lead to excessive network bandwidth consumption, especially in screen sharing scenarios, which are difficult to effectively manage.
By layering the video data on the cloud desktop server, the basic layer image frame and the enhancement layer image frame are generated, and different image frames are sent according to the different network status of the terminal. For weak network terminals, only basic layer image frames are sent to reduce bandwidth consumption; for non-weak network terminals, basic layer image frames and enhancement layer image frames are sent to ensure high frame rates and smooth video playback.
It effectively reduces the bandwidth consumption of video data, especially in weak network environments, avoids the problem of limited video data transmission, and can still ensure high frame rate and smooth video playback in non-weak network environments.
Smart Images

Figure CN120201211A_ABST
Abstract
Description
[0001] This application claims the priority of the Chinese patent application filed with the State Intellectual Property Office on December 22, 2023, with application number 202311791366.0 and invention name “A method and device for deduplication of remote desktop image data”, the entire contents of which are incorporated by reference in this application. Technical Field
[0002] The present application relates to the field of cloud computing, and in particular to a method for processing video data and related equipment. Background Art
[0003] With the development of computer technology, cloud desktop services have been widely used. Cloud desktop services are desktop services based on cloud computing. Users can realize desktop virtualization by logging in to the purchased cloud desktop services through terminals. In addition, cloud desktop services can also realize screen sharing between multiple terminals, and display the same desktop content on different terminals.
[0004] In the related technical solutions, for the transmission of video data, there is a strong dependency between the video frames to be encoded processed by the cloud desktop server, and frame extraction operations are not supported, which makes the number of frames per second (fps) of the video data transmitted by the cloud desktop server to the client high, resulting in increased bandwidth occupancy. Summary of the invention
[0005] The present application provides a video data processing method and related equipment for reducing bandwidth.
[0006] In the first aspect, the present application provides a method for processing video data, which is applied to a cloud desktop server, where the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the screen presented by the cloud desktop server, that is, multiple terminal screens are shared. In other words, the video processing method provided by the present application is applied in the screen sharing scenario of the cloud desktop service. The video processing method includes:
[0007] The cloud desktop server processes an image frame based on the attribute information included in the video data to obtain the video data to be encoded. Among them, the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block. That the image block is the same as the historical image block means that the information included in the image block is the same as the information included in the historical image block, that is, the picture displayed by the image block is the same as the picture displayed by the historical image block. The historical image block refers to the image block obtained by the cloud desktop server processing the historical video data. By performing hierarchical encoding on the video data to be encoded, a base layer image frame and M groups of enhancement layer image frames are obtained, where M is a positive integer. Specifically, the cloud desktop server can perform temporal hierarchical encoding on the video data to be encoded based on the temporal scalability included in the scalability of video encoding. Among them, the base layer image frame retains the information of the lowest frame rate and can be decoded independently. The decoding of the enhancement layer image frame depends on the base layer image frame. After decoding, the enhancement layer data plus the decoded base layer data has a higher frame rate. In other words, the M groups of enhancement layer image frames of the base layer image frame are all obtained by decimating the image frames obtained after hierarchical encoding of the video data. That is, whether it is the base layer image frame or the enhancement layer image frame, they are all partial image frames in the video data to be encoded. The cloud desktop server sends different image frames to different terminals according to the different network states of multiple terminals. For example, the base layer image frame and the first attribute information of the base layer image frame are sent to the first terminal in a weak network state among multiple terminals, and the first attribute information indicates the image block attributes included in the base layer image frame. The base layer image frame, the first attribute information, N groups of enhancement layer image frames, and the second attribute information of the N groups of enhancement layer image frames are sent to the second terminal in a non-weak network state among multiple terminals. The second attribute information indicates the image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M.
[0008] In this application, after the cloud desktop server performs hierarchical encoding on the video data to be encoded to obtain a base layer image frame and M groups of enhancement layer image frames, it sends different image frames to terminals with different network states. For the first terminal in a weak network state, the cloud desktop server sends the base layer image frame, which is obtained by decimating the video data to be encoded, realizing the decimation function and reducing the bandwidth consumption. In addition, for the second terminal in a non-weak network state, the cloud desktop server sends the base layer image frame and the enhancement layer image frame, which neither burdens the transmission bandwidth nor ensures that the second terminal obtains video data with a higher frame rate and displays a smoother video.
[0009] In some alternative implementations of the first aspect, the cloud desktop server processes the image frames based on the attribute information of the image frames included in the video data, including: processing the image blocks in the image frames with the attribute information of the base layer hit attribute or the enhanced layer hit attribute into the background color. This is because both the base layer hit attribute and the enhanced layer hit attribute indicate that the corresponding image blocks are the same as the historical image blocks, that is, these image blocks are the image blocks that have been transmitted before and can be encoded no more.
[0010] In this application, based on the attribute information of the image frames included in the video data, the cloud desktop server can process the image blocks in the image frames that are the same as the historical image blocks into the background color, reducing the image content for video encoding and further reducing the video encoding bandwidth.
[0011] In some alternative implementations of the first aspect, the cloud desktop server includes a first cache pool and a second cache pool. The first cache pool is used to store the image block identifiers with the attribute information of the base layer hit attribute or the base layer new attribute, and the second cache pool is used to store the image block identifiers with the attribute information of the enhanced layer hit attribute or the enhanced layer new attribute. Both the base layer new attribute and the enhanced layer new attribute indicate that the corresponding image blocks are different from the historical image blocks. Then, when the desktop server preprocesses the image frames included in the video data, it can compare with the first cache pool and / or the second cache pool according to the attribute information of the image frames included in the video data to determine whether the image blocks in the image frames included in the video data are the same as the historical image blocks.
[0012] In this application, storing the image block identifiers with different attributes in the cache pool of the cloud desktop server provides a basis for the preprocessing of the image frames included in the video data and provides technical support for the implementation of the technical solution of this application.
[0013] In some alternative implementations of the first aspect, when the attribute information of the image frames in the video data indicates that the attribute of the target image block in the image frame is the base layer new attribute, the identifier of the target image block is stored in the first cache pool.
[0014] In this application, the cloud desktop server can also update the first cache pool to provide a basis for the preprocessing of subsequent image data, which is beneficial to identifying the historical image blocks in the subsequent image data, thereby reducing the content that needs to be encoded.
[0015] In some alternative implementations of the first aspect, after processing the image frames based on the attribute information of the image frames included in the video data, the cloud desktop server can also process the second cache pool. Specifically, the cloud desktop server deletes the image block identifiers with the attribute information of the enhanced layer hit attribute or the enhanced layer new attribute in the M groups of enhanced layer image frames in the second cache pool.
[0016] In this application, since encoding the enhanced layer image data does not affect the encoding of the base layer image data, after preprocessing the image frames included in the video data, the cloud desktop server deletes the image block identifiers with enhanced layer attributes in the image frames of the video data in the second cache pool, which will not affect other image data and can also reduce the storage resources used by the cloud desktop server.
[0017] In a second aspect, this application provides a method for processing video data, which is applied to a client, and the client runs on a terminal. The terminal is connected to the cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the pictures presented by the cloud desktop server. The cloud desktop server is used to perform hierarchical encoding on the video data to be encoded, obtaining a base layer image frame and M groups of enhanced layer image frames. The base layer image frame can be decoded independently and is obtained by extracting frames from the image frames obtained by hierarchical encoding. The decoding of the enhanced layer image frames depends on the base layer image frame. The video data to be encoded is processed by the cloud desktop server based on the attribute information of the image frames included in the video data, and the attribute information of the image frames is used to indicate whether the image blocks included in the image frames are the same as the historical image blocks, where M is a positive integer.
[0018] The client receives the base layer image frame and the first attribute information of the base layer image frame from the cloud desktop server, and the first attribute information indicates the image block attributes included in the base layer image frame. The base layer image frame retains the information of the lowest frame rate and can be decoded independently. The client manages a third cache pool, and the third cache pool is used to store the image block identifiers and image content information whose attribute information includes the base layer hit attribute. The client processes the first image data obtained by decoding the base layer image frame based on the third cache pool and the first attribute information.
[0019] Alternatively, the client receives the base layer image frame, the first attribute information, N groups of enhanced layer image frames, and the second attribute information of the N groups of enhanced layer image frames from the cloud desktop server. The second attribute information indicates the image block attributes included in the N groups of enhanced layer image frames, and N is a positive integer less than or equal to M. The decoding of the enhanced layer image frames depends on the base layer image frame, and the decoded enhanced layer data plus the decoded base layer data has a higher frame rate. The client manages a fourth cache pool, and the fourth cache pool is used to store the image block identifiers and image content information whose attribute information includes the enhanced layer hit attribute. The client processes the first image data obtained by decoding the base layer image frame and the N groups of second image data obtained by decoding the N groups of enhanced layer image frames based on the third cache pool, the fourth cache pool, the first attribute information, and the second attribute information.
[0020] In this application, the client may receive different types of image frames sent by the cloud desktop server. The client manages the relevant information for storing image blocks with different attributes. Then, corresponding processing can be performed for different types of image frames, improving the practicality of the technical solution of this application.
[0021] In some optional implementation manners of the second aspect, regardless of whether the terminal is in a weak network state or a non-weak network state, the client can receive the base layer image frame and the first attribute information. The client processes the first image data based on the third cache pool and the first attribute information, including: if the first attribute information indicates that the attribute of the first image block in the first image data is the base layer hit attribute, then the client determines a second image block with the same identifier as the first image block from the third cache pool, and fills the image content information of the first image block with the image content information of the second image block. Among them, the base layer hit attribute indicates that the corresponding image block is the same as the historical image block. That is to say, for the image block in the base layer image frame that is the same as the historical image block, when the client decodes, this image block is actually the background color, reducing the decoding bandwidth. After decoding, it can be restored by filling with the image content information stored in the third cache pool according to the first attribute information. Additionally, it should be noted that if the terminal device is in a weak network state and only receives the base layer image frame and does not receive the enhancement layer image frame, then after the client processes the first image data, the video data to be displayed is obtained and can be displayed through the display device of the terminal.
[0022] In this application, for the base layer image frame, after the client decodes, filling the historical image block according to the first attribute information can obtain a complete image frame. That is to say, during the decoding process, this historical image block is actually the background color, reducing the decoding bandwidth.
[0023] In some alternative implementations of the second aspect, the process of the client processing the first image data based on the third cache pool and the first attribute information is similar to the foregoing and will not be elaborated here. The client also processes the N groups of second image data according to the fourth cache pool and the second attribute information, or processes the N groups of second image data according to the third cache pool, the fourth cache pool, and the second attribute information. This depends on the content of the second attribute information. If the attributes of the image blocks indicated by the second attribute information include enhanced layer hit attributes, then the client compares the second image data with the fourth cache pool. If the attributes of the image blocks indicated by the second attribute information include base layer hit attributes, then the client compares the second image data with the third cache pool. That is to say, when the second attribute information indicates that the attribute of the third image block in the N groups of second image data is an enhanced layer hit attribute, the client determines a fourth image block with the same identifier as the third image block from the fourth cache pool, and fills the image content information of the third image block with the image content information of the fourth image block. And / or, when the second attribute information indicates that the attribute of the fifth image block in the N groups of second image data is a base layer hit attribute, the client determines a sixth image block with the same identifier as the fifth image block from the third cache pool, and fills the image content information of the fifth image block with the image content information of the sixth image block. Among them, both the enhanced layer hit attribute and the base layer hit attribute indicate that the corresponding image block is the same as the historical image block. That is to say, for an image block in the enhanced layer image frame that is the same as the historical image block, when the client decodes, this image block is actually the background color, reducing the decoding bandwidth. After decoding, it can be restored by filling with the image content information stored in the third cache pool and / or the fourth cache pool according to the second attribute information. Additionally, it should be noted that if the terminal device is in a non-weak network state and receives the base layer image frame and N groups of enhanced layer image frames, then after the client processes the first image data and the N groups of second image data, the video data to be displayed is obtained and can be displayed through the display device of the terminal.
[0024] In this application, for the enhanced layer image frame, after the client decodes and fills the historical image block according to the second attribute information, a complete image frame can be obtained. That is to say, during the decoding process, this historical image block is actually the background color, reducing the decoding bandwidth.
[0025] In some optional implementations of the second aspect, the client may process the third cache pool. Specifically, if the first attribute information indicates that the attribute of the seventh image block in the first image data is a newly added attribute of the base layer, then the terminal stores the identifier and image content information of the seventh image block in the third cache pool. And / or, if the second attribute information indicates that the attribute of the eighth image block in the N groups of second image data is a newly added attribute of the base layer, then the terminal stores the identifier and image content information of the eighth image block in the third cache pool. Among them, the image block indicated by the enhanced attribute of the base layer is different from the historical image block. In general, the client stores the identifier and image content information of the image block with the attribute information as the newly added attribute of the base layer in the third cache pool.
[0026] In the present application, the client can also update the third cache pool to provide a basis for the processing of subsequent image data, which is conducive to identifying historical image blocks in subsequent image data and splicing the decoded images into a complete image, thereby improving the practicality of the technical solution of the present application.
[0027] In some optional implementations of the second aspect, the client may process the fourth cache pool. Specifically, if the second attribute information indicates that the attribute of the ninth image block in the N groups of second image data is a newly added attribute of the enhancement layer, the client stores the identifier and image content information of the ninth image block in the fourth cache pool.
[0028] In some optional implementations of the second aspect, after the N groups of second image data obtained by decoding the enhanced layer image frame are processed, that is, after the complete enhanced layer image data is obtained, the client may update the fourth buffer pool, including deleting the identifier and image memory information of the image block of the enhanced layer hit attribute or the enhanced layer newly added attribute indicated by the second attribute information in the fourth buffer pool.
[0029] In the present application, since the processing of the enhanced layer image data does not affect the processing of the base layer image data, after the processing of the N groups of second image data obtained by decoding the enhanced layer image frame is completed, the client deletes the identifier and image memory information of the image block with the enhanced layer related attributes corresponding to the second attribute information in the fourth cache pool, which will not affect other image data and can also reduce the storage resources used by the terminal.
[0030] In a third aspect, the present application provides a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the images presented by the cloud desktop server; the cloud desktop server includes:
[0031] The processing unit is used to process the image frame based on the attribute information of the image frame included in the video data to obtain the video data to be encoded, wherein the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block. The video data to be encoded is layered encoded to obtain a base layer image frame and M groups of enhancement layer image frames, wherein the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by layered encoding, and the decoding of the enhancement layer image frame depends on the base layer image frame, and M is a positive integer.
[0032] A transceiver unit is configured to send a base layer image frame and first attribute information of the base layer image frame to a first terminal in a weak network state among multiple terminals, wherein the first attribute information indicates an attribute of an image block included in the base layer image frame. A base layer image frame, the first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames are sent to a second terminal in a non-weak network state among multiple terminals, wherein the second attribute information indicates an attribute of an image block included in the N groups of enhancement layer image frames, wherein N is a positive integer less than or equal to M.
[0033] The cloud desktop server is used to implement the method shown in the aforementioned first aspect or any possible implementation of the first aspect. Its beneficial effects are similar to those of the aforementioned first aspect or any possible implementation of the first aspect, and will not be repeated here.
[0034] In a fourth aspect, the present application provides a client, which runs on a terminal, the terminal is connected to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the pictures presented by the cloud desktop server; the cloud desktop server is used to perform layered encoding on the video data to be encoded to obtain a base layer image frame and M groups of enhancement layer image frames, the base layer image frames are independently decoded and are obtained by extracting the image frames obtained by layered encoding, the decoding of the enhancement layer image frames depends on the base layer image frames, the video data to be encoded is obtained by the cloud desktop server processing the image frames based on the attribute information of the image frames included in the video data, the attribute information of the image frames is used to indicate whether the image blocks included in the image frames are the same as the historical image blocks, and M is a positive integer.
[0035] The client includes: a transceiver unit, which is used to receive a base layer image frame and first attribute information of the base layer image frame from a cloud desktop server, wherein the first attribute information indicates an image block attribute included in the base layer image frame. A processing unit, which is used to process first image data obtained by decoding the base layer image frame based on a third buffer pool and the first attribute information, wherein the third buffer pool is managed by the client, and the third buffer pool is used to store attribute information including an image block identifier and image content information of a base layer hit attribute;
[0036] Alternatively, the client includes a transceiver unit configured to receive a base layer image frame, first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames from a cloud desktop server, where the second attribute information indicates the image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M. A processing unit is configured to process first image data obtained by decoding the base layer image frame and N groups of second image data obtained by decoding the N groups of enhancement layer image frames based on a third cache pool, a fourth cache pool, the first attribute information, and the second attribute information, where the fourth cache pool is managed by the client and is used to store the image block identifiers and image content information whose attribute information includes enhancement layer hit attributes.
[0037] The client is used to implement the method shown in the foregoing second aspect or any possible implementation manner of the second aspect, and its beneficial effects are similar to those of the foregoing second aspect or any possible implementation manner of the second aspect, and will not be elaborated here.
[0038] In a fifth aspect, the present application provides a computing device cluster including at least one computing device, and each computing device includes a processor and a memory; the processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device so that the computing device cluster implements the method disclosed in the first aspect or any possible implementation manner of the first aspect. Its beneficial effects are similar to those of the foregoing first aspect or any possible implementation manner of the first aspect, and will not be elaborated here.
[0039] In a sixth aspect, the present application provides a terminal including a processor and a memory, and the processor stores instructions. When the instructions stored in the memory run on the processor, the method shown in the foregoing second aspect or any possible implementation manner of the second aspect is implemented. Its beneficial effects are similar to those of the foregoing second aspect or any possible implementation manner of the second aspect, and will not be elaborated here.
[0040] In a seventh aspect, the present application provides a computer program product including instructions. When the instructions are run on a computing device cluster, the computing device cluster is enabled to implement the method disclosed in the first aspect and any possible implementation manner of the first aspect. Alternatively, when the instructions are run on a terminal, the method shown in the foregoing second aspect or any possible implementation manner of the second aspect is implemented.
[0041] In an eighth aspect, the present application provides a computer-readable storage medium in which computer program instructions are stored. When the computer program instructions are executed by a computing device cluster, the method shown in the foregoing first aspect or any possible implementation manner of the first aspect is implemented. Alternatively, when the computer program instructions are executed by a terminal or a client, the method shown in the foregoing second aspect or any possible implementation manner of the second aspect is implemented.
[0042] The beneficial effects shown in any one of the seventh aspect and the eighth aspect are similar to those in the first aspect, any possible implementation manner of the first aspect, the second aspect, or any possible implementation manner of the second aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 FIG. is a schematic diagram of a system architecture provided by an embodiment of the present application;
[0044] Figure 2 FIG. is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0045] Figure 3 FIG. is another schematic diagram of an application scenario provided by an embodiment of the present application;
[0046] Figure 4 FIG. is a schematic flowchart of a method for processing video data provided by an embodiment of the present application;
[0047] Figure 5 FIG. is another schematic flowchart of a method for processing video data provided by an embodiment of the present application;
[0048] Figure 6 FIG. is another schematic flowchart of a method for processing video data provided by an embodiment of the present application;
[0049] Figure 7 FIG. is another schematic flowchart of a method for processing video data provided by an embodiment of the present application;
[0050] Figure 8 FIG. is another schematic flowchart of a method for processing video data provided by an embodiment of the present application;
[0051] Figure 9 FIG. is another schematic flowchart of a method for processing video data provided by an embodiment of the present application;
[0052] Figure 10 FIG. is a schematic diagram of a structure of a client provided by an embodiment of the present application;
[0053] Figure 11 FIG. is another schematic diagram of a structure of a terminal provided by an embodiment of the present application;
[0054] Figure 12 FIG. is a schematic diagram of a structure of a cloud desktop server provided by an embodiment of the present application;
[0055] Figure 13 FIG. is a schematic diagram of a structure of a computing device provided by an embodiment of the present application;
[0056] Figure 14A schematic structural diagram of a computing device cluster provided by an embodiment of the present application;
[0057] Figure 15 Another schematic structural diagram of a computing device cluster provided by an embodiment of the present application. Detailed implementation manners
[0058] An embodiment of the present application provides a method for processing video data to reduce bandwidth.
[0059] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will understand that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0060] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that these terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not necessarily have to be limited to those units, but may include other units not clearly listed or inherent to these process, methods, products or devices. Additionally, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or similar expressions below refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0061] First, please refer to Figure 1 , Figure 1 A schematic diagram of the system architecture provided by an embodiment of the present application.
[0062] As Figure 1 shown, the tenant logs in to the cloud platform 30 via the client 10 through the Internet 20 using the account and password registered on the cloud platform 30. The cloud platform 30 manages the infrastructure, and the infrastructure includes multiple data centers set in different regions (regions). Among them, at least one cloud data center is set in each region. For exampleFigure 1 The shown Region 1 includes Cloud Data Center 1 and Cloud Data Center 2, and Region 2 includes Cloud Data Center 3 and Cloud Data Center 4. Each cloud data center is provided with multiple servers, and business instances (including at least one of virtual machines, containers, and dedicated hosts) are running on the servers.
[0063] The cloud platform 30 can provide interfaces related to cloud computing services, such as a configuration page (i.e., an interface) or an API for tenants to access cloud services. The cloud service applied in this application is a cloud desktop service. When a tenant uses this service, the desktop displayed on the client 10 is the cloud desktop provided by the cloud desktop service. In addition, multiple terminals can use the same cloud service simultaneously to achieve screen sharing.
[0064] Next, a simple description of the application scenario of the video data processing method provided by the embodiments of this application will be given. Please refer to Figure 2 and Figure 3 , Figure 2 and Figure 3 Both are schematic diagrams of the application scenarios provided by the embodiments of this application.
[0065] Figure 2 The shown scenario is a "cloud desktop + cloud conference" scenario, and a conference software is running within the business instance of the cloud desktop server. In the traditional technical solution, for a multi-party collaboration (i.e., screen sharing) scenario, after the cloud desktop server encodes the video data, the conference software will perform secondary encoding. Then, the client also needs to perform multiple decodings to be able to display.
[0066] Figure 3 In the shown "point-to-multipoint" or "1 to N" scenario, in the traditional technical solution, after the business instance running within the cloud desktop server encodes the video data, it is transmitted in a broadcast form to the master client and the collaborating clients. That is to say, the data received by each client is the same, which may exceed the bandwidth of terminals in a weak network state, resulting in limited transmission of video data.
[0067] When applying the video data processing method provided by the embodiments of this application to the foregoing scenarios, when encoding the video data, the server side can perform hierarchical encoding based on the attribute information of the image frames included in the video data to obtain multiple bitstreams, adapting to terminals in different network states, thereby achieving the effect of reducing bandwidth. Next, a detailed description will be given in combination with the schematic diagrams:
[0068] Please refer to Figure 4 , Figure 4 is a schematic flowchart of the video data processing method provided by the embodiments of this application, including:
[0069] 401. The cloud desktop server processes the image frames based on the attribute information of the image frames included in the video data to obtain the video data to be encoded. The attribute information of the image frames is used to indicate whether the image blocks included in the image frames are the same as the historical image blocks.
[0070] In the embodiments of the present application, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the pictures presented by the cloud desktop server. That is to say, the screens of the multiple terminals connected to the cloud desktop server are shared, and the multiple terminals are in a collaborative state.
[0071] Before encoding the image frames included in the video data, the cloud desktop server preprocesses the image frames based on the attribute information of each image frame. The attribute information of the image frames indicates whether the image blocks in the image frames are the same as the historical image blocks. Here, the so-called "same" means that the information included in the image block is the same as the information included in the historical image block, that is, the picture displayed by the image block is the same as the picture displayed by the historical image block. In addition, the historical image block refers to the image block obtained by the cloud desktop server processing the historical video data. When the image block is the same as the historical image block, it means that the information included in the image block has been encoded and transmitted. Then, when the cloud desktop server processes the video data to be encoded, it can not encode the image block, thereby reducing the encoding bandwidth.
[0072] In the field of video coding, video data can include various types of image frames. Some image frames can be independently decoded, and the decoding of some image frames depends on other image frames. For example, usually, in the video data with a 2-layer structure, a group of pictures (GOP) can include an inter frame (I frame), a predictive-coded picture frame, and a large P frame. Among them, the attributes of the image blocks included in the I frame and the large P frame are base layer-related attributes, and the attributes of the image blocks included in the small P frame are enhancement layer-related attributes. The image blocks with base layer-related attributes can be independently decoded, and the decoding of the image blocks with enhancement layer-related attributes depends on the image blocks with base layer attributes.
[0073] Generally speaking, processing the image frames based on the attribute information of the image frames included in the video data includes: processing the image blocks with the base layer hit attribute or the enhancement layer hit attribute in the image frames as the background color, and both the base layer hit attribute and the enhancement layer hit attribute indicate that the corresponding image blocks are the same as the historical image blocks.
[0074] The following combines the schematic diagrams to detail the process of processing the image frames based on the attribute information. Please refer to Figures 5 to 7 , Figures 5 to 7 All are the schematic flowcharts of the video data processing method provided by the embodiments of the present application.
[0075] Among them, Figures 5 to 7 in the illustrated embodiments, the preprocessing of a single image frame is taken as an example. Based-add represents the newly added attributes of the base layer, Based-hit represents the hit attributes of the base layer, SVC-add represents the newly added attributes of the enhancement layer, and SVC-hit represents the hit attributes of the base layer.
[0076] In Figure 5 the illustrated embodiments, a single-frame image is divided into 6 image blocks. Among them, the attribute information of the image blocks with the image block identifiers ID4, ID5, and ID6 is the hit attribute of the base layer, which means that these image blocks are the same as the historical image blocks. The cloud desktop server processes these image blocks into the background color to obtain Figure 5 the preprocessed image frame shown, where the gray part in the figure represents the background color.
[0077] In Figure 6 the illustrated embodiments, a single-frame image is divided into 6 image blocks. Among them, the attribute information of the image blocks with the image block identifiers ID4, ID5, and ID6 is the hit attribute of the enhancement layer, which means that these image blocks are the same as the historical image blocks. The cloud desktop server processes these image blocks into the background color to obtain Figure 6 the preprocessed image frame shown, where the gray part in the figure represents the background color.
[0078] Whether Figure 5 or Figure 6 in the illustrated embodiments, the attributes of each image block in a single image frame are related attributes of the same layer. In practical applications, a single image frame may also include related attributes of different layers. Exemplarily, as Figure 7 shown, a single-frame image is divided into 6 image blocks. Among them, the attribute information of the image blocks with the image block identifiers ID1, ID2, and ID3 is the related attributes of the base layer, and the attribute information of the image blocks with the image block identifiers ID4, ID5, and ID6 is the related attributes of the enhancement layer. In addition, the attribute information of the image block with the image block identifier ID3 is the hit attribute of the base layer, and the attribute information of the image blocks with the image block identifiers ID5 and ID6 is the hit attribute of the enhancement layer, indicating that these image blocks are the same as the historical image blocks. The cloud desktop server processes these image blocks into the background color to obtain Figure 6 the preprocessed image frame shown, where the gray part in the figure represents the background color.
[0079] It should be noted that Figures 5 to 7 in the illustrated embodiments, a single-frame image is divided into 6 image blocks as an example. In practical applications, each image frame may also be divided into more or fewer image blocks, and the sizes of the respective image blocks may be all or partially the same, or may all be different, which is not limited herein.
[0080] Based on the foregoing description ofFigures 5 to 7 As can be seen from the description, in the embodiments of the present application, based on the attribute information of the image frame, the historical image blocks in the image frame can be processed into the background color, reducing the image content for video coding and reducing the bandwidth of video coding.
[0081] 402. The cloud desktop server performs hierarchical coding on the video data to be encoded, obtaining a base layer image frame and M groups of enhancement layer image frames. The base layer image frame can be independently decoded and is obtained by extracting frames from the image frames obtained by hierarchical coding. The decoding of the enhancement layer image frames depends on the base layer image frame, where M is a positive integer.
[0082] After the cloud desktop server obtains the video data to be encoded, it can perform temporal hierarchical coding on the video data to be encoded according to the temporal scalability in video coding scalability, obtaining a base layer image frame and M groups of enhancement layer image frames. It can be understood that temporal hierarchical coding includes a frame extraction operation. Then both the base layer image frame and the M groups of enhancement layer image frames are partial image frames in the video data to be encoded, or rather, both the base layer image frame and the M groups of enhancement layer image frames are image frames obtained by extracting frames from the image frames obtained by performing hierarchical coding on the video data to be encoded. Among them, the base layer image frame retains the information of the lowest frame rate and can be independently decoded. The decoding of the enhancement layer image frames depends on the base layer image frame, and the decoded enhancement layer data plus the decoded base layer data has a higher frame rate.
[0083] In addition, the attribute information of the image blocks included in the base layer image frame are all base layer-related attributes, and the attribute information of the image blocks included in the enhancement layer image frames includes enhancement layer-related attributes, or enhancement layer-related attributes and base layer-related attributes, which can be determined according to the dependency relationship between the image blocks in the enhancement layer image frame and the image blocks of other frames.
[0084] 403. The cloud desktop server sends the base layer image frame and the first attribute information of the base layer image frame to the first terminal in a weak network state, where the first attribute information indicates the image block attributes included in the base layer image frame.
[0085] The cloud desktop server can determine whether the terminal connected to the cloud desktop server is in a weak network state through parameters such as network round-trip time (RTT), network jitter, or packet loss rate. The type, quantity, and thresholds of each parameter used to determine whether the terminal is in a weak network or non-weak network state can be defined according to actual service requirements, network conditions, etc., and are not specifically limited here.
[0086] Optionally, a weak network state can be defined based on multiple types of parameters. For example, a network state where the RTT is greater than 100 ms, the network jitter range exceeds ±10 ms, and the packet loss rate is greater than 15% is defined as a weak network state. A network state outside this range is defined as a non-weak network state.
[0087] Optionally, a weak network state can be defined based on a single type of parameter. For example, a network state with an RTT greater than 80 ms is defined as a weak network state. A network state outside this range is a non-weak network state.
[0088] The first terminal in the weak network state has less bandwidth resources. The cloud desktop server sends the base layer image frame and the first attribute information of the base layer image frame to the first terminal. Since the base layer image frame retains the information of the lowest frame rate and can be independently decoded, the first terminal can still decode smooth video data within the limited bandwidth resources. In other words, the base layer image frame sent by the cloud desktop server to the first terminal is actually the result of frame extraction or frame dropping from the encoded complete video data.
[0089] 404. The cloud desktop server sends the base layer image frame, the first attribute information, N groups of enhancement layer image frames, and the second attribute information of the N groups of enhancement layer image frames to the second terminal in the non-weak network state. The second attribute information indicates the image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M.
[0090] For the second terminal in the non-weak network state, the cloud desktop server not only sends the base layer image frame and the first attribute information of the base layer image frame, but also sends N groups of enhancement layer image frames and the second attribute information of the N groups of enhancement layer image frames. Since the decoding of the enhancement layer image frame depends on the base layer image frame, and the decoded enhancement layer data plus the decoded base layer data has a higher frame rate, the video decoded by the second terminal is smoother than the video decoded by the first terminal.
[0091] In addition, among the second terminals in the non-weak network state, the network states of each terminal may be different, and the number of enhancement layer image frames sent by the cloud desktop server can also be different. Exemplarily, assume that the network state of terminal A in the second terminal is better than the network state of terminal B in the second terminal. Then the number of groups of enhancement layer image frames sent by the cloud desktop server to terminal A is greater than the number of groups of enhancement layer image frames sent by the cloud desktop server to terminal B.
[0092] Exemplarily, assume that a network state with an RTT greater than 100 ms, a network jitter range exceeding ±10 ms, and a packet loss rate greater than 15% is defined as a weak network state, and a network state outside this range is a non-weak network state. And the cloud desktop server performs hierarchical encoding on the video data to be encoded, obtaining the base layer image frame and 3 groups of enhancement layer image frames, that is, M = 3.
[0093] If the RTT of terminal 1 is 150 ms, the jitter range is ±15 ms, and the packet loss rate is 16%. The RTT of terminal 2 is 80 ms, the jitter range is ±8 ms, and the packet loss rate is 10%. The RTT of terminal 3 is 50 ms, the jitter range is ±5 ms, and the packet loss rate is 5%. That is, terminal 1 is in a weak network state, both terminal 2 and terminal 3 are in a non-weak network state, and the network state of terminal 3 is better than that of terminal 2.
[0094] Then, the cloud desktop server can send the base layer image frame and the first attribute information to terminal 1; send the base layer image frame, the first attribute information, 1 group of enhancement layer image frames and the second attribute information corresponding to this group of enhancement layer image frames to terminal 2; not only send the base layer image frame and the first attribute information to terminal 3, but also send 2 groups or 3 groups of enhancement layer image frames and the corresponding second attribute information.
[0095] It should be noted that no matter how many groups of enhancement layer image frames the cloud desktop server sends to the second terminal, it will send the attribute information of each enhancement layer image frame to the second terminal. So that the second terminal processes the enhancement layer image frames according to the second attribute information.
[0096] Based on the foregoing description, it can be seen that in this application, after the cloud desktop server hierarchically encodes the to-be-encoded video data to obtain the base layer image frame and M groups of enhancement layer image frames, it sends different image frames to terminals with different network states. For the first terminal in a weak network state, the cloud desktop server sends the base layer image frame, which is obtained by frame extraction of the to-be-encoded video data, realizing the frame extraction function and reducing the bandwidth. In addition, for the second terminal in a non-weak network state, the cloud desktop server sends the base layer image frame and the enhancement layer image frame, which neither burdens the transmission bandwidth nor ensures that the second terminal obtains video data with a higher frame rate and displays a smoother video.
[0097] In some alternative embodiments, the cloud desktop server adopts a hierarchical cache pool, and different levels of cache pools are used to store image block identifiers with different layer attributes. Specifically, the cloud desktop server includes a first cache pool and a second cache pool. The first cache pool is used to store the image block identifiers with the base layer hit attribute or the base layer new attribute, and the second cache pool is used to store the image block identifiers with the enhancement layer hit attribute or the enhancement layer new attribute. Among them, both the base layer hit attribute and the enhancement layer hit attribute indicate that the corresponding image block is the same as the historical image block, and both the base layer new attribute and the enhancement layer new attribute indicate that the corresponding image block is different from the historical image block.
[0098] In this application, storing image block identifiers with different attributes in the cache pool of the cloud desktop server provides a basis for the preprocessing of the image frames included in the video data and provides technical support for the implementation of the technical solution of this application.
[0099] When processing video data, the cloud desktop server can also update the cache pool. This will be described in conjunction with Figures 5 to 7 the following example.
[0100] As Figure 5 shown, the attribute information of the image blocks with image block identifiers ID1, ID2, and ID3 is new attributes added to the base layer, indicating that these image blocks are different from the historical image blocks and are image blocks that have never been transmitted before. Additionally, since these image blocks are in the base layer, the decoding of other image blocks depends on these image blocks. Therefore, the cloud desktop server adds the identifiers of these image blocks to the first cache pool to complete the update of the first cache pool required for this preprocessing.
[0101] As Figure 6 shown, the attribute information of the image blocks with image block identifiers ID1, ID2, and ID3 is new attributes added to the enhancement layer, indicating that these image blocks are different from the historical image blocks and are image blocks that have never been transmitted before. The cloud desktop server adds the identifiers of these image blocks to the second cache pool to complete the update of the second cache pool required for this preprocessing.
[0102] As Figure 7 shown, the attribute information of the image blocks with image block identifiers ID1 and ID2 is new attributes added to the base layer, indicating that these image blocks are different from the historical image blocks and are image blocks that have never been transmitted before. Additionally, since these image blocks are in the base layer, the decoding of other image blocks depends on these image blocks. Therefore, the cloud desktop server adds the identifiers of these image blocks to the first cache pool to complete the update of the first cache pool required for this preprocessing. The attribute information of the image block with image block identifier ID4 is new attributes added to the enhancement layer, indicating that this image block is different from the historical image blocks and is image block that has never been transmitted before. The cloud desktop server adds the identifier of this image block to the second cache pool to complete the update of the second cache pool required for this preprocessing.
[0103] In this application, the cloud desktop server can also update the first cache pool to provide a basis for the subsequent preprocessing of image data, which is beneficial for identifying historical image blocks in the subsequent image data, thereby reducing the content that needs to be encoded.
[0104] In some alternative embodiments, after processing the image frames based on the attribute information of the image frames included in the video data, the method further includes: in the second cache pool, deleting the image block identifiers with attribute information being enhancement layer hit attributes or enhancement layer new attributes among the M groups of enhancement layer image frames. That is to say, after the preprocessing is completed, the cloud desktop server deletes the image block identifiers related to this preprocessing in the second cache pool.
[0105] Exemplarily, in Figure 6 , Figure 7 In the illustrated embodiment, since the image blocks in the second cache pool are of the enhancement layer and the decoding of other image blocks does not depend on these image blocks, after the preprocessing is completed, the cloud desktop server can also delete the image block identifiers (including ID1-6) related to the preprocessing here in the second cache pool.
[0106] In this application, since the encoding of the enhancement layer image data does not affect the encoding of the base layer image data, after the preprocessing of the image frames included in the video data, deleting the image block identifiers of the enhancement layer attributes in the image frames of the video data in the second cache pool by the cloud desktop server will not affect other image data and can also reduce the storage resources used by the cloud desktop server.
[0107] In the foregoing, the operations performed by the cloud desktop server in the video data processing method provided by the embodiments of this application are introduced. In the embodiments of this application, a client runs on the terminal. The cloud desktop server sends encoded data to the terminal. In terms of software, that is, the client running on the terminal obtains the encoded data. Based on the different encoded data obtained, the operations of the client are different, which will be described separately below.
[0108] 1. The client running on a terminal in a weak network state (hereinafter referred to as the first terminal).
[0109] The client manages a third cache pool, and the third cache pool is used to store image block identifiers and image content information whose attribute information includes base layer hit attributes. In other words, the first terminal includes a third cache pool, which is managed by the client. In this application, the management of the cache pool by the client includes adding, deleting, changing mapping relationships, etc. for the identifiers and / or image content information of the image blocks stored in the cache pool, and also includes operations such as deleting the cache pool that can change the cache pool itself or the content stored in the cache pool.
[0110] After the client running on the first terminal obtains the base layer image frame and the first attribute information, it first decodes the base layer image frame to obtain the decoded first image data. Then, according to the third cache pool and the first attribute information, it processes the first image data to obtain the video data to be displayed.
[0111] Among them, processing the first image data according to the third cache pool and the first attribute information includes: if the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block with the same identifier as the first image block from the third cache pool, and filling the image content information of the first image block with the image content information of the second image block. This process can also be understood as an image splicing process, where the base layer hit attribute indicates that the corresponding image block is the same as the historical image block.
[0112] Exemplarily, in conjunction with Figure 8 for illustration, please refer to Figure 8 , Figure 8 which is a schematic flowchart of the method for processing video data provided by an embodiment of the present application. Among them, Figure 8 in the illustrated embodiment, the preprocessing of a single base layer image frame is taken as an example. Based-add represents the newly added attributes of the base layer, and Based-hit represents the hit attributes of the base layer.
[0113] In Figure 8 the illustrated embodiment, a single-frame base layer image is divided into 6 image blocks. Among them, the attribute information of the image blocks with the image block identifiers ID4, ID5, and ID6 is the hit attribute of the base layer, which means that these image blocks are the same as the historical image blocks. When the cloud desktop server encodes this base layer image frame, these image blocks are the background color. Then, after the client running on the first terminal decodes this base layer image frame, these image blocks are still the background color. The image in the upper left corner of the figure represents the decoded image, that is, the first image data.
[0114] The client running on the first terminal determines, according to the identifiers of these image blocks, the image blocks with the same identifiers as these image blocks in the third cache pool, and the image content information corresponding to these image blocks. In addition, in the third cache pool, the position information of the image blocks is also stored, and this position information indicates the position of the image blocks in the image frame, so as to ensure the accurate position of the image frame splicing.
[0115] Based on this, the client running on the first terminal can obtain the position information and image content information of the image blocks with the image block identifiers ID4, ID5, and ID6 from the third cache pool, fill them into the corresponding image blocks, and obtain the filled image frame. Performing similar operations on each base layer image frame will obtain the image to be displayed.
[0116] In the present application, for the base layer image frame, after the client decodes it, the historical image blocks are filled according to the first attribute information, and the complete image frame can be obtained. That is to say, during the decoding process, this historical image block is actually the background color, reducing the decoding bandwidth.
[0117] Second, the client running on the terminal in a non-weak network state (hereinafter referred to as the second terminal).
[0118] The client manages the third cache pool and the fourth cache pool. The third cache pool is used to store the image block identifiers and image content information whose attribute information includes the hit attributes of the base layer, and the fourth cache pool is used to store the image block identifiers and image content information whose attribute information includes the hit attributes of the enhancement layer. In other words, the second terminal includes a third cache pool and a fourth cache pool.
[0119] After the client running on the second terminal obtains the base layer image frame and the first attribute information, it first decodes the base layer image frame to obtain the decoded first image data, and then processes the first image data according to the third cache pool and the first attribute information. The specific process is similar to the process of the client running on the first terminal processing the base layer image frame, which will not be elaborated here.
[0120] The client running on the second terminal also processes N groups of enhancement layer image frames, including: first decoding the N groups of enhancement layer image frames to obtain N groups of second image data. Then processing the N groups of second image data according to the fourth cache pool and the second attribute information. Or, processing the N groups of second image data according to the third cache pool, the fourth cache pool and the second attribute information.
[0121] It can be understood that for the client running on the second terminal, after receiving the base layer image frame and N groups of enhancement layer image frames, decoding and processing these image frames to obtain the first image data and M groups of second image data, and then processing the first image data and M groups of second image data based on the corresponding attribute information, the video data to be displayed is obtained. Compared with the video data to be displayed on the first terminal, the frame rate of the video data to be displayed on the second terminal is higher and the video playback is smoother.
[0122] Among them, the process of processing the N groups of second image data includes: if the second attribute information indicates that the attribute of the third image block in the N groups of second image data is the enhancement layer hit attribute, then determine the fourth image block with the same identifier as the third image block from the fourth cache pool, and fill the image content information of the third image block with the image content information of the fourth image block. And / or, if the second attribute information indicates that the attribute of the fifth image block in the N groups of second image data is the base layer hit attribute, then determine the sixth image block with the same identifier as the fifth image block from the third cache pool, and fill the image content information of the fifth image block with the image content information of the sixth image block. Among them, both the enhancement layer hit attribute and the base layer hit attribute indicate that the corresponding image block is the same as the historical image block. Generally speaking, for the historical data block in the decoded image data, based on the attribute information, obtain the image content information of the image block from the corresponding cache pool and fill it.
[0123] Exemplarily, in combination with Figure 9 for illustration, please refer to Figure 9 , Figure 8 which is the schematic flowchart of the video data processing method provided by the embodiment of the present application. Among them, Figure 9In the illustrated embodiment, the preprocessing of a single enhanced layer image frame is taken as an example. Based-add represents the newly added attributes of the base layer, Based-hit represents the hit attributes of the base layer, SVC-add represents the newly added attributes of the enhanced layer, and SVC-hit represents the hit attributes of the base layer.
[0124] In Figure 9 In the illustrated embodiment, a single-frame enhanced layer image is divided into 6 image blocks. Among them, the attribute information of the image block with the image block identifier ID3 is the hit attribute of the base layer, and the attribute information of the image blocks with ID5 and ID6 is the hit attribute of the enhanced layer, which means that these image blocks are the same as the historical image blocks. When the cloud desktop server encodes this enhanced layer image frame, these image blocks are the background color. Then, when the client running on the second terminal decodes this image frame, these image blocks are still the background color. The image in the upper left corner of the figure represents the decoded image, that is, the second image data.
[0125] The client running on the second terminal determines, according to the identifier ID3 of the base layer image block, the image block with the same identifier as this image block and the corresponding image content information in the third cache pool. In addition, in the third cache pool, the position information of the image block is also stored, and this position information indicates the position of the image block in the image frame, so as to ensure the accurate position of the image frame splicing.
[0126] The client running on the second terminal determines, according to the identifiers ID5 and ID6 of the enhanced layer image blocks, the image blocks with the same identifiers as these two image blocks and the corresponding image content information of the two image blocks respectively in the fourth cache pool. In addition, in the fourth cache pool, the position information of the image block is also stored, and this position information indicates the position of the image block in the image frame, so as to ensure the accurate position of the image frame splicing.
[0127] Based on this, the client running on the second terminal can obtain the position information and image content information of the image block with the image block identifier ID3 from the third cache pool; and obtain the position information and image content information of the image blocks with the image block identifiers ID5 and ID6 from the fourth cache pool. Then, fill the image content information in the cache pool into the corresponding image blocks to obtain the filled image frame. Perform similar operations on each enhanced layer image frame, and the client running on the second terminal will obtain M groups of reprocessed second image data. Combining with the reprocessed first image data corresponding to the base layer image frame, the video data to be displayed is obtained.
[0128] It should be noted that Figure 9 In the illustrated embodiment, it is taken as an example that the image blocks of a single enhanced layer image frame include the base layer related attributes and the enhanced layer related attributes. In actual applications, the image blocks of a single enhanced layer image frame can only include the enhanced layer related attributes, and the specific details are not limited here.
[0129] In this application, for the enhanced layer image frame, after the client decodes it, the historical image blocks are filled according to the second attribute information, and then a complete image frame can be obtained. That is to say, during the decoding process, the historical image blocks are actually the background color, which reduces the decoding bandwidth.
[0130] Based on the relevant descriptions of the decoding process above, the client / terminal can adopt a hierarchical cache pool to isolate the basic layer related attributes from the enhanced layer related attributes. At the same time, the basic layer image frame depends only on the third cache pool storing the basic layer related attributes for decoding and reprocessing. This ensures that even if the enhanced layer image frame is discarded, the remaining bitstream can be normally decoded on the premise that the basic layer image frame bitstream is completely transmitted, realizing the frame extraction function and achieving the anti-weak network effect.
[0131] In some optional embodiments, after processing the decoded image data based on the attribute information of the image frame, the client can also process the cache pool. This will be described in combination with Figure 8 and Figure 9 by way of example.
[0132] Optionally, when the first attribute information indicates that the attribute of the seventh image block in the first image data is a newly added attribute of the basic layer, it means that the seventh image block is an image block that has never been transmitted before. In addition, since this image block is of the basic layer, the decoding of other image blocks depends on the seventh image block. Therefore, the client stores the identifier and image content information of the seventh image block in the third cache pool. Optionally, the client can also store the position information of the seventh image block in the third cache pool, and this position information indicates the position of the seventh image block in the first image data.
[0133] Exemplarily, as Figure 8 shown, the attribute information of the image blocks with image block identifiers ID1, ID2, and ID3 is a newly added attribute of the basic layer. The client adds the identifiers, image content information, and position information of these image blocks to the third cache pool to complete the update of the third cache pool required for this reprocessing.
[0134] Optionally, when the second attribute information indicates that the attribute of the eighth image block in the N groups of second image data is a newly added attribute of the basic layer, it means that the eighth image block is different from the historical image blocks and is an image block that has never been transmitted before. In addition, since this image block is of the basic layer, the decoding of other image blocks depends on the eighth image block. Therefore, the client stores the identifier and image content information of the eighth image block in the third cache pool. Optionally, the client can also store the position information of the eighth image block in the third cache pool, and this position information indicates the position of the eighth image block in the second image data.
[0135] Exemplarily, asFigure 9 As shown, the attribute information of the image blocks with image block identifiers ID1 and ID2 is the newly added attribute of the base layer. The client adds the identifiers, image content information, and location information of these image blocks to the third cache pool, completing the update to the third cache pool required for this reprocessing.
[0136] In this application, the client can also update the third cache pool, providing a basis for the subsequent processing of image data, facilitating the identification of historical image blocks in the subsequent image data, facilitating the splicing of the decoded images into a complete image, and enhancing the practicality of the technical solution of this application.
[0137] Optionally, when the second attribute information indicates that the attribute of the ninth image block in the N groups of second image data is the newly added attribute of the enhancement layer, it means that the ninth image block is an image block that has never been transmitted before. The client stores the identifier and image content information of the ninth image block in the fourth cache pool. Optionally, the client can also store the location information of the ninth image block in the fourth cache pool, and this location information indicates the position of the ninth image block in the second image data.
[0138] Exemplarily, as Figure 9 shown, the attribute information of the image block with image block identifier ID4 is the newly added attribute of the enhancement layer. The client adds the identifier, image content information, and location information of this image block to the fourth cache pool, completing the update to the fourth cache pool required for this reprocessing.
[0139] In some alternative embodiments, after reprocessing the second image data, the client can further process the fourth cache pool. Generally speaking, the client can delete the identifiers, image memory information, and location information of the image blocks with the hit attributes or newly added attributes of the enhancement layer indicated by the second attribute information in the fourth cache pool. This is because the image blocks with enhancement layer-related attributes do not affect the processing of image blocks with other layer-related attributes during decoding and reprocessing. Deleting the information of the relevant attributes will not affect the subsequent decoding and reprocessing processes either.
[0140] Exemplarily, in Figure 9 the shown embodiment, since the image blocks in the fourth cache pool are of the enhancement layer and the decoding of other image blocks does not depend on these image blocks, after the reprocessing (i.e., filling) is completed, the client can also delete the image block identifiers (including ID4 - 6) related to this reprocessing in the fourth cache pool.
[0141] In some alternative embodiments, considering that the image blocks related to the enhancement layer attributes do not affect the processing of the image blocks related to other layer attributes during decoding and reprocessing. Then, for the attribute information indicating the newly added attributes of the enhancement layer, the identifiers and image content information of the image blocks corresponding to the attribute information may not be added to the fourth cache pool, further reducing memory occupancy.
[0142] In this application, since the processing of the enhancement layer image data does not affect the processing of the base layer image data, after the processing of the N groups of second image data obtained by decoding the enhancement layer image frames is completed, the client deletes the identifiers and image memory information of the image blocks related to the enhancement layer attributes corresponding to the second attribute information in the fourth cache pool, which will not affect other image data and can also reduce the storage resources used by the client.
[0143] Based on the foregoing description, it can be seen that in the video data processing method provided by the embodiments of this application, the cloud desktop server can encode the video data to obtain multiple bitstreams and distribute different bitstreams for different network states.
[0144] For Figure 2 In the application scenario of "cloud desktop + cloud conference" shown, the cloud desktop server encodes the video data to obtain the base layer image frames and the enhancement layer image frames. If the network states of the cloud conference server and each client are all non-weak network states, then the cloud desktop server can send the base layer image frames and the enhancement layer image frames. If the cloud conference server is in a weak network state, then the cloud desktop server can send the base layer image frames to implement the frame extraction function and reduce the bandwidth.
[0145] For Figure 3 In the application scenario of the collaborative desktop shown, the cloud desktop server encodes the video data to obtain the base layer image frames and the enhancement layer image frames, sends the base layer image frames to the clients in the weak network state to implement the frame extraction function and reduce the bandwidth, and sends the base layer image frames and the enhancement layer image frames to the clients in the non-weak network state.
[0146] In addition, in the related art, there is also a relationship of mutual reference between local macroblocks of previous and next frames, which is called Ref (the current frame macroblock fully refers to a local image of the previous frame) and RefAdd (the current frame macroblock refers to a local image of the previous frame, and there is also newly added image content). In the technical solution of the present application, by defining the attribute information of the image frame in the video data to be encoded, this relationship is changed from "previous frame" to "previous base layer frame", that is, the enhanced layer frame image cannot participate in the calculation of Ref / RefAdd to prevent the occurrence of a scenario where subsequent frames depend on the enhanced frame. At the same time, in the image frame, the identification and image content information of the image block related to the base layer attributes can be added to the corresponding cache pool (the first cache pool, the third cache pool); the identification and image content information related to the enhanced layer attributes can be added to the corresponding cache pool (the second cache pool, the fourth cache pool). In addition, from the perspective of reducing the amount of calculation, it is limited to the content indicated by the relevant attributes used for encoding or decoding of this frame, and it is also possible not to add the cache pool.
[0147] See also Figure 10 , Figure 10 A schematic diagram of the structure of the client provided in the embodiment of the present application. The client provided in the embodiment of the present application runs on a terminal, the terminal is connected to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the images presented by the cloud desktop server; the cloud desktop server is used to perform layered encoding on the video data to be encoded, and obtain a base layer image frame and M groups of enhanced layer image frames, the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by layered encoding, the decoding of the enhanced layer image frame depends on the base layer image frame, the video data to be encoded is obtained by the cloud desktop server processing the image frame based on the attribute information of the image frame included in the video data, the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block, and M is a positive integer.
[0148] In some optional implementations, the client 1000 includes a transceiver unit 1001, which is used to receive a base layer image frame and first attribute information of the base layer image frame from a cloud desktop server, wherein the first attribute information indicates an image block attribute included in the base layer image frame. A processing unit 1002 is used to process the first image data obtained by decoding the base layer image frame based on a third cache pool and the first attribute information, wherein the third cache pool is managed by the client, and the third cache pool is used to store attribute information including an image block identifier and image content information of a base layer hit attribute.
[0149] In some alternative embodiments, a transceiver unit 1001 is configured to receive a base layer image frame, first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames from a cloud desktop server, where the second attribute information indicates the image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M. A processing unit 1002 is configured to process first image data obtained by decoding the base layer image frame and N groups of second image data obtained by decoding the N groups of enhancement layer image frames based on a third cache pool, a fourth cache pool, the first attribute information, and the second attribute information, where the fourth cache pool is managed by the client, and the fourth cache pool is used to store the image block identifiers and image content information whose attribute information includes an enhancement layer hit attribute.
[0150] In some alternative embodiments, the processing unit 1002 is specifically configured to: if the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determine a second image block in the third cache pool that has the same identifier as the first image block, where the base layer hit attribute indicates that the corresponding image block is the same as a historical image block. Fill the image content information of the first image block with the image content information of the second image block.
[0151] In some alternative embodiments, the processing unit 1002 is specifically configured to: if the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determine a second image block in the third cache pool that has the same identifier as the first image block, where the base layer hit attribute indicates that the corresponding image block is the same as a historical image block. Fill the image content information of the first image block with the image content information of the second image block. If the second attribute information indicates that the attribute of the third image block in the N groups of second image data is an enhancement layer hit attribute, determine a fourth image block in the fourth cache pool that has the same identifier as the third image block, where the enhancement layer hit attribute indicates that the corresponding image block is the same as a historical image block. Fill the image content information of the third image block with the image content information of the fourth image block. And / or, if the second attribute information indicates that the attribute of the fifth image block in the N groups of second image data is a base layer hit attribute, determine a sixth image block in the third cache pool that has the same identifier as the fifth image block. Fill the image content information of the fifth image block with the image content information of the sixth image block.
[0152] In some alternative embodiments, the processing unit 1002 is further configured to: if the first attribute information indicates that the attribute of the seventh image block in the first image data is a base layer new attribute, store the identifier and image content information of the seventh image block in the third cache pool. And / or, if the second attribute information indicates that the attribute of the eighth image block in the N groups of second image data is a base layer new attribute, store the identifier and image content information of the eighth image block in the third cache pool. Wherein, the image block indicated by the base layer enhancement attribute is different from the historical image block
[0153] In some alternative embodiments, the processing unit 1002 is further configured to: if the second attribute information indicates that the attribute of the ninth image block in the N groups of second image data is an enhanced layer new attribute, store the identifier and image content information of the ninth image block in the fourth cache pool.
[0154] In some alternative embodiments, the processing unit 1002 is further configured to: in the fourth cache pool, delete the identifiers and image memory information of the image blocks with enhanced layer hit attributes or enhanced layer new attributes indicated by the second attribute information. Among them, the image blocks indicated by the enhanced layer new attributes are different from the historical image blocks.
[0155] The client 1000 is used to implement the foregoing Figure 2 and Figure 3 in the embodiments shown, the master client and / or the collaborating client, Figures 4 to 9 the operations performed by the client running on the first terminal and / or the client running on the second terminal in the embodiments shown are not described herein again.
[0156] Please refer to Figure 11 , Figure 11 , which is a schematic structural diagram of a terminal provided by an embodiment of the present application. The terminal 1100 includes a processor 1101, a memory 1102, a communication interface 1103, and a bus 1104. Among them, the processor 1101, the memory 1102, and the communication interface 1103 communicate through the bus 1104, and can also communicate through other means such as wireless transmission. The memory 1102 stores program codes, and the processor 1101 can call the program codes stored in the memory 1102 to execute the foregoing Figure 2 and Figure 3 in the embodiments shown, the master client and / or the collaborating client, Figures 4 to 9 the operations performed by the client running on the first terminal and / or the client running on the second terminal in the embodiments shown are not described herein again.
[0157] It should be understood that in the embodiments of the present application, the processor 1101 may be a CPU, and the processor 1101 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0158] The memory 1102 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1101. The memory 1102 may also include a non-volatile random access memory. For example, the memory 1102 may also store information about the device type.
[0159] The memory 1102 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0160] In addition to including a data bus, the bus 1104 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, all kinds of buses are labeled as the bus 1104 in the figure. The bus 1140 may be a Peripheral Component Interconnect Express (PCIe) bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. The bus 1140 may be divided into an address bus, a data bus, a control bus, etc.
[0161] The terminal 1100 may further include one or more communication interfaces and one or more operating systems, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0162] Please refer to Figure 12 , Figure 12 , which is a schematic structural diagram of the cloud desktop server provided by the embodiment of the present application. Among them, the cloud desktop server 1200 is connected to multiple terminals, and the screens of the multiple terminals display the pictures presented by the cloud desktop server 1200.
[0163] In some optional embodiments, the cloud desktop server 1200 includes a processing unit 1201 and a transceiver unit 1202.
[0164] The processing unit 1201 is configured to process the image frames based on the attribute information of the image frames included in the video data to obtain the video data to be encoded. The attribute information of the image frames is used to indicate whether the image blocks included in the image frames are the same as the historical image blocks. Perform hierarchical encoding on the video data to be encoded to obtain a base layer image frame and M groups of enhancement layer image frames. Among them, the base layer image frame is independently decoded and is obtained by extracting frames from the image frames obtained by hierarchical encoding. The decoding of the enhancement layer image frames depends on the base layer image frame, and M is a positive integer.
[0165] The transceiver unit 1202 is configured to send the base layer image frame and the first attribute information of the base layer image frame to the first terminal in a weak network state among the multiple terminals. The first attribute information indicates the image block attributes included in the base layer image frame. Send the base layer image frame, the first attribute information, N groups of enhancement layer image frames and the second attribute information of the N groups of enhancement layer image frames to the second terminal in a non-weak network state among the multiple terminals. The second attribute information indicates the image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M.
[0166] In some optional embodiments, the processing unit 1201 is specifically configured to: process the image blocks with the attribute information of the base layer hit attribute or the enhancement layer hit attribute in the image frames as the background color. Both the base layer hit attribute and the enhancement layer hit attribute indicate that the corresponding image blocks are the same as the historical image blocks.
[0167] In some alternative embodiments, the cloud desktop server 1200 includes a first cache pool and a second cache pool. The first cache pool is used to store the image block identifiers with the attribute information being the base layer hit attribute or the base layer new attribute, and the second cache pool is used to store the image block identifiers with the attribute information being the enhancement layer hit attribute or the enhancement layer new attribute. Both the base layer new attribute and the enhancement layer new attribute indicate that the corresponding image block is different from the historical image block.
[0168] In some alternative embodiments, the processing unit 1201 is further configured to: in the second cache pool, delete the image block identifiers with the attribute information being the enhancement layer hit attribute or the enhancement layer new attribute in the M groups of enhancement layer image frames.
[0169] Among them, both the processing unit 1201 and the transceiver unit 1202 can be implemented by software or by hardware. Exemplarily, next, taking the processing unit 1201 as an example, the implementation manner of the processing unit 1201 is introduced. Similarly, the implementation manner of the transceiver unit 1202 can refer to the implementation manner of the processing unit 1201.
[0170] As an example of a software functional unit, the processing unit 1201 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the processing unit 1201 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ), or in different AZs. Each AZ includes one data center or multiple geographically proximate data centers. Among them, generally, one region may include multiple AZs.
[0171] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Among them, generally, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.
[0172] As an example of a hardware functional unit, the processing unit 1201 may include at least one computing device, such as a server, etc. Alternatively, the processing unit 1201 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0173] The multiple computing devices included in the processing unit 1201 may be distributed in the same region or in different regions. The multiple computing devices included in the processing unit 1201 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the processing unit 1201 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0174] It should be noted that different steps in the data processing method are respectively implemented by the processing unit 1201 and the transceiver unit 1202 to implement all the functions of the cloud desktop server 1200. The cloud desktop server 1200 is used for the operations performed by the cloud desktop server in the foregoing Figures 2 to 9 embodiments shown to implement the video data processing method applied to the cloud desktop server provided in the embodiments of the present application, which will not be elaborated here.
[0175] Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a computing device provided in the embodiments of the present application. The computing device 1300 includes a processor 1301, a communication interface 1302, a bus 1303, and a memory 1304. Among them, the processor 1301, the communication interface 1302, and the memory 1304 communicate with each other through the bus 1303. In practical applications, communication may also be achieved by other means such as wireless transmission, and specific details are not limited here.
[0176] The computing device 1300 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1300.
[0177] The processor 1301 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0178] The communication interface 1302 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1300 and other devices or communication networks.
[0179] The bus 1303 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 13 only one line is shown here, but it does not mean that there is only one bus or one type of bus. The bus 1303 may include a path for transmitting information between various components of the computing device 1300 (for example, the memory 1304, the processor 1301, and the communication interface 1302).
[0180] The memory 1304 may include volatile memory, such as random access memory (RAM). The memory 1304 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0181] The memory 1304 stores executable program code, and the processor 1301 executes the executable program code to respectively implement the functions of the aforementioned processing unit 1201 and transceiver unit 1202, thereby implementing a method for processing video data applied to a cloud desktop server. That is, the memory 1304 stores instructions for executing a method for processing video data applied to a cloud desktop server.
[0182] Alternatively, executable code is stored in the memory 1304, and the processor 1301 executes the executable code to implement the functions of the foregoing processing unit 1201 and transceiver unit 1202 respectively, thereby implementing a method for processing video data applied to a cloud desktop server. That is, instructions for executing a method for processing video data applied to a cloud desktop server are stored on the memory 1304.
[0183] An embodiment of this application also provides a computing device cluster, which includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center.
[0184] Please refer to Figure 14 and Figure 15 , Figure 14 and Figure 15 Both are schematic structural diagrams of the computing device cluster provided by the embodiments of this application.
[0185] As Figure 14 shown, the computing device cluster includes at least one computing device 1300. Instructions for executing the method for processing video data applied to the cloud desktop server provided by the embodiments of this application may be stored in the memory 1304 of one or more computing devices 1300 in the computing device cluster.
[0186] In some possible implementation manners, partial instructions for executing a data processing method may also be stored in the memory 1304 of one or more computing devices 1300 in the computing device cluster respectively. In other words, a combination of one or more computing devices 1304 may jointly execute instructions for executing a method for processing video data applied to a cloud desktop server.
[0187] It should be noted that the memories 1304 in different computing devices 1300 in the computing device cluster may store different instructions, which are respectively used to execute partial functions of the data processing device. That is, the instructions stored in the memories 1304 of different computing devices 1300 may implement the functions of one or more of the processing unit 1201 and the transceiver unit 1202.
[0188] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 15 Shows a possible implementation manner. As Figure 15As shown, two computing devices 1300A and 1300B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this possible implementation, the memory 1304 in the computing device 1300A stores instructions for performing the functions of the transceiver unit 1202. At the same time, the memory 1304 in the computing device 1300B stores instructions for performing the functions of the processing unit 1201.
[0189] Figure 15 The connection method between the computing device clusters shown can be considered in the video data processing method applied to the cloud desktop server provided in this application. The processing operations and operations other than the processing operations are separated and executed. That is, therefore, it is considered that the functions of the processing unit 1201 are executed by the computing device 1300B, and the functions of the transceiver unit 1202 are executed by the computing device 1300A.
[0190] It should be understood that Figure 15 the functions of the computing device 1300A shown in can also be completed by multiple computing devices 1300. Similarly, the functions of the computing device 1300B can also be completed by multiple computing devices 1300.
[0191] The embodiments of this application also provide another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to Figure 14 and Figure 15 the connection method of the described computing device cluster, which will not be elaborated here.
[0192] The embodiments of this application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computer device, it causes at least one computer device to execute the above-mentioned video data processing method.
[0193] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute the above-mentioned video data processing method.
[0194] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.
[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for processing video data, characterized in that: The method is applied to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals are used to display the pictures presented by the cloud desktop server; the method includes: Based on attribute information of an image frame included in the video data, the image frame is processed to obtain video data to be encoded, wherein the attribute information of the image frame is used to indicate whether an image block included in the image frame is the same as a historical image block; Performing hierarchical coding on the video data to be coded, obtaining a base layer image frame and M groups of enhancement layer image frames, wherein the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by hierarchical coding, the decoding of the enhancement layer image frame depends on the base layer image frame, and M is a positive integer; Sending the base layer image frame and first attribute information of the base layer image frame to a first terminal in a weak network state among the multiple terminals, where the first attribute information indicates attributes of image blocks included in the base layer image frame; The base layer image frame, the first attribute information, N groups of enhancement layer image frames and the second attribute information of the N groups of enhancement layer image frames are sent to the second terminal in a non-weak network state among the multiple terminals, wherein the second attribute information indicates the image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M.
2. The method according to claim 1, characterized in that The step of processing the image frame based on the attribute information of the image frame included in the video data comprises: An image block whose attribute information in the image frame is a base layer hit attribute or an enhanced layer hit attribute is processed as a background color, and both the base layer hit attribute and the enhanced layer hit attribute indicate that the corresponding image block is the same as the historical image block.
3. The method according to claim 1 or 2, characterized in that: The cloud desktop server includes a first cache pool and a second cache pool, the first cache pool is used to store image block identifiers whose attribute information is a base layer hit attribute or a base layer newly added attribute, and the second cache pool is used to store image block identifiers whose attribute information is an enhancement layer hit attribute or an enhancement layer newly added attribute, and the base layer newly added attribute and the enhancement layer newly added attribute both indicate that the corresponding image block is different from the historical image block.
4. The method according to claim 3, characterized in that: After processing the image frame based on the attribute information of the image frame included in the video data, the method further includes: In the second buffer pool, image block identifiers whose attribute information is enhancement layer hit attributes or enhancement layer newly added attributes in the M groups of enhancement layer image frames are deleted.
5. A method for processing video data, characterized in that: The method is applied to a client, the client runs on a terminal, the terminal is connected to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the images presented by the cloud desktop server; The cloud desktop server is used to perform hierarchical encoding on the video data to be encoded, and obtain a base layer image frame and M groups of enhanced layer image frames, wherein the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by hierarchical encoding, and the decoding of the enhanced layer image frame depends on the base layer image frame, and the video data to be encoded is obtained by the cloud desktop server processing the image frame based on the attribute information of the image frame included in the video data, and the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block, and M is a positive integer; The method comprises: Receiving the base layer image frame and first attribute information of the base layer image frame from the cloud desktop server, wherein the first attribute information indicates attributes of image blocks included in the base layer image frame; Based on a third buffer pool and the first attribute information, processing the first image data obtained by decoding the base layer image frame, wherein the third buffer pool is managed by the client, and the third buffer pool is used to store attribute information including an image block identifier and image content information of a base layer hit attribute; or, Receiving the base layer image frame, the first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames from the cloud desktop server, wherein the second attribute information indicates image block attributes included in the N groups of enhancement layer image frames, where N is a positive integer less than or equal to M; Based on the third cache pool, the fourth cache pool, the first attribute information and the second attribute information, the first image data obtained by decoding the base layer image frame and the N groups of second image data obtained by decoding the N groups of enhancement layer image frames are processed, wherein the fourth cache pool is managed by the client, and the fourth cache pool is used to store attribute information including image block identifiers and image content information of enhancement layer hit attributes.
6. The method according to claim 5, characterized in that The processing, based on the third buffer pool and the first attribute information, of first image data obtained by decoding the base layer image frame includes: If the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block having the same identifier as the first image block from the third cache pool, the base layer hit attribute indicating a historical image block; The image content information of the first image block is filled as the image content information of the second image block.
7. The method according to claim 5, characterized in that The processing, based on the third buffer pool, the fourth buffer pool, the first attribute information, and the second attribute information, of first image data obtained by decoding the base layer image frame and second image data obtained by decoding the N groups of enhancement layer image frames includes: If the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block having the same identifier as the first image block from the third cache pool, the base layer hit attribute indicating that the corresponding image block is the same as the historical image block; Filling the image content information of the first image block with the image content information of the second image block; If the second attribute information indicates that the attribute of the third image block in the N groups of second image data is an enhancement layer hit attribute, determining a fourth image block having the same identifier as the third image block from the fourth cache pool, and the image block corresponding to the enhancement layer hit attribute indication is the same as the historical image block; Filling the image content information of the third image block as the image content information of the fourth image block; and / or, If the second attribute information indicates that the attribute of the fifth image block in the N groups of second image data is a base layer hit attribute, determining a sixth image block having the same identifier as the fifth image block from the third cache pool; The image content information of the fifth image block is filled as the image content information of the sixth image block.
8. The method according to any one of claims 5 to 7, characterized in that The method further comprises: If the first attribute information indicates that the attribute of the seventh image block in the first image data is a newly added attribute of the base layer, the identifier and image content information of the seventh image block are stored in the third cache pool; and / or, If the second attribute information indicates that the attribute of the eighth image block in the N groups of second image data is a new attribute added to the base layer, the identifier and image content information of the eighth image block are stored in the third cache pool; wherein the image block indicated by the base layer enhanced attribute is different from the historical image block.
9. The method according to claim 7, characterized in that: The method further comprises: If the second attribute information indicates that the attribute of the ninth image block in the N groups of second image data is a new attribute of the enhancement layer, the identifier and image content information of the ninth image block are stored in the fourth cache pool, wherein the image block indicated by the new attribute of the enhancement layer is different from the historical image block.
10. The method according to any one of claims 5 to 9, characterized in that The method further comprises: In the fourth buffer pool, the identifier and image memory information of the image block with the enhancement layer hit attribute or the enhancement layer newly added attribute indicated by the second attribute information are deleted.
11. A cloud desktop server, characterized in that: The cloud desktop server is connected to a plurality of terminals, and the screens of the plurality of terminals are used to display the images presented by the cloud desktop server; The cloud desktop server comprises: a processing unit, configured to process the image frame based on attribute information of the image frame included in the video data to obtain the video data to be encoded, wherein the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block; The processing unit is further used to perform hierarchical encoding on the video data to be encoded, to obtain a base layer image frame and M groups of enhancement layer image frames, wherein the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by hierarchical encoding, the decoding of the enhancement layer image frame depends on the base layer image frame, and M is a positive integer; A transceiver unit, configured to send the base layer image frame and first attribute information of the base layer image frame to a first terminal in a weak network state among the multiple terminals, wherein the first attribute information indicates an attribute of an image block included in the base layer image frame; The transceiver unit is also used to send the base layer image frame, the first attribute information, N groups of enhancement layer image frames and second attribute information of the N groups of enhancement layer image frames to a second terminal in a non-weak network state among the multiple terminals, wherein the second attribute information indicates image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M.
12. The cloud desktop server according to claim 11, characterized in that: The processing unit is specifically used for: An image block whose attribute information in the image frame is a base layer hit attribute or an enhanced layer hit attribute is processed as a background color, and both the base layer hit attribute and the enhanced layer hit attribute indicate that the corresponding image block is the same as the historical image block.
13. The cloud desktop server according to claim 11 or 12, characterized in that: The cloud desktop server includes a first cache pool and a second cache pool, the first cache pool is used to store image block identifiers whose attribute information is a base layer hit attribute or a base layer newly added attribute, and the second cache pool is used to store image block identifiers whose attribute information is an enhancement layer hit attribute or an enhancement layer newly added attribute, and the base layer newly added attribute and the enhancement layer newly added attribute both indicate that the corresponding image block is different from the historical image block.
14. The cloud desktop server according to claim 13, characterized in that: The processing unit is further used for: In the second buffer pool, image block identifiers whose attribute information is enhancement layer hit attributes or enhancement layer newly added attributes in the M groups of enhancement layer image frames are deleted.
15. A client, characterized in that: The client runs on a terminal, the terminal is connected to a cloud desktop server, the cloud desktop server is connected to multiple terminals, and the screens of the multiple terminals display the images presented by the cloud desktop server; The cloud desktop server is used to perform hierarchical encoding on the video data to be encoded, and obtain a base layer image frame and M groups of enhanced layer image frames, wherein the base layer image frame is independently decoded and is obtained by extracting the image frame obtained by hierarchical encoding, and the decoding of the enhanced layer image frame depends on the base layer image frame, and the video data to be encoded is obtained by the cloud desktop server processing the image frame based on the attribute information of the image frame included in the video data, and the attribute information of the image frame is used to indicate whether the image block included in the image frame is the same as the historical image block, and M is a positive integer; The client comprises: A transceiver unit, configured to receive the base layer image frame and first attribute information of the base layer image frame from the cloud desktop server, wherein the first attribute information indicates an attribute of an image block included in the base layer image frame; a processing unit, configured to process first image data obtained by decoding the base layer image frame based on a third buffer pool and the first attribute information, wherein the third buffer pool is managed by the client, and the third buffer pool is used to store attribute information including an image block identifier and image content information of a base layer hit attribute; Alternatively, the client comprises: a transceiver unit, configured to receive the base layer image frame, the first attribute information, N groups of enhancement layer image frames, and second attribute information of the N groups of enhancement layer image frames from the cloud desktop server, wherein the second attribute information indicates image block attributes included in the N groups of enhancement layer image frames, and N is a positive integer less than or equal to M; A processing unit is used to process the first image data obtained by decoding the base layer image frame and the N groups of second image data obtained by decoding the N groups of enhancement layer image frames based on the third cache pool, the fourth cache pool, the first attribute information and the second attribute information, wherein the fourth cache pool is managed by the client, and the fourth cache pool is used to store attribute information including image block identifiers and image content information of enhancement layer hit attributes.
16. The client according to claim 15, characterized in that: The processing unit is specifically used for: If the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block having the same identifier as the first image block from the third cache pool, the base layer hit attribute indicating a historical image block; The image content information of the first image block is filled as the image content information of the second image block.
17. The client according to claim 15, wherein the processing unit is specifically configured to: If the first attribute information indicates that the attribute of the first image block in the first image data is a base layer hit attribute, determining a second image block having the same identifier as the first image block from the third cache pool, the base layer hit attribute indicating that the corresponding image block is the same as the historical image block; Filling the image content information of the first image block with the image content information of the second image block; If the second attribute information indicates that the attribute of the third image block in the N groups of second image data is an enhancement layer hit attribute, determining a fourth image block having the same identifier as the third image block from the fourth cache pool, and the image block corresponding to the enhancement layer hit attribute indication is the same as the historical image block; Filling the image content information of the third image block as the image content information of the fourth image block; and / or, If the second attribute information indicates that the attribute of the fifth image block in the N groups of second image data is a base layer hit attribute, determining a sixth image block having the same identifier as the fifth image block from the third cache pool; The image content information of the fifth image block is filled as the image content information of the sixth image block.
18. The client according to any one of claims 15 to 17, characterized in that: The processing unit is also used to: If the first attribute information indicates that the attribute of the seventh image block in the first image data is a newly added attribute of the base layer, storing the identifier and image content information of the seventh image block in the third cache pool; and / or, If the second attribute information indicates that the attribute of the eighth image block in the N groups of second image data is a new attribute added to the base layer, the identifier and image content information of the eighth image block are stored in the third cache pool; wherein the image block indicated by the base layer enhanced attribute is different from the historical image block.
19. The client according to claim 17, characterized in that: The processing unit is further used for: If the second attribute information indicates that the attribute of the ninth image block in the N groups of second image data is a newly added attribute of the enhancement layer, the identifier and image content information of the ninth image block are stored in the fourth cache pool.
20. The client according to any one of claims 15 to 19, characterized in that: The processing unit is also used to: In the fourth cache pool, the identifier and image memory information of the image block with the enhanced layer hit attribute or the enhanced layer newly added attribute indicated by the second attribute information are deleted, wherein the image block indicated by the enhanced layer newly added attribute is different from the historical image block.
21. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 4.
22. A terminal, characterized in that: comprising a processor coupled to a memory; Instructions are stored in the memory, and when the instructions are executed on the processor, the method according to any one of claims 5 to 10 is implemented.
23. A computer program product comprising instructions, characterized in that When the instruction is executed by a computing device cluster, the computing device cluster executes the method as described in any one of claims 1 to 4; or, when the instruction is executed by a terminal, the terminal executes the method as described in any one of claims 5 to 10.
24. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes computer program instructions, and when the computer program instructions are executed by a computing device cluster, the method according to any one of claims 1 to 4 is implemented; or when the computer program instructions are executed by a terminal or a client, the method according to any one of claims 5 to 10 is implemented.