Content generation method, device, system and equipment and computer readable storage medium
By receiving and decompressing the bitstream using generative artificial intelligence technology, content features corresponding to content description information are generated, solving the problem of low compression rate when the amount of cloud content data is large, improving compression rate and transmission efficiency, and ensuring the quality of generated content.
Patent Information
- Application Number
- CN202410551193.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-06
- Publication Date
- 2025-11-07
AI Technical Summary
In generative artificial intelligence technology, when the amount of content data generated in the cloud is large, the compression rate is low, which leads to longer transmission time and affects the quality of the generated content.
By receiving and decompressing the bitstream, content features corresponding to content description information are generated. Techniques such as inverse quantization and entropy decoding are used to improve compression ratio and transmission efficiency, and reduce transmission latency.
It improves the compression rate and transmission efficiency of the bitstream, shortens the transmission latency, and reduces the impact on the quality of the generated content.
Smart Images

Figure CN120915958A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, more specifically, to a content generation method, device, system, equipment and computer readable storage medium. BACKGROUND
[0002] In the generative artificial intelligence technology (AIGC, Artificial Intelligence Generated Content) multi-modal generation scene, the cloud-deployed AICG multi-modal generation model generates text, images, videos, audio and other content corresponding to the description information according to the description information input by the user end. And return the generated content to the user end.
[0003] When the data volume of the content is large, or in order to shorten the data transmission time, the cloud usually needs to compress the generated content before returning it to the user end. Taking text-to-image as an example, the user end sends text data to the cloud, the cloud generates content based on the text data, and obtains an image corresponding to the text data. And compress and encode the image to generate a bit stream, and send the bit stream to the user end. The user end decodes the bit stream to obtain a reconstructed image.
[0004] Usually, the compression rate is low when compressing the content. When the data volume of the content is large, the amount of data to be compressed increases, and the data volume of the compressed data is still large, resulting in long transmission delay and affecting the quality of the generated image. SUMMARY
[0005] The embodiments of the present application provide a content generation method, device, system, equipment and computer readable storage medium to improve the compression rate, reduce the transmission loss and improve the quality of the generated content.
[0006] In a first aspect, the embodiments of the present application provide a content generation method, which comprises receiving a code stream, decompressing the code stream to obtain content features corresponding to content description information. The reconstructed content corresponding to the content description information is generated according to the content features. The data volume of the content features is less than the data volume of the content corresponding to the content description information.
[0007] Based on the first aspect, since the data amount of the content feature is less than the data amount of the content corresponding to the content description information, the data amount of the code stream is less than the data amount of the code stream formed by directly compressing the content in the related art. Compared with the way of compressing the generated content in the related art, the compression rate of the code stream can be improved in the embodiments of the present application. Moreover, the data amount of the feature is less than the data amount of the content, which can improve the compression efficiency. In addition, the data amount of the code stream is small in the embodiments of the present application, the bandwidth occupation is small, the transmission efficiency can be improved, the transmission delay is shortened, and the loss of the code stream caused by the transmission of the code stream is reduced, the transmission delay is shortened, and the influence on the quality of the reconstructed content is reduced.
[0008] In a possible implementation, the code stream is decompressed to obtain the content feature corresponding to the content description information, and the implementation is specifically: the code stream is inverse quantized to obtain the content feature corresponding to the content description information.
[0009] Based on this possible implementation, the content generation rate is improved by inverse quantizing the code stream for fast decompression.
[0010] In a possible implementation, the code stream is decompressed to obtain the content feature corresponding to the content description information, and the implementation is specifically: the code stream is entropy decoded to obtain the decoded feature; and the decoded feature is inverse quantized to obtain the content feature corresponding to the content description information.
[0011] In this way, the accuracy of the decompressed content feature can be ensured by inverse quantization combined with entropy decoding.
[0012] In a possible implementation, the implementation is specifically: before receiving the code stream, the prior feature is obtained. The prior feature includes a quantization parameter and a distribution feature of the content feature. When the code stream is decompressed, the code stream is entropy decoded according to the distribution feature to obtain the decoded feature; and the decoded feature is inverse quantized according to the quantization parameter to obtain the content feature corresponding to the content description information.
[0013] In this way, the decompression performance is improved by the prior feature, the quality of the decompressed content feature is improved, and the data quality of the finally generated reconstructed content is ensured.
[0014] In a possible implementation, the code stream is decompressed to obtain the content feature corresponding to the content description information, and the implementation is specifically: the code stream is decompressed to obtain the dimension-reduced feature of the content feature corresponding to the content description information. The reconstructed content corresponding to the content description information is generated according to the dimension-reduced feature.
[0015] Based on this possible implementation, the corresponding reconstructed content is generated by the dimension-reduced feature, which can reduce the content generation time and improve the content generation efficiency.
[0016] In a possible implementation, the reconstructed content corresponding to the content description information is generated according to the content features, and specifically, the content features are processed to obtain processed features, and the reconstructed content corresponding to the content description information is generated according to the processed features.
[0017] In this way, the data quality of the reconstructed content is improved by reducing the noise in the decompressed features. Furthermore, the data quality of the generated reconstructed content is ensured.
[0018] In a possible implementation, the content features are processed to obtain processed features, and specifically, the content features are input into a diffusion model to obtain the processed features.
[0019] Based on this possible implementation, the decompressed content features are processed by the diffusion model, and in this way, the data quality of the reconstructed content is improved by reducing the noise in the decompressed content features through the diffusion model.
[0020] In a possible implementation, specifically, before receiving the code stream, a content generation request is sent to the server, and the content generation request includes the content description information.
[0021] Based on this possible implementation, by sending the content generation request to the server, the adaptation degree between the generated reconstructed content and the content description information is improved, and the interactivity with the server is also improved.
[0022] In a second aspect, an embodiment of the present application provides a content generation method. The method generates content features corresponding to content description information according to the content description information. The content features are compressed to obtain a code stream. The data amount of the content features is less than the data amount of content corresponding to the content description information.
[0023] Based on the second aspect, the content features corresponding to the content description information are directly compressed, which is different from the compression of the generated content in the related art. Since the data amount of the code stream is less than the data amount of the code stream formed by directly compressing the content in the related art, the compression rate of the code stream can be improved compared with the compression of the generated content in the related art. Moreover, the data amount of the features is less than the data amount of the content, which can improve the compression efficiency.
[0024] In a possible implementation, the content features are compressed to obtain a code stream, and specifically, the content features are quantized to obtain the code stream.
[0025] Based on this possible implementation, the content feature compression efficiency and the content generation efficiency are improved by quantization processing. Moreover, the data amount of the code stream can be reduced by quantization processing, and the transmission efficiency is improved.
[0026] In a possible implementation, the content feature is quantized to obtain the code stream, and the implementation is specifically as follows: when the network bandwidth between the server and the client is greater than or equal to a bandwidth threshold, the content feature is quantized to obtain the code stream.
[0027] Since the quantization manner can obtain higher compression efficiency, the transmission efficiency can be improved. However, in the case of small network bandwidth, there may be a transmission loss problem. Therefore, in this possible implementation, the quantization is performed when the network bandwidth between the server and the client is greater than or equal to the bandwidth threshold, the corresponding compression manner is flexibly selected according to the network bandwidth, the adaptation degree between the compression manner and the network bandwidth is improved, and thus the compression efficiency is improved and the compression quality is ensured.
[0028] In a possible implementation, the content feature is quantized to obtain the code stream, and the implementation is specifically as follows: when the network bandwidth between the server and the client is greater than or equal to a bandwidth threshold, the content feature is quantized to obtain the code stream.
[0029] Since the quantization and entropy coding combined manner can obtain a smaller compression rate, the data amount of the code stream obtained by compression is smaller. Therefore, in this possible implementation, when the network bandwidth between the server and the client is less than the bandwidth threshold, the adaptation degree between the compression manner and the network bandwidth is improved by the quantization and entropy coding combined manner, so as to improve the compression efficiency and ensure the compression quality. Moreover, the quantization and entropy coding combined manner can reduce feature loss, and thus reduce the error between the content feature corresponding to the content description information after decompression and the content feature, and improve the data quality of the reconstructed content.
[0030] In a possible implementation, the content feature is compressed to obtain the code stream, and the implementation is specifically as follows: the feature is input into a hyper-prior network to obtain a prior feature of the content feature; the prior feature includes a quantization parameter and a distribution feature of the content feature; the content feature is quantized according to the quantization parameter to obtain a quantized feature; and the quantized feature is entropy coded according to the distribution feature to obtain the code stream.
[0031] Based on this possible implementation, the feature compression is guided by the prior feature of the content feature, and the redundant information in the content feature corresponding to the content description information is removed. In this way, the compression performance is improved by the prior feature, the compression quality is improved, and thus the data quality of the finally generated content data is ensured.
[0032] In a possible implementation, the content feature is compressed to obtain the code stream, and the implementation is specifically as follows: the content feature is dimensionally reduced to obtain a dimensionally reduced feature of the content feature corresponding to the content description information; and the dimensionally reduced feature is compressed to obtain the code stream.
[0033] Based on the possible implementation manner, the data amount of the code stream can be further reduced by dimension reduction processing, and the compression rate and compression efficiency of the code stream are improved. In addition, since the data amount of the code stream generated based on the dimension reduction feature is small, the bandwidth occupation is small, and the transmission efficiency is improved, and then the transmission loss is reduced.
[0034] In a possible implementation manner, the method specifically comprises: receiving a content generation request sent by the client before generating the content feature corresponding to the content description information according to the content description information, wherein the content generation request comprises the content description information.
[0035] In this way, by receiving the content generation request sent by the client, the adaptation degree between the generated reconstructed content and the content description information is improved, and the interactivity with the client is also improved.
[0036] In a possible implementation manner, the method specifically comprises: obtaining the stored content description information before generating the content feature corresponding to the content description information according to the content description information.
[0037] In this way, by obtaining the stored content description information, the interaction with the client is reduced, and the efficiency of generating the content feature by the server is improved.
[0038] In a possible implementation manner, the method specifically comprises: generating the content feature corresponding to the content description information according to the content description information and a generative model based on a diffusion model.
[0039] In this way, the efficiency of obtaining the content feature is improved by the diffusion model, and the rationality of the content feature is also improved, so as to ensure the data quality of the finally generated reconstructed content.
[0040] In a possible implementation manner, the method specifically comprises: the content description information comprises at least one of text, image, audio or video.
[0041] Based on the possible implementation manner, multi-modal content generation can be implemented.
[0042] In a third aspect, an embodiment of the present application provides a content generation method, in which a server generates a content feature corresponding to content description information according to the content description information, compresses the content feature to obtain a code stream, and sends the code stream to a client. The data amount of the content feature is less than that of the content corresponding to the content description information. The client decompresses the code stream to obtain the content feature corresponding to the content description information. The client generates reconstructed content corresponding to the content description information according to the content feature.
[0043] According to the content generation method provided in the third aspect, the server determines content features corresponding to the content description information according to the content description information, compresses the content features to obtain a code stream, and feeds back the code stream to the client, so that the client decodes the code stream to generate reconstructed content corresponding to the content description information. The quality of the generated content is ensured by reasonably generating the content features. Since the data amount of the content features is much smaller than that of the content, the data amount of the code stream obtained by compressing the content features is smaller than that of the code stream obtained by compressing the content, and the compression rate of the code stream is effectively improved. In addition, since the data amount of the code stream is small, the bandwidth occupation is small, the code stream can be quickly transmitted, the transmission delay is shortened, and the loss of the code stream caused by the transmission of the code stream is reduced, the transmission delay is shortened, and the influence on the quality of the generated image is reduced.
[0044] In a fourth aspect, the embodiments of the present application provide a content generation system, which comprises a server and a client.
[0045] The server is configured to generate content features corresponding to the content description information according to the content description information, compress the content features to obtain a code stream, and send the code stream to the client. The data amount of the content features is smaller than that of the content corresponding to the content description information.
[0046] The client is configured to decompress the code stream to obtain the content features corresponding to the content description information, and generate reconstructed content corresponding to the content description information according to the content features.
[0047] In a fifth aspect, the embodiments of the present application provide a content generation device, which comprises a communication module, a decompression module and a generation module.
[0048] The communication module is configured to receive a code stream. The decompression module is configured to decompress the code stream to obtain content features corresponding to the content description information. The generation module is configured to generate reconstructed content corresponding to the content description information according to the content features.
[0049] In a possible implementation, the decompression module is configured to inverse quantize the code stream to obtain the content features corresponding to the content description information.
[0050] In a possible implementation, the decompression module is configured to entropy decode the code stream to obtain decoded features, and inverse quantize the decoded features to obtain the content features corresponding to the content description information.
[0051] In a possible implementation, the decompression module is further configured to obtain prior features. The code stream is entropy decoded according to distribution features contained in the prior features to obtain decoded features. The decoded features are inverse quantized according to quantization parameters in the prior features to obtain the content features corresponding to the content description information.
[0052] In a possible implementation, the content generation apparatus is specifically implemented as: a decompression module, configured to decompress the code stream to obtain a reduced dimension feature of the content feature corresponding to the content description information; and a generation module, configured to generate the reconstructed content corresponding to the content description information according to the reduced dimension feature.
[0053] In a possible implementation, the content generation apparatus is specifically implemented as: a generation module, configured to process the content feature to obtain a processed feature, and generate the reconstructed content corresponding to the content description information according to the processed feature.
[0054] In a possible implementation, the content generation apparatus is specifically implemented as: a generation module, configured to input the content feature into a diffusion model to obtain the processed feature.
[0055] In a possible implementation, the content generation apparatus is specifically implemented as: a communication module, further configured to send a content generation request to the server. Optionally, the content generation request contains the content description information.
[0056] It should be noted that the content generation apparatus of the fifth aspect can be a computer device, a server or a cloud server, or can be a chip (system) or other components or assemblies that can be arranged in the computer device, the server or the cloud server, and the present application does not limit this.
[0057] In addition, the technical effects of the content generation apparatus of the fifth aspect can refer to the technical effects of the content generation method of the first aspect, which will not be repeated here.
[0058] In a sixth aspect, an embodiment of the present application provides a content generation apparatus, which includes a feature generation module and a compression module.
[0059] The feature generation module is configured to generate a content feature corresponding to content description information according to the content description information, and the data amount of the feature is less than the data amount of the content corresponding to the content description information. The compression module is configured to compress the content feature to obtain a code stream.
[0060] In a possible implementation, the content generation apparatus is specifically implemented as: a compression module, configured to quantize the content feature to obtain the code stream.
[0061] In a possible implementation, the content generation apparatus is specifically implemented as: a compression module, configured to quantize the content feature to obtain the code stream when the network bandwidth between the server and the client is greater than or equal to a bandwidth threshold.
[0062] In a possible implementation, the content generation apparatus is specifically implemented as: a compression module, configured to quantize the content feature to obtain a quantized feature when the network bandwidth between the server and the client is less than a bandwidth threshold, and to perform entropy encoding on the quantized feature to obtain the code stream.
[0063] In a possible implementation, the compression module is specifically configured to input the content feature into the hyper-prior network to obtain a prior feature of the content feature. The content feature is quantized according to a quantization parameter contained in the prior feature to obtain a quantized feature. The quantized feature is entropy encoded according to a distribution feature contained in the prior feature to obtain the code stream.
[0064] In a possible implementation, the compression module is specifically configured to perform dimension reduction processing on the content feature to obtain a dimension-reduced feature of the content feature corresponding to the content description information. The dimension-reduced feature is compressed to obtain the code stream.
[0065] In a possible implementation, the content generation apparatus further includes a communication module. The communication module is configured to receive a content generation request sent by a client. The content generation request contains the content description information.
[0066] In a possible implementation, the feature generation module is specifically configured to obtain the stored content description information.
[0067] In a possible implementation, the communication module is further configured to send the code stream to the client.
[0068] In a possible implementation, the feature generation module is specifically configured to generate the content feature corresponding to the content description information according to the content description information and a generative model based on a diffusion model.
[0069] In a possible implementation, the content description information includes at least one of a text, an image, an audio, or a video.
[0070] It should be noted that the content generation apparatus of the sixth aspect can be a computer device, a server, or a cloud server, or can be a chip (system) or other components or assemblies that can be arranged in the computer device, the server, or the cloud server, and the present application does not limit this.
[0071] In addition, the technical effects of the content generation apparatus of the sixth aspect can refer to the technical effects of the content generation method of the second aspect, which will not be repeated here.
[0072] In a seventh aspect, the embodiments of the present application provide a content generation apparatus. The content generation apparatus includes a processor and a memory. The memory is configured to store a computer program. The processor is configured to execute the computer program stored in the memory, so that the content generation apparatus executes the content generation method in any one of the possible implementations of the first aspect, or executes the content generation method in any one of the possible implementations of the second aspect, or executes the content generation method in any one of the possible implementations of the third aspect.
[0073] In the present application, the content generation apparatus of the seventh aspect can be a computer device, a server or a cloud server, or a chip or chip system arranged inside the computer device, the server or the cloud server.
[0074] In the eighth aspect, the embodiments of the present application provide a computing device cluster, which includes at least one computing device. Each computing device includes a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method in the first aspect or any possible implementation manner of the first aspect, or performs the method in the second aspect or any possible implementation manner of the second aspect, or performs the method in the third aspect or any possible implementation manner of the third aspect.
[0075] In the ninth aspect, the embodiments of the present application provide a computer program product, when the instructions are executed by the computing device cluster, so that the computing device cluster performs the method in the first aspect or any possible implementation manner of the first aspect, or performs the method in the second aspect or any possible implementation manner of the second aspect, or performs the method in the third aspect or any possible implementation manner of the third aspect.
[0076] In the tenth aspect, the embodiments of the present application provide a computer readable storage medium, which includes computer program instructions, when the computer program instructions are executed by the computing device cluster, the computing device cluster executes the instructions in the computer program stored in the computer readable storage medium, so as to perform the method in the first aspect or any possible implementation manner of the first aspect, or perform the method in the second aspect or any possible implementation manner of the second aspect, or perform the method in the third aspect or any possible implementation manner of the third aspect.
[0077] The technical effects brought by any possible implementation manner of the seventh aspect to the tenth aspect can refer to the technical effects brought by different implementation manners of the first aspect to the third aspect. Here, no longer be repeated.
[0078] On the basis of the implementation manners of the above aspects, the present application can be further combined to provide more implementation manners. BRIEF DESCRIPTION OF DRAWINGS
[0079] Figure 1 It is a structural schematic diagram of the DALL-E 2 model in the related art;
[0080] Figure 2An image generation task end cloud transmission schematic diagram in the related art;
[0081] Figure 3 A VAE-based image compression model schematic diagram in the related art;
[0082] Figure 4 A system architecture schematic diagram of a content generation method provided by an embodiment of the present application;
[0083] Figure 5A A flow schematic diagram of a content generation method provided by an embodiment of the present application;
[0084] Figure 5B Another flow schematic diagram of a content generation method provided by an embodiment of the present application;
[0085] Figure 6 A visual feature generation schematic diagram provided by an embodiment of the present application;
[0086] Figure 7 A text semantic feature generation schematic diagram provided by an embodiment of the present application;
[0087] Figure 8 A feature compression flow schematic diagram based on prior features provided by an embodiment of the present application;
[0088] Figure 9 A rate distortion performance schematic diagram of a feature compression flow based on prior features and a JPEG compression mode provided by an embodiment of the present application;
[0089] Figure 10 A first structure schematic diagram of a content generation model provided by an embodiment of the present application;
[0090] Figure 11 A second structure schematic diagram of a content generation model provided by an embodiment of the present application;
[0091] Figure 12 A third structure schematic diagram of a content generation model provided by an embodiment of the present application;
[0092] Figure 13 A structure schematic diagram of a first subnetwork and a second subnetwork provided by an embodiment of the present application;
[0093] Figure 14 A fourth structure schematic diagram of a content generation model provided by an embodiment of the present application;
[0094] Figure 15 A fifth structure schematic diagram of a content generation model provided by an embodiment of the present application;
[0095] Figure 16 A structure schematic diagram of a content generation apparatus provided by an embodiment of the present application;
[0096] Figure 17 Another structural schematic diagram of a content generation apparatus provided by an embodiment of the present application is provided.
[0097] Figure 18 Another structural schematic diagram of a content generation apparatus provided by an embodiment of the present application is provided.
[0098] Figure 19 A structural schematic diagram of a computing device cluster provided by an embodiment of the present application is provided. DETAILED DESCRIPTION
[0099] First, the terms involved in the embodiments of the present application are explained.
[0100] Content generation technology: refers to a technology for automatically generating text, image, video and other content by using artificial intelligence algorithms. For example, content generation is performed by a generation model.
[0101] When the generation model processes content generation of different modalities, the content generation technology can be referred to as a multi-modal content generation technology.
[0102] For example, "text-to-image", "text-to-video" and "image-to-image".
[0103] Among them, text-to-image can refer to a generation model converting input text data into an image.
[0104] Text-to-video can refer to a generation model converting input text data into a video.
[0105] Image-to-image can refer to a generation model generating a new image based on an input image. The new image can be an image inconsistent with the style of the input image. For example, the input image is a printed font image, and the new image is a handwritten font image, etc. Or, the new image can be the input image after image restoration processing. Or, the new image can be the input image after image degradation processing.
[0106] Generation model: refers to a model for generating content by learning the distribution of input data. For example, DALL-E 2 model and diffusion model, etc.
[0107] Taking the DALL-E 2 model as an example, the DALL-E 2 model is a generation model based on the transformer architecture. The DALL-E 2 model is mainly used to generate images according to text data.
[0108] The DALL-E 2 model typically extracts text features from text data using a multimodal machine learning model (Contrastive Language–Image Pretraining, CLIP).
[0109] CLIP learns by contrasting pairs of images and text, mapping images and their associated text descriptions to a shared multi-dimensional model. This allows CLIP to extract visual content corresponding to different linguistic descriptions and to extract corresponding linguistic descriptions from images. Therefore, CLIP is used in both the training and application of the DALL-E 2 model to improve the image quality of generated images.
[0110] like Figure 1 As shown in Figure (a), during the training phase of the DALL-E 2 model, images are used as input to extract image features. These features are then converted into text features using CLIP, and the DALL-E 2 model is trained using these converted text features. Figure 1 As shown in Figure (b), during the image generation stage, CLIP text features are input into an autoregressive or diffusion prior to generate image features, which are then used to modulate the diffusion decoder to finally generate the image.
[0111] Taking the diffusion model as an example, the diffusion model can also be called the diffusion generation model. The diffusion model includes two processes: the forward process and the reverse process.
[0112] The forward process, also known as the diffusion process, refers to the gradual addition of Gaussian noise to the data until the data becomes random noise. The reverse process, on the other hand, uses Gaussian noise sampled from a standard normal distribution to gradually remove noise and obtain noise-free data.
[0113] In content generation technology, the reverse process of the diffusion model is generally used to reconstruct the content corresponding to the input data from the input data.
[0114] Multimodal content generation technology primarily relies on generative models. These models are computationally expensive and complex. To avoid the computational burden on client-side computing resources, the generative models are deployed in the cloud. Content generation is then achieved through interaction between the cloud and the client. This generates significant demands for edge-to-cloud data transfer. Figure 2As shown, the client sends text data to the cloud. The cloud inputs the text data into a generation model to generate an image. The cloud then compresses the generated image to generate a bitstream, and transmits the bitstream to the client. The client decompresses the bitstream to recover the reconstructed image.
[0115] Currently, the cloud can compress the generated image through various image compression methods. For example, Joint Photographic Experts Group (JPEG), High Efficiency Image File Format (HEIF), or Variational Auto-encoder (VAE) based end-to-end compression technology.
[0116] Taking the VAE based end-to-end image compression technology as an example, as shown in Figure 3 The VAE includes an encoding unit, an Arithmetic Encoder (AE) unit, an Arithmetic Decoder (AD) unit, an entropy model, and a decoding unit.
[0117] When compressing the generated content based on the VAE end-to-end image compression technology, the encoding unit is used to encode the input generated content to generate encoded data.
[0118] The entropy model is used to obtain the probability distribution of the generated content.
[0119] The AE unit is used to encode the encoded data into a bitstream according to the probability distribution of the generated content.
[0120] The AD unit is used to decode the bitstream into encoded data according to the probability distribution of the generated content.
[0121] The decoding unit is used to decode the decompressed encoded data to obtain the decompressed generated content.
[0122] In practical applications, the compression efficiency of JPEG, HEIF, and other technologies is low. The compression rate of the VAE based end-to-end image compression technology is low. When the resolution of the image generated by the cloud is large, the amount of data to be compressed increases, and a lot of time is spent on compressing the image. Moreover, the amount of bitstream data after compression is large, resulting in transmission delay. Moreover, when the network bandwidth between the cloud and the client is insufficient, transmission loss will occur, which will affect the image quality of the final reconstructed image. In addition, when the cloud generates other modal generated content, such as text data, other compression technologies need to be used, increasing the complexity of the model.
[0123] Based on this, in order to improve the compression rate, compression efficiency, reduce transmission loss, and at the same time be suitable for multi-modal generated content, the embodiment of the application provides a content generation method. Compared with directly compressing the content corresponding to the content description information, which results in a low compression rate, the content generation method provided by the embodiment of the application is that the server determines the content feature corresponding to the content description information according to the content description information, compresses the content feature to obtain a code stream, and feeds back the code stream to the client, so that the client decodes the code stream to generate the reconstructed content corresponding to the content description information. By reasonably generating the content feature, the quality of the generated content is ensured, and since the data amount of the content feature is much smaller than the data amount of the content, the data amount of the code stream obtained by compressing the content feature is smaller than the data amount of the code stream obtained by compressing the content, thereby effectively improving the compression rate of the code stream. In addition, since the data amount of the code stream is small, the occupation of the bandwidth is small, the code stream can be quickly transmitted, the transmission delay is shortened, and the code stream loss caused by the transmission of the code stream is reduced, the transmission delay is shortened, and the influence on the quality of the generated image is reduced.
[0124] It should be noted that the content generation method provided by the embodiment of the application can be applied to a multi-modal content generation scene. The content generation method provided by the embodiment of the application can also be used for data transmission services in other scenes, such as video image transmission in a live scene, data transmission in a storage scene, etc. The embodiment of the application does not limit this.
[0125] As shown in Figure 4 , the system architecture of the content generation method provided by the embodiment of the application is shown in Figure 4 . The system architecture of the content generation method shown includes a server 200, a network 300 and a client 100.
[0126] The client 100 and the server 200 are connected through the network 300. The network 300 provides a communication link medium between the client 100 and the server 200. The network 300 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, etc. before being published to the server 200.
[0127] In a possible implementation, the client 100 can be a browser, an application, a web application such as an H5 application or a light application, or a cloud application, and the like. The client 100 is obtained based on software development of a corresponding service provided by the server 200. The client 100 can be deployed in a terminal device, and needs to depend on a terminal device or a certain APP in the device for running, and the like. The terminal device can have a display screen and support information browsing, and the like. For example, the terminal device can be a mobile phone, a tablet computer, a personal computer, and the like. Various other types of applications can be configured in the terminal device, for example, man-machine dialogue applications, model training applications, text data applications, web browser applications, shopping applications, search applications, cloud desktop applications, cloud computer applications, and the like.
[0128] The server 200 can include servers providing various services. For example, a server providing a data forwarding service for a plurality of clients 100. For example, a server providing a training service for a model used on the client 100. For example, a server generating content for data sent by the client 100, and the like. It should be noted that the server 200 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server for a distributed system, or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host combined with artificial intelligence technology. Among them, the cloud server is used to provide cloud servers, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms, and the like. Basic cloud computing services.
[0129] In the embodiment of the present application, in the multi-modal content generation scenario, the server 200 is configured to obtain content description information, execute the content generation method provided in the embodiment of the present application to generate a code stream, and return the code stream to the client 100.
[0130] The client 100 is configured to receive the code stream, execute the content generation method provided in the embodiment of the present application to generate the code stream, and obtain reconstructed content.
[0131] In an example, the client 100 can be configured to send the content description information to the server 200.
[0132] In another example, the server 200 can read the content description information from other storage devices.
[0133] In another example, the server 200 can read the content description information from a storage area of the server 200.
[0134] In a first possible implementation, in actual application, the system architecture can include a server 200 and a plurality of clients 100. The plurality of clients 100 can establish a communication connection through the server 200. In a multi-modal content generation scenario, the server 200 executes the content generation method provided in the embodiments of the present application to generate a code stream and provide a content generation service to the plurality of clients 100.
[0135] Alternatively, the plurality of clients 100 can respectively serve as a sending end or a receiving end, and the plurality of clients 100 can realize communication through the server 200. A user can interact with the server 200 through the client 100 to receive data sent by other clients 100 or send data to other clients 100, and the like.
[0136] For example, in multi-modal content generation, a user can send content description information to the server 200 through a client A. The server generates features corresponding to the content description information, and pushes the code stream compressed from the features corresponding to the content description information to a client B.
[0137] In a second possible implementation, the client 100 can have similar functions as the server 200, and thus execute the content generation method provided in the embodiments of the present application to generate a code stream.
[0138] For example, in multi-modal content generation, a user can send content description information to the server 200 through a client A. The server pushes the code stream compressed from the features corresponding to the content description information to a client B. The client B generates features corresponding to the content description information, and pushes the code stream compressed from the features corresponding to the content description information to the server 200. The server 200 returns the code stream pushed by the client B to the client A.
[0139] For another example, in multi-modal content generation, a user can send content description information to a client B through a client A. The client B generates features corresponding to the content description information, and pushes the code stream compressed from the features corresponding to the content description information to the client A.
[0140] For another example, the client A generates features corresponding to the content description information based on the content description information, and pushes the code stream compressed from the features corresponding to the content description information to the client B.
[0141] The content generation method provided in the basic application embodiments is introduced as follows. Figure 4 The content generation method provided in the embodiments of the present application is generally generated by the server 200. Here, the client 100 and the server 200 execute the content generation method provided in the embodiments of the present application as an example.
[0142] As shown in FIG. 1, the content generation method provided in the embodiments of the present application includes the following steps. Figure 5A Figure 5A is a flowchart of a content generation method provided by an embodiment of the present application. The content generation method shown in the flowchart includes S510 to S550.
[0143] S510, the client sends a content generation request to the server.
[0144] The content generation request carries content description information.
[0145] In a possible implementation, the content description information is used to describe the generated content. For example, the content description information includes but is not limited to text data, image data, etc. For another example, the content description information includes but is not limited to text features, image features, etc.
[0146] Taking the content description information including text features and image features as an example. The client obtains initial information, performs preliminary content generation on the initial information, and obtains at least one of text features or image features. The client sends a content generation request to the server based on at least one of the text features or the image features. The initial information includes but is not limited to text data, image data, etc.
[0147] Taking the content description information including but not limited to text data, image data, etc. as an example, when the content description information is text data, the content generation request can be used to instruct the server to perform image generation, text translation, text summary generation, text writing, code generation, etc. When the content description information is image data, the content generation request can be used to instruct the server to perform video generation, image style transfer, image restoration, image degradation, text generation, etc.
[0148] The image style transfer can refer to converting images of different styles, such as image color conversion, image background conversion, handwritten font image conversion to printed font image, etc. The image restoration can be image quality. The image degradation can refer to generating a blurred image, generating a noise image, etc.
[0149] Correspondingly, when the content description information includes text data and image data, the content generation request can be used to instruct the server to perform image recognition, image text synthesis, etc.
[0150] The image recognition can be to recognize a target object indicated by the text data in the input image. For example, when the target object indicated by the text data is a cat, the server recognizes the cat in the input image and labels the image region where the cat is located in the input image.
[0151] The image text synthesis can refer to performing image style transfer, image restoration, image degradation, etc. on the input image according to the text data. Or, adding subtitles in the input image according to the text data, etc.
[0152] In a possible implementation, the client obtains content description information, and sends a content generation request to the sender based on the content description information. In actual application, the content description information can be obtained in various ways, which are selected according to actual conditions. The embodiments of the present application do not limit this. In an example, the content description information can be read from other data acquisition devices or databases. In another example, the content description information input by the user can be accepted.
[0153] In S520, the server generates content features corresponding to the content description information according to the content description information.
[0154] In the embodiments of the present application, the content features corresponding to the content description information can indicate the context information of the content corresponding to the content description information.
[0155] For example, the content corresponding to the content description information is image data, and the content features corresponding to the content description information include but are not limited to visual features, image semantic features and style features, etc. The visual features represent the overall texture information of the content description information, and the image semantic features represent the semantic information of the content description information. The style features represent the image style information of the content description information. The image data includes a single image, an image sequence or a video image.
[0156] For another example, the content corresponding to the content description information is text, and the content features corresponding to the content description information include but are not limited to text semantic features and text style features. The text semantic features represent the context semantic information of the content description information, and the text style features represent the tone style of the content description information. For example, in the case of image-to-text, the server generates features according to the input image to obtain the image text semantic features. For example, in the case of text summarization, the server generates features according to the input text data to obtain the text semantic features and the text style features in the text data.
[0157] For another example, the content description information is text data and image data, and the content features corresponding to the content description information include image-text fusion features. The image-text fusion features are used to represent one or more of the overall texture information of the image, the semantic information of the image, the image style information, the context semantic information of the text and the text style features.
[0158] In the embodiments of the present application, the server generates the content features corresponding to the content description information by feature generation on the content description information.
[0159] In a first possible implementation, the server is deployed with a generation model, and the server inputs the content description information into an encoder to generate features through the generation model.
[0160] The generation model can be a diffusion generation model. Alternatively, the generation model can be a large language model. Alternatively, the generation model can be a neural network model built based on a transformer.
[0161] In an example, to improve the feature generation effect of the content generation model, the generation model needs to be trained.
[0162] For example, after the generation model is deployed on the server, the encoder is trained by the sample description information and the sample content corresponding to the sample description information, and a training loss output by the generation model is obtained. The model parameters of the generation model are adjusted according to the training loss. When the generation model meets the convergence condition, the generation model training is stopped.
[0163] The model parameters include weights of the generation model, etc. The convergence condition can be that the training loss of the generation model is less than or equal to a loss threshold. Alternatively, the convergence condition can be that the number of generation model training is greater than or equal to a training number threshold.
[0164] For another example, before the generation model is deployed on the server, the initial model is trained by the sample description information and the sample content corresponding to the sample description information, and a training loss output by the initial model is obtained. The model parameters of the initial model are adjusted according to the training loss. When the content generation model meets the convergence condition, the initial model training is stopped, and the generation model is obtained.
[0165] In some embodiments, the generation model includes a content generation network layer. The content generation network layer is used for feature generation on input content description information to obtain content features corresponding to the content description information.
[0166] In other embodiments, the generation network includes a content generation network layer and a data encoding network layer.
[0167] The data encoding network layer is used for feature embedding of the content description information to convert the content description information into a vector code that can be processed by the content generation encoder.
[0168] The feature embedding can refer to converting the content description information into a vector code.
[0169] For example, when the content description information includes text data, the feature embedding can be converting the content description information into a word vector code.
[0170] For another example, when the content description information includes text data, the feature embedding can be converting the content description information into an image code.
[0171] In an example, the data encoding network layer can be used for feature embedding of text data. Alternatively, it is used for feature embedding of image data.
[0172] S530, the server compresses the content feature to obtain a code stream, and sends the code stream to the client.
[0173] Since the content feature has a smaller data amount than the generated content, the embodiment of the present application directly compresses the content feature corresponding to the content description information after obtaining the content feature corresponding to the content description information, thereby avoiding the problem of low compression efficiency caused by compressing the generated content.
[0174] In the embodiment of the present application, the server can compress the content feature corresponding to the content description information by using a lossless compression method to obtain a code stream. Or the server can compress the content feature corresponding to the content description information by using a lossy compression method to obtain a code stream. The embodiment of the present application is not limited in comparison. In actual application, the server can select the lossless compression method or the lossy method according to actual business requirements.
[0175] S540, the client decompresses the code stream to obtain the content feature corresponding to the content description information.
[0176] Since, in the case that the server compresses the content feature corresponding to the content description information by using the lossy compression method, there is a compression error between the decompressed feature and the content feature corresponding to the content description information, and the compression error is less than or equal to a first compression error threshold. In the case that the server compresses the content feature corresponding to the content description information by using the lossless compression method, the decompressed feature is the same as the content feature corresponding to the content description information, or the compression error between the decompressed feature and the content feature corresponding to the content description information is less than or equal to a second compression error threshold. Wherein, the first compression error threshold is greater than the second compression error threshold. Therefore, the feature after the client decompresses the code stream can be approximately the content feature corresponding to the content description information generated by the server. In this way, the reconstructed content generated based on the decompressed feature can meet the content description information.
[0177] S550, the client generates reconstructed content corresponding to the content description information according to the content feature.
[0178] In the embodiment of the present application, the reconstructed content is the content data generated by the content description information according to the content generation request.
[0179] For example, taking the content description information as image data as an example, the reconstructed content can refer to the content description information after style migration. For example, when the content description information is a handwritten font image, the reconstructed content can be a printed font image. Or the reconstructed content can refer to the content description information after image restoration processing. For example, when the content description information is a foggy image, the reconstructed content can be a de-fogged image.
[0180] Exemplarily, taking the content description information as text data, the reconstructed content can refer to an image generated based on the content description information.
[0181] Exemplarily, taking the content description information as including image data and text data, the reconstructed content can be a video image with subtitles added.
[0182] In a first possible implementation, a content generation decoder is deployed in the client. The client inputs the content feature into the content generation decoder, and generates the content feature through the content generation decoder to obtain the reconstructed content corresponding to the content description information.
[0183] The content generation decoder can be a decoder based on VAE.
[0184] In an example, the client can be deployed with multiple content generation decoders. Each content generation decoder is used to generate reconstructed content of different modalities. For example, the client is deployed with an image generation decoder, a text generation decoder, a video generation decoder, and the like. The client selects a corresponding content generation decoder according to the modality to which the content description information belongs and the content generation request. For example, when the content description information is text data and the content generation request is a generation request indicating image generation, the client inputs the content feature into the image generation decoder.
[0185] In yet another example, the content generation decoder deployed in the client is integrated with multiple decoding units. Each decoding unit is used to generate reconstructed content of different modalities.
[0186] Based on Figure 5A With the embodiments provided, in content generation, the server directly compresses the content feature corresponding to the content description information. Compared with the manner of compressing the generated content in the related art, since the data amount of the code stream is smaller than the data amount of the code stream formed by directly compressing the content in the related art, compared with the manner of compressing the generated content in the related art, the compression rate of the code stream can be improved in the embodiments of the present application. Moreover, the data amount of the feature is smaller than the data amount of the content, and the compression efficiency can be improved. In addition, the data amount of the code stream is small in the embodiments of the present application, the bandwidth occupation is small, the transmission efficiency can be improved, and thus the transmission loss can be reduced.
[0187] In a possible implementation, multiple content features corresponding to different content description information are stored in the server. The server sends the code stream to the client based on the request of the client. The client generates the content corresponding to the feature according to the code stream.
[0188] As Figure 5B shown, Figure 5B is another flowchart of the content generation method provided by the embodiments of the present application, and the content generation method shown includes S610 to S640:
[0189] S610, the client sends a feature acquisition request to the server.
[0190] The feature acquisition request carries description information corresponding to the content to be generated. The description information can be the content description information in S510.
[0191] S620, the server sends a code stream to the client.
[0192] In a first possible implementation, the server acquires a code stream matching the description information based on the description information, and sends the code stream to the client.
[0193] For example, the server can acquire multiple content description information, generate content features corresponding to each content description information according to S520, and compress the content features corresponding to each content description information according to S530 to form a code stream. When the server receives the feature acquisition request, it acquires a code stream matching the description information from the stored multiple code streams based on the description information carried by the feature acquisition request.
[0194] In a second possible implementation, the server acquires a content feature corresponding to the feature acquisition request based on the description information, compresses the content feature corresponding to the feature acquisition request to obtain a code stream, and sends the code stream to the client.
[0195] In the embodiments of the present application, the server can acquire multiple content description information, generate content features corresponding to each content description information according to S520, and acquire a content feature matching the description information carried by the feature acquisition request from the stored multiple content features corresponding to the content description information based on the description information carried by the feature acquisition request when the server receives the feature acquisition request. The content feature matching the description information carried by the feature acquisition request is used as the content feature corresponding to the feature acquisition request.
[0196] In an example, the server can receive content description information sent by other clients, generate content features corresponding to the content description information, and store the content description information and the content features corresponding to the content description information in association.
[0197] In another example, the server can read pre-stored content description information.
[0198] In another example, the server can generate content description information. For example, the server generates multiple content description information and content features corresponding to the content description information through a generation model.
[0199] S630, the client decompresses the code stream to obtain the content feature corresponding to the feature acquisition request.
[0200] In the embodiment of the application, the client can decompress the code stream according to the above step S540 to obtain the content feature corresponding to the feature acquisition request, which is not described herein.
[0201] S640, the client generates content according to the content feature to obtain the reconstructed content.
[0202] In the embodiment of the application, the client can generate content according to the above step S550 to obtain the reconstructed content, which is not described herein.
[0203] Based on Figure 5B According to the provided embodiment, the server pre-generates content features corresponding to a plurality of different content description information. The client directly requests the server for the features, and the server does not need to generate the content features in real time based on the content description information sent by the client during content generation, thereby improving the content generation efficiency. Moreover, the code stream sent by the server to the client is obtained by compressing the features. Since the data amount of the code stream is smaller than that of the code stream formed by directly compressing the content in the related art, compared with the compression method of the generated content in the related art, the embodiment of the application can improve the compression rate of the code stream. Moreover, the data amount of the features is smaller than that of the content, which can improve the compression efficiency. In addition, the data amount of the code stream is small in the embodiment of the application, which occupies a small bandwidth, and can improve the transmission efficiency, thereby reducing the transmission loss.
[0204] Next, the embodiment provided by the application will be taken as an example to introduce the acquisition method of the content feature corresponding to the content description information. Figure 5A
[0205] In the embodiment of the application, the content generation request also carries at least one generation condition corresponding to the content description information. The server generates the content feature corresponding to the content description information according to the at least one generation condition corresponding to the content description information. In this way, the content feature corresponding to the actual generation requirement is obtained through the generation condition. The unnecessary content feature is avoided, and the data processing amount in the content feature generation stage is reduced. Moreover, the generation of the content feature is guided by the generation condition, and it is ensured that the finally generated reconstructed content meets the actual generation requirement.
[0206] The generation condition is used to indicate the content generation requirement. In an example, the generation condition can be a synthesis generation condition, an augmentation generation condition, a restoration generation condition, or an imitation generation condition, etc.
[0207] The synthesis generation condition can be referred to as a synthesis mode. For example, the synthesis of an image into a video, the synthesis of an image and noise, etc.
[0208] The augmented generation condition can be referred to as an augmented mode. For example, text generation, etc.
[0209] The restoration generation condition can be referred to as a restoration mode. For example, image restoration, text restoration, etc.
[0210] The imitation generation condition can be referred to as an imitation mode. For example, converting a printed text image into a handwritten font image, etc.
[0211] In a possible implementation, the server can generate, according to at least one generation condition corresponding to the content description information, a content feature corresponding to the content description information from the content description information.
[0212] Taking the content description information as image data as an example, when the generation condition includes a synthesis generation condition, the server generates an image semantic feature from the content description information.
[0213] When the generation condition includes an augmented generation condition, the server generates a visual feature from the content description information.
[0214] When the generation condition includes a restoration generation condition, the server generates an image semantic feature and a visual feature from the content description information.
[0215] When the generation condition includes an imitation generation condition, the server generates an image semantic feature and a style feature from the content description information.
[0216] It should be noted that when the content description information corresponds to multiple generation conditions, the server generates, according to each generation condition, a content feature corresponding to the content description information corresponding to each generation condition from the content description information. For example, taking the content description information as image data as an example, when the generation condition includes a restoration generation condition and an imitation generation condition, the server generates an image semantic feature, a style feature, and a visual feature from the content description information.
[0217] It should be noted that in actual applications, the server can obtain the at least one generation condition corresponding to the content description information in multiple ways. The corresponding way can be selected according to the actual application scene. The embodiments of the present application do not limit this. In a first possible implementation, the at least one generation condition corresponding to the content description information can be read from other data acquisition devices or databases. In a second possible implementation, the at least one generation condition corresponding to the content description information input by the user can be received.
[0218] Taking the content description information as image data and the content feature corresponding to the content description information as a visual feature as an example, as shown in Figure 6 Figure 6 is a schematic diagram of generating visual features provided by an embodiment of the present application. The content description information is input into an image encoder to obtain initial visual features of the content description information. Based on the initial visual features, image block indexes of the content description information are obtained. The image block indexes are embedded and coded to obtain embedded visual features of the content description information. Based on the embedded visual features and the initial visual features, visual features of the content description information are obtained.
[0219] For example, the embedded visual features and the initial visual features are subjected to pooling processing to obtain the visual features of the content description information.
[0220] For example, the embedded visual features and the initial visual features are subjected to pooling processing to obtain the visual features of the content description information. Figure 7 Figure 7 is a schematic diagram of generating text semantic features provided by an embodiment of the present application. The content description information is input into a text encoder to obtain text data encoding. Character indexes of the first data content are obtained. The character indexes of the first data content are embedded and coded to obtain embedded semantic features of the content description information. Based on the embedded semantic features and the text data encoding, text semantic features of the content description information are obtained.
[0221] For example, the embedded semantic features and the text data encoding are subjected to linear projection to obtain the text semantic features of the content description information.
[0222] In an embodiment of the present application, when the content description information includes content data of multiple modalities, the content data of each modality in the content description information can be encoded respectively to obtain content data encoding of each modality. The content data encoding of each modality is fused to obtain fused data encoding. The fused data encoding is input into a generation model based on a diffusion model for feature generation to obtain content features corresponding to the content description information.
[0223] For example, the content description information includes text data and image data. The text data is encoded by a text encoder to obtain text data encoding. The image data is encoded by an image encoder to obtain image data encoding. The text data encoding and the image data encoding are spliced to obtain fused data encoding.
[0224] Next, taking a lossy compression mode as an example, the compression mode of the content features corresponding to the content description information is introduced.
[0225] In a first possible implementation, the server can quantize the content features corresponding to the content description information to convert the content features corresponding to the content description information into a code stream. Through quantization, the content feature compression efficiency is improved, and the content generation efficiency is improved. Moreover, through quantization, the data amount of the code stream can be reduced, and the transmission efficiency is improved.
[0226] The quantization processing mode includes, but is not limited to, uniform quantization, affine quantization, vector quantization, and the like.
[0227] Taking uniform affine quantization as an example, the quantization points corresponding to the integer values are made to coincide with the content feature distribution interval corresponding to the content description information as much as possible according to the following formula (1), to obtain the code stream.
[0228]
[0229] wherein round() represents performing a rounding operation, q represents the quantized feature, x represents the content feature, and s represents the quantization parameter. In formula (1), the quantization parameter s can be determined by the range of the feature values in the content feature corresponding to the content description information and the range of the integer values.
[0230] Taking vector quantization as an example, a plurality of feature elements in the content feature corresponding to the content description information are combined into a vector, each vector is coded by a codebook, and the content feature corresponding to the content description information is converted into a code stream file.
[0231] In the first example, the server can perform quantization processing on the content feature corresponding to the content description information, to obtain the code stream.
[0232] In the second example, the server can perform down-sampling on the content feature corresponding to the content description information, to obtain down-sampled features. The server performs quantization processing on the down-sampled features, to obtain the code stream.
[0233] In the third example, the server can perform coding on the content feature corresponding to the content description information, to obtain content feature coding. The content feature coding is quantized to obtain discrete content features corresponding to the content description information, and the discrete content features corresponding to the content description information are down-sampled to obtain the code stream.
[0234] In the embodiments of the present application, there are multiple ways to code the content feature corresponding to the content description information. In actual applications, a corresponding coding mode can be selected according to a business scenario to code the content feature corresponding to the content description information, such as predictive coding, transform coding, and the like.
[0235] For example, taking predictive coding as an example, the server predicts the content feature corresponding to the content description information to obtain a predicted value of the content feature. The predicted value of the feature is compared with the actual value of the feature to obtain a prediction error. The prediction error is coded to obtain the content feature coding.
[0236] In the second possible implementation, since the quantization processing can lose features and the quantization processing is irreversible, which will cause an error between the decompressed features and the content features corresponding to the content description information. Therefore, to reduce the error between the decompressed features and the content features corresponding to the content description information and improve the data quality of the reconstructed content, the server performs quantization processing on the content features corresponding to the content description information to obtain discrete content features. The discrete content features are entropy encoded to obtain a code stream. In this way, the feature loss is reduced by adding entropy encoding. Furthermore, the error between the decompressed features and the content features corresponding to the content description information is reduced, and the data quality of the reconstructed content is improved.
[0237] In actual applications, the entropy encoding techniques include, but are not limited to, Huffman encoding, arithmetic encoding, run-length encoding (RLE), context-based adaptive variable length coding (CAVLC), context-based adaptive binary arithmetic coding (CABAC), and asymmetric numeral systems (ANS), etc.
[0238] For example, the server obtains a probability distribution of the discrete content features by ANS. The discrete content features are encoded based on the probability distribution to obtain a code stream.
[0239] In an example, the server can perform quantization processing on the content features corresponding to the content description information to obtain discrete content features according to the first possible implementation of the compression mode described above. The embodiments of the present application do not repeat here.
[0240] In the third possible implementation, since the quantization mode can obtain higher compression efficiency, the transmission efficiency can be improved. However, in the case of small network bandwidth, there can be transmission loss. The compression time of the combination of quantization and entropy encoding is longer than that of the quantization mode, and the combination of quantization and entropy encoding can obtain greater compression rate, and the data amount of the code stream obtained by compression is smaller. Therefore, to ensure the compression efficiency and improve the data quality of the reconstructed content, the server selects the quantization mode or the combination of quantization and entropy encoding to compress the content description information based on the network bandwidth between the server and the client after obtaining the content features corresponding to the content description information. In this way, the corresponding compression mode is selected flexibly through the network bandwidth, the adaptation degree between the compression mode and the network bandwidth is improved, and thus the compression efficiency is improved and the compression quality is ensured.
[0241] In an example, the server obtains a network bandwidth between the server and the client. The network bandwidth is compared with a bandwidth threshold. If the network bandwidth is less than the bandwidth threshold, the content feature corresponding to the content description information is quantized to obtain discrete content features. The discrete content features are entropy encoded to obtain the code stream. If the network bandwidth is greater than or equal to the bandwidth threshold, the content feature corresponding to the content description information is quantized to obtain the code stream.
[0242] In yet another example, the compression efficiency and throughput of the quantization manner are greater than those of the combination of quantization and entropy encoding. In a service scenario with real-time requirements, the quantization manner is selected for compression, which can reduce the waiting time of the client. For example, in a live service scenario, the video transmission time can be shortened by the quantization manner, and the video stuttering and frame dropping problems caused by the increase in compression time can be avoided. Therefore, to further improve the applicability and flexibility of the content generation method, the server obtains a network bandwidth between the server and the client and a transmission delay of the client. When the network bandwidth and the transmission delay meet a first preset requirement, the content feature corresponding to the content description information is quantized to obtain the code stream. When the network bandwidth and the transmission delay meet a second preset requirement, the content feature corresponding to the content description information is quantized to obtain discrete content features. The discrete content features are entropy encoded to obtain the code stream.
[0243] The transmission delay is used to indicate the real-time requirement of the client for the generated content. It can be understood that the higher the real-time requirement of the client for the generated content, the smaller the corresponding transmission delay. Correspondingly, the lower the real-time requirement of the client for the generated content, the larger the corresponding transmission delay.
[0244] The preset requirement is used to indicate a transmission delay threshold and a bandwidth threshold. The bandwidth threshold indicated by the first preset requirement is greater than the bandwidth threshold indicated by the second preset requirement. The transmission delay threshold indicated by the first preset requirement is greater than the transmission delay threshold indicated by the second preset requirement.
[0245] In a fourth possible implementation manner, the server performs dimension reduction processing on the content feature corresponding to the content description information to obtain a dimension-reduced feature of the content feature corresponding to the content description information. The dimension-reduced feature is compressed to obtain the code stream.
[0246] In an example, the server can perform dimension reduction processing on the content feature corresponding to the content description information by a principal component analysis manner. In yet another example, the server can perform dimension reduction processing on the content feature corresponding to the content description information by a content generation decoder.
[0247] Next, the compression manner of the content feature corresponding to the content description information is introduced by taking the combination of quantization and entropy encoding as an example.
[0248] In the embodiments of the present application, in order to further reduce feature compression loss and improve the data quality of the finally generated content data. In a possible implementation, the server obtains the prior feature of the content feature corresponding to the content description information when compressing the content feature corresponding to the content description information based on the combination of quantization and entropy coding. The prior feature is used to guide the feature compression and remove the redundant information in the content feature corresponding to the content description information. In this way, the compression performance is improved by using the prior feature, the compression quality is improved, and the data quality of the finally generated reconstructed content is ensured.
[0249] The prior feature quantization parameter and the distribution feature of the content feature corresponding to the content description information are obtained.
[0250] As shown in Figure 8 , the prior feature quantization parameter and the distribution feature of the content feature corresponding to the content description information are obtained. Figure 8 is a feature compression process based on prior features provided by the embodiments of the present application. The feature compression process based on prior information includes S531 to S533.
[0251] S531, the server generates a prior feature based on a content feature corresponding to content description information.
[0252] In the embodiments of the present application, the server obtains the quantization parameter, the variance and the mean by capturing the distribution of the feature value in the content feature corresponding to the content description information, and then obtains the prior feature corresponding to the content feature corresponding to the content description information.
[0253] In a possible implementation, the server can capture the distribution of the feature value in the content feature corresponding to the content description information by using a hyper-prior network to obtain the prior feature corresponding to the content feature corresponding to the content description information.
[0254] In an example, the hyper-prior network includes a first sub-network, a quantizer sub-network and a second sub-network.
[0255] The server encodes the content feature corresponding to the content description information through the first sub-network to generate a latent representation feature. The latent representation feature is quantized and entropy encoded through the quantizer sub-network to generate a discrete latent representation feature. The discrete latent representation feature is decoded through the second sub-network to generate the prior feature corresponding to the content feature corresponding to the content description information.
[0256] S532, quantize the content feature corresponding to the content description information according to the quantization parameter to obtain a discrete content feature.
[0257] In a possible implementation, the server quantizes the content feature corresponding to the content description information according to the quantization parameter according to formula (2) below, to obtain discrete content features.
[0258]
[0259] wherein h is a code rate parameter, and the code rate parameter is used to indicate a compression code rate.
[0260] It should be noted that in actual application, there are multiple ways to obtain the code rate parameter. The code rate parameter can be selected according to actual conditions. The embodiments of the present application do not limit this. In an example, the code rate parameter can be read from other data acquisition equipment or a database. In another example, the code rate parameter can be obtained through a feature map. For example, the server generates a feature map through a BETA model for the content feature corresponding to the content description information. The code rate parameter is extracted from the feature map.
[0261] In another possible implementation, the server can set a quantizer according to the quantization parameter and the code rate parameter. The quantizer is used to quantize the content feature corresponding to the content description information, to obtain discrete content features.
[0262] For example, the quantization parameter is adjusted to adjust the quantizer. The quantizer is used to quantize the content feature corresponding to the content description information, to obtain discrete content features.
[0263] In another possible implementation, the server obtains the code rate parameter. The code rate parameter and the quantization parameter are used to quantize the content feature corresponding to the content description information, to obtain discrete content features.
[0264] S533, entropy encoding the discrete content features according to the distribution features, to obtain a code stream.
[0265] In a possible implementation, the server obtains the probability distribution information of the feature value of the content feature corresponding to the content description information according to the mean and the variance. The discrete content features are entropy encoded according to the probability distribution information of the feature value of the content feature corresponding to the content description information, to obtain a code stream.
[0266] The Gaussian distribution model can be established through the mean and the variance, and the probability distribution information of the feature value of the content feature corresponding to the content description information can be obtained through the Gaussian distribution model.
[0267] In the embodiments of the present application, the feature quality of the content features generated by the client is improved. A second sub-network is deployed in the client. The server sends discrete latent representation features to the client when sending the code stream to the client. The client decodes the discrete latent representation features through the second sub-network to obtain prior features. The code stream is generated according to the prior features to obtain the content features corresponding to the content description information.
[0268] Based on Figure 8 In the embodiments provided, in feature compression, the redundancy between the content features corresponding to the content description information is fully captured through the prior features of the content features corresponding to the content description information. In this way, the compression performance is improved through the prior features, the compression quality is improved, and the data quality of the finally generated reconstructed content is ensured.
[0269] For example, taking image data as the reconstructed content, as shown in Figure 9 , the prior feature-based feature compression process provided by the embodiments of the present application and the rate-distortion performance based on the JPEG compression mode are shown in Figure 9 As can be seen from Figure 9 , the compression efficiency of the prior feature-based feature compression process (GFC) is more than 10 times that of the compression efficiency based on the JPEG compression mode. Moreover, the peak signal-to-noise ratio (PSNR) of the generated image based on the prior feature-based feature compression process can reach 65db, while the maximum peak signal-to-noise ratio of the generated image based on the JPEG compression mode is 46db. It can be seen that the prior feature-based feature compression process provided by the embodiments of the present application can improve the compression quality while ensuring the compression efficiency, and approaches the lossless compression quality.
[0270] In addition, when the prior feature-based feature compression process and the JPEG compression mode have similar compression code rates, the image generated by the prior feature-based feature compression process provided by the embodiments of the present application has a square root visual effect compared with the JPEG compression mode. Moreover, when the prior feature-based feature compression process and the JPEG compression mode have similar visual effects, the prior feature-based feature compression process has a lower compression rate than the JPEG compression mode. For example, the compression rate of the prior feature-based feature compression process is 0.22, and the image quality of the compression rate of the JPEG compression mode is 4.03.
[0271] Next, the way of decompressing the code stream by the client is introduced.
[0272] In a first possible implementation, the code stream is obtained according to a content feature corresponding to the service-end compressed content description information. The client can decompress the code stream according to inverse operation of the compression manner of the content feature corresponding to the content description information to obtain the content feature.
[0273] In a first example, the client can perform inverse quantization on the code stream to obtain the content feature.
[0274] The inverse quantization is used to restore the code stream to continuous content features. In some embodiments, the inverse quantization manner includes uniform inverse quantization, affine inverse quantization, vector inverse quantization, or the like.
[0275] For example, taking the uniform inverse quantization as an example, the code stream can be linearly restored to a floating-point value by using the following formula (3) to obtain the content feature.
[0276] x = (q - z) * s (3)
[0277] In some embodiments of the present application, the client can directly perform inverse quantization on the code stream to obtain the content feature.
[0278] Alternatively, the client can perform inverse quantization on the code stream to obtain encoded data. The encoded data is decoded to obtain the content feature, for example, by using a prediction coding, transform coding, or the like.
[0279] Alternatively, the client can perform inverse quantization on the code stream to obtain down-sampled features. The down-sampled features are interpolated to obtain the content feature.
[0280] In a second example, the client performs entropy decoding on the code stream to obtain decoded features. The client performs inverse quantization on the decoded features to obtain the content feature.
[0281] For example, the client obtains a probability distribution of the content feature corresponding to the content description information. The code stream is entropy decoded according to the probability distribution to obtain decoded features. The client performs inverse quantization on the decoded features according to the inverse quantization manner to obtain the content feature.
[0282] In a third example, the client obtains prior features. The client performs entropy decoding on the code stream according to distribution features in the prior features to obtain decoded features. The client performs inverse quantization on the decoded features according to quantization parameters in the prior features to obtain the content feature.
[0283] For example, the client obtains probability distribution information of the content feature corresponding to the content description information based on the distribution features in the prior features. The code stream is decoded according to the probability distribution information to obtain decoded features. The client performs inverse quantization on the decoded features according to the quantization parameters in the prior features to obtain the content feature.
[0284] In an embodiment, the client can suggest a Gaussian distribution model according to the mean and variance in the prior feature, and obtain the probability distribution information of the feature value of the content feature corresponding to the content description information through the Gaussian distribution model.
[0285] In an embodiment, the client can perform dequantization processing on the decoded feature according to the quantization parameter according to the following formula (4) to obtain the content feature.
[0286] x = p * s * h (4)
[0287] Wherein, p represents the decoded feature.
[0288] In the embodiments of the present application, there are multiple ways for the client to obtain the prior feature. For example, in an example, the client can accept the prior feature sent by the server. For another example, in another example, the client receives the discrete latent representation feature sent by the server, and decodes the discrete latent representation feature to obtain the prior feature. For example, the second subnetwork in the hyper-prior network is deployed in the client, and the second subnetwork is used to decode the discrete latent representation feature to generate the prior feature corresponding to the content feature corresponding to the content description information.
[0289] In a second possible implementation, the code stream is obtained by compressing the dimension-reduced feature of the content feature corresponding to the content description information by the server. The client can refer to the decompression manner in the first possible implementation to decompress the code stream to obtain the dimension-reduced feature of the content feature corresponding to the content description information. The embodiments of the present application do not repeat the description.
[0290] Next, the method for generating content by the client is introduced.
[0291] In a possible implementation, the client can directly generate the reconstructed content corresponding to the content description information according to the content feature.
[0292] In a second possible implementation, to improve the data quality of the reconstructed content, the decompressed content feature can be processed through a diffusion model to obtain a processed feature. The reconstructed content corresponding to the content description information is generated according to the processed feature. In this way, the diffusion model is used to reduce the noise in the decompressed feature, thereby improving the data quality of the reconstructed content. Further, the data quality of the generated reconstructed content is ensured.
[0293] For example, the decompressed content feature can be processed to reduce noise. For another example, the decompressed content feature can be processed to generate features,
[0294] Wherein, the decompressed content feature can be processed through the reverse process of the diffusion model.
[0295] In an example, the client can input the processed feature content to the content generation decoder to obtain the reconstructed content, referring to S550 described above.
[0296] The diffusion model is trained by the sample content data, the sample content feature of the sample content data, and the noise sample data. The noise sample data is obtained by adding sample noise to the sample content feature. The sample content feature is obtained by performing feature extraction on the sample content data.
[0297] In a possible implementation of the present application, the client directly displays the reconstructed content to the user after generating the reconstructed content.
[0298] In another possible implementation of the present application, the client receives content generation requirement information input by the user after generating the reconstructed content, and judges whether the currently generated reconstructed content meets the content generation requirement information. If the currently generated reconstructed content meets the content generation requirement information, the client directly displays the reconstructed content to the user. If not, the client sends a content generation request to the server again based on the content generation requirement information and the reconstructed content. The server generates a new code stream based on the content generation request sent again. The client generates a new reconstructed content based on the new code stream. This process is repeated until the generated reconstructed content meets the content generation requirement information, and the reconstructed content meeting the content generation requirement information is displayed to the user.
[0299] The content generation requirement information includes, for example, a modality of the reconstructed content, a quantity of the reconstructed content, a data quality of the reconstructed content, and the like. The specific selection is made according to actual conditions, which is not limited in the embodiments of the present application. For example, when the reconstructed content is an image, the content generation requirement information includes image definition, image quantity, and the like.
[0300] The above mainly describes the content generation method provided by the embodiments of the present application from the perspective of interaction between the client and the server. In the embodiments of the present application, the content generation method described in the above embodiments can be implemented by a model. To better implement the content generation method provided by the embodiments of the present application, the embodiments of the present application further provide a content generation model for implementing the above content generation method based on the above method embodiments.
[0301] Next, the structure of the content generation model is introduced.
[0302] In a first possible implementation, as shown in FIG. 1, the content generation model includes a generation network 101, a compression unit 103, a decompression unit 104, and a decoding network 102. Figure 10
[0303] The generation network 101 includes a content encoding network layer and a plurality of encoding network layers of different modalities.
[0304] As Figure 10 shown, the encoding network layer in the generation network 101 includes a text encoding unit, an image encoding unit, and a text+image encoding unit. The text encoding unit, the image encoding unit, and the text+image encoding unit are respectively used to encode the input data to obtain the text data encoding, the image data encoding, and the fusion encoding corresponding to the text encoding unit, the image encoding unit, and the text+image encoding unit respectively.
[0305] The content encoding network layer is used to generate features for the input encoded data to obtain content features. In an example, the content encoding network layer can be a diffusion generative model.
[0306] The decoding network 102 includes a plurality of decoding units of different modalities. As Figure 11 shown, the decoding network 102 includes an image decoding unit, a video decoding unit, and a text+image decoding unit. Among them, the plurality of decoding units of different modalities are used to generate reconstructed content based on the input content features.
[0307] In the embodiments of the present application, the generation network 101 and the compression unit 103 are deployed in the server. The decoding network 102 and the decompression unit 104 are deployed in the client.
[0308] In an example, the compression unit 103 can include a self-encoding unit, an arithmetic encoding unit. Correspondingly, the decompression unit 104 includes a self-decoding unit, an arithmetic decoding unit.
[0309] In another example, the compression unit 103 can also include an encoding unit in Figure 3 , an arithmetic encoding (AutoEncoder, AE) unit, and an entropy model. Correspondingly, the decompression unit 104 includes an arithmetic decoding (Auto Decoder, AD) unit, an entropy model, and a decoding unit in Figure 3 .
[0310] Among them, the generation network 101 is used to generate features for the input content description information to obtain content features corresponding to the content description information, and input the content features corresponding to the content description information into the compression unit 103.
[0311] The compression unit 103 is used to compress the content features corresponding to the content description information to obtain a code stream. The code stream output by the compression unit 103 is input into the decompression unit.
[0312] The decompression unit 104 is used to decompress the code stream to obtain the content features corresponding to the content description information, and input the content features corresponding to the content description information into the decoder.
[0313] Decoding network 102 is used to generate content corresponding to the content description information based on the content features corresponding to the content description information, and output the reconstructed content.
[0314] In this embodiment of the application, to improve the data quality of content generation, it is necessary to train the generation network 101, compression unit 103, decompression unit 104, and decoding network 102. During the training phase, an appropriate training method can be selected according to the actual application scenario.
[0315] For example, in a first possible implementation, the generator network 101, the compression unit 103, the decompression unit 104, and the decoder network 102 can be trained alternately.
[0316] In the second possible implementation, the generator network 101 and the decoder network 102 can be trained as a single model. The compression unit 103 and the decompression unit 104 can be trained as a separate model. After the generator network 101, decoder network 102, compression unit 103, and decompression unit 104 are trained, they can be combined to form a [model name missing]. Figure 11 The content generation model shown.
[0317] In the third possible implementation, the pre-trained content generation model can be deployed on the corresponding device. Sample content data from the business scenario is then used to adjust the model parameters of the pre-trained content generation model, resulting in the final content generation model.
[0318] based on Figure 10 The content generation model provided in this application, compared to related technologies, adds a decompression unit 104 and a compression unit 103 between the generation network 101 and the decoding network 102, and distributes the decompression unit 104 and the compression unit 103 on the client and server sides. By directly compressing the content features output by the generation network 101 through the compression unit 103, the compression ratio of the bitstream can be improved. Furthermore, the data volume of the features is smaller than the data volume of the content, which improves compression efficiency. In addition, the bitstream data volume in this application embodiment is small, resulting in low bandwidth consumption, improved transmission efficiency, and reduced transmission losses.
[0319] In the second possible implementation, such as Figure 11 As shown, the content generation model includes a generation network 101 and a decoding network 102.
[0320] like Figure 11 As shown, Figure 11is a second structural schematic diagram of the content generation model provided in the embodiments of the present application. The decoding network 102 includes a compression unit 103, a decompression unit 104, and a content decoding unit.
[0321] In the embodiments of the present application, the compression unit 103 in the decoding network 102 is deployed on the server side. The decompression unit 104 and the content decoding unit are deployed on the client side.
[0322] With respect to Figure 11 the content generation model provided in the embodiments of the present application, Figure 11 the content generation model provided in the embodiments of the present application integrates the compression unit 103 and the decompression unit 104 in the decoding network 102, and provides a decoding network 102 with feature compression and decompression functions.
[0323] With respect to the decoder in the generation model in the related art, Figure 11 the decoding network 102 provided in the embodiments of the present application can convert the features output by the generation network 101 into a code stream with high compression rate and high restorability, thereby providing an efficient end-to-cloud transmission function for the entire generation task.
[0324] In a third possible implementation, to improve the feature compression effect and reduce the redundancy between features, in the content generation model provided in the embodiments of the present application, a hyper-prior network 105 is added in the decoding network 102. The prior features extracted by the hyper-prior network 105 improve the feature compression efficiency. Figure 11
[0325] As shown in Figure 12 , Figure 12 is a third structural schematic diagram of the content generation model provided in the embodiments of the present application. As shown in Figure 12 (a) of FIG. 1, the hyper-prior network 105 is also provided in the decoding network 102. The features output by the generation network 101 are respectively input into the hyper-prior network 105 and the compression unit 103. The output of the hyper-prior network 105 is respectively input into the compression unit 103 and the decompression unit 104.
[0326] As shown in Figure 12 (b) of FIG. 1, the hyper-prior network 105 includes a first sub-network 1051, a quantization sub-network 1052, and a second sub-network 1053.
[0327] The first sub-network 1051 generates latent representation features based on the input content features, and inputs the latent representation features into the quantization sub-network 1052. The quantization sub-network 1052 discretizes the input latent representation features, and outputs discrete latent representation features. The discrete latent representation features are input into the second sub-network 1053 for decoding to obtain prior features.
[0328] The first sub-network 1051 includes multiple convolutional layers connected in sequence. An activation function is placed between two adjacent convolutional layers. The second sub-network 1053 includes multiple deconvolutional layers and convolutional layers connected in sequence. Activation functions are placed between adjacent deconvolutional layers and between convolutional layers and deconvolutional layers. The activation function may be ReLU.
[0329] like Figure 13 As shown in Figure (a), the first sub-network 1051 includes three convolutional layers connected in sequence: convolutional layer 1, convolutional layer 2, and convolutional layer 3. Convolutional layer 1 has a 3x3 kernel size and a stride of 1. Convolutional layers 2 and 3 have the same size, both having a 5x5 kernel and a stride of 2. Activation functions are set between convolutional layers 1 and 2, and between convolutional layers 2 and 3.
[0330] like Figure 13 As shown in Figure (b), the second sub-network 1053 includes deconvolution layer 1, deconvolution layer 2, and convolutional layer 3. Deconvolution layer 1 and deconvolution layer 2 are both deconvolution layers with a kernel size of 5*5 and a stride of 2. Convolutional layer 3 is a convolutional layer with a kernel size of 3*3 and a stride of 1. An activation function is set between deconvolution layer 1 and deconvolution layer 2. A cloud activation function is set between deconvolution layer 2 and convolutional layer 3.
[0331] In one example, the super-prior network 105 is deployed on the server. After generating the bitrate, the server sends the prior features and bitrate to the client. For example... Figure 12 As shown in Figure (a).
[0332] In another instance, the first subnetwork 1051, the quantization subnetwork 1052, and the second subnetwork 1053 are deployed on the server side. The second subnetwork 1053 is deployed on the client side. Figure 12 (Not shown in the diagram). After generating the bitstream, the server sends the discrete latent representation features and the bitstream to the client. The client decodes the discrete latent representation features through the second sub-network 1053 to generate prior features.
[0333] Compared to the decoder in the generative model in related technologies, Figure 12 In the provided content generation model, the decoding network 102 can compress the content features output by the generation network 101. And in the feature compression process, Figure 12 The provided content generation model can reduce redundancy between features and improve feature compression by utilizing prior features extracted by the super-prior network 105.
[0334] When the amount of data corresponding to the content features generated on the server side is large, compressing the content features corresponding to the content description information will result in low compression efficiency. Furthermore, deploying the entire decoding network 102 on the client side may lead to content generation failures or low content generation efficiency if the client's computing power is insufficient. Deploying the entire decoding network 102 on the server side will result in low compression efficiency and significant transmission loss. Therefore, in the fourth possible implementation, in the above... Figure 12 Based on the provided content generation model, the decoding network 102 is split into a first content decoding unit and a second content decoding unit.
[0335] like Figure 14 As shown in Figure (a), the content features output by the generator network 101 are input to the first content decoding unit. The dimensionality-reduced features of the content features output by the first content decoding unit are input to the compression unit 103. The bitstream output by the compression unit 103 is input to the decompression unit 104. The dimensionality-reduced features of the content features output by the decompression unit 104 are input to the second content decoding unit. The second content decoding unit outputs the reconstructed content.
[0336] like Figure 14 As shown in Figure (b), the content features output by the generator network 101 are input to the first content decoding unit. The dimensionality-reduced features output by the first content decoding unit are input to the super-prior network 105 and the compression unit 103, respectively. The prior features output by the super-prior network 105 are input to the compression unit 103 and the decompression unit 104, respectively. The compression unit 103 compresses the dimensionality-reduced features based on the prior features to obtain the bitstream. The bitstream is input to the decompression unit 104. The decompression unit 104 decompresses the bitstream based on the input prior features and outputs the dimensionality-reduced features of the content features.
[0337] In one example, the first content decoding unit is deployed on the server side, and the second generation decoding unit is deployed on the client side.
[0338] based on Figure 14 The provided content generation model splits the content decoding unit and deploys it on both the server and client sides, reducing the computational power consumption on the client side. Furthermore, it avoids the problem of low compression efficiency when directly compressing the content features corresponding to the content description information by compressing the low-dimensional content features, which are the data volume of the content description information.
[0339] In the fifth possible implementation, to improve the data quality of the reconstructed content, a diffusion model is added to the decoding network 102 based on any of the above content generation models, and the content features are processed through the diffusion model.
[0340] For example, as shown in Figure 15 , Figure 15 is a fifth structure diagram of a content generation model provided by an embodiment of the present application. The diffusion model is arranged after the decompression unit 104. The diffusion model is used to perform noise reduction processing on the features output by the decompression unit 104.
[0341] The diffusion model is deployed in the client.
[0342] Based on Figure 15 the embodiments provided by the present application, the feature quality of the generated content feature is improved by adding the diffusion model to perform noise reduction processing on the decompressed features. The data quality of the generated reconstructed content is further ensured.
[0343] The above mainly describes the hardware structure and / or software modules corresponding to each function of the content generation system. Those skilled in the art can easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed in the present application, the present application can be realized in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in hardware or computer software driven hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0344] Embodiments of the present application can group the functional modules of the content generation system according to the above method examples. For example, each functional module can be grouped according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be realized in the form of hardware or software functional module. It should be noted that the grouping and naming of the modules in the embodiments of the present application are illustrative, and are only a kind of logical grouping. Actual implementation can have another grouping manner.
[0345] For example, the content generation system can be named as a content generation device. As shown in Figure 16 , Figure 16 is a structure diagram of a content generation device provided by an embodiment of the present application. The content generation device 16 shown includes a communication module 161, a decoding module 162 and a generation module 163.
[0346] The communication module 161 is used to send a content generation request and / or accept a code stream. For example, the communication module 161 performs S510 in the above Figure 5A .
[0347] The decoding module 162 is used to decompress the code stream to obtain the content feature corresponding to the content description information. For example, the decoding module 162 performs the aboveFigure 5A In S540.
[0348] The generating module 163 is configured to generate the reconstructed content corresponding to the content description information according to the content features. For example, the generating module 163 performs S550 in the above method. Figure 5A
[0349] The communication module 161, the decoding module 162, and the generating module 163 can be implemented by software or by hardware. For example, the implementation of the communication module 161 is described below, and the implementation of the decoding module 162 and the generating module 163 can be similar to that of the communication module 161.
[0350] As an example of a software functional unit, the communication module 161 can be code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the communication module 161 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers running the code can be distributed in the same region, or in different regions. Further, the multiple hosts / virtual machines / containers running the code can be distributed in the same availability zone (AZ), or in different AZs. Each AZ includes one data center or multiple data centers with similar geographical locations. Generally, one region can include multiple AZs.
[0351] Similarly, the multiple hosts / virtual machines / containers running the code can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Generally, one VPC is set in one region, and communication between two VPCs in the same region or between VPCs in different regions needs to be set through a communication gateway in each VPC to realize the interconnection between the VPCs.
[0352] As an example of a hardware functional unit, the communication module 161 can include at least one computing device, such as a server or the like. Alternatively, the communication module 161 can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), and the like. The PLD can be implemented by a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0353] The plurality of computing devices included in the communication module 161 can be distributed in the same region or in different regions. The plurality of computing devices included in the communication module 161 can be distributed in the same AZ or in different AZs. Similarly, the plurality of computing devices included in the feature compression module 173 can be distributed in the same VPC or in multiple VPCs. The plurality of computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0354] It should be noted that in other embodiments, the communication module 161 can be used to perform any step in the content generation method. The decoding module 162 can be used to perform any step in the content generation method. The generation module 163 can be used to perform any step in the content generation method. The steps responsible for the communication module 161, the decoding module 162, and the generation module 163 can be specified as needed. The entire function of the content generation apparatus 16 is achieved by the communication module 161, the decoding module 162, and the generation module 163 respectively implementing different steps in the content generation method.
[0355] The embodiments of the present application also provide a content generation apparatus 17. As shown in the figure, the content generation apparatus 17 includes a communication module 171, a feature generation module 172, and a compression module 173. Figure 17
[0356] The feature generation module 172 is configured to generate content features corresponding to the content description information according to the content description information. For example, the feature generation module 172 performs S520 in the above method. Figure 5A
[0357] The compression module 173 is configured to compress the content features to obtain a bitstream. For example, the compression module 173 performs S530 in the above method.Figure 5A The S530 in the middle.
[0358] Communication module 171 also transmits a code stream. For example, communication module 171 performs the above... Figure 5A The operation of sending the bitstream in the process.
[0359] The communication module 171, feature generation module 172, and compression module 173 can be implemented in software or in hardware. The implementation methods of the communication module 171, feature generation module 172, and compression module 173 can refer to the implementation methods of the communication module 161, decoding module 162, and generation module 163 in the above-mentioned content generation device 16; these will not be elaborated upon in this embodiment.
[0360] This application also provides a content generation device 18 for the above-described content generation method. The content generation device 18 may be a computer device, a server, or a cloud server, or it may be a chip (system) or other component or part that can be disposed in a computer device, server, or cloud server. This application does not limit the scope of the application.
[0361] like Figure 18 As shown, the content generation apparatus 18 may include a processor. Optionally, the content generation apparatus 18 may also include a memory and a communication interface. The processor is coupled to the memory and the communication interface, for example, they may be connected via a communication bus.
[0362] like Figure 18 As shown, the content generation apparatus 18 includes a bus 182, a processor 184, a memory 186, and a communication interface 188. The processor 184, the memory 186, and the communication interface 188 communicate with each other via the bus 182. The content generation apparatus 18 can be a server or a terminal device. It should be understood that this application does not limit the number of processors 184 and memory 186 in the content generation apparatus 18.
[0363] The bus 182 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 18 The bus 182 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 182 may include a path for transmitting information between various components of the content generation apparatus 18 (e.g., memory 186, processor 184, communication interface 188).
[0364] The processor 184 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), among other processors.
[0365] In the present application, the processor 184 performs the above-mentioned Figure 5A method shown in the figure. For example, based on the content description information carried by the content generation request, the content feature corresponding to the content description information is generated. The compressed feature obtains the code stream. The code stream is decompressed to obtain the content feature corresponding to the content description information. The content corresponding to the content description information is generated according to the content feature corresponding to the content description information.
[0366] The memory 186 can include a volatile memory, such as a random access memory (RAM). The processor 184 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0367] The memory 186 stores executable program codes, and the processor 184 executes the executable program codes to respectively implement the functions of the above-mentioned communication module 161, decoding module 162 and generation module 163, so as to implement the content generation method. Alternatively, the processor 184 executes the executable program codes to respectively implement the functions of the above-mentioned communication module 171, feature generation module 172 and compression module 173, so as to implement the content generation method, that is, the memory 186 stores instructions for executing the content generation method.
[0368] In the embodiments of the present application, the memory 186 stores code streams, content description information, content features corresponding to the content description information, features and reconstructed content, etc.
[0369] The communication interface 188 uses a transceiver module such as, but not limited to, a network interface card and a transceiver, to realize the communication between the content generation device 18 and other devices or communication networks.
[0370] The content generation method disclosed in the above method embodiments can be applied in the processor 184 or implemented by the processor 184. The processor 184 can be an integrated circuit chip with signal processor capability.
[0371] In implementation, each step of the above method can be completed by integrated logic circuit of hardware in the processor 184 or instruction in the form of software. The processor 184 described above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete electronic tube or transistor logic device, a discrete hardware component. Each method, step and logic block disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as a hardware coding processor for execution, or a combination of hardware and software modules in the coding processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 186, and the processor 184 reads the information in the memory 186, and combines the hardware to complete the steps of the above method.
[0372] In a possible implementation, the processor 184 can also be used to execute the content generation method, and the specific implementation can refer to the embodiments provided by the above content generation method. The embodiments of the present application will not be repeated here.
[0373] In the embodiments of the present application, the chip system can be composed of a chip, or can include a chip and other discrete devices.
[0374] The embodiments of the present application also provide a computing device 19 for executing the above content generation method.
[0375] In an example, the computing device 19 can include a content generation apparatus 16 as shown in Figure 16 , which includes a communication module 161, a decoding module 162 and a generation module 163.
[0376] In another example, the computing device 19 can include a content generation apparatus 17 as shown in Figure 17 , which includes a communication module 171, a feature generation module 172 and a compression module 173.
[0377] In another example, the computing device 19 can have the same structure as the content generation apparatus 18 as shown in Figure 18 , which includes a bus 182, a processor 184, a memory 186 and a communication interface 188.
[0378] The embodiments of the present application also provide a computing device cluster 20 for performing the content generation method.
[0379] In an example, the computing device cluster 20 can include a content generation apparatus 16 as shown in Figure 16 The content generation apparatus 16 includes a communication module 161, a decoding module 162 and a generation module 163.
[0380] In an example, the computing device cluster 20 can include a content generation apparatus 17 as shown in Figure 17 The content generation apparatus 17 includes a communication module 171, a feature generation module 172 and a compression module 173.
[0381] In yet another example, the computing device cluster 20 can include at least one content generation apparatus 18 as shown in Figure 18 The content generation apparatus 18 includes a bus 182, a processor 184, a memory 186 and a communication interface 188. The processor 184, the memory 186 and the communication interface 188 communicate with each other through the bus 182.
[0382] In yet another example, the computing device cluster 20 includes at least one computing device 19 as shown in Figure 19 The computing device 19 includes a bus 182, a processor 184, a memory 186 and a communication interface 188. The processor 184, the memory 186 and the communication interface 188 communicate with each other through the bus 182. The computing device 19 can be a server or a terminal device.
[0383] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be a software or program product containing instructions, which can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device is caused to perform the content generation method.
[0384] For example, when the computer program product is run on at least one computing device, the at least one computing device is caused to perform the content generation method as shown in Figure 5A
[0385] The embodiments of the present application further provide a computer readable storage medium. All or part of the processes in the above method embodiments can be instructed by a computer program to relevant hardware to complete, the program can be stored in the above computer readable storage medium, and the program can include the processes of the above method embodiments when executed. The computer readable storage medium can be the terminal of any of the preceding embodiments, such as an internal storage unit including a data transmission end and / or a data receiving end, for example, a hard disk or a memory of the terminal. The above computer readable storage medium can also be an external storage device of the terminal, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal. Further, the above computer readable storage medium can include both the internal storage unit and the external storage device of the terminal. The above computer readable storage medium is used to store the above computer program and other programs and data required by the terminal. The above computer readable storage medium can also be used to temporarily store data that has been output or will be output.
[0386] It should be noted that the terms "first" and "second" and the like in the specification, claims and drawings of the present application are used to distinguish different objects, and are not used to describe a particular order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.
[0387] It should be understood that in the present application, "at least one" means one or more, "multiple" means two or more, "at least two" means two or three and more, and "and / or" is used to describe the association relationship of the associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are three cases of only A, only B and A and B at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b and c can be single or multiple.
[0388] It should be understood that, in the embodiments of the present application, "B corresponding to A" means that B is associated with A. For example, B can be determined according to A. It should also be understood that determining B according to A does not mean that B is determined only according to A, but B can also be determined according to A and / or other information. In addition, "connection" appearing in the embodiments of the present application refers to various connection modes such as direct connection or indirect connection to achieve communication between devices, and the embodiments of the present application do not make any limitation thereto.
[0389] Through the description of the above embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the grouping of the above functional modules is exemplified, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is grouped into different functional modules to complete all or part of the functions described above.
[0390] In several embodiments provided in the present application, it should be understood that the disclosed content generation device and method can be implemented by other ways. For example, the content generation device embodiments described above are only schematic, for example, the grouping of modules or units is only a logical function grouping, and actual implementation can have another grouping manner, for example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0391] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit.
[0392] If the integrated unit is realized in the form of software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application essentially or say the part that makes contributions to the prior art or all or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium, including a plurality of instructions to make a device, such as a single-chip microcomputer, chip or processor execute all or part of the steps of the embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, ROM, RAM, magnetic disk or optical disk and various storage program codes.
[0393] The above merely provides the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of the changes or replacements within the technical range disclosed by the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A content generation method characterized by, The method comprises: receiving a code stream; decompressing the code stream to obtain a content feature corresponding to content description information, a data amount of the content feature being less than a data amount of content corresponding to the content description information; generating reconstructed content corresponding to the content description information according to the content feature.
2. The method of claim 1, wherein, The decompressing the code stream to obtain a content feature corresponding to content description information comprises: dequantizing the code stream to obtain the content feature corresponding to the content description information.
3. The method of claim 1, wherein, The decompressing the code stream to obtain a content feature corresponding to content description information comprises: entropy decoding the code stream to obtain a decoded feature; dequantizing the decoded feature to obtain the content feature corresponding to the content description information.
4. The method of claim 3, wherein, Before the receiving a code stream, the method further comprises: obtaining a prior feature, the prior feature comprising a quantization parameter and a distribution feature of the content feature; The decompressing the code stream to obtain a content feature corresponding to content description information comprises: entropy decoding the code stream according to the distribution feature to obtain a decoded feature; dequantizing the decoded feature according to the quantization parameter to obtain the content feature corresponding to the content description information.
5. The method according to any one of claims 1 to 4, characterized in that, The decompressing the code stream to obtain a content feature corresponding to content description information comprises: decompressing the code stream to obtain a reduced-dimension feature of the content feature corresponding to the content description information; The generating reconstructed content corresponding to the content description information according to the content feature comprises: generating the reconstructed content corresponding to the content description information according to the reduced-dimension feature.
6. The method according to any one of claims 1 to 5, characterized in that, The generating reconstructed content corresponding to the content description information according to the content feature comprises: processing the content feature to obtain a processed feature; generating the reconstructed content corresponding to the content description information according to the processed feature.
7. The method of claim 6, wherein, The processing the content feature to obtain a processed feature comprises: inputting the content feature into a diffusion model to obtain the processed feature.
8. The method according to any one of claims 1 to 7, characterized in that, Before the receiving a code stream, the method further comprises: sending a content generation request to a server, the content generation request containing the content description information.
9. A content generation method characterized by, The method comprises: generating a content feature corresponding to content description information according to the content description information, a data amount of the content feature being less than a data amount of content corresponding to the content description information; compressing the content feature to obtain a code stream.
10. The method of claim 9, wherein, The compressing the content feature to obtain a code stream comprises: quantizing the content feature to obtain the code stream.
11. The method of claim 10, wherein, The quantizing the content feature to obtain the code stream comprises: when a network bandwidth between a server and a client is greater than or equal to a bandwidth threshold, quantizing the content feature to obtain the code stream.
12. The method of claim 10, wherein, The quantizing the content feature to obtain the code stream comprises: when a network bandwidth between a server and a client is less than a bandwidth threshold, quantizing the content feature to obtain a quantized feature; entropy encoding the quantized feature to obtain the code stream.
13. The method of claim 9, wherein, The compressing the content feature to obtain a code stream comprises: input the content feature into a hyper-prior network to obtain a prior feature of the content feature; the prior feature comprises a quantization parameter and a distribution feature of the content feature; quantize the content feature according to the quantization parameter to obtain a quantized feature; entropy encode the quantized feature according to the distribution feature to obtain the code stream.
14. The method according to any one of claims 9 to 13, characterized in that, The method comprises: dimension reduction processing the content feature to obtain a dimension-reduced feature of the content feature corresponding to the content description information; compressing the dimension-reduced feature to obtain the code stream.
15. The method according to any one of claims 9 to 14, characterized in that, Before the content feature corresponding to the content description information is generated according to the content description information, the method further comprises: receiving a content generation request sent by a client, the content generation request containing the content description information.
16. The method according to any one of claims 9 to 14, characterized in that, Before the content feature corresponding to the content description information is generated according to the content description information, the method further comprises: obtaining the stored content description information.
17. The method according to any one of claims 9 to 16, characterized in that, After the content feature is compressed to obtain the code stream, the method further comprises: sending the code stream to the client.
18. The method according to any one of claims 9 to 17, characterized in that, The content feature corresponding to the content description information is generated according to the content description information and a generative model based on a diffusion model. The content description information comprises at least one of text, image, audio or video.
19. The method according to any one of claims 9 to 18, characterized in that, The method comprises:
20. A content generation method characterized by comprising: The server generates a content feature corresponding to the content description information according to the content description information, and compresses the content feature to obtain a code stream; the data amount of the content feature is less than that of the content corresponding to the content description information; The server sends the code stream to the client; The client decompresses the code stream to obtain the content feature corresponding to the content description information; The client generates reconstructed content corresponding to the content description information according to the content feature. Before the server generates a content feature corresponding to the content description information according to the content description information and compresses the content feature to obtain a code stream, the method further comprises:
21. The method of claim 20, wherein, The client sends a content generation request to the server, the content generation request containing the content description information. The content generation system comprises a server and a client; 22. A content generation system characterized by, The server is configured to generate a content feature corresponding to the content description information according to the content description information, and compress the content feature to obtain a code stream; the data amount of the content feature is less than that of the content corresponding to the content description information; The server is further configured to send the code stream to the client; The client is configured to decompress the code stream to obtain the content feature corresponding to the content description information, and generate reconstructed content corresponding to the content description information according to the content feature. The server is deployed with a hyper-prior network; 23. The system of claim 22, wherein, The server is further configured to input the content feature into a hyper-prior network to obtain a prior feature of the content feature; the prior feature comprises a quantization parameter and a distribution feature of the content feature; quantize the content feature according to the quantization parameter, to obtain a quantized feature; perform entropy coding on the quantized feature according to the distribution feature, to obtain the code stream.
24. A content generating apparatus characterized by comprising: The content generation apparatus is arranged at a client side, and comprises: a communication module configured to receive the code stream; a decompression module configured to decompress the code stream to obtain a content feature corresponding to the content description information; a generation module configured to generate a reconstructed content corresponding to the content description information according to the content feature.
25. A content generating apparatus characterized by comprising: The content generation apparatus is arranged at a server side, and comprises: a feature generation module configured to generate a content feature corresponding to the content description information according to the content description information, wherein a data amount of the content feature is less than a data amount of a content corresponding to the content description information; a compression module configured to compress the content feature to obtain the code stream.
26. A content generating apparatus characterized by comprising: The apparatus comprises a processor coupled to the memory; The memory is configured to store a computer program; The processor is configured to execute the computer program stored in the memory, so that the apparatus performs the method according to any one of claims 1 to 21.
27. A cluster of computing devices, characterized in that, The apparatus comprises at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the cluster of computing devices performs the method according to any one of claims 1 to 21.
28. A computer program product comprising instructions, wherein: The instructions, when executed by the cluster of computing devices, cause the cluster of computing devices to perform the method according to any one of claims 1 to 21.
29. A computer-readable storage medium, characterized in that, The computer program instructions, when executed by the cluster of computing devices, cause the cluster of computing devices to perform the method according to any one of claims 1 to 21.