Video encoding method and apparatus, storage medium, and electronic device
Patent Information
- Application Number
- CN202610487003.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-14
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-04-14
AI Technical Summary
[0006]本申请实施例提供了一种视频编码方法及装置、存储介质、电子设备,以至少解决相关技术无法对单个独立码流进行编码得到多种分辨率视频,导致视频传输资源耗费大、编码灵活性差的技术问题
[0017]在本申请实施例中,解码端依据第一视频的目标分辨率对第一视频的第一压缩特征表示的低分辨率空间维度上执行逐维插值操作,精确恢复其与第一视频完全一致的高度和宽度,以在极低计算开销下完成特征分辨率对齐,为后续混合模型提供时空精准对齐的条件输入,显著提升视频生成质量的稳定性和收敛效率。接着,利用预先训练好的混合模型对第一恢复特征表示进行分析,得到第一视频对应的多个分辨率视频,其中,所述混合模型中包括级联的多个视频生成模型,且第一个所述视频生成模型的输入为所述第一恢复特征表示,第一个所述视频生成模型之后的每个所述视频生成模型的输入均为前一个所述视频生成模型输出的分辨率视频和所述第一恢复特征表示。整个方案实现从压缩语义表示到不同分辨率视频重建的高效、鲁棒、无失真映射,进而解决了相关技术无法对单个独立码流进行编码得到多种分辨率视频,导致视频传输资源耗费大、编码灵活性差的技术问题。
Smart Images

Figure CN122027798B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video coding technology, and more specifically, to a video coding method and apparatus, storage medium, and electronic device. Background Technology
[0002] With the popularization of video applications, user terminal devices (such as mobile phones, tablets, smart TVs, etc.) and network environments (4G, 5G, Wi-Fi, low-bandwidth areas, etc.) are highly heterogeneous, posing a severe challenge to the adaptability of video encoding technologies.
[0003] While traditional Scalable Video Coding (SVC) methods can achieve resolution adaptation through multi-layer bitstreams, they still rely on a manually designed framework based on transforms (such as discrete cosine transform) and hierarchical quantization. This means that the encoder must generate and transmit an independent bitstream for each resolution level, resulting in high bitrate redundancy, significantly increased storage overhead, and exponentially increased transmission bandwidth requirements. This makes it difficult to meet the high-efficiency transmission needs of mobile edge computing and low-bandwidth scenarios.
[0004] Meanwhile, generative video coding methods that have emerged in recent years, such as end-to-end coding frameworks based on generative adversarial networks or variational autoencoders, while significantly outperforming traditional scalable video coding methods in compression efficiency, heavily rely on a preset fixed output resolution during the generation process—the target size is locked during training and cannot be dynamically adjusted during inference. This results in the same bitstream only being able to output video of a single resolution, failing to adapt to the diverse display capabilities of user terminal devices. More importantly, these methods lack robustness to transmission impairments. Once the bitstream experiences packet loss or bit errors in the channel, the reconstructed video is highly susceptible to severe distortions such as blockiness, texture blurring, or spatiotemporal inconsistencies.
[0005] There is currently no effective solution to the above problems. Summary of the Invention
[0006] This application provides a video encoding method, apparatus, storage medium, and electronic device to at least solve the technical problem that related technologies cannot encode a single independent bitstream to obtain videos of multiple resolutions, resulting in high video transmission resource consumption and poor encoding flexibility.
[0007] According to one aspect of the embodiments of this application, a video encoding method is provided, comprising: acquiring a routing data packet sent by an encoding end, wherein the routing data packet includes at least: a target resolution of a first video and a first compressed feature representation corresponding to the first video; upsampling the first compressed feature representation according to the target resolution to obtain a first restored feature representation; and analyzing the first restored feature representation using a pre-trained hybrid model to obtain multiple resolution videos corresponding to the first video, wherein the hybrid model includes multiple cascaded video generation models, and the input of the first video generation model is the first restored feature representation, and the input of each video generation model after the first video generation model is the resolution video output by the previous video generation model and the first restored feature representation.
[0008] According to another aspect of the embodiments of this application, a video encoding apparatus is also provided, comprising: an acquisition module, configured to acquire a routing data packet sent by an encoding end, wherein the routing data packet includes at least: a target resolution of a first video and a first compressed feature representation corresponding to the first video; an upsampling module, configured to upsample the first compressed feature representation according to the target resolution to obtain a first restored feature representation; and an encoding module, configured to analyze the first restored feature representation using a pre-trained hybrid model to obtain multiple resolution videos corresponding to the first video, wherein the hybrid model includes multiple cascaded video generation models, and the input of the first video generation model is the first restored feature representation, and the input of each video generation model after the first video generation model is the resolution video output by the previous video generation model and the first restored feature representation.
[0009] In an exemplary embodiment, the above-described apparatus is further configured to train a hybrid model by means of: constructing an initial model composed of multiple cascaded video generation models; acquiring multiple sets of training sample data, wherein each set of training sample data includes: a second video and a second compressed feature representation corresponding to the second video; and iteratively training the initial model using the multiple sets of training sample data to obtain a hybrid model.
[0010] In an exemplary embodiment, the above-described apparatus is further configured to acquire multiple sets of training sample data by means of the following method: acquiring multiple second videos, wherein each second video has a different resolution; for each second video, encoding and compressing the second video to obtain a corresponding low-dimensional feature representation, performing channel expansion on the low-dimensional feature representation to obtain a corresponding enhanced feature representation, performing multi-level downsampling on the enhanced feature representation to obtain a corresponding compressed enhanced feature representation, and performing masking processing on the compressed enhanced feature representation to obtain a second compressed feature representation of the second video; and the multiple sets of training sample data are composed of the multiple second videos and the second compressed feature representation of each second video.
[0011] In an exemplary embodiment, the above-described apparatus is further configured to perform multi-level downsampling on the enhanced feature representation to obtain a corresponding compressed enhanced feature representation, including: compressing the time dimension, height dimension, and width dimension of the enhanced feature representation using a three-dimensional convolutional layer with a stride of a first value to obtain a low-degree compressed enhanced feature representation; fusing the result of downsampling the enhanced feature representation using a three-dimensional convolutional layer with a stride of the first value with the result of feature enhancement of the low-degree compressed enhanced feature representation to obtain a first intermediate compressed enhanced feature representation; and using a three-dimensional convolutional layer with a stride of a second value to compress the time dimension, height dimension, and width dimension of the first intermediate compressed enhanced feature representation. The first intermediate compressed enhanced feature representation is obtained by compressing the features separately. The result of downsampling the first intermediate compressed enhanced feature representation using a 3D convolutional layer with a stride of the second value is fused with the result of feature enhancement of the intermediate compressed enhanced feature representation to obtain the second intermediate compressed enhanced feature representation. The height and width dimensions of the second intermediate compressed enhanced feature representation are compressed using a 3D convolutional layer with a stride of the third value to obtain the highly compressed enhanced feature representation. The result of downsampling the enhanced feature representation using a 3D convolutional layer with a stride of the product of the first, second, and third values is fused with the highly compressed enhanced feature representation to obtain the compressed enhanced feature representation.
[0012] In an exemplary embodiment, the above-described apparatus is further configured to perform masking processing on the compression enhancement feature representation to obtain a second compressed feature representation of the second video by means of: performing masking processing on the compression enhancement feature representation according to preset masking information to obtain a corresponding mask tensor, wherein the masking information includes at least one of the following: number of occlusion blocks, spatial size of each occlusion block, temporal span, and channel span; multiplying the mask tensor element-wise with the compression enhancement feature representation to obtain a second compressed feature representation of the second video.
[0013] In an exemplary embodiment, the above-described apparatus is further configured to utilize a pre-trained hybrid model to analyze the first restored feature representation to obtain multiple resolution videos corresponding to the first video, including: for the first video generation model in the hybrid model, inputting the first restored feature representation into the first video generation model to obtain the first resolution video output by the first video generation model; for each video generation model in the hybrid model after the first video generation model, upsampling the second resolution video output by the previous video generation model of the current video generation model to obtain a third resolution video, wherein the resolution of the third resolution video is the same as the resolution of the resolution video output by the current video generation model; inputting the third resolution video and the first restored feature representation into the current video generation model to obtain the second resolution video output by the current video generation model, wherein the resolution of the second resolution video output by the current video generation model is higher than the resolution of the second resolution video output by the previous video generation model of the current video generation model.
[0014] In an exemplary embodiment, the above-described apparatus is further configured to upsample the first compressed feature representation according to the target resolution by the following method to obtain a first restored feature representation, including: sampling and interpolating the first compressed feature representation in the height and width dimensions to obtain a first restored feature representation with the same spatial size as the target resolution of the first video.
[0015] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, which stores a computer program, wherein the computer-readable storage medium, when executed by a processor, performs the steps in any of the above method embodiments.
[0016] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the steps of any of the above method embodiments through the computer program.
[0017] In this embodiment, the decoding end performs a dimensional interpolation operation on the low-resolution spatial dimension of the first compressed feature representation of the first video according to the target resolution of the first video, accurately restoring its height and width to be completely consistent with the first video. This achieves feature resolution alignment with extremely low computational overhead, providing the conditional input for precise spatiotemporal alignment of the subsequent hybrid model, significantly improving the stability and convergence efficiency of video generation quality. Next, the first restored feature representation is analyzed using a pre-trained hybrid model to obtain multiple resolution videos corresponding to the first video. The hybrid model includes multiple cascaded video generation models, with the input of the first video generation model being the first restored feature representation, and the input of each subsequent video generation model being the resolution video output by the previous video generation model and the first restored feature representation. The entire scheme achieves efficient, robust, and lossless mapping from compressed semantic representation to video reconstruction at different resolutions, thereby solving the technical problem that related technologies cannot encode a single independent bitstream to obtain multiple resolution videos, resulting in high video transmission resource consumption and poor encoding flexibility. Attached Figure Description
[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0019] Figure 1 This is a flowchart illustrating an optional video encoding method according to an embodiment of this application;
[0020] Figure 2 This is a schematic diagram of an optional training sample data acquisition process according to an embodiment of this application;
[0021] Figure 3 This is a schematic diagram of an optional compression enhancement feature representation determination process according to an embodiment of this application;
[0022] Figure 4 This is a schematic diagram of an optional video encoding process according to an embodiment of this application;
[0023] Figure 5 This is a structural block diagram of an optional video encoding apparatus according to an embodiment of this application;
[0024] Figure 6 This is a structural block diagram of an optional electronic device according to an embodiment of this application. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0026] It should be noted that the terms "first," "second," etc., used in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] According to an embodiment of this application, a video encoding method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0028] Figure 1 This is a flowchart illustrating a video encoding method according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0029] Step S102: Obtain the routing data packet sent by the encoding end, wherein the routing data packet includes at least: the target resolution of the first video and the first compression feature representation corresponding to the first video;
[0030] Step S104: Upsample the first compressed feature representation according to the target resolution to obtain the first restored feature representation;
[0031] Step S106: Analyze the first restored feature representation using a pre-trained hybrid model to obtain multiple resolution videos corresponding to the first video. The hybrid model includes multiple cascaded video generation models, and the input of the first video generation model is the first restored feature representation. The input of each video generation model after the first video generation model is the resolution video output by the previous video generation model and the first restored feature representation.
[0032] Optionally, in the embodiments of this application, the above-mentioned encoding end is a computing node that performs video compression and semantic feature extraction. It is usually deployed on the video source side (such as running on a server, edge node or front-end encoding device), and is responsible for receiving the first video to be transmitted. It maps the high-dimensional spatiotemporal data into a potential representation with a high compression ratio through a variational autoencoder and a downsampling module, and encapsulates and generates a single routed data packet for network transmission in combination with the target resolution requirements.
[0033] Optionally, in this embodiment, the aforementioned routing data packet is a structured data unit that encapsulates the target resolution identifier of the first video to be transmitted, the corresponding first compression feature representation, the fixed value of the time dimension, and other optional metadata fields at the encoding end, and transmits it to the decoding end through the network protocol stack. It does not contain the original pixel data (i.e., the first video), but instead carries unique semantic conditions and reconstruction instructions in a lightweight format. During transmission, this data packet can be routed by network devices according to priority or quality of service policies, but its content structure remains consistent across all decoding ends, ensuring that regardless of the receiving end's device, the video content at the corresponding target resolution can be restored according to preset mapping rules.
[0034] Optionally, in the embodiments of this application, the target resolution is the spatial resolution specification that the decoding end needs to output when reconstructing the video. It is dynamically selected by the encoding end based on network bandwidth, terminal display capabilities, or user preferences. Its expression form is an integer pair or a standardized resolution level, such as 854×480, 1280×720, 1920×1080, etc., or corresponding to video resolution labels such as 240p, 480p, 720p, 1080p, etc.
[0035] Optionally, in this embodiment, the first compressed feature representation is a potential representation obtained by highly compressing the first video (i.e., the original video) at the encoding end. This first compressed feature representation does not contain any pixel data, but only records all reconstruction clues such as the structure, motion, and appearance of the first video. For example, when the first video is a 5-second video of a person speaking indoors, its corresponding first compressed feature representation may contain key semantic clues such as the person's facial contours, head posture changes, lip movement trends, and background static structure, but omit pixel-level texture details.
[0036] Optionally, in this embodiment, the first restored feature representation is a semantic representation of the spatial size corresponding to the target resolution of the first video, which is restored by the decoding end after expanding the spatial dimension of the first compressed feature representation according to the target resolution.
[0037] Optionally, in this embodiment, the hybrid model described above is a deep learning model architecture composed of multiple cascaded video generation models. These video generation models are specifically designed for different generation stages or different resolution levels, achieving progressive generation from low resolution to high resolution. The video generation model is a deep learning-based conditional generation model that uses compressed feature representations as conditional inputs and reconstructs high-quality videos that are spatiotemporally coherent and semantically consistent from random noise or low-quality representations through progressive denoising. In this embodiment, the video generation model can be a diffusion model based on a stream matching architecture, a generative adversarial network, or a time-aware autoregressive model, etc.
[0038] Optionally, in this embodiment, the aforementioned multiple resolution videos are a set of video sequences with different spatial sizes and visual sharpness, reconstructed step-by-step from the first restored feature representation by the decoding end. These include, but are not limited to, low-resolution (e.g., 240p, 360p), medium-resolution (e.g., 720p), and high-resolution (e.g., 1080p, 4K) scales. These resolution videos share the same temporal structure, motion trajectory, and semantic content, differing only in detail richness, texture sharpness, and spatial sampling density.
[0039] For example, step S102 above can be understood as follows: the encoding end first highly compresses the first video (e.g., compresses it to 1 / 8 of the original spatial resolution of the first video) to obtain a first compressed feature representation, wherein the data stream of the first compressed feature representation is much smaller than the sum of the multiple bitstreams required by the first video or traditional multi-layer scalable coding; the encoding end then performs quantization and entropy coding on the first compressed feature representation to convert it into a compact binary bitstream, and encapsulates the binary bitstream and the target resolution of the first video into a single routed data packet through a standard network transmission protocol (such as User Datagram Protocol or Transmission Control Protocol), and sends it to the decoding end via the channel. After receiving the single routed data packet, the decoding end parses the single routed data packet to obtain the binary bitstream and the target resolution, and performs inverse entropy decoding and dequantization operations on the binary bitstream to recover the first compressed feature representation consistent with that of the encoding end.
[0040] For example, step S104 above can be understood as follows: the decoding end expands the spatial dimension of the first compressed feature representation according to the target resolution, thereby obtaining the first restored feature representation that matches the first video. This vector reconstructs the complete spatiotemporal semantic structure suitable for model input while maintaining a high compression ratio semantic representation.
[0041] For example, step S106 above can be understood as follows: the decoding end drives multiple cascaded video generation models based on stream matching sequentially by using the first restored feature representation as a global conditional input. Specifically: the first video generation model uses only the first restored feature representation as a conditional input to initially generate a low-resolution video from Gaussian noise, completing the initialization of the overall motion structure and semantic framework; subsequently, the input of each subsequent video generation model is the resolution video output by the previous video generation model and the first restored feature representation as conditional inputs, where the former provides an initial frame sequence for local detail enhancement, and the latter serves as a global semantic constraint, aligning spatiotemporal semantics through a conditional attention mechanism to ensure that the generated content maintains structural consistency during resolution enhancement. For example, if the hybrid model contains three video generation models, then the first video generation model will use the first restored feature representation as the sole conditional input and generate a low-resolution video as the starting point for subsequent video reconstruction stages. Subsequently, each video generation model takes the higher-resolution video output from the previous stage video generation model and the first restored feature representation as joint conditional inputs. By dynamically focusing on semantic priors, detailed information is gradually injected, and progressive reconstruction from low resolution to medium resolution and then to high resolution is completed in sequence. Each video generation model focuses on different generation tasks such as basic structure restoration, motion consistency enhancement, and high-frequency texture refinement, forming a collaboratively optimized multi-scale generation chain.
[0042] In this embodiment, the decoding end performs a dimensional interpolation operation on the low-resolution spatial dimension of the first compressed feature representation of the first video according to the target resolution of the first video, accurately restoring its height and width to be completely consistent with the first video. This achieves feature resolution alignment with extremely low computational overhead, providing the conditional input for precise spatiotemporal alignment of the subsequent hybrid model, significantly improving the stability and convergence efficiency of video generation quality. Next, the first restored feature representation is analyzed using a pre-trained hybrid model to obtain multiple resolution videos corresponding to the first video. The hybrid model includes multiple cascaded video generation models, with the input of the first video generation model being the first restored feature representation, and the input of each subsequent video generation model being the resolution video output by the previous video generation model and the first restored feature representation. The entire scheme achieves efficient, robust, and lossless mapping from compressed semantic representation to video reconstruction at different resolutions, thereby solving the technical problem that related technologies cannot encode a single independent bitstream to obtain multiple resolution videos, resulting in high video transmission resource consumption and poor encoding flexibility.
[0043] As an optional approach, the first compressed feature representation is upsampled according to the target resolution to obtain the first restored feature representation, including: sampling and interpolating the first compressed feature representation in the height and width dimensions to obtain the first restored feature representation with the same spatial size as the target resolution of the first video.
[0044] Optionally, in the embodiments of this application, the height dimension is the number of pixel rows in the vertical direction of each frame of the first video, that is, the vertical resolution of the image frame, which corresponds to the height axis H.
[0045] Optionally, in the embodiments of this application, the aforementioned width dimension is the number of pixel rows in the horizontal direction of each frame of the first video, that is, the horizontal resolution of the image frame, which corresponds to the width axis W.
[0046] For example, the above technical solution can be understood as follows: the decoding end independently performs upsampling operations on the first compressed feature representation in both the height and width dimensions to obtain a first restored feature representation with the same spatial size as the target resolution of the first video. The specific process is as follows: First, in the height and width dimensions, bilinear or bicubic interpolation algorithms are used to enlarge the two-dimensional spatial slices of each image frame, gradually restoring the spatial resolution to H×W. After the interpolation operation, a two-dimensional convolutional layer is used to perform feature correction and detail sharpening on the interpolation results to eliminate the blurring effect introduced by interpolation and enhance semantic consistency. Finally, the interpolation outputs of the two dimensions are fused to form the first restored feature representation. This vector, while maintaining the original compressed semantic structure, achieves accurate reconstruction of the resolution dimension, providing a conditional input aligned with the space of the first video for subsequent hybrid models.
[0047] In this embodiment, the decoding end performs independent and aligned dimensional interpolation operations on the low-resolution spatial dimension of the first compressed feature representation to accurately restore its height and width scales, which are completely consistent with the first video. This avoids non-physical interpolation artifacts introduced by dimensional coupling in 3D interpolation, ensuring the synchronous restoration of motion spatial structure and semantic consistency. Simultaneously, feature resolution alignment is achieved with extremely low computational overhead, providing spatiotemporally accurate alignment inputs for subsequent hybrid models. This significantly improves the stability and convergence efficiency of video generation quality, achieving an efficient, robust, and distortion-free mapping from compressed representation to high-resolution reconstruction.
[0048] As an optional approach, the training process of the above hybrid model includes:
[0049] Step S11: Construct an initial model consisting of multiple cascaded video generation models;
[0050] Step S12: Obtain multiple sets of training sample data, wherein each set of training sample data includes: the second video and the second compressed feature representation corresponding to the second video;
[0051] Step S13: Iteratively train the initial model using multiple sets of training sample data to obtain a hybrid model.
[0052] As an optional solution, it can be followed Figure 2 The process shown obtains multiple sets of training sample data, including:
[0053] Step S121: Obtain multiple second videos, each with a different resolution;
[0054] Step S122: For each second video, the second video is encoded and compressed to obtain the corresponding low-dimensional feature representation. The low-dimensional feature representation is channel-expanded to obtain the corresponding enhanced feature representation. The enhanced feature representation is downsampled at multiple levels to obtain the corresponding compressed enhanced feature representation. The compressed enhanced feature representation is then masked to obtain the second compressed feature representation of the second video.
[0055] Step S123: Multiple sets of training sample data are composed of multiple second videos and the second compressed feature representation of each second video.
[0056] Optionally, in the embodiments of this application, the second video mentioned above is a video used to construct training samples during the training process. It includes, but is not limited to, high-definition video clips from public video datasets or actual collections, covering different scenes, motion speeds, lighting conditions and resolution specifications, so as to enhance the model's generalization ability to diverse inputs.
[0057] Optionally, in this embodiment, the aforementioned low-dimensional feature representation is a compact coded feature representation with high semantic condensation and strong robustness generated by compressing the high-dimensional latent representation obtained from feature extraction of the second video. Specifically, after encoding the original second video using a scalable video coding method, its latent feature representation in the low-dimensional latent space can be obtained; then, conditional extraction is performed on the latent feature representation to map it into a low-dimensional compressed representation. Typically, the low-dimensional compressed representation can compress the number of channels and the spatial resolution (i.e., the height dimension and the width dimension).
[0058] Optionally, in this embodiment, the enhanced feature representation is a high-dimensional feature representation obtained by expanding the channel dimension of the low-dimensional feature representation of the second video. The purpose is to improve the expressive power of the features, enabling them to carry richer semantic and structural information. It should be noted that this representation is consistent with the latent representation of the second video (i.e., the latent representation obtained by mapping the second video from the pixel domain space to the low-dimensional latent space) in both spatial and temporal dimensions, with only the number of channels increased. It serves as the input basis for subsequent downsampling operations, enhancing the model's ability to model complex motions and details. In the three-dimensional convolutional feature representation, the number of channels is the number of feature maps along the channel dimension. Each channel is an independent two-dimensional (spatial) or three-dimensional (spatial) feature response map, used to encode certain semantic features of the video (such as edges, motion, texture, color changes, etc.). Therefore, the number of channels determines the model's expressive capacity and feature richness.
[0059] Optionally, in this embodiment, the downsampling described above reduces the spatial resolution of the feature representation in the three dimensions of time, height, and width through convolutional operations, thereby compressing the data size and extracting high-level semantics. For example, a 3×3×3 three-dimensional convolutional layer with a stride of 2 can reduce the feature size to half of its original size in the time, height, and width directions, respectively.
[0060] Optionally, in the embodiments of this application, the above-mentioned compression enhancement feature representation is a low-dimensional spatiotemporal representation obtained by gradually reducing the resolution of its time, height and width dimensions through multi-level downsampling operations, thereby significantly reducing the amount of data while preserving the core semantic structure.
[0061] Optionally, in the embodiments of this application, the above-mentioned second compression feature representation is a sparse and anti-interference final compression representation formed after applying random masking to the compression enhancement feature representation. Its dimension is consistent with the compression enhancement feature representation, but some spatiotemporal regions are set to zero to simulate transmission damage or information loss.
[0062] As an optional approach, channel expansion is performed on the low-dimensional feature representation to obtain the corresponding enhanced feature representation. This includes: using a 1×1×1 three-dimensional convolutional layer to perform channel expansion projection on the low-dimensional feature representation to obtain the enhanced feature representation, wherein the weight parameters of the 1×1×1 three-dimensional convolutional layer are... , bias is Therefore, the expression for the enhanced feature representation above can be written as: X represents the low-dimensional feature representation of the second video. Here, h represents the extended hidden dimension (e.g., h=256 or 512, and usually satisfies h greater than the number of channels), and T represents the time dimension (in frames).
[0063] As an optional solution, the decoding end can be achieved through, for example... Figure 3The architecture shown performs multi-level downsampling on the augmented feature representation to obtain the corresponding compressed augmented feature representation, including:
[0064] Step 1: Use a 3D convolutional layer with a stride of the first value to compress the time dimension, height dimension, and width dimension of the enhanced feature representation respectively, to obtain a low-compressed enhanced feature representation;
[0065] The second step is to fuse the result of downsampling the enhanced feature representation using a 3D convolutional layer with a stride of the first value with the result of feature enhancement of the low-compressed enhanced feature representation to obtain the first intermediate compressed enhanced feature representation. Then, the time dimension, height dimension, and width dimension of the first intermediate compressed enhanced feature representation are compressed using a 3D convolutional layer with a stride of the second value to obtain the medium-compressed enhanced feature representation.
[0066] Step 3: The result of downsampling the first intermediate compressed and enhanced feature representation using a 3D convolutional layer with a stride of the second value is fused with the result of feature enhancement of the moderately compressed and enhanced feature representation to obtain the second intermediate compressed and enhanced feature representation. The height and width dimensions of the second intermediate compressed and enhanced feature representation are compressed using a 3D convolutional layer with a stride of the third value to obtain the highly compressed and enhanced feature representation.
[0067] Step 4: The result of downsampling the enhanced feature representation using a 3D convolutional layer with a stride product of the first, second, and third values is fused with the highly compressed enhanced feature representation to obtain the compressed enhanced feature representation.
[0068] For example, the first step described above can be understood as follows: the decoding end first applies three parallel three-dimensional convolutional compression branches to enhance the feature representation. In the time, height, and width dimensions, each branch uses the same stride value k (e.g., k=2), thus compressing the time, height, and width dimensions of the enhanced feature representation to a factor of k, resulting in a low-compressed enhanced feature representation. This method achieves dimensional decoupled local structure extraction of enhanced feature representations, avoids semantic ambiguity caused by traditional global downsampling, improves the independent expressive power of features in each dimension, and lays a high-fidelity foundation for subsequent fine fusion.
[0069] For example, the second step described above can be understood as follows: the decoding end also utilizes a three-dimensional convolutional layer with a stride of the first value to enhance the feature representation. Downsampling is performed (e.g., if k=2, this corresponds to a 2x downsampling operation), and 1×1×1 convolution (channels unchanged) is used to achieve spatial dimensionality reduction, resulting in a low-degree compressed and enhanced feature representation. Simultaneously, an intermediate processing layer is used to enhance the feature representation for low-degree compression. Perform feature enhancement to obtain The intermediate processing layer can be a lightweight three-dimensional residual block used to nonlinearly enhance the local integration after downsampling. Therefore, the expression for the intermediate processing layer can be written as: ,in, This represents a 3×3×3 three-dimensional convolutional layer with a stride of 1 and padding of 1; ReLU represents the activation function, and X represents the feature representation input to the intermediate processing layer. Then, and Element-wise addition yields the first intermediate compressed and enhanced feature representation. This approach achieves complementary enhancement of coarse-grained structure and fine-grained semantics, solves the problem of high-frequency information loss in a single downsampling path, and significantly improves the ability of compressed representations to preserve local motion and texture details. Finally, three parallel 3D convolutional compression branches are applied to the first intermediate compressed and enhanced feature representation. In terms of time, height, and width dimensions, each branch uses the same step size, which is the second value. l (like l =2), thereby compressing the time, height, and width dimensions of the first intermediate compressed and enhanced feature representation to 1, respectively. l This yields a moderately compressed enhancement feature representation. .
[0070] For example, the third step above can be understood as: the decoding end uses a three-dimensional convolutional layer with a stride of the second value to compress and enhance the first intermediate feature representation. Perform downsampling (e.g.) l If the value is 2, it corresponds to a 2x downsampling operation, and a 1×1×1 convolution (channels unchanged) is used to achieve spatial dimensionality reduction, resulting in... Meanwhile, an intermediate processing layer is used to enhance the feature representation of moderate compression. Perform feature enhancement to obtain The expression for this intermediate processing layer has already been given above and will not be repeated here. Then, and Element-wise addition yields the second intermediate compressed enhanced feature representation. This approach achieves synergistic optimization of two-stage progressive compression and semantic enhancement. While gradually reducing resolution, it continuously injects early high-level semantics through residual connections, effectively suppressing information decay in deep networks and enhancing the generalization and robustness of features. Finally, two parallel 3D convolutional compression branches are applied to the second intermediate compression enhancement feature representation. The height and width dimensions of the two branches both use the same stride, a third value n (e.g., n=2), thus compressing the height and width dimensions of the second intermediate compressed enhanced feature representation to a factor of n, resulting in a height compressed enhanced feature representation. This achieves targeted compression of spatial redundancy while preserving the integrity of the temporal dimension, avoiding the problem of video temporal continuity being broken due to premature time downsampling, and ensuring that subsequent generation models can reconstruct dynamic content based on stable temporal cues.
[0071] For example, the fourth step above can be understood as: the decoding end uses a 1×1×1 three-dimensional convolutional layer to highly compress and enhance the feature representation. Spatial dimensionality reduction is performed, and the feature representation is enhanced by using a 3D convolutional layer with a stride product of the first, second, and third values. The downsampling results are then fused to obtain a compressed and enhanced feature representation, i.e. The compressed and enhanced feature representation output in this way has both "low resolution" and "high semantic density" characteristics, which not only meets the requirement of extremely high compression ratio of single bitstream, but also provides a robust and highly expressive conditional input for multi-resolution generation at the decoding end.
[0072] In the above technical solution, the decoding end achieves structured, reversible, and damage-resistant compression of the video's latent representation through a three-dimensional convolutional compression architecture that combines layering, multidimensionality, and residual fusion, without introducing multiple bitstreams or relying on external interpolation or super-resolution networks.
[0073] In an exemplary embodiment, the second video is a 16-frame color video with a resolution of 480×854 (i.e., height×width) and a frame rate of 24fps. Therefore, the decoding end can obtain the corresponding second compression feature representation according to the following steps:
[0074] Step 1: Receive the low-dimensional feature representation of the second video sent by the encoding end. The encoding end uses an encoder to compress the spatial dimension from 480×854 to 60×107, with 64 channels, while keeping the temporal dimension unchanged (still 16 frames) to preserve the temporal dynamics.
[0075] Step 2: Apply a 1×1×1 3D convolutional layer to the encoder output. Channel expansion is performed while maintaining the spatiotemporal dimension, increasing the number of channels from 64 to 128 to obtain enhanced feature representations. .
[0076] Step 3: Enhanced Feature Representation Three-level downsampling and residual fusion are performed to obtain compressed and enhanced feature representations. .
[0077] Specifically, the implementation process of the three-level downsampling and residual fusion in the third step above is as follows:
[0078] (1) Enhance the feature representation by using a 3D convolutional layer with a stride of 2 (the local receptive field size is 3 in the time, height and width dimensions, and a zero-value pixel layer is added at the boundary of each dimension so that the size of the feature representation output after the convolution operation remains unchanged in each dimension) Compression is performed to obtain a low-compression enhanced feature representation. .
[0079] (2) Representation of enhanced features Perform a 2x downsampling to obtain Simultaneously, an intermediate processing layer is used to enhance the feature representation of low-degree compression. Processing is performed to obtain ,Will and Feature fusion is performed to obtain the first intermediate compressed and enhanced feature representation. .
[0080] (3) The first intermediate feature representation is compressed and enhanced by using a three-dimensional convolutional layer with a stride of 2 (the local receptive field size is 3 in the time, height and width dimensions, and a layer of zero-value pixels is added at the boundary of each dimension so that the size of the feature representation output after the convolution operation remains unchanged in each dimension). Compression is performed to obtain a moderately compressed and enhanced feature representation. .
[0081] (4) Representation of enhanced features Perform 4x downsampling to obtain Simultaneously, an intermediate processing layer is used to enhance the feature representation of moderate compression. Processing is performed to obtain ,Will and Feature fusion is performed to obtain the second intermediate compressed and enhanced feature representation. .
[0082] (5) The second intermediate feature representation is compressed and enhanced by using a three-dimensional convolutional layer with a temporal stride of 1 and a spatial stride of 2 (the local receptive field size in the temporal, height and width dimensions is 3, and a zero-value pixel layer is added at the boundary of each dimension so that the size of the feature representation output after the convolution operation remains unchanged in each dimension). Compression is performed to obtain a highly compressed and enhanced feature representation. .
[0083] (6) Enhanced feature representation Perform 8x downsampling and project the channel to C=32 to obtain Simultaneously, a 1×1×1 three-dimensional convolutional layer is used to highly compress and enhance the feature representation. The number of channels C is projected to C=32. Additionally, because... The time dimension is 2, while The time dimension is 4, so upsampling can be used to... Adjusting the time dimension to 4, we get Finally, the channel number is projected onto the highly compressed enhancement feature with C=32 and... Feature fusion is performed to obtain a compressed and enhanced feature representation. .
[0084] As an optional approach, the compressed and enhanced feature representation is masked to obtain the second compressed feature representation of the second video, including:
[0085] Step 1: Perform masking processing on the compressed and enhanced feature representation based on the preset masking information to obtain the corresponding masking tensor. The masking information includes at least one of the following: number of masking blocks, spatial size of each masking block, time span, and channel span.
[0086] Step 2: Multiply the mask tensor element-wise with the compressed and enhanced feature representation to obtain the second compressed feature representation of the second video.
[0087] Optionally, in the embodiments of this application, the aforementioned masking block is a local region randomly defined in the five-dimensional space of the compressed and enhanced feature representation, used to simulate the local loss or damage that may occur during the transmission or storage of information. Essentially, it is a five-dimensional sub-feature representation formed by extending a continuous three-dimensional spatiotemporal block in the channel dimension. Its size and position are dynamically and randomly generated during the training process to enhance the model's adaptability to incomplete inputs.
[0088] For example, the first step above can be understood as: randomly sampling the number of occlusion blocks according to a preset parameter range ( , This indicates the minimum number of occlusion blocks. (Indicates the maximum number of occlusion blocks), and the span of each occlusion block in the time dimension ( , This represents the minimum time span of the occlusion block. (Indicates the maximum time span of the occlusion block), and its spatial dimensions in the height and width dimensions. , , This indicates the minimum height and minimum width of the occlusion block. Indicates the maximum height and maximum width of the occlusion block, and its coverage area in the channel dimension ( , Indicates the minimum channel span of the shading block. The maximum channel span of the masking block is represented by these parameters. Based on these parameters, multiple irregular, non-overlapping three-dimensional spatiotemporal blocks are generated in the five-dimensional space of the compressed and enhanced feature representation, and extended to the channel dimension to form a corresponding mask tensor. This mask tensor is a binary feature representation, where the masked region is assigned a value of zero, and the unmasked region is assigned a value of one. , , Let represent the starting channel span value and the ending channel span value of the i-th occlusion block, respectively. Let represent the start and end time span values of the i-th occlusion block, respectively. Let represent the starting height and ending height of the i-th occlusion block, respectively. These represent the starting width and ending width values of the i-th occlusion block, respectively.
[0089] For example, the second step above can be understood as: performing element-wise multiplication of the generated mask tensor with the compressed enhanced feature representation. This operation forces the feature values of the masked region to zero without changing the feature representation structure, while retaining the original information of the unmasked region, thereby outputting the second compressed feature representation after information sparsification.
[0090] In this embodiment, the decoder constructs a highly variable and realistic five-dimensional masking pattern by randomly sampling the number of masking blocks, the spatial size of the masking blocks, the temporal span, and the channel span. This exposes the compressed and enhanced feature representation to diverse information-deficient scenarios during the training phase, forcing the hybrid model to learn to recover semantically coherent video content even under incomplete and unstructured interference, significantly improving the intrinsic resistance of the potential representation to transmission impairments. Furthermore, lossless, differentiable, and spatially accurate feature masking is achieved through element-wise multiplication, ensuring complete compatibility between the perturbation process and network training, avoiding semantic distortion caused by artificial interpolation or noise injection. This mechanism does not rely on external error correction codes or retransmission mechanisms; it achieves intrinsic enhancement of the robustness of the encoded representation solely through adaptive regularization during the training process, thereby maintaining an extremely high compression ratio.
[0091] Furthermore, the initial model is iteratively trained using the multiple sets of training sample data obtained above. In each round of iterative training, a step-by-step training strategy from low to high resolution is adopted to ensure that each level of visual generation model only learns the "incremental enhancement" capability. This design avoids the problems of "chain error accumulation" and "end-to-end training instability" while maintaining the separation of responsibilities and the rationality of incremental enhancement of each level of the model.
[0092] Therefore, this application adopts a staged, conditionally dependent, and pre-stage frozen joint training strategy. During training, each video generation model uses the output of the previous video generation model as one of the input conditions and combines it with a unified compressed feature representation to learn how to progressively enhance the resolution. This sequential training of each video generation model can form a "condition-driven progressive generation pipeline," ensuring that the model output results are stable and interpretable.
[0093] Furthermore, during the model training process described above, the model obtained in each training round can be loaded into memory. For example, the raw data of the model can be loaded from non-volatile memory into volatile memory so that the processor can run the model. The raw data of the model refers to the unprocessed data, which typically includes the model's parameters and structural data. The structural data can be the computational relationships based on the parameters, such as the forward propagation computational relationships between intermediate layers and between neurons. Specifically, the structural data can include the model's structure-related code, such as the code used to perform related calculations between intermediate layers and between neurons.
[0094] In one implementation, a region for loading the model can be partitioned in memory, which may include a structure data storage area and a parameter storage area. The structure data storage area stores structure-related code, and the parameters referenced by it can be pointed to by pointers to the addresses of specific parameters in the parameter storage area. During model training, it may be necessary to frequently update the parameters, in which case the parameter values in the parameter storage area can be updated.
[0095] As an alternative approach, a pre-trained hybrid model is used to analyze the first recovered feature representation to obtain multiple resolution videos corresponding to the first video, including:
[0096] Step 1: For the first video generation model in the hybrid model, input the first restored feature representation into the first video generation model to obtain the first resolution video output by the first video generation model;
[0097] Step 2: For each video generation model after the first video generation model in the hybrid model, upsample the second-resolution video output by the previous video generation model to obtain a third-resolution video, wherein the resolution of the third-resolution video is the same as the resolution of the video output by the current video generation model; input the third-resolution video and the first restored feature representation into the current video generation model to obtain the second-resolution video output by the current video generation model, wherein the resolution of the second-resolution video output by the current video generation model is higher than the resolution of the second-resolution video output by the previous video generation model.
[0098] For example, the first step above can be understood as follows: the decoder inputs the first restored feature representation into the first video generation model in the hybrid model. The model uses the first restored feature representation as a semantic constraint to gradually remove noise from the initial Gaussian noise, reconstruct the video sequence with the lowest target resolution, and outputs the first resolution video, which is the video with the lowest resolution among the resolution videos output by all video generation models in the hybrid model.
[0099] For example, the second step described above can be understood as follows: For subsequent video generation models in the hybrid model (such as the second video generation model, the third video generation model, etc.), the second-resolution video output by the previous video generation model of the current video generation model is used as the structure and motion prior, and the first restored feature representation is used as the semantic consistency anchor point. Therefore, the second-resolution video output by the previous video generation model of the current video generation model is upsampled to expand its resolution to the resolution of the video output by the current video generation model, resulting in a third-resolution video. The third-resolution video and the first restored feature representation are then merged into an enhancement conditional input to drive the generation of a higher-resolution output result.
[0100] Through the embodiments of this application, the decoding end analyzes the first recovered feature representation based on a hybrid expert model of shared semantic conditions and cascaded incremental enhancement. This allows each subsequent video generation model to supplement the missing high-frequency details and structural information by combining a unified high-order feature representation with the low-resolution video output by the previous video generation model, rather than reconstructing from scratch or repeatedly encoding multiple layers of bitstream. This solves the technical problems of traditional scalable video coding requiring the transmission of multiple layers of bitstream to reconstruct multi-resolution video at the decoding end, resulting in large storage and bandwidth overhead, and existing generative coding methods being unable to achieve adaptive multi-resolution output under a single bitstream.
[0101] The video encoding method in the embodiments of this application will be described below with reference to optional examples.
[0102] If the encoding end has already processed a raw 1080p (1920×1080) video, it generates a unique compressed feature representation. This feature indicates that it contains only about 0.05% of the original video data, but carries its complete semantic structure, motion trajectory, and key texture information, and is sent to the decoding end as the sole transmission bitstream.
[0103] The decoding end deploys three upsampling modules and a hybrid model, which includes three cascaded visual generation models for outputting 480p, 720p, and 1080p video, respectively. Therefore, if the decoding end receives a routing packet (containing compressed feature representation) sent by the encoding segment... And after its target resolution of 1080p, it can be processed as follows: Figure 4 The flowchart shown generates videos in 480p, 720p, and 1080p resolutions as needed:
[0104] The decoding end can first use the first upsampling module to represent the compressed features. Upsampling is performed to restore the feature representation to the same spatial dimensions as the original 1080p video. .
[0105] If the decoder detects that the terminal device's screen resolution (such as a low-performance mobile phone) is limited to 480p and the network bandwidth is only 1.5Mbps, in order to respond quickly, the decoder can start only the first video generation model and restore the feature representation. The input is fed into the first video generation model that has been started, resulting in a 240p video output by the first video generation model. Then, it is upsampled to 480p through bilinear interpolation to obtain a 480p video.
[0106] If the decoding end detects that the maximum resolution of the terminal device screen (such as a tablet) is 720p and the network bandwidth is sufficient (such as 5Mbps), then the decoding end can start the first video generation model and the second video generation model, and recover the feature representation. The input is fed into the first video generation model that has been started, resulting in a 240p video output by the first video generation model; then the 240p video is upsampled to obtain a 720p video, and the 720p video and the restored feature representation are then combined. The input is fed into the second video generation model that has already been started, resulting in a 720p video.
[0107] If the decoding end detects that the resolution limit of the terminal device screen (such as a 4K display) is 4K, and the network bandwidth is sufficient (such as 200Mbps), then the decoding end can start three video generation models and recover the feature representation. The input is fed into the first video generation model that has been started, resulting in a 240p video output by the first video generation model; then the 240p video and the restored feature representation are combined. The input is fed into the second video generation model that has already been started, resulting in a 720p video; the 720p video is then upsampled to obtain a 1080p video, and the 1080p video and the restored feature representation are then combined. The input is fed into the third video generation model that has been started, resulting in a 1080p video.
[0108] Thus, the decoding end can achieve "one-time transmission, multi-end adaptation, and high-quality reconstruction" at the same bitrate, generating high-quality 480p, 720p, and 1080p videos on different devices respectively, without having to pre-encode the corresponding bitstream for each resolution at the encoding end, greatly reducing bitstream storage and transmission bandwidth.
[0109] In summary, the embodiments of this application innovatively realize a generative video coding closed loop of single-stream drive, multi-resolution adaptation, and high-quality reconstruction through the generation mechanism of "condition sharing and cascading enhancement," which is significantly better than the resource waste of traditional multi-stream schemes and the rigid resolution limitations of existing generation models.
[0110] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0112] According to another aspect of the embodiments of this application, a video encoding apparatus for implementing the video encoding method provided in the above embodiments is also provided, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0113] Figure 5 This is a structural block diagram of an optional video encoding apparatus according to an embodiment of this application, such as... Figure 5 As shown, the video encoding device includes:
[0114] The acquisition module 52 is used to acquire the routing data packet sent by the encoding end, wherein the routing data packet includes at least: the target resolution of the first video and the first compression feature representation corresponding to the first video;
[0115] Upsampling module 54 is used to upsample the first compressed feature representation according to the target resolution to obtain the first recovered feature representation;
[0116] The encoding module 56 is used to analyze the first restored feature representation using a pre-trained hybrid model to obtain multiple resolution videos corresponding to the first video. The hybrid model includes multiple cascaded video generation models, and the input of the first video generation model is the first restored feature representation. The input of each video generation model after the first video generation model is the resolution video output by the previous video generation model and the first restored feature representation.
[0117] In an exemplary embodiment, the above-described apparatus is further configured to train a hybrid model by means of: constructing an initial model composed of multiple cascaded video generation models; acquiring multiple sets of training sample data, wherein each set of training sample data includes: a second video and a second compressed feature representation corresponding to the second video; and iteratively training the initial model using the multiple sets of training sample data to obtain a hybrid model.
[0118] In an exemplary embodiment, the above-described apparatus is further configured to acquire multiple sets of training sample data by means of the following method: acquiring multiple second videos, wherein each second video has a different resolution; for each second video, encoding and compressing the second video to obtain a corresponding low-dimensional feature representation, performing channel expansion on the low-dimensional feature representation to obtain a corresponding enhanced feature representation, performing multi-level downsampling on the enhanced feature representation to obtain a corresponding compressed enhanced feature representation, and performing masking processing on the compressed enhanced feature representation to obtain a second compressed feature representation of the second video; and the multiple sets of training sample data are composed of the multiple second videos and the second compressed feature representation of each second video.
[0119] In an exemplary embodiment, the above-described apparatus is further configured to perform multi-level downsampling on the enhanced feature representation to obtain a corresponding compressed enhanced feature representation, including: compressing the time dimension, height dimension, and width dimension of the enhanced feature representation using a three-dimensional convolutional layer with a stride of a first value to obtain a low-degree compressed enhanced feature representation; fusing the result of downsampling the enhanced feature representation using a three-dimensional convolutional layer with a stride of the first value with the result of feature enhancement of the low-degree compressed enhanced feature representation to obtain a first intermediate compressed enhanced feature representation; and using a three-dimensional convolutional layer with a stride of a second value to compress the time dimension, height dimension, and width dimension of the first intermediate compressed enhanced feature representation. The first intermediate compressed enhanced feature representation is obtained by compressing the features separately. The result of downsampling the first intermediate compressed enhanced feature representation using a 3D convolutional layer with a stride of the second value is fused with the result of feature enhancement of the intermediate compressed enhanced feature representation to obtain the second intermediate compressed enhanced feature representation. The height and width dimensions of the second intermediate compressed enhanced feature representation are compressed using a 3D convolutional layer with a stride of the third value to obtain the highly compressed enhanced feature representation. The result of downsampling the enhanced feature representation using a 3D convolutional layer with a stride of the product of the first, second, and third values is fused with the highly compressed enhanced feature representation to obtain the compressed enhanced feature representation.
[0120] In an exemplary embodiment, the above-described apparatus is further configured to perform masking processing on the compression enhancement feature representation to obtain a second compressed feature representation of the second video by means of: performing masking processing on the compression enhancement feature representation according to preset masking information to obtain a corresponding mask tensor, wherein the masking information includes at least one of the following: number of occlusion blocks, spatial size of each occlusion block, temporal span, and channel span; multiplying the mask tensor element-wise with the compression enhancement feature representation to obtain a second compressed feature representation of the second video.
[0121] In an exemplary embodiment, the above-described apparatus is further configured to utilize a pre-trained hybrid model to analyze the first restored feature representation to obtain multiple resolution videos corresponding to the first video, including: for the first video generation model in the hybrid model, inputting the first restored feature representation into the first video generation model to obtain the first resolution video output by the first video generation model; for each video generation model in the hybrid model after the first video generation model, upsampling the second resolution video output by the previous video generation model of the current video generation model to obtain a third resolution video, wherein the resolution of the third resolution video is the same as the resolution of the resolution video output by the current video generation model; inputting the third resolution video and the first restored feature representation into the current video generation model to obtain the second resolution video output by the current video generation model, wherein the resolution of the second resolution video output by the current video generation model is higher than the resolution of the second resolution video output by the previous video generation model of the current video generation model.
[0122] In an exemplary embodiment, the above-described apparatus is further configured to upsample the first compressed feature representation according to the target resolution by the following method to obtain a first restored feature representation, including: sampling and interpolating the first compressed feature representation in the height and width dimensions to obtain a first restored feature representation with the same spatial size as the target resolution of the first video.
[0123] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0124] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein the program executes the steps in any of the above method embodiments when it is run.
[0125] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, ROMs, RAMs, portable hard drives, magnetic disks, or optical disks.
[0126] According to another aspect of the embodiments of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor is configured to perform the steps of any of the method embodiments described above via the computer program. In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0127] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0128] According to another aspect of the embodiments of this application, a computer program product is also provided, the computer program product including a computer program / instructions containing program code for performing the method shown in the flowchart.
[0129] Figure 6 A structural block diagram of an electronic device for implementing embodiments of this application is shown. Figure 6 As shown, the electronic device 60 may include one or more processors 602 (shown as 602a, 602b, ..., 602n in the figure) (processor 602 may include, but is not limited to, a processing device such as a microprocessor or programmable logic device), a memory 604 for storing data, and a transmission device 606 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 6 The structure shown is for illustrative purposes only and does not limit the structure of the computer system described above. For example, electronic device 60 may also include... Figure 6 The more or fewer components shown, or having the same Figure 6 The different configurations shown.
[0130] It should be noted that the aforementioned one or more processors 602 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element of the electronic device 60. As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0131] The memory 604 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the video encoding method in this embodiment. The processor 602 executes various functional applications and data processing by running the software programs and modules stored in the memory 604, thereby implementing the video encoding method of the aforementioned application. The memory 604 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 604 may further include memory remotely located relative to the processor 602, and these remote memories can be connected to the electronic device 60 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0132] The transmission device 606 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the electronic device 60. In one example, the transmission device 606 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 606 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0133] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows the user to interact with the user interface of the electronic device 60.
[0134] It should be noted that, Figure 6 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0135] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0136] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A video encoding method, characterized in that, include: Obtain the routing data packet sent by the encoding end, wherein the routing data packet includes at least: the target resolution of the first video and the first compression feature representation corresponding to the first video; The first compressed feature representation is upsampled according to the target resolution to obtain the first recovered feature representation; The first restored feature representation is analyzed using a pre-trained hybrid model to obtain multiple resolution videos corresponding to the first video. The hybrid model includes multiple cascaded video generation models, and the input of the first video generation model is the first restored feature representation. The input of each video generation model after the first video generation model is the resolution video output by the previous video generation model and the first restored feature representation. The hybrid model is obtained by iterative training using multiple sets of training sample data, and each set of training sample data includes: a second video and a second compressed feature representation corresponding to the second video. The process of acquiring multiple sets of training sample data includes: acquiring multiple second videos, each with a different resolution; for each second video, encoding and compressing the second video to obtain a corresponding low-dimensional feature representation; performing channel expansion on the low-dimensional feature representation to obtain a corresponding enhanced feature representation; performing multi-level downsampling on the enhanced feature representation to obtain a corresponding compressed enhanced feature representation; and performing masking processing on the compressed enhanced feature representation to obtain a second compressed feature representation of the second video; multiple sets of training sample data are composed of multiple second videos and the second compressed feature representation of each second video. The enhanced feature representation is downsampled at multiple levels to obtain a corresponding compressed enhanced feature representation, including: compressing the time, height, and width dimensions of the enhanced feature representation using a 3D convolutional layer with a stride of a first value to obtain a low-compressed enhanced feature representation; fusing the result of downsampling the enhanced feature representation using a 3D convolutional layer with a stride of the first value with the result of feature enhancement of the low-compressed enhanced feature representation to obtain a first intermediate compressed enhanced feature representation; and then compressing the time, height, and width dimensions of the first intermediate compressed enhanced feature representation using a 3D convolutional layer with a stride of a second value to obtain a medium-compressed enhanced feature representation. Feature representation: The result of downsampling the first intermediate compressed enhanced feature representation using a 3D convolutional layer with a stride of the second value is fused with the result of feature enhancement of the moderately compressed enhanced feature representation to obtain a second intermediate compressed enhanced feature representation. The height and width dimensions of the second intermediate compressed enhanced feature representation are compressed using a 3D convolutional layer with a stride of the third value to obtain a highly compressed enhanced feature representation. The result of downsampling the enhanced feature representation using a 3D convolutional layer with a stride equal to the product of the first, second, and third values is fused with the highly compressed enhanced feature representation to obtain the compressed enhanced feature representation.
2. The method according to claim 1, characterized in that, The training process of the hybrid model includes: Construct an initial model consisting of multiple cascaded video generation models; Acquire multiple sets of training sample data, wherein each set of training sample data includes: a second video and a second compressed feature representation corresponding to the second video; The initial model is iteratively trained using multiple sets of training sample data to obtain the hybrid model.
3. The method according to claim 1, characterized in that, The compression enhancement feature representation is masked to obtain the second compression feature representation of the second video, including: The compressed and enhanced feature representation is masked according to the preset mask information to obtain the corresponding mask tensor, wherein the mask information includes at least one of the following: number of masking blocks, spatial size of each masking block, time span, and channel span; The mask tensor is multiplied element-wise with the compression enhancement feature representation to obtain the second compression feature representation of the second video.
4. The method according to claim 1, characterized in that, The first recovered feature representation is analyzed using a pre-trained hybrid model to obtain multiple resolution videos corresponding to the first video, including: For the first video generation model in the hybrid model, the first restored feature representation is input into the first video generation model to obtain the first resolution video output by the first video generation model; For each video generation model in the hybrid model after the first video generation model, the second-resolution video output by the previous video generation model of the current video generation model is upsampled to obtain a third-resolution video, wherein the resolution of the third-resolution video is the same as the resolution of the video output by the current video generation model; the third-resolution video and the first restored feature representation are input into the current video generation model to obtain the second-resolution video output by the current video generation model, wherein the resolution of the second-resolution video output by the current video generation model is higher than the resolution of the second-resolution video output by the previous video generation model of the current video generation model.
5. The method according to claim 1, characterized in that, Upsampling the first compressed feature representation according to the target resolution yields a first recovered feature representation, including: The first compressed feature representation is sampled and interpolated in the height and width dimensions to obtain a first restored feature representation with the same spatial size as the target resolution of the first video.
6. A video encoding device, characterized in that, include: The acquisition module is used to acquire the routing data packet sent by the encoding end, wherein the routing data packet includes at least: the target resolution of the first video and the first compression feature representation corresponding to the first video; The upsampling module is used to upsample the first compressed feature representation according to the target resolution to obtain the first recovered feature representation; An encoding module is used to analyze the first restored feature representation using a pre-trained hybrid model to obtain multiple resolution videos corresponding to the first video. The hybrid model includes multiple cascaded video generation models, and the input of the first video generation model is the first restored feature representation. The input of each video generation model after the first video generation model is the resolution video output by the previous video generation model and the first restored feature representation. The hybrid model is obtained by iterative training using multiple sets of training sample data, and each set of training sample data includes: a second video and a second compressed feature representation corresponding to the second video. The process of acquiring multiple sets of training sample data includes: acquiring multiple second videos, each with a different resolution; for each second video, encoding and compressing the second video to obtain a corresponding low-dimensional feature representation; performing channel expansion on the low-dimensional feature representation to obtain a corresponding enhanced feature representation; performing multi-level downsampling on the enhanced feature representation to obtain a corresponding compressed enhanced feature representation; and performing masking processing on the compressed enhanced feature representation to obtain a second compressed feature representation of the second video; multiple sets of training sample data are composed of multiple second videos and the second compressed feature representation of each second video. The enhanced feature representation is downsampled at multiple levels to obtain a corresponding compressed enhanced feature representation, including: compressing the time, height, and width dimensions of the enhanced feature representation using a 3D convolutional layer with a stride of a first value to obtain a low-compressed enhanced feature representation; fusing the result of downsampling the enhanced feature representation using a 3D convolutional layer with a stride of the first value with the result of feature enhancement of the low-compressed enhanced feature representation to obtain a first intermediate compressed enhanced feature representation; and then compressing the time, height, and width dimensions of the first intermediate compressed enhanced feature representation using a 3D convolutional layer with a stride of a second value to obtain a medium-compressed enhanced feature representation. Feature representation: The result of downsampling the first intermediate compressed enhanced feature representation using a 3D convolutional layer with a stride of the second value is fused with the result of feature enhancement of the moderately compressed enhanced feature representation to obtain a second intermediate compressed enhanced feature representation. The height and width dimensions of the second intermediate compressed enhanced feature representation are compressed using a 3D convolutional layer with a stride of the third value to obtain a highly compressed enhanced feature representation. The result of downsampling the enhanced feature representation using a 3D convolutional layer with a stride equal to the product of the first, second, and third values is fused with the highly compressed enhanced feature representation to obtain the compressed enhanced feature representation.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the device containing the computer-readable storage medium executes the video encoding method according to any one of claims 1 to 5 by running the computer program.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the video encoding method according to any one of claims 1 to 5 through the computer program.
Citation Information
Patent Citations
Method and a device for generating an image super-resolution model
CN109872276A
Encoding and Decoding Method, and Apparatus
US20240223790A1