Generative video coding and decoding method and device and electronic equipment

By using generative video coding methods, utilizing the core bitstream and auxiliary information segments with semantic and motion features, combined with a multidimensional stable representation system and a generative diffusion model, stable video reconstruction in weak network environments is achieved. This solves the problems of video quality collapse and continuity interruption, and improves the viewing experience and system stability under extreme network conditions.

CN122069352AActive Publication Date: 2026-05-19CHINA TELECOM CORP LTD
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA TELECOM CORP LTD
Filing Date
2026-04-17
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing video coding technologies lack a stable reconstruction mechanism for lost bitstreams in weak network environments, leading to video quality collapse or continuous interruption, and making it impossible to achieve stable decoding without bandwidth increments.

Method used

A generative video coding method is adopted. By determining the core bitstream of semantic and motion features of the original video data, as well as auxiliary information segments, a second bitstream is generated. When the packet loss rate is less than or equal to a preset threshold, the missing features are reconstructed through a multidimensional stable representation system. When the packet loss rate is greater than the preset threshold, the missing features are reconstructed through a generative diffusion model.

Benefits of technology

It achieves highly robust video reconstruction without bandwidth increment or transmission latency increase, significantly improving the viewing experience and system stability under extreme network conditions, and solving the problem of video quality collapse or continuous interruption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069352A_ABST
    Figure CN122069352A_ABST
Patent Text Reader

Abstract

The invention discloses a generative video coding and decoding method and device and electronic equipment. The method comprises the following steps: determining a first code stream of original video data and a corresponding auxiliary information segment; generating a second code stream according to the first code stream and the auxiliary information segment; determining a packet loss rate of the second code stream after channel transmission, reconstructing missing features in the second code stream through a multi-dimensional stable representation system under the condition that the packet loss rate is smaller than or equal to a preset threshold value to obtain a first reconstructed code stream, and outputting a video picture according to the first reconstructed code stream; and under the condition that the packet loss rate is greater than a preset threshold value, reconstructing missing features in the second code stream through a generative diffusion model to obtain a second reconstructed code stream, and outputting a video picture according to the second reconstructed code stream. According to the method and the device, the technical problem of video image quality collapse or continuity interruption caused by lack of a stable reconstruction mechanism after code stream loss in a weak network environment in a video coding technology in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video encoding and decoding technology, and more specifically, to a generative video encoding and decoding method, apparatus, and electronic device. Background Technology

[0002] As video communication is increasingly used in emergency rescue, ocean operations, polar scientific research and other scenarios, the demand for video transmission in weak network environments continues to increase, and stable video transmission is the key to ensuring decision-making and execution.

[0003] Current mainstream video coding technologies (such as H.265 / HEVC, AV1, and H.266 / VVC) are all based on pixel-level compression and predictive coding principles. They achieve video data compression and reconstruction by quantizing residual blocks, transmitting motion vectors, and macroblock information. Their encoding process highly depends on the accurate transmission of the complete bitstream. The decoding end needs to receive and reconstruct the original pixel data block by block. Any loss of bitstream will lead to reconstruction failure, manifesting as visual degradation phenomena such as blurry images, stuttering, and black screens. In real-world network environments, especially in weak network scenarios such as ocean communication, mountain surveillance, emergency rescue, and polar scientific expeditions, the packet loss rate often exceeds 5% or even reaches over 10%. Traditional video coding technologies show a significant decline in subjective quality when the packet loss rate exceeds 1%, making it difficult to maintain basic video continuity and availability.

[0004] To mitigate the impact of packet loss, the industry often employs data retransmission or redundant coding mechanisms. However, data retransmission increases transmission latency and cannot meet the needs of real-time video communication. Redundant coding, on the other hand, increases the bitstream size, further consuming already scarce bandwidth resources, which contradicts the low bandwidth requirements of weak network environments.

[0005] In recent years, generative artificial intelligence technology has made breakthroughs in the application of video coding, achieving an upgrade from pixel compression to feature coding. However, its architecture design still focuses on improving compression efficiency, without being specifically optimized for packet loss scenarios. It lacks the ability to structurally reconstruct the bitstream after loss and cannot achieve stable decoding without bandwidth increments.

[0006] There is currently no effective solution to the above problems. Summary of the Invention

[0007] This application provides a generative video encoding and decoding method, apparatus, and electronic device to at least solve the technical problem of video quality collapse or continuous interruption caused by the lack of a stable reconstruction mechanism for bitstream loss in video encoding technology in weak network environments.

[0008] According to one aspect of the embodiments of this application, a generative video coding and decoding method is provided, comprising: determining a first bitstream of original video data, wherein the first bitstream is a core bitstream containing semantic features and motion features of the original video data; determining an auxiliary information segment corresponding to the original video data, and generating a second bitstream based on the first bitstream and the auxiliary information segment, wherein the auxiliary information segment includes at least video parameter information and stable representation information of the original video data, and the stable representation information is a structured prior feature used to constrain the reconstruction process at the decoding end; determining the packet loss rate of the second bitstream after transmission through the channel, and when the packet loss rate is less than or equal to a preset threshold, reconstructing the missing features in the second bitstream through a multidimensional stable representation system to obtain a first reconstructed bitstream, and outputting a video frame based on the first reconstructed bitstream, wherein the multidimensional stable representation system is used to represent the multidimensional representation constraints corresponding to the reconstructed missing features; and when the packet loss rate is greater than the preset threshold, reconstructing the missing features in the second bitstream through a generative diffusion model to obtain a second reconstructed bitstream, and outputting a video frame based on the second reconstructed bitstream.

[0009] Optionally, determining the first bitstream of the original video data includes: receiving the original video data; extracting features from the original video data to obtain semantic features and motion features corresponding to the original video data; quantizing and encoding the semantic features and motion features respectively to obtain semantic feature encoding units corresponding to the semantic features and motion feature encoding units corresponding to the motion features; and fusing the semantic feature encoding units and motion feature encoding units to obtain the first bitstream corresponding to the original video data.

[0010] Optionally, determining the auxiliary information segment corresponding to the original video data includes: determining the video parameter information of the original video data, wherein the video parameter information includes at least one of the following: the resolution, frame rate, and scene type identifier of the original video data; generating stable representation information based on the first bitstream, wherein the vector dimension of the stable representation information is lower than the vector dimension of the first bitstream; determining the auxiliary information segment corresponding to the original video data based on the video parameter information and the stable representation information, wherein the auxiliary information segment carries a unique identifier bit.

[0011] Optionally, determining the packet loss rate of the second bitstream after transmission through the channel includes: separating the second bitstream after transmission through the channel based on the identifier bit to obtain a separated first bitstream and an auxiliary information segment; performing an integrity check on the first bitstream obtained from the separation of the second bitstream to obtain a check result; if the check result indicates that the check passed, directly outputting the video frame corresponding to the original video data based on the video parameter information in the auxiliary information segment; if the check result indicates that the check failed, determining the packet loss rate and missing characteristics of the second bitstream.

[0012] Optionally, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system, including: when the packet loss rate is less than or equal to a preset threshold, constructing a multidimensional stable representation system containing scene structure representation, dynamic continuity representation and visual prior representation based on the stable representation information in the auxiliary information segment; and reconstructing the missing features in the second bitstream through the multidimensional stable representation system.

[0013] Optionally, the missing features in the second bitstream are reconstructed using a multidimensional stable representation system, including: determining the original location of the missing features based on scene structure constraints in the multidimensional stable representation system; determining the motion trajectory of the missing features based on dynamic continuity constraints in the multidimensional stable representation system; and determining the local texture features of the second bitstream based on visual prior constraints in the multidimensional stable representation system; determining the temporal correlation of the second bitstream, wherein the temporal correlation is used to reflect the temporal consistency between semantic features and motion features in the second bitstream; and reconstructing the missing features in the second bitstream based on the original location, motion trajectory, local texture features, and temporal correlation.

[0014] Optionally, the missing features in the second bitstream are reconstructed using a generative diffusion model, including: when the packet loss rate is greater than a preset threshold, the missing features in the second bitstream are reconstructed using a multidimensional stable representation system to obtain a third bitstream, wherein the third bitstream also contains target missing features that cannot be recovered by the multidimensional stable representation system; the third bitstream and auxiliary information segments are processed using a generative diffusion model to reconstruct the target missing features in the third bitstream.

[0015] Optionally, the third bitstream and auxiliary information segment are processed by a generative diffusion model to reconstruct the target missing features in the third bitstream, including: generating a fusion feature vector based on video information parameters and scene type identifiers in the third bitstream and auxiliary information segment; determining the semantic constraints of the auxiliary information segment and the spatiotemporal constraints of the original video data, and processing the fusion feature vector based on the semantic constraints and spatiotemporal constraints through the noise prediction network in the generative diffusion model to obtain a missing feature compensation vector; and reconstructing the target missing features based on the missing feature compensation vector.

[0016] According to another aspect of the embodiments of this application, a generative video coding method is also provided, comprising: determining a first bitstream of original video data, wherein the first bitstream is a core bitstream containing semantic features and motion features of the original video data; determining an auxiliary information segment corresponding to the original video data, and generating a second bitstream based on the first bitstream and the auxiliary information segment, wherein the auxiliary information segment includes at least video parameter information and stability representation information of the original video data, and the stability representation information is a structured prior feature used to constrain the reconstruction process at the decoding end; and transmitting the second bitstream to the decoding end for processing.

[0017] Optionally, determining the first bitstream of the original video data includes: receiving the original video data; extracting features from the original video data to obtain semantic features and motion features corresponding to the original video data; quantizing and encoding the semantic features and motion features respectively to obtain semantic feature encoding units corresponding to the semantic features and motion feature encoding units corresponding to the motion features; and fusing the semantic feature encoding units and motion feature encoding units to obtain the first bitstream corresponding to the original video data.

[0018] According to another aspect of the embodiments of this application, a generative video decoding method is also provided, comprising: receiving a second bitstream transmitted by an encoding end and determining the packet loss rate of the second bitstream, wherein the second bitstream is generated based on a first bitstream and auxiliary information segments, the first bitstream being a core bitstream containing semantic features and motion features of the original video data, and the auxiliary information segments including at least video parameter information and stable representation information of the original video data, the stable representation information being structured prior features used to constrain the reconstruction process at the decoding end; when the packet loss rate is less than or equal to a preset threshold, reconstructing the missing features in the second bitstream through a multidimensional stable representation system to obtain a first reconstructed bitstream, and outputting video images based on the first reconstructed bitstream, wherein the multidimensional stable representation system is used to represent the multidimensional representation constraints corresponding to the reconstructed missing features; when the packet loss rate is greater than the preset threshold, reconstructing the missing features in the second bitstream through a generative diffusion model to obtain a second reconstructed bitstream, and outputting video images based on the second reconstructed bitstream.

[0019] Optionally, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system, including: when the packet loss rate is less than or equal to a preset threshold, constructing a multidimensional stable representation system containing scene structure representation, dynamic continuity representation, and visual prior representation based on the stable representation information in the auxiliary information segment; determining the original position of the missing features based on the scene structure constraints in the multidimensional stable representation system, determining the motion trajectory of the missing features based on the dynamic continuity constraints in the multidimensional stable representation system, and determining the local texture features of the second bitstream based on the visual prior constraints in the multidimensional stable representation system; determining the temporal correlation of the second bitstream, wherein the temporal correlation is used to reflect the temporal consistency relationship between semantic features and motion features in the second bitstream; and reconstructing the missing features in the second bitstream based on the original position, motion trajectory, local texture features, and temporal correlation.

[0020] Optionally, the generative diffusion model reconstructs the missing features in the second bitstream, including: when the packet loss rate is greater than a preset threshold, reconstructing the missing features in the second bitstream through a multidimensional stable representation system to obtain a third bitstream, wherein the third bitstream still contains target missing features that the multidimensional stable representation system cannot recover; generating a fusion feature vector based on the third bitstream, video information parameters in the auxiliary information segment, and scene type identifier; determining the semantic constraints of the auxiliary information segment and the spatiotemporal constraints of the original video data, and processing the fusion feature vector based on the semantic constraints and spatiotemporal constraints through the noise prediction network in the generative diffusion model to obtain a missing feature compensation vector; and reconstructing the target missing features based on the missing feature compensation vector.

[0021] According to another aspect of the embodiments of this application, a generative video encoding and decoding apparatus is also provided, comprising: a determining module, configured to determine a first bitstream of original video data, wherein the first bitstream is a core bitstream containing semantic features and motion features of the original video data; a generating module, configured to determine an auxiliary information segment corresponding to the original video data, and generate a second bitstream based on the first bitstream and the auxiliary information segment, wherein the auxiliary information segment includes at least video parameter information and stable representation information of the original video data, and the stable representation information is a structured prior feature used to constrain the reconstruction process at the decoding end; a first reconstruction module, configured to determine the packet loss rate of the second bitstream after transmission through the channel, and, if the packet loss rate is less than or equal to a preset threshold, reconstruct the missing features in the second bitstream through a multidimensional stable representation system to obtain a first reconstructed bitstream, and output video images based on the first reconstructed bitstream, wherein the multidimensional stable representation system is used to represent the multidimensional representation constraints corresponding to the reconstructed missing features; and a second reconstruction module, configured to, if the packet loss rate is greater than a preset threshold, reconstruct the missing features in the second bitstream through a generative diffusion model to obtain a second reconstructed bitstream, and output video images based on the second reconstructed bitstream.

[0022] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, the non-volatile storage medium including a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-described generative video encoding / decoding, generative video encoding, or generative video decoding method by running the computer program.

[0023] According to another aspect of the embodiments of this application, a computer program product is also provided, including computer instructions that, when executed by a processor, implement the above-described generative video encoding / decoding, generative video encoding, or generative video decoding methods.

[0024] In this embodiment, a first bitstream of the original video data is determined, wherein the first bitstream is a core bitstream containing semantic and motion features of the original video data; an auxiliary information segment corresponding to the original video data is determined, and a second bitstream is generated based on the first bitstream and the auxiliary information segment, wherein the auxiliary information segment includes at least video parameter information and stable representation information of the original video data, and the stable representation information is a structured prior feature used to constrain the reconstruction process at the decoding end; the packet loss rate of the second bitstream after transmission through the channel is determined, and if the packet loss rate is less than or equal to a preset threshold, the missing features in the second bitstream are reconstructed through a multi-dimensional stable representation system to obtain a first reconstructed bitstream, and the video image is output based on the first reconstructed bitstream, wherein the multi-dimensional stable representation system... A multi-dimensional stable representation system is used to represent the multi-dimensional representation constraints corresponding to the missing features in the reconstruction. When the packet loss rate is greater than a preset threshold, the missing features in the second bitstream are reconstructed through a generative diffusion model to obtain the second reconstructed bitstream. The video image is then output based on the second reconstructed bitstream, achieving the goal of outputting visually coherent and natural video images based on a dual reconstruction path mechanism. This achieves highly robust video reconstruction without bandwidth increment or transmission delay increase, significantly improving the viewing experience and system stability under extreme network conditions. Furthermore, it solves the technical problem of video quality collapse or continuous interruption caused by the lack of a stable reconstruction mechanism for bitstream loss in video coding technology in weak network environments. Attached Figure Description

[0025] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0026] Figure 1 This is a hardware structure diagram of a computer terminal for implementing generative video coding and decoding, generative video coding, or generative video decoding methods according to an embodiment of this application.

[0027] Figure 2 This is a flowchart of a generative video encoding and decoding method according to an embodiment of this application;

[0028] Figure 3 This is a flowchart illustrating the technical architecture of a generative video encoding and decoding method according to an embodiment of this application.

[0029] Figure 4 This is a schematic diagram of a hierarchical completion and reconstruction process according to an embodiment of this application;

[0030] Figure 5 This is a flowchart of a generative video coding method according to an embodiment of this application;

[0031] Figure 6 This is a flowchart of a generative video decoding method according to an embodiment of this application;

[0032] Figure 7 This is a structural diagram of a generative video encoding and decoding apparatus according to an embodiment of this application. Detailed Implementation

[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0035] First, some nouns or terms that appear in the explanation of the embodiments of this application shall be interpreted as follows:

[0036] GVC (Generative Video Coding) refers to encoding video content into high-level abstract representations (tokens) of semantic and motion features based on generative artificial intelligence technology, rather than traditional pixel or residual data, thereby achieving low-bandwidth, highly intelligent video compression and reconstruction.

[0037] AI Flow (Artificial Intelligence Flow): also known as "AI Flow Network", is a new video transmission theory system. Its core concept is to upgrade video transmission from a "bandwidth-intensive" mode that relies on bandwidth resources to a "intelligence-intensive" mode that relies on intelligent reasoning capabilities. Through intelligent feature extraction at the encoding end and intelligent reconstruction at the decoding end, it achieves efficient and robust video transmission in weak network environments.

[0038] Token: In this application, it refers to the high-level feature coding unit of the video after feature extraction. It is divided into semantic token and motion token, which respectively represent static information such as scene content and object structure in the video, and temporal information such as object motion trajectory and dynamic changes. It is a lightweight expression form that replaces traditional macroblocks or DCT coefficients.

[0039] Stable representation: refers to the structured prior features used to constrain the video reconstruction process at the decoding end, including scene structure representation, dynamic continuity representation and visual prior representation, which respectively provide prior knowledge of spatial layout, motion inertia and visual perception laws, to ensure that the reconstructed picture conforms to the laws of the real world in terms of structure and dynamics.

[0040] Diffusion model: A generative artificial intelligence model that gradually generates high-quality image or video frames from random noise by simulating noise addition and inverse denoising processes. In this application, it serves as the ultimate completion method in scenarios with severe packet loss, combining semantic and motion features to generate details and textures of missing images, preventing image collapse.

[0041] To address the issue of poor video transmission quality in related technologies, this application provides a generative video encoding and decoding method, a generative video encoding method, and a generative video decoding method. This method can be implemented in... Figure 1 The computer terminal shown is described below.

[0042] The generative video encoding / decoding, generative video encoding, or generative video decoding method embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computer terminal for implementing generative video coding and decoding, generative video coding, or generative video decoding methods is shown. Figure 1 As shown, the computer terminal 10 may include one or more processors (shown as 102a, 102b, ..., 102n in the figure) (the processor may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions connected via wired and / or wireless networks. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0043] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0044] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the generative video encoding / decoding, generative video encoding, or generative video decoding methods in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the aforementioned generative video encoding / decoding, generative video encoding, or generative video decoding methods. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0045] The transmission module 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission module 106 may be a radio frequency (RF) module, used for wireless communication with the Internet.

[0046] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.

[0047] It should be noted here that, in some optional embodiments, the above... Figure 1The computer terminal shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computer terminal.

[0048] In the above operating environment, the embodiments of this application provide an embodiment of a generative video encoding and decoding, generative video encoding, or generative video decoding method. It should be noted that the steps shown in the flowcharts in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0049] Figure 2 This is a flowchart of a generative video encoding and decoding method according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:

[0050] Step S202: Determine the first bitstream of the original video data, wherein the first bitstream is the core bitstream containing the semantic features and motion features of the original video data.

[0051] Step S204: Determine the auxiliary information segment corresponding to the original video data, and generate a second bitstream based on the first bitstream and the auxiliary information segment. The auxiliary information segment includes at least the video parameter information and stable characterization information of the original video data. The stable characterization information is a structured prior feature used to constrain the reconstruction process at the decoding end.

[0052] Step S206: Determine the packet loss rate of the second bitstream after transmission through the channel. If the packet loss rate is less than or equal to a preset threshold, reconstruct the missing features in the second bitstream through a multidimensional stable representation system to obtain a first reconstructed bitstream, and output the video image based on the first reconstructed bitstream. The multidimensional stable representation system is used to represent the multidimensional representation constraints corresponding to the reconstructed missing features.

[0053] Step S208: When the packet loss rate is greater than a preset threshold, the missing features in the second bitstream are reconstructed using a generative diffusion model to obtain the second reconstructed bitstream, and the video image is output based on the second reconstructed bitstream.

[0054] Through the above steps S202 to S208, the goal of outputting visually coherent and natural video images based on the dual reconstruction path mechanism is achieved. This realizes highly robust video reconstruction without bandwidth increment or transmission delay increase, significantly improving the viewing experience and system stability under extreme network conditions. It also solves the technical problem of video quality collapse or continuous interruption caused by the lack of a stable reconstruction mechanism for bitstream loss in video coding technology in weak network environments.

[0055] Figure 3 This is a flowchart illustrating the technical architecture of a generative video encoding and decoding method according to an embodiment of this application, such as... Figure 3 As shown, this generative encoding and decoding method with packet loss resistance mainly involves three core modules: a generative tokenization encoding module, a packet loss-resistant enhanced bitstream construction module, and a generative completion decoding module. The generative tokenization encoding module encodes the original video into a core bitstream (i.e., the first bitstream mentioned above) with a dual-token structure containing semantic and motion features. This module abandons traditional pixel compression and achieves intelligent encoding with low bandwidth and high semantic density. The packet loss-resistant enhanced bitstream construction module embeds ≤5% lightweight auxiliary information (including video parameter information and stable representation information) into the core bitstream to construct a packet loss-resistant enhanced bitstream without additional bandwidth. The generative completion decoding module performs bitstream parsing and packet loss detection, and performs hierarchical completion and reconstruction according to the degree of packet loss, such as calling a multi-dimensional stable representation system for reconstruction or generative diffusion model completion, achieving intelligent reconstruction capability that outputs coherent and clear video even with packet loss exceeding 10%. The following section combines... Figure 3 The technical implementation logic described above provides a detailed explanation of steps S202 to S208.

[0056] In step S202 above, determining the first bitstream of the original video data includes: receiving the original video data; extracting features from the original video data to obtain semantic features and motion features corresponding to the original video data; quantizing and encoding the semantic features and motion features respectively to obtain semantic feature encoding units corresponding to the semantic features and motion feature encoding units corresponding to the motion features; and fusing the semantic feature encoding units and motion feature encoding units to obtain the first bitstream corresponding to the original video data.

[0057] In this embodiment, the generative tokenization encoding module is the core unit of the encoding end. Based on AIFlow theory, it abandons the traditional pixel-level compression method and adopts a feature extraction + dual tokenization encoding approach to achieve efficient video encoding and compression. The specific process analysis is as follows:

[0058] First, the system receives raw video sequences (raw video data) from the front-end acquisition device. It should be noted that this raw video sequence supports multiple resolutions and frame rate formats, including standard definition, high definition, and ultra-high definition, and is compatible with video source inputs from various application scenarios, including but not limited to mobile terminals, vehicle-mounted terminals, emergency communication equipment, and offshore monitoring systems, meeting video requirements in diverse terminal environments.

[0059] Secondly, spatiotemporal analysis was performed on the original video sequence to separate the following two key high-level features:

[0060] (1) Semantic features: used to characterize the static content information of the video, including but not limited to scene type (such as sea, mountain, indoor), main object (such as people, vehicles, ships), environmental structure (such as buildings, terrain, lighting conditions), etc.

[0061] (2) Motion characteristics: used to characterize the dynamic temporal information of the video, including but not limited to the target motion trajectory, displacement amplitude, speed change, inter-frame dynamic frequency, etc., covering rigid body and non-rigid body motion modes.

[0062] Finally, the extracted semantic and motion features are mapped to a discrete feature space to generate structured semantic tokens (i.e., the semantic feature encoding units mentioned above) and motion tokens (i.e., the motion feature encoding units mentioned above). These are then encapsulated into a core bitstream (i.e., the first bitstream mentioned above) using a preset protocol, achieving a paradigm shift from pixel-level compression to semantic-level encoding. This core bitstream can be directly fed into the packet loss-resistant enhanced bitstream construction module for processing without additional encoding conversion costs. This core bitstream eliminates redundant pixel information, significantly reducing its size compared to traditional encoded bitstreams, and can efficiently adapt to the transmission needs of weak networks and low bandwidth.

[0063] In the above process, by constructing a core bitstream that contains only semantic and motion feature encoding and does not carry any pixel information, the encoding end can achieve feature expression at an extremely low bit rate, providing accurate feature guidance for subsequent hierarchical completion and reconstruction based on packet loss.

[0064] In step S204 above, determining the auxiliary information segment corresponding to the original video data includes: determining the video parameter information of the original video data, wherein the video parameter information includes at least one of the following: the resolution, frame rate, and scene type identifier of the original video data; generating stable representation information based on the first bitstream, wherein the vector dimension of the stable representation information is lower than the vector dimension of the first bitstream; determining the auxiliary information segment corresponding to the original video data based on the video parameter information and the stable representation information, wherein the auxiliary information segment carries a unique identifier bit.

[0065] In this embodiment, the packet loss-resistant enhanced bitstream construction module is another core unit of the encoding end. It adds lightweight auxiliary information to the core bitstream to construct an enhanced transmission bitstream with packet loss robustness (i.e., the aforementioned second bitstream). Without increasing bandwidth overhead, it provides parsable and reusable auxiliary reconstruction information to the decoding end, supporting stable decoding under high packet loss conditions. The specific process is as follows:

[0066] First, it receives the core bitstream containing semantic tokens and motion tokens output by the self-generating tokenization encoding module.

[0067] Secondly, auxiliary information segments are embedded on the core bitstream, including the following two types of lightweight auxiliary information:

[0068] (1) Video parameter information, including basic video parameters such as video resolution, frame rate, encoding timestamp and scene type marker, which are used by the decoding end for context awareness and model parameter adaptation;

[0069] (2) Stable representation (seed) information, based on the low-dimensional feature vector generated by the core bitstream, is used to represent the prior and motion consistency patterns of the global structure of the video, and serves as the initial reference benchmark for the decoding end to build a multi-level stable representation system. Its vector dimension is significantly smaller than that of the core bitstream.

[0070] It should be noted that this auxiliary information segment is extremely small, accounting for less than 5% of the core bitstream, and will not significantly increase the overall bitstream volume. It consumes no additional bandwidth and meets the bandwidth constraints in weak network scenarios.

[0071] Subsequently, a structured encapsulation format of "core bitstream + auxiliary information segment" is adopted to generate a packet loss-resistant enhanced bitstream (i.e., the second bitstream mentioned above). The auxiliary information segment is placed at a fixed position at the beginning or end of the bitstream and is supplemented with a unique identifier field to achieve rapid identification and separation.

[0072] Finally, the encapsulated second bitstream is transmitted through a weak network channel, completing the entire processing flow at the encoding end.

[0073] In the above process, the packet loss-resistant enhanced bitstream maintains the original low bandwidth characteristics while providing reliable semantic clues for the decoding end, providing necessary prior support for subsequent multi-level stable representation reconstruction and diffusion model completion. Its lightweight design does not increase the transmission burden and adapts to the core requirements of weak network environments.

[0074] In this embodiment, the generative completion decoding module is the core unit of the decoding end, and its completion and reconstruction process is as follows: Figure 4As shown, the process mainly consists of the following three stages: First, the anti-packet loss enhanced bitstream received from the weak network channel is parsed to separate the core bitstream from the auxiliary information segment, and packet loss is detected. The lost semantic tokens and motion tokens and their packet loss degree are located through a verification mechanism. Second, for scenarios with a packet loss rate ≤10%, a multi-dimensional stable representation system containing scene structure, dynamic continuity, and visual priors is constructed using stable representation information in the auxiliary information. The missing features are then constrained by stable representation and reconstructed with temporal consistency. For the extreme case of a packet loss rate >10%, a pre-trained generative diffusion model is further introduced. Combined with the recovered features and scene priors, the semantic details and texture content of the missing frames are generated to achieve extreme completion.

[0075] Optionally, the bitstream parsing and packet loss detection include: separating the second bitstream after transmission through the channel based on the identifier bit to obtain the separated first bitstream and auxiliary information segment; performing integrity verification on the first bitstream obtained from the separation of the second bitstream to obtain the verification result; if the verification result indicates that the verification passed, directly outputting the video frame corresponding to the original video data based on the video parameter information in the auxiliary information segment; if the verification result indicates that the verification failed, determining the packet loss rate and missing characteristics of the second bitstream.

[0076] Specifically, based on specific identifier bits in the auxiliary information segments of the packet loss-resistant enhanced bitstream, the core bitstream (semantic tokens and motion tokens) and auxiliary information segments (including resolution, frame rate, scene type markings, and stable representation seeds) can be separated. Subsequently, bitstream verification algorithms such as sequence number verification, redundancy check codes, or timestamp consistency analysis are used to perform frame-by-frame integrity checks on the core bitstream (i.e., the first bitstream after separation), accurately locating the positions of lost or damaged tokens (i.e., missing features), quantifying the packet loss rate and distribution pattern, and providing adaptive selection for subsequent completion strategies. Quantization basis: When the verification passes, the basic video image is output directly based on the video parameter information (such as resolution, frame rate, chroma sampling format, quantization parameters, etc.) in the auxiliary information segment. At this stage, no reconstruction algorithm needs to be called. Only the static parameters preset by the encoder and the default generation template cached by the decoder are used to quickly restore the low-resolution but semantically coherent visual output. When the verification fails, the packet loss rate calculation and missing feature localization are started, thereby avoiding the accidental triggering of high-energy-consuming multi-dimensional stable representation system or generative diffusion model reconstruction in the absence of packet loss, and significantly reducing the computing power consumption and response latency of the decoder.

[0077] In the above process, through a hierarchical response mechanism of verification, decision-making, and reconstruction, an intelligent decoding strategy of direct output without loss and reconstruction only when loss occurs is achieved in a weak network and high packet loss environment. This not only ensures the availability of the picture under extreme network conditions, but also achieves zero-latency and efficient output under ideal transmission conditions.

[0078] In step S206 above, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system, including: when the packet loss rate is less than or equal to a preset threshold, constructing a multidimensional stable representation system containing scene structure representation, dynamic continuity representation and visual prior representation based on the stable representation information in the auxiliary information segment; and reconstructing the missing features in the second bitstream through the multidimensional stable representation system.

[0079] Optionally, the missing features in the second bitstream are reconstructed using a multidimensional stable representation system, including: determining the original location of the missing features based on scene structure constraints in the multidimensional stable representation system; determining the motion trajectory of the missing features based on dynamic continuity constraints in the multidimensional stable representation system; and determining the local texture features of the second bitstream based on visual prior constraints in the multidimensional stable representation system; determining the temporal correlation of the second bitstream, wherein the temporal correlation is used to reflect the temporal consistency between semantic features and motion features in the second bitstream; and reconstructing the missing features in the second bitstream based on the original location, motion trajectory, local texture features, and temporal correlation.

[0080] In this embodiment, for mild to moderate packet loss scenarios with a packet loss rate not exceeding 10% (preset threshold), a feature reconstruction mechanism based on multi-dimensional prior constraints is prioritized to achieve efficient and stable video content recovery with low computational overhead. The specific process analysis is as follows:

[0081] First, using the stable representation information in the auxiliary information segment, a multidimensional stable representation system is initialized, which consists of the following three types of structured priors:

[0082] (1) Scene structure representation: Static spatial topological constraints are constructed based on video basic parameters (such as scene type labels) and global context features to maintain the rationality of object position, background structure and visual hierarchy;

[0083] (2) Dynamic continuity representation: smoothness prior constraint based on motion token temporal modeling is used to constrain the physical continuity and consistency of the target motion trajectory and suppress unnatural jumps or jitters;

[0084] (3) Visual prior representation: Visual feature prior constraints are constructed based on the pre-trained visual semantic distribution model to ensure the consistency of texture statistical characteristics and edges of video images, and to ensure that the reconstructed region conforms to the distribution of the real scene at the visual perception level.

[0085] Secondly, the incomplete bitstream after packet loss detection is input into a multidimensional stable representation system. The missing features are subjected to stable representation constraints and temporal consistency modeling to recover the complete semantic token and motion token, thus obtaining the first reconstructed bitstream.

[0086] Optionally, the implementation of stable representation constraints and temporal consistency modeling is as follows: First, based on scene structure constraints, the spatial location of missing features is semantically located to ensure that their layout in the image conforms to the scene topology logic; second, the motion trajectory of missing frames is inferred by combining dynamic continuity constraints, and the temporal displacement path of the target is recovered by using optical flow smoothness and kinematic priors; at the same time, based on visual prior constraints, the statistical distribution and edge structure of local textures are modeled to fill in the missing content of texture details; then, the temporal correlation between semantic features and motion features in the video sequence is analyzed to ensure that the two maintain strong consistency in the temporal domain and avoid misalignment caused by the disconnect between semantics and motion; finally, by integrating the four-dimensional constraints of original position, motion trajectory, local texture and temporal correlation, the missing features are jointly optimized to achieve high-fidelity, physically reasonable and perceptually natural feature reconstruction.

[0087] Finally, based on the first reconstructed bitstream after completion, a stable video image is output, ensuring that the video maintains scene integrity in structure, achieves smooth transition in dynamics, and has no obvious blur, block effect or distortion visually, ensuring that the subjectively acceptable continuous playback quality is achieved under medium and low packet loss rates.

[0088] The above process aims to guide feature completion with structured priors, avoid blind generative interpolation, thereby improving the accuracy of reconstruction results and providing a high-quality, low-error initial reconstruction bitstream for subsequent diffusion model completion.

[0089] In step S208 above, the missing features in the second bitstream are reconstructed through a generative diffusion model, including: when the packet loss rate is greater than a preset threshold, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system to obtain a third bitstream, wherein the third bitstream also contains target missing features that the multidimensional stable representation system cannot recover; the third bitstream and auxiliary information segments are processed through a generative diffusion model to reconstruct the target missing features in the third bitstream.

[0090] Optionally, the third bitstream and auxiliary information segment are processed by a generative diffusion model to reconstruct the target missing features in the third bitstream, including: generating a fusion feature vector based on video information parameters and scene type identifiers in the third bitstream and auxiliary information segment; determining the semantic constraints of the auxiliary information segment and the spatiotemporal constraints of the original video data, and processing the fusion feature vector based on the semantic constraints and spatiotemporal constraints through the noise prediction network in the generative diffusion model to obtain a missing feature compensation vector; and reconstructing the target missing features based on the missing feature compensation vector.

[0091] In this embodiment, for severe packet loss scenarios with a packet loss rate exceeding 10%, an extreme completion mechanism based on a generative diffusion model is activated as the ultimate guarantee against packet loss reconstruction, enabling semantic-level restoration and detail reconstruction of visual content even under conditions of severe token loss. The specific process analysis is as follows:

[0092] First, the reconstructed bitstream sequence (i.e., the third bitstream mentioned above) after being represented by a multidimensional stable characterization is used as the input to the generative model.

[0093] It should be noted that, in scenarios with high packet loss rates, although the third bitstream has achieved preliminary semantic completion in terms of scene structure, dynamic continuity and visual priors, it still suffers from irrecoverable target loss features due to excessive information loss, broken temporal correlations or complete loss of high-frequency details. For example, it may manifest as semantic ambiguity in local areas, texture collapse or discontinuous motion trajectories.

[0094] Secondly, a generative diffusion model, combined with video information parameters and scene type identifiers from the auxiliary information segment, performs generative redrawing and detail completion on the missing target features, achieving feature expansion and detail reconstruction to obtain the second reconstructed bitstream. The specific implementation is as follows: the low-dimensional feature representation in the third bitstream is cross-modal aligned and fused with the structured metadata in the auxiliary information segment. A unified fused feature vector is generated through a feature embedding network, which simultaneously encodes the local state, global scene semantics, and spatiotemporal context of the reconstructed content. Subsequently, in the noise prediction network of the diffusion model, a conditional guidance field is constructed based on the semantic constraints extracted from the auxiliary information segment (such as the target appearance distribution corresponding to scene categories) and the spatiotemporal constraints of the original video data (such as inter-frame motion consistency and object topological continuity), serving as a priori control signal for the denoising process; noise prediction... The network iteratively performs denoising operations across multiple time steps. Through self-attention mechanisms and cross-frame correlation modeling, it predicts and eliminates noise components in the fused feature vector layer by layer, outputting a missing feature compensation vector that is highly consistent with the distribution of the real video. This compensation vector accurately locates the semantic increment and texture reconstruction direction of the missing region in the feature space, and its dimension is aligned with the feature space of the third bitstream to ensure seamless superposition. The missing feature compensation vector is fused with the third bitstream pixel by pixel or feature by feature addition to complete the generative reconstruction of the target missing features, outputting a complete and high-fidelity video feature sequence, which is the second reconstructed bitstream.

[0095] Finally, based on the completed second reconstructed bitstream, a stable video image is output, with excellent subjective visual quality. Even under extreme packet loss conditions, the understandability and viewing continuity of the image can still be maintained.

[0096] In the above process, it does not rely on the direct redundancy of the original pixel data at all. Instead, it achieves a qualitative recovery from incomplete structure to complete perception through semantic prior and spatiotemporal modeling-driven generative reasoning. This significantly improves the visual continuity and detail integrity in heavy packet loss scenarios and provides a high-quality reconstruction guarantee for video communication in extreme communication environments.

[0097] In the embodiments of this application, the generative AI model system involved, including the video semantic and motion feature extraction network, the multidimensional stable representation system and the generative diffusion model, all adopt an end-to-end joint training framework for pre-training. Its core strategy is "scene-driven data construction + multi-dimensional loss constraint optimization objective" to ensure that the model has robust feature completion and video reconstruction capabilities in real weak network and high packet loss environments.

[0098] First, in the dataset construction phase, a pair training dataset with high fidelity, multiple scenarios, and multiple packet loss levels was constructed. This dataset consists of high packet loss token streams (or historical high packet loss token streams) simulated under weak network conditions and corresponding original lossless high-quality video sequences, covering various weak network scenarios such as ocean, mountain, emergency rescue, and polar scientific expeditions. It also includes data samples with different packet loss levels of 1%-20%, which can effectively ensure the scenario adaptability of the model.

[0099] Secondly, at the model input level, the damaged token bitstream is used as the main input, while the basic video parameters and stable representation information are used as auxiliary inputs. Through a learnable embedding layer, they are uniformly mapped to a shared feature space, realizing the semantic alignment and fusion of multi-source heterogeneous information.

[0100] In terms of training objective design, a multi-task joint loss function was adopted, consisting of the following three key losses: (1) adversarial loss, which uses a discriminator network to distinguish the distribution differences between the reconstructed video and the real video, forcing the generated results to tend towards realism in terms of pixel statistics and texture details, thereby improving visual realism; (2) perceptual loss, which extracts high-level semantic features, minimizes the distance between the reconstructed frame and the real frame in the semantic feature space, strengthens subjective visual quality, and avoids excessive smoothing or structural distortion; (3) temporal consistency loss, which uses optical flow estimation or inter-frame feature alignment modules to constrain the continuity of motion trajectories and semantic changes between adjacent frames, suppressing jitter, jumps, or semantic breaks caused by packet loss. The three are weighted and fused to form a comprehensive optimization objective, achieving synergistic optimization of realism, perceptual quality, and dynamic coherence.

[0101] Finally, through a large-scale distributed training architecture, iterative optimization is performed on massive paired samples. Adaptive learning rate scheduling and gradient pruning strategies are adopted to continuously improve the model's generalization ability and reconstruction accuracy for different packet loss modes. After training convergence, this generative AI model can accurately learn the non-linear mapping relationship from "partially missing token bitstream + auxiliary semantic labels" to "complete high-quality video". It has the ability to perform hierarchical reconstruction of mild, moderate and severe packet loss scenarios, providing highly reliable AI kernel support for the entire encoding-transmission-decoding chain in actual deployment.

[0102] In this embodiment, by integrating AI Flow theory, feature-based encoding centered on semantics and motion tokens is achieved, abandoning traditional pixel-level redundant transmission and constructing a lightweight enhanced bitstream architecture without increasing bandwidth or latency. Simultaneously, a dual anti-packet-loss decoding mechanism of "multi-dimensional stable representation constraint reconstruction + generative diffusion model limit completion" is proposed, enabling the output of coherent, clear, and semantically consistent high-quality video even in scenarios with packet loss rates exceeding 10%, breaking through the quality limit of 1% packet loss rate in traditional encoding. This system balances high robustness with low resource consumption, seamlessly adapting to typical weak network scenarios such as ocean communication, mountain rescue, and polar scientific expeditions, and is deeply compatible with future communication architectures such as 6G networks and vehicle-mounted edge terminals, providing a highly reliable, scalable technical foundation for next-generation intelligent video communication.

[0103] Figure 5 This is a flowchart of a generative video coding method according to an embodiment of this application, such as... Figure 5 As shown, this method is applied to the encoding end and includes the following steps:

[0104] Step S502: Determine the first bitstream of the original video data, wherein the first bitstream is the core bitstream containing the semantic features and motion features of the original video data.

[0105] Step S504: Determine the auxiliary information segment corresponding to the original video data, and generate a second bitstream based on the first bitstream and the auxiliary information segment. The auxiliary information segment includes at least the video parameter information and stable characterization information of the original video data. The stable characterization information is a structured prior feature used to constrain the reconstruction process at the decoding end.

[0106] Step S506: The second bitstream is transmitted to the decoding end for processing.

[0107] Through steps S502 to S506 above, video is compressed into a core bitstream with both semantic and motion features under extremely low bandwidth overhead (auxiliary information ratio ≤ 5%), and structured prior constraint information is embedded, laying the foundation for the intelligent reconstruction capability of the decoding end and effectively getting rid of the dependence of traditional pixel-level encoding on the complete bitstream.

[0108] Optionally, determining the first bitstream of the original video data includes: receiving the original video data; extracting features from the original video data to obtain semantic features and motion features corresponding to the original video data; quantizing and encoding the semantic features and motion features respectively to obtain semantic feature encoding units corresponding to the semantic features and motion feature encoding units corresponding to the motion features; and fusing the semantic feature encoding units and motion feature encoding units to obtain the first bitstream corresponding to the original video data.

[0109] Figure 6 This is a flowchart of a generative video decoding method according to an embodiment of this application, such as... Figure 6 As shown, this method is applied to the decoding end and includes the following steps:

[0110] Step S602: Receive the second bitstream transmitted from the encoding end and determine the packet loss rate of the second bitstream. The second bitstream is generated based on the first bitstream and auxiliary information segments. The first bitstream is a core bitstream containing semantic features and motion features of the original video data. The auxiliary information segments include at least video parameter information and stability representation information of the original video data. The stability representation information is a structured prior feature used to constrain the reconstruction process at the decoding end.

[0111] Step S604: When the packet loss rate is less than or equal to a preset threshold, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system to obtain a first reconstructed bitstream, and the video image is output based on the first reconstructed bitstream. The multidimensional stable representation system is used to represent the multidimensional representation constraints corresponding to the reconstructed missing features.

[0112] Step S606: When the packet loss rate is greater than a preset threshold, the missing features in the second bitstream are reconstructed using a generative diffusion model to obtain the second reconstructed bitstream, and the video image is output based on the second reconstructed bitstream.

[0113] Through the above steps S602 to S606, intelligent hierarchical reconstruction without retransmission or additional bandwidth consumption is achieved. In the case of mild packet loss, features are accurately recovered through structured prior constraints. In the case of severe packet loss, the image is fully restored through generative diffusion model. This breaks through the quality limit of traditional video coding technology when the packet loss rate exceeds 1%, ensuring video continuity and subjective clarity in different packet loss scenarios.

[0114] Optionally, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system, including: when the packet loss rate is less than or equal to a preset threshold, constructing a multidimensional stable representation system containing scene structure representation, dynamic continuity representation, and visual prior representation based on the stable representation information in the auxiliary information segment; determining the original position of the missing features based on the scene structure constraints in the multidimensional stable representation system, determining the motion trajectory of the missing features based on the dynamic continuity constraints in the multidimensional stable representation system, and determining the local texture features of the second bitstream based on the visual prior constraints in the multidimensional stable representation system; determining the temporal correlation of the second bitstream, wherein the temporal correlation is used to reflect the temporal consistency relationship between semantic features and motion features in the second bitstream; and reconstructing the missing features in the second bitstream based on the original position, motion trajectory, local texture features, and temporal correlation.

[0115] Optionally, the generative diffusion model reconstructs the missing features in the second bitstream, including: when the packet loss rate is greater than a preset threshold, reconstructing the missing features in the second bitstream through a multidimensional stable representation system to obtain a third bitstream, wherein the third bitstream still contains target missing features that the multidimensional stable representation system cannot recover; generating a fusion feature vector based on the third bitstream, video information parameters in the auxiliary information segment, and scene type identifier; determining the semantic constraints of the auxiliary information segment and the spatiotemporal constraints of the original video data, and processing the fusion feature vector based on the semantic constraints and spatiotemporal constraints through the noise prediction network in the generative diffusion model to obtain a missing feature compensation vector; and reconstructing the target missing features based on the missing feature compensation vector.

[0116] According to embodiments of this application, a generative video encoding and decoding apparatus is provided. It should be noted that the generative video encoding and decoding apparatus of this application can be used to execute the generative video encoding and decoding method provided in the embodiments of this application. The generative video encoding and decoding apparatus provided in the embodiments of this application will be described below.

[0117] Figure 7 This is a structural diagram of a generative video encoding and decoding apparatus provided according to an embodiment of this application. Figure 7 As shown, the device includes:

[0118] The determination module 70 is used to determine the first bitstream of the original video data, wherein the first bitstream is the core bitstream containing the semantic features and motion features of the original video data;

[0119] The generation module 72 is used to determine the auxiliary information segment corresponding to the original video data and generate a second bitstream based on the first bitstream and the auxiliary information segment. The auxiliary information segment includes at least the video parameter information and stable characterization information of the original video data. The stable characterization information is a structured prior feature used to constrain the reconstruction process at the decoding end.

[0120] The first reconstruction module 74 is used to determine the packet loss rate of the second bitstream after transmission through the channel. When the packet loss rate is less than or equal to a preset threshold, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system to obtain the first reconstructed bitstream, and the video image is output based on the first reconstructed bitstream. The multidimensional stable representation system is used to represent the multidimensional representation constraints corresponding to the reconstructed missing features.

[0121] The second reconstruction module 76 is used to reconstruct the missing features in the second bitstream using a generative diffusion model when the packet loss rate is greater than a preset threshold, to obtain the second reconstructed bitstream, and to output video images based on the second reconstructed bitstream.

[0122] Through the determination module, generation module, first reconstruction module and second reconstruction module in the above-mentioned generative video encoding and decoding device, the goal of outputting visually coherent and natural video images based on the dual reconstruction path mechanism is achieved. This realizes highly robust video reconstruction under the premise of no bandwidth increment and no increase in transmission latency, significantly improving the viewing experience and system stability under extreme network conditions. In turn, it solves the technical problem of video quality collapse or continuous interruption caused by the lack of a stable reconstruction mechanism for bitstream loss in video encoding technology in weak network environments.

[0123] In the generative video encoding and decoding apparatus provided in this application embodiment, the determining module is further configured to receive raw video data; extract features from the raw video data to obtain semantic features and motion features corresponding to the raw video data; quantize and encode the semantic features and motion features respectively to obtain a semantic feature encoding unit corresponding to the semantic features and a motion feature encoding unit corresponding to the motion features; and fuse the semantic feature encoding unit and the motion feature encoding unit to obtain a first bitstream corresponding to the raw video data.

[0124] In the generative video encoding and decoding apparatus provided in this application embodiment, the generation module is further configured to determine video parameter information of the original video data, wherein the video parameter information includes at least one of the following: resolution, frame rate and scene type identifier of the original video data; generate stable representation information based on the first bitstream, wherein the vector dimension of the stable representation information is lower than the vector dimension of the first bitstream; determine an auxiliary information segment corresponding to the original video data based on the video parameter information and the stable representation information, wherein the auxiliary information segment carries a unique identifier bit.

[0125] In the generative video encoding and decoding apparatus provided in this application embodiment, the first reconstruction module is further configured to separate the second bitstream after transmission through the channel based on the identifier bit to obtain the separated first bitstream and auxiliary information segment; perform integrity verification on the first bitstream obtained by separating the second bitstream to obtain the verification result; if the verification result indicates that the verification passed, directly output the video image corresponding to the original video data based on the video parameter information in the auxiliary information segment; if the verification result indicates that the verification failed, determine the packet loss rate and missing characteristics of the second bitstream.

[0126] In the generative video encoding and decoding apparatus provided in this application embodiment, the first reconstruction module is further configured to construct a multi-dimensional stable representation system including scene structure representation, dynamic continuity representation and visual prior representation based on the stable representation information in the auxiliary information segment when the packet loss rate is less than or equal to a preset threshold; and reconstruct the missing features in the second bitstream through the multi-dimensional stable representation system.

[0127] In the generative video encoding and decoding apparatus provided in this application embodiment, the first reconstruction module is further configured to determine the original position of the missing feature based on the scene structure constraints in the multidimensional stable representation system, determine the motion trajectory of the missing feature based on the dynamic continuity constraints in the multidimensional stable representation system, and determine the local texture features of the second bitstream based on the visual prior constraints in the multidimensional stable representation system; determine the temporal correlation of the second bitstream, wherein the temporal correlation is used to reflect the temporal consistency relationship between the semantic features and motion features in the second bitstream; and reconstruct the missing feature in the second bitstream based on the original position, motion trajectory, local texture features and temporal correlation.

[0128] In the generative video encoding and decoding apparatus provided in this application embodiment, the second reconstruction module is further used to reconstruct the missing features in the second bitstream through a multidimensional stable representation system when the packet loss rate is greater than a preset threshold, thereby obtaining a third bitstream, wherein the third bitstream also contains target missing features that cannot be recovered by the multidimensional stable representation system; and to reconstruct the target missing features in the third bitstream by processing the third bitstream and auxiliary information segments through a generative diffusion model.

[0129] In the generative video encoding and decoding apparatus provided in this application embodiment, the second reconstruction module is further used to generate a fusion feature vector based on the third bitstream, video information parameters in the auxiliary information segment, and scene type identifier; determine the semantic constraints of the auxiliary information segment and the spatiotemporal constraints of the original video data, and process the fusion feature vector based on the semantic constraints and spatiotemporal constraints through the noise prediction network in the generative diffusion model to obtain a missing feature compensation vector; and reconstruct the target missing features based on the missing feature compensation vector.

[0130] This application also provides an electronic device, including: a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above-described generative video encoding / decoding, generative video encoding, or generative video decoding methods.

[0131] It should be noted that the aforementioned electronic equipment is used to perform Figure 2 The generative video encoding and decoding method shown Figure 5 The generative video coding method shown or Figure 6 The generative video decoding method shown above, therefore, the relevant explanations and descriptions in the generative video encoding and decoding, generative video encoding or generative video decoding methods also apply to this electronic device, and will not be repeated here.

[0132] This application also provides a non-volatile storage medium, which includes a stored computer program, wherein the device containing the non-volatile storage medium executes the above-described generative video encoding / decoding, generative video encoding, or generative video decoding methods by running the computer program.

[0133] It should be noted that the aforementioned non-volatile storage media is used for execution. Figure 2 The generative video encoding and decoding method shown Figure 5 The generative video coding method shown or Figure 6 The generative video decoding method shown above, therefore, the relevant explanations and descriptions in the generative video encoding and decoding, generative video encoding or generative video decoding methods also apply to this non-volatile storage medium, and will not be repeated here.

[0134] This application also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described generative video encoding / decoding, generative video encoding, or generative video decoding methods.

[0135] It should be noted that the above-mentioned computer program product is used to execute Figure 2 The generative video encoding and decoding method shown Figure 5 The generative video coding method shown or Figure 6 The generative video decoding method shown above, therefore, the relevant explanations and descriptions in the above generative video encoding and decoding, generative video encoding or generative video decoding methods also apply to this computer program product, and will not be repeated here.

[0136] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0137] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0138] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0139] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0140] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0141] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0142] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A generative video coding and decoding method, characterized in that, include: A first bitstream of the original video data is determined, wherein the first bitstream is a core bitstream containing the semantic features and motion features of the original video data; An auxiliary information segment corresponding to the original video data is determined, and a second bitstream is generated based on the first bitstream and the auxiliary information segment. The auxiliary information segment includes at least video parameter information and stability characterization information of the original video data. The stability characterization information is a structured prior feature used to constrain the reconstruction process at the decoding end. The packet loss rate of the second bitstream after transmission through the channel is determined. If the packet loss rate is less than or equal to a preset threshold, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system to obtain a first reconstructed bitstream. The video image is then output based on the first reconstructed bitstream. The multidimensional stable representation system is used to represent the multidimensional representation constraints corresponding to the reconstruction of the missing features. If the packet loss rate is greater than the preset threshold, the missing features in the second bitstream are reconstructed using a generative diffusion model to obtain a second reconstructed bitstream, and video images are output based on the second reconstructed bitstream.

2. The method according to claim 1, characterized in that, Determine the first bitstream of the raw video data, including: Receive raw video data; Feature extraction is performed on the original video data to obtain semantic features and motion features corresponding to the original video data; The semantic features and the motion features are quantized and encoded respectively to obtain semantic feature encoding units corresponding to the semantic features and motion feature encoding units corresponding to the motion features; By fusing the semantic feature encoding unit and the motion feature encoding unit, a first bitstream corresponding to the original video data is obtained.

3. The method according to claim 1, characterized in that, Determining the auxiliary information segment corresponding to the original video data includes: Determine the video parameter information of the original video data, wherein the video parameter information includes at least one of the following: the resolution, frame rate, and scene type identifier of the original video data; Stable representation information is generated based on the first bitstream, wherein the vector dimension of the stable representation information is lower than the vector dimension of the first bitstream; Based on the video parameter information and the stability characterization information, an auxiliary information segment corresponding to the original video data is determined, wherein the auxiliary information segment carries a unique identifier.

4. The method according to claim 3, characterized in that, Determining the packet loss rate of the second bitstream after transmission through the channel includes: Based on the aforementioned identifier, the second bitstream after transmission through the channel is separated to obtain the separated first bitstream and auxiliary information segment; The integrity of the first bitstream obtained by separating the second bitstream is verified, and the verification result is obtained. If the verification result indicates that the verification is successful, the video frame corresponding to the original video data is directly output based on the video parameter information in the auxiliary information segment. If the verification result indicates that the verification failed, the packet loss rate and missing features of the second bitstream are determined.

5. The method according to claim 4, characterized in that, The missing features in the second bitstream are reconstructed using a multidimensional stable representation system, including: When the packet loss rate is less than or equal to the preset threshold, a multi-dimensional stable representation system including scene structure representation, dynamic continuity representation and visual prior representation is constructed based on the stable representation information in the auxiliary information segment. The missing features in the second bitstream are reconstructed using the multidimensional stable representation system.

6. The method according to claim 5, characterized in that, Reconstructing the missing features in the second bitstream using the multidimensional stable representation system includes: The original position of the missing feature is determined based on the scene structure constraints in the multidimensional stable representation system; the motion trajectory of the missing feature is determined based on the dynamic continuity constraints in the multidimensional stable representation system; and the local texture features of the second bitstream are determined based on the visual prior constraints in the multidimensional stable representation system. Determine the temporal correlation of the second bitstream, wherein the temporal correlation is used to reflect the temporal consistency relationship between semantic features and motion features in the second bitstream; The missing features in the second bitstream are reconstructed based on the original position, the motion trajectory, the local texture features, and the temporal correlation.

7. The method according to claim 1, characterized in that, Reconstructing the missing features in the second bitstream using a generative diffusion model includes: When the packet loss rate is greater than the preset threshold, the missing features in the second bitstream are reconstructed through the multidimensional stable representation system to obtain the third bitstream, wherein the third bitstream also contains target missing features that the multidimensional stable representation system cannot recover. The third bitstream and the auxiliary information segment are processed by the generative diffusion model to reconstruct the target missing features in the third bitstream.

8. The method according to claim 7, characterized in that, The third bitstream and the auxiliary information segment are processed using the generative diffusion model to reconstruct the target missing features in the third bitstream, including: A fusion feature vector is generated based on the third bitstream, the video information parameters in the auxiliary information segment, and the scene type identifier; The semantic constraints of the auxiliary information segment and the spatiotemporal constraints of the original video data are determined, and the fused feature vector is processed by the noise prediction network in the generative diffusion model according to the semantic constraints and the spatiotemporal constraints to obtain the missing feature compensation vector. The target missing features are reconstructed based on the missing feature compensation vector.

9. A generative video coding method, characterized in that, include: A first bitstream of the original video data is determined, wherein the first bitstream is a core bitstream containing the semantic features and motion features of the original video data; An auxiliary information segment corresponding to the original video data is determined, and a second bitstream is generated based on the first bitstream and the auxiliary information segment. The auxiliary information segment includes at least video parameter information and stability characterization information of the original video data. The stability characterization information is a structured prior feature used to constrain the reconstruction process at the decoding end. The second bitstream is transmitted to the decoding end for processing.

10. The method according to claim 9, characterized in that, Determine the first bitstream of the raw video data, including: Receive raw video data; Feature extraction is performed on the original video data to obtain semantic features and motion features corresponding to the original video data; The semantic features and the motion features are quantized and encoded respectively to obtain semantic feature encoding units corresponding to the semantic features and motion feature encoding units corresponding to the motion features; By fusing the semantic feature encoding unit and the motion feature encoding unit, a first bitstream corresponding to the original video data is obtained.

11. A generative video decoding method, characterized in that, include: The second bitstream transmitted by the encoding end is received, and the packet loss rate of the second bitstream is determined. The second bitstream is generated based on the first bitstream and auxiliary information segments. The first bitstream is a core bitstream containing semantic features and motion features of the original video data. The auxiliary information segments include at least video parameter information and stability representation information of the original video data. The stability representation information is a structured prior feature used to constrain the reconstruction process of the decoding end. When the packet loss rate is less than or equal to a preset threshold, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system to obtain a first reconstructed bitstream, and video images are output based on the first reconstructed bitstream. The multidimensional stable representation system is used to represent the multidimensional representation constraints corresponding to the reconstruction of the missing features. If the packet loss rate is greater than the preset threshold, the missing features in the second bitstream are reconstructed using a generative diffusion model to obtain a second reconstructed bitstream, and video images are output based on the second reconstructed bitstream.

12. The method according to claim 11, characterized in that, The missing features in the second bitstream are reconstructed using a multidimensional stable representation system, including: When the packet loss rate is less than or equal to the preset threshold, a multi-dimensional stable representation system including scene structure representation, dynamic continuity representation and visual prior representation is constructed based on the stable representation information in the auxiliary information segment. The original position of the missing feature is determined based on the scene structure constraints in the multidimensional stable representation system; the motion trajectory of the missing feature is determined based on the dynamic continuity constraints in the multidimensional stable representation system; and the local texture features of the second bitstream are determined based on the visual prior constraints in the multidimensional stable representation system. Determine the temporal correlation of the second bitstream, wherein the temporal correlation is used to reflect the temporal consistency relationship between semantic features and motion features in the second bitstream; The missing features in the second bitstream are reconstructed based on the original position, the motion trajectory, the local texture features, and the temporal correlation.

13. The method according to claim 11, characterized in that, Generative diffusion models reconstruct missing features in the second bitstream, including: When the packet loss rate is greater than the preset threshold, the missing features in the second bitstream are reconstructed through the multidimensional stable representation system to obtain the third bitstream, wherein the third bitstream also contains target missing features that the multidimensional stable representation system cannot recover. A fusion feature vector is generated based on the third bitstream, the video information parameters in the auxiliary information segment, and the scene type identifier; The semantic constraints of the auxiliary information segment and the spatiotemporal constraints of the original video data are determined, and the fused feature vector is processed by the noise prediction network in the generative diffusion model according to the semantic constraints and the spatiotemporal constraints to obtain the missing feature compensation vector. The target missing features are reconstructed based on the missing feature compensation vector.

14. A generative video encoding and decoding apparatus, characterized in that, include: The determining module is used to determine the first bitstream of the original video data, wherein the first bitstream is a core bitstream containing the semantic features and motion features of the original video data; A generation module is used to determine the auxiliary information segment corresponding to the original video data, and generate a second bitstream based on the first bitstream and the auxiliary information segment. The auxiliary information segment includes at least video parameter information and stability characterization information of the original video data. The stability characterization information is a structured prior feature used to constrain the reconstruction process at the decoding end. The first reconstruction module is used to determine the packet loss rate of the second bitstream after transmission through the channel. When the packet loss rate is less than or equal to a preset threshold, the missing features in the second bitstream are reconstructed through a multidimensional stable representation system to obtain a first reconstructed bitstream, and video images are output based on the first reconstructed bitstream. The multidimensional stable representation system is used to represent the multidimensional representation constraints corresponding to the reconstruction of the missing features. The second reconstruction module is used to reconstruct the missing features in the second bitstream using a generative diffusion model when the packet loss rate is greater than the preset threshold, to obtain the second reconstructed bitstream, and to output video images based on the second reconstructed bitstream.

15. An electronic device, characterized in that, include: A memory and a processor, wherein the memory is used to store program instructions; The processor, connected to the memory, is used to execute the generative video encoding and decoding method according to any one of claims 1 to 8, the generative video encoding method according to any one of claims 9 to 10, or the generative video decoding method according to any one of claims 11 to 13.

16. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored computer program, wherein the device containing the non-volatile storage medium executes the generative video coding and decoding method according to any one of claims 1 to 8, the generative video coding method according to any one of claims 9 to 10, or the generative video decoding method according to any one of claims 11 to 13 by running the computer program.