Method, apparatus and electronic device for processing video stream
Patent Information
- Application Number
- CN202610560100.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-24
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-04-24
AI Technical Summary
[0005]本申请实施例提供了一种视频流的处理方法、装置及电子设备,以至少解决由于未考虑随机访问需求,进行随机访问时依赖前序视频片段解码,导致视频跳转延时较大的技术问题
[0021]在本申请实施例中,采用接收随机访问操作请求,其中,随机访问操作请求至少用于携带需要访问的原始视频流中的目标视频片段标识,目标视频片段标识所指示的视频片段包括:对原始视频流按照视频随机访问需求信息划分得到的多个视频片段中的片段,每个视频片段对应一个视频片段标识以及至少用于标识视频片段起始位置的入口点标识信息;依据随机访问操作请求确定目标令牌,其中,目标令牌通过对目标视频片段进行编码后得到,目标令牌携带有目标视频片段的目标入口点标识信息;依据目标令牌携带的目标入口点标识信息对目标令牌进行解码,得到目标视频片段的方式,通过接收随机访问操作请求后依据该请求确定目标令牌,而目标令牌是通过对目标视频片段进行编码后得到。目标令牌携带有目标视频片段的目标入口点标识信息,可以实现进行随机访问时不依赖前序视频片段解码,目标视频片段为对原始视频流按照视频随机访问需求信息划分得到的多个视频片段中目标视频片段标识指示的片段,进而解决了由于未考虑随机访问需求,进行随机访问时依赖前序视频片段解码,导致视频跳转延时较大的技术问题。
Smart Images

Figure CN122093619B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video stream processing, and more specifically, to a method, apparatus, and electronic device for processing video streams. Background Technology
[0002] To support user interactions such as quick navigation, channel switching, and dragging during video playback, related technologies generally employ a group of pictures (GOP)-based random access coding mechanism, using a hybrid coding architecture of keyframes (I-frames) and predicted frames (P-frames / B-frames) to achieve random access functionality. However, due to the large amount of keyframe data, complex coding, and frequent insertion affecting compression ratio, this method struggles to balance compression efficiency with access response speed.
[0003] Building upon this foundation, deep learning-based video compression technologies have gradually emerged. Some methods combine deep learning models to encode the original video stream using an encoder and reconstruct the content at the decoding end, thus improving compression efficiency. However, these methods employ the periodic segmentation strategy of GOPs, dividing video units into fixed frame numbers and treating each unit as an independent compression block. Under this structure, although each unit no longer depends on I-frames, the decoding end still needs to completely reconstruct the video from the beginning of each unit, lacking an independent, jumpable entry mechanism. Therefore, in random access scenarios, the following drawbacks exist: First, the GOP period is generally too long, and entry points are sparse, making it difficult to meet the needs of high-frequency interaction; second, the decoding process often relies on the temporal context of previous frames, making it impossible to achieve independent startup at any entry point, requiring buffering and warm-up, which can easily cause stuttering or reconstruction delays; third, the GOP length is often set as a single static value, without considering the random access needs of specific application scenarios (such as video-on-demand, live streaming, and broadcast television), ultimately making it difficult to balance response speed and bandwidth efficiency, resulting in significant latency for random video access.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides a video stream processing method, apparatus, and electronic device to at least solve the technical problem that video jump delay is large due to the reliance on decoding preceding video segments when performing random access without considering random access requirements.
[0006] According to one aspect of the embodiments of this application, a video stream processing method is provided, comprising: receiving a random access operation request, wherein the random access operation request is at least used to carry a target video segment identifier, the video segment indicated by the target video segment identifier includes: a segment from a plurality of video segments obtained by dividing the original video stream according to video random access requirement information, each video segment corresponding to a video segment identifier and at least entry point identifier information used to identify the start position of the video segment; determining a target token based on the random access operation request, wherein the target token is obtained by encoding the target video segment, and the target token carries the target entry point identifier information of the target video segment; and decoding the target token based on the target entry point identifier information carried by the target token to obtain the target video segment.
[0007] Optionally, the original video stream is divided into multiple video segments according to the video random access requirement information in the following manner: obtaining the video random access requirement information, wherein the video random access requirement information is used at least to indicate the allowed latency threshold when performing video random access in the target application scenario; determining the target video frame threshold from multiple video frame thresholds based on the video random access requirement information, wherein the multiple video frame thresholds correspond to different numbers of video frames; dividing the original video stream into segments according to the target video frame threshold to obtain multiple initial video segments; and determining the multiple initial video segments carrying entry point identification information as multiple video segments.
[0008] Optionally, any one of the multiple video segments can be determined as follows: The starting frame is obtained from the multiple video frames contained in the first initial video segment, wherein the first initial video segment is any one of the multiple initial video segments; visual features are extracted from the starting frame using a variational autoencoder to obtain a first high-dimensional feature map of the starting frame, and the first high-dimensional feature map is determined as the entry point of the first initial video segment; the position information corresponding to the starting frame is obtained, and the set consisting of the position information corresponding to the starting frame and the entry point of the first initial video segment is determined as the first entry point identifier information of the first initial video segment; the first initial video segment carrying the first entry point identifier information is determined as the video segment corresponding to the first initial video segment.
[0009] Optionally, the method further includes: encoding the target video segment to obtain a target token, including: extracting visual features from the initial target video segment using a variational autoencoder to obtain a second high-dimensional feature map corresponding to the initial target video segment, wherein the initial target video segment is the initial video segment corresponding to the target video segment; performing dimensionality reduction processing on the second high-dimensional feature map using a target encoding strategy to obtain a target learnable token, wherein the target encoding strategy is determined based on the number of video frames of the initial target video segment; and encapsulating the target learnable token with the target entry point identifier information of the target video segment to obtain the target token.
[0010] Optionally, the second high-dimensional feature map is subjected to dimensionality reduction processing using a target encoding strategy to obtain a target learnable token. This includes: when the number of video frames of the initial target video segment is greater than or equal to a first preset threshold, the second high-dimensional feature map is subjected to dimensionality reduction processing using a first encoding strategy to obtain a target learnable token. The first encoding strategy divides the target video segment into multiple sub-segments in sequence and performs dimensionality reduction processing on the multiple sub-segments sequentially based on the dependencies between the multiple sub-segments. When the number of video frames of the initial target video segment is less than the first preset threshold, the second encoding strategy is subjected to dimensionality reduction processing using a second encoding strategy to obtain a target learnable token. The second encoding strategy performs dimensionality reduction processing on the target video segment in one step.
[0011] Optionally, determining the target token based on the random access operation request includes: obtaining the target video segment identifier carried in the random access operation request; sending a target token request instruction to the transmission link based on the target video segment identifier; receiving a response message corresponding to the target token request instruction returned by the transmission link, wherein the response message carries at least the target token; and extracting the target token from the response message.
[0012] Optionally, the target token is decoded based on the target entry point identifier information carried by the target token to obtain the target video segment, including: parsing the target token to obtain the target entry point identifier information and the target learnable token carried by the target token; performing feature upscaling processing on the target learnable token to obtain the restored high-dimensional feature map; performing feature fusion processing on the restored high-dimensional feature map and the first high-dimensional feature map of the starting frame carried by the target entry point identifier information to obtain the fused feature map; performing noise reduction processing on the fused feature map to obtain the target fused feature map; and decoding the target fused feature map to obtain the target video segment.
[0013] According to another aspect of the embodiments of this application, a video stream processing method is also provided, comprising: receiving an original video stream; dividing the original video stream according to video random access request information to obtain multiple video segments, wherein each video segment corresponds to a video segment identifier and at least entry point identifier information for identifying the start position of the video segment; encoding the multiple video segments respectively to obtain multiple tokens; receiving a target token request instruction sent by a decoding end to a transmission link, wherein the target token request instruction is used to indicate the target video segment identifier carried in the random access operation request; determining a target token from the multiple tokens, wherein the target token is a token obtained by encoding the video segment indicated by the target video segment identifier from the multiple tokens; and sending a response message corresponding to the target token request instruction to the decoding end through the transmission link, wherein the response message carries at least the target token.
[0014] Optionally, the original video stream is divided according to video random access requirement information to obtain multiple video segments, including: obtaining video random access requirement information, wherein the video random access requirement information is used at least to indicate the allowed latency threshold when performing video random access in the target application scenario; determining a target video frame threshold from multiple video frame thresholds based on the video random access requirement information, wherein the multiple video frame thresholds correspond to different numbers of video frames; dividing the original video stream into segments according to the target video frame threshold to obtain multiple initial video segments; and determining the multiple initial video segments carrying entry point identification information as multiple video segments.
[0015] Optionally, any one of the multiple video segments can be determined as follows: The starting frame is obtained from the multiple video frames contained in the first initial video segment, wherein the first initial video segment is any one of the multiple initial video segments; visual features are extracted from the starting frame using a variational autoencoder to obtain a first high-dimensional feature map of the starting frame, and the first high-dimensional feature map is determined as the entry point of the first initial video segment; the position information corresponding to the starting frame is obtained, and the set consisting of the position information corresponding to the starting frame and the entry point of the first initial video segment is determined as the first entry point identifier information of the first initial video segment; the first initial video segment carrying the first entry point identifier information is determined as the video segment corresponding to the first initial video segment.
[0016] According to another aspect of the embodiments of this application, a video stream processing method is also provided, comprising: receiving a random access operation request, wherein the random access operation request is at least used to carry a target video segment identifier in the original video stream to be accessed, the video segment indicated by the target video segment identifier includes: a segment from a plurality of video segments obtained by dividing the original video stream according to video random access requirement information, each video segment corresponding to a video segment identifier and at least entry point identifier information used to identify the start position of the video segment; obtaining the target video segment identifier carried in the random access operation request, and sending a request instruction for requesting a target token to the transmission link according to the target video segment identifier, wherein the target token is obtained by encoding the target video segment indicated by the target video segment identifier, and the target token carries the target entry point identifier information of the target video segment; receiving a response message corresponding to the request instruction returned by the transmission link, and extracting the target token from the response message; and decoding the target token according to the target entry point identifier information carried by the target token to obtain the target video segment.
[0017] According to another aspect of the embodiments of this application, a video stream processing apparatus is also provided, comprising: a receiving module, configured to receive a random access operation request, wherein the random access operation request is configured to carry at least a target video segment identifier in the original video stream to be accessed, the video segment indicated by the target video segment identifier including: a segment from a plurality of video segments obtained by dividing the original video stream according to video random access requirement information, each video segment corresponding to a video segment identifier and at least entry point identifier information for identifying the start position of the video segment; a determining module, configured to determine a target token based on the random access operation request, wherein the target token is obtained by encoding the target video segment, and the target token carries the target entry point identifier information of the target video segment; and a decoding module, configured to decode the target token based on the target entry point identifier information carried by the target token to obtain the target video segment.
[0018] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, wherein a program is stored in the non-volatile storage medium, and the program controls the device where the non-volatile storage medium is located to execute the above-mentioned video stream processing method when it runs.
[0019] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the program executes the above-described video stream processing method when it runs.
[0020] According to another aspect of the embodiments of this application, a computer program product is also provided, including computer instructions, which, when executed by a processor, implement the above-described video stream processing method.
[0021] In this embodiment, a random access operation request is received, wherein the random access operation request carries at least a target video segment identifier in the original video stream to be accessed. The video segment indicated by the target video segment identifier includes: a segment from multiple video segments obtained by dividing the original video stream according to the video random access requirement information, each video segment corresponding to a video segment identifier and at least entry point identifier information used to identify the start position of the video segment; a target token is determined according to the random access operation request, wherein the target token is obtained by encoding the target video segment, and the target token carries the target entry point identifier information of the target video segment; the target token is decoded according to the target entry point identifier information carried by the target token to obtain the target video segment. The target token is determined according to the random access operation request after receiving it, and the target token is obtained by encoding the target video segment. The target token carries the target entry point identification information of the target video segment, which enables random access without relying on the decoding of preceding video segments. The target video segment is the segment indicated by the target video segment identifier among multiple video segments obtained by dividing the original video stream according to the random access requirement information. This solves the technical problem of large video jump delay caused by relying on the decoding of preceding video segments when performing random access due to the lack of consideration for random access requirements. Attached Figure Description
[0022] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0023] Figure 1 This is a hardware structure block diagram of a computer terminal for implementing a video stream processing method according to an embodiment of this application;
[0024] Figure 2 This is a flowchart of a first video stream processing method provided according to an embodiment of this application;
[0025] Figure 3 This is a schematic diagram of the architecture of a video stream processing system provided according to an embodiment of this application;
[0026] Figure 4 This is a flowchart of a second video stream processing method provided according to an embodiment of this application;
[0027] Figure 5 This is a flowchart of a third video stream processing method provided according to an embodiment of this application;
[0028] Figure 6 This is a flowchart of a fourth video stream processing method provided according to an embodiment of this application;
[0029] Figure 7 This is a schematic diagram of a video stream processing device provided according to an embodiment of this application. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0031] The information collected in this application embodiment is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant regions, and necessary confidentiality measures have been taken. It does not violate public order and good morals, and provides corresponding operation entry points for users to choose to authorize or reject the automated decision results. If the user chooses to reject, the process will proceed to the expert decision-making process.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained below:
[0034] Diffusion Transformer (DiT): DiT is used for noise reduction during the video compression and decoding stage to improve video restoration quality.
[0035] Group of Pictures (GOP): A GOP is the smallest unit for implementing random access. The starting frame (I-frame) of a GOP provides a complete image reference, and subsequent frames are encoded through inter-frame prediction (based on the previous frame or frames before and after), thereby achieving high compression efficiency.
[0036] Variational Autoencoder (VAE): A VAE is used for feature extraction, initial compression, and video reconstruction during the decoding stage of video clips. A VAE is a generative model consisting of an encoder and a decoder. The encoder maps video clips to distribution parameters in the latent space, outputting a high-dimensional feature map as input for subsequent compression. The decoder reconstructs the original video frames from the denoised and upsampled latent variables.
[0037] A learnable token is a compact feature representation of a video clip obtained after processing by a VAE and a compression downsampling module. It is core data in the video streaming process.
[0038] Random Access Encoding: An encoding method that supports random jumps and fast start-up based on GOP cycles. Its core is to introduce accessible points during the encoding process to balance compression efficiency and accessibility.
[0039] Encoder / Decoder: The encoder is responsible for segmenting the original video into video segments and generating tokens, while the decoder is responsible for restoring and decoding the tokens to output the target video.
[0040] In related technologies, the original video stream is encoded by an encoder, and the content is reconstructed at the decoding end, which improves compression efficiency. However, this type of method uses the periodic segmentation strategy of GOP, that is, dividing the video into units with a fixed number of frames and treating each unit as an independent compression block. Under this structure, although each unit no longer depends on I-frames, the decoding end still needs to completely reconstruct from the beginning position of each unit, lacking an independent jumpable entry mechanism. Therefore, in random access scenarios, the following defects exist: First, the GOP period is generally too long, and the entry points are sparse, making it difficult to meet the needs of high-frequency interaction; second, the decoding process often depends on the temporal context of the preceding frame, making it impossible to achieve independent start at any entry point, requiring buffering and warm-up, which can easily cause stuttering or reconstruction delays; third, the GOP length is mostly set as a single static value, without considering the random access needs of specific application scenarios (such as video-on-demand, live broadcast, and broadcast television), ultimately making it difficult to balance response speed and bandwidth efficiency, resulting in a large video random access jump delay. Therefore, there is a technical problem of large video jump delays due to the failure to consider random access needs and reliance on the decoding of preceding video segments during random access. To address this issue, relevant solutions are provided in the embodiments of this application, which are described in detail below.
[0041] According to an embodiment of this application, an embodiment of a video stream processing method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0042] The methods and embodiments provided in this application can be executed on a computer terminal or similar computing device. Figure 1 A hardware block diagram of a computer terminal for implementing a video stream processing method is shown. Figure 1 As shown, the computer terminal 10 may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1The different configurations shown.
[0043] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As involved in the embodiments of this application, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).
[0044] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the video stream processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned video stream processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0045] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0046] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.
[0047] Under the above operating environment, embodiments of this application provide a video stream processing method, such as... Figure 2 The diagram shown is a flowchart of a first video stream processing method according to an embodiment of this application, including:
[0048] Step S202: Receive a random access operation request.
[0049] In the technical solution provided in step S202, the random access operation request is used to carry at least a target video segment identifier. The video segment indicated by the target video segment identifier includes: a segment (i.e., the target video segment) among multiple video segments obtained by dividing the original video stream according to the video random access requirement information. Each video segment corresponds to a video segment identifier and at least an entry point identifier information used to identify the starting position of the video segment.
[0050] Step S204: Determine the target token based on the random access operation request.
[0051] In the technical solution provided in step S204, the target token is obtained by encoding the target video segment, and the target token carries the target entry point identification information of the target video segment.
[0052] Step S206: Decode the target token based on the target entry point identifier information carried by the target token to obtain the target video segment.
[0053] Steps S202-S206 above involve receiving a random access request carrying a target video segment identifier, determining a target token generated from the encoded target video segment based on the identifier, and imbuing the target token with the entry point identifier information of its corresponding video segment. This allows the decoding process to be started independently based on this entry point identifier information during the decoding phase, without relying on any preceding video segment data or state. This mechanism enables each video segment to use its own entry point as the decoding starting point, achieving completely independent random access. This completely overcomes the inherent limitation of traditional generative video compression architectures, which require sequential decoding due to inter-frame dependencies and cannot skip to other segments. Therefore, it effectively solves the problem that generative video compression architectures cannot support independent random access per video segment, achieving efficient, accurate, low-latency random access and rapid positioning of video segments without increasing storage overhead or encoding complexity.
[0054] In some embodiments of this application, in order to achieve high compression and high-fidelity reconstruction of the video stream, the above steps S202-S206 are performed by a video stream processing system (which is an end-to-end generative video compression system). Figure 3This is a schematic diagram of the architecture of a video stream processing system provided according to an embodiment of this application. The generative video compression system of this application is composed of five functional modules working together: a video segmentation module 301, a variational autoencoder (VAE) 303, a compression downsampling module (Compressor Down) 305, a compression upsampling module (Compressor Up) 307, and a diffusion transformer denoising module (DiT) 309. According to function, the generative video compression system can be divided into two parts: an encoding end and a decoding end. The encoding end includes modules such as the video segmentation module 301, the encoder in the variational autoencoder (VAE) 303, and the compression downsampling module (Compressor Down) 305. The video stream processing path of the encoding end of the generative video compression system is video segmentation, VAE encoding, dimensionality reduction by the compression downsampling module until the target token is generated. The decoding end includes modules such as a decoder for the Variational Autoencoder (VAE303), a Compressor Up module (Compressor Up) 307, and a DiT (Differentiation Transformer) denoising module (DiT) 309. The video stream processing path at the decoding end is: dimensionality upsampling by the Compressor Up module, DiT denoising, and VAE decoding to reconstruct video frames. The decoding end and the encoding end interact through a transmission link.
[0055] The system architecture comprises several modules: a compression downsampling module, which compresses the high-dimensional feature representation output by the VAE into a learnable token with an entry point identifier; a compression upsampling module, which restores the high-dimensional feature representation from the learnable token; and a diffusion Transformer denoising module, independent of the VAE, which denoises the high-dimensional feature representation restored by the compression upsampling module to improve reconstruction quality. These modules form a closed-loop processing chain, collectively achieving low bitrate transmission, high-fidelity reconstruction, and millisecond-level random access. This system architecture breaks away from the traditional coding paradigm that relies on I-frames and inter-frame prediction, achieving end-to-end, reference-free, and randomly accessible video compression within a generative model framework, providing a new technical path for scenarios such as video-on-demand, live streaming, and broadcast television.
[0056] In the technical solution provided in step S202, each time a new original video stream is received, the original video stream is divided into multiple video segments according to the video random access requirement information in the following manner: obtaining the video random access requirement information, wherein the video random access requirement information is used at least to indicate the allowed latency threshold when performing video random access in the target application scenario; determining the target video frame threshold from multiple video frame thresholds based on the video random access requirement information, wherein the multiple video frame thresholds correspond to different numbers of video frames; dividing the original video stream into segments according to the target video frame threshold to obtain multiple initial video segments; and determining the multiple initial video segments carrying entry point identification information as multiple video segments.
[0057] It should be noted that this video random access requirement information is configured by the user or business system and at least indicates the allowable latency threshold for random video access under the target application scenario. Target application scenarios include, but are not limited to: video-on-demand, broadcast television, video conferencing, or live online streaming. Different application scenarios have significantly different allowable latency thresholds for random access. For example, broadcast television scenarios require channel switching latency to be no more than 0.6 seconds, video-on-demand scenarios allow drag-and-drop response latency to be no more than 1.8 seconds, while low-bandwidth video-on-demand scenarios, where compression efficiency is prioritized, can tolerate a maximum latency of 5.64 seconds. The allowable latency thresholds for random video access under the target application scenario include, but are not limited to, the drag-and-drop latency threshold for video-on-demand and the channel switching latency threshold for broadcast television.
[0058] In the above steps, multiple initial video segments carrying entry point identifier information are determined as multiple video segments. Taking any one of these video segments as an example, the process of obtaining the video segment corresponding to the initial video segment is explained as follows: Any one of the multiple video segments is determined as follows: The starting frame is obtained from the multiple video frames contained in the first initial video segment, where the first initial video segment is any one of the multiple initial video segments; visual features are extracted from the starting frame using a variational autoencoder to obtain the first high-dimensional feature map of the starting frame, and this first high-dimensional feature map is determined as the entry point of the first initial video segment; the position information corresponding to the starting frame is obtained, and the set consisting of the position information corresponding to the starting frame and the entry point of the first initial video segment is determined as the first entry point identifier information of the first initial video segment; the first initial video segment carrying the first entry point identifier information is determined as the video segment corresponding to the first initial video segment. This method establishes the initial segment carrying the entry point identifier information as a randomly accessible video segment, ensuring that each video segment accurately matches the latency requirements of the application scenario. This achieves a fast location and smooth playback experience after the user clicks to jump, without significantly reducing the video compression rate.
[0059] After obtaining the entry point identifier information required for all initial video segments using the above method, multiple initial video segments, each carrying its own entry point identifier information, are identified as multiple video segments. Once multiple video segments are obtained, the entire process of segmenting the initial video stream is complete. It should be noted that the segmentation process of the initial video stream is a prerequisite for step S202. For random access operation requests of the same original video stream, this step does not need to be repeated; only one video segmentation and encoding process is required for the same original video stream. For all subsequent random access operation requests (such as dragging and switching channels) of the same original video stream, the generated video segments and their entry point identifier information are directly reused without re-segmentation or encoding. This ensures that when a user triggers random access, only lightweight token retrieval and decoding are required, achieving a millisecond-level response.
[0060] The following section provides a further explanation of the process of segmenting the initial video stream.
[0061] In some embodiments of this application, the segmentation of the initial video stream is performed by the encoding end of a generative video compression system using the following steps:
[0062] Step 1: Obtain video random access requirement information and the original video stream (with a fixed frame rate, such as 25fps) through the video segmentation module at the encoding end of the generative video compression system. Based on this random access requirement information, determine the target video frame threshold from multiple video frame thresholds. These multiple video frame thresholds can be flexibly adapted to the needs of different application scenarios. The video frame thresholds are discrete values determined based on scenario testing and compression efficiency modeling of the target scenario, and each video frame corresponds to a theoretical latency. For example, the multiple video frame thresholds are divided into: 13 frames, 29 frames, 45 frames, 77 frames, and 141 frames, corresponding to theoretical latencies of 0.52 seconds, 1.16 seconds, 1.80 seconds, 3.08 seconds, and 5.64 seconds at a frame rate of 25fps, respectively.
[0063] Step 2: Determine the target video frame threshold from multiple video frame thresholds. Compare the allowable latency threshold for the target application scenario with the theoretical latency corresponding to each video frame threshold. Select the smallest video frame threshold that satisfies the latency constraint (not exceeding the corresponding latency threshold) and has the highest compression efficiency as the target video frame threshold. For example, for a video-on-demand scenario, when the allowable latency threshold for random video access in the target application scenario is 1.5 seconds, 29 frames are selected as the target video frame threshold because its latency (1.16 seconds ≤ 1.5 seconds) is better than the smaller 13 frames. For a broadcast television scenario, 13 frames (corresponding to a theoretical latency of 0.52 seconds) can be selected to balance switching latency and compression efficiency.
[0064] Step 3: After determining the target video frame threshold, the original video stream is periodically segmented. During segmentation, starting from the first frame, the original video stream is divided into multiple initial video clips according to the target video frame threshold. That is, starting from frame 1, every N frames (N = target video frame threshold) constitute one initial video clip. During the segmentation process, if the number of remaining frames at the end of the original video stream is less than the size of one target video frame threshold, it is treated as an independent initial video clip to ensure the integrity of the segmentation. Each initial video clip is logically the smallest independently processable unit, independent of any encoded information from preceding or succeeding segments (such as motion vectors or inter-frame prediction), providing structural guarantees for unbuffered random access.
[0065] After obtaining multiple initial video segments through the periodic segmentation of the original video stream in step 3, the crucial stage of assigning independent decoding capabilities to each initial video segment begins. This stage is a necessary step to achieve reference-frame-independent random access in the generative video compression architecture. Traditional coding standards achieve random access by periodically inserting I-frames (keyframes), but I-frames have large data volumes and complex encoding. In order to enable the decoding end of the generative video compression system to independently, stably, and with high fidelity reconstruct the starting frame content of the target video segment without any preceding video segment information, this application generates and binds entry point identifier information (EPI) for each initial video segment in step 4. The initial video segment carrying the entry point identifier information serves as the final structure (multiple video segments) of the original video stream segmentation process. The following uses the first initial video segment (which is any one of the multiple initial video segments) as an example to illustrate the execution process of step 4.
[0066] Step 4: Obtain the starting frame (the first frame among multiple video frames) from the multiple video frames contained in the first initial video segment. Extract visual features from the starting frame using a variational autoencoder to obtain the first high-dimensional feature map of the starting frame: Input the starting frame into the encoder of the VAE, which consists of multiple layers of 3D convolutional and non-linear activation layers. Extract visual features from the starting frame using the VAE encoder to obtain the high-dimensional feature map of the starting frame (the first high-dimensional feature map). The first high-dimensional feature map fully represents the semantic structure, texture, and edge information of the starting frame. This first high-dimensional feature map is determined as the entry point of the first initial video segment.
[0067] After obtaining the first high-dimensional feature map, the position information corresponding to the starting frame is obtained: the absolute timestamp or frame number of the starting frame in the original video stream is obtained, and a learnable position encoding layer (an embedding module that automatically learns the vector representation corresponding to the frame number through training) is used to map the absolute timestamp or frame number of the starting frame in the original video stream into a position vector, and this position vector is determined as the position information corresponding to the starting frame. Finally, the set consisting of the position information corresponding to the starting frame and the entry point of the first initial video segment is determined as the first entry point identifier information of the first initial video segment, and the first initial video segment carrying the first entry point identifier information is determined as the video segment corresponding to the first initial video segment.
[0068] Step 4 can identify multiple video segments. Each video segment is an initial video segment carrying entry point identification information. The entry point identification information contains two parts: the position information of the starting frame and the high-dimensional feature map of the starting frame. The video segments obtained after the original video stream segments are divided have independent decoding initialization capabilities. That is, after receiving a random access operation request, the decoding end only needs to obtain the target token corresponding to the video segment that the random access operation request wants to access in order to complete the decoding start, without buffering any data of the preceding video segments.
[0069] To achieve efficient transmission and compression of video content, directly transmitting the entire video segment would result in excessively large data volumes, failing to meet the transmission constraints of application scenarios such as low bandwidth, high concurrency, and real-time switching. Therefore, a mechanism must be introduced to compress the entire video segment into an extremely low-dimensional representation without relying on inter-frame prediction. This involves encoding multiple initial video segments into a compact, learnable target token. This learnable token serves as a semantic summary of the video content, preserving the core visual structure and motion features of the segment while compressing the data volume to the hundreds of bytes level, and still being effectively utilized to reconstruct high-quality video. Finally, the learnable token and entry point identifier information that identifies the decryption entry point are jointly encapsulated into a complete token. Multiple video segments from the original video stream must be encoded into tokens. Each encoded token contains two parts: entry point identifier information and the learnable token obtained from encoding the initial video segment. The following uses a target video segment as an example to illustrate the process of encoding the target video segment to obtain the target token.
[0070] There are several ways to encode a target video segment to obtain a target token. For example, a variational autoencoder can be used to extract visual features from the initial target video segment, resulting in a second high-dimensional feature map corresponding to the initial target video segment. The initial target video segment is then used as the target video segment itself. A target encoding strategy is applied to the second high-dimensional feature map for dimensionality reduction, yielding a target learnable token. This target encoding strategy is determined based on the number of video frames in the initial target video segment, and the target learnable token is a low-dimensional semantically compressed representation of the initial target video segment. Finally, the target learnable token is encapsulated with the target entry point identifier information of the target video segment to obtain the target token. The target encoding strategy includes a first encoding strategy and a second encoding strategy.
[0071] There are several ways to implement the target learnable token by using a target encoding strategy to reduce the dimensionality of the second high-dimensional feature map. For example, when the number of video frames in the initial target video segment is greater than or equal to a first preset threshold, a first encoding strategy is used to reduce the dimensionality of the second high-dimensional feature map to obtain the target learnable token. In this first encoding strategy, the target video segment is sequentially divided into multiple sub-segments, and the dimensionality of these sub-segments is reduced sequentially based on the dependencies between them. When the number of video frames in the initial target video segment is less than the first preset threshold, a second encoding strategy is used to reduce the dimensionality of the second high-dimensional feature map to obtain the target learnable token. In this second encoding strategy, the dimensionality reduction of the target video segment is performed all at once. The two strategies are automatically switched through a frame number threshold, avoiding redundant computation caused by applying complex temporal modeling to short segments and preventing information loss caused by single compression of long segments. This achieves adaptive and collaborative optimization of compression efficiency and encoding latency without changing the target token structure and entry point identification information, significantly improving the decoding response speed and reconstruction quality consistency of segments of different lengths in random access scenarios.
[0072] In some embodiments of this application, the encoder of the VAE at the encoding end of the aforementioned generative video compression system extracts visual features from the initial target video segment (the initial video segment corresponding to the target video segment that does not carry entry point identification information), obtaining a second high-dimensional feature map corresponding to the initial target video segment. This encoder consists of multiple layers of 3D convolutional and nonlinear activation layers. This second high-dimensional feature map, without any compression, completely preserves the semantic structure, texture details, and motion gradient information of each frame in the initial target video segment, and its dimension is directly related to the number of frames in the initial video segment.
[0073] To balance compression efficiency and reconstruction quality, and to adapt to the encoding requirements of video segments of different lengths, an encoding strategy is dynamically selected for dimensionality reduction based on the frame number of the initial target video segment, resulting in a learnable token for the target. This process is divided into two cases:
[0074] In scenario one, when the number of video frames in the initial target video segment is greater than or equal to a first preset threshold (e.g., 45 frames), the compression upsampling module of the generative video compression system uses a first encoding strategy to reduce the dimensionality of the second high-dimensional feature map, obtaining a learnable token for the target: The initial target video segment is sequentially divided into multiple sub-segments, each containing a fixed number of frames (e.g., 15 frames). After division into multiple sub-segments, dimensionality reduction is performed on each sub-segment according to the dependencies between them: First, the first sub-segment in the multiple sub-segments is compressed and reduced in dimensionality using the convolutional layer and linear projection layer of the compression upsampling module, obtaining the sub-segment token of the first sub-segment. The compression upsampling module consists of a 1×1×1 convolutional layer (for channel compression) and a linear projection layer (for dimensionality reduction to 128 dimensions). Based on the temporal dependencies between sub-segments, a hierarchical prediction structure is adopted: The compression process of the i-th sub-segment (which can be any sub-segment from the second to the last among multiple sub-segments) uses the sub-segment token of the (i-1)-th sub-segment as conditional input. This token is injected into the compression and dimensionality reduction process of the current sub-segment (the i-th sub-segment) through concatenation or a lightweight cross-attention mechanism, resulting in the sub-segment token of the i-th sub-segment. Finally, all sub-segment tokens are concatenated in chronological order to obtain the concatenated token, which is the target learnable token.
[0075] In the second scenario, when the number of video frames in the initial target video segment is less than the first preset threshold, a second encoding strategy is used to reduce the dimensionality of the second high-dimensional feature map to obtain a target learnable token. Under this strategy, instead of segmenting into sub-segments, the entire second high-dimensional feature map is directly input into the compression upsampling module. The compression upsampling module performs compression and dimensionality reduction through its convolutional and linear projection layers to obtain the target learnable token.
[0076] After generating the target learnable token, its validity is verified by performing the following checks: Integrity check: Confirming that no zero or abnormal values appear in the target learnable token to ensure numerical stability; Dimension check: Verifying that the dimension of the target learnable token conforms to preset rules (when the number of video frames in the initial target video segment is greater than or equal to the first preset threshold, the target learnable token dimension is 128×N (N is the number of sub-segments); when the number of video frames in the initial target video segment is less than the first preset threshold, the target learnable token dimension is 128 dimensions). Only when all checks pass is the target learnable token marked as valid, and the process proceeds to the next encapsulation step. The target learnable token and the target entry point identifier information of the target video segment are encapsulated to obtain the target token. The encapsulated target token contains the target entry point identifier information, the target video segment identifier, and the target learnable token.
[0077] In some embodiments of this application, multiple video segments obtained by dividing the original video stream are encoded using the same encoding method as the target video segment to obtain the target token, resulting in a token encoded for each video segment.
[0078] In step S202, a random access operation request triggered by the user is received. This random access operation request carries the identifier of the target video segment in the original video stream that needs to be accessed, i.e., the video segment identifier of the target video segment to be accessed. The random access operation request is the request information for a random access operation, which includes, but is not limited to, on-demand video dragging and broadcast television channel switching.
[0079] After receiving the random access operation request triggered by the user, proceed to step S204. There are several ways to determine the target token based on the random access operation request in step S204, such as: obtaining the target video segment identifier carried in the random access operation request; sending a target token request instruction to the transmission link based on the target video segment identifier; receiving a response message corresponding to the target token request instruction returned by the transmission link, wherein the response message carries at least the target token; and extracting the target token from the response message.
[0080] In some embodiments of this application, the decoding end of the generative video compression system receives a random access operation request and obtains the target video segment identifier carried in the random access operation request. It then sends a target token request instruction to the transmission link. This target token request instruction carries the target video segment identifier and is used to request a target token from the encoding end. The transmission link is the data transmission link between the decoding end and the encoding end of the generative video compression system. After receiving the target token request instruction, the encoding end returns a response message corresponding to the target token request instruction to the decoding end through the transmission link. This response message carries a pre-encapsulated target token. The transmission process does not require the transmission of original video data, effectively improving transmission efficiency and significantly reducing the amount of data transmitted during random access.
[0081] In the technical solution provided in step S206, there are several ways to decode the target token based on the target entry point identifier information carried by the target token to obtain the target video segment. For example: parse the target token to obtain the target entry point identifier information and the target learnable token carried by the target token; perform feature upscaling processing on the target learnable token to obtain the restored high-dimensional feature map; perform feature fusion processing on the restored high-dimensional feature map and the first high-dimensional feature map of the starting frame carried by the target entry point identifier information to obtain the fused feature map; perform denoising processing on the fused feature map to obtain the target fused feature map; and decode the target fused feature map to obtain the target video segment. This step significantly improves the reconstruction accuracy and visual coherence of independent video segments without contextual dependencies, and completely solves the reconstruction distortion problem caused by segment isolation in generative video compression architectures.
[0082] The following provides a detailed explanation of step S206.
[0083] Step S206.1: The target token is parsed by the decoding end of the generative video compression system to obtain the target entry point identification information and the target learnable token carried by the target token.
[0084] Step S206.2 involves performing feature upsampling on the target learnable token using the compression upsampling module at the decoding end to obtain a restored high-dimensional feature map. The compression upsampling module is a neural network structure symmetrical to the compression downsampling module, consisting of multiple transposed convolutional layers and nonlinear activation functions. The restored high-dimensional feature map contains the global semantics of the video content, but due to compression and dimensionality reduction losses, the features of the starting frame have significant semantic shifts (such as brightness distortion and structural blurring) compared to the original video starting frame, and cannot be directly used for high-quality reconstruction. Therefore, to align the decoding starting state with the real video content, step S206.3 is performed. This step uses the true latent representation of the starting frame (the first high-dimensional feature map of the starting frame) saved during the encoding stage as a semantic anchor point to dynamically correct the initial reconstruction state at the decoding end.
[0085] Step S206.3: Perform feature fusion between the restored high-dimensional feature map and the first high-dimensional feature map of the starting frame carried by the target entry point identification information, obtain the restored features corresponding to the starting frame in the restored high-dimensional feature map, calculate the feature similarity (e.g., cosine similarity) between the restored features and the first high-dimensional feature map, and determine that the current restored starting frame has serious distortion and needs to be semantically corrected if the feature similarity is less than a preset threshold (e.g., 0.8). Replace the features at the starting frame position in the restored high-dimensional feature map with the first high-dimensional feature map to obtain the fused feature map.
[0086] Step S206.4: The fused feature map is denoised using the Diffusion Transformer Denoising (DiT) module at the decoding end to eliminate noise introduced during compression, optimize feature quality, and obtain the target fused feature map. This module uses a pre-trained DiT to denoise the fused feature map. The denoising process does not change the semantic features of the fused feature map that have already been corrected by the real features of the entry points; it only performs fine-tuning on local areas such as texture details, edge sharpness, and color consistency. Each denoising step is adaptively adjusted based on the temporal correlation between the current frame and adjacent frames in the fused feature map to ensure that the continuity is not disturbed. The dimensions of the target fused feature map are (C, T, H), where C is the number of channels, T is the number of video frames in the initial target video segment, and H is the spatial resolution of the target fused feature map.
[0087] Step S206.5: Decode the target fusion feature map to obtain the target video segment: The decoder of the variational autoencoder (VAE) at the decoding end uses the initial frame position indicated by the position information in the target entry point identifier as the decoding starting point to decode the target fusion feature map and obtain the target video segment. This decoder is a deep neural network co-optimized with the encoder during the training phase. The decoder structure includes multiple upsampling convolutional layers, residual connection blocks, and nonlinear activation units. Its design goal is to accurately map high-dimensional latent features (target fusion feature map) back to pixel space. During the decoding process, the VAE decoder decodes the features of multiple video frames in the target fusion feature map sequentially according to the time sequence, starting from the initial frame. The processing flow is as follows:
[0088] Starting from the initial frame, the features of each video frame in the target fusion feature map are progressively expanded in spatial resolution to the target output resolution through multiple upsampling convolutional layers, resulting in expanded features. These expanded features are then input to multiple residual connection blocks (which restore high-frequency textures, edge sharpness, and local details lost during compression while preserving the overall feature structure). A non-linear activation unit maps the number of channels from C to RGB three channels (3 channels), ultimately yielding the decoded frame segments of each video frame in the target fusion feature map. The decoder sequentially combines all frame segments into the target video segment. The entire decoding process can be started directly from the entry point without buffering the preceding video segments, enabling rapid playback.
[0089] In some embodiments of this application, for specific application scenarios, such as latency-sensitive scenarios like broadcast television, a token priority transmission strategy is adopted to improve response speed and achieve low-latency transmission optimization. The token priority transmission strategy is implemented as follows: when a random access operation request instructs the user to trigger a channel switching operation, the token corresponding to the currently playing video segment is identified, and its transmission priority is set to the highest. In this way, the transmission link prioritizes sending the highest priority token to the decoding end for decoding, ensuring it arrives at the decoding end in the shortest possible time, thereby reducing the waiting time during channel switching and achieving rapid start-up. This strategy does not change the structure of the token itself, nor does it add extra data volume; it only adjusts the transmission scheduling order to achieve efficient response to critical content.
[0090] In some embodiments of this application, when receiving a random access operation request instructing the user to perform continuous playback or jump across video segments, the decoding end, based on the independent encoding characteristics of each video segment, completes seamless splicing of video frames, further ensuring the smoothness of random access. Specific implementation methods are as follows:
[0091] Timing synchronization between video segments: The decoding end arranges the restored video frames according to the time sequence of the original video based on the timing information in the video segment identifier carried by each video segment, ensuring that the playback sequence is completely consistent with the original video and avoiding frame order disorder.
[0092] Optimized splicing transition: For situations where multiple video segments need to be decoded simultaneously, at the junction of two adjacent video segments, the last frame of the previous video segment and the first frame of the next segment are used as transition references during splicing. A smooth transition algorithm is adopted (for example, splicing the decoded video segments by inter-frame interpolation or brightness and color smoothing) to eliminate playback stuttering caused by inter-frame differences and improve the continuous playback experience.
[0093] Dynamic Adaptation and Adjustment: The system acquires the target user's interactive behavior. When frequent dragging or channel switching is detected, the system obtains the user's historical behavior patterns. Based on these patterns (which reflect the user's past dragging or channel switching habits), the generative video compression system dynamically adjusts the request priority and decoding order of subsequent video segment tokens for that user. For example, it prioritizes preloading tokens for adjacent video segments indicating upcoming playback based on historical behavior patterns, and initiates the decoding process for these patterns earlier, thereby reducing stuttering caused by waiting for data and improving overall response stability.
[0094] In related technologies, the hybrid coding architecture based on "keyframes + prediction frames" relies on periodically inserted I-frames as entry points for random access. I-frames have large data volumes and high coding complexity, significantly increasing bandwidth consumption. The method in this application is based on a generative video compression system with a DiT generative model architecture. It uses video segments as basic units and generates tokens as transmission carriers through a VAE+ compression upsampling module. This eliminates the need for large I-frames, significantly reducing the amount of transmitted data while maintaining the number of entry points. In related technologies, random access entry points (I-frames) rely on reference information from preceding GOPs or a complete intra-frame prediction process. After the decoder jumps, it needs to buffer some subsequent frames for smooth playback, resulting in significant latency. In this application, video segments have completely independent decoding capabilities through entry point identification information. The decoder does not need to rely on any preceding video segment data and can achieve millisecond-level decoding of the target token without buffering. In related technologies, GOP length settings require a rigid trade-off between compression efficiency and access latency. Long GOPs have high compression efficiency but large jump latency, while short GOPs have low latency but lower compression ratio. The method in this application embodiment controls random access latency within the threshold required by the scenario without sacrificing compression efficiency. The decoding end can directly start decoding from the entry point of any video segment without relying on the information of the preceding video segments, realizing bufferless random jump and fast playback. The scenario-based adaptation of video segment size ensures the access experience under different application scenarios.
[0095] Figure 4 This is a flowchart of a second video stream processing method provided according to an embodiment of this application. The method is applied to the encoding end of the aforementioned generative video compression system and includes:
[0096] Step S402: Receive the raw video stream.
[0097] Step S404: Divide the original video stream according to the video random access requirement information to obtain multiple video segments.
[0098] Each video segment corresponds to a video segment identifier and at least an entry point identifier used to identify the starting position of the video segment.
[0099] Step S406: Encode multiple video segments separately to obtain multiple tokens.
[0100] Step S408: Receive the target token request instruction sent by the decoding end to the transmission link, wherein the target token request instruction is used to indicate the target video segment identifier carried in the random access operation request.
[0101] Step S410: Determine the target token from the multiple tokens. The target token is the token obtained by encoding the video segment indicated by the target video segment identifier from the multiple tokens.
[0102] Step S412: Send a response message corresponding to the target token request instruction to the decoding end via the transmission link, wherein the response message carries at least the target token.
[0103] In the technical solution provided by S404, the original video stream is divided into multiple video segments according to the video random access requirement information in the following way: obtaining the video random access requirement information, wherein the video random access requirement information is used to at least indicate the allowed latency threshold when performing video random access in the target application scenario; determining the target video frame threshold from multiple video frame thresholds based on the video random access requirement information, wherein the multiple video frame thresholds correspond to different numbers of video frames; dividing the original video stream into segments according to the target video frame threshold to obtain multiple initial video segments; and determining the multiple initial video segments carrying entry point identification information as multiple video segments.
[0104] The following method is used to determine any one of multiple video segments: Obtain the starting frame from the multiple video frames contained in the first initial video segment, where the first initial video segment is any one of the multiple initial video segments; extract visual features from the starting frame using a variational autoencoder to obtain a first high-dimensional feature map of the starting frame, and determine the first high-dimensional feature map as the entry point of the first initial video segment; obtain the position information corresponding to the starting frame, and determine the set consisting of the position information corresponding to the starting frame and the entry point of the first initial video segment as the first entry point identifier information of the first initial video segment; determine the first initial video segment carrying the first entry point identifier information as the video segment corresponding to the first initial video segment.
[0105] The technical solution provided in step S406, taking the encoding of a target video segment to obtain a target token as an example, illustrates the process of encoding multiple video segments separately to obtain multiple tokens: Encoding a target video segment to obtain a target token includes: extracting visual features from an initial target video segment using a variational autoencoder to obtain a second high-dimensional feature map corresponding to the initial target video segment, wherein the initial target video segment is the initial video segment corresponding to the target video segment; performing dimensionality reduction processing on the second high-dimensional feature map using a target encoding strategy to obtain a target learnable token, wherein the target encoding strategy is determined based on the number of video frames of the initial target video segment; and encapsulating the target learnable token with the target entry point identifier information of the target video segment to obtain the target token.
[0106] The process of using a target encoding strategy to reduce the dimensionality of the second high-dimensional feature map to obtain a target learnable token is as follows: When the number of video frames of the initial target video segment is greater than or equal to a first preset threshold, the first encoding strategy is used to reduce the dimensionality of the second high-dimensional feature map to obtain a target learnable token. The first encoding strategy divides the target video segment into multiple sub-segments in sequence and performs dimensionality reduction on the multiple sub-segments sequentially according to the dependencies between the multiple sub-segments. When the number of video frames of the initial target video segment is less than the first preset threshold, the second encoding strategy is used to reduce the dimensionality of the second high-dimensional feature map to obtain a target learnable token. The second encoding strategy performs a one-time dimensionality reduction on the target video segment.
[0107] It should be noted that the specific implementation process of steps S402 to S412 has been explained in the embodiments corresponding to steps S202-S206, and the relevant explanations of steps S202-S206 apply to steps S402-S412, and will not be repeated here.
[0108] Figure 5 This is a flowchart of a third video stream processing method provided according to an embodiment of this application. The method is applied to the decoding end of the aforementioned generative video compression system and includes:
[0109] Step S502: Receive a random access operation request, wherein the random access operation request is used to carry at least the identifier of the target video segment in the original video stream to be accessed. The video segment indicated by the target video segment identifier includes: a segment from multiple video segments obtained by dividing the original video stream according to the video random access requirement information, each video segment corresponding to a video segment identifier and at least entry point identifier information used to identify the starting position of the video segment.
[0110] Step S504: Obtain the target video segment identifier carried in the random access operation request, and send a request instruction to the transmission link to request the target token based on the target video segment identifier.
[0111] In the technical solution provided in step S504, the request instruction for requesting the target token is the aforementioned target token request instruction. The target token is obtained by encoding the target video segment indicated by the target video segment identifier, and the target token carries the target entry point identifier information of the target video segment.
[0112] Step S506: Receive the response message corresponding to the request instruction returned by the transmission link and extract the target token from the response message.
[0113] The response message must carry at least the target token.
[0114] Step S508: Decode the target token based on the target entry point identifier information carried by the target token to obtain the target video segment.
[0115] There are several ways to decode the target token based on the target entry point identifier information carried by the target token to obtain the target video segment. For example: parse the target token to obtain the target entry point identifier information and the target learnable token; perform feature upscaling on the target learnable token to obtain the restored high-dimensional feature map; perform feature fusion on the restored high-dimensional feature map and the first high-dimensional feature map of the starting frame carried by the target entry point identifier information to obtain the fused feature map; perform denoising on the fused feature map to obtain the target fused feature map; and decode the target fused feature map to obtain the target video segment.
[0116] The aforementioned learnable token is obtained by encoding the initial target video segment at the encoding end. The encoding end divides the original video stream into multiple video segments according to the video random access requirement information in the following ways: obtaining the video random access requirement information, wherein the video random access requirement information is used to at least indicate the allowed latency threshold when performing video random access in the target application scenario; determining the target video frame threshold from multiple video frame thresholds based on the video random access requirement information, wherein the multiple video frame thresholds correspond to different numbers of video frames; dividing the original video stream into segments according to the target video frame threshold to obtain multiple initial video segments; and determining the multiple initial video segments carrying entry point identification information as multiple video segments. The initial target video segment is the initial video segment corresponding to the target video segment, and the initial target video segment and the target video segment have the same video segment identifier. The target video segment is the initial target video segment carrying the target entry point identification information.
[0117] The target learnable token is a low-dimensional semantic compressed representation used to characterize the initial target video segment. The target learnable token is obtained by: extracting visual features from the initial target video segment using a variational autoencoder to obtain a second high-dimensional feature map corresponding to the initial target video segment, where the initial target video segment is the initial video segment corresponding to the target video segment; and then performing dimensionality reduction on the second high-dimensional feature map using a target encoding strategy to obtain the target learnable token, where the target encoding strategy is determined based on the number of video frames in the initial target video segment.
[0118] It should be noted that the specific implementation process of steps S502 to S508 has been explained in the embodiments corresponding to steps S202-S206, and the relevant explanations of steps S202-S206 apply to steps S502-S508, and will not be repeated here.
[0119] Figure 6This is a flowchart of the fourth video stream processing method provided in the embodiments of this application. After the original video is input, the video segmentation module performs scene requirement analysis, determines the video segment size, and performs periodic slicing (corresponding to the above-mentioned division of the original video stream into multiple initial video segments according to the video random access requirement information). The second step involves configuring video segment-level random access entry points, determining the entry point based on the starting frame, and independently decoding and initializing (corresponding to the above-mentioned determination of multiple initial video segments carrying entry point identifier information into multiple video segments). The third step involves video segment encoding and token generation through VAE encoding and Compressor Down to perform long video segment hierarchical prediction and short video segment rapid generation (corresponding to the above-mentioned process of encoding the target video segment to obtain the target token). The fourth step involves selective token transmission through video segment-level encapsulation and random access request priority scheduling (corresponding to the above-mentioned token priority transmission strategy). Video segment-level decoding and video restoration are performed through Compressor Up → DiT denoising → VAE decoding (corresponding to the process of decoding the target token based on the target entry point identifier information carried by the target token to obtain the target video segment). At the same time, through timing synchronization (corresponding to the timing synchronization between video segments), smooth transition (corresponding to the splicing transition optimization), and dynamic adjustment (corresponding to the dynamic adaptation adjustment), seamless splicing of multiple video segments and experience optimization are achieved.
[0120] Figure 7 This is a schematic diagram of a video stream processing apparatus according to an embodiment of this application, comprising:
[0121] The receiving module 702 is used to receive a random access operation request, wherein the random access operation request is used to carry at least the identifier of the target video segment in the original video stream to be accessed, and the video segment indicated by the target video segment identifier includes: a segment from multiple video segments obtained by dividing the original video stream according to the video random access requirement information, each video segment corresponds to a video segment identifier and at least entry point identifier information used to identify the starting position of the video segment.
[0122] The determination module 704 is used to determine the target token based on the random access operation request. The target token is obtained by encoding the target video segment and carries the target entry point identification information of the target video segment.
[0123] The decoding module 706 is used to decode the target token based on the target entry point identification information carried by the target token to obtain the target video segment.
[0124] It should be noted that, Figure 7 The video stream processing device shown is used to perform Figure 2 The video stream processing method shown is therefore Figure 2 The explanations and descriptions in the video stream processing method also apply to the video stream processing device, and will not be repeated here.
[0125] It should be noted that the modules in the aforementioned video stream processing device can be program modules (e.g., a set of program instructions that implement a specific function) or hardware modules. For the latter, they can take the following forms, but are not limited to them: each of the aforementioned modules is represented by a processor, or the functions of each of the aforementioned modules are implemented by a processor.
[0126] This application also provides a non-volatile storage medium, which includes a stored program, wherein, when the program is running, it controls the device where the non-volatile storage medium is located to execute the video stream processing method of any of the above embodiments.
[0127] This application also provides an electronic device, which includes a processor for running a program, wherein the video stream processing method of any of the above embodiments is executed when the program is running.
[0128] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the video stream processing method of any of the above embodiments.
[0129] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0130] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0131] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0132] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0133] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0134] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for processing a video stream, characterized in that, include: Receive a random access operation request, wherein the random access operation request is at least used to carry a target video segment identifier, and the video segment indicated by the target video segment identifier includes: a segment from multiple video segments obtained by dividing the original video stream according to video random access requirement information, each video segment corresponding to a video segment identifier and at least entry point identifier information used to identify the starting position of the video segment; A target token is determined based on the random access operation request, wherein the target token is obtained by encoding the target video segment and carries the target entry point identification information of the target video segment; The target token is decoded based on the target entry point identification information carried by the target token to obtain the target video segment; The original video stream was divided into multiple video segments according to the video random access requirement information in the following manner: Obtain the video random access requirement information, wherein the video random access requirement information is at least used to indicate the allowed latency threshold when performing random video access in the target application scenario; The target video frame threshold is determined from multiple video frame thresholds based on the video random access requirement information, wherein the multiple video frame thresholds correspond to different numbers of video frames. The original video stream is divided into segments according to the target video frame threshold to obtain multiple initial video segments; Multiple initial video segments carrying entry point identification information are identified as the multiple video segments.
2. The method according to claim 1, characterized in that, The method further includes determining any one of the plurality of video segments by: Obtain the starting frame from the plurality of video frames contained in the first initial video segment, wherein the first initial video segment is any one of the plurality of initial video segments; Visual features are extracted from the starting frame by a variational autoencoder to obtain a first high-dimensional feature map of the starting frame, and the first high-dimensional feature map is determined as the entry point of the first initial video segment. Obtain the position information corresponding to the starting frame, and determine the set consisting of the position information corresponding to the starting frame and the entry point of the first initial video segment as the first entry point identifier information of the first initial video segment; The first initial video segment carrying the first entry point identifier information is determined as the video segment corresponding to the first initial video segment.
3. The method according to claim 2, characterized in that, The method further includes: encoding the target video segment to obtain the target token, including: Visual features are extracted from the initial target video segment by a variational autoencoder to obtain a second high-dimensional feature map corresponding to the initial target video segment, wherein the initial target video segment is the initial video segment corresponding to the target video segment; The second high-dimensional feature map is subjected to dimensionality reduction processing using a target encoding strategy to obtain a target learnable token, wherein the target encoding strategy is determined based on the number of video frames of the initial target video segment; The target learnable token is encapsulated with the target entry point identifier information of the target video segment to obtain the target token.
4. The method according to claim 3, characterized in that, The second high-dimensional feature map is subjected to dimensionality reduction using a target encoding strategy to obtain target learnable tokens, including: When the number of video frames of the initial target video segment is greater than or equal to a first preset threshold, the second high-dimensional feature map is reduced in dimensionality using a first encoding strategy to obtain the target learnable token. The first encoding strategy divides the initial target video segment into multiple sub-segments in sequence and performs dimensionality reduction on the multiple sub-segments in sequence according to the dependency relationship between the multiple sub-segments. If the number of video frames in the initial target video segment is less than the first preset threshold, a second encoding strategy is used to reduce the dimensionality of the second high-dimensional feature map to obtain the target learnable token. The second encoding strategy performs a one-time dimensionality reduction on the initial target video segment.
5. The method according to claim 1, characterized in that, Determining the target token based on the random access operation request includes: Obtain the target video segment identifier carried in the random access operation request; A target token request instruction is sent to the transmission link based on the target video segment identifier; Receive a response message corresponding to the target token request instruction returned by the transmission link, wherein the response message carries at least the target token; Extract the target token from the response message.
6. The method according to claim 3, characterized in that, The target token is decoded based on the target entry point identifier information carried by the target token to obtain the target video segment, including: The target token is parsed to obtain the target entry point identifier information carried by the target token and the target learnable token; The target learnable token is subjected to feature upsizing processing to obtain a restored high-dimensional feature map; The restored high-dimensional feature map and the first high-dimensional feature map of the starting frame carried by the target entry point identification information are subjected to feature fusion processing to obtain the fused feature map; The fused feature map is then denoised to obtain the target fused feature map. The target fusion feature map is decoded to obtain the target video segment.
7. A method for processing a video stream, characterized in that, include: Receive the raw video stream; The original video stream is divided according to the random access requirement information of the video to obtain multiple video segments, wherein each video segment corresponds to a video segment identifier and at least an entry point identifier for identifying the start position of the video segment; The multiple video segments are encoded separately to obtain multiple tokens; The decoder receives a target token request instruction sent to the transmission link, wherein the target token request instruction is used to indicate the target video segment identifier carried in the random access operation request; A target token is determined from the plurality of tokens, wherein the target token is a token obtained by encoding the video segment indicated by the target video segment identifier from the plurality of tokens; The response message corresponding to the target token request instruction is sent to the decoding end through the transmission link, wherein the response message carries at least the target token; The original video stream is divided according to the video random access demand information to obtain multiple video segments, including: Obtain the video random access requirement information, wherein the video random access requirement information is at least used to indicate the allowed latency threshold when performing random video access in the target application scenario; The target video frame threshold is determined from multiple video frame thresholds based on the video random access requirement information, wherein the multiple video frame thresholds correspond to different numbers of video frames. The original video stream is divided into segments according to the target video frame threshold to obtain multiple initial video segments; Multiple initial video segments carrying entry point identification information are identified as the multiple video segments.
8. The method according to claim 7, characterized in that, Any one of the plurality of video segments is determined in the following manner: Obtain the starting frame from the plurality of video frames contained in the first initial video segment, wherein the first initial video segment is any one of the plurality of initial video segments; Visual features are extracted from the starting frame by a variational autoencoder to obtain a first high-dimensional feature map of the starting frame, and the first high-dimensional feature map is determined as the entry point of the first initial video segment. Obtain the position information corresponding to the starting frame, and determine the set consisting of the position information corresponding to the starting frame and the entry point of the first initial video segment as the first entry point identifier information of the first initial video segment; The first initial video segment carrying the first entry point identifier information is determined as the video segment corresponding to the first initial video segment.
9. A method for processing a video stream, characterized in that, include: A random access operation request is received, wherein the random access operation request is at least used to carry a target video segment identifier in the original video stream to be accessed, and the video segment indicated by the target video segment identifier includes: a segment from multiple video segments obtained by dividing the original video stream according to video random access requirement information, each video segment corresponding to a video segment identifier and at least entry point identifier information used to identify the start position of the video segment; the original video stream is divided into multiple video segments according to the video random access requirement information in the following manner: obtaining the video random access requirement information, wherein the video random access requirement information is at least used to indicate the allowed latency threshold when performing video random access in the target application scenario; determining a target video frame threshold from multiple video frame thresholds according to the video random access requirement information, wherein the multiple video frame thresholds correspond to different numbers of video frames; dividing the original video stream into segments according to the target video frame threshold to obtain multiple initial video segments; and determining the multiple initial video segments carrying entry point identifier information as the multiple video segments. Obtain the target video segment identifier carried in the random access operation request, and send a request instruction for requesting a target token to the transmission link based on the target video segment identifier. The target token is obtained by encoding the target video segment indicated by the target video segment identifier, and the target token carries the target entry point identifier information of the target video segment. Receive the response message corresponding to the request instruction returned by the transmission link, and extract the target token from the response message; The target token is decoded based on the target entry point identifier information carried by the target token to obtain the target video segment.
10. The method according to claim 9, characterized in that, The target token is decoded based on the target entry point identifier information carried by the target token to obtain the target video segment, including: The target token is parsed to obtain the target entry point identifier information and the target learnable token carried by the target token; The target learnable token is subjected to feature upsizing processing to obtain a restored high-dimensional feature map; The restored high-dimensional feature map and the first high-dimensional feature map of the starting frame carried by the target entry point identification information are subjected to feature fusion processing to obtain the fused feature map; The fused feature map is then denoised to obtain the target fused feature map. The target fusion feature map is decoded to obtain the target video segment.
11. A video stream processing apparatus, characterized in that, include: A receiving module is configured to receive a random access operation request, wherein the random access operation request carries at least a target video segment identifier in the original video stream to be accessed, and the video segment indicated by the target video segment identifier includes: a segment from multiple video segments obtained by dividing the original video stream according to video random access requirement information, each video segment corresponding to a video segment identifier and at least entry point identifier information for identifying the start position of the video segment; the original video stream is divided into multiple video segments according to the video random access requirement information in the following manner: obtaining the video random access requirement information, wherein the video random access requirement information is at least used to indicate the allowed latency threshold when performing random video access in the target application scenario; determining a target video frame threshold from multiple video frame thresholds based on the video random access requirement information, wherein the multiple video frame thresholds correspond to different numbers of video frames; dividing the original video stream into segments according to the target video frame threshold to obtain multiple initial video segments; and determining the multiple initial video segments carrying entry point identifier information as the multiple video segments. The determination module is used to determine the target token based on the random access operation request, wherein the target token is obtained by encoding the target video segment and carries the target entry point identification information of the target video segment; The decoding module is used to decode the target token based on the target entry point identification information carried by the target token to obtain the target video segment.
12. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a program, wherein when the program is executed, it controls the device where the non-volatile storage medium is located to perform the video stream processing method according to any one of claims 1 to 6, or the video stream processing method according to any one of claims 7 to 8, or the video stream processing method according to any one of claims 9 to 10.
13. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, performs a video stream processing method according to any one of claims 1 to 6, or a video stream processing method according to any one of claims 7 to 8, or a video stream processing method according to any one of claims 9 to 10.
14. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the video stream processing method according to any one of claims 1 to 6, or the video stream processing method according to any one of claims 7 to 8, or the video stream processing method according to any one of claims 9 to 10.
Citation Information
Patent Citations
Method and device for processing encoded video data, and method and device for generating encoded video data
CN107005704A
Techniques for random access point indication and picture output in encoded video stream
CN114287132A