Video stream transmission method and system, video stream sending end and video stream receiving end

Through the encoder-decoder collaborative design, combining compression methods to reduce frame resolution and reduce pixel color bit depth, and a diffusion recovery model is built on the receiving end, solving the problem of both compression rate and recovery quality in video streaming, achieving efficient video streaming and recovery effects.

CN120281912APending Publication Date: 2025-07-08SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510400982.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing video encoding technology cannot take into account the compression rate of the sending end of the video stream and the recovery quality of the receiver. Traditional methods focus on the encoder efficiency and ignore the recovery ability of the decoder, resulting in excessive loss of video information and lack of universality and flexibility in the neural recovery model.

Method used

The encoder-decoder collaborative design is adopted to compress by reducing frame resolution and reducing pixel color bit depth, and a diffusion recovery model is built on the receiving end, optimizing the encoder's compression conditions to improve the decoder's recovery capability.

Benefits of technology

It significantly improves the compression efficiency and recovery quality of video streaming, realizes high compression rate at the sending end of the video stream and high visual quality recovery at the receiving end, and is compatible with a variety of codecs and adapts to different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281912A_ABST
    Figure CN120281912A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of video transmission, and provides a video stream transmission method and system, a video stream sending end and a video stream receiving end, and the method comprises the steps that the video stream sending end obtains video original data; performing compression processing on the plurality of original video frames according to a preset compression condition to obtain video compression data corresponding to the video original data, and sending the video compression data to a video stream receiving end; the video stream receiving end receives the video compressed data sent by the video stream sending end; performing resolution recovery processing on the plurality of compressed video frames through fast bilinear interpolation to obtain a first recovered video frame; performing channel dimension splicing processing on the plurality of first recovered video frames and the initial Gaussian noise to obtain a splicing result; and performing image recovery processing on the splicing result through a diffusion recovery model to obtain video recovery data corresponding to the video compression data. The video stream compression ratio of the video stream sending end can be maximized, and the video stream recovery quality of the video stream receiving end can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of video transmission, and particularly relates to a video stream transmission method and system, a video stream sender, and a video stream receiver. Background Art

[0002] In recent years, Internet video traffic has increased rapidly. In order to adapt to the uplink bandwidth limitation of the sender, videos are usually compressed and transmitted at a relatively low quality.

[0003] However, traditional video coding technologies usually only focus on the compression efficiency at the encoder end. In this way, too much loss of video information after compression may occur, exceeding the recovery range of the decoder. At the same time, most traditional video coding technologies are also optimized based on the traditional rate-distortion theory, that is, to minimize distortion at a given bit rate. However, this optimization method of pursuing low distortion does not always bring high perceptual quality. Summary of the Invention

[0004] Embodiments of this application provide a video stream transmission method and system, a video stream sender, and a video stream receiver, aiming to solve the problem that the video stream transmission method in related technologies cannot take into account both the video stream compression ratio at the video stream sender and the video stream recovery quality at the video stream receiver.

[0005] In a first aspect, embodiments of this application provide a video stream transmission method, which is applied to a video stream sender. The method includes:

[0006] Obtain video original data, where the video original data includes a plurality of original video frames;

[0007] According to preset compression conditions, perform compression processing on the plurality of original video frames to obtain video compression data corresponding to the video original data, where the compression processing includes reducing frame resolution processing and reducing pixel color bit depth processing;

[0008] Send the video compression data corresponding to the video original data to a video stream receiver, so that the video stream receiver performs recovery processing on the video compression data to obtain video recovery data corresponding to the video compression data.

[0009] In a possible implementation manner of the first aspect, the preset compression conditions include a first preset compression condition and a second preset compression condition.

[0010] The step of performing compression processing on the plurality of original video frames according to the preset compression conditions to obtain video compression data corresponding to the video original data includes:

[0011] Performing a frame resolution reduction process on the multiple original video frames according to the first preset compression condition to obtain first video compression data, where the first video compression data includes multiple first compressed video frames;

[0012] Performing a pixel color bit depth reduction process on the multiple first compressed video frames according to the second preset compression condition to obtain the video compression data corresponding to the video original data.

[0013] In a possible implementation manner of the first aspect, the performing a frame resolution reduction process on the multiple original video frames according to the first preset compression condition to obtain first video compression data includes:

[0014] Dividing each original video frame in the video original data into multiple initial macroblocks according to a preset scaling factor;

[0015] Based on a Gaussian blur algorithm, smoothing the features at the intersections of the multiple initial macroblocks to obtain multiple processed macroblocks corresponding to the multiple initial macroblocks;

[0016] Calculating the weighted average value of all pixels in each processed macroblock in each original video frame, and compressing each processed macroblock to the weighted average value to obtain the first compressed video frame corresponding to each original video frame;

[0017] Arranging the multiple first compressed video frames in the order of the multiple original video frames in the video original data to obtain the first video compression data.

[0018] In a possible implementation manner of the first aspect, the performing a pixel color bit depth reduction process on the multiple first compressed video frames according to the second preset compression condition to obtain the video compression data corresponding to the video original data includes:

[0019] Performing a vector quantization process on the color values of all pixels in the multiple first compressed video frames to obtain pixel vectors corresponding to all pixels in the multiple first compressed video frames;

[0020] Performing a clustering analysis on all the pixel vectors in the multiple first compressed video frames, and converting full-bit depth pixels in the multiple first compressed video frames into low-bit depth pixels to obtain second compressed video frames corresponding to the multiple first compressed video frames;

[0021] Arranging the multiple second compressed video frames in the order of the multiple original video frames in the video original data to obtain the video compression data corresponding to the video original data.

[0022] Second aspect, an embodiment of the present application provides a video stream transmission method, which is applied to a video stream receiving end. The method includes:

[0023] Receiving video compression data sent by a video stream sending end, where the video compression data is obtained by the video stream sending end performing compression processing on multiple original video frames in the original video data according to preset compression conditions, and the video compression data includes multiple compressed video frames;

[0024] Performing resolution restoration processing on multiple compressed video frames through fast bilinear interpolation to obtain first restored video frames corresponding to the multiple compressed video frames;

[0025] Performing splicing processing on multiple first restored video frames and initial Gaussian noise in the channel dimension to obtain a splicing result;

[0026] Performing image restoration processing on the splicing result through a diffusion restoration model to obtain video restoration data corresponding to the video compression data; wherein, multiple restoration parameters in the diffusion restoration model are optimized according to the preset compression conditions in the video stream sending end.

[0027] In a possible implementation manner of the second aspect, the performing image restoration processing on the splicing result through a diffusion restoration model to obtain video restoration data corresponding to the video compression data includes:

[0028] Performing image restoration processing on the splicing result through the diffusion restoration model to generate second restored video frames corresponding to the multiple first restored video frames;

[0029] Arranging multiple second restored video frames in the order of multiple compressed video frames in the video compression data to obtain the video restoration data corresponding to the video compression data.

[0030] In a possible implementation manner of the second aspect, before the performing image restoration processing on the splicing result through a diffusion restoration model to obtain video restoration data corresponding to the video compression data, the method includes:

[0031] Obtaining a training data set, where the training data set includes multiple original video frames and compressed video frames corresponding to the multiple original video frames;

[0032] Constructing an initial diffusion restoration model and initializing multiple restoration parameters in the initial diffusion restoration model through a normal distribution function;

[0033] Gradually adding Gaussian noise to multiple original video frames according to a preset attenuation coefficient and preset time steps to generate compressed video frames corresponding to the multiple original video frames;

[0034] Based on the current state of the compressed video frame, the noise level, and the preset compression conditions in the video stream sender, obtain the predicted noise in the compressed video frame, and remove the predicted noise in the compressed video frame to obtain the original video frame corresponding to the compressed video frame;

[0035] By using the gradient descent algorithm to minimize the difference value between the predicted noise and the Gaussian noise, optimize multiple recovery parameters in the initial diffusion recovery model to obtain the diffusion recovery model.

[0036] In a third aspect, an embodiment of the present application provides a video stream transmission system, including: a video stream sender and a video stream receiver, where,

[0037] There is a communication connection between the video stream sender and the video stream receiver;

[0038] The video stream sender executes the video stream transmission method described in any one of the above first aspects; the video stream receiver executes the video stream transmission method described in any one of the above second aspects.

[0039] In a fourth aspect, an embodiment of the present application provides a video stream sender, including:

[0040] An acquisition module, configured to acquire video original data, where the video original data includes multiple original video frames;

[0041] A compression module, configured to perform compression processing on multiple original video frames according to preset compression conditions to obtain video compression data corresponding to the video original data, where the compression processing includes reducing frame resolution processing and reducing pixel color bit depth processing;

[0042] A sending module, configured to send the video compression data corresponding to the video original data to a video stream receiver, so that the video stream receiver performs recovery processing on the video compression data to obtain video recovery data corresponding to the video compression data.

[0043] In a fifth aspect, an embodiment of the present application provides a video stream receiver, including:

[0044] A video receiving module, configured to receive video compression data sent by a video stream sender, where the video compression data is obtained by the video stream sender performing compression processing on multiple original video frames in video original data, and the video compression data includes multiple compressed video frames;

[0045] A first restoration module, configured to perform resolution restoration processing on the multiple compressed video frames through fast bilinear interpolation to obtain first restored video frames corresponding to the multiple compressed video frames;

[0046] A splicing processing module, configured to perform splicing processing on the multiple first restored video frames and initial Gaussian noise in the channel dimension to obtain a splicing result;

[0047] A second restoration module, configured to perform image restoration processing on the splicing result through a diffusion restoration model to obtain video restoration data corresponding to the video compression data; wherein, multiple restoration parameters in the diffusion restoration model are optimized according to the preset compression conditions in the video stream sending end.

[0048] In a sixth aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the video stream transmission method described in any one of the above is implemented.

[0049] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the video stream transmission method described in any one of the above is implemented.

[0050] In an eighth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is enabled to execute the video stream transmission method described in any one of the first aspects above.

[0051] The beneficial effects of the embodiments of the present application compared with the prior art are:

[0052] An embodiment of the present application provides a video stream transmission method. The video stream sender obtains original video data, where the original video data includes multiple original video frames; performs compression processing on the multiple original video frames according to preset compression conditions to obtain video compression data corresponding to the original video data, where the compression processing includes reducing frame resolution processing and reducing pixel color bit depth processing; sends the video compression data corresponding to the original video data to the video stream receiver, so that the video stream receiver performs restoration processing on the video compression data to obtain video restoration data corresponding to the video compression data. It can reduce the frame resolution and the pixel color bit depth, improve the compression efficiency of the video stream sender, and thus significantly reduce the transmission bit rate of the video; then, the video stream receiver receives the video compression data sent by the video stream sender, where the video compression data includes multiple compressed video frames; performs resolution restoration processing on the multiple compressed video frames through fast bilinear interpolation to obtain the first restored video frames corresponding to the multiple compressed video frames; performs channel dimension splicing processing on the multiple first restored video frames and initial Gaussian noise to obtain a splicing result; performs image restoration processing on the splicing result through a diffusion restoration model to obtain video restoration data corresponding to the video compression data; where multiple restoration parameters in the diffusion restoration model are optimized according to the preset compression conditions in the video stream sender. The video stream receiver can perceive the compression conditions of the resolution and color of the video stream through the constructed diffusion restoration model, so as to provide a stronger visual quality restoration ability. This method can maximize the video stream compression rate of the video stream sender, and at the same time, can also improve the video stream restoration quality of the video stream receiver. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0054] Figure 1 It is a schematic structural diagram of a video stream transmission system provided by an embodiment of the present application;

[0055] Figure 2 It is a schematic flowchart of a video stream transmission method provided by an embodiment of the present application;

[0056] Figure 3 It is a schematic flowchart of another video stream transmission method provided by an embodiment of the present application;

[0057] Figure 4 It is a schematic structural diagram of a video stream sender provided by an embodiment of the present application;

[0058] Figure 5 It is a schematic structural diagram of a video stream receiving end provided by an embodiment of the present application;

[0059] Figure 6 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Specific implementation manners

[0060] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are put forward in order to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0061] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0062] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0063] As used in the specification of the present application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0064] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0065] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but rather mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0066] In recent years, Internet video traffic has grown rapidly. To adapt to the uplink bandwidth limitations of senders, videos are usually compressed and transmitted at a relatively low quality. There are the following three levels of deficiencies in the existing related technologies:

[0067] First, traditional video coding technologies usually only focus on the compression efficiency of the video stream sender (encoder), while the recovery process of the video stream receiver (decoder) is regarded as an independent task. If the encoder does not consider the recovery ability of the decoder, it may lead to excessive loss of compressed video information, thus exceeding the recovery range of the decoder, resulting in a lack of active cooperation between the video stream sender and the video stream receiver.

[0068] Second, traditional rate - distortion theory mainly focuses on the objective distortion of signals, such as mean - square error, etc. These objective metrics are not always consistent with human subjective perceptual quality. Especially in the field of video compression, the human eye's perception of visual information is highly non - linear and complex. Therefore, simply pursuing low distortion is not sufficient to guarantee high perceptual quality. However, most of the existing related technologies are optimized based on traditional rate - distortion theory, that is, pursuing to minimize distortion at a given bitrate. This optimization method does not always bring high perceptual quality.

[0069] Third, existing neural recovery models are often designed in a coupled manner with specific video codecs or tasks. Since different video codecs have different compression mechanisms and characteristics, a neural recovery model designed for a specific codec may not be able to adapt to the compression methods of other codecs, lacking generality and unable to be generalized to different video codecs or application scenarios. In addition, different application scenarios have different requirements for video quality. Therefore, a more general and flexible recovery model is needed to adapt to different needs.

[0070] Therefore, considering the essential characteristics of video stream transmission, this application adopts the idea of ​​encoder-decoder collaborative design. For the video stream sender (encoder), by reducing the frame resolution of the video stream and reducing the pixel color bit depth, the compression efficiency of the video stream sender is improved, thereby significantly reducing the video transmission bit rate. At the same time, for the video stream receiver (decoder), a diffusion recovery model is constructed, and the model is made to perceive the resolution and color conditions of the video stream, thereby providing a stronger visual quality recovery capability. Achieve an ideal balance between bit rate saving and quality recovery.

[0071] See also Figure 1 , Figure 1 The video stream transmission system 100 of the embodiment of the present application comprises: a video stream transmitter 101 and a video stream receiver 102, wherein the video stream transmitter 101 and the video stream receiver 102 are connected in communication.

[0072] like Figure 1 In the embodiment, the video stream sending end 101 may include an encoder (Encoder) for compressing the original high-definition video to obtain a compressed video. The compression method may include reducing the frame resolution and reducing the pixel color bit depth. The encoder is responsible for executing and sending the compressed video to the video stream receiving end 102; the video stream receiving end 102 may include a decoder (Decoder) for receiving the compressed video sent by the video stream sending end 101, and enhancing the video quality of the compressed video through the diffusion recovery model constructed in the decoder to restore it to a high-definition video. The encoder (Encoder) in the video stream sending end 101 and the decoder (Decoder) in the video stream receiving end 102 work in collaboration.

[0073] It is understandable that a video stream transmission system provided by an embodiment of the present application designs an encoder at the sending end with the goal of minimizing transmission traffic and minimizing transmission delay, thereby optimizing the performance of network transmission. At the same time, a diffusion recovery model is deployed at the receiving end to enhance the video quality recovery capability of the decoder and improve the visual viewing experience of the video. Through the synchronous design and optimization of the encoder and decoder, the core performance indicator of the video stream transmission service, the "rate-distortion trade-off", is significantly improved, that is, the video stream compression rate at the sending end is maximized (thereby minimizing the transmission traffic demand and saving communication bandwidth), while improving the quality of the restored video at the receiving end, striving to achieve a visual effect similar to that of directly transmitting the original high-definition video.

[0074] It should be noted that in the benchmark test of a video stream transmission system provided by this application using public cloud services, compared with traditional video streaming standards (such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC), it can effectively save up to 5 - 22 times the bit rate, and the recovery quality is also significantly better than existing neural enhancement methods.

[0075] The compression method of the video stream by the video stream sender 101 and the recovery method of the video stream by the video stream receiver 102 in the video stream transmission system 100 are specifically described as follows.

[0076] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a video stream transmission method provided by an embodiment of this application. By way of example and not limitation, this method can be applied to a terminal device, such as a video stream sender, and the video stream sender can be a terminal device such as a smartphone or a smart TV. The encoder in the video stream sender can compress the high-definition video to obtain a low-quality compressed video. This video stream transmission method includes:

[0077] S11. Obtain the original video data, where the original video data includes multiple original video frames.

[0078] S12. According to preset compression conditions, perform compression processing on the multiple original video frames to obtain video compression data corresponding to the original video data, where the compression processing includes reducing the frame resolution processing and reducing the pixel color bit depth processing.

[0079] S13. Send the video compression data corresponding to the original video data to the video stream receiver so that the video stream receiver performs recovery processing on the video compression data to obtain video recovery data corresponding to the video compression data.

[0080] The original video data refers to the original video file without any compression or processing, and can include a series of consecutive original video frames. The original video frame is a single static image in the original video data and is the smallest time unit that constitutes the video. For example, a video with 20 frames per second contains 20 original video frames per second. Each original video frame can contain complete pixel information (such as RGB values, brightness, chrominance, etc.) and timestamp information, and the complete picture and timing of the video can be restored through the pixel information and timestamp information.

[0081] Video compression data is the video data obtained after the original video data has been compressed. Compression processing can be understood as a process of reducing the amount of video data through an algorithm in order to reduce the bandwidth and storage space requirements for video transmission. In this embodiment, the compression processing may include reducing the frame resolution processing and reducing the pixel color bit depth processing. Reducing the frame resolution processing reduces the video quality by reducing the number of pixels in each frame of the video, thereby reducing the size of the video. For example, reducing a video frame with a resolution of 3840×2160 to a video frame with a resolution of 1920×1080. Reducing the pixel color bit depth processing reduces the video quality by reducing the precision of the color information of each pixel, thereby reducing the size of the video. For example, reducing the RGB value of each pixel from 24 bits to 8 bits. The preset compression conditions are a series of compression rules preset when compressing the original video data. Such as the scaling factor set when performing the reducing frame resolution processing.

[0082] The video stream receiving end is a terminal device that receives the video compression data sent by the video stream sending end and decodes the video compression data. The video stream receiving end can be a terminal device such as a smart phone or a smart TV.

[0083] Recovery processing can be understood as a process of decoding the video compression data and restoring it to a quality close to that of the original video data. The recovery processing may include methods such as decompression, denoising processing, and interpolation processing. Video recovery data is the video data obtained after the recovery processing. This video recovery data is very close to the original video data, but there may also be pixel details that cannot be fully restored due to compression losses.

[0084] Specifically, after the video stream sending end obtains the original video data containing multiple original video frames, it performs compression processing at two levels: reducing the frame resolution processing and reducing the pixel color bit depth processing, reducing the number of pixels in each frame and the bit depth of the pixel color, thereby obtaining the video compression data; after obtaining the video compression data, it sends the video compression data through network transmission (such as the Internet, local area network) to the video stream receiving end, so that the video stream receiving end performs recovery processing on the video compression data and restores it to obtain the video recovery data, completing the efficient transmission of the video stream.

[0085] It can be understood that the embodiment of the present application provides a video stream transmission method. The video stream sender obtains original video data, where the original video data includes multiple original video frames; compresses the multiple original video frames according to preset compression conditions to obtain video compression data corresponding to the original video data, where the compression process includes reducing the frame resolution process and reducing the pixel color bit depth process; sends the video compression data corresponding to the original video data to the video stream receiver, so that the video stream receiver performs a recovery process on the video compression data to obtain video recovery data corresponding to the video compression data. By performing compression processing on the original video data at two levels of reducing the frame resolution and reducing the pixel color bit depth, the compression efficiency of the video stream sender is improved, thereby significantly reducing the transmission bit rate of the video.

[0086] In a possible implementation manner, the preset compression conditions include a first preset compression condition and a second preset compression condition.

[0087] Compressing the multiple original video frames according to the preset compression conditions to obtain video compression data corresponding to the original video data includes:

[0088] Performing a frame resolution reduction process on the multiple original video frames according to the first preset compression condition to obtain first video compression data, where the first video compression data includes multiple first compressed video frames.

[0089] Performing a pixel color bit depth reduction process on the multiple first compressed video frames according to the second preset compression condition to obtain the video compression data corresponding to the original video data.

[0090] The preset compression conditions may include a first preset compression condition and a second preset compression condition. The first preset compression condition and the second preset compression condition respectively correspond to different compression processes. The first preset compression condition may be compression parameters corresponding to the frame resolution reduction process, such as a scaling factor. The second preset compression condition may be compression parameters corresponding to the pixel color bit depth reduction process, such as a target bit depth.

[0091] The first video compression data is intermediate compression data obtained after reducing the frame resolution of multiple original video frames in the original video data. Correspondingly, the first video compression data contains multiple first compressed video frames (i.e., video frames after reducing the frame resolution). The video compression data is the final compression data obtained after reducing the pixel color bit depth of the first compressed video frames in the first video compression data, or is the final compression data obtained after reducing the frame resolution and reducing the pixel color bit depth of the original video data.

[0092] Further, performing a frame resolution reduction process on the multiple original video frames according to the first preset compression condition to obtain first video compression data includes:

[0093] According to a preset scaling factor, each original video frame in the original video data is divided into a plurality of initial macroblocks.

[0094] Based on the Gaussian blur algorithm, the features at the junctions of the plurality of initial macroblocks are smoothed to obtain a plurality of processed macroblocks corresponding to the plurality of initial macroblocks.

[0095] Calculate the weighted average value of all pixels within each processed macroblock in each original video frame, and compress each processed macroblock to the weighted average value to obtain a first compressed video frame corresponding to each original video frame.

[0096] Arrange the plurality of first compressed video frames in the order of the plurality of original video frames in the original video data to obtain first video compression data.

[0097] Specifically, first, a preset scaling factor s is preset. The encoder in the video stream sender can divide the original video frame into a series of initial macroblocks with an area of s×s. Among them, the preset scaling factor s can be understood as a scaling ratio parameter preset when reducing the frame resolution of the original video frame. The initial macroblock is the smallest processing unit after the original video frame is divided based on the preset scaling factor. In this embodiment, the preset scaling factor s can be set to s = 2, s = 4, and the specific value of the preset scaling factor s is not limited. Generally, the preset scaling factor is not set too large to avoid excessive compression of the video.

[0098] After dividing the original video frame into a series of initial macroblocks with an area of s×s, the features at the junctions of the initial macroblocks are smoothed by the Gaussian blur algorithm to reduce the block effect that may occur after compression, thereby obtaining the processed macroblocks after processing. Among them, the Gaussian blur algorithm refers to a smoothing filter algorithm based on the Gaussian function. Through the Gaussian blur algorithm, the features at the junctions of the original macroblocks can be smoothed to reduce block effects such as false edges and mosaics that may occur after compression. The processed macroblock is the macroblock obtained after processing the initial macroblock by the Gaussian blur algorithm. Compared with the original macroblock, the features at its junction are smoothed, and the transition between each macroblock is more natural.

[0099] After obtaining the processed macroblocks, within each processed macroblock, calculate the weighted average of all the pixels therein, and shrink the processed macroblock to this weighted average to achieve compression of the processed macroblock, thereby obtaining the first compressed video frame. At this time, the original video frame is reduced to 1 / s of its original size in both the width dimension and the height dimension. The weighted average can be understood as the value obtained by weighted summing and averaging all the pixels according to their respective weights, and this weighted average can represent the color information of the entire processed macroblock. The first compressed video frame is a single video frame obtained after the original video frame undergoes macroblock segmentation, smoothing processing, and compression, and its resolution is reduced according to the preset scaling factor s.

[0100] Finally, arrange all the first compressed video frames in the order of the original video frames in the original video data to obtain the first video compression data. At this time, the first video compression data has a reduced resolution compared to the original video data, but the original color bit depth has not been changed at this time.

[0101] In one example, assume that the original video frame has a resolution of 1920×1080 and the preset scaling factor is 2. Then the original video frame is divided into 960×540 initial macroblocks with an area of 2×2 pixels. Perform Gaussian blur algorithm processing on each initial macroblock, then calculate the weighted average of the pixels in each processed macroblock, and shrink the processed macroblock to this weighted average to obtain the first compressed video frame with a resolution of 960×540. The first compressed video frame is reduced by 4 times compared to the original video frame. Then, arrange all the first compressed video frames in the order of the original video frames to obtain the first video compression data.

[0102] It should be understood that as the number of pixels in the width and height dimensions of the video frame decreases, compared with the original video frame, the resolution can achieve a frame size reduction of s squared times.

[0103] Furthermore, perform pixel color bit depth reduction processing on multiple first compressed video frames according to the second preset compression condition to obtain the video compression data corresponding to the original video data, including:

[0104] Perform vector quantization processing on the color values of all the pixels in multiple first compressed video frames to obtain pixel vectors corresponding to all the pixels in multiple first compressed video frames.

[0105] Perform clustering analysis on all the pixel vectors in multiple first compressed video frames, and convert the full-bit-depth pixels in multiple first compressed video frames into low-bit-depth pixels to obtain second compressed video frames corresponding to multiple first compressed video frames.

[0106] Arrange multiple second compressed video frames in the order of multiple original video frames in the original video data to obtain the video compression data corresponding to the original video data.

[0107] For common video standards, video frames are usually organized in 8-bit color space and RGB channels, and the color space of pixels in the frame is represented by a total of 24 bits. When describing visually close colors, pixel vectors along the RGB channels have similar distributions. Therefore, this embodiment reduces the bit depth of video pixel colors from the perspective of vector quantization, clusters all pixels according to the numerical value of the color, and converts full-bit-depth pixels into low-bit-depth pixels.

[0108] Vectorization can be understood as the process of converting the color value of a pixel into a vector form. Cluster analysis is an exploratory analysis method, such as K-means clustering algorithm, hierarchical clustering, etc., which classifies the research object based on the characteristics of the research object, so that the differences between individuals in the same category are relatively small and the similarities are relatively large, and the differences between individuals in different categories are large and the similarities are small. In this embodiment, cluster analysis is used to group similar pixel vectors, that is, to divide the color space into multiple clusters, each cluster corresponding to a representative color.

[0109] Full-bit depth pixels are pixels that use all available binary bits to represent color information, such as 24-bit color space, which has rich colors and smooth transitions, but the amount of data is large. Low-bit depth pixels are pixels that use fewer binary bits to represent color information, such as 16-bit color space, which has a smaller amount of data and may cause color discontinuities.

[0110] Specifically, first, all the first compressed video frames in the first video compression data are traversed, and the color value of each pixel in the first compressed video frame is extracted, and the color value of each pixel is converted into a vector form, so as to obtain the pixel vectors corresponding to all pixels. Then, all the pixel vectors in each video frame are clustered and analyzed, and all the pixel vectors are divided into multiple clusters (each cluster corresponds to a color), and the color value of each pixel is replaced with the representative color of the cluster to which it belongs, so as to realize the conversion of full-bit depth pixels to low-bit depth pixels, that is, to obtain the second compressed video frame. Finally, all the second compressed video frames are arranged in the order of the original video frames to obtain the final video compression data. At this time, the frame resolution of the video compression data is reduced and the color bit depth is reduced relative to the original video data.

[0111] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0112] See also Figure 3 , Figure 3It is a schematic flowchart of a video stream transmission method provided by an embodiment of the present application. As an example rather than a limitation, this method can be applied to a terminal device, such as a video stream receiving end, and the video stream receiving end can be a terminal device such as a smart phone or a smart TV. The video stream receiving end is a terminal device that receives the video compression data sent by the video stream sending end and decodes the video compression data through a decoder to obtain video restoration data. The method includes:

[0113] S21. Receive the video compression data sent by the video stream sending end, where the video compression data is obtained by the video stream sending end performing compression processing on multiple original video frames in the video original data according to a preset compression condition, and the video compression data includes multiple compressed video frames.

[0114] S22. Perform resolution restoration processing on multiple compressed video frames through fast bilinear interpolation to obtain first restored video frames corresponding to the multiple compressed video frames.

[0115] S23. Perform splicing processing on the multiple first restored video frames and the initial Gaussian noise in the channel dimension to obtain a splicing result.

[0116] S24. Perform image restoration processing on the splicing result through a diffusion restoration model to obtain video restoration data corresponding to the video compression data; where multiple restoration parameters in the diffusion restoration model are optimized according to the preset compression condition in the video stream sending end.

[0117] In step S21, the video compression data is the video data sent by the video stream sending end, and the video compression data is obtained by the video sending end performing compression processing (such as reducing the frame resolution and reducing the pixel color bit depth) on the video original data according to a preset compression condition.

[0118] In step S22, fast bilinear interpolation is an image interpolation algorithm that can be used to restore the resolution of compressed video frames. Through fast bilinear interpolation, the size of the compressed video frame can be enlarged, but the details and colors of the image are not restored. The first restored video frame is the video frame after the compressed video frame is processed by fast bilinear interpolation, and the resolution of the first restored video frame is higher than that of the compressed video frame.

[0119] In step S23, the initial Gaussian noise is a random noise whose probability distribution follows a Gaussian function. During the process of image restoration of the video compression data, the initial Gaussian noise is added to the first restored video frame to simulate the information lost during the compression process. The result of splicing and combining the first restored video frame and the initial Gaussian noise in the channel dimension is the splicing result. When the initial Gaussian noise simulates the information lost during compression, the simulation direction is guided by adding the first restored video frame.

[0120] In step S24, the diffusion recovery model is a deep learning-based image recovery model. By simulating the noise diffusion and denoising processes, it attempts to minimize the gap between the predicted noise and the real noise, and gradually restore the compressed image to a clear image. Among them, the recovery parameters are the parameters optimized in the diffusion recovery model according to preset compression conditions (such as preset scaling factors, bit depths, etc.). The recovery parameters can control the intensity and details of the recovery process.

[0121] Specifically, first, the video stream receiving end receives the video compression data sent by the video stream sending end. The decoder in the video stream receiving end captures and extracts all the compressed video frames. Then, through fast bilinear interpolation, multiple compressed video frames are calculated to improve the resolution of the compressed video, and the first recovered video frame is obtained. Then, the decoder randomly generates initial Gaussian noise, and the first recovered video frame and the initial Gaussian noise are concatenated in the channel dimension as the generation seed to obtain a concatenation result. Finally, the concatenation result is input into the diffusion recovery model, and the diffusion recovery model uses the concatenation result to perform image recovery processing to restore the compressed video frame to a high-quality recovered video frame, thereby obtaining video recovery data.

[0122] It can be understood that the embodiment of the present application provides a video stream transmission method. The video stream receiving end receives the video compression data sent by the video stream sending end, where the video compression data includes multiple compressed video frames. Through fast bilinear interpolation, the resolution of multiple compressed video frames is restored to obtain the first recovered video frames corresponding to the multiple compressed video frames. The multiple first recovered video frames and the initial Gaussian noise are concatenated in the channel dimension to obtain a concatenation result. Through the diffusion recovery model, the concatenation result is subjected to image recovery processing to obtain the video recovery data corresponding to the video compression data. Among them, the multiple recovery parameters in the diffusion recovery model are optimized according to the preset compression conditions in the video stream sending end. The video stream receiving end can perceive the compression conditions of the resolution and color of the video stream through the constructed diffusion recovery model, thereby providing a stronger visual quality recovery ability. Furthermore, while maximizing the video stream compression rate of the video stream sending end, the recovery quality of the video stream at the video stream receiving end can be improved.

[0123] In a possible implementation manner, performing image recovery processing on the concatenation result through the diffusion recovery model to obtain the video recovery data corresponding to the video compression data includes:

[0124] Performing image recovery processing on the concatenation result through the diffusion recovery model to generate the second recovered video frames corresponding to the multiple first recovered video frames.

[0125] Arranging the multiple second recovered video frames in the order of the multiple compressed video frames in the video compression data to obtain the video recovery data corresponding to the video compression data.

[0126] The second restored video frame is the video frame obtained after the quality of the stitching result is restored by the diffusion restoration model. Based on the first restored video frame, the image quality of the video frame is further improved, the details are richer, and the noise is less. The video restoration data is the final video data obtained by arranging all the second restored video frames in the order of all the compressed video frames in the video compression data. The image quality of the video restoration data is close to that of the original video data.

[0127] Specifically, first, the diffusion restoration model is used to gradually denoise the stitching result to restore the image quality and generate the second restored video frame; then, all the second restored video frames are arranged in the order of all the compressed video frames in the video compression data, and the arranged frame sequence is saved as the video restoration data.

[0128] In a possible implementation manner, before the image of the stitching result is restored by the diffusion restoration model to obtain the video restoration data corresponding to the video compression data, the video stream transmission method includes:

[0129] Obtain a training data set, where the training data set includes a plurality of original video frames and the compressed video frames corresponding to the plurality of original video frames.

[0130] Construct an initial diffusion restoration model, and initialize a plurality of restoration parameters in the initial diffusion restoration model through a normal distribution function.

[0131] According to a preset attenuation coefficient and a preset time step, gradually add Gaussian noise to the plurality of original video frames to generate the compressed video frames corresponding to the plurality of original video frames.

[0132] Based on the current state of the compressed video frame, the noise level, and the preset compression condition in the video stream sender, obtain the predicted noise in the compressed video frame, and remove the predicted noise in the compressed video frame to obtain the original video frame corresponding to the compressed video frame.

[0133] Use the gradient descent algorithm to minimize the difference value between the predicted noise and the Gaussian noise, and optimize a plurality of restoration parameters in the initial diffusion restoration model to obtain the diffusion restoration model.

[0134] Before image restoration by the diffusion restoration model, it is necessary to construct an initial diffusion restoration model and train the initial diffusion restoration model to obtain the final diffusion restoration model. In this embodiment, the training of the initial diffusion restoration model can adopt a dual mode of forward diffusion and reverse diffusion. Therefore, as an enhancement module, the diffusion restoration model can restore the video quality by perceiving the resolution-color compression of the encoder in the video stream sender. The decoder in the video stream receiver limits the computational complexity of the diffusion restoration model to adapt to the video streaming media environment.

[0135] It should be noted that the training of the initial diffusion recovery model simulates the degradation process of data from high quality to low quality (i.e., noise addition) through a forward diffusion process. This process is to generate a series of intermediate state data, which will be used in the subsequent reverse diffusion process. The forward diffusion process includes two aspects: gradual noise addition and generation of degradation states. Specifically: (1) Gradual noise addition: Starting from the original high-quality video frames (i.e., the original video frames), Gaussian noise is gradually added to them. The amount of noise added at each step is controlled by a preset attenuation coefficient (or called noise level). As the number of steps increases, the amount of added noise also gradually increases until the original video frames are completely degraded into pure noise. (2) Generation of degradation states: After adding noise at each step, a new video frame in a degraded state is generated. These degraded states form a degradation sequence from high quality to low quality (i.e., compressed video frames).

[0136] The reverse diffusion process is the core of training the initial diffusion recovery model, aiming to learn how to gradually restore the original high-quality video frames (original video frames) from the degraded noisy frames (compressed video frames). The reverse diffusion process can include three aspects: noise prediction and removal, model optimization, and gradual restoration. Specifically: (1) Noise prediction and removal: At each step of the reverse diffusion process, the diffusion recovery model needs to predict the noise contained in the current state (i.e., the video frame at the current noise level) and remove it. This prediction process is based on the state of the current frame, the noise level, and a conditional signal (including the preset compression conditions of the encoder, such as preset scaling factors and color bit depth information). (2) Model optimization: The mean squared error (MSE) is used as the loss function, that is, the difference between the noise predicted by the model and the actually added noise is calculated, and the difference is minimized through a gradient descent algorithm (such as the Adam optimizer) to optimize multiple recovery parameters of the model. Among them, the gradient descent algorithm is an optimization algorithm that can minimize the difference value between the noise predicted by the model and the actually added noise. (3) Gradual restoration: By repeatedly performing the above noise prediction and removal process, the diffusion recovery model can gradually restore the degraded noisy frames to the original high-quality video frames.

[0137] The specific process of training the initial diffusion recovery model is as follows:

[0138] (1) Obtain the training dataset: Prepare a dataset containing original high-quality video frames and corresponding compressed low-quality video frames, that is, multiple original video frames and multiple compressed video frames corresponding to the original video frames. These data will be used to train the diffusion recovery model.

[0139] (2) Model Initialization: Construct an initial diffusion recovery model and initialize multiple recovery parameters of the diffusion model using the normal distribution function. These recovery parameters determine the behavior of the model during noise prediction and removal. Among them, the initial diffusion recovery model is an untrained diffusion recovery model, and its recovery parameters can be initialized using the normal distribution function.

[0140] (3) Iterative Training: In each iteration, perform the forward diffusion process to generate video frames in the degraded state, that is, gradually add Gaussian noise to multiple original video frames according to the preset attenuation coefficient and preset time steps to generate compressed video frames corresponding to the multiple original video frames; then perform the reverse diffusion process to train the model, that is, based on the current state of the compressed video frame, the noise level, and the preset compression conditions in the video stream sender, obtain the predicted noise in the compressed video frame, and remove the predicted noise in the compressed video frame to obtain the original video frame corresponding to the compressed video frame, and update the recovery parameters of the model through the loss function, thereby obtaining the trained diffusion recovery model. This process is repeated until the model converges, that is, the value of the loss function no longer decreases significantly. Through the above training process, the diffusion recovery model can learn how to recover high-quality video content from compressed low-quality video frames.

[0141] Among them, the preset attenuation coefficient is a parameter preset to control the attenuation speed of the noise amount during the noise addition process. The preset attenuation coefficient can determine the intensity of noise addition and the time step, etc. The preset time step is the number of time steps preset during the noise addition process, which can determine the fineness of noise addition and the total amount of noise. The predicted noise is the noise in the compressed video frame predicted by the diffusion recovery model based on the current state of the compressed video frame, the noise level, and the preset compression conditions in the video stream sender.

[0142] It should be noted that after obtaining the trained diffusion recovery model, it is also necessary to evaluate and adjust the diffusion recovery model. During the training process, regularly evaluate the performance of the model on the validation set. According to the evaluation results, adjust the hyperparameters of the model (such as the learning rate, noise level, etc.) to prevent overfitting and improve the generalization ability of the model.

[0143] It should be noted that in order to reduce the computational complexity and improve the inference speed, the diffusion recovery model can be compressed. That is, remove unimportant neurons or connections in the diffusion recovery model, reduce the number of parameters and computational amount of the diffusion recovery model, thereby reducing the computational cost of inference while maintaining high performance.

[0144] Therefore, the process of recovering high-quality video content from compressed low-quality video frames through the diffusion recovery model can be as follows:

[0145] (1) Initialization: Sample initial Gaussian noise from an isotropic Gaussian distribution as the generation seed.

[0146] (2) Conditional concatenation: Concatenate the upsampled compressed video frames and the initial Gaussian noise along the channel dimension, and use the resulting concatenated result as the input to the diffusion recovery model.

[0147] (3) Step-by-step denoising: Gradually remove the noise through a multi-step reverse diffusion process to recover high-quality video frames. Here, the reverse diffusion process is usually between 10 steps and 100 steps, which can be adjusted according to the difficulty of image recovery. Increase the number of steps of the reverse diffusion process when image recovery is difficult, and decrease the number of steps of the reverse diffusion process when image recovery is easy.

[0148] (4) Iterative optimization: In each step of the reverse diffusion process, use the diffusion recovery model to predict and remove the noise, and obtain the final high-quality video frames through multiple iterations.

[0149] It should be noted that a video stream transmission method provided by an embodiment of the present application has the following advantages:

[0150] (1) By combining diffusion enhancement technology, the compression efficiency and recovery quality in the video stream are significantly improved, breaking the limitations of the traditional neural video stream paradigm and bringing new breakthroughs to the technical development in this field.

[0151] (2) Co-optimization of the encoder and decoder: This method not only focuses on the compression efficiency of the encoder in the video stream sending end, but also perceives the compression conditions of the encoder through the diffusion recovery process of the decoder in the video stream receiving end, realizing the co-optimization of the encoder and decoder, and effectively solving the problem of the lack of active cooperation between the encoder and decoder in existing methods.

[0152] (3) High-efficiency compression and recovery performance: The test results show that while maintaining high perceptual quality, this method achieves a nearly one-order-of-magnitude improvement in the compression ratio, that is, under the same bandwidth conditions, higher-quality video content can be transmitted through this video stream transmission method. At the same time, compared with previous methods, this method has achieved a significant improvement in the recovery quality. For example, the Video Multi-method Assessment Fusion (VMAF) score reaches above 93, and the Mean Opinion Score (MOS) has increased by 23%. It shows that this method has obvious advantages in recovering video details and overall perceptual quality.

[0153] (4) Wide applicability and compatibility. As a general video stream enhancement transmission technology (compress first, then transmit, and recover after receiving), this method can be compatible with existing video codecs (such as H.265, VP9, and AV1), enabling this method to be used as a plug-and-play module to upgrade existing video stream transmission methods.

[0154] (5) This method takes into account the real-time requirements in the video stream. By optimizing the inference process of the diffusion recovery model, the time cost for recovering each frame is significantly reduced, enabling it to meet the needs of real-time video streams. This design fully considers the problem of limited resources on the client side (used to receive compressed video and perform recovery), and makes full use of the powerful computing capabilities of the media server.

[0155] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0156] A video stream transmission method corresponding to the above embodiments Figure 4 The schematic structural diagram of a video stream sending end provided by an embodiment of this application is shown. For the sake of convenience of description, only the parts related to the embodiments of this application are shown.

[0157] Refer to Figure 4 , the video stream sending end 101 of this embodiment includes:

[0158] An acquisition module 31, configured to acquire video original data, where the video original data includes a plurality of original video frames.

[0159] A compression module 32, configured to perform compression processing on the plurality of original video frames according to preset compression conditions to obtain video compression data corresponding to the video original data, where the compression processing includes reducing frame resolution processing and reducing pixel color bit depth processing.

[0160] A sending module 33, configured to send the video compression data corresponding to the video original data to the video stream receiving end, so that the video stream receiving end performs recovery processing on the video compression data to obtain video recovery data corresponding to the video compression data.

[0161] It can be understood that in this embodiment, the video stream sending end 101 obtains the original video data through the obtaining module 31. Among them, the original video data includes a plurality of original video frames; the compression module 32 performs compression processing on the plurality of original video frames according to the preset compression conditions to obtain the video compression data corresponding to the original video data. Among them, the compression processing includes reducing the frame resolution processing and reducing the pixel color bit depth processing; the sending module 33 sends the video compression data corresponding to the original video data to the video stream receiving end, so that the video stream receiving end performs restoration processing on the video compression data to obtain the video restoration data corresponding to the video compression data. By performing compression processing on the original video data at two levels of reducing the frame resolution and reducing the pixel color bit depth, the compression efficiency of the video stream sending end is improved, thereby significantly reducing the transmission bit rate of the video.

[0162] Further, the preset compression conditions include the first preset compression condition and the second preset compression condition, and the compression module 32 includes:

[0163] The first compression sub-module is used to perform frame resolution reduction processing on the plurality of original video frames according to the first preset compression condition to obtain the first video compression data. Among them, the first video compression data includes a plurality of first compressed video frames.

[0164] The second compression sub-module is used to perform pixel color bit depth reduction processing on the plurality of first compressed video frames according to the second preset compression condition to obtain the video compression data corresponding to the original video data.

[0165] Further, the first compression sub-module includes:

[0166] The splitting unit is used to split each original video frame in the original video data into a plurality of initial macroblocks according to the preset scaling factor.

[0167] The processing unit is used to smooth the features at the junctions of the plurality of initial macroblocks based on the Gaussian blur algorithm to obtain a plurality of processed macroblocks corresponding to the plurality of initial macroblocks.

[0168] The compression unit is used to calculate the weighted average value of all pixels in each processed macroblock in each original video frame, and compress each processed macroblock to the weighted average value to obtain the first compressed video frame corresponding to each original video frame.

[0169] The first sorting unit is used to arrange the plurality of first compressed video frames in the order of the plurality of original video frames in the original video data to obtain the first video compression data.

[0170] Further, the second compression sub-module includes:

[0171] A quantization unit is configured to perform vector quantization on the color values of all pixels in multiple first compressed video frames to obtain pixel vectors corresponding to all pixels in the multiple first compressed video frames.

[0172] A clustering unit is configured to perform clustering analysis on all pixel vectors in multiple first compressed video frames, convert full-bit-depth pixels in the multiple first compressed video frames into low-bit-depth pixels, and obtain second compressed video frames corresponding to the multiple first compressed video frames.

[0173] A second sorting unit is configured to sort multiple second compressed video frames in the order of multiple original video frames in the original video data to obtain video compressed data corresponding to the original video data.

[0174] It should be noted that for the information interaction, execution process, etc. among the modules in the above video stream sender 101, since they are based on the same concept as the method embodiment of this application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.

[0175] Corresponding to a video stream transmission method in another embodiment above, Figure 5 FIG. shows a schematic structural diagram of a video stream receiver provided by an embodiment of this application. For ease of description, only parts related to the embodiment of this application are shown.

[0176] Referring to Figure 5 , the video stream receiver 102 of this embodiment includes:

[0177] A video receiving module 41 is configured to receive video compressed data sent by a video stream sender. The video compressed data is obtained by the video stream sender compressing multiple original video frames in the original video data according to preset compression conditions, and the video compressed data includes multiple compressed video frames.

[0178] A first restoration module 42 is configured to perform resolution restoration processing on multiple compressed video frames through fast bilinear interpolation to obtain first restored video frames corresponding to the multiple compressed video frames.

[0179] A splicing processing module 43 is configured to perform splicing processing on multiple first restored video frames and initial Gaussian noise in the channel dimension to obtain a splicing result.

[0180] A second restoration module 44 is configured to perform image restoration processing on the splicing result through a diffusion restoration model to obtain video restored data corresponding to the video compressed data; multiple restoration parameters in the diffusion restoration model are optimized according to the preset compression conditions in the video stream sender.

[0181] It can be understood that in this embodiment, the video stream receiving end 102 receives the video compression data sent by the video stream sending end 101 through the video receiving module 41. Among them, the video compression data includes multiple compressed video frames; the first restoration module 42 performs resolution restoration processing on the multiple compressed video frames through fast bilinear interpolation to obtain the first restored video frames corresponding to the multiple compressed video frames; the splicing processing module 43 performs splicing processing on the multiple first restored video frames and the initial Gaussian noise in the channel dimension to obtain a splicing result; the second restoration module 44 performs image restoration processing on the splicing result through a diffusion restoration model to obtain the video restoration data corresponding to the video compression data; among them, the multiple restoration parameters in the diffusion restoration model are optimized according to the preset compression conditions in the video stream sending end. The video stream receiving end 102 perceives the compression conditions of the resolution and color of the video stream through the constructed diffusion restoration model, so as to provide a stronger visual quality restoration ability. Furthermore, while maximizing the video stream compression rate of the video stream sending end, the restoration quality of the video stream at the video stream receiving end can be improved.

[0182] Further, the second restoration module 44 includes:

[0183] A restoration unit, configured to perform image restoration processing on the splicing result through a diffusion restoration model to generate second restored video frames corresponding to the multiple first restored video frames.

[0184] A third arrangement unit, configured to arrange the multiple second restored video frames in the order of the multiple compressed video frames in the video compression data to obtain the video restoration data corresponding to the video compression data.

[0185] Further, the video stream receiving end 102 includes:

[0186] A sample acquisition module, configured to acquire a training data set, where the training data set includes multiple original video frames and compressed video frames corresponding to the multiple original video frames.

[0187] A model construction module, configured to construct an initial diffusion restoration model and initialize multiple restoration parameters in the initial diffusion restoration model through a normal distribution function.

[0188] A forward diffusion module, configured to gradually add Gaussian noise to the multiple original video frames according to a preset attenuation coefficient and preset time steps to generate compressed video frames corresponding to the multiple original video frames.

[0189] A reverse diffusion module, configured to obtain the predicted noise in the compressed video frame based on the current state of the compressed video frame, the noise level, and the preset compression conditions in the video stream sending end, and remove the predicted noise in the compressed video frame to obtain the original video frame corresponding to the compressed video frame.

[0190] The model optimization module is used to optimize multiple recovery parameters in the initial diffusion recovery model by using the gradient descent algorithm to minimize the difference value between the predicted noise and the Gaussian noise, so as to obtain the diffusion recovery model.

[0191] It should be noted that the information interaction, execution process, etc. among the modules in the above video stream receiving end 102 are based on the same concept as another method embodiment of the present application. For the specific functions and the technical effects brought about, please refer to the method embodiment part, and details will not be repeated here.

[0192] The embodiment of the present application also provides a terminal device, such as Figure 6 shown Figure 6 is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Referring to Figure 6 , the terminal device 6 of this embodiment includes: a memory 61, a processor 62, and a computer program stored in the memory 61 and executable on the processor 62. When the processor 62 executes the computer program, the steps in the above-mentioned various method embodiments are implemented.

[0193] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0194] The embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal is enabled to execute the steps in the above-mentioned various method embodiments.

[0195] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, USB flash drive, mobile hard disk, magnetic disk or optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0196] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0197] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0198] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0199] The unit described as a separation component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0200] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A video stream transmission method, applied to a video stream sender, characterized in that, Including: Obtain the original video data, where the original video data includes a plurality of original video frames; According to preset compression conditions, perform compression processing on the plurality of original video frames to obtain video compression data corresponding to the original video data, where the compression processing includes reducing frame resolution processing and reducing pixel color bit depth processing; Send the video compression data corresponding to the original video data to a video stream receiving end, so that the video stream receiving end performs restoration processing on the video compression data to obtain video restoration data corresponding to the video compression data.

2. The video stream transmission method according to claim 1, characterized in that, The preset compression conditions include a first preset compression condition and a second preset compression condition, The performing compression processing on the plurality of original video frames according to preset compression conditions to obtain video compression data corresponding to the original video data includes: Perform reducing frame resolution processing on the plurality of original video frames according to the first preset compression condition to obtain first video compression data, where the first video compression data includes a plurality of first compressed video frames; Perform reducing pixel color bit depth processing on the plurality of first compressed video frames according to the second preset compression condition to obtain the video compression data corresponding to the original video data.

3. The video stream transmission method according to claim 2, wherein The performing reducing frame resolution processing on the plurality of original video frames according to the first preset compression condition to obtain first video compression data includes: According to a preset scaling factor, divide each original video frame in the original video data into a plurality of initial macroblocks; Based on a Gaussian blur algorithm, perform smoothing processing on the features at the junctions of the plurality of initial macroblocks to obtain a plurality of processed macroblocks corresponding to the plurality of initial macroblocks; Calculate the weighted average value of all pixels in each processed macroblock in each original video frame, and compress each processed macroblock to the weighted average value to obtain the first compressed video frame corresponding to each original video frame; Arrange the plurality of first compressed video frames in the order of the plurality of original video frames in the original video data to obtain the first video compression data.

4. The video stream transmission method according to claim 3, wherein The performing reducing pixel color bit depth processing on the plurality of first compressed video frames according to the second preset compression condition to obtain the video compression data corresponding to the original video data includes: Perform vector quantization processing on the color values of all pixels in the plurality of first compressed video frames to obtain pixel vectors corresponding to all pixels in the plurality of first compressed video frames; Perform clustering analysis on all the pixel vectors in the plurality of first compressed video frames, and convert full-bit-depth pixels in the plurality of first compressed video frames into low-bit-depth pixels to obtain second compressed video frames corresponding to the plurality of first compressed video frames; Arrange the plurality of second compressed video frames in the order of the plurality of original video frames in the original video data to obtain the video compression data corresponding to the original video data.

5. A video stream transmission method, applied to a video stream receiving end, characterized in that, Including: Receive the video compression data sent by the video stream sender, where the video compression data is obtained by the video stream sender through compressing multiple original video frames in the original video data according to preset compression conditions, and the video compression data includes multiple compressed video frames; Perform resolution restoration processing on multiple compressed video frames through fast bilinear interpolation to obtain multiple first restored video frames corresponding to the compressed video frames; Perform splicing processing on multiple first restored video frames and initial Gaussian noise in the channel dimension to obtain a splicing result; Perform image restoration processing on the splicing result through a diffusion restoration model to obtain video restoration data corresponding to the video compression data; where multiple restoration parameters in the diffusion restoration model are optimized according to the preset compression conditions in the video stream sender.

6. The video stream transmission method according to claim 5, characterized in that The performing image restoration processing on the splicing result through a diffusion restoration model to obtain video restoration data corresponding to the video compression data includes: Perform image restoration processing on the splicing result through the diffusion restoration model to generate multiple second restored video frames corresponding to the first restored video frames; Arrange multiple second restored video frames in the order of multiple compressed video frames in the video compression data to obtain the video restoration data corresponding to the video compression data.

7. The video stream transmission method according to claim 6, characterized in that, Before performing image restoration processing on the splicing result through the diffusion restoration model to obtain video restoration data corresponding to the video compression data, the method includes: Obtain a training data set, where the training data set includes multiple original video frames and compressed video frames corresponding to the multiple original video frames; Construct an initial diffusion restoration model and initialize multiple restoration parameters in the initial diffusion restoration model through a normal distribution function; Gradually add Gaussian noise to multiple original video frames according to a preset attenuation coefficient and preset time steps to generate compressed video frames corresponding to the multiple original video frames; Based on the current state of the compressed video frame, the noise level, and the preset compression conditions in the video stream sender, obtain the predicted noise in the compressed video frame, and remove the predicted noise in the compressed video frame to obtain the original video frame corresponding to the compressed video frame; Optimize multiple restoration parameters in the initial diffusion restoration model through a gradient descent algorithm to minimize the difference value between the predicted noise and the Gaussian noise, and obtain the diffusion restoration model.

8. A video stream transmission system, characterized in that, Includes: A video stream sender and a video stream receiver, where There is a communication connection between the video stream sender and the video stream receiver; The video stream sender executes the video stream transmission method according to any one of claims 1-4; the video stream receiver executes the video stream transmission method according to any one of claims 5-7.

9. A video stream sender, characterized in that Includes: An acquisition module for acquiring original video data, where the original video data includes multiple original video frames; A compression module, configured to perform compression processing on a plurality of the original video frames according to preset compression conditions to obtain video compression data corresponding to the original video data, wherein the compression processing includes reducing frame resolution processing and reducing pixel color bit depth processing; A sending module, configured to send the video compression data corresponding to the original video data to a video stream receiving end, so that the video stream receiving end performs recovery processing on the video compression data to obtain video recovery data corresponding to the video compression data.

10. A video stream receiving end, characterized in that, Comprising: A video receiving module, configured to receive video compression data sent by a video stream sending end, wherein the video compression data is obtained by the video stream sending end performing compression processing on a plurality of original video frames in the original video data, and the video compression data includes a plurality of compressed video frames; A first recovery module, configured to perform resolution recovery processing on a plurality of the compressed video frames through fast bilinear interpolation to obtain a plurality of first recovered video frames corresponding to the compressed video frames; A splicing processing module, configured to perform splicing processing on a plurality of the first recovered video frames and initial Gaussian noise in a channel dimension to obtain a splicing result; A second recovery module, configured to perform image recovery processing on the splicing result through a diffusion recovery model to obtain video recovery data corresponding to the video compression data; wherein a plurality of recovery parameters in the diffusion recovery model are optimized according to the preset compression conditions in the video stream sending end.

Citation Information

Patent Citations

  • Video data transmission method and system based on deep learning model, medium and equipment

    CN112565777A

  • Image restoration method and device

    CN116805290A