Video live broadcast system and method

Through the encoding and decoding technology of lightweight video rescaling neural network, an efficient encoding and decoding solution is generated, which solves the contradiction between bandwidth and computing resources in high-definition video live broadcast, and realizes high-definition video live broadcast under low bandwidth.

CN120281907APending Publication Date: 2025-07-08XIAYAN TECH (SHANDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410029599.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-06
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing live video broadcast technology is difficult to maintain high-definition image quality while taking into account the needs of low bandwidth transmission and low computing resources, resulting in the problem of playback delay and decoding difficulties.

Method used

A lightweight video rescaling neural network is adopted to generate high-resolution and low-resolution video image sequence pairs, initial training is performed and convolutional layer bias terms are pruned, and configured as a source-side encoder and a receiving-side decoder to achieve efficient encoding and decoding.

Benefits of technology

On the basis of maintaining high-definition picture quality, bandwidth requirements and computing resources are reduced, real-time live broadcast of high-definition videos is realized, and user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281907A_ABST
    Figure CN120281907A_ABST
Patent Text Reader

Abstract

The invention provides a video live broadcast method, which comprises the following steps of: acquiring input image data from an acquisition module, making a sequence pair of an input video image and a low-resolution video image, constructing a lightweight video re-scaling neural network based on the sequence pair, and performing initial training; pruning the network after the initial training is completed through a feature map mask mechanism, and closing a convolution layer bias item; performing fine tuning training on the network of which the convolutional layer bias terms are closed, dividing the network after fine tuning training into a source end neural network encoder network and a receiving end neural network decoder network, configuring the source end neural network encoder network to a corresponding source end, and configuring the receiving end neural network decoder network to a corresponding source end; and the receiving end neural network decoder network configures a receiving end.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of video encoding and decoding technology, and more particularly to a system and method for live video broadcasting. Background Art

[0002] With the development of Internet technology, live video broadcasting has become more and more popular, which has led to various user demands for live broadcasting services, especially requirements for clarity and latency. As a result, high-definition video live broadcasting technology is also facing some challenges, which are mainly reflected in the following two aspects: First, bandwidth requirements. Compared with traditional video live broadcasting technology, the transmission / reception of high-definition video (such as 8K video) data requires a larger bandwidth, and low-bandwidth transmission and high-quality video reproduction must be taken into account. Since the current mainstream video live broadcasting technology cannot take into account low-bandwidth transmission, it can only maintain picture quality at the cost of increasing transmission bandwidth, making high-definition video live broadcasting difficult; second, processing and storage requirements. In traditional technologies, processing and storing high-definition videos requires larger computing resources and computing power, which makes the live broadcasting system (such as live broadcasting platform, etc.) bear greater pressure on video stream processing and distribution. At the same time, since most clients are limited by their own computing power and bandwidth, playback delays and decoding failures may occur, thereby reducing the efficiency of video live broadcasting. Summary of the invention

[0003] A solution of the present invention is completed in view of the above situation, and its purpose is to provide a system and method for live video broadcasting to efficiently encode and decode videos, so as to achieve real-time live broadcast of high-definition videos with lower bandwidth while maintaining the original picture quality.

[0004] In one embodiment of the present disclosure, a method for live video broadcast is provided, the method comprising the following steps: acquiring input image data from an acquisition module; generating a sequence pair of a high-resolution video image and a low-resolution video image; constructing a lightweight video rescaling neural network based on the sequence pair, and performing initial training; pruning the network after the initial training through a feature map mask mechanism, and turning off the convolutional layer bias item; performing fine-tuning training on the network with the convolutional layer bias item turned off; dividing the network after the fine-tuning training into a source-end neural network encoder network and a receiving-end neural network decoder network, and configuring the source-end neural network encoder network at the corresponding source end, and configuring the receiving-end neural network decoder network at the receiving end.

[0005] In yet another embodiment of the present disclosure, a video live streaming system is provided. The system includes at least one source end, at least one transmission channel, and at least one receiving end, and causes the system to perform the following steps: obtain input image data from the source end; generate a sequence pair of high-resolution video images and low-resolution video images; based on the sequence pair, construct a lightweight video rescaling neural network and perform initial training; through a feature map masking mechanism, prune the network after the initial training is completed, and turn off the convolutional layer bias term; for the network with the convolutional layer bias term turned off, perform fine-tuning training; divide the network after the fine-tuning training is completed into a source-end neural network encoder network and a receiving-end neural network decoder network, and configure the source-end neural network encoder network at the corresponding source end, and configure the receiving-end neural network decoder network at the receiving end.

[0006] According to one aspect of the present invention, the video live streaming system and method provided by the embodiments of the present disclosure can perform efficient encoding and decoding of videos, so as to achieve real-time live streaming of high-definition videos with a relatively low bandwidth while maintaining the clarity of the original picture quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] When reading the following detailed description in conjunction with the accompanying drawings, various aspects of the present disclosure can be best understood. The various features are not drawn to scale, and for the sake of clear discussion, the sizes of various features can be arbitrarily increased or decreased. Moreover, it is obvious that the drawings in the following description are only embodiments of the present disclosure, and those skilled in the art can obtain other drawings according to the provided drawings without creative efforts. Figure 1 It is a schematic structural diagram of a video live streaming system provided by some embodiments of the present disclosure. Figure 2 It is a schematic structural diagram of the source end in a video live streaming system provided by some embodiments of the present disclosure. Figure 3 It is a schematic structural diagram of the receiving end in a video live streaming system provided by some embodiments of the present disclosure. Figure 4 It is a diagram showing an example of the schematic configuration of the information compression module at the source end in a video live streaming system provided by some embodiments of the present disclosure. Figure 5 It is a diagram showing an example of the schematic configuration of the information extraction module at the receiving end in a video live streaming system provided by some embodiments of the present disclosure. Figure 6 It is a diagram showing an example of the schematic configuration of the encoder module at the source end in a video live streaming system provided by some embodiments of the present disclosure. Figure 7A diagram showing an example of the schematic composition of the super-resolution reconstruction module at the receiving end in the video live streaming system provided in some embodiments of the present disclosure. Figure 8 A schematic diagram of the architecture of the lightweight video rescaling network in the video live streaming system provided in some embodiments of the present disclosure. Figure 9 A flowchart of the method for video live streaming provided in some embodiments of the present disclosure. Detailed implementation manners

[0008] The following disclosure includes specific information related to the embodiments in the present disclosure. The accompanying drawings and the corresponding detailed disclosure are directed to exemplary embodiments.

[0009] However, the present disclosure is not limited to these exemplary embodiments. Those skilled in the art will think of other variations and embodiments of the present disclosure.

[0010] Unless otherwise indicated, similar or corresponding elements in the drawings may be indicated by similar or corresponding reference numerals. The drawings and illustrations are generally not drawn to scale and are not intended to correspond to actual relative sizes.

[0011] For the purposes of consistency and ease of understanding, in the exemplary drawings, similar features are identified by reference numerals (although not shown in some examples). However, the features in different embodiments may vary in other aspects and should not be narrowly limited to what is shown in the drawings.

[0012] It should be noted that "at least one" in the present disclosure means one or more, and "a plurality" means two or more than two. "And / or" describes the relationship between associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural.

[0013] In the embodiments of the present disclosure, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present disclosure should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0014] Those skilled in the art can understand that the "source end" and "receiving end" in this disclosure include not only devices with video transceiver capabilities, which are devices with video transceivers that do not have transmission capabilities, but also devices with receiving and transmitting hardware, which are devices with receiving and transmitting hardware that can perform two-way transmission on a two-way transmission channel. Such devices may include: communication devices such as personal computers and tablet computers, which have single-line displays or multi-line displays or cellular or other communication devices without multi-line displays; personal communication systems (PCS), which can combine voice, data processing, fax, and / or data communication capabilities; personal digital assistants (PDAs), which may include radio frequency receivers, pagers, Internet / intranet access, web browsers, notepads, calendars, and / or global positioning system (GPS) receivers; traditional laptop and / or palm-held computers or other devices, which are traditional laptop and / or palm-held computers or other devices with and / or including radio frequency receivers. The "source end" and "receiving end" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to operate locally and / or in a distributed manner at any other location on the earth and / or in space. The "source end" and "receiving end" used herein can also be communication terminals, Internet access terminals, music / video playback terminals, such as PDAs, mobile Internet devices (MIDs), and / or mobile phones with music / video playback functions, or can also be devices such as smart TVs and set-top boxes.

[0015] The hardware referred to by names such as the "source end" and "receiving end" in this disclosure is essentially an electronic device with the equivalent capabilities of a personal computer, which is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0016] It should be noted that the "source end" and "video transmission medium" referred to in this disclosure can also be referred to as the "server end", and this "server end" can similarly be extended to apply to the case of a server cluster. According to the network deployment principle understood by those skilled in the art, the various servers should be logically divided. Physically, these servers can either be independent of each other but can be called through an interface, or can be integrated into a physical computer or a set of computer clusters. Those skilled in the art should understand this variation and should not be restricted by this when implementing the network deployment method of this disclosure.

[0017] To solve the above problems, the embodiments of this disclosure provide a video live broadcast system and method, which can perform efficient encoding and decoding of videos, so as to achieve real-time live broadcast of high-definition videos with a lower bandwidth while maintaining the clarity of the original picture quality.

[0018] Figure 1 The structural schematic diagram of the video live broadcast system 100 provided by some embodiments of this disclosure is shown. In Figure 1 it, the video live broadcast system 100 may include a source end 110, a receiving end 130, and a video transmission medium 120.

[0019] The source end 110 includes any device configured to encode video data by obtaining data information from a collection device and transmit the encoded video data to the video transmission medium 120. The receiving end 130 can be a destination device, and the receiving end 130 includes any device configured to receive the encoded video data via the video transmission medium 120 and decode the encoded video data.

[0020] The video transmission medium 120 can also be a video transmission channel (Video Transmission Channel, VTC), and the video transmission channel can include a computer system interface capable of storing a compatible video bitstream on a storage device or receiving a compatible video bitstream from a storage device. For example: the video transmission channel can include a chipset that supports the Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, proprietary bus protocols, Universal Serial Bus (USB) protocols, I2C, or any other logical and physical structure that can be used to interconnect peer devices.

[0021] The source end 110 may include a collection module and a display module. The receiving end 130 may include an information extraction module, a super-resolution reconstruction module, and a display module. The source end 110 may be a video encoder, and the receiving end 130 may be a video decoder.

[0022] The source end 110 and / or the receiving end 130 may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, or other electronic devices. Figure 1 An example of the source end 110 and the receiving end 130 is shown. The source end 110 and the receiving end 130 may include more or fewer components than those shown, or have different configurations of various components.

[0023] As Figure 1 shown, the collection module 111 may include a video capture device for capturing new videos, a video archive for storing previously captured videos, and / or a video feed interface for receiving videos from a video content provider. The collection module 111 may generate computer graphics-based data as the source video, or generate a combination of live videos, archived videos, and computer-generated videos as the source video. The video capture device may be a charge-coupled device (CCD) image sensor, a complementary metal-oxide-semiconductor (CMOS) image sensor, a camera, or the like.

[0024] Figure 2Schematic diagram of the structure of the source end 110 in the video live broadcast system provided by some embodiments of the present disclosure. The source end 110 further includes an encoder module 112 and an information compression module 113. Each of the encoder module 112 and the information compression module 113 can be implemented as any one of a variety of suitable encoder / decoder circuits, such as one or more microprocessors, a central processing unit (CPU), a graphic processing unit (GPU), a system on chip (SoC), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the disclosed methods. In at least some embodiments, each of the encoder module 112 and the information compression module 113 can be included in one or more encoders and decoders, each of which can be integrated as part of a combined encoder / decoder (CODEC) in the device.

[0025] Figure 3Schematic diagram of the structure of the receiving end 130 in the video live broadcast system provided by some embodiments of the present disclosure. The receiving end 130 includes: an information extraction module 131, a super-resolution reconstruction module 132, and a display module. The information extraction module 131 and the super-resolution reconstruction module 132 can each be implemented as any one of a variety of suitable encoder / decoder circuits, such as one or more microprocessors, central processing unit (CPU), graphic processing unit (GPU), system on chip (SoC), digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array (FPGA), discrete logic, software, hardware, firmware, or any combination thereof. When implemented in part in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the disclosed methods. In at least some embodiments, each of the information extraction module and the super-resolution reconstruction module can be included in one or more encoders and decoders, each of which can be integrated as part of a combined encoder / decoder (CODEC) in the device.

[0026] Reference Figures 1 to 3 , it should be noted that the display module may include a display, which uses liquid crystal display (LCD) technology, plasma display technology, organic light emitting diode (OLED) display technology, or light emitting polymer display (LPD) technology, and in other embodiments, other display technologies are used. The display module may include a high-definition display or an ultra-high-definition display. In this application, for the sake of simplicity, such a display module is not explicitly illustrated.

[0027] Figure 4A diagram showing a schematic configuration of an information compression module 113 of a source end 110 in a video live broadcast system provided by some embodiments of the present disclosure. The information compression module 113 may include an image encoder (Encoder_1), a statistical information encoder (Encoder_2), a subtractor, and a multiplexer (Muxer). The information compression module 113 receives video data and encodes the video data to output the encoded video, and the video data may be a bitstream or a data stream. The image encoder and the statistical information encoder may encode the image and the statistical information through a video compression method, where the video compression method may be a video image encoding and decoding method specified in standards such as Motion-Join Photographic Experts Group (M-JPEG), H.265 / High Efficiency Video Coding (HEVC), AVS2, AV1, etc. An intra-frame encoding method or an inter-frame encoding method may be used, which is not limited herein. And the encoded data streams are respectively output, and through a multiplexer for integrating the data streams output by the two encoders, the integrated compressed information I (in formats such as MP4, AVI, etc.) is output C , so as to achieve more efficient transmission.

[0028] In some embodiments of this case, the image encoder may be a 4K image encoder with a 4K resolution, or any one of image encoders with a resolution lower than 8K. When the image encoder is a 4K image encoder, the information compression module 113 may use the M-JPEG method to perform 4K image encoding on a part of the information in the downsampled image I LR obtained by the encoder module 112, and at the same time integrate another part of the information in the downsampled image I LR with the statistical information S obtained by the encoder module 112, and integrate the statistical information S obtained by the encoder module 112 with the information in the downsampled image I LR through a subtractor and input it into the statistical information encoder, and further integrate the data streams output from the 4K image encoder and the statistical information encoder through a multiplexer.

[0029] Figure 4 This is only an example because there are multiple ways to perform encoding on images, videos, or information data. Figure 4 Each box shown in may represent one or more processes, methods, or subroutines to be executed.

[0030] Figure 4The order of the boxes in [description] is merely illustrative and may be changed. Additional boxes may be added or fewer boxes may be used without departing from the present disclosure.

[0031] Figure 5 FIG. [description] is a diagram showing an example of a schematic configuration of an information extraction module 131 of a receiver 130 in a video live streaming system provided for some embodiments of the present disclosure. The information extraction module 131 may include an image decoder (Decoder_1), a statistical information decoder (Decoder_2), an adder, and a demultiplexer (DeMuxer). The information extraction module 131 receives the compressed video data and decodes the video data to output the decoded video. The video data may be a bitstream or a data stream. The demultiplexer receives the compressed information I in a specific format C , and decouples (or de-couples) the compressed information, and then inputs the decoupled information into the image decoder and the statistical information decoder respectively for decoding. After the decoding is completed, a part of the video data of the decoded image decoder and the video data of the statistical information decoder are input into the adder, integrated by the adder, and the integrated data and another part of the video data of the decoded image decoder are respectively input into a super-resolution reconstruction module 132 to perform video reconstruction.

[0032] In some embodiments of the present case, the image decoder may be a 4K image decoder with a 4K resolution or any one of image decoders with a resolution lower than 8K. When the image decoder is a 4K image decoder, the information extraction module 131 may transmit the compressed information decoupled by the demultiplexer to the 4K image decoder and the statistical information decoder respectively to perform the decoding process. Among them, the decoding process may be an M-JPEG decoding process.

[0033] See Figure 4 and Figure 5 , it should be noted that the subtractor and the adder may be filters respectively. The filters may include a deblocking filter, a sample adaptive offset (SAO) filter, a bilateral filter, and / or an adaptive loop filter (ALF) to remove block artifacts from the reconstructed block. In addition to the deblocking filter, SAO filter, bilateral filter, and ALF, other filters (intra-loop or post-loop) may also be used. For the sake of brevity, such filters are not explicitly shown.

[0034] Figure 6A diagram showing a schematic configuration of an encoder module 112 of a source end 110 in a video live streaming system provided in some embodiments of the present disclosure. In some embodiments of the present disclosure, the encoder module 112 may include a data transformation unit 141, a channel attention unit 142, and a plurality of feature extraction units. Among them, the data transformation unit 141 may perform transforms such as discrete cosine transform (DCT), discrete sine transform (DST), adaptive multiple transform (AMT), mode-dependent non-separable secondary transform (MDNSST), hypercube-givens transform (HyGT), signal-dependent transform, Karhunen-Loéve transform (KLT), wavelet transform, integer transform, subband transform, or conceptually similar transforms.

[0035] In some embodiments of the present disclosure, as Figure 6 shown, the data transformation unit 141 may be a discrete wavelet transform unit U DWT , and the encoder module 112 may include a discrete wavelet transform unit U DWT , a channel attention unit U CA , and at least three feature extraction units U FE . Specifically, the discrete wavelet transform unit U DWT may include a discrete wavelet transform layer (Discrete Wavelet transform layer, DWT layer) L DWT and at least one convolution layer (convolution layer) L C0 . First, the input image may be defined as where H, W, and C respectively represent the height, width, and number of channels of each video frame feature map, and the input image I HR may be a video image with high resolution, such as: 8K video image or a video image with a resolution higher than 8K, or may also be any high-definition video frame or high-definition still image other than 8K. Then, the input image I HR is input to the discrete wavelet transform layer for multi-level transformation, and the generated image data is downsampled to obtain a feature map, and the feature map is generated into a low-frequency feature F Low and a high-frequency feature F HighTwo parts. Among them, in some embodiments of the present disclosure, a bicubic interpolation method can be used to generate a downsampled low-resolution image by a factor of two, such as a 4K image. In addition, in some embodiments of the present disclosure, a bilinear interpolation method can also be used for downsampling processing. The specific implementation manner of downsampling is not limited in this application. At the same time, the downsampling processing manner of the present disclosure can implement the algorithm through applications such as Python, Ruby, MATLAB, Simulink, Stateflow, Visual Basic, JavaScript, or a combination thereof, but this application is not limited to the above applications. Some of these applications can be compiled and executed on virtual machines such as the Java virtual machine and the Dalvik virtual machine.

[0036] Further, in some embodiments of the present disclosure, the image body information of the feature map can be generated by a first-level discrete wavelet transform. The image body information may include information on regions in the input image where the brightness or grayscale value changes slowly. The image body information is set as the low-frequency feature, that is, Next, the detailed information such as the edge contour and texture in the feature map can be set as the high-frequency feature, that is, The high-frequency feature may include a horizontal high-frequency feature, a vertical high-frequency feature, and a diagonal high-frequency feature. Among them, the horizontal high-frequency feature can be defined as The vertical high-frequency feature can be defined as The diagonal high-frequency feature can be defined as

[0037] In some embodiments of the present disclosure, a Convolutional Neural Networks (CNN) can be used in a system for video live streaming. The CNN can include a convolution layer, a pooling layer, and a full-connection layer. Among them, multiple convolution layers can be alternately arranged with multiple pooling layers. After the convolution layer, it can be followed by another convolution layer or a pooling layer. The convolution layer can be used to perform a convolution operation on an input matrix composed of pixel data of an image to be recognized, so as to extract features of the image to be recognized. The pooling layer can be used immediately after the convolution layer to perform a pooling operation on the output matrix of the previous layer. Common processing methods include max-pooling, mean-Pooling, Global Average Pooling (GlobalAvgPool), and Stochastic-pooling. The pooling layer can also be called a subsampling layer to sample the features extracted by the convolution layer, and it can also be considered as simplifying the output matrix of the convolution layer.

[0038] In addition, the convolution layer can include a non-linear layer, a linear layer, or an activation function layer, such as a Sigmoid activation layer. The non-linear layer or the linear layer can be connected after the convolution layer to perform non-linear or linear mapping processing on the information output by the convolution layer. And / or, the non-linear layer or the linear layer is connected after the pooling layer to perform non-linear or linear mapping processing on the information output by the pooling layer.

[0039] In some embodiments of the present disclosure, the convolution layer L C0 can include at least one 2D convolution layer of a 3×3 matrix and at least one Sigmoid activation layer. By initially compressing the low-frequency feature F Low the scale of the feature map can be reduced and the number of channels of the feature map can be expanded to 33.

[0040] Furthermore, as Figure 6 shown, in some embodiments of the present disclosure, the horizontal high-frequency feature F High can be obtained from the high-frequency feature F H , the vertical high-frequency feature F V and the diagonal high-frequency feature F D . The obtained horizontal high-frequency feature F H , vertical high-frequency feature F V and diagonal high-frequency feature F D are respectively input into the channel attention unit U CA142, wherein the channel attention unit U CA 142 may include at least one global average pooling layer, at least one linear layer, and at least one Sigmoid activation layer. In some embodiments of the present disclosure, first, after vector concatenation of the horizontal high-frequency feature F H , vertical high-frequency feature F V and diagonal high-frequency feature F D , a global average pooling operation is performed. Specifically, the global average pooling operation can be defined as follows: M = GlobalAvgPool(Concatenate(F H , F V , F D )); wherein, Concatenate() represents the vector concatenation operator; GlobalAvgPool() represents the global average pooling operation; represents the feature value on each channel, that is, the average value of the feature map corresponding to this channel. Then, the feature value corresponding to each channel (i.e., the above M feature value) is input into the linear layer, and the process of calculating the weights of each channel is performed by using Gradient Descent (GD). In some embodiments of the present disclosure, this process can be defined as follows: w = σ(W·M + b); wherein, W represents the learnable weight matrix, b is the bias term, and σ is the Sigmoid activation function, that is, by multiplying the feature of each channel by the corresponding weight coefficient, the statistical information S can be obtained. Specifically, after vector concatenation of the horizontal high-frequency feature F H , vertical high-frequency feature F V and diagonal high-frequency feature F D , an addition operation is performed with the information obtained after multiplying the feature of each channel by the corresponding weight coefficient to obtain the statistical information S. In some embodiments of the present disclosure, this process can be defined as follows: wherein, Add() represents the element-wise addition operator; represents the element-wise multiplication.

[0041] In addition, the above channel attention unit may include network layers such as a pooling layer, a linear layer, an activation function layer, and a dimension expansion. The number of each layer included in the channel attention unit can be set to a suitable size according to needs, and the present application does not limit this.

[0042] Such as Figure 6As shown, in some embodiments of the present disclosure, the encoder module includes three of the feature extraction units U FE , where the feature extraction unit U FE may include at least two 5×5 2D convolutional layers and at least one Relu activation layer, and a residual connection can be introduced in each of the three feature extraction units U FE to avoid gradient degradation or gradient vanishing. The image features obtained after being processed by the convolutional layer L C0 , that is, the low-frequency feature F Low are sequentially input into the three feature extraction units U FE to further extract deep features to obtain more abstract information. Then, the deep features of the low-resolution video image can be further downsampled to obtain a downsampled low-resolution image, and this image is the compressed image, and the compressed image can be defined as In addition, the low-resolution video image can be a 4K video image or a video image with a resolution lower than 8K video image.

[0043] Then, a vector concatenation operation is performed on the compressed image I LR and the statistical information S, and the obtained concatenated vector is input into the super-resolution reconstruction module 132. Among them, the vector concatenation operation can be a concatenation method such as horizontal concatenation or vertical concatenation, and the present application does not limit this concatenation method.

[0044] Figure 7 FIG. is an example diagram of the schematic composition of the super-resolution reconstruction module 132 of the receiving end 130 in the video live broadcast system provided by some embodiments of the present disclosure. In some embodiments of the present disclosure, the super-resolution reconstruction module 132 includes at least one feature extraction unit U FE and at least one image reconstruction unit U RC . In the super-resolution reconstruction module 132, the feature extraction unit U FE may include at least two 5×5 2D convolutional layers and a Relu activation layer, as Figure 6 and Figure 7 shown. A residual connection can be introduced in each of the feature extraction units in the super-resolution reconstruction module 132 to avoid gradient degradation or gradient vanishing. The image reconstruction unit U RC may include at least one 3×3 2D convolutional layer, at least one Sigmoid activation layer, and at least one transposed convolution (Transpose Convolution, Transpose Conv) layer L TC .

[0045] In some embodiments of the present disclosure, the transposed convolutional layer can also be a de-convolution layer, or an inverse layer of a convolutional layer, or an inverse convolutional layer.

[0046] As Figure 7 shown, in an embodiment of the present disclosure, the vector concatenation feature information obtained after performing a vector concatenation operation on the downsampled image I LR and the statistical information S is sequentially input into Figure 7 the feature extraction unit U FE therein to further extract shallow features and / or deep features. Further, the obtained image feature information is input into the image reconstruction unit U RC therein. In the image reconstruction unit U RC the image feature information is subjected to transposed convolution processing. Specifically, the transposed convolution processing can be defined as follows: I′ HR = L TC (σ(Conv2D(U FE (Concatenate(I LR , S))))); wherein, Concatenate() represents a vector concatenation operator; σ is a Sigmoid activation function; thereby obtaining a reconstructed image, which can be defined as

[0047] Figure 8 is a schematic diagram of the architecture of a lightweight video rescaling network in the video live broadcast system provided by some embodiments of the present disclosure. In some embodiments of the present disclosure, the lightweight video rescaling neural network can be configured to generate a network model through the encoder module 112 and the super-resolution reconstruction module 132. The lightweight video rescaling neural network is constructed by using a training set, and the lightweight video rescaling neural network can be initially trained using the number of global feature channels for the model.

[0048] Specifically, in some embodiments of the present disclosure, the constructed model can be initially trained using the number of global feature channels. Except that the output channel number of the last convolutional layer of the two modules can be set to 3, the output channel numbers of the remaining convolutional layers can be fixed to 33 during the initial training.

[0049] As Figure 8 shown, in some embodiments of the present disclosure, the initial training of the network can be supervised by using the mean squared error (MSE). Specifically, the supervision processing can be defined as: Among them, f() represents a low-frequency processing operation, g() represents a high-frequency processing operation, specifically expressed as h() represents a super-resolution reconstruction operation, specifically expressed as follows:

[0050] Figure 9 It is a flowchart of the video live streaming method provided by some embodiments of the present disclosure. The video live streaming method in this embodiment can be applied to Figure 1 and 8 the video live streaming system shown in. Among them, the video live streaming method includes the following steps:

[0051] S100, generate a sequence pair of high-resolution video images and low-resolution video images.

[0052] In some embodiments of the present disclosure, as Figures 1 - 4 shown, the original data of the input video or image can be obtained through the acquisition module. The input video or image can be a high-resolution video image conforming to the H.265 / HEVC standard, such as: 8K video image. The original data of the input video or image can be sliced into single-frame pictures, and the bicubic interpolation method can be used to generate a two-fold downsampled low-resolution video image, such as: 4K video image, so that the downsampled 4K video image frames and 8K video image frames form a sequence pair. At the same time, the generated sequence pair (or dataset) can be divided into a training set, a validation set, and a test set according to a ratio of 6:2:2 for training, validating, and testing the network respectively. In some embodiments of the present disclosure, the sequence pair may also include information such as reference image frames and / or matching coding blocks and / or inter-frame prediction markers.

[0053] S200, construct a lightweight video rescaling neural network and perform initial training.

[0054] As Figure 8 shown, in some embodiments of the present disclosure, the training set can be used to construct a lightweight video rescaling neural network. Among them, the lightweight video rescaling neural network at least includes an encoder module 112 and a super-resolution reconstruction module 132, and the constructed network model is initially trained using the global feature channel number. Thus, encoding compression and super-resolution reconstruction can be trained simultaneously. Since the input image video data contains information on the main features of the image and statistical information on the image detail features, it is possible to ensure that the source image retains as much of the original image information as possible after being compressed, and at the same time ensure the quality of the reconstructed video, thereby reducing the requirement for transmission bandwidth in video transmission.

[0055] S300, prune the network after the initial training is completed through a feature map masking mechanism, and turn off the bias terms of the convolutional layers.

[0056] In some embodiments of the present disclosure, based on the feature map masking mechanism, pruning operations are performed on the convolutional layers of the network model after the initial training is completed by using the method of matching pursuit.

[0057] Specifically, in some embodiments of the present disclosure, all possible masks can be determined for the convolutional layer to be pruned. In other words, by masking any one channel in the feature map each time, different masks are obtained. Then, based on all the obtained masks, by applying any one of all the masks, the Peak Signal-to-Noise Ratio (PSNR) of the network model on the validation set is calculated to measure the image reconstruction quality. Among them, the mask that has the least impact on the PSNR is used as the final mask for the corresponding convolutional layer. After the final mask is established, the mask of this convolutional layer is fixed, and the remaining masks in the convolutional layer that have not been established are determined in the same way, repeating this process until the number of channels in all convolutional layers is reduced to 32. After the number of channels in the network model is reduced to 32, the bias terms of all convolutional layers are turned off, and the network model after the bias terms are turned off is fine-tuned on the dataset, thereby ensuring the accuracy of the converted network model, enabling the constructed lightweight rescaling neural network model to be fully optimized, reducing the complexity and computational amount of the network model on the premise of ensuring video image compression and reconstruction quality, so that the model can achieve better reconstruction quality with lower computing power, and thus highly restoring the true texture details of the source-end video image by performing real-time encoding and decoding processing, enabling users to obtain a high-quality experience.

[0058] In some embodiments of the present disclosure, the dataset may include a training set, a validation set, and a test set, or may include at least one of the training set, the validation set, and the test set for constructing the network model.

[0059] S400, divide the network after the fine-tuning training is completed into a source-end neural network encoder (Source Neural Network Encoder, SNNE) and a receiver-end neural network decoder (Receiver Neural Network Decoder, RNND), and configure them on the source end and the receiver end respectively.

[0060] In some embodiments of the present disclosure, the network after the fine-tuning training can be allocated. Specifically, as Figures 1 - 3As shown, the acquisition module, encoder module, information compression module, and video transmission medium can be configured as SNNE; the video transmission medium, information extraction module, and super-resolution reconstruction module can be configured as RNND. In other embodiments of the present disclosure, at least the encoder module and the information compression module can also be configured as SNNE; at least the information extraction module and the super-resolution reconstruction module can also be configured as RNND. Then, the configured SNNE is used as the source end, and the configured RNND is used as the receiving end.

[0061] To verify the effect of the present invention, the performance of the proposed network model was confirmed and compared through experiments. Hereinafter, Table 1 shows the comparison results of the image quality of one embodiment of the present disclosure and three other network model structures.

[0062] Table 1

[0063] Specifically, the above embodiment in Table 1 uses the public datasets DIV2K and Urban100. At the same time, the public datasets DIV2K and Urban100 are used as the training set. During the training process, MSE is used as the loss function, and the number of iterations is set to 800,000 times. Set5, Set14, and BSD100 are respectively general public datasets. Set5, Set14, and BSD100 are used as the test sets. The specific test results of the trained network tested on the Set5, Set14, and BSD100 datasets respectively (as shown in Table 1) are quantitatively compared with the test results of three advanced models FSRCNN, EDSR, and SwinIR in Table 1 in terms of three performance evaluation indicators: PSNR, Structural Similarity (SSIM), and Average Run time (ART). Based on the experimental results of this embodiment above, in this embodiment, on 8 RTX4090 graphics cards, the single-frame image processing time for RNND to reconstruct a 4K video into an 8K video is about 16 ms. This further proves that the present invention is superior to existing methods in terms of performance evaluation indicators and improves the video reconstruction effect.

[0064] Furthermore, various solutions of the video live broadcast system and method of the embodiment of the present disclosure are described.

[0065] In an embodiment of the present disclosure, a method for video live streaming is provided. The method includes the following steps: obtaining input image data from an acquisition module, creating a sequence pair of an input video image and a low-resolution video image, constructing a lightweight video rescaling neural network based on the sequence pair, and performing initial training; pruning the network after the initial training is completed through a feature map masking mechanism, and turning off the convolutional layer bias term; performing fine-tuning training on the network with the convolutional layer bias term turned off, dividing the network after the fine-tuning training is completed into a source-side neural network encoder network and a receiver-side neural network decoder network, and configuring the source-side neural network encoder network at the corresponding source side, and configuring the receiver-side neural network decoder network at the receiver side.

[0066] In another embodiment of the present disclosure, a system for video live streaming is provided. The system includes at least one source side, at least one transmission channel, and at least one receiver side, and causes the system to perform the following steps: obtaining input image data from the source side, creating a sequence pair of an input video image and a low-resolution video image, constructing a lightweight video rescaling neural network based on the sequence pair, and performing initial training; pruning the network after the initial training is completed through a feature map masking mechanism, and turning off the convolutional layer bias term; performing fine-tuning training on the network with the convolutional layer bias term turned off, dividing the network after the fine-tuning training is completed into a source-side neural network encoder network and a receiver-side neural network decoder network, and configuring the source-side neural network encoder network at the source side, and configuring the receiver-side neural network decoder network at the receiver side.

[0067] As described above, the embodiments of the present invention have been described in detail with reference to the accompanying drawings. However, the specific configuration is not limited to this embodiment, and also includes design changes and the like within the scope not departing from the gist of the present invention. In addition, one aspect of the present invention can be variously changed within the scope shown in the technical solution, and embodiments obtained by appropriately combining the technical solutions disclosed in different embodiments are also included in the technical scope of the present invention. In addition, it also includes a configuration obtained by mutually replacing elements having the same effect among the elements described in the above respective embodiments.

[0068] Industrial applicability

[0069] One aspect of the present invention can be used, for example, in a video live streaming system, a transmission device (such as a portable mobile phone, etc.), an integrated circuit (such as an image processing chip), or a program, etc.

Claims

1. A method for video live broadcast, characterized in that, The method includes the following steps: Obtain input image data from the acquisition module; Generate a sequence pair of high-resolution video images and low-resolution video images; Based on the sequence pair, construct a lightweight video rescaling neural network and perform initial training; Through the feature map masking mechanism, prune the network after the initial training is completed, and turn off the convolutional layer bias term; Perform fine-tuning training on the network with the convolutional layer bias term turned off; Divide the network after the fine-tuning training is completed into a source-side neural network encoder network and a receiving-side neural network decoder network, and configure the source-side neural network encoder network on the corresponding source side; configure the receiving-side neural network decoder network on the corresponding receiving side.

2. The method according to claim 1, wherein The method further includes: Slice the original data of the input image into single-frame pictures; Based on all the sliced single-frame pictures, generate downsampled low-resolution video image frames through downsampling operations using bicubic interpolation; Generate a sequence pair based on the low-resolution video image frames and the high-resolution video image frames and use the sequence pair for training.

3. The method according to claim 1, wherein The method further includes: The lightweight video rescaling neural network is initially trained using the global number of feature channels.

4. The method according to claim 1, wherein The method further includes: Perform a pruning operation on the network after the initial training is completed through matching pursuit.

5. The method according to claim 1, characterized in that, The method further includes: Perform a pruning operation on the network after the initial training is completed through matching pursuit until the number of channels in the convolutional layers of the network is reduced to 32, and turn off the convolutional layer bias term; After the number of channels is reduced to 32, turn off the bias terms of all convolutional layers and perform fine-tuning training on the network with the bias terms turned off.

6. A video live broadcast system, characterized in that, The system includes: At least one source side, at least one transmission channel, and at least one receiving side, enabling the system to perform the following steps: Obtain input image data from the source side; Generate a sequence pair of high-resolution video images and low-resolution video images; Based on the sequence pair, construct a lightweight video rescaling neural network and perform initial training; Through the feature map masking mechanism, prune the network after the initial training is completed, and turn off the convolutional layer bias term; Perform fine-tuning training on the network with the convolutional layer bias term turned off; Divide the network after the fine-tuning training is completed into a source-side neural network encoder network and a receiving-side neural network decoder network, and configure the source-side neural network encoder network on the corresponding source side, and configure the receiving-side neural network decoder network on the receiving side.

7. The system according to claim 6, wherein The source side at least includes an acquisition module, an encoder module, and an information compression module; wherein The acquisition module acquires the original data of the video image and sends the original data to the encoder module and the information compression module to encode the sent information based on a video compression method; The receiving side at least includes an information extraction module and a super-resolution reconstruction module; Wherein The information extraction module receives the compressed and encoded video data, decodes the video data, and outputs the decoded video image data to the super-resolution reconstruction module.

8. The system according to claim 7, wherein: The encoder module includes a data transformation unit, a channel attention unit, and a plurality of feature extraction units.

9. The system according to claim 8, wherein: The number of the plurality of feature extraction units is at least three.

10. A system for video live broadcast, wherein The system includes one or more modules configured to execute all combinations of claims 1 to 5.