Video processing method and system

The super-resolution neural network is built through the three-dimensional sparse attention mechanism and self-similarity coefficient, which solves the problem of device resolution limiting of high-resolution video transmission and playback, and realizes effective reconstruction and detail retention of high-definition videos.

CN120278880APending Publication Date: 2025-07-08XIAYAN TECH (SHANDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410029487.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-06
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, high-resolution video transmission and playback are limited by device resolution, and traditional super-scoring methods cannot effectively restore high-frequency information of video images, resulting in large high-definition video transmission code stream and loss of detailed features.

Method used

A super-resolution neural network is built using a three-dimensional sparse attention mechanism and self-similarity coefficient. By generating high-resolution video sequences, the super-resolution reconstruction is carried out by generating high-resolution video sequences, and feature extraction and reconstruction is performed using window self-attention, sparse self-attention and mixed attention mechanisms.

Benefits of technology

Effectively reduce the high-definition video transmission code stream, avoid loss of detailed features after compression and decoding, and generate higher-definition video images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278880A_ABST
    Figure CN120278880A_ABST
Patent Text Reader

Abstract

The present invention provides a video processing method comprising the steps of: generating a high resolution video and a sequence pair corresponding to a low resolution video corresponding to the high resolution; importing the three-dimensional sparse attention mechanism into a super-resolution neural network; building a model of the super-resolution neural network based on the self-similarity coefficient; a low-resolution video sequence is input to the model of the super-resolution neural network to generate a super-resolution video sequence. According to the mode, the high-definition video transmission code stream is reduced, the loss of detail features after compression decoding is avoided, and thus a higher-definition image is effectively obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of video processing technologies, and in particular, to a video processing method and system. Background Art

[0002] With the development of video processing technologies, users' demands for video experience are constantly increasing. They not only need more realistic visual effects but also gradually have more information requirements for the details of the picture. However, transmitting or playing high-resolution videos usually requires a large bitstream. In traditional devices, due to different supported resolutions, the import or retention of high-resolution video transmission or playback is restricted.

[0003] Generally, the resolution of video images is restricted by conditions such as video image acquisition, imaging speed, and hardware storage. Moreover, in many existing video systems, the captured images and videos are of low resolution. For example, the images and videos captured by digital cameras and video surveillance systems are usually of low resolution. Therefore, in order to obtain high-resolution images or videos, super-resolution (SR) methods are needed to utilize the acquired low-resolution images or videos to reconstruct high-resolution images or videos. In the prior art, low-resolution video content can be upsampled and other methods to enhance it to high-definition picture quality. However, traditional super-resolution methods, such as interpolation, cannot effectively restore the high-frequency information in video images. Summary of the Invention

[0004] One solution of the present invention is completed in view of the above situation, and its purpose is to provide a video processing method and system to reduce the bitstream of high-definition video transmission, avoid the loss of detailed features after compression and decoding, and thus effectively obtain a higher-definition image.

[0005] In an embodiment of the present disclosure, a video processing method is provided, characterized in that the video processing method includes the following steps: generating a sequence pair of a high-resolution video and a low-resolution video corresponding to the high-resolution; introducing a three-dimensional sparse attention mechanism into a super-resolution neural network; building a model of the super-resolution neural network based on a self-similarity coefficient; and inputting the low-resolution video sequence into the model of the super-resolution neural network to generate a super-resolution video sequence.

[0006] In yet another embodiment of the present disclosure, a video processing system is provided, characterized in that the video processing system at least includes a server, a receiving end, and a transmission medium. Among them, the server acquires a high-resolution video, generates a low-resolution video from the high-resolution video to obtain a sequence pair corresponding to the high-resolution video and the low-resolution video corresponding to the high-resolution, and sends the sequence pair to the receiving end via the transmission medium; the receiving end performs the following steps: importing a three-dimensional sparse attention mechanism into a super-resolution neural network; building a model of the super-resolution neural network based on a self-similarity coefficient; inputting a low-resolution video sequence into the model of the super-resolution neural network to generate a super-resolution video sequence.

[0007] According to one solution of the present invention, the video processing method and system provided by the embodiments of the present disclosure can reduce the transmission bitstream of high-definition videos, avoid the loss of detailed features after compression and decoding, and thus effectively obtain a higher-definition image. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The various aspects of the present disclosure can be best understood when the following detailed description is read in conjunction with the accompanying drawings. The various features are not drawn to scale, and for the sake of discussion clarity, the dimensions of the various features can be arbitrarily increased or decreased. Moreover, it is obvious that the drawings in the following description are only embodiments of the present disclosure, and those skilled in the art can obtain other drawings according to the provided drawings without creative efforts.

[0009] Figure 1 It is a schematic structural diagram of a video processing system provided by some embodiments of the present disclosure.

[0010] Figure 2 It is a schematic flow diagram of a window self-attention unit provided by some embodiments of the present disclosure.

[0011] Figure 3 It is a schematic diagram of the window division features of a window self-attention unit provided by some embodiments of the present disclosure.

[0012] Figure 4 It is a schematic diagram of the window division features of a shifted dense self-attention unit provided by some embodiments of the present disclosure.

[0013] Figure 5 It is a schematic diagram of feature fusion of a video processing method provided by some embodiments of the present disclosure.

[0014] Figure 6 It is a schematic flow diagram of a sparse self-attention unit provided by some embodiments of the present disclosure.

[0015] Figure 7Flow diagram of the grid self-attention unit in the sparse self-attention unit provided by some embodiments of the present disclosure.

[0016] Figure 8 Schematic diagram of the grid division features of the grid self-attention unit provided by some embodiments of the present disclosure.

[0017] Figure 9 Schematic diagram of feature fusion of the video processing method provided by some embodiments of the present disclosure.

[0018] Figure 10 Schematic diagram of the structure of the super-resolution module in the video processing system provided by some embodiments of the present disclosure.

[0019] Figure 11 Flow diagram of the video processing method provided by some embodiments of the present disclosure.

[0020] Figure 12 Schematic diagram of the comparison of the results of a visual experiment in a public dataset (Urban100) for one embodiment of the present disclosure. Detailed implementation manners

[0021] The following disclosure includes specific information related to the embodiments in the present disclosure. The accompanying drawings and the corresponding detailed disclosure are directed to exemplary embodiments. However, the present disclosure is not limited to these exemplary embodiments. Those skilled in the art will think of other variations and embodiments of the present disclosure. Unless otherwise indicated, similar or corresponding elements in the drawings may be indicated by similar or corresponding reference numerals. The drawings and illustrations are generally not drawn to scale and are not intended to correspond to actual relative sizes.

[0022] For the purpose of consistency and ease of understanding, in the exemplary drawings, similar features are identified by reference numerals (although not shown in some examples). However, the features in different embodiments may be different in other aspects, and thus should not be narrowly limited to what is shown in the drawings.

[0023] It should be noted that "at least one" in the present disclosure means one or more, and "a plurality" means two or more than two. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The terms "first", "second", etc. (if any) in the specification, claims, and drawings of the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0024] Those skilled in the art of the present technology can understand that the "client", "receiver", and "server" in the present disclosure include not only devices with video transceiver capabilities, which are devices with video transceivers without transmission capabilities, but also devices with receiving and transmitting hardware, which are devices with receiving and transmitting hardware capable of two-way transmission on a two-way transmission channel. Such devices may include: communication devices such as personal computers and tablet computers, which have single-line displays or multi-line displays or cellular or other communication devices without multi-line displays; personal communication systems (PCS), which can combine voice, data processing, fax, and / or data communication capabilities; personal digital assistants (PDAs), which may include radio frequency receivers, pagers, Internet / intranet access, web browsers, notepads, calendars, and / or global positioning system (GPS) receivers; traditional laptop and / or palm computers or other devices, which are traditional laptop and / or palm computers or other devices with and / or including radio frequency receivers. The "client", "receiver", and "server" used herein can be portable, transportable, installed in a vehicle (air, sea, and / or land), or suitable for and / or configured to run locally and / or run in a distributed manner at any other location on the earth and / or in space. The "client", "receiver", and "server" used herein can also be communication terminals, Internet access terminals, music / video playback terminals, such as PDAs, mobile Internet devices (MIDs), and / or mobile phones with music / video playback functions, and can also be devices such as smart TVs and set-top boxes.

[0025] The hardware referred to by names such as "client", "receiver", and "server" in the present disclosure is essentially an electronic device with the equivalent capabilities of a personal computer, which is a hardware device with the necessary components disclosed by the von Neumann principle, including a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. The computer program is stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input / output devices to complete specific functions.

[0026] It should be noted that the "server" mentioned in this disclosure can similarly be extended to apply to the case of a server cluster. According to the network deployment principle understood by those skilled in the art, the various servers should be a logical division. Physically, these servers can either be independent of each other but can be called through an interface, or integrated into a single physical computer or a set of computer clusters. Those skilled in the art should understand this variation and should not be restricted by this in the implementation manner of the network deployment method of this disclosure.

[0027] Figure 1 FIG. 4 shows a schematic structural diagram of a video processing system 100 provided by some embodiments of the present disclosure. In Figure 1 this case, the video processing system 100 may include a server 110, a receiving end 130, and a transmission medium 120.

[0028] The server 110 may be a video encoder and / or decoder. The server 110 may also be any device such as a source end or a sending end that is capable of transmitting or sending image and / or video data. The server 110 may also include any device configured to encode video data based on data information obtained from a specific acquisition device and transmit the encoded video data to the transmission medium 120.

[0029] The receiving end 130 may be a terminal device or a client device. The receiving end 130 includes any device configured to receive the encoded video data via the transmission medium 120 and decode the encoded video data.

[0030] The server 110 may at least include a downsampling module 111 and an encoding module 121. The receiving end may at least include a decoding module 131 and a super-resolution module 132. The server 110 may be a video encoder and / or a video decoder, and the receiving end 130 may be a video decoder and / or a video encoder.

[0031] In addition, the server 110 and / or the receiving end 130 may be a mobile phone, a tablet computer, a desktop computer, a laptop computer, or other electronic devices. Figure 1 FIG. 5 shows an example of the server 110 and the receiving end 130. The server 110 and the receiving end 130 may include more or fewer components than those shown, or have different configurations of various components.

[0032] The transmission medium 120 may also be referred to as a video or image transmission medium. The transmission medium may be a Video Transmission Channel (VTC), and the video transmission channel may include a computer system interface capable of storing a compatible video bitstream on a storage device or receiving a compatible video bitstream from a storage device. For example, the video transmission channel may include a chipset supporting Peripheral Component Interconnect (PCI) and Peripheral Component Interconnect Express (PCIe) bus protocols, proprietary bus protocols, Universal Serial Bus (USB) protocols, I2C, or any other logical and physical structure that can be used to interconnect peer devices.

[0033] As Figure 1 shown, the server 110 can process information obtained from fixed collection devices. Specifically, the collection devices may include a video capture device for capturing new videos, a video archive for storing previously captured videos, and / or a video feed interface for receiving videos from video content providers. The collection devices may generate computer graphics-based data as the source video, or generate a combination of live videos, archived videos, and computer-generated videos as the source video. The video capture device may be a charge-coupled device (CCD) image sensor, a complementary metal-oxide-semiconductor (CMOS) image sensor, or a camera, etc.

[0034] In some embodiments of the present disclosure, the server 110 may include either the downsampling module 111 or the encoding module 121, or may at least include the downsampling module 111 and the encoding module 121. The downsampling module 111 may be a downsampling module or a module with equivalent or similar functions. The server 110 may also be referred to as an encoder.

[0035] In some embodiments of the present disclosure, the receiving end 130 may include either the decoding module 131 or the super-resolution module 132, or may at least include the decoding module 131 and the super-resolution module 132. The super-resolution module 132 may also be a video reconstruction module. The receiving end 130 may also be referred to as a decoder.

[0036] In some embodiments of the present disclosure, as Figure 1As shown, the server 110 encodes the received video information through a video encoding method and inputs it to the transmission medium 120. The transmission medium 120 transmits the encoded information to the receiving end 130. The receiving end 130 first performs decoding processing on the encoded information through a video decoding method, and then reconstructs the decoded video information.

[0037] It should be noted that the encoding module 121 and the decoding module 131 can each be implemented as any one of a variety of suitable encoder / decoder circuits, such as one or more microprocessors, central processing unit (CPU), graphic processing unit (GPU), system on chip (SoC), digital signal processor (DSP), application specific integrated circuit (ASIC), field programmable gate array (FPGA), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the disclosed methods. In at least some embodiments, each of the encoder module and the information compression module can be included in one or more encoders and decoders, each of which can be integrated as part of a combined encoder / decoder (CODEC) in the device.

[0038] In some embodiments of the present disclosure, during the process of performing high-resolution processing on the acquired video information, the input and feature extraction of the image bitstream can be analyzed. Specifically, the bitstream data of the image and the features of the image can be analyzed by combining the Attention mechanism and convolutional networks (such as CNN, RNN). Among them, the Attention mechanism can be the Self-attention mechanism.

[0039] As Figure 2 shown, Figure 2Flow diagram of the window self-attention unit 140 provided for some embodiments of the present disclosure. The window self-attention unit 140 may include at least one dense self-attention unit, a layer normalization unit (Layer Normalization, Layer Norm) (hereinafter may be simply referred to as: Layer Norm), and a multi-layer perceptron (MLP) unit.

[0040] In some embodiments of the present disclosure, the window self-attention unit 140 may at least include a (shifted) dense self-attention unit 142, a first Layer Norm unit 141, a second Layer Norm unit 143, and an MPL unit 144. At the same time, in each unit, embodiments of the present disclosure may introduce a residual connection to avoid gradient vanishing.

[0041] In some embodiments of the present disclosure, as Figure 2 shown, in the dense self-attention unit, the feature map of the input image may be first divided into non-overlapping dense regions according to a certain window size, where the window size may be set to 16×16, or may be any other size. Regarding the window size, this application does not limit it.

[0042] Figure 3 Schematic diagram of the window division characteristics of the window self-attention unit 140 provided for some embodiments of the present disclosure. In some embodiments of the present disclosure, the input feature map may be defined as The window is divided according to the size K×K to form non-overlapping dense regions. Here, the formed dense regions may be defined as where H, W, and C respectively represent the height, width, and number of feature channels of the feature map of each video frame of the input video.

[0043] Figure 4 Schematic diagram of the window division characteristics of the shifted dense self-attention unit 142 provided for some embodiments of the present disclosure. In some embodiments of the present disclosure, a window configuration different from that of the previous layer (such as: window size or shift configuration) is used to shift any pixel of the window of the input feature map, and the feature map is segmented by the shifted window segmentation method.

[0044] In some embodiments of the present disclosure, the shifted window segmentation method may be large window shifted segmentation or small window shifted segmentation. As Figure 4 shown, in some embodiments of the present disclosure, the window may be shifted Pixels are used to generate non - overlapping dense regions. Then, regions smaller than the K×K window size in the dense regions are determined, and the regions smaller than the K×K window size are stitched together.

[0045] In some embodiments of the present disclosure, as Figure 4 shown, the window can be shifted by pixels, that is: as Figure 4 A, B, C, D in. Then A, B, C, D are respectively shifted to positions symmetric about the feature map. Specifically, A can be shifted to the lower - right position; B and C can be shifted to the lower position; D and E can be shifted to the right position, and then stitched together respectively to generate a region with a window size of K×K. The generated feature map can be defined as for further use in the self - attention mechanism.

[0046] Figure 5 This is a schematic diagram of feature fusion of the video processing method provided by some embodiments of the present disclosure. In some embodiments of the present disclosure, as Figure 2 shown, the features obtained by dense self - attention or shifted dense self - attention Figure X m or X sm , through the Layer Norm unit, are normalized or linearly transformed to calculate the window self - attention, so as to obtain the features after window self - attention. Among them, the feature map obtained by the dense self - attention or shifted dense self - attention mechanism can also be defined as X (s)m .

[0047] In some embodiments of the present disclosure, the features obtained by dense self - attention or shifted dense self - attention Figure X (s)m are respectively converted into the Q (Query), K (Key), and V (Value) matrices in the attention mechanism through linear transformation. The matrices can be rows, columns, or vectors, where In some embodiments of the present disclosure, as Figure 2 and Figure 5 shown, in the Layer Norm unit, the converted Q, K, and V matrices are subjected to SoftMax normalization processing, and the normalized feature information is added to the features obtained by dense self - attention or shifted dense self - attention Figure X (s)m to obtain the features obtained after calculating the window self - attention for network reconstruction. Specifically, the features obtained after window self - attention can be calculated using Q, K, and V, and the features can be defined as follows: where b represents the learnable relative position encoding, XA It represents the features obtained after computing window self-attention. Then, the obtained X A is input into the second Layer Norm unit 143 and the MLP unit 144. The second Layer Norm unit 143 and the MLP unit 144 perform a linear combination on the input features to obtain the multi-layer features of the window self-attention unit. The multi-layer features can be defined as follows: X D = LN(MLP) + X A ; Thus, the specific detailed features of the feature map can be further obtained, facilitating the refinement of network reconstruction.

[0048] Figure 6 The flowchart of the sparse self-attention unit 150 provided by some embodiments of the present disclosure. In some embodiments of the present disclosure, the sparse self-attention unit 150 may at least include a grid self-attention unit, a Layer Norm unit, and an MLP unit. The sparse self-attention unit 150 may also at least include a grid self-attention unit 152, a first Layer Norm unit 151, a second Layer Norm unit 153, and an MPL unit 154. At the same time, in each unit, embodiments of the present disclosure can introduce residual connections to avoid gradient disappearance.

[0049] In some embodiments of the present disclosure, in order to obtain the feature data of the grid self-attention unit 152, it can be obtained at least based on dense self-attention and shifted dense self-attention mechanisms, or it can be obtained based on either the dense self-attention or the shifted dense self-attention mechanism.

[0050] Figure 7 The flowchart of the grid self-attention unit 152 in the sparse self-attention unit 150 provided by some embodiments of the present disclosure. In some embodiments of the present disclosure, in the grid self-attention unit, first, the input features Figure X can be divided into two parts by channel. One part is defined as The other part is defined as After that, the divided F W can also be divided into two parts by channel and are respectively input into the dense self-attention unit and the shifted dense self-attention unit for calculation, and feature maps F WA and F WSA are generated; at the same time, F G is input into the grid self-attention unit to generate the feature map F GA . In some embodiments of the present disclosure, the grid self-attention unit can perform operations such as Figure 8 .

[0051] Figure 8 Schematic diagram of the grid partitioning features of the grid self-attention unit 152 provided in some embodiments of the present disclosure. As Figure 8 shown, the grid self-attention unit 152 can process the input features Figure X and the feature map F G , partition the area according to a certain grid size, and the partitioned area can be a sparse area. Among them, the grid size can be defined as G×G, and the present application does not limit the grid size.

[0052] In some embodiments of the present disclosure, in the grid self-attention module, the features Figure X and F G can be evenly divided into distributed sparse areas according to the grid size G×G, and through linear transformation, the global interaction features are obtained.

[0053] Figure 9 Schematic diagram of the feature fusion process of the video processing method provided in some embodiments of the present disclosure. In some embodiments of the present disclosure, based on the distributed sparse areas, the global interaction features can be generated by linearly mapping the features Figure X , and the global interaction features can be defined as At the same time, by linearly mapping the feature map F G , based on the attention mechanism, the attention interaction features are generated, and the attention interaction features can be respectively defined as G Q , G K , G V , and As Figure 9 shown, the generated G Q , G K , G V and G are input into the hybrid attention unit, and the hybrid attention unit can be defined as F H . Specifically, by normalizing the global interaction feature G with one of the attention interaction features, that is, the grid feature G Q , the texture feature similarity coefficient between different regions of the feature map is generated. Further, the global feature G and other attention interaction features G K and G V are subjected to self-attention calculation to generate a feature map fused based on the self-similarity coefficient, thereby improving the effect of video reconstruction.

[0054] In some embodiments of the present disclosure, in the hybrid attention unit, based on the attention mechanism and normalization processing, the following model can be used to obtain the similarity fusion feature map, and the model can be defined as: F H= SoftMax(G × G Q ) × (SoftMax(G × G K ) × G V ); where G contains the global features of the image; F G contains the local features of the image. Specifically, after grid division, the network can model according to the local self-similarity of the image. Through the global feature G and the grid feature G Q , perform SoftMax calculation to generate the texture feature similarity coefficient between different regions of the feature map. Then, through the global feature G and the grid features G K and G V , perform self-attention calculation, which can fuse the global feature G and the grid feature F G . Then, multiply the fused feature by the self-similarity coefficient to generate a feature map fused based on the self-similarity coefficient, so as to form more comprehensive feature detail information and further improve the video reconstruction effect.

[0055] As Figure 5 and Figure 9 shown, in the embodiments of the present disclosure, the feature maps generated by the dense self-attention unit, the shifted dense self-attention unit, and the hybrid attention unit can be dimensionally stacked by channel to generate a comprehensive feature map As Figure 5 shown, input the generated feature map into the second Layer Norm unit and the MLP unit to generate the features of the sparse attention unit for super-resolution video reconstruction.

[0056] In some embodiments of the present disclosure, splicing, dimensional stacking, or residual connection can be implemented by subtractors, adders, filters, etc. The dimensional stacking can also be stacking methods such as vector splicing, identity mapping, and residual operation. Among them, the subtractors, adders, and filters can include deblocking filters, sample adaptive offset (SAO) filters, bilateral filters, and / or adaptive loop filters (ALF) to remove blocking artifacts from the reconstructed blocks. In addition to deblocking filters, SAO filters, bilateral filters, and ALF, other filters (inside or after the loop) can also be used. For the sake of brevity, such filters are not explicitly illustrated.

[0057] Figure 10Schematic diagram of the structure of the super-resolution module in the video processing system provided by some embodiments of the present disclosure. The super-resolution module can be divided into at least a shallow feature extraction module, a deep feature extraction module, and a video image reconstruction module. In the super-resolution module, super-resolution video image reconstruction is performed through the shallow feature extraction module, the deep feature extraction module, and the video image reconstruction module.

[0058] In some embodiments of the present disclosure, the shallow feature extraction module can perform a convolution operation. Specifically, the shallow feature extraction module can include at least a convolution kernel of size 3×3. The shallow feature extraction module generates a feature map with 180 channels from a picture with 3 channels by using 180 convolution kernels, thereby completing the extraction of shallow features. Specifically, the feature extraction formula can be defined as follows: F L =Conv(X in ); where, X in represents the input picture with 3 channels, Conv(·) represents the convolution operation, and F L represents the feature map with 180 channels obtained after the convolution operation.

[0059] In some embodiments of the present disclosure, the deep feature extraction module can be composed of a plurality of consecutive window self-attention units and a plurality of sparse self-attention units, or can be composed of a plurality of consecutive window self-attention units and at least one sparse self-attention unit, or can be composed of at least one window self-attention unit and at least one sparse self-attention unit.

[0060] In some embodiments of the present disclosure, a complete window processing layer can be formed by every 2 consecutive (shifted) dense self-attention. Every 3 window processing layers and a sparse self-attention processing unit can form a residual hybrid attention layer. 180 channels can be used in the residual hybrid attention layer, and residual connection is used to avoid gradient disappearance. During the deep feature extraction process, the residual hybrid attention layer can be repeated 6 times. The deep feature extraction can be defined as follows: where, F in represents the feature input to each hybrid attention layer; represents the i-th window self-attention unit, and H G-MSA (·) represents the sparse self-attention unit; represents the hybrid attention layer; F D represents the feature finally obtained after deep feature extraction.

[0061] In some embodiments of the present disclosure, the video image reconstruction module may generate a high-resolution image by performing convolution and / or recombination between multiple channels on the low-resolution feature map obtained by performing shallow feature extraction.

[0062] In some embodiments of the present disclosure, video image reconstruction may be performed using any one or any combination of transposed convolution, Pixel Shuffle, Down and UP Sampling, Meta-Upscale, and CAPAFE.

[0063] In some embodiments of the present disclosure, video image reconstruction may be performed by pixel recombination, i.e., in the Pixel Shuffle manner. Specifically, the process of the video reconstruction may be defined as follows: HR = H P (F D ); where, H P (·) represents Pixel Shuffle upsampling, and HR represents the image finally obtained through super-resolution reconstruction.

[0064] Figure 11 is a schematic flow diagram of a video processing method provided in some embodiments of the present disclosure. Some embodiments of the present disclosure can directly or indirectly complete video image reconstruction through the Figure 11 steps shown, and obtain a higher-definition video image.

[0065] As Figure 11 shown, in step S100, a high-resolution video and a sequence pair corresponding to the low-resolution video corresponding to the high-resolution are generated.

[0066] In some embodiments of the present disclosure, the high-resolution video image received by the server is sliced into single-frame pictures. Among them, the high-resolution video image may be a high-resolution video image conforming to video standards such as H.265 / HEVC, for example: 8K video image. By using a specified application program to process the sliced images, an image sequence pair is formed between the sliced video images and the high-resolution video images, and codec processing is performed, and the processed information (such as elements such as low-resolution video sequences or decoded information) is sent to the receiving end for super-resolution reconstruction. Among them, the application program may be an application program such as Python, Ruby, MATLAB, Simulink, Stateflow, Visual Basic, JavaScript or a combination thereof, and the present application is not limited to the above application programs.

[0067] In some embodiments of the present disclosure, MATLAB can be used to generate a downsampled image by interpolation. Among them, the interpolation method can be the nearest neighbor method, bilinear interpolation, bicubic interpolation, etc. The present application is not limited to the above interpolation methods.

[0068] In some embodiments of the present disclosure, the bicubic interpolation kernel in MATLAB can be used to generate a two-fold downsampled image. The generated two-fold downsampled image and the 8K video frame form an image sequence pair. Then, based on the formed image sequence pair, through the Figure 1 encoding / decoding module shown in the figure, video compression and / or generation of the formed image sequence are performed. Among them, the video compression and generation methods can be video image encoding and decoding methods specified in standards such as motion still image (or frame-by-frame) compression, (Motion-Join Photographic Experts Group, M-JPEG), H.265 / High Efficiency Video Coding (HEVC), AVS2, AV1, etc. It can also use intra-frame coding methods or inter-frame coding methods, which are not limited here.

[0069] As Figure 11 shown, in step S200, a three-dimensional sparse attention mechanism is introduced into the video super-resolution neural network.

[0070] In some embodiments of the present disclosure, a self-similarity model can be introduced into the three-dimensional sparse attention mechanism for the establishment of the video super-resolution neural network. Among them, as Figures 2 to 9 shown, the three-dimensional sparse attention mechanism can at least include the processes performed by the window self-attention unit and the sparse self-attention unit.

[0071] Step S300, construct a self-similarity video super-resolution network model. As Figure 10 shown, the super-resolution network model is constructed through a shallow feature extraction module, a deep feature extraction module, and a video image reconstruction module, and the super-resolution network model is introduced into the super-resolution module.

[0072] Step S400, input the low-resolution video sequence into the video super-resolution network model to generate a super-resolution video sequence, so as to obtain the true detail features of the original high-resolution video and present a higher-definition video image.

[0073] In some embodiments of the present disclosure, as Figure 1 and Figure 8As shown, the linear parameter information (i.e., Linear parameter) can be included in the decoding module for decoding to generate the linear parameter information corresponding to the low-resolution video sequence and the super-resolution neural network. Finally, the low-resolution video is super-resolution reconstructed through the established super-resolution neural network to obtain the high-resolution video sequence, thereby obtaining an ultra-high-definition video image.

[0074] To verify the effect of the present invention, the performance of the proposed network model was confirmed and compared through experiments. Hereinafter, Table 1 shows the comparison results of the video image quality of one embodiment of the present disclosure and three other network model structures.

[0075] Table 1

[0076] Specifically, the above embodiment in Table 1 adopted the public datasets DIV2K and Flickr2K. At the same time, the public datasets DIV2K and Flickr2K were used as the training set. During the training process, L1 was used as the loss function, and the number of iterations was set to 500,000 times. Set5, Set14, B100, Urban100, and Manga109 are respectively general public datasets. The specific test results (as shown in Table 1) of the network after training on the Set5, Set14, B100, Urban100, and Manga109 datasets were quantitatively compared with the test results of the three advanced models FSRCNN, EDSR, and SwinIR in Table 1 in terms of two performance evaluation indicators: PSNR and Structural Similarity (SSIM).

[0077] Figure 12 This is a comparison diagram of the visual experimental results of one embodiment of the present disclosure in the public dataset (Urban100). Through comparison, it can be seen that in pictures with high self-similarity, this embodiment can reconstruct pictures with richer texture details and present higher-definition images.

[0078] The video processing method and system of the embodiments of the present disclosure are described in various aspects.

[0079] In one embodiment of the present disclosure, a video processing method is provided, characterized in that the video processing method includes the following steps: generating a high-resolution video and a sequence pair corresponding to a low-resolution video corresponding to the high-resolution; introducing a three-dimensional sparse attention mechanism into a super-resolution neural network; building a model of the super-resolution neural network based on a self-similarity coefficient; and inputting the low-resolution video sequence into the model of the super-resolution neural network to generate a super-resolution video sequence.

[0080] In another embodiment of the present disclosure, a video processing system is provided, characterized in that the video processing system at least includes a server, a receiving end, and a transmission medium. Among them, the server obtains a high-resolution video, generates a low-resolution video from the high-resolution video to obtain a high-resolution video and a sequence pair corresponding to the low-resolution video corresponding to the high-resolution, and sends the sequence pair to the receiving end via the transmission medium; the receiving end performs the following steps: introducing a three-dimensional sparse attention mechanism into a super-resolution neural network; building a model of the super-resolution neural network based on a self-similarity coefficient; and inputting the low-resolution video sequence into the model of the super-resolution neural network to generate a super-resolution video sequence.

[0081] The video processing method and system provided by the embodiments of the present disclosure can reduce the high-definition video transmission bitstream and avoid the loss of detailed features after compression and decoding, so as to effectively obtain a higher-definition image.

[0082] As described above, the embodiments of the present invention have been described in detail with reference to the accompanying drawings. However, the specific configuration is not limited to this embodiment, and also includes design changes and the like that do not depart from the gist of the present invention. In addition, one aspect of the present invention can be variously changed within the scope shown in the technical solution, and embodiments obtained by appropriately combining the technical solutions disclosed in different embodiments are also included in the technical scope of the present invention. In addition, it also includes a configuration obtained by mutually replacing elements having the same effect among the elements described in the above embodiments.

[0083] Industrial applicability

[0084] One aspect of the present invention can be used, for example, in a video processing system, a transmission device (such as a portable mobile phone, etc.), an integrated circuit (such as an image processing chip), or a program.

Claims

1. A video processing method, characterized in that The video processing method includes the following steps: Generate a sequence pair corresponding to a high-resolution video and a low-resolution video corresponding to the high resolution; Introduce a three-dimensional sparse attention mechanism into the super-resolution neural network; Build a model of the super-resolution neural network based on the self-similarity coefficient; Input the low-resolution video sequence into the model of the super-resolution neural network to generate a super-resolution video sequence.

2. The video processing method according to claim 1, wherein The method further includes: Slice the high-resolution video into single-frame pictures, and use bicubic interpolation to generate the downsampled low-resolution video by a factor of two; Based on the low-resolution video, generate a sequence pair corresponding to the high-resolution video through upsampling and downsampling.

3. The video processing method according to claim 1, wherein The three-dimensional sparse attention mechanism at least includes the processes performed by the window self-attention unit and the sparse self-attention unit.

4. The video processing method according to claim 3, wherein The window self-attention unit includes at least one dense self-attention unit, a layer normalization (Layer Norm) unit, and a multi-layer perceptron (MLP) unit, wherein The dense self-attention unit is used to divide a specific feature map into dense regions according to a certain window size to generate a first feature map, wherein the dense regions are non-overlapping regions; Perform normalization processing on the first feature map through the LayerNorm unit and the MLP unit.

5. The video processing method according to claim 3, wherein The window self-attention unit includes at least one dense self-attention unit, a layer normalization (Layer Norm) unit, and a multi-layer perceptron (MLP) unit; The sparse self-attention unit includes at least one grid self-attention unit, a Layer Norm unit, and an MLP unit. The method further includes: Divide a specific feature map into first feature information and second feature information according to channels; Input the first feature information into the grid self-attention unit to generate third feature information; Input the second feature information into the dense self-attention unit to generate fourth feature information; Perform dimensional stacking on the third feature and the fourth feature.

6. The video processing method according to claim 5, wherein The method further includes: Divide the feature map and the first feature information into sparse regions according to a certain grid size to generate a second feature map, wherein the sparse regions are uniformly distributed regions; Perform a linear transformation on the second feature map to generate a global interaction feature, and perform a linear transformation on the first feature information to generate an attention interaction feature, wherein the attention interaction feature contains three pieces of feature information; Perform normalization processing on either the global interaction feature or the attention interaction feature to obtain the self-similarity coefficient of different regions, and generate a fused feature map of the self-similarity coefficient.

7. The video processing method according to claim 1, wherein The super-resolution neural network includes a shallow feature extraction module, a deep feature extraction module, and a video image reconstruction module; The method further includes: Obtain the shallow features of the fused feature map from the shallow feature extraction module; Obtain the deep features of the fused feature map from the deep feature extraction module; Build the model of the super-resolution neural network based on the information of the shallow features and the information of the deep features.

8. The video processing method according to claim 7, wherein The shallow feature extraction module includes at least a convolutional kernel of size 3×3; The deep feature extraction module includes a plurality of consecutive window self-attention units and sparse self-attention units.

9. The video processing method according to claim 8, wherein The method further includes: The shallow feature extraction module generates a feature map with 180 channels from a picture with 3 channels by using 180 convolutional kernels to obtain the shallow features of a specific feature map.

10. The video processing method according to claim 8, wherein The method further includes: Construct each two consecutive (shifted) dense self-attention units into a window processing layer; Construct every three window processing layers and a sparse self-attention processing unit into a residual hybrid attention layer. Through cyclic processing, 180 channels are used in each residual hybrid attention layer, and deep features are generated through residual connection.

11. The video processing method according to claim 5, wherein The method further includes: The window self-attention unit includes a dense self-attention unit, a shifted dense self-attention unit, a layer normalization (Layer Norm) unit, and a multi-layer perceptron (MLP) unit.

12. A video processing system, wherein The video processing system includes at least a server, a receiver, and a transmission medium, wherein The server obtains a high-resolution video, generates a low-resolution video from the high-resolution video to obtain a sequence pair of the high-resolution video and the low-resolution video corresponding to the high resolution, and sends the sequence pair to the receiver via the transmission medium; The receiver performs the following steps: Import a three-dimensional sparse attention mechanism into the super-resolution neural network; Build the model of the super-resolution neural network based on the self-similarity coefficient; Input the low-resolution video sequence into the model of the super-resolution neural network to generate a super-resolution video sequence.

13. The video processing system according to claim 12, wherein The receiver is configured to include one or more modules that execute all combinations of claims 1 to 11.