Video enhancement method, related device and computer program product
By extracting features from the target network model and performing multi-frame alignment processing, combined with multi-parallel convolutional branches and equivalent convolutional kernel weight merging techniques, the problem of efficient and real-time removal of video coding artifacts was solved, achieving high-quality video enhancement.
Patent Information
- Application Number
- CN202510984438.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-10-24
AI Technical Summary
Existing technologies struggle to efficiently, in real-time, and with high quality remove artifacts introduced by video encoding and compression; traditional methods cannot simultaneously meet the requirements of efficiency and quality.
Feature extraction, multi-frame alignment, and image enhancement are performed using a target network model. By utilizing multiple parallel convolutional branches and equivalent convolutional kernel weight merging techniques, efficient fusion and enhancement of multi-frame features are achieved.
While reducing computational complexity, it effectively removes artifacts, improves video quality, and achieves efficient, real-time, and high-quality video enhancement effects.
Smart Images

Figure CN120833274A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of image processing, and particularly relates to a video enhancement method and device, electronic equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] This section is intended to provide background information to facilitate a better understanding of embodiments of the present disclosure. Information in this section is not admitted to be prior art.
[0003] With the rapid growth of Internet video data, in order to control the storage and transmission cost of video, a higher compression rate is usually adopted when encoding the video. Generally, a lossy compression algorithm is adopted in the process of video encoding. These encoding methods, while reducing the size of the video, will introduce video artifacts caused by compression, such as blocking effect, ringing effect, flicker effect and mosquito noise, etc. These artifacts will seriously degrade the video quality and affect the user's viewing experience. Therefore, removing video artifacts caused by encoding compression is a very valuable technology.
[0004] However, the traditional video de-compression artifact method has the technical bottleneck of being unable to simultaneously meet efficient, real-time and high-quality enhancement. SUMMARY
[0005] The purpose of the present disclosure is to provide a video enhancement method and device, electronic equipment, computer readable storage medium and computer program product, which can efficiently, in real time and with high quality, remove video artifacts to enhance the video.
[0006] Other characteristics and advantages of the present disclosure will become apparent from the following detailed description, or will be learned by practice of the present disclosure.
[0007] The embodiment of the present disclosure provides a video enhancement method, comprising: obtaining a target frame and at least one reference frame of the target frame from a video to be enhanced, the at least one reference frame comprising at least one of N frames before the target frame and M frames after the target frame, N being a frame number greater than or equal to 1, and M being an integer greater than or equal to 1; performing feature extraction on the target frame and the at least one reference frame through a feature extraction module of a target network model to obtain target frame features corresponding to the target frame and reference frame features corresponding to each reference frame; performing multi-frame alignment processing on the target frame features and the reference frame features through a multi-frame alignment module of the target network model to obtain multi-frame alignment features; performing image enhancement processing on the multi-frame alignment features through an image enhancement module of the target network model to obtain image enhancement features; and adding the image enhancement features and the target frame pixel by pixel to obtain an enhanced target frame; wherein the image enhancement module comprises at least one representative feature extraction submodule, the representative feature extraction submodule comprises a plurality of parallel convolution branches, and each convolution branch comprises at least one convolution kernel; wherein the image enhancement module performs image enhancement processing on the multi-frame alignment features to obtain image enhancement features, comprising: merging convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight; and performing convolution processing on the multi-frame alignment features through the equivalent convolution kernel weight to obtain equivalent convolution features, so as to determine the image enhancement features according to the equivalent convolution features.
[0008] In some embodiments, the equivalent convolution kernel weight comprises a first equivalent weight, the plurality of branches comprise a first branch and a second branch, the first branch comprises a first convolution kernel, and the second branch comprises a second convolution kernel; wherein merging the convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight comprises: performing zero value padding on a first convolution kernel weight of the first convolution kernel and a second convolution kernel weight of the second convolution kernel to make the matrix sizes of the padded first convolution kernel weight and the padded second convolution kernel weight the same; and adding the first convolution kernel weight and the second convolution kernel weight channel by channel to determine the first equivalent weight.
[0009] In some embodiments, the equivalent convolution kernel weight comprises a second equivalent weight, the plurality of branches comprise a third branch, the third branch comprises a third convolution kernel and a fourth convolution kernel in series; wherein merging the convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight comprises: performing zero value padding on a third convolution kernel weight of the third convolution kernel and a fourth convolution kernel weight of the fourth convolution kernel to make the matrix sizes of the padded third convolution kernel weight and the padded fourth convolution kernel weight the same; and performing spatial convolution on the third convolution kernel weight and the fourth convolution kernel weight channel by channel to determine the second equivalent weight.
[0010] In some embodiments, the method further comprises: obtaining a training target frame and at least one training reference frame of the training target frame from the training video; performing image enhancement processing on the training target frame and the at least one training reference frame by the target network model to obtain training image enhancement features; adding the training image enhancement features and the training target frame pixel by pixel to obtain an enhanced training target frame; determining a target loss function according to the enhanced training target frame; and adjusting parameters of the target network model by the target loss function, wherein the parameters of the plurality of parallel convolution branches in the image enhancement module are adjusted respectively.
[0011] In some embodiments, the multi-frame alignment module comprises at least two alignment sub-modules, including a first alignment sub-module and a second alignment sub-module; wherein the multi-frame alignment processing of the target frame features and the respective reference frame features by the multi-frame alignment module of the target network model comprises: performing multi-frame alignment processing on the target frame features and the respective reference frame features by the first alignment sub-module to obtain first alignment features; and performing multi-frame alignment processing on the target frame features and the respective reference frame features by the second alignment sub-module to obtain second alignment features; the convolution kernel sizes in the first alignment sub-module and the second alignment sub-module are different; and the first alignment features and the second alignment features are fused to obtain the multi-frame alignment features.
[0012] In some embodiments, the first alignment sub-module comprises a first bias prediction sub-structure and a first convolution sub-structure; wherein the multi-frame alignment processing of the target frame features and the respective reference frame features by the first alignment sub-module to obtain first alignment features comprises: performing bias prediction processing on the target frame features and at least one reference frame by the first bias prediction sub-structure to determine a first offset of actual sampling points of the convolution sub-structure on the respective reference frame features compared with standard sampling points; wherein the standard sampling points are determined according to the convolution kernel of the convolution sub-structure; the first convolution sub-structure determines first actual sampling points according to the standard sampling points and the first offset; the first convolution sub-structure performs convolution processing on the respective reference frame features according to the first actual sampling points to obtain first aligned reference frame features; and the first aligned reference frame features and the target frame features are fused to obtain the first alignment features.
[0013] The embodiments of the present disclosure provide a video enhancement device, which comprises an image acquisition module, a feature extraction module, a multi-frame alignment feature determination module, an image enhancement module, and a pixel addition module.
[0014] The image acquisition module is configured to acquire a target frame and at least one reference frame of the target frame from a video to be enhanced, the at least one reference frame including at least one of N previous frames of the target frame and M subsequent frames of the target frame, N being a frame number greater than or equal to 1, and M being an integer greater than or equal to 1; the feature extraction module is configured to perform feature extraction on the target frame and the at least one reference frame by a feature extraction module of a target network model to obtain target frame features corresponding to the target frame and reference frame features corresponding to each reference frame; the multi-frame alignment feature determination module is configured to perform multi-frame alignment processing on the target frame features and the reference frame features by a multi-frame alignment module of the target network model to obtain multi-frame alignment features; the image enhancement module is configured to perform image enhancement processing on the multi-frame alignment features by an image enhancement module of the target network model to obtain image enhancement features; and the pixel addition module is configured to add the image enhancement features and the target frame pixel by pixel to obtain an enhanced target frame. The image enhancement module includes at least one representative feature extraction sub-module, and the representative feature extraction sub-module includes a plurality of parallel convolution branches, each convolution branch including at least one convolution kernel. The image enhancement module performs image enhancement processing on the multi-frame alignment features to obtain image enhancement features, including: combining convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight; and performing convolution processing on the multi-frame alignment features by using the equivalent convolution kernel weight to obtain equivalent convolution features, so as to determine the image enhancement features according to the equivalent convolution features.
[0015] The electronic device includes a memory and a processor. The memory is configured to store computer program instructions. The processor is configured to invoke the computer program instructions stored in the memory to implement the video enhancement method.
[0016] The computer-readable storage medium stores computer program instructions for implementing the video enhancement method.
[0017] The computer program product or the computer program includes computer program instructions stored in a computer-readable storage medium. The computer program instructions are read from the computer-readable storage medium, and a processor executes the computer program instructions to implement the video enhancement method.
[0018] The video enhancement method, device, electronic device, computer readable storage medium and computer program product provided by the embodiments of the present disclosure can, on the one hand, effectively suppress the propagation of inter-frame artifacts by using reference frames to supplement the spatio-temporal domain information of target frames through multi-frame alignment processing, thereby removing artifacts in high quality and improving the video quality of the target video; on the other hand, the multi-frame alignment features are extracted through multi-receptive field feature extraction by multiple parallel convolution branches, and the effects of video enhancement are improved again through the synergistic effect of the multiple parallel convolution branches, that is, the global blocking effect suppression of the large receptive field branch and the local texture repair of the small receptive field branch; in addition, the design of the equivalent convolution kernel can reduce the computational complexity. In summary, the above method can realize the joint optimization of artifact elimination effect, temporal consistency and computational efficiency while reducing the computational complexity, and realize efficient, real-time and high-quality video enhancement.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and do not limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0020] The drawings incorporated into the specification and forming part of the specification, show embodiments consistent with the present disclosure, and together with the specification, serve to explain the principles of the present disclosure. It is obvious that the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.
[0021] Figure 1 A scene schematic diagram of a video enhancement method or a video enhancement device that can be applied to the embodiments of the present disclosure is shown.
[0022] Figure 2 is a flowchart of a video enhancement method according to an exemplary embodiment.
[0023] Figure 3 is a structural schematic diagram of a target network model according to an exemplary embodiment.
[0024] Figure 4 is a structural schematic diagram of a feature extraction module according to an exemplary embodiment.
[0025] Figure 5 is a structural schematic diagram of an image enhancement module according to an exemplary embodiment.
[0026] Figure 6 is a flowchart of a weight merging method according to an exemplary embodiment.
[0027] Figure 7 is a structural schematic diagram of a feature extraction sub-module according to an exemplary embodiment.
[0028] Figure 8 is a structural diagram of reconstructing convolution according to an example embodiment.
[0029] Figure 9 is a flow chart of a weight merging method according to an example embodiment.
[0030] Figure 10 is a flow chart of a model training method according to an example embodiment.
[0031] Figure 11 is a flow chart of a multi-frame alignment method according to an example embodiment.
[0032] Figure 12 is a structural diagram of a multi-frame alignment module according to an example embodiment.
[0033] Figure 13 is a flow chart of a multi-frame alignment method according to an example embodiment.
[0034] Figure 14 is a block diagram of a video enhancement device according to an example embodiment.
[0035] Figure 15 shows a structural diagram of an electronic device suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0036] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments can, however, be implemented in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example embodiments to those skilled in the art. Like reference numerals refer to like elements throughout the figures, and description of the same or similar elements can be omitted.
[0037] Those skilled in the art know that the embodiments of the present disclosure can be a system, a device, an apparatus, a method or a computer program product. Therefore, the present disclosure can be embodied in the form of a complete hardware, a complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0038] The features, structures or characteristics described in the present disclosure can be incorporated in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the present disclosure. One skilled in the relevant art will recognize, however, that the techniques of the present disclosure can be practiced without one or more of the specific details, or with other methods, components, devices, steps, etc. In other instances, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0039] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0040] The accompanying drawings are merely schematic illustrations of the present disclosure, in which the same reference numerals refer to the same or similar parts, and thus repeated description thereof will be omitted. Some of the block diagrams shown in the drawings do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0041] The flowcharts shown in the drawings are merely exemplary illustrations, and do not necessarily include all contents and steps, nor are they necessarily executed in the order described. For example, some steps can be further divided, and some steps can be combined or partially combined, so the actual execution order can be changed according to the actual situation.
[0042] In the description of the present disclosure, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this document is merely a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, "one or more" means one or more, and "multiple" means two or more. "First", "second", etc. do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different; the terms "include", "contain" and "have" are used to mean open inclusion and mean that in addition to the listed elements / components / etc. there can be other elements / components / etc.
[0043] In order to enable a more clear understanding of the above-mentioned purposes, features and advantages of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, and it should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0044] It should be noted that in the technical solutions of the present disclosure, the collection, collection, updating, analysis, processing, use, transmission, storage, etc. of user personal information are in line with the relevant legal regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain user personal information security and network security.
[0045] First, some terms related to the embodiments of the present disclosure will be explained below to facilitate understanding by those skilled in the art.
[0046] Coding artifact: refers to a visual or auditory distortion phenomenon caused by the imperfection of the coding algorithm or the loss of information in the compression process during digital signal processing, image / video compression or data transmission.
[0047] The foregoing introduces some concepts related to the embodiments of the present disclosure, and the technical features related to the embodiments of the present disclosure are introduced below.
[0048] At present, there are mainly the following technical solutions in the industry for the artifact problem generated in the video coding process:
[0049] (1) Traditional filter-based post-processing method. In order to reduce the block effect and artifacts generated in the compression process, an in-loop filter module is usually integrated in the encoder, such as a deblocking filter and a sample adaptive offset. Although these in-loop filtering methods can reduce artifacts and improve video subjective quality to some extent, they mainly rely on local pixel statistics or simple linear filtering, and it is difficult to fully exploit more complex spatial structures and temporal information, so it is still difficult to completely eliminate coding artifacts in complex scenes. In practical applications, residual artifacts often occur, and more efficient and intelligent post-processing enhancement methods are needed to supplement and optimize.
[0050] (2) Deep learning-based image processing method. Convolutional neural network-based image enhancement methods are widely used in artifact elimination. This kind of method can effectively extract spatial features and improve the visual quality of single-frame images, but it only processes single frames and does not consider the temporal continuity of video, making it difficult to eliminate artifacts that propagate across frames, and may affect the temporal consistency of video.
[0051] (3) Multi-frame based enhancement method. To further improve the artifact removal effect, some studies try to use multi-frame information to fuse the time domain features through 3D convolution, optical flow estimation, etc. Although this kind of method performs excellently in artifact removal and detail restoration, the model is usually complex, with large parameter quantity and high computational resource consumption, which is difficult to meet the real-time processing and low-power deployment requirements in actual scenarios.
[0052] To solve the above problems, the method proposed by the present application considers the efficient fusion of spatial and temporal information in design, realizes low computational resource consumption through a lightweight network structure, and significantly improves the artifact removal effect, video temporal consistency and processing efficiency, overcoming the limitations of traditional filtering, single-frame deep learning and multi-frame enhancement methods, and being more suitable for high-quality, real-time video enhancement requirements in actual scenarios.
[0053] The example embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0054] Figure 1 A scene schematic diagram of a video enhancement method or a video enhancement device that can be applied to the embodiments of the present disclosure is shown.
[0055] Reference is made to Figure 1 which shows a schematic diagram of an implementation environment provided by an example embodiment of the present disclosure.
[0056] As shown in Figure 1 , the system architecture 100 can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a communication link medium between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc.
[0057] A user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Among them, the terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, wearable devices, virtual reality devices, smart home devices, etc.
[0058] The server 105 can be a server that provides various services, such as a background management server that provides support for the devices operated by the user using the terminal devices 101, 102, 103. The background management server can analyze and process the received request data, etc., and feed back the processing result to the terminal device.
[0059] The server can be a standalone physical server, a server cluster composed of multiple physical servers, or a distributed system, and can also be a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms, and the like, and the present disclosure does not limit the same.
[0060] The server 105 can obtain, for example, a target frame and at least one reference frame of the target frame from a video to be enhanced, the at least one reference frame including at least one of the first N frames of the target frame and the last M frames of the target frame, N being a frame number greater than or equal to 1, and M being an integer greater than or equal to 1. The server 105 can perform feature extraction on the target frame and the at least one reference frame, for example, by a feature extraction module of the target network model, to obtain target frame features corresponding to the target frame and reference frame features corresponding to each reference frame. The server 105 can perform multi-frame alignment processing on the target frame features and the reference frame features, for example, by a multi-frame alignment module of the target network model, to obtain multi-frame alignment features. The server 105 can perform image enhancement processing on the multi-frame alignment features, for example, by an image enhancement module of the target network model, to obtain image enhancement features. The server 105 can add the image enhancement features and the target frame pixel by pixel, for example, to obtain an enhanced target frame. The server 105 can include, for example, at least one representative feature extraction submodule in the image enhancement module, the representative feature extraction submodule including a plurality of parallel convolution branches, each convolution branch including at least one convolution kernel. The image enhancement module performs image enhancement processing on the multi-frame alignment features to obtain image enhancement features, including merging the convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight. The equivalent convolution kernel weight is used to perform convolution processing on the multi-frame alignment features to obtain equivalent convolution features, so as to determine the image enhancement features based on the equivalent convolution features.
[0061] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the system is only illustrative, and the server 105 can be a server of one entity, or can be composed of multiple servers. According to actual needs, the system can have any number of terminal devices, networks, and servers.
[0062] Those skilled in the art can know that Figure 1 The number of terminals, servers, networks, and network-side devices in the system is only illustrative, and according to actual needs, the system can have any number of terminals, networks, and servers. The embodiments of the present disclosure do not limit the same.
[0063] Under the above system architecture, the video enhancement method provided in the embodiments of the present disclosure can be executed by any electronic device with computing processing capability.
[0064] Figure 2 is a flowchart of a video enhancement method according to an exemplary embodiment. The method provided in the embodiments of the present disclosure can be executed by any electronic device with computing processing capability, for example, the method can be executed by the server or the terminal device in the above embodiments, or can be executed by the server and the terminal device together. In the following embodiments, the server is taken as an example for illustration, but the present disclosure is not limited thereto. Figure 1
[0065] Referring to Figure 2 , the video enhancement method provided in the embodiments of the present disclosure can include the following steps.
[0066] Step S202, obtaining a target frame and at least one reference frame of the target frame from a video to be enhanced, the at least one reference frame including at least one of the first N frames before the target frame and the last M frames after the target frame, N being a frame number greater than or equal to 1, and M being an integer greater than or equal to 1.
[0067] In some embodiments, the video to be enhanced can include a plurality of images, one frame can be selected from the plurality of images as the target frame, and at least one of the first N frames before the target frame and the last M frames after the target frame can be selected from the plurality of images as the reference frame of the target frame.
[0068] The above-mentioned video to be enhanced can be a real-time video being decoded, and the above-mentioned target frame can be a video currently being decoded or just decoded in the real-time video.
[0069] In the present embodiment, the first N frames before the target frame and the last M frames after the target frame can be taken as an example for explanation and illustration. Wherein, N and M can both be 3.
[0070] Step S204, performing feature extraction on the target frame and the at least one reference frame through a feature extraction module of a target network model, to obtain target frame features corresponding to the target frame and reference frame features corresponding to each reference frame.
[0071] In some embodiments, the target network model can be a network model for image enhancement proposed in the present application.
[0072] Figure 3 is a structural schematic diagram of a target network model according to an exemplary embodiment.
[0073] As Figure 3 As shown, the target network model can include a feature extraction module 301, a multi-frame alignment module 302, and a quality enhancement module 303.
[0074] The feature extraction network module 301 can be configured to perform feature extraction on an image (e.g., a target frame or a reference frame).
[0075] Figure 4 FIG. 3 is a structural schematic diagram of a feature extraction module according to an example embodiment.
[0076] The feature extraction module 301 can be, for example, a U-Net architecture as shown in FIG. 3. Figure 4 As a classical encoder-decoder structure, the U-Net can effectively capture feature information at different levels. Specifically, the input target frame and its adjacent reference frames are first subjected to multi-layer convolution and down-sampling operations to gradually extract low-dimensional deep features. Subsequently, these features are gradually restored through a symmetric up-sampling path, while combining the features of the corresponding layers of the encoder with the decoder through a skip connection to retain more spatial detail information.
[0077] At step S206, the target frame features and the reference frame features are subjected to multi-frame alignment processing by the multi-frame alignment module of the target network model to obtain multi-frame alignment features.
[0078] In some embodiments, the multi-frame alignment can be performed based on a feature point matching method, for example, by Scale-Invariant Feature Transform (SIFT) to perform multi-frame alignment; the multi-frame alignment can also be performed based on a light flow method; the multi-frame alignment can also be performed based on a phase correlation method or an optimization loss function method; the present application does not limit this.
[0079] In some embodiments, the multi-frame alignment can also be performed by the multi-frame alignment method proposed in the present application, which will be described in detail in Figure 13 , and the present embodiment will not be described again.
[0080] At step S208, the multi-frame alignment features are subjected to image enhancement processing by the image enhancement module of the target network model to obtain image enhancement features.
[0081] In some embodiments, the image enhancement module further includes a pixel rearrangement structure, a pixel shuffling structure, and a feature fusion structure Figure 5The feature fusion structure includes at least one representative feature extraction submodule; the image enhancement module performs image enhancement processing on the multi-frame alignment features to obtain image enhancement features, including: performing down-sampling processing on the multi-frame alignment features through the pixel rearrangement structure to obtain down-sampled features; performing feature fusion processing on the down-sampled features through the feature fusion structure to obtain fused features; and performing up-sampling processing on the fused features through the pixel shuffling structure to obtain the image enhancement features.
[0082] In some embodiments, the feature fusion structure in the image enhancement module can include at least one representative feature extraction submodule, which can include a plurality of parallel convolution branches, each convolution branch including at least one convolution kernel.
[0083] Figure 5 FIG. 1 is a structural schematic diagram of an image enhancement module according to an exemplary embodiment.
[0084] As shown in FIG. 1, the image enhancement module can include a feature fusion structure, which can include at least one feature extraction submodule (as shown by modules 1-6 in FIG. 1). Figure 5 Figure 5 As shown in FIG. 1, the image enhancement module can include a feature fusion structure, which can include at least one feature extraction submodule (as shown by modules 1-6 in FIG. 1).
[0085] As shown in FIG. 1, the image enhancement module can include a feature fusion structure, which can include at least one feature extraction submodule (as shown by modules 1-6 in FIG. 1). Figure 5 As shown in FIG. 1, the image enhancement module can include a feature fusion structure, which can include at least one feature extraction submodule (as shown by modules 1-6 in FIG. 1).
[0086] In some embodiments, the image enhancement module performs image enhancement processing on the multi-frame alignment features to obtain image enhancement features, which can include: merging the convolution kernel weights of the convolution kernels on the multiple branches to obtain an equivalent convolution kernel weight; and performing convolution processing on the multi-frame alignment features through the equivalent convolution kernel weight to obtain equivalent convolution features, so as to determine the image enhancement features according to the equivalent convolution features.
[0087] In some embodiments, when the convolution results of the multiple branches are directly added at the output end, and the convolution kernels of all the branches have the same input / output channel number, stride, and padding, the weights at the same positions can be directly summed and processed. For example, if there are two branch convolution kernel weights W1 and W2, the merged equivalent convolution kernel weight Weq = W1 + W2.
[0088] In some embodiments, when the outputs of the branches are spliced in the channel dimension, and the convolution kernels of all the branches have the same spatial size (such as 3x3), but the output channels can be different; the weights can be spliced along the output channel dimension, which is equivalent to a larger convolution kernel.
[0089] Of course, other convolution kernel weight merging methods are also within the protection scope of the present application, and the present application does not limit this.
[0090] In addition, there can be only one convolution kernel on each of the above convolution branches, or there can be multiple convolution kernels in series, and the present application does not limit this.
[0091] In some embodiments, the convolution kernels (or equivalent convolution kernels) on the above multiple parallel convolution branches can be the same, different, or partially the same, and the present application does not limit this.
[0092] It can be understood that if the convolution kernels on the branches are different, the receptive fields corresponding to the branches are also different. Among them, the large receptive field branch globally suppresses the block effect, and the small receptive field branch locally repairs the texture. Through the above scheme, the multi-branch design can comprehensively perceive the features of the multi-frame alignment features.
[0093] The above image enhancement module first dynamically merges the convolution kernel weights of the multiple parallel convolution branches into a single equivalent convolution kernel, which has the following technical effects.
[0094] (1) Reduced calculation: merging N parallel convolution operations into 1 equivalent convolution, the theoretical calculation complexity is reduced to 1 / N of the original scheme; memory access optimization: avoiding repeated storage and reading of multi-branch feature maps, reducing memory bandwidth occupation; hardware friendliness: the equivalent convolution kernel can be directly mapped to a standard convolution unit, compatible with the instruction set of mainstream AI accelerators (such as NPU / GPU).
[0095] (2) Dynamic receptive field integration: the equivalent convolution kernel automatically inherits the multi-scale characteristics of each branch (such as containing both 5x5 global perception and 3x3 local refinement capability); weight adaptation: learning the optimal merging strategy through end-to-end training (experiments show that the equivalent kernel retains small kernel characteristics in texture areas and activates large kernel weights in flat areas). In summary, through the equivalent merging of convolution kernel weights, the multi-branch structure is converted into a single standard convolution under the guarantee of mathematical equivalence, which not only retains the advantages of multi-receptive field feature extraction, but also significantly reduces the calculation load, solving the core pain points of traditional multi-branch network calculation redundancy and poor hardware adaptability.
[0096] Step S210, pixel-by-pixel addition of the image enhancement feature and the target frame to obtain an enhanced target frame.
[0097] Through the above method, image features with efficient spatio-temporal information fusion can be obtained, so that the image enhancement feature fully highlights the key position (such as the artifact position), and the artifact position of the target frame can be supplemented by referring to the reference frame.
[0098] Then, by pixel-by-pixel addition of the image enhancement features and the target frame, the key positions in the target frame pixels can be enhanced, such as the positions of the target frame that are blocked or blurred.
[0099] The above method can, on the one hand, use reference frames to supplement the spatio-temporal domain information of the target frame through multi-frame alignment processing, effectively suppress the propagation of inter-frame artifacts, and thus effectively remove artifacts and improve the video quality of the target video. On the other hand, the multi-frame alignment features are extracted by multiple parallel convolution branches with multi-receptive field, and the synergistic effect of the multiple parallel convolution branches further improves the video enhancement effect. In addition, the design of the equivalent convolution kernel can reduce the computational complexity. In summary, the above method can jointly optimize the artifact elimination effect, temporal consistency and computational efficiency while reducing the computational complexity, and achieve efficient, real-time and high-quality video enhancement.
[0100] The above embodiment proposes a lightweight video enhancement network for encoding artifact elimination, and the overall network structure can be as shown in Figure 3 The network mainly includes three parts: a feature extraction module, a multi-frame alignment module and a quality enhancement module. The feature extraction module takes the target frame and its adjacent reference frames as input, and contains seven frames in total, of which the fourth frame in the middle is the target frame, and the first three frames and the last three frames are reference frames. After the feature extraction module, rich spatial feature information can be obtained. Then, the extracted spatial feature information is input into the multi-frame alignment module together with the target frame and the reference frames. In this module, the inter-frame context information can be aligned, and the result is finally transmitted to the quality enhancement module. In the quality enhancement module, the features are processed by a series of convolutional neural networks, and the learned enhancement residual is added to the target frame pixel by pixel, and finally the enhanced target frame is obtained.
[0101] The above embodiment proposes a lightweight, re-parameterized spatio-temporal data fusion video enhancement network for the problem of video quality degradation caused by video encoding, which combines temporal and spatial domain information, effectively eliminates encoding artifacts while ensuring efficient real-time inference and small model parameter amount, and improves the visual subjective quality of the video.
[0102] Figure 6 is a flowchart of a weight merging method according to an example embodiment.
[0103] In some embodiments, the equivalent convolution kernel weights can include first equivalent weights.
[0104] Figure 7 is a structural schematic diagram of a feature extraction submodule according to an example embodiment.
[0105] As Figure 7 indicated, the feature extraction submodule can include a reconstruction convolution, an extended convolution, and an attention mechanism, etc. Among them, the extended convolution can include a plurality of serial convolution kernels.
[0106] In the training process of the target network model, the plurality of serial convolution kernels of the extended convolution can independently perform parameter iteration. However, in the inference process of the target network model, the plurality of serial convolution kernels of the extended convolution can be merged to generate an equivalent convolution kernel, thereby reducing the amount of calculation and improving the calculation efficiency.
[0107] In some embodiments, if the plurality of serial convolution kernels of the extended convolution are of the same size, the weights of the plurality of convolution kernels can be directly added to obtain the plurality of serial convolution kernels of the extended convolution; if the plurality of serial convolution kernels of the extended convolution are of different sizes, the weights of the respective convolution kernels can be respectively padded by zero to make the sizes of the padded convolution kernel weights the same, and then the padded convolution kernel weights are added to obtain the plurality of serial convolution kernels of the extended convolution.
[0108] In some embodiments, the representative feature extraction submodule described above can include at least one reconstruction convolution.
[0109] Figure 8 is a structural schematic diagram of a reconstruction convolution according to an exemplary embodiment.
[0110] Referring to Figure 8 , a reconstruction convolution can include a plurality of branches, and one branch can include one or more convolution kernels. Among them, a branch can include parallel convolution and / or serial convolution kernels, in short, the serial and parallel convolution kernels can be distributed on a branch. For example, a branch can include two parallel sub-branches, one sub-branch includes two parallel convolution kernels, another sub-branch includes two grandchild branches, one grandchild branch includes two serial convolution kernels, and the other grandchild branch can include two parallel convolution kernels, and so on, which will not be described herein. Among them, the merging between branches can be connection merging, can also be splicing merging, can also be weighted merging, etc., which will not be described herein.
[0111] In some embodiments, the plurality of branches described above can include a first branch (such as the branch indicated by 801 in Figure 8 ) and a second branch (such as the branch shown by 802 in Figure 8 ), the first branch can include a first convolution kernel, and the second branch can include a second convolution kernel.
[0112] Referring to Figure 6 , the weight merging method described above can include the following steps.
[0113] Step S602, the first convolution kernel weight of the first convolution kernel and the second convolution kernel weight of the second convolution kernel are zero-padded, so that the matrix size of the padded first convolution kernel weight and the padded second convolution kernel weight is the same.
[0114] As shown in Figure 8 , the first convolution kernel and the second convolution kernel are different in size, then the first convolution kernel and the second convolution kernel can be zero-padded respectively, so that the matrix size of the padded first convolution kernel weight and the padded second convolution kernel weight is the same (for example, both are padded to 3x3 size).
[0115] Step S604, the first convolution kernel weight and the second convolution kernel weight are added channel by channel to determine the first equivalent weight.
[0116] Figure 9 is a flowchart of a weight merging method according to an exemplary embodiment.
[0117] In some embodiments, the equivalent convolution kernel weight includes a second equivalent weight, and the plurality of branches includes a third branch (as indicated by 803 in FIG. 8), and the third branch includes a third convolution kernel and a fourth convolution kernel in series.
[0118] In some embodiments, if the third convolution kernel and the fourth convolution kernel are the same in size (such as the two convolution kernels shown in 803 in Figure 8 , the third convolution kernel weight and the fourth convolution kernel weight are directly added to obtain the second equivalent weight.
[0119] If the third convolution kernel and the fourth convolution kernel are different in size, the third convolution kernel weight of the third convolution kernel and the fourth convolution kernel weight of the fourth convolution kernel are zero-padded, so that the matrix size of the padded third convolution kernel weight and the padded fourth convolution kernel weight is the same.
[0120] Referring to Figure 9 , the above weight merging method can include the following steps.
[0121] Step S902, the third convolution kernel weight of the third convolution kernel and the fourth convolution kernel weight of the fourth convolution kernel are zero-padded, so that the matrix size of the padded third convolution kernel weight and the padded fourth convolution kernel weight is the same.
[0122] Step S904, the third convolution kernel weight and the fourth convolution kernel weight are spatially convolved channel by channel to determine the second equivalent weight.
[0123] It should be noted that the convolution kernel weights on the above multiple branches are independently updated in the training phase of the target network model, and the convolution kernels on each branch are only merged for processing in the model inference phase.
[0124] In summary, the representative feature extraction submodule is a multi-scale feature fusion module, and a module structure diagram is as shown in Figure 7 The module adopts a multi-branch structure (a multi-branch as shown in Figure 8 ), extracts features with different receptive fields through multiple parallel convolution kernels of different sizes, and realizes efficient feature fusion through local residual connection. In addition, the representative feature extraction submodule decouples the training process from the inference process, effectively avoiding the additional parameters and computational burden brought by the multi-branch structure. At the end of the representative feature extraction submodule, an attention mechanism is also introduced to improve the discriminability and relevance of the features. The attention mechanism can enable the model to dynamically focus on important areas and suppress irrelevant background noise, thereby improving the accuracy of feature expression. Specifically, the attention module weights the importance of different feature maps, enabling the network to focus on key information in complex scenes and significantly improving the robustness and generalization ability of the model.
[0125] The quality enhancement module in the above embodiment mainly consists of pixel rearrangement (pixel unshuffling (downsampling) and pixel shuffling (upsampling)) and a feature fusion module, and a structure thereof is as shown in Figure 5 First, the input is subjected to a pixel unshuffling (downsampling) operation to obtain downsampled features, effectively reducing the computational load of the model. Pixel unshuffling reduces the spatial resolution of the input while redistributing the spatial information to more channels, thereby maintaining the total amount of features. In this way, while maintaining the number of channels in the subsequent convolution layer, the computational load of the convolution operation is significantly reduced due to the smaller spatial size of the feature map, improving the efficiency of the model. Subsequently, the input is subjected to a convolution layer to adjust the number of channels to meet the needs of subsequent processing. Next, the features are fed into six representative feature extraction submodules. These modules transform and refine the features through a multi-path structure, further improving the feature extraction capability of the network. After multi-branch processing, the features of multiple branches are fused through splicing operation, and then subjected to a convolution operation again to integrate the fused features, further improving the quality and consistency of the features. Finally, the module restores the spatial dimension of the image through pixel shuffling operation, and adds the learned residual information to the target frame to obtain the final enhancement result.
[0126] Figure 10 is a flowchart of a model training method according to an example embodiment.
[0127] Referring to Figure 10 , the above model training method can include the following steps.
[0128] Step S1002, obtaining a training target frame and at least one training reference frame of the training target frame from the training video.
[0129] In some embodiments, the training video can be a video in which artifacts are added in a known clear video.
[0130] After obtaining the training target frame and at least one training reference frame of the training target frame from the training video, the clear frame can be obtained from the above clear frame according to the position of the target frame in the training video.
[0131] Step S1004, obtaining training image enhancement features by performing image enhancement processing on the training target frame and the at least one training reference frame through the target network model.
[0132] In some embodiments, the training image enhancement features can be determined according to the image enhancement feature determination method shown in Figures 2-9 The above embodiment will not be described again.
[0133] It should be noted that in the inference process, the multiple branches in the image enhancement module need to be merged; and in the training process, the parameters of the multiple branches in the image enhancement module need to be updated separately.
[0134] Step S1006, adding the training image enhancement features and the training target frame pixel by pixel to obtain an enhanced training target frame.
[0135] Step S1008, determining a target loss function according to the enhanced training target frame.
[0136] In some embodiments, the target loss function can be determined according to the enhanced training target frame and the clear frame.
[0137] Step S1010, adjusting the parameters of the target network model through the target loss function, wherein the parameters of the multiple parallel convolution branches in the image enhancement module are adjusted respectively.
[0138] Through the above method, the parameters of the multiple branches of the image enhancement module can be adjusted respectively, so that different feature information of the image is perceived through the convolution kernels of different receptive fields—large receptive field branch global inhibition blocking effect, small receptive field branch local repair texture.
[0139] Figure 11 is a flowchart of a multi-frame alignment method according to an example embodiment.
[0140] Figure 12 is a structural schematic diagram of a multi-frame alignment module according to an example embodiment.
[0141] In some embodiments, the multi-frame alignment module can include at least two alignment sub-modules (such as the two modules indicated by 1201 and 1202 in FIG. 12), and the at least two alignment sub-modules include a first alignment sub-module (such as the module indicated by 1201 in FIG. 12) and a second alignment sub-module (such as the module indicated by 1202 in FIG. 12). Figure 12 In some embodiments, the multi-frame alignment module can include at least two alignment sub-modules (such as the two modules indicated by 1201 and 1202 in FIG. 12), and the at least two alignment sub-modules include a first alignment sub-module (such as the module indicated by 1201 in FIG. 12) and a second alignment sub-module (such as the module indicated by 1202 in FIG. 12). Figure 12 In some embodiments, the multi-frame alignment module can include at least two alignment sub-modules (such as the two modules indicated by 1201 and 1202 in FIG. 12), and the at least two alignment sub-modules include a first alignment sub-module (such as the module indicated by 1201 in FIG. 12) and a second alignment sub-module (such as the module indicated by 1202 in FIG. 12). Figure 12 In some embodiments, the multi-frame alignment module can include at least two alignment sub-modules (such as the two modules indicated by 1201 and 1202 in FIG. 12), and the at least two alignment sub-modules include a first alignment sub-module (such as the module indicated by 1201 in FIG. 12) and a second alignment sub-module (such as the module indicated by 1202 in FIG. 12).
[0142] Referring to FIG. 12, Figure 11 the multi-frame alignment method described above can include the following steps.
[0143] In step S1102, the target frame feature and each reference frame feature are processed by the first alignment sub-module for multi-frame alignment to obtain first alignment features.
[0144] In step S1104, the target frame feature and each reference frame feature are processed by the second alignment sub-module for multi-frame alignment to obtain second alignment features; the kernel sizes of the first alignment sub-module and the second alignment sub-module are different.
[0145] In step S1106, the first alignment features and the second alignment features are fused to obtain multi-frame alignment features.
[0146] The method described above can perform multi-frame alignment through two branches respectively, and then perform feature merging. Through this double-path design, the comprehensiveness and accuracy of multi-frame alignment are significantly improved while maintaining the computing efficiency.
[0147] Figure 13 FIG. 13 is a flowchart of a multi-frame alignment method according to an example embodiment.
[0148] In some embodiments, the first alignment sub-module (such as the module indicated by 1201 in FIG. 12) includes a first bias prediction sub-structure (such as the convolution structure indicated by 12011 in FIG. 12) and a first convolution sub-structure (such as the convolution structure indicated by 1202 in FIG. 12). Figure 12 In some embodiments, the first alignment sub-module (such as the module indicated by 1201 in FIG. 12) includes a first bias prediction sub-structure (such as the convolution structure indicated by 12011 in FIG. 12) and a first convolution sub-structure (such as the convolution structure indicated by 1202 in FIG. 12). Figure 12 In some embodiments, the first alignment sub-module (such as the module indicated by 1201 in FIG. 12) includes a first bias prediction sub-structure (such as the convolution structure indicated by 12011 in FIG. 12) and a first convolution sub-structure (such as the convolution structure indicated by 1202 in FIG. 12). Figure 12 In some embodiments, the first alignment sub-module (such as the module indicated by 1201 in FIG. 12) includes a first bias prediction sub-structure (such as the convolution structure indicated by 12011 in FIG. 12) and a first convolution sub-structure (such as the convolution structure indicated by 1202 in FIG. 12).
[0149] Referring to FIG. 12, Figure 13 the multi-frame alignment method described above can include the following steps.
[0150] In step S1302, the target frame feature and at least one reference frame are processed by the first bias prediction sub-structure for bias prediction to determine a first offset of the actual sampling points of the convolution sub-structure on each reference frame feature compared to the standard sampling points; wherein the standard sampling points are determined according to the kernel of the convolution sub-structure.
[0151] In deep learning, the standard sampling points of a convolution kernel refer to the pixels or feature points covered by the convolution kernel at each position of the convolution kernel sliding on an input feature map (or image) during a convolution operation. The positions and number of the sampling points are determined by the size of the convolution kernel, the stride, the padding, and the dilation rate. For example, for a convolution kernel with a size of k x k (e.g., 3 x 3), the standard sampling points are k x k adjacent pixels / feature points covered on the input feature map with a fixed stride and without dilation.
[0152] Generally, the positions of the same object in the target frame and the reference frame are the same, but if there is an artifact, the position of the object in the target frame can be offset.
[0153] Through the calculation of the first offset, it can be determined how much the position of the object in the reference frame is offset relative to the position in the target frame.
[0154] In step S1304, the first convolution substructure determines the first actual sampling points according to the standard sampling points and the first offset.
[0155] In some embodiments, each standard sampling point can be added with the corresponding first offset to determine the actual sampling point corresponding to the standard sampling point.
[0156] Wherein, the standard sampling points are regularly distributed, and the actual sampling points determined according to the first offset can be irregularly distributed.
[0157] Through the above method, the positions of the same object in the reference frame and the target frame can be aligned.
[0158] In step S1306, the first convolution substructure respectively convolves each reference frame feature according to the first actual sampling points to obtain the first aligned reference frame features.
[0159] Through the convolution of each reference frame according to the actual sampling points, the same object in different reference frames can be aligned, so that the features of the same position extracted finally correspond to the same object.
[0160] In step S1308, the first aligned reference frame features are fused with the target frame features to obtain the first alignment features.
[0161] Through the above method, the objects in each reference frame can be aligned with the target frame before extracting the features, so that the target frame can be supplemented with key information through the reference frame.
[0162] Reference Figure 13Those skilled in the art can understand how the second alignment sub-module works in the embodiments shown, and thus the present application will not be described here.
[0163] Reference Figure 12 The input in the multi-frame alignment module comes from the feature extraction module and the continuous frames (target frame + reference frame) of the model input. For the output of the feature extraction module, offset can be extracted through convolution operations with kernel sizes of 1 and 3. Then, these offsets are input into the corresponding convolution structure together with the continuous multiple frames for convolution processing. Since there is a high correlation between the continuous frames, the offset prediction of a certain frame can utilize the information of other frames, thereby providing additional reference information to enhance the image quality (by mining the direct spatio-temporal correlation of continuous frames, the system can make up for the deficiency of single-frame information when predicting the offset of the current frame by means of the feature information of adjacent frames). Next, the alignment features obtained through the two paths can be spliced and combined, and an attention mechanism can be introduced to weight these features. This mechanism can enable the network to adaptively assign weights according to the importance of the alignment features, thereby optimizing the information fusion effect. Finally, the fused features are input into the quality enhancement module through convolution operation to further improve the overall quality of the image.
[0164] The following will describe the experimental effects of the video enhancement method composed of the above-mentioned modules. Figures 2-13
[0165] In the training phase, 64x64 segments can be randomly cropped from the original and corresponding compressed videos as training samples. In order to make full use of these samples, data enhancement techniques (such as rotation or flipping) are further applied. The initial value of the learning rate is set to 9 -4 , and remains unchanged throughout the training process. The mean squared error can be selected as the loss function. In the evaluation process, quality enhancement can be applied only to the Y channel in the YUV space. The incremental peak signal-to-noise ratio (ΔPSNR) is used to evaluate the quality enhancement performance. Next, the effects of the present application will be shown.
[0166] Table 1 shows the PSNR (peak signal-to-noise ratio) improvement of different methods on test videos. It can be seen that the method of the present application achieves a higher average PSNR on test videos, showing a more excellent quality enhancement effect compared to other comparative methods. This indicates that the method proposed in the present application has strong ability in video compression distortion recovery.
[0167] The following will explain the terms involved in Table 1.
[0168] AR-CNN (Artifact Reduction Convolutional Neural Network) is a deep learning-based image / video compression artifact reduction method, which is specifically designed to fix the image quality degradation caused by lossy compression (e.g., JPEG, H.264, etc.).
[0169] DnCNN (Denoising Convolutional Neural Network) is a deep learning-based image denoising method.
[0170] RNAN (Residual Non-local Attention Network) is a deep learning-based image restoration network, mainly used for image super-resolution, denoising, deblurring, etc.
[0171] RNAN (Residual Non-Local Attention Network) is a deep learning-based image restoration network, mainly used for image super-resolution (Super-Resolution), denoising (Denoising), and compression artifact reduction (Compression Artifact Reduction) tasks.
[0172] MFQE (Multi-Frame Quality Enhancement) is a multi-frame information-based video quality enhancement method, mainly used to improve the visual quality of compressed videos.
[0173] Table 1 Results of different methods under test video ΔPSNR (dB)
[0174]
[0175]
[0176] Table 2 and Table 3 respectively show the comparison of the parameter quantity and the inference speed of the method of the present application and other methods under different resolutions. It can be seen that, except for MFQE2.0, the model of the present application is significantly smaller in parameter quantity than other methods. In terms of inference speed, the present application is tested on NVIDIA RTX3090 graphics card, and the results show that under different resolutions, the method of the present application is significantly superior to other video enhancement methods in inference speed. Taking 980P resolution as an example, under the premise of ensuring the average quality improvement, compared with MFQE2.0, the inference speed of the method of the present application is improved by 18.9%, which proves the universality and practical application value of the method of the present application. Overall, the method of the present application performs outstandingly in terms of calculation efficiency, quality improvement and resource consumption, and is very suitable for practical scenes such as coding distortion elimination.
[0177] Table 2 Comparison of parameter quantity of different method models
[0178]
[0179] Table 3 Comparison of inference speed (IPS) of different methods under different resolutions
[0180]
[0181] It should be particularly pointed out that each step in each of the above video enhancement methods can be crossed, replaced, added, deleted and reduced. Therefore, these reasonable permutations and combinations of the video enhancement method should also belong to the protection scope of the present disclosure, and the protection scope of the present disclosure should not be limited to the above-mentioned embodiments.
[0182] Based on the same inventive concept, the present disclosure also provides a video enhancement device, as follows. Since the principles of solving problems of the device embodiments are similar to those of the above-mentioned method embodiments, the implementation of the device embodiments can be referred to the implementation of the above-mentioned method embodiments, and the repeated parts will not be repeated.
[0183] Figure 14 is a block diagram of a video enhancement device according to an exemplary embodiment. Referring to Figure 14 , the video enhancement device 1400 provided by the embodiments of the present disclosure can include an image acquisition module 1401, a feature extraction module 1402, a multi-frame alignment feature determination module 1403, an image enhancement module 1404 and a pixel addition module 1405.
[0184] The image acquisition module 1401 can be configured to acquire a target frame and at least one reference frame of the target frame from the video to be enhanced, the at least one reference frame including at least one of the first N frames of the target frame and the last M frames of the target frame, N being a frame number greater than or equal to 1, and M being an integer greater than or equal to 1; the feature extraction module 1402 can be configured to perform feature extraction on the target frame and the at least one reference frame by using a feature extraction module of the target network model, to obtain target frame features corresponding to the target frame and reference frame features corresponding to each reference frame; the multi-frame alignment feature determination module 1403 can be configured to perform multi-frame alignment processing on the target frame features and the reference frame features by using a multi-frame alignment module of the target network model, to obtain multi-frame alignment features; the image enhancement module 1404 can be configured to perform image enhancement processing on the multi-frame alignment features by using an image enhancement module of the target network model, to obtain image enhancement features; and the pixel addition module 1405 can be configured to add the image enhancement features and the target frame pixel by pixel, to obtain an enhanced target frame; the image enhancement module includes at least one representative feature extraction sub-module, the representative feature extraction sub-module includes a plurality of parallel convolution branches, each convolution branch includes at least one convolution kernel, the image enhancement module performs image enhancement processing on the multi-frame alignment features to obtain the image enhancement features, including: combining convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight; and performing convolution processing on the multi-frame alignment features by using the equivalent convolution kernel weight to obtain equivalent convolution features, so as to determine the image enhancement features according to the equivalent convolution features.
[0185] It should be noted that the image acquisition module 1401, the feature extraction module 1402, the multi-frame alignment feature determination module 1403, the image enhancement module 1404, and the pixel addition module 1405 correspond to S202-S29 in the method embodiment, and have the same examples and application scenarios as the corresponding steps, but are not limited to the disclosure of the above method embodiment. It should be noted that the above modules as part of the device can be executed in a computer system such as a group of computer executable instructions.
[0186] In some embodiments, the equivalent convolution kernel weight includes a first equivalent weight, the plurality of branches includes a first branch and a second branch, the first branch includes a first convolution kernel, and the second branch includes a second convolution kernel; and the image enhancement module 1404 can include a first padding sub-module and a first equivalent weight determination sub-module.
[0187] The first padding submodule can be configured to perform zero-value padding on the first convolution kernel weight of the first convolution kernel and the second convolution kernel weight of the second convolution kernel, so that the matrix size of the padded first convolution kernel weight and the padded second convolution kernel weight is the same. The first equivalent weight determination submodule can be configured to add the first convolution kernel weight and the second convolution kernel weight channel by channel to determine the first equivalent weight.
[0188] In some embodiments, the equivalent convolution kernel weight includes a second equivalent weight, and the plurality of branches includes a third branch, and the third branch includes a third convolution kernel and a fourth convolution kernel in series. The image enhancement module 1404 can include a second padding submodule and a spatial convolution submodule.
[0189] The second padding submodule can be configured to perform zero-value padding on the third convolution kernel weight of the third convolution kernel and the fourth convolution kernel weight of the fourth convolution kernel, so that the matrix size of the padded third convolution kernel weight and the padded fourth convolution kernel weight is the same. The spatial convolution submodule can be configured to perform spatial convolution on the third convolution kernel weight and the fourth convolution kernel weight channel by channel to determine the second equivalent weight.
[0190] In some embodiments, the video enhancement device 1400 can further include a training reference frame determination module, a training image enhancement feature determination module, an enhanced training target frame determination module, a loss function determination module, and a parameter adjustment module.
[0191] The training reference frame determination module is configured to obtain a training target frame and at least one training reference frame of the training target frame from a training video. The training image enhancement feature determination module is configured to perform image enhancement processing on the training target frame and the at least one training reference frame by using the target network model to obtain a training image enhancement feature. The enhanced training target frame determination module is configured to add the training image enhancement feature and the training target frame pixel by pixel to obtain an enhanced training target frame. The loss function determination module is configured to determine a target loss function according to the enhanced training target frame. The parameter adjustment module is configured to adjust the parameters of the target network model by using the target loss function, and the plurality of parallel convolution branches in the image enhancement module are adjusted respectively.
[0192] In some embodiments, the multi-frame alignment module includes at least two alignment submodules, including a first alignment submodule and a second alignment submodule. The multi-frame alignment feature determination module 1403 can include a first alignment feature determination submodule, a second alignment feature obtaining submodule, and a feature fusion submodule.
[0193] The first alignment feature determination submodule can be configured to perform multi-frame alignment processing on the target frame feature and each reference frame feature through a first alignment submodule to obtain the first alignment feature; the second alignment feature determination submodule can be configured to perform multi-frame alignment processing on the target frame feature and each reference frame feature through a second alignment submodule to obtain the second alignment feature; the convolution kernel sizes in the first alignment submodule and the second alignment submodule are different; and the feature fusion submodule is configured to perform feature fusion on the first alignment feature and the second alignment feature to obtain the multi-frame alignment feature.
[0194] In some embodiments, the first alignment submodule includes a first bias prediction substructure and a first convolution substructure; and the first alignment feature determination submodule can include a first offset determination unit, a first actual sampling point determination unit, a first aligned reference frame feature determination unit, and a first alignment feature determination unit.
[0195] The first offset determination unit can be configured to perform bias prediction processing on the target frame feature and at least one reference frame through the first bias prediction substructure to determine a first offset of an actual sampling point of the convolution substructure on each reference frame feature compared with a standard sampling point; the standard sampling point is determined according to a convolution kernel of the convolution substructure; the first actual sampling point determination unit can be configured to determine a first actual sampling point according to the standard sampling point and the first offset; the first aligned reference frame feature determination unit can be configured to perform convolution processing on each reference frame feature according to the first actual sampling point to obtain a first aligned reference frame feature; and the first alignment feature determination unit can be configured to perform feature fusion on the first aligned reference frame feature and the target frame feature to obtain the first alignment feature.
[0196] Since the functions of the apparatus 1400 have been described in detail in the corresponding method embodiments, the present disclosure will not be repeated here.
[0197] The modules and / or submodules and / or units described in the embodiments of the present disclosure can be implemented in the form of software or hardware. The described modules and / or submodules and / or units can also be arranged in a processor. In some cases, the names of these modules and / or submodules and / or units do not constitute a limitation on the modules and / or submodules and / or units themselves.
[0198] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0199] Further, the above-described diagrams are merely schematic illustrations of processes included in the method according to the exemplary embodiments of the present disclosure, and are not intended for limiting purposes. It is readily understood that the processes shown in the above-described diagrams do not indicate or limit the time sequence of the processes. In addition, it is readily understood that the processes can be executed synchronously or asynchronously, for example, in a plurality of modules.
[0200] Figure 15 A structural schematic of an electronic device suitable for implementing the embodiments of the present disclosure is shown. It should be noted that Figure 15 The electronic device 1500 shown is merely an example, and should not bring any limitation to the functions and usage range of the embodiments of the present disclosure.
[0201] As Figure 15 shown, the electronic device 1500 includes a central processing unit (CPU) 1501, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1502 or programs loaded from a storage section 1508 into a random access memory (RAM) 1503. In the RAM 1503, various programs and data required for the operation of the electronic device 1500 are also stored. The CPU 1501, the ROM 1502, and the RAM 1503 are connected to each other through a bus 1504. An input / output (I / O) interface 1505 is also connected to the bus 1504.
[0202] The following components are connected to the I / O interface 1505: an input part 1506 including a keyboard, a mouse, etc.; an output part 1507 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 1508 including a hard disk, etc.; and a communication part 1509 including a network interface card such as a LAN card, a modem, etc. The communication part 1509 performs communication processing via a network such as the Internet. A drive 159 is also connected to the I / O interface 1505 as necessary. A removable media 1510 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 159 as necessary, so that a computer program read out therefrom is installed in the storage part 1508 as necessary.
[0203] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing computer program instructions for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication part 1509, and / or installed from the removable media 1510. When the computer program is executed by the central processing unit (CPU) 1501, the above-described functions defined in the system of the present disclosure are executed.
[0204] It should be noted that the computer-readable storage medium shown in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer-readable signal medium can include a data signal carrying computer-readable computer program instructions in a baseband or as part of a carrier wave. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate or transmit programs for use by or in conjunction with an instruction execution system, device or apparatus. The computer program instructions contained on the computer-readable storage medium can be transmitted in any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0205] As another aspect, the disclosure also provides a computer readable storage medium, which can be included in the device described in the above embodiments, or can exist independently without being assembled into the device. The computer readable storage medium carries one or more programs, which, when executed by the device, enable the device to implement functions including: obtaining a target frame and at least one reference frame of the target frame from a video to be enhanced, the at least one reference frame including at least one of N previous frames of the target frame and M subsequent frames of the target frame, N being a frame number greater than or equal to 1, and M being an integer greater than or equal to 1; performing feature extraction on the target frame and the at least one reference frame by a feature extraction module of a target network model to obtain target frame features corresponding to the target frame and reference frame features corresponding to each reference frame; performing multi-frame alignment processing on the target frame features and the reference frame features by a multi-frame alignment module of the target network model to obtain multi-frame alignment features; performing image enhancement processing on the multi-frame alignment features by an image enhancement module of the target network model to obtain image enhancement features; and adding the image enhancement features and the target frame pixel by pixel to obtain an enhanced target frame; wherein the image enhancement module includes at least one representative feature extraction submodule, the representative feature extraction submodule includes a plurality of parallel convolution branches, and each convolution branch includes at least one convolution kernel; wherein the image enhancement module performs image enhancement processing on the multi-frame alignment features to obtain the image enhancement features, including: merging convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight; and performing convolution processing on the multi-frame alignment features by the equivalent convolution kernel weight to obtain equivalent convolution features, so as to determine the image enhancement features according to the equivalent convolution features.
[0206] According to an aspect of the disclosure, a computer program product or computer program is provided, which includes computer program instructions stored in a computer readable storage medium. The computer program instructions are read from the computer readable storage medium, and a processor executes the computer program instructions to implement the method provided in various optional implementation manners of the above embodiments.
[0207] Through the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solutions of the embodiments of the disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of computer program instructions to make an electronic device (which can be a server or a terminal device, etc.) execute the method according to the embodiments of the disclosure.
[0208] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the disclosure be construed as including any variations, uses, or adaptations of the specific embodiments following, including equivalents thereof, which are within the spirit and scope of the disclosure. The specification and examples given are considered exemplary only, with the true scope and spirit of the disclosure being indicated by the following claims.
[0209] It is understood that the present disclosure is not limited to the details of construction, drawings, or method of implementation shown here, but that the present disclosure intends to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
Claims
1. A method of video enhancement, characterized by, The method comprises: obtaining a target frame and at least one reference frame of the target frame from a video to be enhanced, the at least one reference frame comprising at least one of N frames before the target frame and M frames after the target frame, N being an integer greater than or equal to 1, and M being an integer greater than or equal to 1; performing feature extraction on the target frame and the at least one reference frame by a feature extraction module of a target network model to obtain target frame features corresponding to the target frame and reference frame features corresponding to each reference frame; performing multi-frame alignment processing on the target frame features and the reference frame features by a multi-frame alignment module of the target network model to obtain multi-frame alignment features; performing image enhancement processing on the multi-frame alignment features by an image enhancement module of the target network model to obtain image enhancement features; pixel-by-pixel adding the image enhancement features and the target frame to obtain an enhanced target frame; wherein the image enhancement module comprises at least one representative feature extraction submodule, and each representative feature extraction submodule comprises a plurality of parallel convolution branches, and each convolution branch comprises at least one convolution kernel; wherein the image enhancement module performs image enhancement processing on the multi-frame alignment features to obtain image enhancement features, comprising: merging convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight; performing convolution processing on the multi-frame alignment features by the equivalent convolution kernel weight to obtain equivalent convolution features, so as to determine the image enhancement features according to the equivalent convolution features.
2. The method of claim 1, wherein, The equivalent convolution kernel weight comprises a first equivalent weight, the plurality of branches comprise a first branch and a second branch, the first branch comprises a first convolution kernel, and the second branch comprises a second convolution kernel; wherein merging the convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight comprises: zero-value padding a first convolution kernel weight of the first convolution kernel and a second convolution kernel weight of the second convolution kernel so that the matrix sizes of the padded first convolution kernel weight and the padded second convolution kernel weight are the same; adding the first convolution kernel weight and the second convolution kernel weight channel by channel to determine the first equivalent weight.
3. The method of claim 1, wherein, The equivalent convolution kernel weight comprises a second equivalent weight, the plurality of branches comprise a third branch, the third branch comprises a third convolution kernel and a fourth convolution kernel in series; wherein merging the convolution kernel weights of the convolution kernels on the plurality of branches to obtain an equivalent convolution kernel weight comprises: zero-value padding a third convolution kernel weight of the third convolution kernel and a fourth convolution kernel weight of the fourth convolution kernel so that the matrix sizes of the padded third convolution kernel weight and the padded fourth convolution kernel weight are the same; spatially convolving the third convolution kernel weight and the fourth convolution kernel weight channel by channel to determine the second equivalent weight.
4. The method of claim 1, wherein, The method further comprises: obtaining a training target frame and at least one training reference frame of the training target frame from a training video; performing image enhancement processing on the training target frame and the at least one training reference frame by the target network model to obtain training image enhancement features; add the training image enhancement feature and the training target frame pixel by pixel to obtain an enhanced training target frame; determine a target loss function according to the enhanced training target frame; adjust parameters of the target network model through the target loss function, wherein the parameters of multiple parallel convolution branches in the image enhancement module are adjusted respectively.
5. The method of claim 1, wherein, The multi-frame alignment module includes at least two alignment sub-modules, including a first alignment sub-module and a second alignment sub-module; wherein the target frame feature and each reference frame feature are processed by the multi-frame alignment module of the target network model to obtain a multi-frame alignment feature, including: the target frame feature and each reference frame feature are processed by the first alignment sub-module to obtain a first alignment feature; the target frame feature and each reference frame feature are processed by the second alignment sub-module to obtain a second alignment feature; the convolution kernel sizes in the first alignment sub-module and the second alignment sub-module are different; the first alignment feature and the second alignment feature are fused to obtain the multi-frame alignment feature.
6. The method of claim 5, wherein, The first alignment sub-module includes a first bias prediction sub-structure and a first convolution sub-structure; wherein the target frame feature and each reference frame feature are processed by the first alignment sub-module to obtain a first alignment feature, including: the target frame feature and at least one reference frame are processed by the first bias prediction sub-structure to determine a first offset of actual sampling points of the convolution sub-structure on each reference frame feature compared with standard sampling points; wherein the standard sampling points are determined according to the convolution kernel of the convolution sub-structure; the first convolution sub-structure determines a first actual sampling point according to the standard sampling point and the first offset; the first convolution sub-structure respectively convolves each reference frame feature according to the first actual sampling point to obtain a first aligned reference frame feature; the first aligned reference frame feature and the target frame feature are fused to obtain the first alignment feature.
7. The method of claim 1, wherein, The image enhancement module further includes a pixel rearrangement structure, a pixel shuffling structure and a feature fusion structure, and the feature fusion structure includes the at least one representative feature extraction sub-module; wherein the multi-frame alignment feature is processed by the image enhancement module to obtain an image enhancement feature, including: the multi-frame alignment feature is down-sampled by the pixel rearrangement structure to obtain a down-sampled feature; the down-sampled feature is fused by the feature fusion structure to obtain a fused feature; the fused feature is up-sampled by the pixel shuffling structure to obtain the image enhancement feature.
8. An electronic device, comprising: including: a memory and a processor; the memory is used to store computer program instructions; the processor calls the computer program instructions stored in the memory to implement the video enhancement method of any one of claims 1-7.
9. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions, when executed by a processor, implement the video enhancement method according to any one of claims 1-7.
10. A computer program product comprising computer program instructions stored in a computer readable storage medium, characterized in that, The computer program instructions, when executed by a processor, implement the method according to any one of claims 1-7.
Citation Information
Patent Citations
Real-time video processing method, device, equipment, medium and product
CN116128730A
Video quality enhancement method and device, video quality compression method and device, terminal and medium
CN117714714A
System and method for burst image restoration and enhancement
US20240135496A1
Cited By
Super-resolution method and system for reconstructing high-quality 4K video through low-image description
CN122027758A
A super-resolution method and system for reconstructing high-quality 4K video from low-quality descriptions
CN122027758B