A video coding reference picture reconstruction method based on lookup table enhancement and adaptive fusion and related equipment
Patent Information
- Application Number
- CN202610969590.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-01
AI Technical Summary
1.传统重采样或固定滤波重建的复杂度较低、实现稳定,但其对高频纹理、边缘和压缩失真的恢复能力有限,容易产生过平滑的重建结果
[0021]本申请实施例至少包括以下有益效果:本申请提供一种基于查找表增强与自适应融合的视频编解码参考图像重建方法及相关设备,该方案根据输入的待重建图像,构建基准重建图像;根据所述基准重建图像,确定亮度基准分量和色度基准分量;通过查找表增强模块对所述亮度基准分量进行增强映射处理,将得到的基准重建结果进行残差自适应融合处理得到亮度输出结果;对所述色度基准分量采用旁路策略处理,以色度基准分量作为色度输出结果;根据所述亮度输出结果和所述色度输出结果,构建得到最终输出图像。本申请能够利用查找表替代在线神经网络推理,又能够保留基准重建路径的稳定性,并通过自适应融合控制增强强度,同时对色度分量采用更稳健的旁路策略。
Smart Images

Figure CN122676010A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method and related equipment for video codec reference image reconstruction based on lookup table enhancement and adaptive fusion. Background Technology
[0002] With the development of ultra-high definition, low latency, adaptive bitrate, and multi-terminal video services, reconstructed images in modern video codec systems can serve as both the final output and the input for subsequent prediction, quality enhancement, or reference buffering. To maintain stable performance under different bitrates, resolutions, and device complexities, codecs typically incorporate resampling, scale restoration, loop filtering, neural network enhancement, or other learning-based reconstruction modules.
[0003] When relying solely on traditional interpolation or fixed filtering for scale restoration or reconstruction enhancement, the output image typically exhibits good determinism and low complexity, but its ability to compensate for high-frequency textures, edge details, and compression distortion is limited. Learning-based enhancement methods can further improve some reconstruction quality, such as using NNSR after RPR to restore details in the reconstructed image, but online neural network inference significantly increases decoding latency, memory access pressure, and deployment complexity.
[0004] Lookup table methods offer a low-complexity implementation path that trades storage for online computation. Existing methods have demonstrated that local neural network responses can be cached offline as lookup tables, with some convolution computations replaced by indexing, table lookup, and interpolation in the online stage. However, these methods are primarily geared towards general image super-resolution or image restoration tasks and do not fully consider issues such as baseline reconstruction constraints, encoding / decoding consistency, component differences, and reference image caching in the video encoding / decoding reconstruction chain.
[0005] Existing resampling, learning-based augmentation, and lookup table augmentation schemes have the following main shortcomings in encoding / decoding reconstruction scenarios: 1. Traditional resampling or fixed-filter reconstruction has low complexity and is stable, but its ability to recover high-frequency textures, edges and compression distortion is limited, and it is easy to produce overly smooth reconstruction results.
[0006] 2. Online neural network augmentation can improve some reconstruction quality, but it requires repeated model inference in the encoder or decoder, resulting in high runtime, storage access, and hardware deployment costs at the decoding end.
[0007] 3. Existing LUT methods are not designed for encoding / decoding reconstruction links. If they are directly used to replace the baseline reconstruction results, unstable corrections may be introduced in low-activity areas or areas with strong compression distortion, affecting rate-distortion performance and reference image consistency.
[0008] 4. In video formats such as YUV 4:2:0, the luminance and chrominance components have different sampling densities and distortion characteristics. If the same LUT enhancement strategy is used for both luminance and chrominance, the chrominance index may deteriorate due to insufficient chrominance samples or model mismatch. Summary of the Invention
[0009] The main objective of this application is to propose a video codec reference image reconstruction method and related equipment based on lookup table enhancement and adaptive fusion. This method can use lookup tables to replace online neural network inference, while preserving the stability of the baseline reconstruction path. It also controls the enhancement intensity through adaptive fusion and adopts a more robust bypass strategy for the chroma components.
[0010] To achieve the above objectives, one aspect of this application proposes a video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion, comprising: Based on the input image to be reconstructed, construct a baseline reconstructed image; Reconstruct the image based on the aforementioned reference, and determine the luminance reference component and the chrominance reference component; The luminance reference component is enhanced and mapped by a lookup table enhancement module, and the resulting reference reconstruction result is subjected to residual adaptive fusion processing to obtain the luminance output result; wherein, the fusion weight of the residual adaptive fusion processing is determined based on the local variance, local gradient, texture intensity or edge intensity of the luminance reference component; The chromaticity reference component is processed using a bypass strategy, and the chromaticity reference component is used as the chromaticity output result. The final output image is constructed based on the brightness output result and the chromaticity output result.
[0011] In some embodiments, constructing a baseline reconstructed image based on the input image to be reconstructed includes: Acquire the image to be reconstructed for scale restoration or quality enhancement; The image to be reconstructed is processed by any one or more of the following methods: reference image resampling, equivalent resampling, fixed filtering, or benchmark reconstruction, to obtain a benchmark reconstructed image.
[0012] In some embodiments, determining the luminance reference component and chrominance reference component based on the reconstructed image according to the reference includes: The reconstructed reference image is decomposed into one luminance reference component and two chrominance reference components; Based on the baseline reconstruction path, the luminance reference components are subjected to geometric scale recovery and stable reconstruction, and the luminance components with spatially aligned locations are enhanced at the same resolution according to the lookup table module.
[0013] In some embodiments, the step of enhancing the luminance reference component through a lookup table enhancement module and then performing residual adaptive fusion processing on the obtained reference reconstruction result to obtain the luminance output result includes: The brightness enhancement candidates are obtained by enhancing the mapping based on a preset lookup table. The expression for this process is: ,in, Represents enhanced lookup table mapping; Candidates for brightness enhancement; Represents the luminance reference component; For any pixel location, define the corresponding LUT residual, and use the brightness reference component as the fusion reference to calculate the local mean and local variance in the neighborhood of the selected pixel location. Based on the local mean and local variance, the fusion weights are determined, and then the brightness output result is obtained.
[0014] In some embodiments, determining the fusion weights based on the local mean and local variance to obtain the brightness output result includes: The activity coefficient is obtained by normalizing the local variance between the first and second reference thresholds; the expression for this process is: ,in, This represents the second reference threshold; Represents the first reference threshold; Represents the activity level coefficient; Represents local variance; ( ) represents normalization processing; The fusion weights are obtained based on the activity coefficients; the expression for this process is: ,in, Represents the fusion weight; Represents the basic fusion strength; and Indicates the lower and upper limits of the fusion weights; The final brightness output result is calculated based on the fusion weights.
[0015] In some embodiments, the method further includes: During the offline lookup table construction phase: A multi-lookup table collaborative structure is adopted as a local enhancement model, including standard sampling, dilated sampling, and Y-shaped sampling; Expanding the equivalent receptive field through two-stage cascading; Configure a preset scale so that each final entry of the LUT outputs a brightness enhancement response aligned with the spatial location of the input pixels; After training, local input combinations are enumerated or uniformly sampled, and the model output is exported offline as a lookup table. Fine-tuning can be performed in the lookup table format to adapt the table entries to the actual query and interpolation process.
[0016] In some embodiments, the method further includes: During the online reconstruction phase: The encoder or decoder first obtains the reference reconstructed image according to the preset method of the current encoding and decoding system; For the luminance reference component, a lookup table index is constructed based on the current pixel and its neighboring samples, and luminance enhancement candidates are obtained by looking up the table and interpolation; Then, the residual is calculated based on the luminance reference component, and the fusion weight is obtained according to the local variance, and finally the luminance output is obtained.
[0017] Another aspect of this application provides a video codec reference image reconstruction apparatus based on lookup table enhancement and adaptive fusion, comprising: The first module is used to construct a baseline reconstructed image based on the input image to be reconstructed; The second module is used to reconstruct the image based on the reference and determine the luminance reference component and the chrominance reference component; The third module is used to perform enhancement mapping processing on the brightness reference component through the lookup table enhancement module, and perform residual adaptive fusion processing on the obtained reference reconstruction result to obtain the brightness output result; wherein, the fusion weight of the residual adaptive fusion processing is determined according to the local variance, local gradient, texture intensity or edge intensity of the brightness reference component; The fourth module is used to process the chromaticity reference component using a bypass strategy, and use the chromaticity reference component as the chromaticity output result. The fifth module is used to construct the final output image based on the brightness output result and the chromaticity output result.
[0018] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0019] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0020] This application also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned method.
[0021] The embodiments of this application include at least the following beneficial effects: This application provides a video codec reference image reconstruction method and related equipment based on lookup table enhancement and adaptive fusion. This scheme constructs a reference reconstructed image based on the input image to be reconstructed; determines the luminance reference component and chrominance reference component based on the reference reconstructed image; performs enhancement mapping processing on the luminance reference component through a lookup table enhancement module; performs residual adaptive fusion processing on the obtained reference reconstruction result to obtain the luminance output result; employs a bypass strategy to process the chrominance reference component, using the chrominance reference component as the chrominance output result; and constructs the final output image based on the luminance output result and the chrominance output result. This application can utilize lookup tables to replace online neural network inference, while preserving the stability of the reference reconstruction path, controlling the enhancement intensity through adaptive fusion, and employing a more robust bypass strategy for the chrominance component. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application; Figure 2 This is a flowchart of the overall steps provided in the embodiments of this application; Figure 3 This is a flowchart of the overall process for reference image reconstruction based on lookup table enhancement and adaptive fusion provided in the embodiments of this application; Figure 4 This is a flowchart illustrating the offline lookup table construction and online deployment process provided in this application embodiment; Figure 5 This is the consistent reconstruction process in the encoder and decoder provided in the embodiments of this application; Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0024] It is understood that the terms "first," "second," "third," "fourth," etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0027] Before providing a detailed description of the embodiments of this application, some related technologies involved in the embodiments of this application will be described first, as follows: Existing resampling, learning-based augmentation, and lookup table augmentation schemes have the following main shortcomings in encoding / decoding reconstruction scenarios.
[0028] 1. Traditional resampling or fixed-filter reconstruction has low complexity and is stable, but its ability to recover high-frequency textures, edges and compression distortion is limited, and it is easy to produce overly smooth reconstruction results.
[0029] 2. Online neural network augmentation can improve some reconstruction quality, but it requires repeated model inference in the encoder or decoder, resulting in high runtime, storage access, and hardware deployment costs at the decoding end.
[0030] 3. Existing LUT methods are not designed for encoding / decoding reconstruction links. If they are directly used to replace the baseline reconstruction results, unstable corrections may be introduced in low-activity areas or areas with strong compression distortion, affecting rate-distortion performance and reference image consistency.
[0031] 4. In video formats such as YUV 4:2:0, the luminance and chrominance components have different sampling densities and distortion characteristics. If the same LUT enhancement strategy is used for both luminance and chrominance, the chrominance index may deteriorate due to insufficient chrominance samples or model mismatch.
[0032] Therefore, it is necessary to propose a general low-complexity enhancement method for the encoding and decoding reconstruction link, which can both use lookup tables to replace online neural network inference, preserve the stability of the baseline reconstruction path, control the enhancement intensity through adaptive fusion, and adopt a more robust bypass strategy for the chroma component.
[0033] In view of this, this application aims to address the problems of high complexity of learning-based enhancement, instability of direct lookup table replacement, and chroma component adaptability in the encoding / decoding reconstruction chain. It proposes a video encoding / decoding reference image reconstruction method based on lookup table enhancement and adaptive fusion. This method uses the outputs of the RPR or equivalent resampling, fixed filtering, and baseline reconstruction modules as the baseline reconstruction path, uses the lookup table model as a candidate for same-resolution enhancement of the luminance component, and adaptively injects the LUT residuals based on local activity. Simultaneously, it allows the chroma component to directly utilize the RPR or equivalent resampling reconstruction results.
[0034] The video codec reference image reconstruction method and related equipment based on lookup table enhancement and adaptive fusion provided in this application belong to the field of image processing technology. The video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion provided in this application can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited thereto; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the server can also be a node server in a blockchain network; the software can be an application implementing the video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion, but is not limited to the above forms.
[0035] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0036] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0037] like Figure 1 The diagram shown is a schematic representation of an implementation environment provided in an embodiment of this application. (Refer to...) Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected via a network, either wirelessly or via a wired connection, to complete data transmission and exchange.
[0038] Server 101 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0039] Additionally, server 101 can also be a node server in a blockchain network. Blockchain is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms.
[0040] Terminal 102 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc. It can also be a vehicle-mounted terminal of the various device types described above, but is not limited to these. Terminal 102 and server 101 can be directly or indirectly connected via wired or wireless communication, and this embodiment does not impose any limitations.
[0041] Exemplary based on Figure 1 The implementation environment shown in this application embodiment provides a video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion. The following description uses the application of this video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion in server 101 as an example. It can be understood that this method can also be applied in terminal 102.
[0042] Reference Figure 2 , Figure 2 This is a flowchart illustrating a video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion applied to a server, provided as an embodiment of this application. The execution subject of this method can be any of the aforementioned computer devices (including servers or terminals). (Refer to...) Figure 2 The method may include the following steps: S210. Construct a baseline reconstructed image based on the input image to be reconstructed; Image reconstruction is a technique that obtains the three-dimensional shape of an object through digital processing of external measurement data. It is mainly used in fields such as medical imaging, industrial inspection and material analysis. The image to be reconstructed in this embodiment can come from fields such as medical imaging, industrial inspection and material analysis.
[0043] S220. Reconstruct the image based on the reference and determine the luminance reference component and chrominance reference component; S230. The brightness reference component is enhanced and mapped by the lookup table enhancement module, and the obtained reference reconstruction result is subjected to residual adaptive fusion processing to obtain the brightness output result. S240. The chromaticity reference component is processed using a bypass strategy, and the chromaticity reference component is used as the chromaticity output result. LUT (Lookup Table Image Enhancement) is a technique that uses pre-computed lookup tables to replace complex real-time calculations to achieve image color correction, contrast enhancement, and stylized color grading. Its core advantage lies in "trading space for time," transforming mathematical calculations into memory indexes, thereby achieving millisecond-level real-time processing. It is widely used in mobile photography, video post-production, and embedded vision systems.
[0044] S250. Based on the brightness output result and the chromaticity output result, the final output image is constructed.
[0045] It should be noted that LUTs / lookup tables, as a technique for pre-storing input-output mapping relationships and then retrieving them online, already exist in scenarios such as image super-resolution, denoising, filtering, image restoration, and video enhancement. The method in this application mainly embodies the combined application of LUTs and adaptive fusion strategies in the video encoding / decoding and reconstruction chain, and has the following specific characteristics: (1) The lookup table enhancement module is only used to generate brightness enhancement candidates. The final brightness output does not directly replace the reference reconstruction result, but is obtained by adaptive fusion of the reference reconstruction result and the LUT residual. (2) The fusion weight is determined by the local variance, local gradient, texture intensity, edge intensity or other local activity index of the luminance reference component; (3) The encoding and decoding ends obtain the fusion weights based on the same reconstructed samples and the same weight calculation rules, without the need to transmit additional fusion parameters in the bitstream; (4) The method is used as an enhancement module after upsampling, downsampling, scaling transformation, reference image resampling or equivalent benchmark reconstruction processing, and is connected to the reconstructed image enhancement process of video codec reference software, encoder or decoder; (5) The final output image is written into the reference image buffer, used as a reference image for subsequent inter-frame prediction, or used as a decoding output image; (6) The lookup table can be obtained offline by neural networks, filters, statistical mapping, training samples or other local enhancement models, so as to avoid limiting the protection scope to a specific LUT or MuLUT structure.
[0046] In some embodiments, constructing a baseline reconstructed image based on the input image to be reconstructed includes: Acquire the image to be reconstructed for scale restoration or quality enhancement; The image to be reconstructed is processed by any one or more of the following methods: reference image resampling, equivalent resampling, fixed filtering, or benchmark reconstruction, to obtain a benchmark reconstructed image.
[0047] In some embodiments, determining the luminance reference component and chrominance reference component based on the reconstructed image according to the reference includes: The reconstructed reference image is decomposed into one luminance reference component and two chrominance reference components; Based on the baseline reconstruction path, the luminance reference components are subjected to geometric scale recovery and stable reconstruction, and the luminance components with spatially aligned locations are enhanced at the same resolution according to the lookup table module.
[0048] In some embodiments, the step of enhancing the luminance reference component through a lookup table enhancement module and then performing residual adaptive fusion processing on the obtained reference reconstruction result to obtain the luminance output result includes: The brightness enhancement candidates are obtained by enhancing the mapping based on a preset lookup table. The expression for this process is: ,in, Represents enhanced lookup table mapping; Candidates for brightness enhancement; Represents the luminance reference component; For any pixel location, define the corresponding LUT residual, and use the brightness reference component as the fusion reference to calculate the local mean and local variance in the neighborhood of the selected pixel location. Based on the local mean and local variance, the fusion weights are determined, and then the brightness output result is obtained.
[0049] In some embodiments, determining the fusion weights based on the local mean and local variance to obtain the brightness output result includes: The activity coefficient is obtained by normalizing the local variance between the first and second reference thresholds; the expression for this process is: ,in, This represents the second reference threshold; Represents the first reference threshold; Represents the activity level coefficient; Represents local variance; ( ) represents normalization processing; The fusion weights are obtained based on the activity coefficients; the expression for this process is: ,in, Represents the fusion weight; Represents the basic fusion strength; and Indicates the lower and upper limits of the fusion weights; The final brightness output result is calculated based on the fusion weights.
[0050] In some embodiments, the method further includes: During the offline lookup table construction phase: A multi-lookup table collaborative structure is adopted as a local enhancement model, including standard sampling, dilated sampling, and Y-shaped sampling; Expanding the equivalent receptive field through two-stage cascading; Configure a preset scale so that each final entry of the LUT outputs a brightness enhancement response aligned with the spatial location of the input pixels; LUT, or Look-Up Table, is a method that stores pre-calculated, trained, or statistically derived input-output mappings as table entries, and quickly retrieves the corresponding output based on the input index during runtime. LUTs transform complex computations into table lookups, thereby reducing online computation and inference latency. In the field of video encoding and decoding, current standalone LUT research techniques can be used in tasks such as image super-resolution, denoising, filtering, image restoration, and video enhancement to approximate or replace some complex models, filters, or local mapping processes.
[0051] After training, local input combinations are enumerated or uniformly sampled, and the model output is exported offline as a lookup table. Fine-tuning can be performed in the lookup table format to adapt the table entries to the actual query and interpolation process.
[0052] In some embodiments, the method further includes: During the online reconstruction phase: The encoder or decoder first obtains the reference reconstructed image according to the preset method of the current encoding and decoding system; For the luminance reference component, a lookup table index is constructed based on the current pixel and its neighboring samples, and luminance enhancement candidates are obtained by looking up the table and interpolation; Then, the residual is calculated based on the luminance reference component, and the fusion weight is obtained according to the local variance, and finally the luminance output is obtained.
[0053] The following detailed description of the implementation process of this application, with reference to the accompanying drawings, uses a specific application scenario as an example: The overall process of this application is as follows: Figure 3 As shown. Let the input reconstructed image to be scaled or enhanced be... , Representing the RPR or equivalent resampling, fixed filtering, and baseline reconstruction process, the resulting baseline reconstructed image is: (1) in, It can be decomposed into a luminance component and two chrominance components, namely In this application, the baseline reconstruction path is responsible for geometric scale recovery and stable reconstruction, and the lookup table module performs same-resolution enhancement only on the spatially aligned luminance components.
[0054] It should be noted that RPR in this paper refers to Reference Picture Resampling in the field of video encoding and decoding. RPR is a scale-adaptive tool in video coding standards / reference software, used to upsample, downsample, or scale-restore reference images during encoding or decoding reconstruction, enabling reconstructed images of different resolutions to be used for subsequent prediction, reference buffering, or quality enhancement. The RPR output in this application should be understood as a stable reference image obtained through reference image resampling or equivalent reference reconstruction. Subsequent LUT enhancement and adaptive fusion are based on this reference reconstructed image, performing detail correction at the same resolution, rather than being used for 3D scene reconstruction or depth estimation.
[0055] Let the lookup table enhancement mapping be The candidates for brightness enhancement are: (2) To avoid the instability caused by directly replacing the RPR result with the LUT output, this application performs fusion in the residual domain. For pixel position p, the LUT residual is defined as: (3) by For fusion reference, in the neighborhood of position p Calculate the local mean and local variance: (4) (5) Based on local variance at two reference thresholds and The normalization results between them yield the activity coefficient: (6) Further obtain the fusion weight : (7) in, Indicates the basic fusion strength. and These represent the lower and upper limits of the fusion weights. The final brightness output is: (8) in, For pixel bit depth, this application adopts a bypass strategy for the chroma component, that is, it does not perform LUT enhancement on the chroma component, but directly uses the RPR or equivalent reference reconstruction result as the chroma output: (9) The final output image is This output can be used as the decoded output image or written into the reference image buffer in the encoding / decoding reference software for subsequent inter-frame prediction or quality assessment.
[0056] In addition, in the specific implementation process, such as Figure 4 and Figure 5 As shown, the scheme of this application can also be divided into an offline lookup table construction process and an online reconstruction process.
[0057] Specifically, regarding the offline lookup table construction process: such as Figure 4 As shown, in the offline phase, training samples are first generated based on the target encoding / decoding system or equivalent reconstruction process. The reference reconstructed image, after encoding, decoding, and RPR or equivalent resampling, can be used as input, and the corresponding original image, target reconstructed image, or quality-constrained reference image can be used as the supervision target, forming input-target sample pairs with the same resolution. During training, a lookup table allows the model to learn the local enhancement mapping from the reference reconstructed brightness image to the target brightness image.
[0058] In one implementation, a multi-lookup table collaborative structure can be used as a local enhancement model, including complementary local indexing methods such as standard sampling, dilated sampling, and Y-shaped sampling, and the equivalent receptive field is expanded through two-level cascading. Unlike conventional image super-resolution, the scale in this implementation is set to x1, meaning that each final entry of the LUT outputs a brightness enhancement response aligned with the spatial position of the input pixel, without pixel rearrangement-style magnification.
[0059] After training, local input combinations are enumerated or uniformly sampled, and the model output is exported offline as a lookup table. Fine-tuning can then be performed in lookup table form to adapt the table entries to the actual query and interpolation process. Offline training, table export, and fine-tuning do not participate in the online encoding / decoding process, therefore they do not increase the computational complexity of online convolution at the decoding end.
[0060] For online reconstruction processes: such as Figure 5 As shown, in the online phase, the encoder or decoder first obtains the reference reconstructed image according to the RPR, fixed filtering, or equivalent resampling process in the current encoding / decoding system. For the luminance component A lookup table index is constructed based on the current pixel and its neighboring samples. LUT enhancement candidates are obtained through table lookup and interpolation. After that, Calculate residuals based on the benchmark The fusion weights are obtained based on the local variance. The brightness output is obtained according to formula (8). .
[0061] This application does not simply replace the baseline reconstruction output with the LUT output, but rather retains the baseline reconstruction result as the geometric restoration and stable reconstruction path, injecting only the effective detail corrections in the LUT output as residuals. For flat areas, the smaller local variance results in lower fusion weights, reducing over-enhancement or noise amplification; for edge and textured areas, the higher local variance allows the LUT residuals to be preserved more fully, thereby improving the restoration of brightness details.
[0062] For chromaticity components such as Cb and Cr, the online stage directly uses the RPR or equivalent benchmark reconstruction results as output. This bypass strategy can avoid inappropriately applying the local corrections obtained from luminance training to the chromaticity components, and is especially suitable for video formats with fewer chromaticity sampling points, such as YUV 4:2:0.
[0063] Additionally, it should be noted that this application can be integrated into video codec reference software as a resampling enhancement module, reference image generation module, loop post-processing module, or output reconstruction enhancement module. When the existing software calls a neural network super-resolution or other learning-based enhancement model after baseline reconstruction, the online model inference can be replaced with a lookup table enhancement module and an adaptive fusion module. The encoder and decoder sides use the same resampling, lookup table, interpolation, and fusion rules, thus maintaining consistency when the reconstructed image is used as a subsequent reference image.
[0064] In NNVC or VVC-type reference software, specific implementations may include: obtaining a reference reconstructed image through RPR upsampling; calling the LUT brightness enhancement module to obtain LUT candidate images; performing adaptive luminance residual fusion before writing the reconstructed output or reference buffer; and directly copying the RPR resampling results for the Cb and Cr components. This process can also be ported to VTM, ECM, or other video codec software with scaling, reference image resampling, resolution adaptive coding, loop post-processing, or reconstruction enhancement modules.
[0065] Compared to traditional online neural network super-resolution, this application replaces convolutional inference with lookup table access, interpolation, and simple local statistical calculations, significantly reducing online computation. This application employs an RPR baseline plus residual adaptive fusion approach, ensuring that LUT enhancement only works in appropriate regions, improving the stability of encoding / decoding reconstruction. The chroma bypass strategy reduces the performance degradation of chroma components caused by insufficient samples and model mismatch.
[0066] The experimental results in Table 1 below are from the implementation and testing on the NNVC-15.0 reference software. To illustrate the complexity-performance trade-off between this application and the original online neural network super-resolution scheme, Table 1 lists the results of both NNSR and the LUT-AR embodiment of this application relative to the same NNVC-15.0 anchor point. Evaluation metrics include BD-rate, encoding time EncT, and decoding time DecT. A negative BD-rate indicates a reduction in bitrate while maintaining the same objective quality; EncT and DecT are percentages of runtime relative to the anchor point. The experimental results are used to illustrate the effects of the embodiments of this application and do not constitute a limitation on the scope of protection.
[0067] Table 1. Main experimental results of NNSR and LUT-AR on NNVC-15.0
[0068] As shown in Table 1, NNSR exhibits greater luminance BD-rate gains in most configurations, but its DecT is nearly identical to the original anchor runtime. The LUT-AR implementation in this application retains some luminance gains while significantly reducing decoding-side runtime. In the lower bitrate ranges of QP 27, 32, 37, 42, and 47, LUT-AR achieves luminance BD-rates of -2.66% and -3.09% in AI and RA configurations, respectively, with DecTs reduced to 29.52% and 60.97%, respectively. These results demonstrate that this application does not aim to completely replace NNSR with the highest reconstruction gain, but rather focuses on low-complexity deployments, significantly reducing online inference overhead while maintaining effective luminance improvement.
[0069] Table 2 Comparison of Model Complexity
[0070] As shown in Table 2, this embodiment reduces the online computation from 4.574 kMAC / pixel to 0.120 kMAC / pixel by moderately increasing the lookup table storage, demonstrating the technical effect of reducing online complexity by trading storage.
[0071] Compared to existing technologies, this application does not protect the basic structure of the lookup table itself as a separate object. Existing methods have disclosed the idea of caching local neural network responses through lookup tables and using them for general image super-resolution or image restoration. The difference in this application is that it embeds this type of low-complexity enhancement mechanism into the video encoding and decoding reference image reconstruction process, uses RPR or equivalent resampling results as a stable benchmark, performs LUT enhancement at the same resolution only on the spatially aligned luminance components, and obtains the final luminance output through residual fusion driven by local activity, while simultaneously bypassing the chrominance components along the RPR path.
[0072] Compared to schemes that directly connect to online neural network super-resolution after RPR, this application avoids repeatedly performing convolutional inference at the decoding end; compared to ordinary LUT super-resolution schemes, this application does not undertake the task of geometric amplification, but serves to improve and enhance details at the same resolution after encoding and decoding resampling.
[0073] Another aspect of this application provides a video codec reference image reconstruction apparatus based on lookup table enhancement and adaptive fusion, comprising: The first module is used to construct a baseline reconstructed image based on the input image to be reconstructed; The second module is used to reconstruct the image based on the reference and determine the luminance reference component and the chrominance reference component; The third module is used to perform enhanced mapping processing on the brightness reference component through the lookup table enhancement module, and perform residual adaptive fusion processing on the obtained reference reconstruction result to obtain the brightness output result. The fourth module is used to process the chromaticity reference component using a bypass strategy, and use the chromaticity reference component as the chromaticity output result. The fifth module is used to construct the final output image based on the brightness output result and the chromaticity output result.
[0074] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0075] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0076] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0077] Please see Figure 6 , Figure 6 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 to implement the video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion according to the embodiments of this application. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0078] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion.
[0079] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0080] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0081] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0082] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0083] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0084] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0085] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0086] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0087] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0088] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks. The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for reconstructing video codec reference images based on lookup table enhancement and adaptive fusion, characterized in that, include: Based on the input image to be reconstructed, construct a baseline reconstructed image; Reconstruct the image based on the aforementioned reference, and determine the luminance reference component and the chrominance reference component; The luminance reference component is enhanced and mapped by a lookup table enhancement module, and the resulting reference reconstruction result is subjected to residual adaptive fusion processing to obtain the luminance output result; wherein, the fusion weight of the residual adaptive fusion processing is determined based on the local variance, local gradient, texture intensity or edge intensity of the luminance reference component; The chromaticity reference component is processed using a bypass strategy, and the chromaticity reference component is used as the chromaticity output result. The final output image is constructed based on the brightness output result and the chromaticity output result.
2. The video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion according to claim 1, characterized in that, The step of constructing a baseline reconstructed image based on the input image to be reconstructed includes: Acquire the image to be reconstructed for scale restoration or quality enhancement; The image to be reconstructed is processed by any one or more of the following methods: reference image resampling, equivalent resampling, fixed filtering, or benchmark reconstruction, to obtain a benchmark reconstructed image.
3. The video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion according to claim 1, characterized in that, The step of reconstructing the image based on the reference and determining the luminance reference component and chrominance reference component includes: The reconstructed reference image is decomposed into one luminance reference component and two chrominance reference components; Based on the baseline reconstruction path, the luminance reference components are subjected to geometric scale recovery and stable reconstruction, and the luminance components with spatially aligned locations are enhanced at the same resolution according to the lookup table module.
4. The video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion according to claim 1, characterized in that, The step of enhancing the luminance reference component through a lookup table enhancement module, and then performing residual adaptive fusion processing on the obtained reference reconstruction result to obtain the luminance output result includes: The brightness enhancement candidates are obtained by enhancing the mapping based on a preset lookup table. The expression for this process is: ,in, Represents enhanced lookup table mapping; Candidates for brightness enhancement; Represents the luminance reference component; For any pixel location, define the corresponding LUT residual, and use the brightness reference component as the fusion reference to calculate the local mean and local variance in the neighborhood of the selected pixel location. Based on the local mean and local variance, the fusion weights are determined, and then the brightness output result is obtained.
5. The video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion according to claim 4, characterized in that, The step of determining the fusion weights based on the local mean and local variance, and then obtaining the brightness output result, includes: The activity coefficient is obtained by normalizing the local variance between the first and second reference thresholds; the expression for this process is: ,in, This represents the second reference threshold; Represents the first reference threshold; Represents the activity level coefficient; Represents local variance; ( ) represents normalization processing; The fusion weights are obtained based on the activity coefficients; the expression for this process is: ,in, Represents the fusion weight; Represents the basic fusion strength; and Indicates the lower and upper limits of the fusion weights; The final brightness output result is calculated based on the fusion weights.
6. The video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion according to claim 1, characterized in that, The method further includes: During the offline lookup table construction phase: A multi-lookup table collaborative structure is adopted as a local enhancement model, including standard sampling, dilated sampling, and Y-shaped sampling; Expanding the equivalent receptive field through two-stage cascading; Configure a preset scale so that each final entry of the LUT outputs a brightness enhancement response aligned with the spatial location of the input pixels; After training, local input combinations are enumerated or uniformly sampled, and the model output is exported offline as a lookup table. Fine-tuning can be performed in the lookup table format to adapt the table entries to the actual query and interpolation process.
7. The video codec reference image reconstruction method based on lookup table enhancement and adaptive fusion according to claim 1, characterized in that, The method further includes: During the online reconstruction phase: The encoder or decoder first obtains the reference reconstructed image according to the preset method of the current encoding and decoding system; For the luminance reference component, a lookup table index is constructed based on the current pixel and its neighboring samples, and luminance enhancement candidates are obtained by looking up the table and interpolation; Then, the residual is calculated based on the luminance reference component, and the fusion weight is obtained according to the local variance, and finally the luminance output is obtained.
8. A video codec reference image reconstruction apparatus based on lookup table enhancement and adaptive fusion, characterized in that, include: The first module is used to construct a baseline reconstructed image based on the input image to be reconstructed; The second module is used to reconstruct the image based on the reference and determine the luminance reference component and the chrominance reference component; The third module is used to perform enhancement mapping processing on the brightness reference component through the lookup table enhancement module, and perform residual adaptive fusion processing on the obtained reference reconstruction result to obtain the brightness output result; wherein, the fusion weight of the residual adaptive fusion processing is determined according to the local variance, local gradient, texture intensity or edge intensity of the brightness reference component; The fourth module is used to process the chromaticity reference component using a bypass strategy, and use the chromaticity reference component as the chromaticity output result. The fifth module is used to construct the final output image based on the brightness output result and the chromaticity output result.
9. An electronic device, characterized in that, Including the processor and memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.