A method, device and medium for nondestructive enhancement of an ultra-low-illumination original image of cultural relics
By combining residual variational autoencoders and diffusion models with a region cross-attention mechanism, the problems of information loss and pigment shift in ultra-low illumination image enhancement of cultural relics were solved, achieving lossless enhancement and accurate color restoration, which meets the archiving standards for cultural relics.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 浪潮智慧科技有限公司
- Filing Date
- 2026-04-10
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies for enhancing ultra-low illumination images of cultural relics suffer from problems such as overly smoothed images, loss of local archaeological structures, and pigment color shifts, failing to meet the legal standards for archiving cultural relics.
By employing a residual variational autoencoder and a diffusion model combined with a region-based cross-attention mechanism, archive-grade sRGB images are generated through image preprocessing, latent feature encoding, denoising, and decoding, preserving archival information and accurately restoring the colors of mineral pigments.
It achieves lossless enhancement of original images of cultural relics under ultra-low illumination, preserves archaeological information and accurately restores the colors of mineral pigments, avoids damage to cultural relics caused by long exposure to strong light, and meets the requirements of archive-level image quality.
Smart Images

Figure CN122453673A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer vision and digital preservation of cultural relics, and in particular to a non-destructive enhancement method, device and medium for original images of cultural relics under ultra-low illumination. Background Technology
[0002] In the current field of digital preservation of cultural heritage, acquiring high-precision, high-fidelity archive-grade sRGB images of cultural relics is the cornerstone of their research, restoration, display, and transmission. During the digital acquisition of immovable painted cultural relics such as grottoes, tomb murals, and ancient architectural paintings, strong lighting is strictly prohibited on-site to protect these fragile artifacts. This results in the original images being extremely low-light images, unusable directly and requiring image enhancement.
[0003] Existing regression-based low-light enhancement models suffer from the "mean regression" problem, resulting in overly smooth images of cultural relics that fail to preserve subtle archaeological information such as brushstrokes, pigment cracks, and inscriptions in low-light areas. When traditional latent diffusion models (LDM) are directly used to enhance original images of cultural relics in ultra-low light, the use of global cross-attention can lead to the loss of local archaeological structures and the creation of content illusions (such as altering mural texts or smoothing out pigment cracks). Moreover, existing models rely solely on latent space loss during training, which can easily lead to mineral pigment color shifts, resulting in inconsistencies between the enhanced image color and the original pigment color of the cultural relic, failing to meet the legal standards for archiving cultural relics. Summary of the Invention
[0004] This application provides a method, device, and medium for non-destructive enhancement of original images of cultural relics under ultra-low illumination, in order to solve the following technical problems: how to achieve non-destructive enhancement of original images of cultural relics under ultra-low illumination, preservation of archaeological information, and accurate restoration of mineral pigment colors.
[0005] In a first aspect, embodiments of this application provide a lossless enhancement method for original images of cultural relics under ultra-low illumination. The method includes: acquiring a Bayer format image of the cultural relic and performing image preprocessing on the Bayer format image to obtain a linear LRGB image; encoding the LRGB image using a residual variational autoencoder to obtain latent features, wherein a residual connection is provided between corresponding blocks of the encoder and decoder of the residual variational autoencoder, and preset weights are assigned to archaeological feature regions in the LRGB image to preserve core archaeological information; inputting the latent features into a diffusion model to obtain denoised latent features, wherein a region-based cross-attention mechanism is used in the denoising network of the diffusion model to divide the latent features into multiple regions, and independent condition-guided denoising is performed on each region using the latent features of the processed LRGB image as conditions to obtain denoised latent features; and decoding the denoised latent features using the decoder of the residual variational autoencoder to obtain an archive-grade standard sRGB image.
[0006] Secondly, embodiments of this application also provide a non-destructive enhancement device for original images of cultural relics under ultra-low illumination. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a non-destructive enhancement method for original images of cultural relics under ultra-low illumination as described in the first aspect above.
[0007] Thirdly, embodiments of this application also provide a computer storage medium storing computer-executable instructions, which, when executed, implement a non-destructive enhancement method for original images of cultural relics under ultra-low illumination as described in the first aspect above.
[0008] The non-destructive enhancement method, device, and medium for original images of cultural relics under ultra-low illumination provided in this application have the following beneficial effects: In this embodiment, a Bayer format image of the cultural relic is acquired, and the Bayer format image is preprocessed to obtain a linear LRGB image. A residual variational autoencoder is used to encode the LRGB image to obtain latent features. These latent features are then input into a diffusion model to obtain denoised latent features. Finally, the decoder of the aforementioned residual variational autoencoder is used to decode the denoised latent features to obtain an archive-grade sRGB image. This allows for lossless enhancement of the original image of the cultural relic under ultra-low illumination, completely resolving the industry-wide conflict between cultural relic preservation and acquisition accuracy. It eliminates the need for long exposures under strong light, preventing irreversible damage to the cultural relic during the acquisition process. Furthermore, through the residual connections in the residual variational autoencoder and the region-based cross-attention mechanism in the diffusion model, archaeological information fidelity and accurate color reproduction of mineral pigments can be achieved. Attached Figure Description
[0009] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a non-destructive enhancement method for original images of cultural relics under ultra-low illumination, provided in this application embodiment; Figure 2 This is a schematic diagram of the internal structure of a non-destructive enhancement device for original images of cultural relics under ultra-low illumination, provided as an embodiment of this application. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0011] This application provides a non-destructive enhancement method for original images of cultural relics under ultra-low illumination. The technical solution proposed in this application will be described in detail below with reference to the accompanying drawings.
[0012] Figure 1 This is a flowchart illustrating a non-destructive enhancement method for original low-light images of cultural relics, provided as an embodiment of this application. Figure 1 As shown in the embodiment of this application, a non-destructive enhancement method for original images of cultural relics under ultra-low illumination includes the following steps: Step 101: Obtain the Bayer format image of the cultural relic and perform image preprocessing on the Bayer format image to obtain a linear LRGB image.
[0013] In practical applications, ultra-low-light original images of cultural relics refer to Bayer format raw data images captured by professional cameras during the digital acquisition of immovable painted cultural relics such as grottoes, tomb murals, and ancient architectural paintings, under ultra-low-light environments (acquisition illumination ≤ 20 lux) that comply with cultural relic protection standards. These images have extremely low photon counts, contain a large amount of Gaussian and Poisson noise, and have not been fully processed by an image signal processor (ISP). They are the only lossless carrier of the original information of cultural relics. Acquiring Bayer format images allows us to obtain all the original information of cultural relic images, thus laying a solid data foundation for lossless enhancement of cultural relic images.
[0014] In this embodiment, the Bayer format image of the cultural relic can be preprocessed to obtain a linear (LRGB) image. An LRGB image is generally a linear spatial image obtained by preprocessing the original image. The LRGB image is an intermediate image format; Bayer format images cannot be directly enhanced. Converting it to an LRGB image provides a standardized, linearized image that retains the original, lossless information of the cultural relic for subsequent processing.
[0015] Step 102: Encode the LRGB image using a residual variational autoencoder to obtain latent features.
[0016] The residual variational autoencoder has a residual connection between the encoder and decoder blocks, and assigns a preset weight to the archaeological feature regions in the LRGB image to preserve core archaeological information. In this embodiment, the aforementioned residual variational autoencoder (VAE) can be used to encode LRGB images to obtain latent features. Traditional VAEs include an encoder and a decoder, used to map images to a latent space and reconstruct images from the latent space. The residual variational autoencoder used in this application, based on the traditional VAE, introduces residual connections with archaeological feature weights between encoder-decoder blocks. This enhances the fidelity of the original content of cultural relics and prevents the loss or alteration of original archaeological information. Thus, it improves the fidelity of cultural relic content and prevents the loss or alteration of original information.
[0017] Step 103: Input the latent features into the diffusion model to obtain the denoised latent features.
[0018] In the denoising network of the diffusion model, a region-based cross-attention mechanism is used to divide the latent features into multiple regions. The latent features of the processed LRGB image are used as conditions to perform independent condition-guided denoising on each region to obtain the denoised latent features.
[0019] The diffusion model in this embodiment is not a traditional diffusion model, but an adaptively modified diffusion model based on the needs of the task of lossless enhancement of the original image of cultural relics under ultra-low illumination. In the diffusion model, the latent space refers to a low-dimensional, compressed representation space in which the diffusion model can perform forward noise addition and backward noise reduction processes, rather than operating directly in the original high-dimensional pixel space.
[0020] In this embodiment, a specialized region-based cross-attention mechanism can be employed to achieve pixel-level conditional control, eliminating structural tampering. The model cannot freely determine the macroscopic structure of the image; each step of denoising and each completion step is forced to be spatially aligned with the low-quality original input image. This ensures that the enhanced result is consistent with the original artifact in terms of composition, contour, and line position. Moreover, under extremely low illumination, the attention mechanism helps the model identify which fluctuations are real edge or texture signals and which are random noise (irrelevant to the conditional image) among noisy latent features. This ensures that when generating a clear image from noise, it is based on real images, not fictional ones, thus preserving archaeological information and ensuring the generation of clear, high-fidelity images.
[0021] Step 104: Use the decoder of the residual variational autoencoder to decode the denoised latent features to obtain an archive-grade standard sRGB image.
[0022] In practical applications, the denoised latent features are an abstract mathematical language that humans and image processing software cannot understand. The decoder, however, can convert them into a universally recognized and processed image format. Furthermore, latent features are typically low-resolution; the decoder can decode these features through a series of upsampling and deconvolution operations to reconstruct a high-resolution, clear image. Residual connections can automatically inject original details during decoding, ensuring that the final output image includes core archaeological features without any tampering.
[0023] In practical applications, archive-grade sRGB images refer to image files encoded and stored using the standard RGB (Standard Red Green Blue, sRGB) color space, meeting the requirements for long-term preservation and high-quality reproduction. The core of this approach lies in ensuring stable image presentation across different devices and time dimensions through standardized color management. In practice, museums, libraries, and other institutions typically use archive-grade sRGB images to preserve digital records of cultural relics for long-term preservation, ensuring that color information is not lost due to technological advancements. It should be noted that the quality requirements of archive-grade sRGB images (such as color fidelity and detail fidelity) are already incorporated into the parameters of the residual variational autoencoder and the diffusion model as training objectives. This ensures that the reconstructed image meets the quality requirements without requiring additional processing.
[0024] In this embodiment of the application, the training loss function of the diffusion model includes: mineral pigment spectral fidelity loss, which is used to constrain the color deviation of various mineral pigments in the reconstructed image based on the spectral features of a standard mineral pigment color chart.
[0025] In the embodiments of this application, the above-mentioned mineral pigment spectral fidelity loss can be used to suppress pigment color shift and achieve accurate color reproduction of mineral pigments.
[0026] In this embodiment, a Bayer format image of the cultural relic is acquired, and the Bayer format image is preprocessed to obtain a linear LRGB image. A residual variational autoencoder is used to encode the LRGB image to obtain latent features. These latent features are then input into a diffusion model to obtain denoised latent features. Finally, the decoder of the aforementioned residual variational autoencoder is used to decode the denoised latent features to obtain an archive-grade sRGB image. This allows for lossless enhancement of the original image of the cultural relic under ultra-low illumination, completely resolving the industry-wide conflict between cultural relic preservation and acquisition accuracy. It eliminates the need for long exposures under strong light, preventing irreversible damage to the cultural relic during the acquisition process. Furthermore, through the residual connections in the residual variational autoencoder and the region-based cross-attention mechanism in the diffusion model, archaeological information fidelity and accurate color reproduction of mineral pigments can be achieved.
[0027] In one possible implementation, the image preprocessing of the Bayer format image to obtain a linear RGB image includes: Black level subtraction is performed on the lossless RGBG four-channel data of the Bayer format image to obtain linearized original data; A preset digital gain is applied to the linearized raw data to obtain a gain image; The white balance matrix is calibrated based on the standard mineral pigment color chart, and the gain image is then subjected to white balance correction to obtain the corrected image. The corrected image is de-mosaiced using a lossless bilinear interpolation algorithm to obtain a linear LRGB image.
[0028] In the above embodiments, Bayer format ultra-low illumination original images conforming to cultural relic acquisition standards can be read, and their RGBG four-channel lossless data can be extracted. The camera-calibrated black level is subtracted to eliminate the influence of sensor dark current, resulting in linearized original data. The original photon count information is preserved throughout the process without any lossy compression. Then, compliant digital gain adjustment is performed. A preset digital gain (adjusted according to the acquisition illumination; 200-300 times gain is used for typical grotto temple acquisition scenarios) is applied to the linearized original data to brighten the image while ensuring zero noise mean. This preserves the linearity of the original data and avoids information loss during the gain process. Finally, basic ISP processing for mineral pigment adaptation is performed. Cultural relic-specific white balance correction is performed on the gained image. The white balance matrix can be calibrated based on a standard mineral pigment color chart to eliminate the influence of light color temperature on mineral pigment colors, resulting in a corrected image. Then, a lossless bilinear interpolation algorithm is used to de-mosaic the corrected image and merge the RGBG channels to obtain an LRGB image. This not only reduces the gap with the pre-trained diffusion model in the sRGB domain but also completely preserves high-frequency details such as mural lines and pigment particles.
[0029] In one possible implementation, in the residual variational autoencoder, the residual connection is implemented by adjusting the number of output channels of the encoder block through a convolutional layer and then weighting and adding it with the input of the corresponding decoder block. The weighting weight of the archaeological feature region is set to a preset multiple of the weighting weight of the non-archaeological feature region.
[0030] In practical applications, the structure of a residual variational autoencoder includes an encoder and a decoder. The encoder includes four convolutional blocks (Conv2d), and the decoder includes four deconvolutional blocks. Residual connections with archaeological feature weights are added between corresponding blocks of the encoder and decoder. The output of each encoder block is adjusted for the number of channels by the Conv2d layer and then weighted and added to the input of the corresponding decoder block to form a residual connection. This can increase the residual connection weight to a preset multiple (e.g., usually 1.2-1.5 times) for high-frequency archaeological feature areas such as mural lines, inscriptions, and pigment cracks, thus enhancing the preservation of core information.
[0031] In practical applications, the convolutional layer parameters are set as follows: the residual connection Conv2d layer between the encoder and decoder is set as follows: layer 1 (128 input channels → 256 output channels, 3×3 cores, stride 1, padding 1), layer 2 (256 → 512, 3×3 cores, stride 1, padding 1), and layers 3-4 (512 → 512, 3×3 cores, stride 1, padding 1).
[0032] In this way, the LRGB image of the cultural relic is encoded into a latent feature vector, and the core archaeological details of the cultural relic are completely preserved through weighted residual connections. Thus, when decoding, the archive-level sRGB image is reconstructed from the denoised latent features, which can prevent the loss and tampering of the original content.
[0033] In one possible implementation, the region is specifically used based on a cross-attention mechanism for: In each cross-attention layer of the denoising network, the noisy latent features are divided into K non-overlapping regions, and the latent features of each region are converted into N feature tokens; For each region, a query vector Q, a key vector K, and a value vector V are independently computed, where Q comes from the noisy latent features of the current region, and K and V come from the latent features after being processed by the context processor. The attention weights in each region are calculated using the Softmax function, and the value vector V is weighted and summed to obtain the updated region features. All updated region features are concatenated and used as the output of the cross-attention layer.
[0034] In practical applications, the aforementioned diffusion model can use a pre-trained Stable Diffusion V2-1 (a text-to-image generation model) U-Net (a symmetric encoder-decoder convolutional neural network) as the denoising network. The diffusion process is performed in the latent space; Gaussian noise is added to the latent features during the forward diffusion stage, and noise is gradually removed through the denoising network during the backward diffusion stage. This fully reuses the large-scale image generation capability of the pre-trained model without requiring training from scratch.
[0035] In the above embodiments, region-based cross-attention is mainly used for: Region partitioning: The latent features (dimensions H×W×d) of each cross-attention layer of U-Net are divided into K regions. i There are N non-overlapping regions, each of which is converted into N' feature tokens. At each cross-attention layer of the U-Net, the latent feature map (HxW in size, d channels) to be denoised is uniformly divided into K non-overlapping blocks (e.g., each block is an 8x8 local region). Let i represent the i-th region. Each small region contains all the information of the current denoising state in that local space. The above steps form the basis for implementing spatial alignment constraints, forcing the model to focus only on the corresponding local area in the conditional image when processing a local area, rather than arbitrarily acquiring information from the global image. Furthermore, it reduces computational complexity and ensures that within a small region, denoising decisions are made collaboratively based on the information of that region itself and the corresponding conditional region, avoiding interference from irrelevant information at a distance and enhancing the coherence of local detail reconstruction.
[0036] Attention computation: Query (Q), key (K), and value (V) are computed independently for each region, where Q comes from noisy latent features, K and V come from LRGB image latent features processed by the context processor, and the projection matrix (W) is used. Q W K W V This is shared across all regions, so for the i-th region, query Q... i The noisy latent feature regions to be denoised in the current U-Net are processed through a learnable projection matrix W. Q Generate. Key K i Sum V i From the clear LRGB conditional image latent features processed by the context processor, the region that strictly corresponds to the current spatial location is processed by W. K and W V Projection generation. It can represent what region i originally looked like (K) and the detailed information it contained (V) in the original map.
[0037] W Q W K W V Sharing across all regions and all samples means the model learns a general ability to compare, rather than memorizing a specific location. This allows for the establishment of precise local conditioned reflexes, ensuring that decisions regarding region i during denoising depend on information about region i in the original image. This fundamentally eliminates the possibility of tampering. Furthermore, sharing the projection matrix allows the model to focus on learning the semantic correspondences of features.
[0038] Feature Update: Attention weights within a region are calculated using the Softmax function, and the updated latent features of the region are obtained by weighted summation. All region features are then concatenated as the output of this layer. Calculating the attention weights primarily involves calculating the current noisy region (Q...). i ) and conditional region (K) i The correlation weights between feature points within the region are calculated. Areas with high weights indicate a strong correlation between the current noisy feature and the original clear feature, and should be given priority consideration. The weights from the previous step are then used to refine the detailed information V of the conditional region. i Weighted summation yields a new feature incorporating conditional information. This new feature not only reflects the requirements of the current denoising state but also deeply integrates reliable guidance information from the original image. Finally, concatenation achieves flexible and adaptive feature fusion. Thus, the Softmax attention mechanism allows the model to dynamically determine the conditional feature V. iThe system intelligently fuses the data to determine which parts are most useful for denoising. For example, when repairing a line, it focuses more on the edge information in the conditional features. It also maintains the integrity of the feature map by re-splicing all updated regions to obtain an updated latent feature map of the same size as the input, which is then fed into the next layer of U-Net for further processing, achieving seamless information flow.
[0039] In one possible implementation, the network architecture of the context processor is the same as that of the encoder in the denoising network, and during training, the weights are initialized from the corresponding layer of the encoder in the denoising network.
[0040] In practical applications, during denoising, U-Net's encoder extracts features layer by layer, from concrete to abstract. The context processor processes the conditional image in the exact same rhythm and manner, ensuring that at each layer, the conditional features it provides are at the same level of abstraction and spatial resolution as the noise features that U-Net is currently processing. The context processor is initialized by copying weights from the corresponding layer of the pre-trained U-Net encoder. It inherits powerful general visual feature extraction capabilities; the pre-trained U-Net encoder has learned how to extract key features such as semantics, edges, and texture from noise. Now, the context processor possesses the same capabilities from the outset. From the moment training begins, the context processor's understanding of the conditional image is completely consistent with the U-Net encoder's understanding of the noisy image, greatly accelerating training convergence.
[0041] In practical applications, only a small amount of cultural relic data is needed to fine-tune how to apply this general feature extraction capability to process cultural relic images, which can improve the efficiency and quality of training.
[0042] In practical applications, the DDIM sampling algorithm can be used for denoising inference in the diffusion enhancement step, and classifier-independent guidance can be introduced. The guidance weight is set according to the type of cultural relic scene. For example, the weight is set to 2.0 for grotto mural scene, 2.5 for ancient building painted scene, and 2.2 for tomb mural scene.
[0043] In one possible implementation, the formula for calculating the spectral fidelity loss of the mineral pigment is:
[0044] Where C(·) represents the pigment spectral feature map extracted from the input image, x is the input high-resolution image, and D(·) is the decoder of the residual variational autoencoder. Let E be the latent feature after denoising, E be the expectation operator, and || ...||2² be the square of the L2 norm, i.e. the mean square error.
[0045] In the above embodiments, in order to meet the requirements of color fidelity of mineral pigments in cultural relics, spectral fidelity constraints can be added on the basis of the original double loss to solve the problem of color deviation of mineral pigments under ultra-low illumination.
[0046] In practical applications, a three-loss function collaborative design can be used to optimize the diffusion model, with a total loss L=L LDM +λ1*L image +λ2*L spectrum ,in: L LDM The standard loss of the latent diffusion model constrains the noise prediction accuracy of the denoising network in the latent space, as shown in the formula: ε represents standard Gaussian noise, and w t As a weighting factor; L image The decoder reconstruction loss is defined as the difference between the reconstructed image and the real clean image constrained in the sRGB pixel space, and is expressed by the following formula: , λ represents the latent features after denoising, D is the VAE decoder, and λ is the balance coefficient (which can range from 0.1 to 0.5). L spectrum To compensate for the spectral fidelity loss of mineral pigments, the color deviation of various mineral pigments in the reconstructed image can be constrained based on the spectral characteristics of a standard mineral pigment color chart. The input image is a real, clean, and high-resolution reference image of the cultural relic, i.e., a color-accurate archived image, which can be used as the standard answer for training. When training the diffusion model, for each pair of training data (low-quality image → high-resolution real image x), the diffusion model generates an enhancement image. Using C(·) to process the real high-resolution image x and the enhanced image respectively. Perform analysis and output their spectral eigenvectors. Calculate the difference between the two eigenvectors using the mean squared error. This difference is L. spectrum This loss value is used to update the model weights. Through repeated training, the diffusion model (and its encoder and decoder) internalizes this constraint. Thus, during denoising and reconstruction, not only must the image appear clear, but each generated color must also conform to the spectral characteristics of the specific mineral pigment. This fundamentally solves the color shift problem caused by the nonlinear response of the sensor under low light conditions.
[0047] In practical applications, the three-loss synergistic optimization can take into account the latent space denoising effect, pixel space content fidelity and mineral pigment color accuracy, completely suppress pigment color deviation under ultra-low illumination, and meet the color accuracy requirements of digital archiving of cultural relics (color deviation ΔE≤2).
[0048] In one possible implementation, the method further includes: A two-stage strategy is employed when training the residual variational autoencoder and the diffusion model. In the first stage, a small sample of artifact images is used to train the residual variational autoencoder. In the second stage, the backbone weights of the pre-trained diffusion model are frozen, the region cross-attention module and the context processor are initialized and trained, and the diffusion model is jointly optimized using a multi-task loss function. During the training process in the second stage, the latent features of the LRGB image are replaced with Gaussian noise with a preset probability.
[0049] In practical applications, in the first stage, residual VAEs with archaeological feature constraints can be fine-tuned. 100-200 pairs of LRGB images of artifacts from a single cave / scene and archive-grade clean sRGB images can be used, with perceptual loss, patch adversarial loss, and KL divergence as optimization objectives. Training for 3000 epochs with a learning rate of 4e-5 requires only a very small sample size to complete the adaptation. In the second stage, the diffusion enhancement module can be fine-tuned. Pre-trained Stable Diffusion U-Net weights are loaded, and regions are initialized based on cross-attention and context processors, using L... LDM L image With L spectrum The three-loss training was conducted for 3000 epochs, with a learning rate of 2.5e-4 for the context processor and cross-attention layer, and 5e-5 for the remaining layers. During training, 5% of the latent features of the LRGB images could be randomly replaced with Gaussian noise to adapt to classifier-independent guidance. At the same time, the backbone weights of the pre-trained model were frozen, and only the newly added modules and conditional branches were fine-tuned to avoid overfitting.
[0050] In this way, by designing the above two-stage strategy for the small sample characteristics of cultural relics scenarios, we can solve the problem of scarce cultural relics datasets while ensuring the traceability and compliance of the augmentation process.
[0051] In one possible implementation, after obtaining the archive-level sRGB image, the method further includes: Calculate the first hash value corresponding to the Bayer format image; Calculate the second hash value corresponding to the key processing parameters used in the lossless enhancement process; Calculate the third hash value corresponding to the archived sRGB image; Based on the first hash value, the second hash value, and the third hash value, a notarization record is generated and uploaded to the blockchain network for storage.
[0052] In practical applications, the aforementioned key processing parameters may include: digital gain factor, white balance correction matrix, generation parameters corresponding to archaeological feature weights, number of denoising sampling steps, and classifier-independent guiding weights, etc. In this way, the original Bayer format image, preprocessing parameters, model inference parameters, and archive-level sRGB image can be stored on the blockchain, thereby enhancing the traceability and reproducibility of the entire process and meeting the relevant requirements for digital archive management of cultural relics.
[0053] The above are embodiments of the method proposed in this application. Based on the same inventive concept, embodiments of this application also provide a non-destructive enhancement device for original images of cultural relics under ultra-low illumination, the structure of which is as follows: Figure 2 As shown.
[0054] Figure 2 This is a schematic diagram of the internal structure of a non-destructive enhancement device for original low-light images of cultural relics, provided as an embodiment of this application. Figure 2 As shown, the device includes: At least one processor 201; And a memory 202 that is communicatively connected to at least one processor; The memory 202 stores instructions that can be executed by at least one processor. The instructions are executed by at least one processor 201 to enable at least one processor 201 to: perform the above-mentioned non-destructive enhancement method for the original image of the cultural relic under ultra-low light conditions.
[0055] In one possible implementation, the processor is capable of acquiring a Bayer format image of the artifact and performing image preprocessing on the Bayer format image to obtain a linear LRGB image; using a residual variational autoencoder, the LRGB image is encoded to obtain latent features, wherein a residual connection is provided between the encoder and decoder blocks of the residual variational autoencoder, and preset weights are assigned to archaeological feature regions in the LRGB image to preserve core archaeological information; the latent features are input into a diffusion model to obtain denoised latent features, wherein a region-based cross-attention mechanism is used in the denoising network of the diffusion model to divide the latent features into multiple regions, and each region is independently conditionally guided to denoise using the latent features of the processed LRGB image as a condition to obtain denoised latent features; the decoder of the residual variational autoencoder is used to decode the denoised latent features to obtain an archive-grade standard sRGB image.
[0056] Some embodiments of this application provide corresponding to Figure 1 A non-volatile computer storage medium stores computer-executable instructions, which are configured to execute a non-destructive enhancement method for the aforementioned original image of the cultural relic under ultra-low illumination.
[0057] In one possible implementation, the aforementioned computer-executable instructions are configured to acquire a Bayer format image of an artifact and perform image preprocessing on the Bayer format image to obtain a linear LRGB image; use a residual variational autoencoder to encode the LRGB image to obtain latent features, wherein a residual connection is provided between the encoder and decoder blocks of the residual variational autoencoder, and preset weights are assigned to archaeological feature regions in the LRGB image to preserve core archaeological information; the latent features are input into a diffusion model to obtain denoised latent features, wherein a region-based cross-attention mechanism is used in the denoising network of the diffusion model to divide the latent features into multiple regions, and each region is independently conditionally guided to denoise using the latent features of the processed LRGB image as a condition to obtain denoised latent features; use the decoder of the residual variational autoencoder to decode the denoised latent features to obtain an archive-grade standard sRGB image.
[0058] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for IoT devices and media are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0059] The systems, media, and methods provided in this application are one-to-one correspondences. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.
[0060] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0061] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0062] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0063] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0064] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0065] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0066] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0067] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0068] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A non-destructive enhancement method for original images of cultural relics under ultra-low illumination, characterized in that, The method includes: Acquire Bayer format images of cultural relics and perform image preprocessing on the Bayer format images to obtain linear LRGB images; The LRGB image is encoded using a residual variational autoencoder to obtain latent features. The encoder and decoder blocks of the residual variational autoencoder are connected by a residual connection. The archaeological feature regions in the LRGB image are assigned preset weights to preserve core archaeological information. The latent features are input into the diffusion model to obtain the denoised latent features. In the denoising network of the diffusion model, a region-based cross-attention mechanism is used to divide the latent features into multiple regions. The latent features of the processed LRGB image are used as conditions to perform independent condition-guided denoising on each region to obtain the denoised latent features. The latent features are decoded using the decoder of the residual variational autoencoder to obtain an archive-grade standard sRGB image.
2. The method according to claim 1, characterized in that, The step of preprocessing the Bayer format image to obtain a linear RGB image includes: Black level subtraction is performed on the lossless RGBG four-channel data of the Bayer format image to obtain linearized original data; A preset digital gain is applied to the linearized raw data to obtain a gain image; The white balance matrix is calibrated based on the standard mineral pigment color chart, and the gain image is then subjected to white balance correction to obtain the corrected image. The corrected image is de-mosaiced using a lossless bilinear interpolation algorithm to obtain a linear LRGB image.
3. The method according to claim 1, characterized in that, In the residual variational autoencoder, the residual connection is implemented by adjusting the number of output channels of the encoder block through a convolutional layer and then weighting and adding it with the input of the corresponding decoder block. The weighting weight of the archaeological feature region is set to a preset multiple of the weighting weight of the non-archaeological feature region.
4. The method according to claim 1, characterized in that, The region, based on the cross-attention mechanism, specifically includes: In each cross-attention layer of the denoising network, the noisy latent features are divided into K non-overlapping regions, and the latent features of each region are converted into N feature tokens; For each region, a query vector Q, a key vector K, and a value vector V are independently computed, where Q comes from the noisy latent features of the current region, and K and V come from the latent features after being processed by the context processor. The attention weights within each region are calculated using the Softmax function, and the V is then weighted and summed to obtain the updated region features. All updated region features are concatenated and used as the output of the cross-attention layer.
5. The method according to claim 4, characterized in that, The network architecture of the context processor is the same as that of the encoder in the denoising network, and the weights are initialized from the corresponding layer of the encoder in the denoising network.
6. The method according to claim 1, characterized in that, The formula for calculating the spectral fidelity loss of the mineral pigment is as follows: Where C(·) represents the pigment spectral feature map extracted from the input image, x is a high-resolution image, and D(·) is the decoder of the residual variational autoencoder. Let || ... ||2² be the latent feature after denoising, || ... ||2² be the square of the L2 norm, and E be the expectation operator.
7. The method according to claim 1, characterized in that, The method further includes: A two-stage strategy is employed when training the residual variational autoencoder and the diffusion model. In the first stage, a small sample of artifact images is used to train the residual variational autoencoder. In the second stage, the backbone weights of the pre-trained diffusion model are frozen, the region cross-attention module and the context processor are initialized and trained, and the diffusion model is jointly optimized using a multi-task loss function. During the training process in the second stage, the latent features of the LRGB image are replaced with Gaussian noise with a preset probability.
8. The method according to claim 1, characterized in that, After obtaining the archive-level sRGB image, the method further includes: Calculate the first hash value corresponding to the Bayer format image; Calculate the second hash value corresponding to the key processing parameters used in the lossless enhancement process; Calculate the third hash value corresponding to the archived sRGB image; Based on the first hash value, the second hash value, and the third hash value, a notarization record is generated and uploaded to the blockchain network for storage.
9. A non-destructive enhancement device for original images of cultural relics under ultra-low illumination, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform a non-destructive enhancement method for an original image of a cultural relic under ultra-low illumination as described in any one of claims 1-8.
10. A computer storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, they implement a non-destructive enhancement method for the original image of a cultural relic under ultra-low illumination as described in any one of claims 1-8.