Quantum level image reconstruction and dynamic spectral restoration method

By acquiring interlaced pixel information in a single scan and performing guided filtering reconstruction and adaptive fusion, the problems of image sensor dynamic range limitation and linear array scanning timing offset are solved, achieving efficient and high-quality high dynamic range image reconstruction, eliminating spatial misalignment and artifacts, and improving image acquisition efficiency and quality.

CN122289093APending Publication Date: 2026-06-26BEIJING ZHONGKEZHIYAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHONGKEZHIYAN TECH CO LTD
Filing Date
2026-05-29
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, the limited dynamic range of image sensors under single-exposure conditions makes it difficult for images to simultaneously and completely retain effective details in both bright and dark areas. The timing shift of linear array scanning leads to inherent spatial misalignment and information inconsistency between images exposed at different times. Post-processing methods based on feature matching or global registration are difficult to eliminate errors. Pixel weight-based fusion strategies are prone to producing discontinuous fusion, excessive smoothing, or loss of details in the transition between bright and dark areas. The applicability of deep learning models for deployment in industrial settings is also limited.

Method used

By simultaneously acquiring staggered pixel information under different exposure conditions in a single scan, guided filtering reconstruction and adaptive fusion are used to introduce a quantum-level image modeling and improved structural similarity loss constraint image reconstruction network to reconstruct high dynamic range images.

Benefits of technology

It achieves efficient acquisition and high-quality reconstruction of high dynamic range images, eliminates temporal misalignment between images with different exposures, ensures spatial consistency, avoids discontinuities and artifacts in the fusion of bright and dark transition areas, and improves acquisition efficiency and fusion quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289093A_ABST
    Figure CN122289093A_ABST
Patent Text Reader

Abstract

This invention relates to the fields of image processing and industrial inspection, and provides a quantum-level image reconstruction and dynamic spectral restoration method, comprising: acquiring an original row scan image obtained from a single camera scan, containing information on various exposure pixels arranged in an interleaved column pattern; separating the original row scan image into various sparse exposure images based on a preset pixel arrangement rule; using the original row scan image or a preliminary interpolated image as a guide image, performing guided filtering on the various sparse exposure images to fill in missing pixels, generating various spatially aligned and uniformly reconstructed images with different exposures; and performing weighted fusion on the various reconstructed images with different exposures to generate a fused image, wherein the weights are adaptively determined based on the brightness information of the pixel positions. This invention avoids temporal misalignment and registration errors from multi-frame acquisition, improves acquisition efficiency, and effectively suppresses artifacts and enhances details in dark areas through guided filtering and adaptive fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of high dynamic range imaging technology, and in particular to a quantum-level image reconstruction and dynamic spectral restoration method. Background Technology

[0002] In industrial inspection and high dynamic range imaging applications, the inherent limitations of the dynamic range of image sensors under single-exposure conditions make it difficult to simultaneously and completely retain effective detail information in both bright and dark areas. To address this, multi-exposure image fusion technology is commonly used. This involves acquiring multiple images of the same scene with different exposure parameters using an image acquisition device, and then generating a fused image with expanded dynamic range through image registration, weight calculation, and fusion algorithms. In typical implementations based on area scan cameras, acquiring multiple exposure images usually relies on continuous shooting or multiple exposures of the same field of view. The acquired images are then aligned to eliminate displacement errors caused by device shake or target motion, and weights are calculated and superimposed based on indicators such as pixel brightness, contrast, or saturation. However, when the image acquisition method changes from an area scan camera to a camera, the above technical path faces fundamental obstacles. Specifically, cameras rely on scanning motion to achieve two-dimensional imaging. Acquiring multi-exposure information typically requires acquiring multiple lines of scan data of the same target during the scanning process by changing the exposure time or light source intensity, and then reconstructing this data into two-dimensional images corresponding to different exposure parameters. Because linear scan has strict temporal characteristics, the different exposure images obtained by the above method correspond to different moments in the scanning process in the time dimension. There is an unavoidable temporal offset between the exposure times of each frame image, which leads to inherent spatial misalignment and information inconsistency between different exposure images. Post-processing methods based on feature matching or global registration are difficult to fundamentally eliminate such errors.

[0003] Furthermore, even with spatial alignment completed, the currently prevalent pixel-weight-based fusion strategies still exhibit significant shortcomings when handling transitional areas between light and dark regions. When weight allocation is unreasonable, these methods are prone to issues such as discontinuous fusion, excessive smoothing, or loss of detail in transitional areas, and may even introduce significant artifacts, directly reducing the usability of the fused image. In addition, some solutions attempt to incorporate deep learning models to improve fusion quality; however, the performance of these models is highly dependent on the distribution of training data. When the acquisition conditions or scene characteristics differ from the training data, reconstruction bias or detail distortion can easily occur. Moreover, the high requirements for the scale of labeled data and computational resources required for model training further limit their applicability in industrial environments. Summary of the Invention

[0004] Based on this, it is necessary to address the inherent limitations of the dynamic range of image sensors under single-exposure conditions in existing solutions, which makes it difficult to simultaneously and completely retain effective details in both bright and dark areas. Furthermore, scanning time-series offsets lead to inherent spatial misalignment and information inconsistencies between images exposed at different times, and post-processing methods based on feature matching or global registration are insufficient to fundamentally eliminate these errors. Pixel-weight-based fusion strategies are prone to discontinuous fusion, over-smoothing, or loss of detail in transitional bright and dark areas, even introducing significant artifacts. Additionally, the performance of deep learning models is highly dependent on the distribution of training data; when the acquisition conditions or scene characteristics differ from the training data, reconstruction bias or detail distortion can easily occur, and the requirements for the scale of labeled data and computational resources limit their applicability in industrial environments. Therefore, a quantum-level image reconstruction and dynamic spectral restoration method is proposed. This method simultaneously acquires staggered pixel information under different exposure conditions in a single scan and performs guided filtering reconstruction and adaptive fusion. It further introduces quantum-level image modeling and an improved structural similarity loss-constrained image reconstruction network and reconstruction quality scoring mechanism, achieving efficient acquisition and high-quality reconstruction of high dynamic range images.

[0005] This invention provides a quantum-level image reconstruction and dynamic spectral restoration method, comprising: A raw line scan image is acquired by the camera during a single scan. Pixels at different column positions in the raw line scan image correspond to different exposure parameters, so that the raw line scan image contains pixel information under various exposure parameters arranged in an interleaved column manner. Based on a preset pixel arrangement rule, the original line scan image is separated into multiple sparse exposure images, including a first sparse exposure image and a second sparse exposure image. In the first sparse exposure image, the pixel position corresponding to the low exposure pixel information is a missing pixel, and in the second sparse exposure image, the pixel position corresponding to the high exposure pixel information is a missing pixel. Using the original line scan image or the corresponding preliminary interpolated image as the guide image, guide filtering operations are performed on the first sparse exposure image and the second sparse exposure image respectively to fill in the missing pixels in the first sparse exposure image and the second sparse exposure image, and correspondingly generate a spatially aligned first exposure reconstructed image and a second exposure reconstructed image with the same resolution. The first exposure-reconstructed image and the second exposure-reconstructed image are subjected to weighted fusion processing to generate a fused image, wherein the weights participating in the weighted fusion are adaptively determined based on the brightness information at the pixel position.

[0006] In one embodiment, separating the original line scan image into a first sparse exposure image and a second sparse exposure image based on a preset pixel arrangement rule specifically includes: The original line scan image is pixel-filtered using an exposure mask, wherein the first sparse exposure image is obtained by extracting using a first exposure mask, and the second sparse exposure image is obtained by extracting using a second exposure mask. The separation process is represented as follows: , , In the formula Indicates the second exposure sparse image in the row index Column index Pixel value at that location, Indicates the row index of the sparse image in the first exposure. Column index Pixel value at that location, This indicates that the original row scan image is at the row index. Column index Pixel value at that location, These are binary masks corresponding to the second exposure column position and the first exposure column position, respectively.

[0007] In one embodiment, the step of using the original line scan image or the corresponding preliminary interpolated image as a guide image to perform guided filtering operations on the first sparse exposure image and the second sparse exposure image respectively includes: The original line scan image Alternatively, the preliminary interpolation result of the corresponding exposed image can be used as the guide image G, with the second exposed sparse image. Or the first exposure sparse image As input image P, guided filtering reconstruction is performed, so that when filling in missing pixels, the edge information and texture information in the guided image G are used to constrain the estimated direction and magnitude of the missing pixels, so as to maintain the structural consistency between the first exposure reconstructed image and the second exposure reconstructed image in the edge region.

[0008] In one embodiment, the weighted fusion process of the first exposure-reconstructed image and the second exposure-reconstructed image to generate a fused image includes: Weighted fusion is performed according to the following formula: , In the formula, This represents the pixel value at position X in the merged image. These are the pixel values ​​at position X of the second exposure-reconstructed image and the first exposure-reconstructed image, respectively. For adaptive weights.

[0009] In one embodiment, the weights participating in the weighted fusion are adaptively determined based on the brightness information at the pixel location, wherein the adaptive weights are determined using a non-linear mapping form based on the brightness difference: , In the formula, K is the adjustment parameter.

[0010] In one embodiment, after generating the fused image, the method further includes: The reconstructed image from the first exposure after interpolation. Second exposure reconstructed image and the fused image As input data, construct a multi-channel input tensor ; The input tensor is fed into an image reconstruction network for feature extraction and fusion. The image reconstruction network extracts brightness features, edge information and texture structure in images with different exposures through multi-layer convolution operations, and establishes a mapping relationship between different exposures. The image reconstruction network outputs a set of spatially correlated weight maps. This is used to represent the contribution of different exposure information at each pixel location, and a reconstructed image is generated based on this weight mapping: , In the formula, This represents the pixel value of the reconstructed image at position X. This represents the spatial correlation weight mapping at location X. These are the pixel values ​​at position X in the second exposure reconstructed image and the first exposure reconstructed image, respectively.

[0011] In one embodiment, the image reconstruction network incorporates a quantum-level image modeling mechanism during training to perform fine-grained quantization of image brightness information, specifically including: The image brightness space is divided into multiple discrete quantization intervals to construct a quantized brightness representation. ,in This represents the quantization function, used to map continuous brightness to discrete quantum levels; By incorporating quantized brightness information as a constraint into the network optimization, the model learns the fine structural features of the brightness distribution while reconstructing the image.

[0012] In one embodiment, the image reconstruction network employs an improved structural similarity loss function for constraint optimization. This improved structural similarity loss function introduces a window-level maximum structural similarity selection mechanism, including: The image is divided into multiple local windows, and multiple candidate structure similarity values ​​are calculated within each window. Where K represents different exposure information or different feature mapping channels, Indicates the window area; The improved structural similarity is defined as follows: , Based on this improved definition, the overall loss function is constructed as follows: , In the formula, M represents the total number of windows. This loss function enables the model to prioritize the retention of optimal structural information and adaptively select the optimal structural source in different exposure areas.

[0013] In one embodiment, after generating the reconstructed image through the image reconstruction network, a reconstruction quality scoring step is further included to quantitatively evaluate the reliability and accuracy of the reconstruction results. This specifically includes image structure quality scoring and spectral reconstruction accuracy scoring, wherein: The image structure quality score, based on the improved structural similarity index, divides the reconstructed image into multiple local windows. Calculate structural similarity under different exposures or different feature responses within each window. The maximum structural response is selected as the evaluation value for that window to obtain the overall structural quality score. , The spectral reconstruction accuracy score is obtained by comparing the reconstructed spectral information with the reference spectral data and calculating the spectral error index. The structural quality score and the spectral accuracy score are combined to obtain the overall reconstruction quality evaluation result: , In the formula, Indicates the spectral reproduction accuracy score. and These are the weighting coefficients.

[0014] In one embodiment, the original row scan image is acquired by a camera during a single scan, wherein the camera, during the scan, acquires the first row scan image... Row pixel data is represented as: , In the formula, Indicates the scanned row index. Indicates the pixel column index. Indicates scene irradiance, This indicates the exposure time corresponding to the pixel, and the odd and even columns or grouped columns of the original row scan image correspond to different exposures.

[0015] The present invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the quantum-level image reconstruction and dynamic spectral restoration method as described above.

[0016] The aforementioned quantum-level image reconstruction and dynamic spectral restoration method, under camera-based image acquisition conditions, utilizes a technique that acquires information on different exposure pixels arranged in an interleaved column pattern during a single scan. This allows the first and second exposure pixel information to be acquired synchronously within the same row of scanned images. This fundamentally avoids the temporal shift problem caused by multiple independent exposures or scans in traditional schemes, where different exposure images belong to different times. It eliminates the inherent temporal dimension misalignment between different exposure images, thus ensuring spatial consistency of multi-exposure information. Furthermore, by separating the original row-scanned image into a first-exposure sparse image and a second-exposure sparse image according to a preset arrangement rule, and using the original row-scanned image or its preliminary interpolated image as a guide image to perform guided filtering on each sparse image to fill in missing pixels, the interpolation reconstruction process can adaptively constrain the estimation direction and amplitude of missing pixels using edge and texture structure information in the guide image. This avoids interpolation artifacts or edge blurring at the boundary between light and dark areas due to adjacent pixels belonging to different exposure channels, ensuring spatial consistency and detail integrity in the structure preservation between the first and second-exposure reconstructed images. After obtaining the aforementioned spatially aligned and detailed complete exposure image pairs, the weighted fusion weights are adaptively determined based on the brightness information at the pixel location. This allows the bright areas to preferentially adopt the unsaturated information of the second exposure image, and the dark areas to preferentially adopt the higher signal-to-noise ratio information of the first exposure image. Furthermore, the weights are smoothly changed in the transition areas. Thus, based on the exposure interleaving information obtained in a single scan, a coherent process from sparse sampling to complete reconstruction and then to adaptive fusion is completed. This effectively suppresses problems such as discontinuous fusion, excessive smoothing, or loss of detail in the transition areas between bright and dark areas, achieving a dynamic range expansion that balances acquisition efficiency and fusion quality. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the quantum-level image reconstruction and dynamic spectral restoration method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the quantum-level image reconstruction and dynamic spectral restoration method according to an embodiment of the present invention; Figure 3 This is a view showing a partial exposure anomaly in the camera according to an embodiment of the present invention; Figure 4 This is a partial structural view of the metal panel according to an embodiment of the present invention; Figure 5 This is a view of the brightness distribution of the camera image acquisition area according to an embodiment of the present invention; Figure 6 This is a view showing partial details of a camera image according to an embodiment of the present invention; Figure 7 This is an internal structural diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The following is combined with Figures 1-7 The present invention describes a quantum-level image reconstruction and dynamic spectral restoration method.

[0021] like Figure 1 and Figure 2 As shown, in one embodiment, a quantum-level image reconstruction and dynamic spectral restoration method, through the fusion of single-scan acquisition and subsequent pixel-level reconstruction, solves the problems in existing technologies such as the inherent limitations of the dynamic range of image sensors under single-exposure conditions, making it difficult to simultaneously and completely retain effective details of bright and dark areas in the acquired image; the inherent spatial misalignment and information inconsistency caused by the temporal offset of linear array scanning, which post-processing methods based on feature matching or global registration cannot fundamentally eliminate such errors; and the problems that pixel weight-based fusion strategies are prone to producing discontinuous fusion, excessive smoothing, or loss of details or even introducing significant artifacts in bright-dark transition areas. This method achieves efficient acquisition and high-quality reconstruction of high dynamic range images. Its core processing flow mainly includes the following steps: Step S110: Acquire a raw line scan image, which is captured by the camera during a single scan. Pixels at different column positions in the raw line scan image correspond to different exposure parameters, resulting in the raw line scan image containing pixel information under various exposure parameters arranged in an interleaved column pattern.

[0022] In high-speed continuous scanning scenarios such as industrial production lines, the target being measured is often in rapid motion. If traditional area scan cameras or multi-frame acquisition methods are used to acquire images with different exposures, the strict temporal characteristics of linear scans result in an unavoidable temporal offset between different exposure images. This leads to inherent spatial misalignment and information inconsistency, and post-processing methods based on feature matching or global registration are insufficient to fundamentally eliminate such errors. This embodiment utilizes the sensor pixel arrangement structure of the camera to simultaneously acquire pixel information under different exposure conditions during a single scan, thereby forming raw data with staggered exposures. The staggered column arrangement here has crucial physical significance: it means that in the same row of scan data, adjacent or regularly spaced columns of pixels correspond to long and short exposures, respectively. For example, odd-numbered columns correspond to high exposure times while even-numbered columns correspond to low exposure times, or they can be arranged in groups of four, alternating between them. Differences in exposure parameters can be achieved not only by controlling the pixel-level exposure time but also by adjusting the gain coefficient of the corresponding column or the intensity of the light source. By using this spatially staggered acquisition within a single scan, the problem of misalignment in the temporal dimension is completely avoided. This fundamentally eliminates the inherent spatial misalignment and information inconsistency caused by temporal offset in multi-frame fusion, significantly improving data acquisition efficiency and spatiotemporal consistency.

[0023] Step S120: Based on a preset pixel arrangement rule, the original line scan image is separated into a first sparse exposure image and a second sparse exposure image. In the first sparse exposure image, the pixel positions corresponding to low-exposure pixel information are empty pixels, and in the second sparse exposure image, the pixel positions corresponding to high-exposure pixel information are empty pixels.

[0024] Since the original row-scan image is composite data consisting of alternating pixels from the first and second exposures, decoupling and separation are required according to the aforementioned column-interleaved arrangement rule in order to extract pure information from a single exposure channel. The essence of this separation operation is to split the composite image into two independent sub-images based on column position indices. However, this separation process directly leads to the formation of missing pixels: in the extracted first-exposure sparse image, pixel positions originally belonging to the second-exposure column are removed or set to zero, creating data gaps at those positions; similarly, pixel positions originally belonging to the first-exposure column in the second-exposure sparse image also become missing pixels. Therefore, the two separated sparse images exhibit high spatial non-uniformity and sparsity, retaining only the effective pixel values ​​corresponding to the column positions of their respective exposure channels, while a large number of missing pixels urgently need to be filled and reconstructed in subsequent steps.

[0025] Step S130: Using the original line scan image or the corresponding preliminary interpolation image as the guide image, guide filtering operations are performed on the first exposure sparse image and the second exposure sparse image respectively to fill in the missing pixels in the first exposure sparse image and the second exposure sparse image, and correspondingly generate a spatially aligned and consistent first exposure reconstructed image and a second exposure reconstructed image.

[0026] To address the sparse pixel gap problem, traditional methods such as bilinear interpolation or mean interpolation are often used for filling in gaps. Since adjacent pixels often belong to different exposure channels and have significant brightness differences, severe interpolation artifacts and edge blurring are easily generated at the boundary between light and dark areas. This embodiment introduces a guided filtering mechanism, whose structural advantage lies in its ability to utilize the rich edge and texture information in the guide image to constrain the estimation direction and magnitude of the gap pixels. The guide image can be the original line scan image containing a complete structural outline, or an image containing a rough outline obtained after preliminary simple interpolation of the sparse image. During the filtering reconstruction process, the guide image acts as a spatial constraint template, forcing the filling values ​​of the gap pixels to not only transition smoothly numerically but also maintain consistency with the local edge and texture direction geometrically. This reconstruction mechanism based on structural consistency ensures that the final generated first-exposure reconstructed image and second-exposure reconstructed image are not only perfectly aligned in resolution and spatial coordinates but also maintain a high degree of consistency in key structural features such as edge outlines, laying a solid data foundation for subsequent artifact-free fusion.

[0027] Step S140: Perform weighted fusion processing on the first exposure reconstructed image and the second exposure reconstructed image to generate a fused image. The weights involved in the weighted fusion are adaptively determined based on the brightness information at the pixel location.

[0028] The two reconstructed images respectively contain clear details in the dark areas and rich information in the bright areas of the scene, but a single image still cannot cover the full dynamic range. The purpose of weighted fusion is to integrate the advantageous information of both, and adaptive weight determination is the core mechanism for achieving a smooth transition and avoiding local distortion. During the fusion process, the fusion weight at each pixel position is not a fixed constant, but is dynamically calculated based on the local brightness difference or brightness distribution at that position in the first and second exposure reconstructed images. For example, in the bright areas of the scene, the second exposure image often contains more reliable unsaturated information, and the adaptive weight will tend to allocate a larger proportion to the second exposure image; conversely, in the dark areas, the first exposure image has a higher signal-to-noise ratio, and the weight will adaptively tilt towards the first exposure image. Through this brightness-driven adaptive weight allocation, the fused image can automatically select the optimal exposure source in different brightness areas, achieving a smooth and natural transition between bright and dark areas. It effectively suppresses the overexposure and saturation problem in bright areas and significantly enhances the visibility of details in dark areas. It effectively avoids the problems of discontinuous fusion, excessive smoothing, or loss of detail or even significant artifacts that are easily generated in the bright-dark transition area by traditional pixel weight-based fusion strategies. As a result, a fused image with high dynamic range and high-quality visual effects is generated.

[0029] In one embodiment, the specific implementation mechanism of front-end acquisition and pixel-level reconstruction is further refined, forming a complete processing link from data source decoupling to structural fidelity.

[0030] Regarding the acquired raw line scan images, these images are captured by the camera during a single scan. Specifically, during the scan, the camera captures the first... Row pixel data is represented as: , In the formula, Indicates the scanned row index. Indicates the pixel column index. Indicates scene irradiance, This indicates the exposure time corresponding to that pixel. The odd and even columns or grouped columns of the original row scan image correspond to different exposures.

[0031] Specifically, the camera's sensor hardware layout allows for the application of independent exposure control parameters to pixels in different columns. This is achieved through the formula described above. The system can configure exposure time differently based on column index dimensions. For example, in a typical odd-even column interleaving configuration, when... When it is an odd sequence, It is set to a long exposure time (e.g., 10ms) to acquire the first exposure pixel information; when When the sequence is even, The exposure time is set to a short time (e.g., 2ms) to acquire second-exposure pixel information. It should be understood that alternating odd and even columns is only a preferred arrangement rule; in other embodiments, grouped column alternation can also be used, for example, grouping four columns together, with the first three columns using long exposure and the fourth column using short exposure, as long as pixel information with different exposure parameters can be included simultaneously in the same row of data within a single scan. This spatially interleaved acquisition mechanism within a single scan fundamentally eliminates the inherent spatial misalignment and information inconsistency caused by the temporal offset of linear array scanning in traditional multi-frame fusion, providing a highly spatiotemporally consistent original data source for subsequent high-quality pixel-level reconstruction.

[0032] Based on a preset pixel arrangement rule, the original line scan image is separated into a first sparse exposure image and a second sparse exposure image. Specifically, this includes: filtering pixels in the original line scan image using an exposure mask, wherein the first sparse exposure image is obtained by extracting the first exposure mask, and the second sparse exposure image is obtained by extracting the second exposure mask. The separation process is represented as follows: , , In the formula, This indicates that the original row scan image is at the row index. Column index Pixel value at that location, These are binary masks corresponding to the low-exposure column position and the high-exposure column position, respectively.

[0033] Specifically, the above formula achieves precise decoupling of complex interleaved data through the multiplication operation of binary masks. Binary mask The value of is only 0 or 1, and its physical meaning is to perform a hard selection of column positions. Taking alternating odd and even columns as an example, when j is an even column (corresponding to the second exposure), The value is 1 and The value is 0; when j is an odd number (corresponding to the first exposure), The value is 0 and The value is set to 1. Through this binarized mask separation operation, the original line scan image is split into two independent sub-images. However, this separation operation directly leads to the formation of missing pixels: mathematically, when the mask value is 0, the pixel value at the corresponding position is forced to be 0, which means that the other exposure information that originally existed at that position is completely removed, thus leaving data gaps in the spatial distribution. Therefore, the separated first and second exposure sparse images are essentially incompletely sampled sparse matrices, retaining only the effective pixel values ​​at the corresponding column positions of their respective exposure channels, while a large number of missing pixel positions urgently need to be structurally complete and reconstructed in subsequent steps.

[0034] The guided filtering operation is performed using the original line scan image or the corresponding preliminary interpolated image as the guide image. Specifically, this includes: using the original line scan image... Alternatively, the preliminary interpolation result of the corresponding exposed image can be used as the guide image G, with the second exposed sparse image. Or the first exposure sparse image As input image P, guided filtering reconstruction is performed, so that when filling in missing pixels, the edge information and texture information in the guided image G are used to constrain the estimated direction and magnitude of the missing pixels, so as to maintain the structural consistency between the first exposure reconstructed image and the second exposure reconstructed image in the edge region.

[0035] Specifically, to address the sparse gaps resulting from mask separation, traditional methods such as bilinear interpolation or mean interpolation are often used for filling these gaps. Since adjacent pixels often belong to different exposure channels and have significant brightness differences, severe interpolation artifacts and edge blurring are easily generated at the boundary between light and dark areas. This embodiment introduces a guided filtering mechanism, whose core advantage lies in its ability to utilize the rich edge and texture information in the guided image G to constrain the estimation direction and amplitude of the missing pixels. The guided image G can be obtained from two sources: one is by directly using the original line scan image containing the complete structural outline. Another approach is to obtain an image containing a general outline after performing preliminary simple interpolation on the sparse image. At the microscopic level of filtering and reconstruction, guided filtering does not perform non-directional numerical smoothing in the missing regions, but rather treats the guide image G as a spatially constrained template. When missing pixels are located in the edge regions or areas with complex textures, the guided filter adaptively adjusts the filtering coefficients based on the gradient direction and magnitude of that local region in the guide image G. This forces the filling value of the missing pixels to not only smoothly transition numerically with the surrounding valid pixels, but also to maintain strict alignment geometrically with the local edge direction and texture features. This reconstruction mechanism based on structural consistency ensures that the final first-exposure reconstructed image and the second-exposure reconstructed image are not only perfectly aligned in resolution and spatial coordinates, but also maintain a high degree of consistency in key structural features such as edge contours. This effectively avoids artifacts at the boundary between light and dark areas, laying a solid data foundation for subsequent seamless fusion.

[0036] In one embodiment, to address the problem that existing pixel weight-based fusion strategies are prone to producing discontinuous fusion, excessive smoothing, or loss of detail or even introducing significant artifacts in the bright-dark transition region, this embodiment further refines the specific mapping mechanism of the weighted fusion process, providing specific mathematical support and parameter safeguards for determining adaptive weights.

[0037] The first exposure reconstructed image and the second exposure reconstructed image are weighted and fused to generate a fused image. Specifically, the weighted fusion is performed using the following formula: , In the formula, This represents the pixel value at position X in the merged image. These are the pixel values ​​at position X in the second exposure reconstructed image and the first exposure reconstructed image, respectively. For adaptive weights.

[0038] Specifically, this formula essentially establishes a pixel-level blending framework based on spatial location. Adaptive weights The value range of is strictly limited to a closed interval between 0 and 1, and its physical meaning lies in characterizing the contribution ratio of the second exposure reconstructed image to the final fusion result. When When the value approaches 1, the merged pixels primarily inherit information from the reconstructed image from the second exposure. This typically occurs in the highlight areas of the scene because the second exposure image is less prone to overexposure and saturation in such areas, preserving more reliable texture details; conversely, when... When the signal strength approaches 0, the fused pixels primarily rely on the first exposure to reconstruct the image. This corresponds to the dark areas of the scene, where the first exposure image has a higher signal-to-noise ratio and visibility. It should be understood that while this embodiment provides a specific form of the linear weighted mixing described above, in other embodiments, any mechanism that dynamically adjusts the contribution ratio of the two exposure information streams based on local brightness information—such as using a nonlinear combination or a mapping function that introduces a bias term—should be considered to fall within the protection scope of this weighted fusion framework.

[0039] Furthermore, the weights participating in the weighted fusion are adaptively determined based on the brightness information at the pixel location. These adaptive weights are determined using a non-linear mapping based on the brightness difference. , In the formula, K is an adjustment parameter used to control the sensitivity of weight changes, so that better exposure information is adaptively selected in different brightness areas to achieve a smooth transition between bright and dark areas of the image.

[0040] This formula introduces the Sigmoid function as the core mapping mechanism, whose microscopic principle lies in using the brightness difference between the two reconstructed images as the driving source of the weights. In multi-exposure imaging systems, the difference in brightness response of the same scene point under different exposure conditions directly reflects the illumination intensity attribute of that point: in dark areas, the pixel value of the first exposure is significantly greater than that of the second exposure, at which point the brightness difference... For larger positive values, after Sigmoid mapping When the value approaches 0, the system automatically selects high exposure information; in bright areas, the second exposure pixel value may be greater than the first exposure pixel value, which has already exceeded saturation, resulting in a negative brightness difference. When the value approaches 1, the system automatically selects low-exposure information. Choosing the Sigmoid non-linear mapping instead of a simple linear piecewise function has a structural advantage: the Sigmoid function has a smooth, monotonically decreasing characteristic, enabling continuous and gradual weight changes in the bright-dark boundary region. This fundamentally avoids the weight jumps and blending seam artifacts produced by linear piecewise functions at the threshold.

[0041] More importantly, the parameter K acts as a valve controlling the sensitivity of weight changes in the mapping function, and its value directly determines the width and steepness of the blending transition zone. To solidify the parameter's defenses, the range of K values ​​and its physical effects must be explained in detail. For example, when K is 0.5, it represents a low-sensitivity configuration. In this case, the Sigmoid curve is relatively flat, and the transition range of the weights with brightness difference is very wide. This configuration allows the blending process to retain the mixed information of the two exposures simultaneously over a large brightness difference range. While this ensures an extremely smooth and seamless transition, it may not be able to quickly switch to the optimal exposure channel in extremely bright and dark areas, resulting in a slight reduction in local contrast. When K is 5, it represents a high-sensitivity configuration. The Sigmoid curve is steep, and the weights rapidly switch between 0 and 1 near the critical point where the brightness difference is zero. This configuration allows the system to decisively select the locally optimal exposure source, maximizing the preservation of high-contrast details, but the transition range is relatively narrow. In practical industrial inspection applications, the optimal value of K is usually set between 1 and 3 to achieve the best balance between smooth transition and detail preservation.

[0042] To further demonstrate the necessity of the aforementioned parameter value range, a comparative example must be provided to illustrate the serious problems caused by an inappropriate value for K. If K is set to 0, the Sigmoid function degenerates into a constant 1, and the weights become ineffective regardless of changes in brightness difference. The value is always equal to 0.5. This means that the fusion process degenerates into a simple mean averaging of the two reconstructed images, completely losing its adaptive selection capability. This results in overexposed highlights and blurred details in dark areas, failing to leverage the dynamic range expansion advantage of multi-exposure fusion. Conversely, if K takes a maximum value (e.g., approaching infinity), the Sigmoid function degenerates into a hard-switching step function, and the weights... When the brightness difference crosses zero, it jumps instantaneously from 1 to 0. This hard-switching mechanism produces severe brightness abrupt changes and mosaic-like artifacts at the boundary between light and dark areas. Pixels in the transition area either come entirely from the first exposure or entirely from the second exposure, lacking a smooth spatial transition, resulting in a visually obvious stitching boundary in the merged image. This embodiment, through the above comparative examples, clearly excludes the extreme value range of K, proving the irreplaceable nature of reasonably setting the K parameter to achieve a smooth nonlinear transition, thus building a solid numerical defense and mechanistic support for the functional constraints of adaptive weights.

[0043] In one embodiment, a deep learning reconstruction network is further introduced, realizing an algorithmic leap from traditional filtering fusion to network feature fusion, providing a data-driven solution for in-depth mining of image details and improvement of dynamic range expression accuracy.

[0044] After generating the fused image, the first exposure reconstructed image after interpolation is used. Second exposure reconstructed image and fused images As input data, construct a multi-channel input tensor Specifically, the construction logic of the multi-channel input tensor lies in integrating information sources from different levels and exposure dimensions. Second exposure reconstructs the image. Provides reliable texture and unsaturated detail in highlighted areas, and reconstructs the image using the first exposure. It contains high signal-to-noise ratio and visibility information of dark areas, and the preliminary fused image This provides a macroscopic reference base for global brightness balance. The tensor X formed by stitching these three along the channel dimension is equivalent to providing a complete panoramic view of multi-exposure information for the subsequent neural network, avoiding information bias caused by a single input source, and enabling the network to make cross-channel reference comparisons at any time during local feature extraction, thereby establishing a more accurate mapping relationship.

[0045] The input tensor is fed into an image reconstruction network for feature extraction and fusion. This network extracts brightness features, edge information, and texture structures from images with different exposures through multi-layer convolutional operations, establishing mapping relationships between different exposures. Specifically, image reconstruction networks typically employ a Convolutional Neural Network (CNN) architecture. As its multi-layer convolutional kernels slide spatially, they automatically learn and extract visual features from low to high levels. In shallow networks, convolutional operations primarily extract brightness gradients and local edge information; in deeper networks, through the expansion of the receptive field and the abstraction of features, the network begins to capture more complex texture structures and cross-regional brightness distribution patterns. Crucially, the network does not process a single channel in isolation but establishes deep mapping relationships between different exposure conditions through cross-channel convolution and feature fusion mechanisms. This mapping relationship transcends simple pixel-level weighting; it understands complex cross-exposure physical relationships, such as "how the texture of a dark area in the first exposure image corresponds to the noise distribution in the same area in the second exposure image," thereby achieving true structural fidelity and detail enhancement during reconstruction.

[0046] Image reconstruction networks output a set of spatially correlated weight maps. This is used to represent the contribution of different exposure information at each pixel location, and a reconstructed image is generated based on this weight mapping: , In the formula, This represents the pixel value of the reconstructed image at position X. This represents the spatial correlation weight mapping at location X. and These are the pixel values ​​at position X in the second exposure reconstructed image and the first exposure reconstructed image, respectively.

[0047] Specifically, the core defense depth of this embodiment lies in spatially related weight mapping. With the aforementioned adaptive weights The essential difference. It is artificially designed based on a nonlinear mapping of brightness differences (such as the Sigmoid function). Its weight calculation logic is a preset, fixed mathematical formula. While it can achieve a smooth transition between light and dark areas on a macroscopic scale, its responsiveness to complex local textures, weak edges, and cross-exposure structural relationships is limited by the expressive power of the function itself. In this embodiment... It is a spatially relevant weight mapping learned by an image reconstruction network through a data-driven approach. Its micro-mechanism lies in the fact that during training, the network continuously adjusts the convolution parameters through backpropagation of a loss function based on a large number of samples, thereby achieving... The generation of this signal not only relies on the current local brightness difference but also deeply integrates contextual information from the surrounding region, multi-scale feature responses, and structural similarity across exposures. This means that in weakly textured regions at the boundary between light and dark areas, the network may adaptively output a non-monotonic or locally fine-tuned signal. This data-driven weight mapping mechanism precisely preserves key details in a particular exposure, rather than simply performing a coarse-grained smooth transition like the Sigmoid algorithm. This gives the system stronger adaptive capabilities and detail extraction abilities under complex lighting conditions, resulting in a more refined reconstructed image. Significant leaps have been achieved in edge sharpness, texture fidelity, and dynamic range expression accuracy.

[0048] It should be understood that although this embodiment uses a convolutional neural network as a typical architecture for image reconstruction, in other embodiments, the network can also adopt architectures such as residual networks (ResNet), U-Net, or attention mechanism networks, as long as it can extract multi-exposure features and output spatially related weight maps. The functionality is sufficient, and these variations should all be considered to fall within the network feature fusion defense depth established by this invention.

[0049] In one embodiment, addressing the issues that existing deep learning models' performance is highly dependent on the distribution of training data, prone to reconstruction bias or detail distortion when the acquisition conditions or scene characteristics differ from the training data, and limited by the requirements for the scale of labeled data and computing resources, thus restricting their deployment applicability in industrial environments, this embodiment further refines the constraint mechanism of the image reconstruction network during the training phase. By introducing quantum-level image modeling and an improved structural similarity loss function, it fully elucidates the training mechanism from enhancing model expressiveness to improving optimization objectives. This provides micro-mechanical support at the loss function level for preserving the fine structural features of reconstructed images and selecting the optimal structural source. Simultaneously, it effectively improves the model's generalization ability under different acquisition conditions and scene characteristics, reduces reconstruction bias and detail distortion when the acquisition conditions or scene characteristics differ from the training data, lowers the dependence on specific training data distributions and large amounts of labeled training data and computing resources, and enhances its deployment applicability in industrial environments.

[0050] Regarding the quantum-level image modeling mechanism introduced during the training process of the image reconstruction network, this mechanism performs fine-grained quantization processing of image brightness information. Specifically, it involves dividing the image brightness space into multiple discrete quantization intervals and constructing a quantized brightness representation. ,in The quantization function is used to map continuous brightness to discrete quantum levels. Quantized brightness information is introduced as a constraint into the network optimization, enabling the model to learn the fine structural features of the brightness distribution while reconstructing the image.

[0051] Specifically, in the physical representation of natural images, the brightness response is typically a continuous and infinitely subdivided floating-point value. While this continuity aligns with the laws of physical optics, it can easily lead to the model getting stuck in local smoothing or overgeneralization during feature learning in deep networks, thus ignoring subtle but crucial brightness steps and texture undulations. The quantum-level image modeling mechanism introduced in this embodiment aims to forcibly discretize this continuous brightness space, dividing it into multiple discrete quantization intervals with clearly defined boundaries. This is achieved through a quantization function... The mapping operation, originally a continuous reconstructed image The pixel values ​​are forced to normalize to the nearest discrete quantum level, thus forming a quantized brightness representation. The micro-mechanism by which this discretization mapping constrains the network's learning of fine structures lies in the following: During backpropagation optimization, the quantization operation introduces a step-like gradient cutoff effect, forcing the network not only to fit large-scale brightness gradients but also to generate sufficient gradient jumps at the boundaries of each quantization interval to overcome the information loss caused by quantization errors. This forces the network to accurately capture and enhance those weak edge and texture features that cause brightness to cross quantization intervals during reconstruction, thereby objectively constraining and improving the model's ability to learn and represent fine structural features.

[0052] It should be understood that the granularity of the discrete quantization intervals directly determines the resolution limit for fine-structure modeling. For example, in a conventional configuration, the 0-255 brightness range of an 8-bit image can be divided into 256 discrete quantization intervals, each corresponding to one integer quantum level. This division can effectively constrain the network to learn pixel-level brightness step details. In another configuration for finer modeling of high dynamic range images, the brightness space can be divided into 1024 discrete quantization intervals. At this point, the quantization accuracy reaches the sub-pixel level, which can constrain the network to capture even fainter illumination gradations and ultra-fine textures. Although this embodiment provides specific examples of 256 and 1024 levels, in practical applications, the number of quantization intervals can be flexibly set according to the dynamic range requirements and computational resource limitations of the specific scene, as long as the mechanism of forcibly discretizing continuous brightness to constrain fine-structure learning is satisfied.

[0053] Regarding the improved structural similarity loss function used in the image reconstruction network, the network employs this improved function for constraint optimization. This improved function introduces a window-level maximum structural similarity selection mechanism, which includes: dividing the image into multiple local windows and calculating multiple candidate structural similarity values ​​within each window. Where K represents different exposure information or different feature mapping channels, Represents the window region; the improved structural similarity is defined as: , Based on this improved definition, the overall loss function is constructed as follows: , In the formula, M represents the total number of windows. This loss function enables the model to prioritize the retention of optimal structural information and adaptively select the optimal structural source in different exposure areas.

[0054] Specifically, in the training scenario of multi-exposure image reconstruction, the design of the optimization objective directly determines the direction of model convergence and the upper limit of reconstruction quality. Traditional structural similarity loss functions typically employ an averaging mechanism on the responses of all pixels or different channels within a window when calculating the structural similarity of a local window. This averaging mechanism is applicable in single-exposure or low dynamic range scenarios, but it suffers from serious structural defects in multi-exposure fusion scenarios: due to the inherent asymmetry in signal-to-noise ratio and structural sharpness between the first and second exposure images in bright and dark regions, simply averaging the structural similarity of the two within a local window will dilute and average the high similarity in sharp regions with the low similarity in blurred regions. This leads to incorrect optimization guidance during gradient backpropagation, forcing the model to converge towards local compromise and smoothing blur, ultimately causing the reconstructed image to lose crucial details at the boundary between bright and dark areas and in extreme exposure regions.

[0055] To overcome this deficiency, this embodiment introduces a window-level maximum structural similarity selection mechanism. Its microscopic mechanism lies in: within each local window of the image... Internally, the system no longer compares candidate structure similarity values ​​from different exposure information or different feature mapping channels. Instead of calculating the mean, it takes the maximum value. The maximum selection mechanism forcibly selects the response with the highest structural fidelity within a window as the final structural similarity measure for that region. This maximum selection mechanism establishes a rigorous causal chain: because the maximum value operation naturally shields low-similarity responses, it prioritizes the preservation of optimal structural information during gradient propagation. This allows the model to focus only on further improving the structural fidelity of locally optimal regions during optimization, without being dragged down by inferior structural responses. Furthermore, in multi-exposure scenarios, the optimal structural source for different regions is often different; the optimal structure in a dark area may come from the first exposure channel, while the optimal structure in a bright area may come from the second exposure channel. The maximum selection mechanism can adaptively select the optimal structural source in different exposure regions without requiring manual pre-setting of the correspondence between regions and exposures. Based on this, the overall loss function... By aggregating this maximum structural response across all local windows, a globally oriented optimization objective is constructed, fundamentally avoiding the detail blurring and structural weakening problems caused by traditional mean-based mechanisms. This enables reconstructed images to maintain higher clarity and structural consistency under complex lighting conditions.

[0056] It should be understood that although this embodiment uses the maximum value operation within a local window as the core selection mechanism for explanation, other embodiments may also adopt selection mechanisms based on soft maximum values ​​or other nonlinear aggregation functions. As long as the mechanism of prioritizing the preservation of local optimal structural responses and shielding inferior responses is satisfied, they should all be considered to fall within the maximum selection-avoidance mean fuzziness defense depth established by this invention.

[0057] In one embodiment, after reconstructing the image, this embodiment further introduces a reconstruction quality scoring step to quantitatively evaluate the reliability and accuracy of the reconstruction results, thereby demonstrating a complete closed loop from image reconstruction to quality quantification, and providing a scoring algorithm binding for the reliability and accuracy of the system output.

[0058] Specifically, after generating the reconstructed image through the image reconstruction network, a reconstruction quality scoring step is included to quantitatively evaluate the reliability and accuracy of the reconstruction results. This step specifically includes image structure quality scoring and spectral reconstruction accuracy scoring. The image structure quality scoring is based on an improved structural similarity index, which divides the reconstructed image into multiple local windows. Calculate structural similarity under different exposures or different feature responses within each window. The maximum structural response is selected as the evaluation value for that window to obtain the overall structural quality score. , In the formula, M represents the total number of windows. This scoring formula directly reuses the core logic of the improved SSIM loss function from the aforementioned embodiments at the microscopic level, namely, using a maximum selection mechanism to filter out inferior structural responses and prioritize the retention of locally optimal structural information. This reuse is not a simple logical application, but rather has profound defensive implications: during the training phase, the maximum selection mechanism, as the loss function, guides the model to converge towards preserving the optimal structure; while during the inference and evaluation phase, the same mechanism is transformed into a quality scoring metric, ensuring a strict unity between the training optimization objective and the inference evaluation standard in both mathematical definition and physical meaning, avoiding the evaluation bias problem caused by inconsistencies between training loss and evaluation metrics in traditional schemes. Through this scoring method, the system can adaptively select the source of optimal structural information in different brightness regions, thereby more accurately reflecting the fidelity of the reconstructed image in local structures such as edges and textures.

[0059] Furthermore, the spectral reconstruction accuracy score is obtained by comparing the reconstructed spectral information with the reference spectral data and calculating the spectral error index.

[0060] Specifically, the spectral reconstruction in this invention does not rely on data acquired by a dedicated multispectral camera. Instead, it is based on image reconstruction results, establishing a mapping relationship between brightness and spectral information through model learning to estimate and recover spectral information. To quantitatively evaluate the accuracy of this mapping relationship, the system compares the reconstructed spectral vector with a reference spectral vector obtained through a high-precision spectrometer or a standard reference board. Various specific mathematical metrics can be used to calculate the spectral error index. For example, in one implementation focusing on evaluating spectral shape fidelity, the spectral angle error (SAM) is used as an index. The cosine of the angle between two spectral vectors is calculated to measure the degree of spectral distortion; a smaller SAM value indicates more accurate spectral shape reconstruction. In another implementation focusing on evaluating absolute spectral amplitude deviation, the root mean square error (RMSE) is used as an index. The root mean square error of the reconstructed spectrum and the reference spectrum is directly calculated in each band; a smaller RMSE value indicates higher accuracy in spectral intensity reconstruction. It should be understood that although this embodiment lists SAM and RMSE as typical spectral error indicators, in other embodiments, indicators such as mean absolute error (MAE) or relative spectral error can also be used, as long as they can quantify the degree of deviation between the reconstructed spectrum and the reference spectrum. These variations should all be considered to fall within the scope of the spectral restoration accuracy score.

[0061] After obtaining the two independent scores mentioned above, the structural quality score and the spectral accuracy score are fused to obtain the overall reconstruction quality evaluation result: , In the formula, Indicates the spectral reproduction accuracy score. and is a weighting coefficient used to balance the importance of image structure quality and spectral accuracy.

[0062] Specifically, this fusion scoring formula uses weighting coefficients. and This achieves macroscopic control over visual structure fidelity and physical spectral accuracy. To solidify the parameter defenses, it is essential to... and Specific numerical examples and their physical effects are explained in detail. In typical industrial surface defect detection scenarios, the human eye's sensitivity to texture and edges often takes precedence over the accurate reproduction of absolute color. In such cases, [the following can be done]: Set to 0.6, Setting it to 0.4 makes the overall score more reflective of the image's structural clarity, ensuring the visual identifiability of defects; however, in high-end applications such as material composition analysis or high-precision color quantification, even slight shifts in spectral characteristics can directly lead to errors in material classification. In such cases, the score can be adjusted accordingly. Adjusted to 0.4 The value was adjusted to 0.6, assigning a higher weight to spectral accuracy to constrain the model's convergence accuracy in the spectral dimension. It should be understood that the above combination of 0.6 and 0.4 is only a preferred example; in actual deployment, and The value of can be flexibly set according to the priority of structural fidelity and spectral accuracy requirements in specific application scenarios, as long as the requirements are met. The normalization constraints can be applied.

[0063] More importantly, this embodiment establishes a closed-loop logic of scoring and system stability improvement through the aforementioned reconstruction quality scoring steps. The scoring result is not merely a static output label, but rather a dynamic feedback signal that reintegrates into the system's operation. During model training, the overall reconstruction quality evaluation result Q can serve as a reward signal or an additional regularization term for reinforcement learning, guiding the image reconstruction network to perform targeted optimization on weak local regions or spectral bands in subsequent iterations, thereby continuously improving the model's generalization ability and reconstruction accuracy. In practical industrial applications, for the massive amounts of reconstructed images generated through continuous scanning on an assembly line, the system can adaptively filter results based on the overall reconstruction quality evaluation result Q. For example, a quality threshold can be set, outputting only reliable images with Q values ​​higher than the threshold for subsequent detection algorithm analysis, while triggering re-acquisition or parameter adjustment commands for images with Q values ​​lower than the threshold. This closed-loop feedback mechanism fundamentally avoids the interference of poor reconstruction results on downstream decision-making systems, significantly improving the long-term operational stability and reliability of the entire multi-exposure fusion and reconstruction system in complex industrial environments.

[0064] In a specific embodiment, to verify the operational effectiveness and technical advantages of the method provided by the present invention in a real complex industrial environment, this embodiment applies the quantum-level image reconstruction and dynamic spectral restoration method described in the above embodiments to a metal panel surface defect detection scenario on an industrial production line for detailed explanation.

[0065] In this application scenario, the camera is mounted above the assembly line to perform high-speed scanning imaging of continuously conveyed metal panels. (See reference...) Figure 3 and Figure 4 The metal panel and its surrounding mechanical structure have extremely complex brightness distribution characteristics. Figure 3The image shows a view of localized exposure anomalies. The left side of the image is marked as locally underexposed, containing white cylindrical components and black cables. Traditional single-exposure or multi-frame fusion methods, which have errors in registration, are prone to underexposure in such areas, resulting in severe loss of details in dark areas such as cable entanglement and component edges, making structural information difficult to identify. Meanwhile, the central area of ​​the image is marked as locally overexposed, containing densely distributed bolts and complex wiring. Traditional methods are prone to overexposure and saturation in such bright areas, causing details in bright areas such as bolt texture and wiring direction to be obscured by white halos. Figure 4 The image further shows a partial structural view of the metal panel. The main body of the panel has horizontal textures, with circular metal parts distributed horizontally in the middle. The right side contains an embedded rectangular groove structure. These alternating physical structures can easily cause strong brightness steps and local reflections during imaging.

[0066] To address the shortcomings of traditional methods, the method provided in this invention allows the camera to acquire an original row-scan image containing first and second exposure pixel information arranged in a column-interleaved manner during a single scan. Subsequently, through mask separation, guided filtering reconstruction, adaptive weighted fusion, and even the introduction of an image reconstruction network for deep optimization, the final output reconstructed image is obtained by referring to... Figure 5 and Figure 6 It demonstrates a significantly improved visual effect. Figure 5 The image capture area of ​​the camera is shown as a brightness distribution view. The left area, which was originally below the cable, and the right area around the central square metal piece are both marked as having moderate brightness and uniform distribution. This shows that the method of the present invention effectively suppresses overexposure and saturation in the bright areas and significantly enhances the visibility of the dark areas, thus restoring the structural information that was originally lost due to uneven lighting. Figure 6 The image shows a view of local details of a camera image. The area on the right, enclosed by a red box, contains two side-by-side circular components. The area is clearly marked as having moderate brightness and complete details. This demonstrates that the method of this invention achieves a smooth and natural transition in the area where light and dark meet, avoiding the brightness abrupt changes and fusion artifacts common in traditional fusion, and fully preserving subtle but crucial local textures and edge details.

[0067] To objectively demonstrate the improvements of this invention in various metrics, a quantitative comparative analysis is conducted below, using data from the comparison table in the disclosure document, to compare the method of this invention with existing mainstream multi-exposure fusion methods. Under the same acquisition conditions, multi-exposure images of the same metal panel scene were processed, and the comparison methods included traditional weighted fusion methods, the DeepFuse method, and the MEF-Net method. Evaluation metrics cover image quality and spectral accuracy.

[0068] Regarding image quality metrics, referring to the data comparison in Table 1: the peak signal-to-noise ratio (PSNR) of traditional fusion methods is only 18 to 22 dB, the structural similarity index (SSIM) is 0.70 to 0.82, the multi-exposure fusion structural similarity (MEF-SSIM) is 0.85 to 0.90, and the visual information fidelity (VIF) is 0.45 to 0.60, showing low overall performance and poor stability under complex lighting conditions; the DeepFuse method shows improvement under the deep learning framework, with PSNR reaching 20 to 24 dB, SSIM reaching 0.80 to 0.88, MEF-SSIM reaching 0.88 to 0.93, and VIF reaching 0.55 to 0.70, but detail loss still exists in extreme exposure areas; the MEF-Net method further improves through a weight prediction mechanism, with PSNR reaching 22 to 26. The PSNR ranges from 24 to 28 dB, SSIM from 0.90 to 0.96, MEF-SSIM from 0.93 to 0.98, and VIF from 0.65 to 0.80. The method of this invention achieves objective improvements in all these metrics: PSNR reaches 24 to 28 dB, SSIM from 0.90 to 0.96, MEF-SSIM from 0.93 to 0.98, and VIF from 0.75 to 0.90. These data objectively demonstrate that this invention, through guided filtering for structure-fidelity reconstruction, nonlinear adaptive fusion based on brightness differences, and even deep mining of network feature maps, achieves quantifiable gains in suppressing overexposure, enhancing dark areas, and smoothing transitions.

[0069] Table 1

[0071] Table 2

[0072] Regarding spectral accuracy metrics, referring to the data comparison in Table 2: the spectral angle error (SAM) of traditional fusion methods ranges from 0.08 to 0.12, and the root mean square error (RMSE) ranges from 0.05 to 0.08; the SAM of the DeepFuse method ranges from 0.06 to 0.10, and the RMSE ranges from 0.04 to 0.07; the SAM of the MEF-Net method ranges from 0.05 to 0.08, and the RMSE ranges from 0.03 to 0.06; the SAM of the method of this invention is reduced to 0.03 to 0.06, and the RMSE is reduced to 0.02 to 0.05. This improvement in spectral restoration accuracy directly stems from the quantum-level image modeling mechanism introduced in this invention during network training, which constrains the learning of fine structural features of brightness distribution, and the micro-mechanism of the improved SSIM loss function, which prioritizes the retention of optimal structural information and adaptively selects the optimal structural source through a window-level maximum structural similarity selection mechanism. Thus, while preserving the brightness fidelity of the reconstructed image, the mapping relationship established through model learning achieves effective restoration and high-precision recovery of spectral information.

[0073] It should be understood that although this embodiment uses the detection of surface defects on metal panels as a typical industrial scenario for data verification and mechanism explanation, the single-scan interleaved acquisition, guided filter structure fidelity reconstruction, adaptive nonlinear smooth fusion, and data-driven network depth reconstruction and scoring closed-loop mechanism established by this invention are also applicable to other industrial inspection scenarios with high dynamic range imaging requirements and complex lighting challenges, such as color difference detection on printed surfaces and microstructure observation of semiconductor wafers. These variant applications should all be considered to fall within the protection scope of this invention.

[0074] Figure 7 This example illustrates a schematic diagram of the physical structure of an electronic device, which can be a smart terminal. Its internal structure diagram can be as follows: Figure 7 As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the quantum-level image reconstruction and dynamic spectral restoration method of any of the above embodiments.

[0075] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the electronic device to which the present invention is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0076] On the other hand, the present invention also provides a computer storage medium storing a computer program, which, when executed by a processor, implements the quantum-level image reconstruction and dynamic spectral restoration method of any of the above embodiments.

[0077] In another aspect, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and when the processor executes the computer instructions, it implements the quantum-level image reconstruction and dynamic spectral restoration method of any of the above embodiments.

[0078] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory.

[0079] By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0080] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0081] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A quantum-level image reconstruction and dynamic spectral restoration method, characterized in that, include: A raw line scan image is acquired by the camera during a single scan. Pixels at different column positions in the raw line scan image correspond to different exposure parameters, so that the raw line scan image contains pixel information under various exposure parameters arranged in an interleaved column manner. Based on a preset pixel arrangement rule, the original line scan image is separated into multiple sparse exposure images, including a first sparse exposure image and a second sparse exposure image. The exposure parameters of the first sparse exposure image are higher than those of the second sparse exposure image. The pixel positions corresponding to low exposure pixel information in the first sparse exposure image are empty pixels, and the pixel positions corresponding to high exposure pixel information in the second sparse exposure image are empty pixels. Using the original line scan image or the corresponding preliminary interpolated image as the guide image, guide filtering operations are performed on the first sparse exposure image and the second sparse exposure image respectively to fill in the missing pixels in the first sparse exposure image and the second sparse exposure image, and correspondingly generate a spatially aligned first exposure reconstructed image and a second exposure reconstructed image with the same resolution. The first exposure-reconstructed image and the second exposure-reconstructed image are subjected to weighted fusion processing to generate a fused image, wherein the weights participating in the weighted fusion are adaptively determined based on the brightness information at the pixel position.

2. The quantum-level image reconstruction and dynamic spectral restoration method according to claim 1, characterized in that, The separation of the original line scan image into a first exposure sparse image and a second exposure sparse image based on a preset pixel arrangement rule specifically includes: The original line scan image is pixel-filtered using an exposure mask, wherein the first sparse exposure image is obtained by extracting using a first exposure mask, and the second sparse exposure image is obtained by extracting using a second exposure mask. The separation process is represented as follows: , , In the formula Indicates the second exposure sparse image in the row index Column index Pixel value at that location, Indicates the row index of the sparse image in the first exposure. Column index Pixel value at that location, This indicates that the original row scan image is at the row index. Column index Pixel value at that location, These are binary masks corresponding to the second exposure column position and the first exposure column position, respectively.

3. The quantum-level image reconstruction and dynamic spectral restoration method according to claim 1, characterized in that, The step of using the original line scan image or the corresponding preliminary interpolated image as the guide image to perform guided filtering operations on the first sparse exposure image and the second sparse exposure image respectively includes: The original line scan image Alternatively, the preliminary interpolation result of the corresponding exposed image can be used as the guide image G, with the second exposed sparse image. Or the first exposure sparse image As input image P, guided filtering reconstruction is performed, so that when filling in missing pixels, the edge information and texture information in the guided image G are used to constrain the estimated direction and magnitude of the missing pixels, so as to maintain the structural consistency between the first exposure reconstructed image and the second exposure reconstructed image in the edge region.

4. The quantum-level image reconstruction and dynamic spectral restoration method according to claim 1, characterized in that, The step of weighted fusion processing of the first exposure reconstructed image and the second exposure reconstructed image to generate a fused image includes: Weighted fusion is performed according to the following formula: , In the formula, This represents the pixel value at position X in the merged image. These are the pixel values ​​at position X of the second exposure-reconstructed image and the first exposure-reconstructed image, respectively. For adaptive weights.

5. The quantum-level image reconstruction and dynamic spectral restoration method according to claim 4, characterized in that, The weights participating in the weighted fusion are adaptively determined based on the brightness information at the pixel location, wherein the adaptive weights are determined using a non-linear mapping form based on the brightness difference: , In the formula, K is the adjustment parameter.

6. The quantum-level image reconstruction and dynamic spectral restoration method according to any one of claims 1 to 5, characterized in that, After generating the fused image, the process also includes: The reconstructed image from the first exposure after interpolation. Second exposure reconstructed image and the fused image As input data, construct a multi-channel input tensor ; The input tensor is fed into an image reconstruction network for feature extraction and fusion. The image reconstruction network extracts brightness features, edge information and texture structure in images with different exposures through multi-layer convolution operations, and establishes a mapping relationship between different exposures. The image reconstruction network outputs a set of spatially correlated weight maps. This weight is used to represent the contribution of different exposure information to each pixel location, and a reconstructed image is generated based on this weight mapping: , In the formula, This represents the pixel value of the reconstructed image at position X. This represents the spatial correlation weight mapping at location X. These are the pixel values ​​at position X in the second exposure reconstructed image and the first exposure reconstructed image, respectively.

7. The quantum-level image reconstruction and dynamic spectral restoration method according to claim 6, characterized in that, During the training process, the image reconstruction network introduces a quantum-level image modeling mechanism to perform fine-grained quantization of image brightness information, specifically including: The image brightness space is divided into multiple discrete quantization intervals to construct a quantized brightness representation. ,in This represents a quantization function used to map continuous brightness to discrete quantum levels; By incorporating quantized brightness information as a constraint into the network optimization, the model learns the fine structural features of the brightness distribution while reconstructing the image.

8. The quantum-level image reconstruction and dynamic spectral restoration method according to claim 7, characterized in that, The image reconstruction network employs an improved structural similarity loss function for constraint optimization. This improved structural similarity loss function introduces a window-level maximum structural similarity selection mechanism, including: The image is divided into multiple local windows, and multiple candidate structure similarity values ​​are calculated within each window. Where K represents different exposure information or different feature mapping channels, Indicates a window area; The improved structural similarity is defined as follows: , Based on this improved definition, the overall loss function is constructed as follows: , In the formula, M represents the total number of windows. This loss function enables the model to prioritize the retention of optimal structural information and adaptively select the optimal structural source in different exposure areas.

9. The quantum-level image reconstruction and dynamic spectral restoration method according to claim 8, characterized in that, After generating the reconstructed image through the image reconstruction network, a reconstruction quality scoring step is included to quantitatively evaluate the reliability and accuracy of the reconstruction results. This specifically includes image structure quality scoring and spectral reconstruction accuracy scoring, wherein: The image structure quality score, based on the improved structural similarity index, divides the reconstructed image into multiple local windows. Calculate structural similarity under different exposures or different feature responses within each window. The maximum structural response is selected as the evaluation value for that window to obtain the overall structural quality score. , The spectral reconstruction accuracy score is obtained by comparing the reconstructed spectral information with the reference spectral data and calculating the spectral error index. The structural quality score and the spectral accuracy score are combined to obtain the overall reconstruction quality evaluation result: , In the formula, Indicates the spectral reproduction accuracy score. and These are the weighting coefficients.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the quantum-level image reconstruction and dynamic spectral restoration method as described in any one of claims 1 to 9.