Optical image noise reduction method and system for composite imaging assembly
By combining multi-sensor time synchronization, multi-spectral adaptive masking, and dual-blind-spot self-supervised learning with Wiener filters, the problems of multi-sensor noise coupling and strong pixel correlation in composite imaging components are solved, and the imaging quality and target recognition accuracy are improved. It is suitable for autonomous driving, industrial inspection, and medical imaging.
Patent Information
- Application Number
- CN202511031133.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-25
AI Technical Summary
Existing optical image noise reduction technologies have difficulty effectively handling noise coupling and strong pixel correlation between multiple sensors when processing composite imaging components. This causes noise to influence and propagate between sensors, affecting imaging quality and target recognition accuracy, and it is difficult to achieve a balance between real-time performance and edge protection.
A noise coupling matrix is established through a multi-sensor time synchronization mechanism, and a multispectral adaptive mask generation algorithm and a conditional mask convolution block are used for cross-band feature interaction processing. A double-blind spot self-supervised learning algorithm and a block-based Wiener filter are combined for edge detail protection to achieve cross-sensor noise coordination and effective noise reduction.
It improves the imaging quality and target recognition accuracy of composite imaging components in complex environments, ensures image clarity and the integrity of target features, and is suitable for fields such as autonomous driving, industrial inspection, and medical imaging.
Smart Images

Figure CN120525759B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an optical image noise reduction method and system for a composite imaging component. Background Art
[0002] Composite imaging components, as a core technology in modern optoelectronic devices, are widely used in precision optical instruments such as autonomous driving systems, industrial inspection equipment, and medical imaging devices. By integrating multiple heterogeneous sensors, such as visible light sensors, infrared sensors, and laser ranging sensors, they can acquire multispectral and multimodal image information in complex environments. Existing optical image denoising technologies are primarily designed for single sensors, including wavelet-based denoising methods, adaptive filtering techniques, and the recently developed deep learning denoising networks. These methods have demonstrated excellent results in addressing image noise from single sensors, effectively suppressing common noise types such as photon shot noise, thermal noise, and readout noise. Furthermore, advanced denoising algorithms, such as blind spot networks and self-supervised learning methods, achieve superior denoising performance without the need for a clean reference image by hiding noisy image pixels and restoring them using surrounding contextual information.
[0003] However, existing technologies for optical image noise reduction in composite imaging assemblies have significant shortcomings, primarily in effectively addressing noise coupling between multiple sensors. When multiple sensors operate simultaneously, physical coupling phenomena such as electromagnetic interference, thermal coupling, and vibration transmission occur between the different sensors, causing noise to influence and propagate between the sensors. Traditional single-sensor noise reduction methods are unable to establish and utilize this cross-sensor noise correlation. Furthermore, strong pixel correlation exists between images of different wavelengths in composite imaging assemblies. For example, target outlines in visible light images correspond highly to thermal radiation distributions in infrared images and range profiles in lidar. Existing noise reduction methods lack specialized processing mechanisms to address this strong pixel correlation, which can easily destroy cross-band feature correspondences during the noise reduction process, affecting subsequent target recognition and ranging accuracy.
[0004] A deeper technical problem lies in the fact that the application scenarios of composite imaging components place strict demands on real-time performance and edge detail protection. Applications such as industrial automation equipment and intelligent monitoring systems require both rapid response and the integrity of target features. However, existing technologies have difficulty achieving a balance between noise reduction effect, processing speed, and edge protection. Due to the lack of a unified multimodal noise reduction framework, existing methods usually adopt a serial processing approach, first independently reducing the noise of each sensor and then fusing the data. This processing flow not only increases computational complexity and time delay, but also easily loses coordination information between sensors during the independent noise reduction process, resulting in spatial registration deviations, inconsistent features, and other problems in the fused image. Especially in areas of strongly correlated pixels, traditional masking strategies cannot effectively break the correlation between pixels, making it easy for the network to learn noise patterns rather than the true image structure, thereby affecting the imaging quality and reliability of the entire composite imaging component. Summary of the Invention
[0005] The present application provides an optical image noise reduction method and system for a composite imaging component, which is used to solve the noise reduction failure problem caused by multi-sensor noise coupling and strong pixel correlation in the composite imaging component, and improve the imaging quality and target recognition accuracy of the composite imaging component in complex environments.
[0006] In a first aspect, the present application provides an optical image denoising method for a composite imaging component, the optical image denoising method for the composite imaging component comprising: performing geometric correction processing on the original image data of the visible light sensor, infrared sensor and laser ranging sensor in the composite imaging component through a multi-sensor time synchronization mechanism to obtain a multi-sensor noise coupling matrix; performing mask strategy adjustment processing on images of different bands through a multi-spectral adaptive mask generation algorithm according to the multi-sensor noise coupling matrix to obtain a ring mask pattern and a checkerboard mask pattern; inputting the ring mask pattern and the checkerboard mask pattern into a conditional mask convolution block for cross-band feature interaction processing to obtain a multi-scale densely connected feature map; performing weight adaptive adjustment processing on the cross-band coordination loss through a double-blind spot self-supervised learning algorithm according to the multi-scale densely connected feature map to obtain a denoising prediction result; inputting the denoising prediction result into a block-based Wiener filter for edge detail protection processing to obtain a final denoised image of the composite imaging component.
[0007] In a second aspect, the present application provides an optical image noise reduction system for a composite imaging assembly, the optical image noise reduction system for the composite imaging assembly comprising:
[0008] A correction module is used to perform geometric correction processing on the raw image data of the visible light sensor, infrared sensor and laser ranging sensor in the composite imaging component through a multi-sensor time synchronization mechanism to obtain a multi-sensor noise coupling matrix;
[0009] A generation module is used to adjust the mask strategy of images of different bands by using a multispectral adaptive mask generation algorithm according to the multi-sensor noise coupling matrix to obtain a ring mask pattern and a checkerboard mask pattern;
[0010] An interaction module, configured to input the annular mask pattern and the checkerboard mask pattern into a conditional mask convolution block for cross-band feature interaction processing to obtain a multi-scale densely connected feature map;
[0011] An adjustment module is used to perform weight adaptive adjustment processing on the cross-band coordination loss according to the multi-scale densely connected feature map through a double-blind spot self-supervised learning algorithm to obtain a noise reduction prediction result;
[0012] The protection module is used to input the noise reduction prediction result into a block-based Wiener filter to perform edge detail protection processing to obtain a final noise reduction image of the composite imaging component.
[0013] In a third aspect, an optical image noise reduction device of a composite imaging component is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the optical image noise reduction device of the composite imaging component executes the above-mentioned optical image noise reduction method of the composite imaging component.
[0014] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium, which, when executed on a computer, enables the computer to execute the above-mentioned optical image noise reduction method of the composite imaging component.
[0015] In the technical solution provided by this application, the raw image data of the visible light sensor, infrared sensor and laser ranging sensor in the composite imaging component are geometrically corrected through a multi-sensor time synchronization mechanism, and a unified multi-sensor noise coupling matrix is established, which effectively solves the technical problem that the traditional single-sensor noise reduction method cannot handle the cross-sensor noise correlation, so that the noise impact relationship between different sensors can be accurately modeled and quantitatively analyzed. The multi-spectral adaptive mask generation algorithm dynamically adjusts the mask strategy according to the noise characteristics of different bands. The generated annular mask pattern and checkerboard mask pattern can effectively break the strong correlation between pixels in the composite imaging component, avoid the problem of related pixel information leakage caused by the traditional fixed mask strategy, and ensure that the network learns the real image structure rather than the noise pattern. The conditional mask convolution block realizes the deep fusion of multi-sensor data through cross-band feature interaction processing. The generated multi-scale densely connected feature map not only retains the unique information of each band but also establishes the intrinsic connection between bands, overcoming the information loss and computational redundancy caused by the traditional serial processing method. The dual-blind-spot self-supervised learning algorithm adaptively adjusts the weights of cross-band coordination losses. This dual-constraint mechanism ensures that the network learns stable and reliable feature representations without the need for extensive labeled data. This is particularly useful in composite imaging applications where clean reference images are difficult to obtain. A block-based Wiener filter combined with edge detail preservation effectively suppresses noise while accurately protecting important structural edge information. This addresses the technical challenge of traditional noise reduction methods in balancing noise reduction effectiveness and edge preservation, ensuring that composite imaging components can achieve both clear images and maintain the integrity of target features in applications such as industrial inspection and intelligent monitoring.
[0016] The multi-sensor time synchronization mechanism ensures temporal consistency between sensor data, which is crucial for autonomous driving systems and industrial visual inspection equipment that require multimodal information fusion. This prevents target positioning errors and recognition failures caused by time deviation. The introduction of an adaptive mask generation algorithm enables the composite imaging component to intelligently process noise based on the actual noise distribution characteristics, improving its adaptability to complex noise environments compared to traditional methods. This is particularly evident in industrial environments with drastic lighting changes or large temperature fluctuations. The dual-blind-spot self-supervised learning mechanism reduces its reliance on training data, enabling the algorithm to quickly adapt to different application scenarios and equipment configurations, thus paving the way for the industrial application of composite imaging components. A post-processing strategy combining Wiener filtering with edge protection ensures that the noise reduction results meet image quality requirements while preserving key feature information. This has important application value in medical imaging, precision manufacturing, and other fields that require precise target identification and measurement, effectively enhancing the practicality and reliability of composite imaging components in these specialized fields with extremely high requirements for image quality and feature fidelity. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A schematic diagram of an embodiment of a method for reducing optical image noise of a composite imaging assembly in an embodiment of the present application;
[0019] Figure 2 A schematic diagram of an embodiment of an optical image noise reduction system of a composite imaging assembly in an embodiment of the present application;
[0020] Figure 3 1 is a schematic block diagram of the structure of an optical image noise reduction device of a composite imaging assembly in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The embodiments of the present application provide a method and system for optical image noise reduction of a composite imaging assembly. The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or are inherent to these processes, methods, products or devices.
[0022] For ease of understanding, the specific process of the embodiment of the present application is described below. Figure 1 In one embodiment of the present application, a method for reducing optical image noise of a composite imaging assembly includes:
[0023] Step S101: geometrically correcting the raw image data of the visible light sensor, infrared sensor, and laser ranging sensor in the composite imaging component through a multi-sensor time synchronization mechanism to obtain a multi-sensor noise coupling matrix;
[0024] Step S102: performing mask strategy adjustment processing on images of different bands using a multispectral adaptive mask generation algorithm according to the multi-sensor noise coupling matrix to obtain a ring mask pattern and a checkerboard mask pattern;
[0025] Step S103: input the annular mask pattern and the checkerboard mask pattern into the conditional mask convolution block for cross-band feature interaction processing to obtain a multi-scale densely connected feature map;
[0026] Step S104: Adaptively adjust the weights of the cross-band coordination loss using a double-blind-spot self-supervised learning algorithm based on the multi-scale densely connected feature graph to obtain a noise reduction prediction result;
[0027] Step S105 : Input the noise reduction prediction result into a block-based Wiener filter for edge detail protection processing to obtain a final noise reduction image of the composite imaging component.
[0028] It is understandable that the execution subject of the present application can be the optical image noise reduction system of the composite imaging component, or can be a terminal or a server, which is not limited here. The embodiment of the present application is described by taking the server as the execution subject as an example.
[0029] Specifically, the multi-sensor time synchronization mechanism first timestamps the raw image data from the visible light sensor, infrared sensor, and laser ranging sensor. The image data collected by each sensor is aligned along the time axis to eliminate time deviations caused by differences in sensor response times. Synchronization accuracy is controlled within 10 microseconds to ensure data temporal consistency. Next, the synchronized image sequences are geometrically registered using an affine transformation algorithm to eliminate spatial offsets caused by differences in sensor installation positions, with registration errors controlled within 0.5 pixels. The noise characteristics of each sensor are then extracted: photon shot noise and readout noise parameters are primarily extracted for the visible light sensor, thermal noise and dark current noise parameters for the infrared sensor, and speckle noise and coherence noise parameters for the laser ranging sensor. Finally, a 3×3 inter-sensor noise correlation matrix is constructed, with the diagonal elements representing the self-noise intensity of each sensor and the off-diagonal elements representing the noise coupling intensity between sensors. This matrix quantifies the noise interactions among the multiple sensors and forms the multi-sensor noise coupling matrix.
[0030] The multispectral adaptive mask generation algorithm, based on the multisensor noise coupling matrix obtained in step S101, processes image features for the visible, infrared, and laser bands separately through a band feature encoding module. Each band uses a separate convolutional neural network branch for feature extraction, with the number of channels set to 64, 128, and 256, respectively. The dilated convolutional architecture captures spatial correlations at different scales by setting dilation rates of 2, 4, and 8, corresponding to receptive field radii of 5, 9, and 17 pixels, respectively. This generates a multiscale spatial correlation weight matrix. A cross-band attention mechanism calculates feature similarities at corresponding pixel positions across different bands, generating a cross-band correlation matrix ranging from 0 to 1. Correlation strengths exceeding 0.75 are labeled as strongly correlated, between 0.3 and 0.75 as moderately correlated, and below 0.3 as weakly correlated. Based on pixel distribution, strongly correlated regions are masked using a circular mask pattern with an inner diameter of 3 pixels and an outer diameter of 7 pixels. Moderately correlated regions are masked using a checkerboard mask pattern with a 2×2 mask block size. Weakly correlated regions are masked using a random mask pattern with a masking rate of 15%.
[0031] The conditional mask convolution block receives a ring mask pattern and a checkerboard mask pattern from the mask condition judgment unit, performs a logical AND operation to verify the validity of each pixel position, and generates a mask condition control signal. The selective convolution calculation unit dynamically adjusts the convolution kernel weights based on the mask condition control signal, setting the weights corresponding to masked pixels to 0 to ensure that masked pixels do not participate in feature calculation. The densely connected network consists of four scale layers, processing features at scales of 1 / 1, 1 / 2, 1 / 4, and 1 / 8, respectively. Each scale layer contains six convolution blocks, and the outputs of the first five convolution blocks are passed to the last convolution block via skip connections. The 3×3×3 cubic convolution structure simultaneously integrates features in both spatial and band dimensions. The first two dimensions process spatial features, while the third dimension handles inter-band feature interactions. Residual connections ensure that the original information is not lost during feature interactions. The attention gating mechanism assigns weights based on feature importance, with important features set to 1 and less important features adjusted between 0.2 and 0.8.
[0032] The double-blind spot self-supervised learning algorithm applies two different masking strategies to the multi-scale densely connected feature map: the first uses adaptive masking, and the second uses a 20% random masking rate, generating two different input versions. The double-blind spot training network simultaneously predicts the values of the masked pixels in both versions. This dual constraint ensures that the network learns the true image structure rather than noise patterns. The composite loss function consists of four components: reconstruction loss, consistency loss, edge preservation loss, and cross-band coordination loss. The reconstruction loss uses the L1 loss to ensure pixel-level reconstruction accuracy. The consistency loss calculates the difference between double-blind spot predictions using the L2 norm. The edge preservation loss uses the Sobel operator to extract edge information and prevent structural blurring. The cross-band coordination loss calculates the consistency of features at the same spatial location across different bands. Gradient directional cosine similarity is used as the basis for assigning loss weights. Pixel locations with similarity greater than 0.9 are assigned a loss weight of 0.1, those between 0.5 and 0.9 are assigned a weight of 1.0, and those with similarity less than 0.5 are assigned a weight of 2.0. Initially, the reconstruction loss weight is 1.0, and all others are 0.2. This weight is gradually adjusted to 0.8 and 1.0 in later training stages.
[0033] The block-based Wiener filter first evaluates the quality of the denoised prediction results. By calculating the signal-to-noise ratio (SNR), edge preservation index, and structural similarity index, image regions are classified into three categories: high-quality, medium-quality, and low-quality. SNRs greater than 25dB are labeled high-quality, those between 15 and 25dB are labeled medium-quality, and those less than 15dB are labeled low-quality. The image is segmented into 8×8 non-overlapping blocks, and the noise power spectral density and signal power spectral density parameters are independently calculated for each block. Blocks containing edge texture information have their noise power spectral density multiplied by a factor of 0.3 to prevent over-filtering. Blocks in smooth regions retain their original estimated values. The multi-scale Canny operator sets detection thresholds of 0.1, 0.15, and 0.2 to obtain edge information of different intensities. Edges are divided into three categories according to the gradient amplitude: strong edges, medium edges, and weak edges. Strong edges completely retain the original pixel values. Medium edges use a weighted fusion of 0.7 times the filtering result and 0.3 times the original value. Weak edges use a fusion of 0.9 times the filtering result and 0.1 times the original value. When edges are detected in multiple bands at the same spatial location, the protection strength is increased by 20% to ensure that important structural information is fully preserved.
[0034] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0035] The raw image data of the visible light sensor, infrared sensor, and laser ranging sensor are time-synchronized and calibrated based on the timestamp mark, and a multi-sensor synchronized image sequence with a synchronization accuracy of 10 microseconds is obtained.
[0036] The multi-sensor synchronized image sequence is input into the affine transformation algorithm for geometric registration processing, and a spatially aligned image group with a registration error of less than 0.5 pixels is obtained;
[0037] Extracting photon shot noise and readout noise parameters of the visible light sensor in the spatially aligned image group to obtain a visible light noise feature vector;
[0038] The thermal noise, dark current noise and speckle coherence noise parameters of the infrared sensor and laser ranging sensor in the spatially aligned image group are extracted and processed respectively to obtain the infrared noise feature vector and the laser noise feature vector.
[0039] Based on the visible light noise eigenvector, infrared noise eigenvector and laser noise eigenvector, a 3×3 inter-sensor noise correlation calculation process is constructed to obtain the multi-sensor noise coupling matrix.
[0040] Specifically, timestamp information is extracted from the raw image data acquired by each sensor. The timestamp records the precise acquisition moment of each frame, identifying the temporal attributes of the image data with nanosecond precision. The timestamp of the visible light sensor includes the start and end times of exposure, the timestamp of the infrared sensor records the sampling moment of the thermal imaging data, and the timestamp of the laser ranging sensor records the laser pulse emission and echo reception moments. The synchronous calibration algorithm calculates the time difference between the timestamps of each sensor to identify sensor response delays and acquisition time offsets. It then rearranges the image sequence time axis to control the time difference to within 10 microseconds, forming a time-synchronized multi-sensor image sequence in which the visible light image, infrared image, and laser ranging image corresponding to each moment are fully aligned in time. The affine transformation algorithm accepts the multi-sensor synchronized image sequence as input and establishes spatial correspondences between the images of different sensors through feature point matching. The algorithm first detects corner and edge feature points in each sensor's image. It then uses descriptor matching to find corresponding point pairs between different sensor images. Based on these corresponding point pairs, it calculates the affine transformation matrix, which contains four geometric transformation parameters: rotation, translation, scaling, and shearing. The geometric registration process applies the affine transformation matrix to each pixel coordinate in the source image and calculates the new coordinate position of that pixel in the target image. Registration accuracy is measured by calculating the Euclidean distance between corresponding feature points after the transformation. Registration is considered successful when the distance is less than 0.5 pixels. The result is a set of images that are completely aligned in spatial position, where each pixel position corresponds to the same physical point in the different sensor images.
[0041] The photon shot noise and readout noise parameter extraction process analyzes the noise characteristics of visible light sensor images. Photon shot noise is Poisson-distributed noise caused by the randomness of photons arriving at the sensor, and its variance is equal to the signal intensity. Readout noise is additive Gaussian noise introduced by the sensor circuit during signal reading and is independent of signal intensity. The parameter extraction algorithm analyzes the statistical characteristics of pixel values in uniform areas of the image, calculates the mean and variance of the pixel values, and separates the photon shot noise and readout noise components based on a noise model. The photon shot noise intensity is estimated by the ratio of the pixel value variance to the mean, while the readout noise intensity is estimated by the variance of dark pixel regions. These noise parameters are combined to form a visible light noise feature vector, which contains numerical elements such as the photon shot noise variance, the readout noise variance, and the signal-dependent noise coefficient.
[0042] Thermal noise, dark current noise, and speckle coherence noise parameter extraction perform specialized noise analysis for infrared sensors and laser ranging sensors, respectively. Thermal noise is the random noise generated by temperature fluctuations in infrared sensors, dark current noise is the current noise generated by the sensor in the absence of light, speckle noise is the interference pattern noise produced by the interaction of laser coherent light with rough surfaces, and coherence noise is the noise generated by optical path interference in laser systems. Infrared sensor noise extraction analyzes the statistical characteristics of pixels in temperature-stable regions, separates thermal noise and dark current noise components, and calculates the noise power spectral density and correlation function to form an infrared noise feature vector. Laser ranging sensor noise extraction analyzes the repeatability and consistency of distance measurements, calculates the contrast and coherence length of speckle noise, and analyzes the spectral characteristics and spatial distribution of coherence noise to form a laser noise feature vector. This vector contains key parameters such as speckle contrast, coherence length, and noise power density. The three-by-three inter-sensor noise correlation calculation model is constructed based on the visible light noise eigenvector, infrared noise eigenvector, and laser noise eigenvector. The calculation process first normalizes the three noise eigenvectors to eliminate dimensionality differences. Then, the cross-correlation coefficients between the vectors are calculated. The cross-correlation coefficients reflect the degree of linear correlation between the noises of different sensors. During matrix construction, the diagonal elements are set to the self-noise intensity of each sensor, calculated by the modulus of the noise eigenvectors. The off-diagonal elements are set to the noise coupling strength between sensors, determined by the absolute value of the cross-correlation coefficients. The calculation of the noise coupling strength takes into account the physical coupling mechanisms between sensors, including electromagnetic interference, thermal coupling, and vibration transmission. When the noise eigenvectors of two sensors have similar statistical characteristics, the corresponding coupling strength is high, while when the noise characteristics differ significantly, the coupling strength is low. Ultimately, a three-by-three noise coupling matrix is formed to describe the mutual influence of multi-sensor noise.
[0043] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0044] The multi-sensor noise coupling matrix is input into the band feature encoding module to extract independent features of visible light, infrared and laser bands, and a three-branch convolution feature group containing 64, 128 and 256 channels is obtained;
[0045] Based on the three-branch convolutional feature group, the spatial attention weight is calculated through the dilated convolution structure with expansion rates of 2, 4, and 8, and the multi-scale spatial correlation weight matrix with receptive field coverage radius of 5, 9, and 17 pixels is obtained;
[0046] The cross-band attention similarity calculation is performed on the corresponding pixel positions of different bands according to the multi-scale spatial correlation weight matrix to obtain a cross-band correlation matrix with a value range of 0 to 1;
[0047] Based on the correlation strength thresholds of 0.75 and 0.3 in the cross-band correlation matrix, the mask pattern classification and judgment processing of the pixel area were performed to obtain the pixel distribution labels of strong correlation area, medium correlation area and weak correlation area;
[0048] According to the pixel distribution mark, a ring mask pattern with an inner diameter of 3 pixels and an outer diameter of 7 pixels is set for the strong correlation area, a checkerboard mask pattern with a mask block size of 2×2 is set for the medium correlation area, and a random mask processing with a mask rate of 15% is set for the weak correlation area, thus obtaining a ring mask pattern and a checkerboard mask pattern.
[0049] Specifically, the band feature encoding module receives a multi-sensor noise coupling matrix as input. The inter-sensor noise correlation information contained in the matrix guides the weight allocation for feature extraction. The module internally comprises three parallel convolutional neural network branches, each dedicated to processing image features in the visible light band, infrared band, and laser band. The visible light branch uses 64 convolution kernels for feature extraction, each with a size of 3×3. Convolution operations are performed on the image through a sliding window to extract visual features such as edges, texture, and color. The infrared branch uses 128 convolution kernels to extract thermal imaging features such as thermal radiation intensity, temperature gradient, and hot spot distribution. The laser branch uses 256 convolution kernels to extract lidar features such as range information, reflection intensity, and surface roughness. The increasing number of convolution kernels in the three branches reflects the varying complexity of the data in different bands. Laser ranging data contains two dimensions of information, distance and intensity, and therefore requires more feature channels. Infrared data contains temperature information and requires a moderate number of feature channels. The relatively simple visible light data requires fewer feature channels. The independent processing of the three branches ensures that the features of each band do not interfere with each other while maintaining the band-specificity of the features. The dilated convolution architecture calculates spatial attention weights based on a three-branch convolutional feature set. Dilated convolution expands the receptive field without increasing the number of parameters by inserting zero values into the standard convolution kernel. A convolution kernel with a dilation rate of 2 inserts gaps in the original 3×3 kernel, forming a convolution pattern that effectively covers a 5×5 area. A convolution kernel with a dilation rate of 4 covers a 9×9 area, and a convolution kernel with a dilation rate of 8 covers a 17×17 area. Spatial attention weights are determined by analyzing the influence of pixels at different locations on the current pixel. The weights reflect the strength of spatial correlation between pixels. The calculation process first performs spatial pooling on the three-branch features, then calculates the attention score through a fully connected layer, and finally normalizes the weights using a softmax function. Multi-scale processing simultaneously captures both short-range and long-range spatial correlations by applying convolution kernels with three different dilation rates in parallel. Short-range correlations primarily reflect local structural information, while long-range correlations reflect global contextual information. The rows and columns of the weight matrix correspond to the row and column coordinates of the image, respectively, and the matrix element value represents the attention weight at the corresponding position.
[0050] Cross-band attention similarity is calculated based on a multi-scale spatial correlation weight matrix. Similarity is calculated by comparing eigenvectors from different bands at the same spatial location. The similarity is calculated using the cosine similarity method, where the dot product of two eigenvectors is divided by the product of their respective moduli to obtain the similarity value. The calculation process first normalizes the eigenvectors of the visible, infrared, and laser bands to eliminate dimensionality differences. Then, the similarity between the bands is calculated pixel by pixel. The similarity between the visible and infrared bands reflects the correspondence between visible light features and thermal features, the similarity between the visible and laser bands reflects the correspondence between visual features and distance features, and the similarity between the infrared and laser bands reflects the correspondence between thermal features and distance features. Similarity values range from 0 to 1, with 0 indicating complete indifference and 1 indicating complete correlation. The cross-band correlation matrix is constructed by combining the similarity calculation results of the three pairs of bands. The matrix size is the same as the image size, and each pixel position is assigned a correlation value. The correlation value is obtained by weighted averaging the similarities of the three pairs of bands, with the weight determined based on the importance of each band in the current application scenario.
[0051] The mask pattern classification process classifies pixel regions based on correlation strength thresholds in the cross-band correlation matrix. A threshold of 0.75 is used as the strong correlation criterion. When the cross-band correlation value of a pixel location is greater than 0.75, it indicates strong feature correspondence between different bands. Such regions often correspond to important target structures or significant features, and a more sophisticated masking strategy is required to prevent information loss. A threshold of 0.3 is used as the weak correlation criterion. When the correlation value is less than 0.3, it indicates a lack of clear cross-band correspondence, indicating a noise-dominated region or unimportant background area. A random masking strategy will not affect the extraction of important information. Moderate correlation regions correspond to correlation values between 0.3 and 0.75. These regions contain some useful information but are of moderate importance. A regular masking strategy is used to balance information preservation and noise suppression. Classification is achieved by comparing the correlation value with the threshold value on a pixel-by-pixel basis. The result is a pixel distribution label map, in which each pixel is labeled as having strong, moderate, or weak correlation.
[0052] The mask mode settings employ different mask geometries and densities based on pixel distribution markers. The annular mask mode is specifically designed for regions with strong correlation. Its annular structure ensures complete separation between masked pixels and their associated neighbors. An inner radius of 3 pixels defines the minimum mask coverage, while an outer radius of 7 pixels defines the maximum coverage. Pixels within the annular region are set to zero or a special marker value, while pixels outside the annular region retain their original values for subsequent calculations. The checkerboard mask mode targets regions with moderate correlation. It uses 2×2 squares as the basic masking unit. Adjacent masking units are arranged in a checkerboard pattern, with adjacent units alternating to form a black and white pattern similar to a chessboard. The mask density is approximately half the total number of pixels. The random mask mode targets regions with weak correlation. A pseudorandom number generator randomly selects 15% of pixel locations for masking. This randomness ensures a uniform and irregular mask distribution. A masking rate of 15% strikes a balance between information preservation and noise reduction. A masking rate that is too high will lose useful information, while a masking rate that is too low will not effectively suppress noise.
[0053] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0054] Inputting the annular mask pattern and the chessboard mask pattern into the mask condition judgment unit to perform a logic AND operation pixel validity verification process to obtain a mask condition control signal;
[0055] Dynamically adjust the convolution kernel weights in the selective convolution calculation unit based on the mask condition control signal to obtain the conditional convolution kernel parameters with the masked pixel weights being 0;
[0056] The conditional convolution kernel parameters are input into a densely connected network containing 4 scale layers to perform 1 / 1, 1 / 2, 1 / 4, and 1 / 8 scale feature extraction processing, and a multi-scale feature group containing 6 convolution blocks in each scale layer is obtained;
[0057] Based on the multi-scale feature group, a 3×3×3 convolution kernel size stereo convolution structure is used to perform spatial and inter-band feature fusion processing to obtain a cross-band feature interaction matrix;
[0058] According to the cross-band feature interaction matrix, feature importance weights are distributed through residual connection and attention gating mechanism to obtain a multi-scale densely connected feature map.
[0059] Specifically, the mask condition judgment unit receives a circular mask pattern and a checkerboard mask pattern as input, and verifies the validity of each pixel position through a logical AND operation. The logical AND operation is a Boolean operation. When both inputs are true, the output is true, otherwise the output is false. The verification process first converts the circular mask pattern into a binary matrix, where the masked pixel positions are marked as 0 and the unmasked positions are marked as 1. The checkerboard mask pattern is also converted into the corresponding binary matrix. Then, a pixel-by-pixel logical AND operation is performed on the two binary matrices. When a pixel position is not masked in both mask patterns, the output of the position is 1, indicating that the pixel is valid. When the position is masked in any mask pattern, the output is 0, indicating that the pixel is invalid. The mask condition control signal is a binary matrix of the same size as the original image. Each element in the matrix corresponds to the validity status of a pixel position. This signal directly controls whether each pixel participates in the feature calculation in the subsequent convolution operation, effectively solving the noise coupling problem caused by the strong correlation between multi-sensor data in the composite imaging component.
[0060] The selective convolution calculation unit dynamically adjusts the convolution kernel weights based on the mask condition control signal. Traditional convolution operations use the same convolution kernel weights for all pixel positions, while conditional convolution selectively adjusts the weight values based on the mask condition. The dynamic adjustment process is achieved by performing element-wise multiplication of the mask condition control signal and the original convolution kernel weights. When the mask condition control signal is 0 at a certain position, the corresponding convolution kernel weight is set to 0. When the signal is 1, the weight remains the original value. This adjustment ensures that the masked pixels have no effect on the convolution results, while maintaining the normal feature extraction function of the unmasked pixels. The conditional convolution kernel parameters include the original weight matrix and the corresponding mask weight matrix. The mask weight matrix records the validity status of each weight position. In the subsequent feature extraction process, only valid weights participate in the actual convolution calculation, and invalid weights are skipped or set to zero. This selective calculation mechanism effectively avoids the interference of masked pixels on the feature extraction results, ensuring the purity and accuracy of the features.
[0061] A densely connected network uses conditional convolution kernel parameters to extract multi-scale features. The network consists of four scaling layers, each processing image features at the original, half, quarter, and eighth scales. Different scales correspond to different spatial resolutions and receptive fields. Scaling is achieved through pooling: 2×2 pooling kernels reduce the image size to half at half scale, 4×4 pooling kernels reduce the image size to one-quarter of its original size at quarter scale, and 8×8 pooling kernels reduce the image size to one-eighth at eighth scale. Each scaling layer contains six convolutional blocks, which are interconnected using dense connections. The output of each convolutional block is not only passed to the next convolutional block, but also directly to all subsequent blocks via skip connections. This dense connection pattern ensures the efficient flow and reuse of feature information. Low-level features extracted by the preceding convolutional block are fused with high-level features from the following convolutional block via skip connections, forming a comprehensive feature representation containing multi-level information. The multi-scale feature group concatenates feature maps of different scales along the channel dimension to form a feature set containing multi-scale information.
[0062] The stereo convolution architecture fuses spatial and inter-band features based on multi-scale feature groups. The three dimensions of the 3×3×3 convolution kernel correspond to the image's height, width, and number of bands, respectively. The first two dimensions address spatial correlation, while the third dimension addresses feature interactions between different bands. The stereo convolution operation simultaneously captures feature patterns in both spatial and band neighborhoods by sliding the convolution kernel across three-dimensional space. Unlike traditional two-dimensional convolution, which can only process spatial features from a single band, stereo convolution can establish cross-band feature correspondences and discover intrinsic connections between sensor data. The feature fusion process first stacks feature maps from the visible, infrared, and laser bands along the band dimension to form a three-dimensional feature tensor. A stereo convolution kernel is then applied to perform a convolution operation. At each location, the kernel computes the weighted sum of all elements in a 3×3×3 neighborhood. The weights are learned through training and reflect the importance of different location and band combinations. The cross-band feature interaction matrix captures the interaction strength and correlation patterns between different bands at each spatial location after feature fusion.
[0063] Residual connections and an attention gating mechanism assign feature importance weights to the cross-band feature interaction matrix. Residual connections alleviate the vanishing gradient problem in deep networks by directly adding input features to output features, ensuring that original feature information is not lost in the deep network. The attention gating mechanism assigns attention weights by calculating an importance score for each feature channel. The importance score, calculated through global average pooling and fully connected layers, reflects the contribution of each feature channel to the final task. Important features with weights close to 1 are important for the denoising task and should be fully retained. Minor features with weights between 0.2 and 0.8 are considered to have some contribution but limited importance. Irrelevant features with weights close to 0 are considered to be unhelpful for the task and should be suppressed. Weight assignment is achieved by element-wise multiplication of the attention weights with the feature map. High-weight features are amplified, while low-weight features are suppressed. The reweighted feature map retains the most useful information for the denoising task, forming a multi-scale densely connected feature map.
[0064] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0065] Apply adaptive masking and 20% random masking rate to the multi-scale densely connected feature map to generate a double mask version, and obtain the first mask version and the second mask version of the input image pair;
[0066] The first mask version and the second mask version of the input image are input to the double blind spot training network for parallel pixel prediction processing to obtain a dual-constrained prediction result;
[0067] Based on the dual-constraint prediction results, a composite loss function including reconstruction loss, consistency loss, edge protection loss and cross-band coordination loss is constructed to obtain a four-component loss numerical combination.
[0068] According to the combination of the four-component loss values, the loss weight distribution process is set by the gradient direction cosine similarity threshold of 0.9 and 0.5, and the adaptive weight adjustment parameters are obtained with weights of 1.0 and 0.2 in the early stage of training and weights of 0.8 and 1.0 in the late stage of training;
[0069] Based on the adaptive weight adjustment parameters, the network parameters are back-propagated and optimized to update, and the noise reduction prediction results are obtained.
[0070] Specifically, the dual mask version generation process uses two different masking strategies for multi-scale densely connected feature maps. The first masking strategy applies the adaptive mask generated in the previous steps, including a combination of circular mask patterns, checkerboard mask patterns, and random mask patterns. This adaptive mask is designed based on the strength of inter-pixel correlation. A circular mask is used in areas with strong correlation to prevent related pixels from being blocked simultaneously. A checkerboard mask is used in areas with medium correlation to maintain spatial uniformity. A random mask is used in areas with weak correlation to reduce regularity. The second masking strategy uses a fixed random masking rate. A pseudo-random number generator is used to randomly select 20% of the pixel positions in the entire feature map for masking. The random mask does not consider the correlation between pixels and selects pixels purely based on statistical probability. The mask density is uniformly distributed in space. The dual mask generation process first copies the multi-scale densely connected feature map into two identical copies, and then applies different masking strategies to the two copies. The masked pixel positions are set to zero values or special marked values, and the unmasked pixels retain the original feature values, forming a first mask version and a second mask version input image pair with different mask patterns. The two versions contain the same underlying image information but have different visible pixel distribution patterns.
[0071] The double-blind spot training network receives a pair of input images, one masked version and one masked version, and performs parallel pixel-wise prediction. The double-blind spot mechanism is a self-supervised learning strategy that requires the network to simultaneously predict the true values of occluded pixels in both masked versions. However, this prediction process cannot utilize the information inherent in the masked pixels themselves, relying solely on contextual information from surrounding unmasked pixels. This parallel prediction process utilizes two independent network branches, each containing the same network architecture but processing different input data. Within each branch, modules such as convolutional layers, pooling layers, and upsampling layers are used to encode and decode the input features. The encoding process compresses the input features into a low-dimensional representation, while the decoding process restores the low-dimensional representation to the original resolution. The dual-constraint mechanism requires that the two branches maintain consistent predictions for pixels at the same spatial location. When a pixel is occluded in the first masked version but visible in the second, the first branch predicts the pixel value based on contextual information, while the second branch directly uses the true value of the pixel. By comparing the predicted and true values, the network learns the correct pixel relationships. The dual-constraint prediction result contains the prediction outputs of both branches and the corresponding consistency constraints. The composite loss function is constructed by calculating four different types of loss components based on the dual-constrained prediction results. The reconstruction loss measures the reconstruction accuracy by comparing the difference between the predicted pixel values and the true pixel values. The calculation method uses the L1 norm or L2 norm to calculate the average of the pixel-level errors. The reconstruction loss directly reflects the network's ability to restore the image content. The consistency loss constrains the network to learn consistent representations by comparing the prediction differences between the two branches for the same pixel position in double-blind spot training. When the prediction results of the two branches for the same pixel differ significantly, the consistency loss increases, forcing the network to learn more stable and reliable feature representations. The edge protection loss extracts the edge information of the image through the Sobel edge detection operator. The edge structure difference between the predicted image and the real image is compared. The prediction error in the edge area is given a higher weight to ensure that the denoising process does not blur the important structural edges. The cross-band coordination loss is specially designed for composite imaging components. It constrains the coordination relationship of multi-sensor data by calculating the feature consistency of different bands at the same spatial position. When the prediction results of the visible light, infrared, and laser bands at a certain position show unreasonable differences, the loss term increases, prompting the network to learn cross-band correspondences that conform to physical laws. The four-component loss numerical combination is combined into a total loss value through weighted summation.
[0072] The gradient direction cosine similarity threshold is set based on a four-component loss value combination to dynamically assign loss weights. Gradient direction cosine similarity measures the consistency of the optimization direction by calculating the cosine of the angle between the loss function gradient vector and the reference direction vector. A similarity close to 1 indicates that the current optimization direction is highly consistent with the desired direction, while a similarity close to 0 indicates that the optimization direction deviates from the desired direction. A threshold of 0.9 is used as a high consistency judgment standard. When the similarity exceeds this threshold, it indicates that the network optimization direction is correct and stable. In this case, the loss weight is reduced to avoid over-optimization. A threshold of 0.5 is used as a low consistency judgment standard. When the similarity is lower than this threshold, it indicates that the optimization direction is deviating and needs to be constrained. In this case, the loss weight is increased to guide the network to converge in the correct direction. At the beginning of training, the random initialization of network parameters leads to unstable optimization direction. The main loss weight is set to 1.0 to ensure basic reconstruction capability, and the auxiliary loss weight is set to 0.2 to avoid introducing complex constraints too early. In the later stage of training, the network gradually converges, the main loss weight is reduced to 0.8, and the auxiliary loss weight is increased to 1.0 to strengthen the constraints on detailed features and cross-band consistency. The adaptive weight adjustment parameters change dynamically according to the training progress and gradient similarity to ensure that the network can obtain appropriate learning signals at different training stages.
[0073] Backpropagation optimization updates the network parameters using adaptive weight adjustment parameters, calculating gradients and updating them. The backpropagation algorithm uses the chain rule to calculate the partial derivatives of the loss function with respect to each network parameter. Partial derivatives indicate the impact of small changes in the parameters on the loss function. The magnitude of the gradient reflects the importance of the parameter, and the direction of the gradient indicates the direction of the parameter update. Parameter updates are implemented using a gradient descent algorithm, which shifts the parameters in the direction of the negative gradient by a step size equal to the learning rate. The learning rate controls the magnitude of the parameter update. Excessive learning rates lead to unstable training, while excessively low learning rates lead to slow convergence. During the optimization process, a momentum mechanism is used to accumulate historical gradient information, reducing parameter update oscillations and accelerating convergence. An adaptive learning rate adjustment strategy is also employed, dynamically adjusting the learning rate based on the magnitude of the gradient change. When the gradient is large, the learning rate is reduced to maintain stability, while when the gradient is small, the learning rate is increased to accelerate convergence. After multiple rounds of iterative optimization, the network parameters gradually converge to their optimal values, resulting in the final denoised prediction results.
[0074] In a specific embodiment, the step of applying the adaptive mask and a 20% random mask rate to the multi-scale densely connected feature map to generate a double-masked version may specifically include the following steps:
[0075] Input the multi-scale densely connected feature map into the mask version separator for image copy diversion processing to obtain two identical feature map copies;
[0076] Applying the circular mask pattern and the checkerboard mask pattern to the first copy of the two identical feature map copies to perform adaptive mask covering processing to obtain a first mask version image;
[0077] For the second copy of the two identical feature map copies, a pseudo-random number generator is used to randomly select pixel positions with a mask rate of 20% to obtain a random mask position index;
[0078] Performing a mask covering process on the second copy by setting pixel values to zero based on the random mask position index to obtain a second mask version image;
[0079] The first mask version image and the second mask version image are paired, combined and packaged to obtain an input image pair of the first mask version and the second mask version.
[0080] Specifically, the mask version separator receives multi-scale densely connected feature maps for image copy and diversion processing. The separator is a data distribution module that completely copies the input feature map data into two independent memory spaces through memory copy operations, forming two feature map copies with exactly the same data content but different storage addresses. The copying process is implemented by copying data pixel by pixel and channel by channel. Each numerical element of the original feature map is accurately copied to the corresponding position of the two target memory areas. The copying operation does not change the numerical content, spatial distribution or channel structure of the data, ensuring that the two copies are mathematically completely equivalent to the original feature map. The diversion processing outputs the two copied feature map copies to different processing channels respectively. The first copy is marked as the adaptive mask processing channel, and the second copy is marked as the random mask processing channel. The two channels are subsequently processed using different masking strategies. In this way, the separator converts a single input into a dual-path parallel processing data stream, laying the data foundation for the subsequent dual-blind spot training mechanism.
[0081] Adaptive mask coverage processing applies a ring mask pattern and a checkerboard mask pattern to the first feature map copy. The ring mask pattern is designed specifically for areas of strong correlation. By setting a ring-shaped occlusion area with an inner diameter of three pixels and an outer diameter of seven pixels around the target pixel, the masked pixel is completely separated from its strongly correlated neighboring pixels. Pixels inside the ring are set to zero value, while pixels outside the ring retain their original feature values. The checkerboard mask pattern targets areas of moderate correlation and uses two-by-two pixel blocks as the basic masking unit. These blocks are arranged in a black and white chessboard pattern, with adjacent masking units alternating. All pixels within the masked unit are set to zero value, while unmasked units retain their original values. This alternating pattern ensures spatial uniformity and regularity in the mask distribution. The mask covering process is implemented through bit operations, converting the mask pattern into a binary matrix, where 1 represents retained pixels and 0 represents blocked pixels. The binary matrix is then element-wise multiplied with the original feature map. The values of the blocked positions in the multiplication result become zero, and the unblocked positions retain their original values. The feature map after mask covering processing forms the first mask version image, which shows a regular pixel missing pattern in the strong correlation area and the medium correlation area.
[0082] A pseudorandom number generator (PRG) randomly selects mask locations for the second feature map copy. A Pseudorandom Number Generator (PRG) is a computational module that generates approximately random number sequences based on mathematical algorithms. It generates uniformly distributed random numbers using methods such as the linear congruential method or the Mersenne Twister algorithm, with the random numbers ranging from zero to one. The random pixel location selection process first calculates the total number of pixels in the feature map. The PRG then assigns a random value to each pixel location. When the random value is less than 0.2, the pixel location is selected for masking. When the random value is greater than or equal to 0.2, the pixel location remains unchanged. This selection mechanism ensures that approximately 20 percent of the pixel locations are randomly selected. The selected pixel locations exhibit a random spatial distribution, unlike the structured distribution of adaptive masks. A random mask location index is formed by recording the row and column coordinates of the selected pixels. The index list contains information about all pixel locations to be masked. The index format is a collection of coordinate pairs, each pair identifying the precise location of a selected pixel in the feature map.
[0083] The pixel value zero mask overlay process numerically modifies the second copy based on the random mask position index. The overlay process is achieved by traversing each coordinate pair in the index list and setting the pixel value of the corresponding position to zero. The zero operation is a simple and effective masking method. It simulates the complete loss of pixel information by forcing the value of the selected pixel to zero. The zero value will not affect the feature extraction of neighboring pixels in subsequent convolution calculations, thereby achieving a pixel-level information masking effect. The mask overlay process processes each selected position one by one in index order. For multi-channel feature maps, the zero operation is applied to all channels at the position simultaneously to ensure that the pixel is completely blocked in all feature dimensions. After the overlay process is completed, a second mask version image is formed. The image presents a randomly distributed pixel missing pattern in space, with a missing density of approximately 20% of the total number of pixels. The missing position is independent of the feature content and is purely based on statistical probability distribution.
[0084] The paired combination encapsulation process combines the first and second masked versions of the image into a data structure. This encapsulation is achieved by creating a data container containing references to both images. The container structure includes metadata such as a version identifier, an image data pointer, and a mask pattern description. The version identifier distinguishes the two mask versions: the first version is marked as an adaptive mask type, and the second version is marked as a random mask type. This identifier is used to select the appropriate algorithm branch in subsequent processing. The image data pointer points to the memory address where the feature map values are actually stored. This pointer access method avoids duplicate data copies, saving memory space and computation time. The mask pattern description records the specific masking strategy and parameter settings used for each version, including detailed information such as mask type, mask density, and mask distribution pattern. This information provides essential reference for subsequent loss function calculation and gradient backpropagation. After encapsulation, the input image pair data structure is formed. This structure serves as the input to the double-blind spot training network. It contains two image versions with the same underlying content but different mask patterns, providing the necessary data diversity for the self-supervised learning mechanism.
[0085] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0086] The noise reduction prediction results are evaluated for image quality by calculating the signal-to-noise ratio, edge preservation index, and structural similarity index, and quality grading marks are obtained, with the signal-to-noise ratio greater than 25dB in high-quality areas, 15-25dB in medium-quality areas, and less than 15dB in low-quality areas.
[0087] Based on the quality grading mark, the noise reduction prediction results are divided into 8×8 non-overlapping blocks, and the statistical characteristics of the pixels in the blocks are analyzed to obtain the noise power spectrum density and signal power spectrum density parameters of each block;
[0088] According to the noise power spectrum density and signal power spectrum density parameters, the blocks containing edge texture information are multiplied by the attenuation factor 0.3, and the smooth area blocks are kept at the original value to adjust the Wiener filter parameters to obtain the adaptive filter transfer function H(u,v);
[0089] The adaptive filter transfer function H(u,v) is input into the multi-scale Canny operator to perform edge detection with thresholds of 0.1, 0.15, and 0.2, and edge classification processing with strong edge gradient amplitude greater than 50, medium edge gradient amplitude of 20-50, and weak edge gradient amplitude of 10-20 to obtain the graded edge protection weight parameters;
[0090] Based on the hierarchical edge protection weight parameters, the original pixel values in the strong edge area are maintained, the medium edge area is fused with the 0.7 times filtering result and the 0.3 times original value, and the weak edge area is fused with the 0.9 times filtering result and the 0.1 times original value to obtain the final denoised image of the composite imaging component.
[0091] Specifically, the image quality assessment process comprehensively calculates three metrics based on the noise reduction prediction results. The signal-to-noise ratio (SNR) measures image clarity by calculating the ratio of signal power to noise power. The calculation process first separates the useful signal and noise components in the image. The useful signal is obtained through low-pass filtering, and the noise component is obtained by subtracting the useful signal from the original image. The power of the two components is then calculated and the ratio is calculated. The final result is converted to decibels. The edge preservation index evaluates the edge preservation effect by comparing the edge structure similarity of the images before and after noise reduction. The calculation process uses the Sobel operator to extract edge information from the images before and after noise reduction, and then calculates the normalized cross-correlation coefficient between the two edge images. A coefficient closer to 1 indicates better edge preservation. The structural similarity index comprehensively evaluates image quality by comparing image brightness, contrast, and structural information. The calculation process divides the image into several local windows and calculates brightness similarity, contrast similarity, and structural similarity within each window. The three components are then weighted and combined to form the local similarity index. Finally, all local indices are averaged to obtain the global structural similarity index. The quality grading mark is divided into regions according to the signal-to-noise ratio value. When the signal-to-noise ratio of a region exceeds 25 decibels, the region is marked as a high-quality region, indicating that the noise reduction effect is good and the noise level is very low. When the signal-to-noise ratio is between 15 and 25 decibels, it is marked as a medium-quality region, indicating that the noise reduction effect is generally good but there is still some residual noise. When the signal-to-noise ratio is lower than 15 decibels, it is marked as a low-quality region, indicating that the noise reduction effect is poor and the noise is more serious.
[0092] Block segmentation and statistical analysis divide the noise reduction prediction results into non-overlapping 8×8 pixel square blocks based on quality grading. The segmentation process starts from the top left corner of the image and proceeds from left to right and top to bottom. Each block contains 64 pixels, and block boundaries are strictly spaced 8 pixels apart to ensure no overlap or omissions between adjacent blocks. Intra-block pixel statistical analysis calculates the frequency domain characteristics of noise and signal for each block independently. The analysis first applies a two-dimensional discrete Fourier transform to the 8×8 pixel block to convert the spatial signal into a frequency domain representation. The squared amplitude of the frequency domain coefficient represents the power density of that frequency component. The noise power spectral density is obtained by analyzing the high-frequency components. Since noise is typically concentrated in these high-frequency regions, the calculation method is to sort the frequency domain coefficients from high to low frequency and select the squared amplitude of the coefficient in the high-frequency range as the noise power density estimate. The signal power spectral density is obtained by analyzing the mid- and low-frequency components. Since the main structural information of the image is concentrated in the mid- and low-frequency parts, the calculation method is to select the square of the coefficient amplitude of the mid- and low-frequency bands as the signal power density estimate. The noise power spectral density and signal power spectral density parameters of each block are independently calculated through this frequency domain analysis method.
[0093] The Wiener filter parameter adjustment process adaptively designs a filter based on the power spectral density parameters of each block. Wiener filtering is an optimal linear filtering method based on the minimum mean square error criterion. The filter transfer function calculation considers both the signal power spectral density and the noise power spectral density. Edge and texture information is detected by calculating the gradient variance of pixels within a block. When the gradient variance exceeds a set threshold, the block is determined to contain edge or texture information. The noise power spectral density of such blocks is multiplied by a factor of 0.3 to reduce the noise estimate and avoid misidentifying edge texture as noise and over-filtering. For smooth regions, where the gradient variance is small, the original noise power spectral density estimate is retained without adjustment to ensure effective suppression of real noise. The adaptive filter transfer function H(u,v) is calculated as the signal power spectral density divided by the sum of the signal power spectral density and the adjusted noise power spectral density. u and v represent the horizontal and vertical frequency coordinates, respectively. The transfer function has different filtering strengths at different frequencies. Frequency points with high signal-to-noise ratios have larger transfer coefficients to preserve more signal, while frequency points with low signal-to-noise ratios have smaller transfer coefficients to suppress more noise.
[0094] Multi-scale Canny edge detection uses an adaptive filter transfer function to extract and classify edges. The Canny operator is a multi-stage edge detection algorithm consisting of four steps: Gaussian filtering, gradient calculation, non-maximum suppression, and dual-threshold detection. Multi-scale processing captures edge information of varying strengths by setting three different detection thresholds: 0.1, 0.15, and 0.2. A low threshold of 0.1 detects weak edges, including texture details and noise boundaries; a medium threshold of 0.15 detects moderately strong edges, primarily object outlines and structural boundaries; and a high threshold of 0.2 detects strong edges, primarily significant object boundaries and high-contrast structures. Edge classification is quantitatively assessed by calculating the gradient magnitude, calculated using the Sobel operator, which reflects the magnitude of pixel intensity changes. A gradient magnitude greater than 50 is classified as a strong edge, indicating a clear and significant boundary. A gradient magnitude between 20 and 50 is classified as a medium edge, indicating a distinct but moderately significant boundary. A gradient magnitude between 10 and 20 is classified as a weak edge, indicating a blurred boundary or texture detail. The graded edge protection weight parameter sets different protection levels according to the edge strength. A strong edge is assigned a weight of 1.0, indicating full protection; a medium edge is assigned a weight of 0.7, indicating partial protection; and a weak edge is assigned a weight of 0.3, indicating mild protection.
[0095] Adaptive edge protection uses differentiated fusion strategies for different edge types based on a graded edge protection weight parameter. Strong edge regions, due to their high importance and clarity, receive a full protection strategy, maintaining the original pixel values without any filtering, ensuring the preservation of important boundary information. Moderate edge regions employ a weighted fusion strategy, linearly combining the filtered result by 0.7 and the original pixel value by 0.3. This fusion approach removes some noise while preserving most edge information, balancing noise reduction and edge protection. Weak edge regions, primarily containing texture details and noisy edges, receive a filtering-based fusion strategy, combining the filtered result by 0.9 and the original pixel value by 0.1. This prioritizes noise removal while still preserving a small amount of original information and avoiding oversmoothing. Fusion is performed using a pixel-by-pixel weighted average. Each pixel position is assigned a fusion weight based on its edge classification result. The calculation ensures that the sum of all pixel weights equals 1, maintaining overall image brightness balance. The resulting composite imaging component de-noised image effectively suppresses noise while maintaining the clarity of important edge and structural information.
[0096] The above describes the optical image noise reduction method of the composite imaging assembly in the embodiment of the present application. The following describes the optical image noise reduction system of the composite imaging assembly in the embodiment of the present application. Figure 2 In one embodiment of the present application, an optical image noise reduction system of a composite imaging assembly includes:
[0097] A correction module is used to perform geometric correction processing on the raw image data of the visible light sensor, infrared sensor and laser ranging sensor in the composite imaging component through a multi-sensor time synchronization mechanism to obtain a multi-sensor noise coupling matrix;
[0098] A generation module is used to adjust the mask strategy of images of different bands by using a multispectral adaptive mask generation algorithm according to the multi-sensor noise coupling matrix to obtain a ring mask pattern and a checkerboard mask pattern;
[0099] An interaction module, configured to input the annular mask pattern and the checkerboard mask pattern into a conditional mask convolution block for cross-band feature interaction processing to obtain a multi-scale densely connected feature map;
[0100] An adjustment module is used to perform weight adaptive adjustment processing on the cross-band coordination loss according to the multi-scale densely connected feature map through a double-blind spot self-supervised learning algorithm to obtain a noise reduction prediction result;
[0101] The protection module is used to input the noise reduction prediction result into a block-based Wiener filter to perform edge detail protection processing to obtain a final noise reduction image of the composite imaging component.
[0102] above Figure 2 The optical image noise reduction system of the composite imaging component in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The optical image noise reduction device of the composite imaging component in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0103] Reference Figure 3 In an embodiment of the present invention, there is also provided an optical image noise reduction device of a composite imaging component. The optical image noise reduction device of the composite imaging component can be a server, and its internal structure can be as follows: Figure 3 As shown. The optical image noise reduction device of the composite imaging component includes a processor, a memory, a display screen, an input device, a network interface and a database connected via a system bus. Among them, the computer-designed processor is used to provide computing and control capabilities. The memory of the optical image noise reduction device of the composite imaging component includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the optical image noise reduction device of the composite imaging component is used to store the corresponding data in this embodiment. The network interface of the optical image noise reduction device of the composite imaging component is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0104] Those skilled in the art will understand that Figure 3The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention, and does not constitute a limitation on the optical image noise reduction device of the composite imaging assembly to which the solution of the present invention is applied.
[0105] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the optical image noise reduction method of the composite imaging component.
[0106] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0107] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an optical image noise reduction device (which can be a personal computer, server, or network device, etc.) of a composite imaging assembly to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0108] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for reducing optical image noise of a composite imaging component, characterized in that: The method comprises: The raw image data of the visible light sensor, infrared sensor, and laser ranging sensor in the composite imaging component are geometrically corrected through a multi-sensor time synchronization mechanism to obtain a multi-sensor noise coupling matrix. Specifically, the following steps are performed: The raw image data of the visible light sensor, infrared sensor, and laser ranging sensor are time-synchronized and calibrated based on the timestamp mark, and a multi-sensor synchronized image sequence with a synchronization accuracy of 10 microseconds is obtained. Inputting the multi-sensor synchronized image sequence into an affine transformation algorithm for geometric registration processing to obtain a spatially aligned image group with a registration error less than 0.5 pixels; performing photon shot noise and readout noise parameter extraction processing on the visible light sensor in the spatially aligned image group to obtain a visible light noise feature vector; performing thermal noise, dark current noise, and speckle coherence noise parameter extraction processing on the infrared sensor and the laser ranging sensor in the spatially aligned image group, respectively, to obtain an infrared noise feature vector and a laser noise feature vector; Based on the visible light noise feature vector, the infrared noise feature vector and the laser noise feature vector, a 3×3 inter-sensor noise correlation calculation process is constructed to obtain a multi-sensor noise coupling matrix; According to the multi-sensor noise coupling matrix, a mask strategy adjustment process is performed on images of different bands using a multispectral adaptive mask generation algorithm to obtain a ring mask pattern and a checkerboard mask pattern; Inputting the annular mask pattern and the checkerboard mask pattern into a conditional mask convolution block for cross-band feature interaction processing to obtain a multi-scale densely connected feature map; According to the multi-scale densely connected feature map, a weight adaptive adjustment process is performed on the cross-band coordination loss through a double-blind spot self-supervised learning algorithm to obtain a noise reduction prediction result; The noise reduction prediction result is input into a block-based Wiener filter for edge detail protection processing to obtain a final noise reduction image of the composite imaging component.
2. The optical image noise reduction method of a composite imaging assembly according to claim 1, characterized in that: The mask strategy adjustment process is performed on images of different bands using a multispectral adaptive mask generation algorithm according to the multi-sensor noise coupling matrix to obtain a ring mask pattern and a checkerboard mask pattern, including: The multi-sensor noise coupling matrix is input into the band feature encoding module to perform independent feature extraction processing on the visible light, infrared and laser bands to obtain a three-branch convolution feature group containing 64, 128 and 256 channels; Based on the three-branch convolutional feature group, spatial attention weight calculation is performed through a dilated convolution structure with expansion rates of 2, 4, and 8, and a multi-scale spatial correlation weight matrix with a receptive field coverage radius of 5, 9, and 17 pixels is obtained; Performing cross-band attention similarity calculation on pixel positions corresponding to different bands according to the multi-scale spatial correlation weight matrix to obtain a cross-band correlation matrix with a value range of 0 to 1; Based on the correlation strength thresholds of 0.75 and 0.3 in the cross-band correlation matrix, the pixel area is subjected to mask pattern classification and judgment processing to obtain pixel distribution labels of strong correlation area, medium correlation area and weak correlation area; According to the pixel distribution mark, a ring mask pattern with an inner diameter of 3 pixels and an outer diameter of 7 pixels is set for the strong correlation area, a checkerboard mask pattern with a mask block size of 2×2 is set for the medium correlation area, and a random mask processing with a mask rate of 15% is set for the weak correlation area, thereby obtaining a ring mask pattern and a checkerboard mask pattern.
3. The optical image noise reduction method of a composite imaging assembly according to claim 1, characterized in that: The ring mask pattern and the checkerboard mask pattern are input into the conditional mask convolution block for cross-band feature interaction processing to obtain a multi-scale densely connected feature map, including: Inputting the annular mask pattern and the checkerboard mask pattern into a mask condition judgment unit to perform a logic AND operation pixel validity verification process to obtain a mask condition control signal; Dynamically adjusting the convolution kernel weights in the selective convolution calculation unit based on the mask condition control signal to obtain conditional convolution kernel parameters with a masked pixel weight of 0; Inputting the conditional convolution kernel parameters into a densely connected network containing 4 scale layers to perform 1 / 1, 1 / 2, 1 / 4, and 1 / 8 scale feature extraction processing to obtain a multi-scale feature group containing 6 convolution blocks in each scale layer; Based on the multi-scale feature group, spatial and inter-band feature fusion processing is performed through a three-dimensional convolution structure with a convolution kernel size of 3×3×3 to obtain a cross-band feature interaction matrix; According to the cross-band feature interaction matrix, feature importance weight distribution is performed through residual connection and attention gating mechanism to obtain a multi-scale dense connection feature map.
4. The optical image noise reduction method of a composite imaging assembly according to claim 1, characterized in that: The method of performing weight adaptive adjustment processing on the cross-band coordination loss by using a double-blind-spot self-supervised learning algorithm according to the multi-scale densely connected feature map to obtain a noise reduction prediction result includes: Applying adaptive masking and a 20% random masking rate to the multi-scale densely connected feature map to perform a double mask version generation process to obtain a first mask version and a second mask version of the input image pair; Performing parallel pixel prediction processing on the input images of the first mask version and the second mask version into the double blind spot training network to obtain a dual-constraint prediction result; Based on the dual-constraint prediction results, a composite loss function including reconstruction loss, consistency loss, edge protection loss and cross-band coordination loss is constructed for calculation and processing to obtain a four-component loss numerical combination; According to the four-component loss value combination, the loss weight distribution process is set by gradient direction cosine similarity thresholds of 0.9 and 0.5, and adaptive weight adjustment parameters with weights of 1.0 and 0.2 in the early stage of training and 0.8 and 1.0 in the late stage of training are obtained; Based on the adaptive weight adjustment parameters, the network parameters are back-propagated and optimized to update, thereby obtaining a noise reduction prediction result.
5. The optical image noise reduction method of the composite imaging assembly according to claim 4, characterized in that: The step of applying an adaptive mask and a 20% random mask rate to the multi-scale densely connected feature map to perform a dual mask version generation process to obtain an input image pair of a first mask version and a second mask version includes: Inputting the multi-scale densely connected feature map into a mask version separator for image copy shunting processing to obtain two identical feature map copies; Applying a circular mask pattern and a checkerboard mask pattern to perform adaptive mask covering processing on a first copy of the two identical feature map copies to obtain a first mask version image; Performing random pixel position selection processing with a mask rate of 20% on the second copy of the two identical feature map copies using a pseudo-random number generator to obtain a random mask position index; Performing a mask covering process of setting pixel values to zero on the second copy based on the random mask position index to obtain a second mask version image; The first mask version image and the second mask version image are paired, combined and packaged to obtain an input image pair of the first mask version and the second mask version.
6. The optical image noise reduction method of a composite imaging assembly according to claim 1, characterized in that: The step of inputting the noise reduction prediction result into a block-based Wiener filter for edge detail protection processing to obtain a final noise reduction image of the composite imaging component comprises: Performing image quality assessment on the noise reduction prediction results by calculating the signal-to-noise ratio, edge preservation index, and structural similarity index, and obtaining quality grading marks such as a signal-to-noise ratio greater than 25dB in high-quality areas, 15-25dB in medium-quality areas, and less than 15dB in low-quality areas; Dividing the noise reduction prediction results into 8×8 non-overlapping blocks based on the quality grading mark, performing pixel statistical characteristic analysis processing within the blocks, and obtaining noise power spectral density and signal power spectral density parameters of each block; According to the noise power spectrum density and signal power spectrum density parameters, the block containing edge texture information is multiplied by an attenuation factor of 0.3, and the smooth area block is kept at the original value to perform Wiener filter parameter adjustment processing to obtain an adaptive filter transfer function H(u,v); Inputting the adaptive filter transfer function H(u, v) into a multi-scale Canny operator to perform edge detection with thresholds of 0.1, 0.15, and 0.2 and edge classification processing with strong edge gradient amplitude greater than 50, medium edge gradient amplitude of 20-50, and weak edge gradient amplitude of 10-20 to obtain a graded edge protection weight parameter; Based on the hierarchical edge protection weight parameters, edge adaptive protection processing is performed on the strong edge area by maintaining the original pixel value, the medium edge area adopts 0.7 times filtering result and 0.3 times original value fusion, and the weak edge area adopts 0.9 times filtering result and 0.1 times original value fusion to obtain the final noise reduction image of the composite imaging component.
7. An optical image noise reduction system for a composite imaging assembly, characterized in that: The optical image noise reduction method for implementing the composite imaging assembly according to any one of claims 1 to 6, wherein the optical image noise reduction system of the composite imaging assembly comprises: A correction module is used to perform geometric correction processing on the raw image data of the visible light sensor, infrared sensor and laser ranging sensor in the composite imaging component through a multi-sensor time synchronization mechanism to obtain a multi-sensor noise coupling matrix; A generation module is used to adjust the mask strategy of images of different bands by using a multispectral adaptive mask generation algorithm according to the multi-sensor noise coupling matrix to obtain a ring mask pattern and a checkerboard mask pattern; An interaction module, configured to input the annular mask pattern and the checkerboard mask pattern into a conditional mask convolution block for cross-band feature interaction processing to obtain a multi-scale densely connected feature map; An adjustment module is used to perform weight adaptive adjustment processing on the cross-band coordination loss according to the multi-scale densely connected feature map through a double-blind spot self-supervised learning algorithm to obtain a noise reduction prediction result; The protection module is used to input the noise reduction prediction result into a block-based Wiener filter to perform edge detail protection processing to obtain a final noise reduction image of the composite imaging component.
8. An optical image noise reduction device for a composite imaging assembly, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the optical image noise reduction method of the composite imaging component according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the processor is caused to perform the optical image noise reduction method of the composite imaging assembly according to any one of claims 1 to 6.
Citation Information
Patent Citations
Unmanned system image multi-noise interference suppression method based on self-supervised continuous learning
CN117853734A
Infrared image super-resolution reconstruction method based on noise decoupling
CN120198293A