Image deblurring method, system and device of optical communication chip based on convolutional neural network, and medium

By using a multi-scale optical distortion compensation pyramid network and a dual-domain attention mechanism based on convolutional neural networks, the image blurring problem of optical communication chips was solved, achieving efficient and accurate image deblurring processing and improving the automation capability of chip inspection.

CN120976060AActive Publication Date: 2025-11-18湖南奥创普科技有限公司
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511208153.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-18
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

The images of optical communication chips are blurred due to their complex structure and precision manufacturing process. Existing methods are difficult to effectively remove the blur, which affects the accuracy and efficiency of detection.

Method used

A convolutional neural network-based approach is employed, utilizing a multi-scale optical distortion compensation pyramid network, a cascaded channel-priority dual-domain attention mechanism, and a multi-level dilated convolutional group with adaptive dilation rate to pre-extract and enhance optical features, thereby reconstructing a clear image.

Benefits of technology

It improves the efficiency and accuracy of image deblurring processing for optical communication chips, automatically extracts advanced semantic features, enhances subpixel-level detail recognition, and adapts to diverse chip backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976060A_ABST
    Figure CN120976060A_ABST
Patent Text Reader

Abstract

The invention relates to an image deblurring method, system and device for an optical communication chip based on a convolutional neural network, and a medium, and the method comprises the steps: carrying out the optical feature pre-extraction of an input image of the optical communication chip based on an improved pyramid network, and obtaining a receptive field extension feature map through the parallel feature fusion of a multi-size convolution kernel; performing multi-scale feature extraction on the receptive field extended feature map according to an encoder, and introducing a double-domain attention mechanism to perform feature enhancement on an extracted feature vector to obtain a multi-scale feature image; performing resolution self-adaptive receptive field matching fusion on different levels of feature maps of the multi-scale feature image and a decoder by adopting jump connection improved by a cavity convolution group; and inputting the fusion features into a decoder to perform progressive up-sampling processing, and synchronously adopting a multi-scale composite loss function to optimize the multi-level features output by the decoder to reconstruct a clear image of the optical communication chip. According to the invention, the efficiency and precision of deblurring processing of the image of the optical communication chip are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to an image deblurring method, system, device and medium for an optical communication chip based on a convolutional neural network. BACKGROUND

[0002] In the field of optical communication chips, the image quality of the chip is crucial for its performance evaluation and defect detection. However, due to the complex structure and precise manufacturing process of optical communication chips, the images often have blurring problems. This blurring can be caused by a variety of factors, including imperfections in the optical imaging system, small defects on the chip surface, chip stage vibration or laser scanning galvanometer jitter leading to image blurring, optical system defocus or chip surface microstructure height difference causing defocus effect, and sub-micron features such as electrode stripes that are difficult to distinguish in blurred images.

[0003] Blurred images not only reduce the accuracy of chip detection, but also increase the difficulty of subsequent processing. The image deblurring technology for optical communication chips faces many challenges. First, there are many types of chips, with complex appearances and large differences in optical characteristics. Second, the defect size on the chip surface is small, and the contrast difference with the background area is weak, which makes it difficult to accurately identify and separate defects in blurred images. In addition, the instability of the chip manufacturing process leads to a variety of defects, and the same defect may exhibit multiple different feature forms, further increasing the difficulty of deblurring and defect detection.

[0004] Traditional image deblurring methods mainly rely on manual adjustment of parameters and pre-set image processing algorithms. These methods often struggle to adapt to complex blur patterns and diverse chip backgrounds when processing optical communication chip images. Due to the variety of chip image blur types (such as motion blur, defocus blur, microstructure feature loss, etc.), traditional methods cannot effectively distinguish different types of blur, resulting in unsatisfactory deblurring results. In addition, traditional methods have high computational complexity when processing high-resolution images, making it difficult to meet real-time requirements.

[0005] Furthermore, in the existing image taking process, there is a lot of manual intervention, which relies on the experience and subjective judgment of operators to determine whether the image is blurred. Due to the complexity of optical communication chip images, manual parameter adjustment and re-imaging not only consume time and effort, but also make it difficult to ensure the consistency of deblurring results. Especially for the detection of small defects, manual operation is difficult to accurately quantify the degree of blurring, and if the picture is not clear, it is easy to cause missed detection and false detection. SUMMARY

[0006] (I) Technical problems to be solved

[0007] In view of the above-mentioned defects and deficiencies of the prior art, the present application provides an image deblurring method, system, device and medium based on a convolutional neural network for an optical communication chip, which solves the technical problem of low deblurring quality in complex scenes in the prior art.

[0008] (II) Technical solutions

[0009] In order to achieve the above-mentioned purpose, the main technical solutions adopted by the present application include:

[0010] In a first aspect, the present application provides an image deblurring method based on a convolutional neural network for an optical communication chip, comprising:

[0011] Based on the pre-constructed multi-scale optical distortion compensation pyramid network, the optical feature of the input optical communication chip image is pre-extracted, and through the parallel feature fusion of the multi-size convolution kernel, the receptive field expansion feature map with the cross-scale light intensity distribution feature is obtained;

[0012] According to the encoder in the convolutional neural network, the multi-scale feature extraction is performed on the receptive field expansion feature map, and the feature enhancement is performed on the extracted feature vector by introducing the cascaded channel priority dual-domain attention mechanism, so as to obtain the multi-scale feature image of the optical communication chip;

[0013] The improved skip connection of the multi-level adaptive expansion rate is adopted, so that the feature maps of different levels of the multi-scale feature image are adaptively matched and fused with the receptive field of the decoder;

[0014] The fused features are input into the decoder for progressive up-sampling processing, and a multi-scale composite loss function of fused pixel level, feature level and frequency domain reconstruction is simultaneously adopted to optimize the multi-level features output by the decoder, so as to reconstruct the clear image of the optical communication chip.

[0015] Optionally, before the optical feature of the input optical communication chip image is pre-extracted based on the pre-constructed multi-scale optical distortion compensation pyramid network, and through the parallel feature fusion of the multi-size convolution kernel, the receptive field expansion feature map with the cross-scale light intensity distribution feature is obtained, it further comprises:

[0016] The acquired blurred optical communication chip image data is pre-processed for image adjustment and feature enhancement to obtain an image training set;

[0017] A multi-scale optical distortion compensation pyramid network is set at the input layer of the original convolutional neural network, and the improved convolutional neural network is subjected to model pruning, quantization compression and lightweight processing of knowledge distillation;

[0018] The improved convolutional neural network after lightening is trained according to the image training set, and the training result is evaluated in multiple dimensions through a set of multi-scale composite loss functions, and the optimal convolutional neural network is output according to the evaluation result.

[0019] Optionally, the acquired blurred optical communication chip image data is preprocessed through image adjustment and feature enhancement, and an image training set is obtained.

[0020] The acquired optical communication chip image data is blurred;

[0021] The optical communication chip image data after blurring is subjected to image adjustment of multi-size transformation and normalization processing, and normalized image data output in a set of sizes is obtained;

[0022] The normalized image data is subjected to feature enhancement of noise suppression, interested region cropping and data enhancement, and an image training set with enhanced features is obtained.

[0023] Optionally, the training result is evaluated in multiple dimensions through a set of multi-scale composite loss functions, and the optimal convolutional neural network is output according to the evaluation result.

[0024] The mean square error is used to calculate the pixel difference between the predicted image and the actual image output by the convolutional neural network, and a multi-scale pixel difference measurement is realized in combination with a multi-scale content loss function;

[0025] The multi-scale Manhattan distance between the predicted image and the actual image is calculated in the same frequency domain through a multi-scale frequency reconstruction loss function;

[0026] The feature vectors of the predicted image and the actual image at the same position are extracted using a pre-trained convolutional neural network, and the similarity of the extracted feature vectors is compared using a perception loss function to obtain the feature similarity of the predicted image and the actual image;

[0027] The frequency components higher than a set frequency threshold in the frequency spectrum of the predicted image and the actual image are weighted and allocated using a weighted Fourier frequency domain loss function to obtain the frequency domain constraint of the convolutional neural network.

[0028] According to the multi-scale pixel difference value, the multi-scale Manhattan distance, the feature similarity and the frequency domain constraint, the training result is evaluated in multiple dimensions in combination with the weight ratio of multi-dimensional loss evaluation, and the optimal convolutional neural network is output according to the evaluation result.

[0029] Optionally, the input optical communication chip image is pre-extracted based on a pre-constructed multi-scale optical distortion compensation pyramid network, and the receptive field expansion feature map with cross-scale light intensity distribution features is obtained through parallel feature fusion of multi-size convolution kernels.

[0030] Different scale convolution branches are arranged in parallel to extract optical feature response maps with different receptive fields in the optical communication chip image, and the optical feature response maps include local defect features caused by scattering, transmission inhomogeneity or material property difference of the optical structure edge;

[0031] Upsampling operations are performed on each optical feature response map to the original spatial resolution of the optical communication chip image to obtain feature maps corresponding to different original receptive fields;

[0032] The feature maps of different original receptive fields after uniform resolution are cascaded and fused in the channel dimension through the parallel features of multi-size convolution kernels to obtain a receptive field expansion feature map with cross-scale light intensity distribution characteristics.

[0033] Optionally, the encoder in the convolutional neural network performs multi-scale feature extraction on the receptive field expansion feature map, and introduces a cascaded channel-priority dual-domain attention mechanism to perform feature enhancement on the extracted feature vectors to obtain a multi-scale feature image of the optical communication chip, including:

[0034] A multi-scale composite loss function that fuses pixel-level, feature-level and frequency domain reconstruction is used to optimize the encoder to perform multi-scale feature extraction on the receptive field expansion feature map to obtain an initial multi-scale feature image of the optical communication chip;

[0035] According to the cross-scale light intensity distribution characteristics of the receptive field expansion feature map, the contribution weights of each feature channel are dynamically adjusted by combining the channel domain attention module in the dual-domain attention mechanism to adaptively compensate for the feature expression deviation caused by uneven illumination in the initial multi-scale feature image;

[0036] Based on the waveguide structure information of the optical communication chip, the spatial domain attention module in the dual-domain attention mechanism is used to strengthen the pixel features in the waveguide edge region and the adjacent weak texture region in the multi-scale feature image after channel enhancement to obtain a multi-scale feature image with high spatial recognition of the strengthened waveguide structure;

[0037] The channel domain attention module dynamically adjusts the saliency response strength of each feature channel through light intensity gradient guided channel weight mapping; the spatial domain attention module generates a spatial position sensitive attention mask based on the edge gradient distribution characteristics of the distortion region to strengthen the sub-pixel level detail features of the blurred boundary.

[0038] Optionally, a multi-level adaptive dilation rate improved hole convolution group is used to improve the skip connection to perform resolution adaptive receptive field matching and fusion of different level feature maps of the multi-scale feature image and the decoder, including:

[0039] In the decoder and encoder skip connection path, the multi-level adaptive inflation rate of the hollow convolution group is introduced to improve the skip connection mechanism.

[0040] Through the improved skip connection mechanism, the feature vectors of different levels in the multi-scale feature image extracted by the encoder are adaptively matched and fused with the resolution of the receptive field of the decoder, so that multi-scale fusion features capable of capturing local details and global context information are obtained.

[0041] In a second aspect, the embodiment of the present application provides an image deblurring system of an optical communication chip based on a convolutional neural network, comprising:

[0042] An optical feature pre-extraction module is configured to pre-extract optical features from an input optical communication chip image based on a pre-constructed multi-scale optical distortion compensation pyramid network, and to obtain receptive field expansion feature maps with cross-scale light intensity distribution features through parallel feature fusion of multi-size convolution kernels.

[0043] An encoding module is configured to extract multi-scale features from the receptive field expansion feature maps by an encoder in the convolutional neural network, and to enhance the extracted feature vectors by introducing a cascaded channel-priority dual-domain attention mechanism, so as to obtain multi-scale feature images of the optical communication chip.

[0044] An adaptive receptive field matching module is configured to improve the skip connection by using a multi-level adaptive inflation rate of a hollow convolution group, and to adaptively match and fuse different level feature maps of the multi-scale feature images with the resolution of the receptive field of the decoder.

[0045] A decoding module is configured to input the fused features into the decoder for progressive up-sampling processing, and to simultaneously optimize the multi-level features output by the decoder by using a multi-scale composite loss function of fused pixel-level, feature-level and frequency domain reconstruction, so as to reconstruct a clear image of the optical communication chip.

[0046] In a third aspect, the embodiment of the present application provides an electronic device, comprising:

[0047] A processor;

[0048] A memory is configured to store a method for controlling the above-mentioned image deblurring method of the optical communication chip based on the convolutional neural network.

[0049] In a fourth aspect, the embodiment of the present application provides a computer readable medium having computer executable instructions stored thereon, wherein the executable instructions are executed by a processor to implement the above-mentioned image deblurring method of the optical communication chip based on the convolutional neural network.

[0050] (III) Beneficial effects

[0051] The beneficial effects of the present application are: the present application adopts a convolutional neural network to perform image deblurring processing on an optical communication chip image, which can automatically extract high-level semantic features in the image, without the need for manual design of a feature extractor, thereby greatly improving the efficiency and accuracy of image deblurring processing. Moreover, the present application also performs optical feature pre-extraction on the input optical communication chip image through a pre-constructed multi-scale optical distortion compensation pyramid network to expand the receptive field of the cross-scale light intensity distribution. Furthermore, the present application introduces a cascaded channel-priority dual-domain attention mechanism in the encoder part of the convolutional neural network to screen the extracted features, thereby improving the sub-pixel level detail feature recognition of the optical communication chip image. Finally, the present application adopts a multi-level adaptive dilated rate hollow convolution group to enhance the multi-resolution detail feature processing when performing a skip connection, which can further improve the efficiency and accuracy of deblurring processing. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 A flowchart of an image deblurring method for an optical communication chip based on a convolutional neural network is provided for an embodiment of the present application.

[0053] Figure 2 A structure diagram of a convolutional neural network is provided for an embodiment of the present application.

[0054] Figure 3 A structure diagram of an improved pyramid is provided for an embodiment of the present application.

[0055] Figure 4 A structure diagram of a dual-domain attention mechanism is provided for an embodiment of the present application.

[0056] Figure 5 A structure diagram of a channel domain attention module is provided for an embodiment of the present application.

[0057] Figure 6 A structure diagram of a spatial domain attention module is provided for an embodiment of the present application.

[0058] Figure 7 A structure diagram of a hollow convolution group is provided for an embodiment of the present application.

[0059] Figures 8 to 13 A visual effect diagram before and after six-group optical communication chip image deblurring processing is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0060] In order to better explain the present application and facilitate understanding, the present application will be described in detail below through specific embodiments in combination with the accompanying drawings.

[0061] Before that, in order to facilitate understanding of the technical solutions provided by the present application, some basic information related to the technical solutions of the present application will be introduced first.

[0062] PSNR (Peak Signal-to-Noise Ratio): PSNR is a calculation based on mean square error (MSE), the larger the value of PSNR, the better the image quality, that is, the smaller the difference between images.

[0063] SSIM (Structural Similarity): SSIM takes into account that the human eye is more sensitive to structural information of an image, and it evaluates the similarity of images from three aspects of brightness, contrast and structure. The closer the value of SSIM to 1, the more similar the images.

[0064] Reference Figures 1-7 As shown in the figure, the image deblurring method for optical communication chips based on the convolutional neural network proposed in the embodiment of the application comprises the following steps: first, optical feature pre-extraction is performed on the input optical communication chip image based on a pre-constructed multi-scale optical distortion compensation pyramid network, parallel feature fusion of multi-size convolution kernels is performed, and a receptive field expansion feature map with cross-scale light intensity distribution features is obtained; then, multi-scale feature extraction is performed on the receptive field expansion feature map according to an encoder in the convolutional neural network, a cascaded channel-priority dual-domain attention mechanism is introduced to perform feature enhancement on the extracted feature vectors, and a multi-scale feature image of the optical communication chip is obtained; then, a multi-level adaptive dilution rate hollow convolution group improved jump connection is used to perform resolution adaptive receptive field matching fusion on the feature maps at different levels of the multi-scale feature image and a decoder; finally, the fused features are input into the decoder for progressive up-sampling processing, and a multi-scale composite loss function of fused pixel level, feature level and frequency domain reconstruction is simultaneously used to optimize the multi-level features output by the decoder, and a clear image of the optical communication chip is reconstructed.

[0065] The convolutional neural network is used for image deblurring processing of the optical communication chip in the embodiment, which can automatically extract high-level semantic features in the image, without the need for manual design of a feature extractor, thereby greatly improving the efficiency and accuracy of image deblurring processing. Moreover, the pre-constructed multi-scale optical distortion compensation pyramid network is used to perform optical feature pre-extraction on the input optical communication chip image, so as to expand the receptive field of cross-scale light intensity distribution. Furthermore, the cascaded channel-priority dual-domain attention mechanism is introduced in the encoder part of the convolutional neural network to screen the extracted features, thereby improving the sub-pixel level detail feature recognition of the optical communication chip image. Finally, the multi-level adaptive dilution rate hollow convolution group is used to enhance the multi-resolution detail feature processing in the jump connection, which can further improve the efficiency and accuracy of deblurring processing.

[0066] For better understanding of the above technical solutions, the exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a clearer, more thorough understanding of the present application and to convey the scope of the present application to those skilled in the art.

[0067] Specifically, referring to Figure 1 the present embodiment proposes a method for deblurring an optical communication chip image based on Figure 2 the convolutional neural network shown in the figure, which can include:

[0068] S100, optical feature pre-extraction is performed on the input optical communication chip image based on the pre-constructed multi-scale optical distortion compensation pyramid network, and a receptive field expansion feature map with cross-scale light intensity distribution features is obtained through parallel feature fusion of multi-size convolution kernels.

[0069] In the present embodiment, a multi-scale optical distortion compensation pyramid network with the structure shown in Figure 3 the figure is used to perform optical feature pre-extraction on the input optical communication chip image: after the RGB picture with the size of H*W*3 is input and the convolution kernels (1*1, 3*3, 5*5, 7*7) of different receptive fields are used to obtain feature maps, the feature maps are up-sampled and fused, and finally the features of different scales are spliced in the channel dimension to output a receptive field expansion feature map with the size of H*W*7. In the present embodiment, the optical feature pre-extraction is performed on the input optical communication chip image based on the pre-constructed multi-scale optical distortion compensation pyramid network, so as to expand the receptive field of the cross-scale light intensity distribution and further improve the extraction efficiency of the feature vector by the encoder in the subsequent convolutional neural network.

[0070] In the present embodiment, before step S100, method steps G100 to G300 can also be performed:

[0071] G100, performing image adjustment and feature enhancement preprocessing on the acquired blurred optical communication chip image data to obtain an image training set.

[0072] Further, in view of the problems of different image blur degrees, small amount of blurred and clear paired data, and various noise interferences, the present embodiment adopts a unique image preprocessing method, which includes a series of image adjustment and feature enhancement image preprocessing operations such as scale transformation, normalization, noise suppression, ROI (region of interest) cropping, and data enhancement, so as to improve the quality of the image, preserve the integrity of the key information as much as possible, and expand the amount of defect data. Specifically, step G100 can include the following sub-steps G110 to G130:

[0073] G110, the acquired optical communication chip image data is processed.

[0074] G120, the image adjustment of the blurred optical communication chip image data is processed by multi-size transformation and normalization, and the normalized image data output according to the set size set is obtained.

[0075] Size transformation: the image size aspect ratio of the optical communication chip is 5:1, and in the process of deep learning training, the length and width of the image will be scaled to the picture of equal proportion size. In order to adapt to this scaling change, the original data is expanded and scaled in this embodiment, and different chip images also have 3:1 and 4:1 cases, so the original image needs to consider these forms when performing size transformation. For example, the image of 1500*300 size, when the size is transformed and expanded, the picture size of the data set is composed of 1200*300, 900*300, 300*900, 512*512, 256*256, 1024*1024, etc.

[0076] Normalization: there are many types of optical communication chips, and the thickness and surface roughness of the chips are different due to unstable chip process in the production process, which leads to changes in color and brightness in imaging. Even the same batch of chips, there are certain differences in imaging. Therefore, the embodiment adopts Z-score normalization technology for image preprocessing, so as to stabilize the training process and improve the adaptability of the deblurring algorithm model to different optical communication chip images. Specifically, the Z-score normalization method includes: first, the image is converted into a gray scale, and then the data points of the gray scale image are converted into a standard normal distribution with a mean of 0 and a standard deviation of 1. The calculation formula (1) is as follows:

[0077]

[0078] In formula (1), X is the original data point, μ is the mean of the gray scale image, and σ1 is the standard deviation of the data set.

[0079] G130, the normalized image data is processed by noise suppression, interested region cropping and data enhancement, and the image training set with enhanced features is obtained.

[0080] Noise suppression: In actual industrial detection scenarios, the optical communication images obtained by the camera have Gaussian noise or Poisson noise to some extent. According to the different degrees of noise, slight Gaussian filtering (σ2 = 0.5-1.0) is performed to reduce noise interference and avoid excessive smoothing to avoid loss of high-frequency details in the noise suppression process. The specific implementation method is: by using the convolution operation of the Gaussian function (normal distribution function) on the image, the smoothing processing of the image is realized. The core of Gaussian filtering is to use the weight of the Gaussian function to weight the average of the pixels in the image, so that the noise in the image is smoothed out, while the edge information of the image is preserved as much as possible. The calculation formula (2) is as follows:

[0081]

[0082] In formula (2), x and y are the coordinates of the pixels in the image, σ2 is the standard deviation of the Gaussian function, which determines the width of the Gaussian kernel and the smoothing degree, and exp is the natural exponential function

[0083] Region of interest cropping: Since the detection area of the optical communication chip is usually concentrated in a specific position, the different areas of the chip are first located by the Canny / Sobel edge detection or template matching method, and the irrelevant other areas are removed to focus on the ROI (region of interest). After cropping, the image is input into the deep learning network to reduce the calculation amount, improve the training efficiency and the robustness of the model. The cropped image can be further resized to a fixed size of 256*256 or 512*512 to adapt to the input requirements of different use scenarios.

[0084] Data enhancement: The enhancement methods for data quality include image contrast adjustment, horizontal edge and vertical edge enhancement superposition, high-frequency information enhancement, low-frequency information reduction, etc. The enhancement methods for data quantity include horizontal flipping, vertical flipping, translation, rotation (non-90 degrees), saturation adjustment, local nonlinear deformation, random partial region filling, etc.

[0085] G200, a multi-scale optical distortion compensation pyramid network is set at the input layer of the original convolutional neural network, and the generated improved convolutional neural network is subjected to model pruning, quantization compression and knowledge distillation lightweight processing.

[0086] Model pruning includes channel pruning and structured pruning. Channel pruning: based on the importance evaluation of convolution channels at different gradients, redundant convolution channels are removed; structured pruning: the key channels of the multi-scale Inception module are preferentially retained to avoid damaging the multi-scale feature fusion.

[0087] Quantization compression: convert FP32 (32-bit floating point) weights to INT8 (8-bit integer), reduce memory occupancy, and improve model inference speed; use Debit quantization, i.e. QAT (quantization-aware training) to reduce precision loss.

[0088] Knowledge distillation: use UNet (convolutional neural network) as a teacher model to guide the model to learn more compact feature representation.

[0089] G300 trains the improved lightweight convolutional neural network according to the image training set, and evaluates the training results in multiple dimensions through the set multi-scale composite loss function, and outputs the optimal convolutional neural network according to the evaluation results.

[0090] Further, step G300 can include the following sub-steps G310 to G150

[0091] G310, calculate the pixel difference between the predicted image and the actual image output by the convolutional neural network using mean square error, and realize multi-scale pixel difference measurement combined with multi-scale content loss function.

[0092] Mean square error is used to measure the pixel difference between the predicted image and the real clear image, which can effectively optimize the overall similarity of the image, and the calculation formula (3) is as follows:

[0093]

[0094] In formula (3), y i is the true value, the predicted value, and n is the sample number.

[0095] The multi-scale content loss function can better capture the details and structural information of the image by calculating the difference between the predicted image and the real image in multiple scales, and the calculation formula (4) is as follows:

[0096]

[0097] In formula (4), and represent the features of the real image and the predicted image in the kth scale respectively, and λ k is the weight coefficient.

[0098] G320, calculate the multi-scale Manhattan distance between the predicted image and the actual image in the same frequency domain through the multi-scale frequency reconstruction loss function.

[0099] The multi-scale frequency reconstruction loss function can effectively improve the reconstruction ability of the model to high-frequency information and reduce the difference in frequency space by measuring the L1 (Manhattan distance) distance between the multi-scale real image and the deblurring image in the frequency domain, and the calculation formula (5) is as follows:

[0100]

[0101] In formula (5), F is a Fourier transform.

[0102] G330, the feature vectors of the predicted image and the actual image at the same position are extracted by using a pre-trained convolutional neural network, and the extracted feature vectors are compared in similarity by using a perception loss function, to obtain the feature similarity of the predicted image and the actual image.

[0103] The perception loss function optimizes the model by comparing the similarity of the predicted image and the actual image in the feature space, and the pre-trained convolutional neural network is used to extract the features in the embodiment, which can better reflect the perception characteristics of the human visual system, and the calculation formula (6) is as follows:

[0104]

[0105] In formula (6), VGG is a pre-trained convolutional neural network.

[0106] G340, the frequency components higher than the set frequency threshold in the frequency spectrum of the predicted image and the actual image are weighted and distributed by using a weighted Fourier frequency domain loss function, to obtain the frequency domain constraint of the convolutional neural network.

[0107] The weighted Fourier frequency domain loss function processes the frequency spectrum after Fourier transform in a weighted manner, so that the pre-trained model pays more attention to high-frequency components, and by designing a weight function based on the distance of the center point, the learning ability of the model to high-frequency information can be enhanced, and the calculation formula (7) is as follows:

[0108]

[0109] In formula (7), W i is a frequency-based weight.

[0110] G350, according to the multi-scale pixel difference value, the multi-scale Manhattan distance, the feature similarity and the frequency domain constraint, the training results are evaluated in multiple dimensions by combining the weight ratio of multi-dimensional loss evaluation, and the optimal convolutional neural network is output according to the evaluation results.

[0111] The combination of mean square error, multi-scale content loss, multi-scale frequency reconstruction loss, perception loss and weighted Fourier frequency domain loss is used to optimize the overall similarity and detail information of the image, which can further improve the generalization ability of the pre-trained model, and the total loss function (8) is as follows:

[0112] L total =a1L MSE +a2L MSC +a3LMSFR +a4L PL +a5L WFL (8)

[0113] In formula (8), a1, a2, a3, a4, and a5 represent weight coefficients of different losses.

[0114] In this embodiment, step S100 can include the following sub-steps S110 to S130:

[0115] S110, using a plurality of different scale convolution branches arranged in parallel, respectively extracting optical feature response maps with different receptive fields in the optical communication chip image, the optical feature response maps including local defect features caused by edge scattering, transmission unevenness, or material property differences of optical structures.

[0116] S120, performing an up-sampling operation on each optical feature response map to the original spatial resolution of the optical communication chip image to obtain feature maps corresponding to different original receptive fields.

[0117] S130, through the parallel features of multi-size convolution kernels, cascading and fusing the feature maps of different original receptive fields after uniform resolution in the channel dimension to obtain a receptive field expansion feature map with cross-scale light intensity distribution features.

[0118] S200, performing multi-scale feature extraction on the receptive field expansion feature map according to the encoder in the convolutional neural network, and introducing a cascaded channel-priority dual-domain attention mechanism to perform feature enhancement on the extracted feature vectors to obtain a multi-scale feature image of the optical communication chip.

[0119] In this embodiment, referring to the cascaded channel-priority dual-domain attention mechanism structure diagram shown in Figure 4 , step S200 can include the following sub-steps S210 to S230:

[0120] S210, using a multi-scale composite loss function that fuses pixel-level, feature-level, and frequency domain reconstruction to optimize the encoder to perform multi-scale feature extraction on the receptive field expansion feature map to obtain an initial multi-scale feature image of the optical communication chip.

[0121] S220, according to the cross-scale light intensity distribution features of the receptive field expansion feature map, combining the channel domain attention module in the dual-domain attention mechanism to dynamically adjust the contribution weights of each feature channel to adaptively compensate for the feature expression deviation caused by uneven illumination in the initial multi-scale feature image.

[0122] Further, referring to the cascaded channel-priority dual-domain attention mechanism structure diagram shown in Figure 5The channel domain attention module of the structure shown dynamically adjusts the response intensity of each feature channel by light intensity gradient guided channel weight mapping, and sends the channel weighted feature vector and the original feature vector to the spatial domain attention module after fusion under the channel attention mechanism.

[0123] S230, based on the waveguide structure information of the optical communication chip, the spatial domain attention module in the dual-domain attention mechanism is used to strengthen the pixel features in the waveguide edge region and the adjacent weak texture region in the multi-scale feature image enhanced by the channel, and obtain a multi-scale feature image with high recognition degree in the spatial dimension.

[0124] Further, with reference to Figure 6 The spatial domain attention module of the structure shown generates a spatial position sensitive attention mask based on the edge gradient distribution characteristics of the distortion region, and strengthens the sub-pixel level detail features of the blurred boundary. First, the spatial dimension weight distribution is performed on the channel fused features to generate spatial weighted features, and then the spatial weighted features and the original input features are fused again to output the final features.

[0125] S300, the improved skip connection of the multi-level adaptive dilation rate group of the hole convolution is used to adaptively match and fuse the features of different levels of the multi-scale feature image with the receptive field of the decoder.

[0126] In this embodiment, through the skip connection, the detail features extracted in the encoder can be directly transmitted to the decoder, thereby making up for the possible information loss in the upsampling process. At the same time, without increasing the size and parameter amount of the convolution kernel, the receptive field is expanded by adjusting the dilation rate to improve the processing efficiency of the multi-scale features such as the microstructure, edge and texture of the optical communication chip, so that the convolution branch can capture both local details and global context information. The skip connection is improved by using the multi-level adaptive dilation rate group of the hole convolution in this embodiment, and the dilation rate of the hole convolution is adaptively adjusted according to the feature distribution (such as target scale and texture complexity) of the feature image of the optical communication chip, so as to obtain the optimal expanded receptive field, prevent too small to capture long distance dependency, and prevent too large to make the sampling points too sparse, resulting in loss of local information. The hole convolution group is used for reference Figure 7 As shown in the figure, (a) represents a common 3*3 convolution, and the receptive field is 3*3. (b) represents the hole convolution with a dilation rate of 2, and the receptive field of each red dot is 3*3, and the total receptive field is 7*7. (c) represents the hole convolution with a dilation rate of 4, and the receptive field of each red dot is 5*5, and the total receptive field is 15*15.

[0127] In this embodiment, step S300 can include the following sub-steps S310-S320:

[0128] S310. In the skip connection path between the decoder and encoder, a multi-level dilated convolutional group with adaptive dilation rate is introduced to improve the skip connection mechanism.

[0129] S320: Through an improved skip connection mechanism, feature vectors at different levels in the multi-scale feature image extracted by the encoder are fused with the decoder through resolution-adaptive receptive field matching to obtain multi-scale fused features that can simultaneously capture local details and global contextual information.

[0130] S400: The fused features are input to the decoder and progressively upsampled. Simultaneously, a multi-scale composite loss function that fuses pixel-level, feature-level, and frequency-domain reconstructions is used to optimize the multi-level features output by the decoder, thereby reconstructing a clear image of the optical communication chip.

[0131] This embodiment proposes an image reconstruction framework for optical communication chips based on multi-scale feature fusion and composite loss optimization. Through a progressive upsampling structure in the decoder, combined with multi-modal monitoring signals at the pixel, feature, and frequency domains, high-fidelity reconstruction of the chip's micro / nano structure is achieved. This method effectively solves the detail blurring problem of traditional super-resolution techniques in high-noise, low-contrast industrial scenarios.

[0132] In one specific embodiment, the optical communication chip image deblurring method described in steps S100 to S400 is used to... Figures 8 to 13 Six sets of optical communication chip images were deblurred. The performance of the deblurring results in terms of PSNR and SSIM was analyzed, and the visual effects were combined to evaluate the deblurring results. Figures 8 to 11 In the image, the left side is the original image, and the right side is the deblurred image. Figures 12 to 13 In the image, the original image is shown at the top, and the image after blurring is shown at the bottom.

[0133] By comparing the visual effects of the six groups of optical communication chip images before and after image processing, it can be clearly seen that there is a significant difference before and after processing. In the image before processing, the details of the optical communication chip are blurred, the overall contrast of the edge and internal structure of the chip is low, which makes it difficult to clearly present the key features and defect content of the chip, and there are many noises in the image, which interfere with the observation of the details of the chip, making it difficult to distinguish the microstructure and features of the chip. In contrast, the image after the deblurring processing of the embodiment method shows a significant improvement. The processed image has rich levels and a significantly improved contrast, making it possible to clearly distinguish each part of the chip, the edge profile of the chip becomes sharp, the internal structure and micro features (such as circuit lines, solder joints, etc.) are clearly visible, and the details are more richly expressed. At the same time, the noise is effectively suppressed, and the high-frequency information (such as fine texture and structure) of the image is in good balance with the noise, and the overall image quality is significantly improved. In addition, the processed image is also more natural and accurate in color representation, which can more truly reflect the actual appearance of the optical communication chip. This improvement in image quality not only helps the visual inspection and quality evaluation of the chip, but also provides a more reliable basis for subsequent image analysis, fault detection and automation detection. Overall, the embodiment method performs well in improving the image quality of the optical communication chip, with significant results, and can provide a more reliable basis for visual inspection, quality evaluation and automation detection of the chip.

[0134] The image quality indicators (PSNR and SSIM) of the six groups of optical communication chip images are calculated, and the specific indicator data shown in Table 1 is obtained, and the convolutional neural network model achieves excellent performance of PSNR>28dB and SSIM>0.9 on most samples, where PSNR>28dB indicates excellent deblurring effect, and SSIM>0.9 indicates excellent structure similarity and complete detail retention.

[0135] Table 1, PSNR and SSIM indicators on different data sets

[0136] Image number PSNR (dB) SSIM Figure 8 28.00 0.9056 Figure 9 32.94 0.9081 Figure 10 31.48 0.9424 Figure 11 27.21 0.8945 Figure 12 29.18 0.9150 Figure 13 29.35 0.9145

[0137] In addition, the embodiment also proposes an image deblurring system for optical communication chips based on a convolutional neural network, comprising:

[0138] An optical feature pre-extraction module is configured to perform optical feature pre-extraction on an input optical communication chip image based on a pre-constructed multi-scale optical distortion compensation pyramid network, and to obtain a receptive field expansion feature map with cross-scale light intensity distribution features through parallel feature fusion of multi-size convolution kernels.

[0139] The encoding module is configured to perform multi-scale feature extraction on the receptive field expansion feature map according to an encoder in the convolutional neural network, and introduce a cascaded channel-priority dual-domain attention mechanism to perform feature enhancement on the extracted feature vector, so as to obtain a multi-scale feature image of the optical communication chip.

[0140] The adaptive receptive field matching module is configured to perform resolution-adaptive receptive field matching fusion of different level feature maps of the multi-scale feature image and the decoder by using a multi-level adaptive dilation rate improved skip connection of a hollow convolution group.

[0141] The decoding module is configured to input the fused feature into the decoder for progressive up-sampling processing, and simultaneously use a multi-scale composite loss function of fused pixel-level, feature-level and frequency domain reconstruction to optimize multi-level features output by the decoder, so as to reconstruct a clear image of the optical communication chip.

[0142] In addition, the embodiment further provides an electronic device, which comprises a processor and a memory storing a program for controlling the processor to perform the image deblurring method based on the convolutional neural network for the optical communication chip.

[0143] Finally, the embodiment further provides a computer readable medium having computer executable instructions stored thereon, and the executable instructions are executed by the processor to implement the image deblurring method based on the convolutional neural network for the optical communication chip.

[0144] To sum up, the embodiment proposes an image deblurring method, system, device and medium based on a convolutional neural network for optical communication chips. Through efficient deblurring processing, the image quality is significantly improved, thereby providing a clearer and more accurate image basis for subsequent defect detection and classification. First, in view of the characteristics of the optical communication chip, such as multiple types, non-fixed aspect ratio, multiple imaging color brightness changes, and uneven distribution of defect quantity, a special image adjustment and feature enhancement preprocessing method is designed to improve the influence of data problems on model accuracy and robustness. Next, an improved multi-scale optical distortion compensation pyramid network is used to extract features from the input image before the encoder of the convolutional neural network, and the receptive field is expanded. In the encoder part of the convolutional neural network, a dual-domain attention mechanism is introduced to filter the extracted features, enhance important features, and transplant unimportant features. At the same time, a multi-level adaptive dilation rate of the empty convolution is used in the skip connection to enhance the multi-resolution detail feature processing. Finally, the clear image of the optical communication chip is reconstructed through the progressive upsampling processing of the decoder. Moreover, the embodiment also proposes five loss combination forms, which are mean square error, multi-scale content loss, multi-scale frequency reconstruction loss, perceptual loss, and weighted Fourier frequency domain loss, to improve the generalization ability of the model. Therefore, the embodiment can effectively distinguish different types of optical communication chip blur, adapt to complex chip image backgrounds, and improve the efficiency and accuracy of deblurring, thereby providing an efficient and reliable solution for optical communication chip quality detection.

[0145] Since the system / device described in the above embodiments of the present application is used for the method of the above embodiments of the present application, the specific structure and modification of the system / device can be understood by those skilled in the art based on the method described in the above embodiments of the present application, and thus will not be described here. Any system / device used in the method of the above embodiments of the present application belongs to the scope of the present application.

[0146] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media containing computer usable program code (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.).

[0147] The present application is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams can be realized by computer program instructions, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams.

[0148] It should be noted that the words "comprise", "comprising", "include", "including" and "includes" when used in this specification are not used in the definite sense; rather, they are used to specify the presence of stated features, integers, steps or components, but do not preclude the presence or addition of one or more other features, integers, steps, components or groups thereof.

[0149] In addition, it should be pointed out that the description in the present specification of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. The illustrative description of the above terms in the present specification does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. Furthermore, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples, without mutual contradiction.

[0150] Although preferred embodiments of the present application have been described, those skilled in the art can make further changes and modifications to these embodiments after learning the basic inventive concept.

[0151] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application.

Claims

1. A method for image deblurring of an optical communication chip based on a convolutional neural network, characterized in that, The application relates to an optical communication chip image processing method based on a pre-constructed multi-scale optical distortion compensation pyramid network. The method comprises the following steps: The input optical communication chip image is subjected to optical feature pre-extraction based on the pre-constructed multi-scale optical distortion compensation pyramid network, and a receptive field expansion feature map with cross-scale light intensity distribution features is obtained through parallel feature fusion of multi-size convolution kernels; Multi-scale feature extraction is performed on the receptive field expansion feature map according to an encoder in the convolutional neural network, and a cascaded channel-priority dual-domain attention mechanism is introduced to perform feature enhancement on the extracted feature vectors, so as to obtain a multi-scale feature image of the optical communication chip; A multi-level adaptive dilated rate improved hollow convolution group is adopted to improve the skip connection, so that the feature maps at different levels of the multi-scale feature image are adaptively matched and fused with the receptive field of the decoder; 2. The method of claim 1, wherein, The fused features are input into the decoder for progressive up-sampling processing, and a multi-scale composite loss function of fused pixel level, feature level and frequency domain reconstruction is simultaneously adopted to optimize the multi-level features output by the decoder, so that a clear image of the optical communication chip is reconstructed. Before the input optical communication chip image is subjected to optical feature pre-extraction based on the pre-constructed multi-scale optical distortion compensation pyramid network, and a receptive field expansion feature map with cross-scale light intensity distribution features is obtained through parallel feature fusion of multi-size convolution kernels, the method further comprises the following steps: The acquired blurred optical communication chip image data is subjected to image adjustment and feature enhancement preprocessing to obtain an image training set; A multi-scale optical distortion compensation pyramid network is arranged at an input layer of an original convolutional neural network, and the improved convolutional neural network is subjected to model pruning, quantization compression and knowledge distillation lightweight processing; 3. The method of claim 2, wherein, The improved convolutional neural network after the lightweight processing is trained according to the image training set, and the training result is evaluated in multiple dimensions through a set multi-scale composite loss function, and the optimal convolutional neural network is output according to the evaluation result. The acquired blurred optical communication chip image data is subjected to image adjustment and feature enhancement preprocessing to obtain an image training set, which comprises the following steps: The acquired optical communication chip image data is subjected to blurring processing; The optical communication chip image data after the blurring processing is subjected to multi-size transformation and normalization processing of image adjustment to obtain normalized image data output in a set size set; 4. The method of claim 2, wherein, The normalized image data is subjected to noise suppression, interested region cropping and data enhancement of feature enhancement to obtain an image training set with enhanced features. The training result is evaluated in multiple dimensions through a set multi-scale composite loss function, and the optimal convolutional neural network is output according to the evaluation result, which comprises the following steps: The mean square error is used to calculate the pixel difference between a predicted image and an actual image output by the convolutional neural network, and a multi-scale pixel difference measurement is realized in combination with a multi-scale content loss function; A multi-scale frequency reconstruction loss function is used to calculate the multi-scale Manhattan distance between the predicted image and the actual image in the same frequency domain; The pre-trained convolutional neural network is used to extract feature vectors of the predicted image and the actual image at the same position, and a perception loss function is used to compare the similarity of the extracted feature vectors, so as to obtain the feature similarity of the predicted image and the actual image. The frequency domain constraint of the convolutional neural network is obtained by using a weighted Fourier frequency domain loss function to perform weighted distribution on frequency components higher than a set frequency threshold in spectral graphs of the predicted image and the actual image; According to the multi-scale pixel difference value, the multi-scale Manhattan distance, the feature similarity, and the frequency domain constraint, a multi-dimensional loss evaluation weight ratio is combined to evaluate the training result in multiple dimensions, and an optimal convolutional neural network is output according to the evaluation result.

5. The method of claim 1, wherein, Based on the pre-constructed multi-scale optical distortion compensation pyramid network, optical feature pre-extraction is performed on the input optical communication chip image, and through parallel feature fusion of multi-size convolution kernels, a receptive field expansion feature map with cross-scale light intensity distribution characteristics is obtained, including: A plurality of different scale convolution branches are arranged in parallel to extract optical feature response maps with different receptive fields in the optical communication chip image, and the optical feature response maps include local defect features caused by optical structure edge scattering, transmission unevenness, or material property differences; Upsampling operations are performed on each optical feature response map to the original spatial resolution of the optical communication chip image to obtain feature maps corresponding to different original receptive fields; Through parallel features of multi-size convolution kernels, the feature maps of different original receptive fields after uniform resolution are cascaded and fused in the channel dimension to obtain a receptive field expansion feature map with cross-scale light intensity distribution characteristics.

6. The method of claim 1, wherein, According to the encoder in the convolutional neural network, multi-scale feature extraction is performed on the receptive field expansion feature map, and a cascaded channel-priority dual-domain attention mechanism is introduced to perform feature enhancement on the extracted feature vectors to obtain a multi-scale feature image of the optical communication chip, including: A multi-scale composite loss function that fuses pixel-level, feature-level, and frequency domain reconstruction is used to optimize the encoder to perform multi-scale feature extraction on the receptive field expansion feature map to obtain an initial multi-scale feature image of the optical communication chip; According to the cross-scale light intensity distribution characteristics of the receptive field expansion feature map, the contribution weights of each feature channel are dynamically adjusted by the channel domain attention module in the dual-domain attention mechanism to adaptively compensate for the feature expression deviation caused by uneven illumination in the initial multi-scale feature image; Based on the waveguide structure information of the optical communication chip, the spatial domain attention module in the dual-domain attention mechanism is used to strengthen the pixel features in the waveguide edge region and the adjacent weak texture region in the multi-scale feature image after channel enhancement to obtain a multi-scale feature image with strong spatial recognition of the strengthened waveguide structure. The channel domain attention module dynamically adjusts the saliency response strength of each feature channel through light intensity gradient guided channel weight mapping; the spatial domain attention module generates a spatial position sensitive attention mask based on the edge gradient distribution characteristics of the distortion region to strengthen the sub-pixel level detail features of the fuzzy boundary.

7. The method of claim 1, wherein, A multi-level adaptive dilation rate improved hole convolution group is used to improve the skip connection, and different level feature maps of the multi-scale feature image are adaptively matched and fused with the decoder in the receptive field, including: In the decoder and encoder skip connection path, a multi-level adaptive dilation rate hole convolution group is introduced to improve the skip connection mechanism. Through the improved skip connection mechanism, the feature vectors of different levels in the multi-scale feature image extracted by the encoder are adaptively matched and fused with the resolution of the receptive field of the decoder, so that multi-scale fusion features capable of capturing local details and global context information are obtained.

8. An image deblurring system for optical communication chips based on convolutional neural networks, characterized in that, The method comprises the steps of: An optical feature pre-extraction module is configured to pre-extract optical features from an input optical communication chip image based on a pre-constructed multi-scale optical distortion compensation pyramid network, and to obtain receptive field expansion feature maps with cross-scale light intensity distribution features through parallel feature fusion of multi-size convolution kernels; An encoding module is configured to extract multi-scale features from the receptive field expansion feature maps by an encoder in a convolutional neural network, and to enhance the extracted feature vectors by introducing a cascaded channel-priority dual-domain attention mechanism, so as to obtain multi-scale feature images of the optical communication chip; An adaptive receptive field matching module is configured to improve the skip connection by using a multi-level adaptive expansion rate of a group of dilated convolutions, and to adaptively match and fuse the feature maps of different levels of the multi-scale feature images with the resolution of the receptive field of the decoder; A decoding module is configured to input the fused features into the decoder for progressive up-sampling processing, and to simultaneously optimize the multi-level features output by the decoder by using a multi-scale composite loss function of fused pixel-level, feature-level and frequency domain reconstruction, so as to reconstruct a clear image of the optical communication chip.

9. An electronic device, comprising: The method comprises the steps of: A processor; A memory is configured to store executable instructions for controlling the processor to perform the steps of the image deblurring method of the optical communication chip based on the convolutional neural network according to any one of claims 1 to 7.

10. A computer readable medium having stored thereon computer- executable instructions, characterized in that, The executable instructions are executed by the processor to perform the steps of the image deblurring method of the optical communication chip based on the convolutional neural network according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Attention mechanism-based image blind deblurring method and system

    CN111709895A

  • Image super-resolution method based on pyramid fusion attention network

    CN114926343A

  • Image deblurring method based on multi-scale attention feature fusion

    CN117011184A

  • Multi-scale remote sensing image target detection method based on enhanced small target feature extraction

    CN117809200A

  • Road extraction method and device based on cavity pyramid and strip attention

    CN118982757A