Image processing method, electronic equipment and computer readable medium
By using an end-to-end neural network image processing model and multi-scale feature extraction and fusion techniques, Newton's rings are accurately removed, solving the problem of image detail loss in existing technologies and achieving high-quality Newton's ring removal results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIHAO TECHNOLOGY (TIANJIN) CO LTD
- Filing Date
- 2025-08-26
- Publication Date
- 2026-05-15
AI Technical Summary
Existing techniques often result in loss of image detail when removing Newton's rings, and are difficult to completely remove when Newton's rings are mixed with image content or have complex shapes, leading to low image quality.
An end-to-end neural network image processing model is adopted, which includes an encoder and a decoder. The encoder includes a multi-scale feature extraction module and a multi-scale feature fusion module. Through multi-scale feature extraction and fusion, Newton's rings are accurately located and removed while preserving image details.
It improves the accuracy and comprehensiveness of Newton's rings removal, enhances image quality, avoids image blurring and artifacts, and preserves original details.
Smart Images

Figure CN122048738A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to image processing methods, electronic devices, and computer-readable media. Background Technology
[0002] In the field of optical imaging, Newton's rings are a common interference phenomenon, often caused by factors such as thin films and air gaps in the imaging system. They appear as a series of alternating bright and dark circular fringes. Newton's rings can significantly affect image quality.
[0003] In existing technologies, traditional visual processing methods such as frequency domain filtering, spatial domain filtering, edge detection, and morphological processing can be used to remove Newton's rings from images. Alternatively, neural networks can be introduced to predict Newton's ring parameters to assist traditional algorithms in removing Newton's rings from images. However, these methods tend to lead to loss of image details while removing Newton's rings, and Newton's rings are prone to remain when they are severely mixed with image content or have complex and varied shapes, resulting in low image quality. Summary of the Invention
[0004] This application provides an image processing method, electronic device, and computer-readable medium that improve the accuracy and comprehensiveness of Newton's rings removal from images, thereby improving image quality.
[0005] In a first aspect, embodiments of this application provide an image processing method, which includes: acquiring an image to be processed containing Newton's rings interference; inputting the image to be processed into a pre-trained image processing model to obtain a target image with the Newton's rings interference removed, wherein the image processing model includes an encoder and a decoder, and the encoder includes a multi-scale feature extraction module and a multi-scale feature fusion module.
[0006] In a second aspect, embodiments of this application provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in the first aspect.
[0007] Thirdly, embodiments of this application provide a computer-readable medium having a computer program stored thereon that, when executed by a processor, implements the method described in the first aspect.
[0008] Fourthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0009] The image processing method, electronic device, and computer-readable medium provided in this application first acquire an image to be processed containing Newton's rings interference; then, the image to be processed is input into a pre-trained image processing model to obtain a target image with Newton's rings interference removed. The image processing model includes an encoder and a decoder. The encoder includes a multi-scale feature extraction module and a multi-scale feature fusion module. The structure of the encoder and decoder enables the image processing model to achieve end-to-end Newton's ring removal. The multi-scale feature extraction module and multi-scale feature fusion module in the encoder allow the image processing process to simultaneously focus on the overall image and local details, thereby comprehensively utilizing rich feature information to more accurately locate and remove Newton's rings while effectively preserving detailed information in the image. This overcomes the shortcomings of traditional methods that cause image blurring and artifacts due to Newton's ring removal, improving the accuracy and comprehensiveness of Newton's ring removal from images, and thus enhancing the quality of the target image. Attached Figure Description
[0010] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0011] Figure 1 This is a flowchart of an embodiment of the image processing method according to this application;
[0012] Figure 2 This is a schematic diagram of the structure of the image processing model in the image processing method of this application;
[0013] Figure 3 This is a schematic diagram of the structure of an embodiment of the image processing apparatus according to this application;
[0014] Figure 4 This is a schematic diagram of the structure of an electronic device used to implement the embodiments of this application. Detailed Implementation
[0015] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0016] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0017] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.
[0018] In recent years, significant progress has been made in research on technologies based on artificial intelligence, such as computer vision, deep learning, machine learning, image processing, and image recognition. Artificial intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies, and application systems to simulate and extend human intelligence. AI is a comprehensive discipline involving numerous technologies, including chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, and neural networks. Computer vision, as an important branch of AI, specifically enables machines to recognize the world. Computer vision technologies typically include face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, object detection, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, and robot navigation and localization. With the research and advancement of artificial intelligence technology, this technology has been applied in numerous fields, such as security, urban management, traffic management, building management, park management, facial recognition access control, facial recognition attendance, logistics management, warehouse management, robotics, intelligent marketing, computational photography, mobile imaging, cloud services, smart homes, wearable devices, autonomous driving, autonomous driving, smart healthcare, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile internet, live streaming, beautification, makeup, medical aesthetics, and intelligent temperature measurement.
[0019] In the field of image processing, Newton's rings are a common interference phenomenon, often caused by factors such as thin films and air gaps in imaging systems, appearing as a series of alternating bright and dark circular fringes. Newton's rings can significantly impact image quality. Currently, traditional visual processing methods such as frequency domain filtering, spatial domain filtering, edge detection, and morphological processing are commonly used to remove Newton's rings from images, or neural networks are introduced to predict Newton's ring parameters to assist traditional algorithms in the removal process. These methods, while removing Newton's rings, easily lead to loss of image details, and when Newton's rings are severely mixed with image content or have complex and varied shapes, residual Newton's rings are prone to appear, resulting in low image quality. This application provides an image processing method that can overcome the above-mentioned shortcomings and improve image quality.
[0020] Please refer to Figure 1 The diagram illustrates a flow 100 of an embodiment of an image processing method according to this application. The image processing method includes the following steps:
[0021] Step 101: Obtain the image to be processed containing Newton's rings interference.
[0022] In this embodiment, the image to be processed can be any image captured by the imaging device, including but not limited to fingerprint images, palm print images, and face images. The image may contain Newton's rings interference. Newton's rings are concentric circular fringes formed by thin-film interference in optical imaging, manifesting as periodic noise superimposed on the original image.
[0023] In practice, images to be processed can be acquired in various ways. For example, optical imaging devices such as cameras and microscopes can be used to capture images in real time. Alternatively, images that have not yet undergone Newton's ring removal can be read from already captured images. Furthermore, images containing Newton's ring interference can be received from remote devices via network transmission.
[0024] Step 102: Input the image to be processed into a pre-trained image processing model to obtain a target image with Newton's rings interference removed. The image processing model includes an encoder and a decoder. The encoder includes a multi-scale feature extraction module and a multi-scale feature fusion module.
[0025] In this embodiment, the image processing model can be an end-to-end neural network with the function of removing Newton's rings interference from images, which can be pre-trained using machine learning methods. Through pre-training, the image processing model can learn the mapping relationship between images containing Newton's rings interference and images without Newton's rings interference.
[0026] During image processing model training, image pairs from different devices and shooting scenarios can be used for training. This allows the model to learn the feature changes of Newton's rings and the differences in images under different conditions, thereby possessing a certain generalization ability. Each image pair can include an image with Newton's rings interference and an image without Newton's rings interference, and both should be acquired by the same device in the same shooting scenario to ensure the accuracy of the training samples.
[0027] In this embodiment, the image processing model may include an encoder and a decoder, see [link to documentation]. Figure 2 The encoder is the front-end structure of a neural network, used to map the input image to a low-dimensional feature representation through operations such as convolution and downsampling. The decoder is the back-end structure of the neural network, used to reconstruct the target image from the features extracted by the encoder through operations such as upsampling and convolution. For example, the decoder can be a decoder from the U-net architecture.
[0028] The encoder may further include a multi-scale feature extraction module and a multi-scale feature fusion module. The multi-scale feature extraction module can capture features at different scales of the image to be processed through convolutional layers with different receptive fields. For example, it can extract features expressing the overall structure and low-frequency information of the image, features expressing local regional characteristics, and features expressing high-frequency details. Features at different scales can be represented by feature maps. For example, if the image to be processed has a width of w and a height of h, the multi-scale feature extraction module can extract feature maps with scales of w / 2×h / 2, w / 4×h / 4, w / 8×h / 8, and w / 16×h / 16. The smaller the feature scale, the lower the resolution of the feature map, but the more abstract the semantic information.
[0029] The multi-scale feature fusion module is used to spatially align and fuse features at different scales to generate fused features. For example, feature maps at scales of w / 2×h / 2, w / 4×h / 4, w / 8×h / 8, and w / 16×h / 16 are transformed to the same scale, such as w / 16×h / 16, through convolution, and then concatenated along the channel dimension to achieve feature fusion. Through feature fusion, features at different scales can be complementary and correlated, enabling image processing models to comprehensively utilize these features for subsequent image processing, improving the accuracy and comprehensiveness of Newton's rings removal, thereby enhancing the quality of the generated target image.
[0030] In this embodiment, after the image to be processed containing Newton's rings interference is input into the image processing model, feature processing can be performed sequentially through the multi-scale feature extraction module, the multi-scale feature fusion module and the decoder in the encoder, and finally the target image with Newton's rings interference removed is output.
[0031] The method provided in the above embodiments of this application first acquires an image to be processed containing Newton's rings interference; then, the image to be processed is input into a pre-trained image processing model to obtain a target image with Newton's rings interference removed. The image processing model includes an encoder and a decoder. The encoder includes a multi-scale feature extraction module and a multi-scale feature fusion module. The structure of the encoder and decoder enables the image processing model to achieve end-to-end Newton's ring removal. The multi-scale feature extraction module and multi-scale feature fusion module in the encoder allow the image processing process to simultaneously focus on the overall image and local details, thereby comprehensively utilizing rich feature information to more accurately locate and remove Newton's rings while effectively preserving detailed information in the image. This overcomes the shortcomings of traditional methods that cause image blurring and artifacts due to Newton's ring removal, improving the accuracy and comprehensiveness of Newton's ring removal from images, and thus enhancing the quality of the target image.
[0032] In some alternative embodiments, step 102 may be performed as follows:
[0033] Sub-step 1021: Input the image to be processed into the multi-scale feature extraction module to obtain the multi-scale features of the image to be processed.
[0034] Multi-scale features can be represented using feature maps of different sizes. A multi-scale feature extraction module can contain multiple convolutional and pooling layers. The convolutional layers use kernels of different sizes to perform convolution operations on the image to extract features. The pooling layers downsample the feature maps, reducing their resolution and thus extracting features at a larger scale. Extracting features from the image at different scales allows the model to capture the feature representation of Newton's rings at different levels.
[0035] Optionally, see Figure 2 As shown, the multi-scale feature extraction module may include multiple feature extraction sub-modules, and each feature extraction sub-module may include convolutional layers, channel attention layers, spatial attention layers, and pooling layers. The convolutional layers in each feature extraction sub-module may include one or more.
[0036] Convolutional layers are used to perform convolution operations on images to extract image features. Pooling layers can be used to downsample features, reducing the resolution of the feature map. Channel attention layers can use a channel attention mechanism to adjust the weights of the feature map along the channel dimension to emphasize channels that are more important to the current task. Spatial attention layers can use a spatial attention mechanism to weight the feature map along the spatial dimension, highlighting important regions in the feature map, such as the Newton's rings region.
[0037] By leveraging the synergistic effect of convolutional layers, channel attention layers, spatial attention layers, and pooling layers in the multi-scale feature extraction module, image features can be extracted from different scales and levels. This enables the image processing model to more accurately identify and remove Newton's rings while preserving image details. The introduction of channel attention layers and spatial attention layers further enhances the model's focus on Newton's ring features, avoiding over-suppression of non-Newton's ring regions due to indiscriminate processing of all feature channels and all regions in the feature map. This improves the recognition accuracy of Newton's ring regions, prevents erroneous removal of background details, and enhances the removal effect of Newton's rings.
[0038] Sub-step 1022: Input the image to be processed and the multi-scale features into the multi-scale feature fusion module to obtain the fused features.
[0039] The multi-scale feature fusion module can include multiple convolutional layers, allowing feature maps of different sizes to be input into different convolutional layers and converted into feature maps of the same size. These feature maps of the same size can then be concatenated along the channel dimension to obtain fused features. The fused features contain rich feature information about Newton's rings in the image, providing a solid foundation for subsequent removal operations.
[0040] Sub-step 1023: Input the fused features into the decoder to obtain the target image.
[0041] The decoder consists of multiple deconvolutional layers and upsampling layers. Its function is to progressively restore the fused low-resolution feature map to the same size as the image to be processed. During the upsampling process, the resolution of the feature map can be increased through methods such as interpolation, and further processing and refinement of the features can be achieved by combining deconvolution operations. Finally, the decoder outputs a target image with Newton's rings removed. This target image visually removes Newton's rings, restores the original details and structure of the image, and improves the image's clarity and naturalness, making it suitable for various subsequent image analysis and application tasks.
[0042] Since the multi-scale feature extraction module, the multi-scale feature fusion module, and the decoder are capable of extracting multi-scale features from the image to be processed, fusing multi-scale features, and reconstructing the image, integrating these network modules into an image processing model can ensure that the image processing model can efficiently and accurately complete the Newton's ring removal task in practical applications.
[0043] In some alternative embodiments, the image processing model can be trained through the following steps:
[0044] Step 201: Obtain a first sample set. The first sample set includes first image pairs acquired by multiple shooting devices in multiple shooting scenarios. Each first image pair includes a first sample image with Newton's ring interference and a second sample image without Newton's ring interference. The first sample image and the second sample image in the same first image pair are acquired by the same shooting device in the same shooting scenario.
[0045] The first sample set may include a large number of first image pairs. These first image pairs can be acquired using various devices in various shooting scenarios. In different shooting scenarios, at least one of the following is different: shooting environment, subject, shooting angle, shooting parameters, etc.
[0046] Each first image pair may include a first sample image containing Newton's rings interference and a second sample image corresponding to the first sample image without Newton's rings interference. In the same first image pair, the first sample image and the second sample image can be acquired by the same device in the same shooting scene, that is, the shooting device, shooting environment, shooting object, shooting angle, shooting parameters, etc. are all the same.
[0047] Step 202: Input the first sample image into the end-to-end neural network to obtain the predicted image.
[0048] Step 203: Determine the total loss value of the end-to-end neural network based on the second sample image and the predicted image.
[0049] The loss value is the result of the loss function, a non-negative real-valued function that characterizes the difference between the predicted image and the second sample image. Generally, the smaller the loss value, the better the robustness of the model. The loss function can be set according to actual needs.
[0050] Step 204: Based on the total loss value, update the parameters of the end-to-end neural network to obtain the image processing model.
[0051] Specifically, the gradient of the loss value with respect to the model parameters can be obtained using the backpropagation algorithm, and then the model parameters can be updated based on the gradient using the gradient descent algorithm. In practice, the backpropagation algorithm described above is also called the error backpropagation (BP) algorithm, or error inverse propagation algorithm, and it is a learning algorithm suitable for multi-layer neural networks. During the backpropagation process, the partial derivatives of the loss function with respect to the weights of each neuron can be calculated layer by layer, forming the gradient of the loss function with respect to the weight vector, which serves as the basis for modifying the weights. The gradient descent algorithm described above is a commonly used method in the field of machine learning for solving model parameters. When solving for the minimum value of the loss function, the neuron weights can be adjusted based on the calculated gradient using the gradient descent algorithm.
[0052] By acquiring a large number of paired samples and inputting them into an end-to-end neural network, the image processing model can learn the mapping relationship between images with Newton's rings and images without Newton's rings under different devices and shooting scenarios. This makes the image processing model highly adaptable to multiple devices and shooting scenarios, and improves the generalization ability of the image processing model.
[0053] In some optional embodiments, step 203 above may further include the following sub-steps:
[0054] Sub-step 2031: Based on the mean squared error loss function, the second sample image, and the predicted image, determine the first loss value of the end-to-end neural network.
[0055] The Mean Squared Error Loss (MSE) function calculates the squared mean of the pixel-level differences between the predicted and ground truth images, measuring the pixel differences between the predicted and sample images. The first loss value can be obtained by substituting the pixel values of each pixel in the second sample image and the predicted image into the MSE function.
[0056] Sub-step 2032: Based on the structural similarity loss function, the second sample image, and the predicted image, determine the second loss value of the end-to-end neural network.
[0057] The Structural Similarity Loss (SSIM) function measures the similarity between two images in terms of brightness, contrast, and structure. Specifically, the mean, standard deviation, and covariance of the second sample image and the predicted image are first calculated. Then, these values are substituted into the SSIM function to obtain the second loss value.
[0058] Sub-step 2033: Input the second sample image and the predicted image into the feature extraction network to obtain the first feature map and the second feature map.
[0059] The feature extraction network can be configured based on the training task. For example, if the training task is face recognition, the feature extraction network in a face feature extractor or face classifier can be used to extract the first feature map of the second sample image and the second feature map of the predicted image. If the training task is object detection, the feature extraction network in an object detection model can be used to extract the first feature map of the second sample image and the second feature map of the predicted image.
[0060] Sub-step 2034: Based on the perceptual loss function, the first feature map, and the second feature map, determine the third loss value of the end-to-end neural network.
[0061] The perceptual loss function can be the mean squared error loss function mentioned above, or it can be the half-mean squared error (Half-MSE) loss function, Huber loss function, etc., without specific limitations here. The pixel values of each pixel in the first and second feature maps can be substituted into the perceptual loss function to obtain the third loss value.
[0062] Sub-step 2035: Determine the total loss value of the end-to-end neural network based on the first loss value, the second loss value, and the third loss value.
[0063] Here, the first loss value, the second loss value, and the third loss value can be weighted and summed to obtain the total loss value of the end neural network.
[0064] Since mean squared error loss focuses on pixel-level differences, structural similarity loss emphasizes the similarity of image structural information, and perceptual loss considers the differences in high-level semantic features, the above multi-dimensional loss evaluation methods enable image processing models to accurately align pixels while maintaining the consistency of image structure and semantic information during image processing. This allows for better preservation of image details and guarantee of image quality while removing Newton's rings.
[0065] In some optional embodiments, sub-step 2035 may further include the following steps:
[0066] The first step is to identify the Newton's rings region and the non-Newton's rings region in the first sample image. The Newton's rings region is the area in the image that exhibits a concentric ring-like feature with alternating light and dark areas, while the non-Newton's rings region is any other area in the first sample image besides the Newton's rings region.
[0067] In practice, the Newton's rings region can be determined in several ways. As an example, the region can be pre-annotated manually, and the annotation information can be read to determine the Newton's rings region in the first sample image. As another example, the region can be determined using traditional visual processing methods such as edge detection and morphological processing.
[0068] The second step involves determining the fourth loss value of the end-to-end neural network based on the region loss function, the second sample image, the predicted image, the first weight of the Newton's ring region, and the second weight of the non-Newton's ring region, with the first weight being greater than the second weight.
[0069] Specifically, the region loss function is used to calculate the weighted loss between the Newton's rings region and the non-Newton's rings region, and the mean squared error loss function can be used. In practice, the region corresponding to the Newton's rings region in the second sample image can be first designated as the first region, the region corresponding to the non-Newton's rings region in the second sample image can be designated as the second region, the region corresponding to the Newton's rings region in the predicted image can be designated as the third region, and the region corresponding to the non-Newton's rings region in the predicted image can be designated as the fourth region. Then, the pixel value of each pixel in the first region is multiplied by the first weight, and the pixel value of each pixel in the second region is multiplied by the second weight to obtain the second sample image with updated pixel values. Similarly, the pixel value of each pixel in the third region is multiplied by the first weight, and the pixel value of each pixel in the fourth region is multiplied by the second weight to obtain the predicted image with updated pixel values. Finally, the pixel values of each pixel in the second sample image with updated pixel values and the pixel values of each pixel in the predicted image with updated pixel values are input into the region loss function to obtain the fourth loss value.
[0070] The third step is to determine the total loss value of the end-to-end neural network based on the first, second, third, and fourth loss values.
[0071] Here, the first loss value, the second loss value, and the third loss value can be weighted and summed to obtain the total loss value of the end neural network.
[0072] By assigning higher weights to the Newton's rings region, the end-to-end neural network can focus on processing this region. Introducing a region loss function allows the total loss value to more comprehensively and accurately reflect the model's processing effect on the Newton's rings region and the overall image quality, thereby guiding the end-to-end neural network to accurately remove the Newton's rings in this region, improving both the removal effect and image quality.
[0073] In some alternative embodiments, the Newton's rings region and the non-Newton's rings region may be determined by the following steps:
[0074] The first step is to perform a frequency domain transformation on the first sample image to obtain the first frequency domain image. This frequency domain transformation converts the image from the spatial domain to the frequency domain. Methods such as Fourier transform can be used to perform the frequency domain transformation on the first sample image to obtain the first frequency domain image.
[0075] The second step is to determine the target region in the frequency domain map based on the mask image. The mask image is a binary image of the same size as the first frequency domain map, used to mark the target region. The target region can be a circular region to preserve frequencies within a specific radius.
[0076] The third step is to perform filtering or frequency domain enhancement processing on the target region to obtain the second frequency domain map.
[0077] As an example, a bandpass filter can be used to bandpass filter the target area, preserving information within a specific frequency range to retain the mid-frequency components corresponding to Newton's rings while removing low-frequency background information and high-frequency noise. A bandpass filter is a filter that allows signals within a specific frequency range to pass through while attenuating or suppressing signals of other frequencies. This specific frequency range can be obtained in advance through spectral analysis of the image containing Newton's ring interference.
[0078] As another example, a high-pass filter can be first used to filter the target region to remove low-frequency information, which typically corresponds to the background region. Then, a low-pass filter is used to filter the target region after removing the low-frequency information to further remove high-frequency noise, which typically corresponds to image sensor noise, fine particle noise, etc. This achieves the effect of preserving the mid-frequency components of the Newton's rings region while removing low-frequency background information and high-frequency noise.
[0079] As another example, frequency domain enhancement can be performed on the target region. Frequency domain enhancement refers to the operation of amplifying or suppressing specific frequency components of an image in the frequency domain to highlight or weaken certain features.
[0080] The fourth step is to determine the Newton's rings region based on the second frequency domain map, and to determine the region in the first sample image other than the Newton's rings region as the non-Newton's rings region.
[0081] Specifically, an inverse Fourier transform can be performed on the frequency domain image to obtain the processed spatial domain image. In this spatial domain image, the target region has been freed from background and high-frequency noise, resulting in a clearer and more distinct display of the Newton's rings. Based on this spatial domain image, the location and contour of the Newton's rings region can be accurately determined. Based on this location and contour, the Newton's rings region and non-Newton's rings regions in the first sample image can be identified.
[0082] By converting the first sample image to the frequency domain and then performing filtering or frequency domain enhancement processing on the target region, the Newton's ring region can be further highlighted. This frequency domain-based analysis method avoids interference from complex backgrounds in the spatial domain, improves the accuracy of Newton's ring region detection, and provides a more reliable basis for subsequent loss calculation and model optimization.
[0083] In some optional embodiments, the image to be processed can be acquired by a target capturing device or in a target capturing scene. The target capturing device can be any image acquisition device, such as a newly manufactured mobile phone, camera, etc. The target capturing scene can be any specified capturing scene. The image processing model is obtained through training as follows:
[0084] Step 301: Train the end-to-end neural network based on the first sample set to obtain a pre-trained model. The first sample set includes first image pairs acquired by multiple shooting devices in multiple shooting scenarios. Each first image pair includes a first sample image with Newton's ring interference and a second sample image without Newton's ring interference. The first sample image and the second sample image in the same first image pair are acquired by the same shooting device in the same shooting scenario.
[0085] Step 302: Fine-tune the pre-trained model based on the second sample set to obtain the image processing model. The second sample set includes second image pairs acquired by the target shooting device or acquired in the target shooting scene. Each second image pair includes a third sample image with Newton's ring interference and a fourth sample image without Newton's ring interference.
[0086] The training and fine-tuning processes of the pre-trained model can be found in the training process of the end-to-end neural network described above, and will not be repeated here. It should be noted that when the target imaging device acquires the fourth sample image, interfering components that may generate Newton's rings can be removed. For example, the interfering component could be the screen of an electronic device, i.e., a display screen, and the target imaging device could be a camera located below the screen of the electronic device. Alternatively, the target imaging device can be modified to have two cameras. One camera is covered by the screen and is used to acquire the third sample image when photographing the target object above the screen; the other camera is not covered by the screen and is used to acquire the fourth sample image.
[0087] The pre-training process enables the model to learn the diverse characteristics of Newton's rings under different devices and shooting scenarios, giving it good generalization ability. The fine-tuning process further utilizes specific data from the target shooting device or scene to optimize the model parameters, adapting the model to the imaging characteristics of the target device or the Newton's rings features of the target scene. This two-stage training strategy allows the image processing model to maintain its generalization ability while more accurately removing Newton's rings interference under the target shooting device or scene, thus improving image quality.
[0088] Further reference Figure 3 As an implementation of the methods shown in the above figures, this application provides an embodiment of an image processing apparatus, which is similar to... Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0089] like Figure 3 As shown, the image processing device 300 of this embodiment includes: an acquisition unit 301, used to acquire an image to be processed containing Newton's rings interference; and a processing unit 302, used to input the image to be processed into a pre-trained image processing model to obtain a target image with the Newton's rings interference removed. The image processing model includes an encoder and a decoder, and the encoder includes a multi-scale feature extraction module and a multi-scale feature fusion module.
[0090] In some optional implementations of this embodiment, the processing unit 302 is further configured to: input the image to be processed into the multi-scale feature extraction module to obtain multi-scale features of the image to be processed; input the image to be processed and the multi-scale features into the multi-scale feature fusion module to obtain fused features; and input the fused features into the decoder to obtain the target image.
[0091] In some optional implementations of this embodiment, the multi-scale feature extraction module includes multiple feature extraction sub-modules, each of which includes a convolutional layer, a channel attention layer, a spatial attention layer, and a pooling layer.
[0092] In some optional implementations of this embodiment, the image processing model is trained as follows: A first sample set is obtained, comprising first image pairs acquired by multiple imaging devices in multiple shooting scenarios. Each first image pair includes a first sample image with Newton's ring interference and a second sample image without Newton's ring interference. The first sample image and the second sample image in the same first image pair are acquired by the same imaging device in the same shooting scenario. The first sample image is input into an end-to-end neural network to obtain a predicted image. Based on the second sample image and the predicted image, the total loss value of the end-to-end neural network is determined. Based on the total loss value, the parameters of the end-to-end neural network are updated to obtain the image processing model.
[0093] In some optional implementations of this embodiment, determining the total loss value of the end-to-end neural network based on the second sample image and the predicted image includes: determining a first loss value of the end-to-end neural network based on the mean squared error loss function, the second sample image, and the predicted image; determining a second loss value of the end-to-end neural network based on the structural similarity loss function, the second sample image, and the predicted image; inputting the second sample image and the predicted image into a feature extraction network to obtain a first feature map and a second feature map, respectively; determining a third loss value of the end-to-end neural network based on the perceptual loss function, the first feature map, and the second feature map; and determining the total loss value of the end-to-end neural network based on the first loss value, the second loss value, and the third loss value.
[0094] In some optional implementations of this embodiment, determining the total loss value of the end-to-end neural network based on the first loss value, the second loss value, and the third loss value further includes: determining the Newton's ring region and the non-Newton's ring region in the first sample image; determining a fourth loss value of the end-to-end neural network based on the region loss function, the second sample image, the predicted image, the first weight of the Newton's ring region, and the second weight of the non-Newton's ring region, wherein the first weight is greater than the second weight; and determining the total loss value of the end-to-end neural network based on the first loss value, the second loss value, the third loss value, and the fourth loss value.
[0095] In some optional implementations of this embodiment, determining the Newton's ring region and non-Newton's ring region in the first sample image includes: performing a frequency domain transformation on the first sample image to obtain a first frequency domain map; determining the target region in the frequency domain map based on a mask image; performing filtering or frequency domain enhancement processing on the target region to obtain a second frequency domain map; determining the Newton's ring region based on the second frequency domain map, and determining the region in the first sample image other than the Newton's ring region as the non-Newton's ring region.
[0096] In some optional implementations of this embodiment, the image to be processed is acquired by a target shooting device or in a target shooting scene; the image processing model is trained through the following steps: training an end-to-end neural network based on a first sample set to obtain a pre-trained model, wherein the first sample set includes first image pairs acquired by multiple shooting devices in multiple shooting scenes, each first image pair includes a first sample image with Newton's ring interference and a second sample image without Newton's ring interference, and the first sample image and the second sample image in the same first image pair are acquired by the same shooting device in the same shooting scene; fine-tuning the pre-trained model based on a second sample set to obtain the image processing model, wherein the second sample set includes second image pairs acquired by the target shooting device or in the target shooting scene, each second image pair includes a third sample image with Newton's ring interference and a fourth sample image without Newton's ring interference.
[0097] The apparatus provided in the above embodiments of this application first acquires an image to be processed containing Newton's rings interference; then, it inputs the image to be processed into a pre-trained image processing model to obtain a target image with Newton's rings interference removed. The image processing model includes an encoder and a decoder. The encoder includes a multi-scale feature extraction module and a multi-scale feature fusion module. The structure of the encoder and decoder enables the image processing model to achieve end-to-end Newton's ring removal. The multi-scale feature extraction module and the multi-scale feature fusion module in the encoder allow the image processing process to simultaneously focus on both the overall image and local details, thereby comprehensively utilizing rich feature information to more accurately locate and remove Newton's rings while effectively preserving detailed information in the image. This overcomes the shortcomings of traditional methods that cause image blurring and artifacts due to Newton's ring removal, improving the accuracy and comprehensiveness of Newton's ring removal from images, and thus enhancing the quality of the target image.
[0098] This application also provides an electronic device, including one or more processors and a storage device storing one or more programs thereon. When the one or more programs are executed by the one or more processors, the one or more processors implement the above-described image processing method.
[0099] The following is for reference. Figure 4It shows a schematic diagram of the structure of an electronic device used to implement some embodiments of this application. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this application.
[0100] like Figure 4 As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. Processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0101] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, disks, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 4 Each box shown can represent a device or multiple devices as needed.
[0102] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described image processing method.
[0103] In particular, according to some embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined above in the methods of some embodiments of this application.
[0104] This application also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the above-described image processing method.
[0105] It should be noted that the computer-readable medium described in some embodiments of this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0106] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0107] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire an image to be processed containing Newton's rings interference; input the image to be processed into a pre-trained image processing model to obtain a target image with Newton's rings interference removed, wherein the image processing model includes an encoder and a decoder, and the encoder includes a multi-scale feature extraction module and a multi-scale feature fusion module.
[0108] Computer program code for performing operations of some embodiments of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++; and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, or it can be connected to an external computer (e.g., via the Internet using an Internet service provider), including local area networks (LANs) or wide area networks (WANs).
[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0110] The units described in some embodiments of this application can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a first determining unit, a second determining unit, a selecting unit, and a third determining unit. The names of these units do not necessarily limit the specific unit itself.
[0111] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0112] The above description is merely a selection of preferred embodiments of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this application.
Claims
1. An image processing method, characterized in that, The method includes: Obtain the image to be processed that contains Newton's rings interference; The image to be processed is input into a pre-trained image processing model to obtain a target image with the Newton's rings interference removed. The image processing model includes an encoder and a decoder. The encoder includes a multi-scale feature extraction module and a multi-scale feature fusion module.
2. The method according to claim 1, characterized in that, The step of inputting the image to be processed into a pre-trained image processing model to obtain a target image with the Newton's rings interference removed includes: The image to be processed is input into the multi-scale feature extraction module to obtain the multi-scale features of the image to be processed; The image to be processed and the multi-scale features are input into the multi-scale feature fusion module to obtain fused features; The fused features are input into the decoder to obtain the target image.
3. The method according to claim 1, characterized in that, The multi-scale feature extraction module includes multiple feature extraction sub-modules, each of which includes a convolutional layer, a channel attention layer, a spatial attention layer, and a pooling layer.
4. The method according to any one of claims 1-3, characterized in that, The image processing model is trained in the following manner: A first sample set is obtained, which includes first image pairs acquired by multiple shooting devices in multiple shooting scenarios. Each first image pair includes a first sample image with Newton's ring interference and a second sample image without Newton's ring interference. The first sample image and the second sample image in the same first image pair are acquired by the same shooting device in the same shooting scenario. The first sample image is input into an end-to-end neural network to obtain the predicted image; Based on the second sample image and the predicted image, the total loss value of the end-to-end neural network is determined; Based on the total loss value, the parameters of the end-to-end neural network are updated to obtain the image processing model.
5. The method according to claim 4, characterized in that, Determining the total loss value of the end-to-end neural network based on the second sample image and the predicted image includes: Based on the mean squared error loss function, the second sample image, and the predicted image, the first loss value of the end-to-end neural network is determined; Based on the structural similarity loss function, the second sample image, and the predicted image, a second loss value for the end-to-end neural network is determined; The second sample image and the predicted image are respectively input into the feature extraction network to obtain the first feature map and the second feature map; Based on the perceptual loss function, the first feature map, and the second feature map, a third loss value is determined for the end-to-end neural network; The total loss value of the end-to-end neural network is determined based on the first loss value, the second loss value, and the third loss value.
6. The method according to claim 5, characterized in that, The step of determining the total loss value of the end-to-end neural network based on the first loss value, the second loss value, and the third loss value further includes: Determine the Newton's ring region and the non-Newton's ring region in the first sample image; Based on the region loss function, the second sample image, the predicted image, the first weight of the Newton's ring region and the second weight of the non-Newton's ring region, a fourth loss value of the end-to-end neural network is determined, wherein the first weight is greater than the second weight. The total loss value of the end-to-end neural network is determined based on the first loss value, the second loss value, the third loss value, and the fourth loss value.
7. The method according to claim 6, characterized in that, Determining the Newton's rings region and non-Newton's rings region in the first sample image includes: The first sample image is transformed in the frequency domain to obtain a first frequency domain image; The target region in the frequency domain image is determined based on the mask image; The target region is filtered or frequency domain enhancement processed to obtain a second frequency domain map; Based on the second frequency domain map, the Newton's rings region is determined, and the region in the first sample image other than the Newton's rings region is determined as the non-Newton's rings region.
8. The method according to any one of claims 1-3, characterized in that, The image to be processed is acquired by the target shooting device or in the target shooting scene; the image processing model is trained through the following steps: The end-to-end neural network is trained based on the first sample set to obtain a pre-trained model. The first sample set includes first image pairs acquired by multiple shooting devices in multiple shooting scenarios. Each first image pair includes a first sample image with Newton's ring interference and a second sample image without Newton's ring interference. The first sample image and the second sample image in the same first image pair are acquired by the same shooting device in the same shooting scenario. The pre-trained model is fine-tuned based on the second sample set to obtain the image processing model. The second sample set includes second image pairs acquired by the target shooting device or acquired in the target shooting scene. Each second image pair includes a third sample image with Newton's ring interference and a fourth sample image without Newton's ring interference.
9. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the method as claimed in any one of claims 1-8.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-8.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-8.