An autofocus method, system and storage medium for weakly featured targets
By combining the parallel U-Net model of image spectrogram and convolutional neural network, the accuracy and stability of weak feature target autofocus in industrial environments are solved, and efficient and low-cost autofocus effect is achieved.
Patent Information
- Application Number
- CN202510465204.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The prior art autofocus method for weak feature targets in industrial environments lacks accuracy, stability and robustness, and has high system complexity and cost.
Combining image spectrograms and convolutional neural networks, a parallel U-Net network model is built, digital images and their spectrograms are trained, and automatic focus is achieved using improved Dice loss function and frequency domain information-based evaluation function.
Improves autofocus accuracy and reliability of weak feature targets, simplifies system structure and reduces costs.
Smart Images

Figure CN120151654B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an automatic focusing method, system and storage medium for weak feature targets aiming at the problem of image recognition focusing for weak feature images. Background Art
[0002] Weak feature targets are of great significance in the industrial field, especially in scenarios such as quality control, automated inspection, and intelligent manufacturing. Due to the complex industrial environment, target objects may exhibit weak features due to lighting, material, surface characteristics, or background interference, which poses higher requirements for detection and recognition technologies. Due to the lack of obvious texture and edge information, when imaging weak feature targets, traditional contrast detection or phase detection focusing methods may fail, thereby affecting the realization of functions such as product quality inspection, precise measurement, and automatic positioning. Therefore, a reliable automatic focusing method for feature targets is of great significance in the industrial field.
[0003] Currently, traditional weak feature target focusing methods mainly complete the focusing through methods such as contrast detection and phase detection in combination with measures such as structured light and ToF. Although the focusing can be completed, there are still large gaps in terms of accuracy, stability, and robustness, and additional facilities need to be added, resulting in a relatively high system complexity and cost. Summary of the Invention
[0004] To overcome the problems and defects existing in the current focusing on weak feature targets, the present invention combines the image spectrogram and convolutional neural network technology for the automatic focusing problem in the imaging process of weak feature targets. First, an improved parallel U-Net network is used to process the source image, and the digital image and its spectrogram are used as training data to train the network so that it can accurately mark the weak feature regions. Then, an improved automatic focusing evaluation function based on image sharpness and spectrum is used to complete the automatic focusing of the image. The present invention proposes an automatic focusing method for weak feature targets, including the following steps:
[0005] S1. Construct a parallel U-Net network model;
[0006] The model consists of two improved U-Net model branches. The upper branch receives the digital image input, and the lower branch receives the spectrogram input;
[0007] S2. Supervise and train the model;
[0008] Generate the spectrograms of the digital images before and after focusing, and use the images and their spectrograms before and after focusing to supervise and train the parallel U-Net network model in step S1;
[0009] S3. Optimize the network using an improved Dice loss function;
[0010] S4. Implement autofocus based on the evaluation function that fuses frequency-domain information.
[0011] Further, in the improved U-Net model in step S1, the decoder receives both the output of the encoder at the same level and the output of the upper-level encoder.
[0012] Further, a 1×1 convolutional layer is set between the decoding layers for dimension adjustment;
[0013] Further, the dual-branch output feature maps in step S1 are stacked and then fused through a 3×3 convolutional layer and a fully connected layer to output the final result.
[0014] Further, the improved Dice loss function in step S3 is:
[0015]
[0016] where and represent the true region and the predicted region of the target respectively, is the dot product between the pixels of the predicted region and the pixels of the true region, and the dot product results are added up, is the sum of the pixels in their respective corresponding regions, is the spectral component in the original image spectrogram, is the spectral component in the predicted image spectrogram, is the number of spectral components, is the index of the spectrum.
[0017] Further, the implementation of autofocus based on the evaluation function that fuses frequency-domain information is specifically:
[0018]
[0019] where, is the gray value of the pixels in the weak feature region of the image, and are the row and column numbers of the image pixels respectively, and are the indexes of the row and column respectively, is the pixel average value, represents the frequency-domain element value, and represent the frequency components, is a high-pass filter, is the frequency-domain penalty term, where is the fundamental frequency component, is the penalty coefficient, and its value can be taken from 10 −4 to 10 −1 in between.
[0020] The present invention also provides a weak feature target autofocus system, including:
[0021] An image acquisition module, configured to acquire the original image data of the target object;
[0022] A spectrum generation module, configured with a Fourier transform processor, for converting the digital image into a frequency domain map;
[0023] A parallel U-Net processing module, including the above-mentioned parallel U-Net network model;
[0024] A focusing module, configured to implement the above-mentioned autofocus method based on the evaluation function integrating frequency domain information and adjust the focus.
[0025] The present invention also provides a computer-readable storage medium, storing program instructions of the above method.
[0026] Advantages of the present invention:
[0027] Aiming at the autofocus problem in the imaging process of weak feature targets, the present invention combines the image spectrum map with convolutional neural network technology. First, the improved parallel U-Net network is used to process the source image, and the digital image and its spectrum map are used as training data to train the network, enabling it to accurately mark the weak feature regions. Then, the improved autofocus evaluation function based on image sharpness and spectrum is used to complete the autofocus of the image, thus effectively improving the accuracy and reliability of autofocus for weak feature targets, and having a broad application market space and economic value. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1 It is the basic structure diagram of U-Net;
[0030] Figure 2 It is the schematic diagram of the basic structure of the parallel U-Net;
[0031] Figure 3 It is the weak feature image and its spectrum map;
[0032] Figure 4 It is the autofocus effect diagram of the present invention. Detailed Embodiments
[0033] The present application will be described below in conjunction with specific embodiments:
[0034] Embodiment 1:
[0035] This embodiment provides a method for automatically focusing on weak feature targets, including the following steps:
[0036] S1. Construct a parallel U-Net network model; the model consists of two improved U-Net model branches. The upper branch receives digital image input, and the lower branch receives spectrogram input.
[0037] U-Net is a convolutional neural network architecture widely used in image segmentation tasks. It mainly consists of an encoder and a decoder, in a U shape. The basic structure of U-Net is as Figure 1 shown, where the encoder and the decoder correspond to each other. Both the encoder and the decoder are composed of convolutional blocks. Each layer of the encoder and the decoder contains two 3×3 convolutional layers. Each layer of the encoder is connected by a 2×2 downsampling layer, and each layer of the decoder is connected by a 2×2 upsampling layer. The feature maps obtained by each layer of the encoder are simultaneously input into the lower-layer encoder and the decoder at the corresponding level. Compared with the original U-Net neural network model, the present invention makes certain improvements to it and proposes a parallel U-Net model.
[0038] First of all, the decoder not only receives the output of the encoder at the same level, but also receives the output of each upper-level encoder, so as to better improve the decoding ability and quality. A 1×1 convolutional layer is added between each layer of the decoder and the next-layer decoder to adjust the dimension of the input data.
[0039] At the same time, this embodiment proposes a parallel U-Net model, as Figure 2 shown. After paralleling two improved U-Net models, the upper branch receives digital images as input, and the lower branch receives spectrograms as input. By stacking the output results of the two branches of the network and the output feature map of the top layer, and then through the processing of two 3×3 convolutional layers and a fully connected layer, the final output result is obtained.
[0040] S2. Supervise and train the model; generate spectrograms of the digital images before and after focusing, and supervise and train the parallel U-Net network model in step S1 through the images before and after focusing and their spectrograms.
[0041] In this embodiment, the training of the neural network model is completed using both digital images and spectrograms, and precise positioning is completed through the parallel U-Net network model architecture.
[0042] The characteristics of weak feature target imaging mainly include low contrast, no texture or weak texture, etc. In order to achieve accurate focus on weak feature targets, limited features must be highlighted. These features correspond to certain high-frequency areas in the spectrum diagram. The spectrum diagram of the image can be obtained by Fourier transform, and its calculation method is:
[0043]
[0044] in, represents the frequency domain element value, , represents the frequency component, is the image pixel value, Is an imaginary unit. Figure 3 The figure shows a nearly pure color target. Due to the lack of significant features, traditional autofocus algorithms often have difficulty in achieving ideal results. After Fourier transforming the source image and the weak features in it, the spectrum diagram obtained is as follows: Figure 3 shown.
[0045] First, the weak feature parts in the image are marked, and then the camera focal length is manually adjusted to obtain a clear image of the marked part. At this time, a pair of images are obtained, which are set as and .
[0046] Then the two images are Fourier transformed, and the obtained spectrum diagrams are and , now for the same scene, we get two pairs of images, where the input image is and , the images used for supervised training are and , by allowing the neural network to learn the characteristics of weak features from the input original image and spectrum image, and using the supervised clear image and spectrum image as targets, the decoder in the neural network is trained, and finally a network model that can accurately detect weak features is obtained.
[0047] S3. Use the improved Dice loss function for network optimization.
[0048] The present invention improves the Dice coefficient loss function to obtain a loss function suitable for the network of the present invention. Dice is a set similarity measurement function, which is usually used to calculate the similarity of two samples. The definition of the Dice loss function is as follows:
[0049]
[0050] in, , respectively represent the true region and the predicted region of the target, is the dot product between the pixels of the predicted region and the pixels of the true region, and the results of the dot product are added together, is the addition of the pixels in their respective corresponding regions.
[0051] The loss function used in the present invention is defined as follows:
[0052]
[0053] where, is the spectral component in the spectral diagram of the original image, is the spectral component in the spectral diagram of the predicted image, is the number of spectral components, is the index of the spectrum. Through the above improvements, the neural network used in the present invention can realize the prediction of an image with clear features from the original weak feature image, laying a foundation for autofocus.
[0054] S4. Realize autofocus based on the evaluation function that fuses frequency domain information.
[0055] Camera autofocus algorithms are usually implemented based on methods such as gradients and image contrast. The gradient-based method uses the edge details in the image to determine the sharpness of the image, thereby completing autofocus, but this method is sensitive to noise; the autofocus method based on image contrast evaluates the sharpness by calculating the differences of all pixel values in a certain region of the image, but it is sensitive to light changes. Based on the contrast evaluation, the present invention fuses the frequency domain evaluation method to complete the automatic and precise autofocus of the weak feature image. The calculation method is as follows:
[0056]
[0057] where, is the gray value of the pixels in the weak feature region recognized in step one, , are the row and column numbers of the image pixels respectively, , are the indexes of the row and column respectively, is the pixel average value, represents the frequency domain element value, , represent the frequency components, is a high-pass filter, is the frequency domain penalty term, where is the fundamental frequency component, is the penalty coefficient, and its value can be taken from 10 −4 to 10 −1 between.
[0058] Example 2:
[0059] This embodiment provides a weak feature target autofocus system for realizing the autofocus of the weak feature target in Embodiment 1, specifically including:
[0060] An image acquisition module for acquiring the original image data of the target object;
[0061] A spectrum generation module configured with a Fourier transform processor for converting a digital image into a frequency domain map;
[0062] A parallel U-Net processing module including the above-mentioned parallel U-Net network model;
[0063] A focusing module for implementing the autofocus method based on the evaluation function integrating frequency domain information and adjusting the focus.
[0064] Example 3:
[0065] This embodiment provides a computer-readable storage medium storing program instructions for implementing the method of Embodiment 1.
[0066] Aiming at the autofocus problem in the imaging process of weak feature targets, the present invention combines the image spectrogram with convolutional neural network technology. First, the improved parallel U-Net network is used to process the source image, and the digital image and its spectrogram are used as training data to train the network so that it can accurately mark the weak feature areas. Then, the improved autofocus evaluation function based on image sharpness and spectrum is used to complete the autofocus of the image, thereby effectively improving the accuracy and reliability of the autofocus of weak feature targets, and having a broad application market space and economic value.
[0067] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0068] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, various equivalent transformations (such as quantity, shape, position, etc.) can be made to the technical solution of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. An automatic focusing method for weak feature targets, characterized in that, It includes the following steps: S1. Construct a parallel U-Net network model; The model consists of two improved U-Net model branches. The upper branch receives digital image input, and the lower branch receives spectrogram input; S2. Conduct supervised training on the parallel U-Net network model; Generate spectrograms of digital images before and after focusing. Use the images before and after focusing and their spectrograms to conduct supervised training on the parallel U-Net network model in step S1; S3. Adopt an improved Dice loss function for network optimization; S4. Achieve autofocus based on an evaluation function that fuses frequency-domain information; The improved Dice loss function in step S3 is: ; wherein and represent the true region and the predicted region of the target respectively, is the dot product between the pixels of the predicted region and the pixels of the true region, and the dot product results are added together, is the sum of the pixels in their respective corresponding regions, is the spectral component in the original image spectrogram, is the spectral component in the predicted image spectrogram, is the number of spectral components, is the index of the spectrum.
2. The automatic focusing method for weak feature targets according to claim 1, wherein: In the improved U-Net model in step S1, the decoder receives the output of the encoder at the same level and the output of the upper-level encoder simultaneously.
3. The automatic focusing method for weak feature targets according to claim 2, wherein: A 1×1 convolutional layer is set between encoders at different levels for dimension adjustment.
4. The automatic focusing method for weak feature targets according to claim 1, characterized in that: The feature maps output by the two branches in step S1 are stacked and then fused through a 3×3 convolutional layer and a fully connected layer to output the final result.
5. According to the method for automatically focusing on weak feature targets described in claim 1, wherein: The implementation of autofocus based on the evaluation function that fuses frequency-domain information is specifically: ; Among them, is the gray value of the pixels in the weak feature region of the image, , are the row and column numbers of the image pixels respectively, , are the indexes of the row and column respectively, is the pixel average value, represents the frequency domain element value, , represent the frequency components, is a high-pass filter, is the frequency domain penalty term, where is the fundamental frequency component, is the penalty coefficient, and its value can be taken between 10 −4 and 10 −1 .
6. An automatic focusing system for weak feature targets, characterized in that It includes: An image acquisition module for acquiring the original image data of the target object; A spectrum generation module configured with a Fourier transform processor for converting digital images into frequency-domain graphs; A parallel U-Net processing module containing the parallel U-Net network model described in any one of claims 1-5; A focusing module for implementing the method for automatically focusing based on the evaluation function that fuses frequency-domain information described in any one of claims 1-5 and adjusting the focus.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are executed by the system, the method described in any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Medical image segmentation method based on federal learning and attention mechanism
CN116245886A
Cascade Transform-based brain tumor segmentation method
CN117036380A