Multi-focus image fusion method and device applied to microscope and medium
By grouping images acquired by a microscope and fusing them with a deep learning network model, clear images are generated and a sharpness map is recorded. This solves the problem of the variable number of images in a microscope scene, improves the image fusion effect, and achieves pixel-level sharpness scoring.
Patent Information
- Application Number
- CN202511007875.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-07
AI Technical Summary
In existing microscope scenarios, the number of images is not fixed. Traditional image sharpness discrimination methods perform poorly in textureless areas, and deep learning models cannot handle the fusion of multiple images, resulting in poor image fusion effects.
By grouping images acquired by a microscope, a deep learning network model with multiple image inputs is constructed. The trained deep learning network model is then used to fuse images, generate clear images, and record a sharpness map to achieve pixel-level sharpness scoring.
It improves image fusion performance, solves the problem of fusion failure in textureless areas by traditional methods, realizes the application of deep learning models with multiple image inputs in microscope scenes, and outputs a final clear image and a pixel-level sharpness rating map.
Smart Images

Figure CN120912449A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of microscope multi-focus image fusion, and particularly relates to a multi-focus image fusion method, device and medium applied to a microscope. BACKGROUND
[0002] In the case of known microscope focal length, the distance between the object and the microscope lens can be determined according to whether the image is clear. For an object with ups and downs on the surface, the object image under different distances is shot, so that the object surface at different positions is sequentially clear, thereby the ups and downs information of the object surface can be obtained, and finally the images under different distances are synthesized into a clear image, that is, multi-focus image fusion. An example is shown in Figure 2 The principle is that the effective focusing depth of design of the microscope lens is relatively small. We think that at a certain distance d between the objective lens and the object, the object surface has ups and downs, and the region of the object satisfying the current focal length is clear in the microscope imaging, and the other object surface not satisfying the current focal length is blurred. According to this principle, we can know that the depth of the clear region of the object is equal to the current focal length. In this way, slightly moving the microscope to change the distance between the current object and the objective lens, another part of the object surface clear image can be obtained. When we obtain a plurality of images under different focal lengths in sequence, a group of clear images of different parts of the object surface can be obtained. We synthesize this group of images to synthesize a clear image of the entire object surface, that is, multi-focus fusion.
[0003] There are two technical routes to realize this set of multi-focus image fusion, the traditional image sharpness discrimination route and the deep learning image fusion generation route. These two routes have their own shortcomings, which we will introduce in turn. The implementation method of the traditional image sharpness discrimination route is to use the calculation of the local area gradient and sharpness of the image to obtain the sharpness of the current area, and select the image with the highest sharpness from the set of images as the final fused image of the current area. This method will produce poor results in areas with less texture, especially when there are some dirty particles in the area with less texture, which will greatly affect the accuracy of the sharpness calculation. The deep learning image fusion route has been studied in recent years with the development of deep learning, but this method relies on training data, and in practical applications, it will encounter the problem of an indefinite number of images. The current deep learning image fusion model inputs two images, the close-up image and the far-focus image, and then fuses them into a clear image, but it does not provide information about the distance of the area (i.e., whether the area is clear in the close-up image or the far-focus image). However, in actual microscope applications, more than two images are usually taken, and dozens or even hundreds of images with different focal lengths are obtained by continuously zooming in. However, only a small part of each image is clear. The number of images is also determined by the size of the target object. In addition, for deep learning models, the number of input images is fixed once the deep learning model is constructed, but the number of images taken by the microscope each time is not fixed. Therefore, deep learning models cannot be directly applied to the microscope scene. SUMMARY
[0004] Technical purpose: In view of the defects in the prior art that the number of images obtained in the microscope scene is not fixed and the current deep learning model only inputs two images, the present application discloses a multi-focus image fusion method, device and medium applied to a microscope, which groups the images collected by the microscope, constructs a multi-image input image fusion model, obtains the final clear image, and provides a pixel-level sharpness score map.
[0005] Technical scheme: In order to achieve the above technical purpose, the present application adopts the following technical scheme.
[0006] A multi-focus image fusion method applied to a microscope, the method comprising: S1, collecting a plurality of sequence images of a to-be-tested object by a microscope, each sequence image comprising a plurality of images with continuously changing distances within a depth of field range; S2, grouping the sequence images, setting a parameter N according to the input image number of the trained deep learning network model, and dividing the images with the same distance in each sequence image into a group, i.e., dividing the S multi-focus images in all sequence images into a plurality of groups according to the fact that each group comprises N images; S3, input each group of images including N images into the trained deep learning network model for image fusion to generate a corresponding fusion generated image, i.e. a clear image; S4, add the generated clear image to the original data set, and record the clarity map corresponding to the clear image; S5, group the images not participating in the last image fusion process again according to S2, and enter S6; the images not participating in the last image fusion process include the clear image generated by the trained deep learning network model last time; S6: input each group of images including N images into the trained deep learning network model for fusion generation; S7: determine whether the number of images not fused is 1, i.e. all images except the image generated by the last fusion have participated in the image fusion process, if yes, proceed to S8; if no, return to S4 and continue to execute; S8: take the clear image obtained by the last image fusion as the final clear image, backtrack the clarity map record in each fusion process according to the clarity map of the final clear image, and find the image number of each pixel in the final clear image that is clearest, i.e. the image sequence number in the sequence image in S1.
[0007] Preferably, in S2, if S / N is not an integer, the last group of images is expanded so that there are N images in the last group participating in the current grouping fusion process; or the last group of images is retained and does not participate in the current grouping fusion process, and enters the next grouping fusion process as the images not participating in the last image fusion process.
[0008] Preferably, the training process of the deep learning network model comprises: S31, constructing a deep learning network model; the deep learning network model comprises N connected feature extraction modules, a feature fusion module, a clarity recognition module and a Decoder image generation module; S32, constructing a sample set of the deep learning network model; the sample set comprises a plurality of sequence images of different objects, each sequence image comprising a plurality of images with continuously changing distances within a range of a depth of field; the target ground truth clear image for training is obtained by manually removing noise points after image fusion according to a current traditional image processing method; and the clarity map required for training is obtained by calculating the similarity between the clear image fused and the original image; S33, training the deep learning network model using the sample set to obtain the trained deep learning network model.
[0009] Preferably, each feature extraction module is configured to perform feature extraction on an input image, and outputs image features; the feature fusion module is connected with the N feature extraction modules, and is configured to perform fusion coding on the input image features, and outputs coded features; the sharpness recognition module is connected with the feature fusion module, and is configured to generate a sharpness map; and the decoder image generation module is connected with the feature fusion module, and is configured to output a sharp image.
[0010] Preferably, the size of the sharpness map is H*W*N, H, W and N represent height, width and channel number respectively; the sharpness map has the same width and height as the original image; and the probability distribution of each pixel point on the sharpness map in N represents the maximum sharpness of the pixel point from which image.
[0011] Preferably, in S33, the total loss function Loss = Loss_map * lambda + Loss_img, lambda is a weight, Loss_map is a loss function of the sharpness map, and Loss_img is a loss function of the sharp image.
[0012] Preferably, in S33, the deep learning network model is trained by using a sample set, including two-stage training, in the first stage, the sample set is directly used for training until the total loss function Loss meets a preset threshold; and in the second stage, the sharp image generated by the deep learning network model trained in the first stage according to the sample set is mixed into the sample set, and the deep learning network model is trained again until the total loss function Loss meets the preset threshold.
[0013] Preferably, the sequence images of different objects in the sample set include collecting a different number of sequence images of each object by using a microscope, collecting sequence images of different positions of the object by setting different depth of field strokes of the microscope according to different parts of the object surface.
[0014] A computer device includes a processor and a memory, the memory stores a computer program, and the processor is configured to execute the computer program to implement the multi-focus image fusion method applied to a microscope according to any one of the above.
[0015] A computer storage medium stores a computer program, and the computer program, when executed on a processor, implements the multi-focus image fusion method applied to a microscope according to the above.
[0016] Beneficial effects: the application adopts a deep learning model fusion manner to improve image fusion effect and solve the fusion failure problem of traditional methods in non-texture areas. The model proposed in the application solves the problem of traditional fusion models only processing two images. The grouping fusion method proposed in the application can truly apply the deep learning model to actual scenes, output the final clear image, and the pixel-level clarity score map, indicating the corresponding input image of the most clear image in the area. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 is a total method flowchart of the embodiment of the application; Figure 2 is an example diagram of multi-focus image fusion; Figure 3 is a structural schematic diagram of the deep learning network model of the embodiment of the application; Figure 4 is a training process schematic diagram of the deep learning network model of the embodiment of the application. DETAILED DESCRIPTION
[0018] In order to enable personnel in the technical field to better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor are within the scope of protection of the present application. EMBODIMENT
[0019] As shown in the accompanying Figure 1 The multi-focus image fusion method applied to a microscope of the embodiment includes: S1, collecting a plurality of sequence images of a to-be-tested article through a microscope, each sequence image including a plurality of images with continuously changing distances within a depth of field range; In the embodiment, the to-be-tested article is placed under the microscope, the upper and lower limits of the height of the article are adjusted to set the upper and lower limit ranges of the depth of field range of the collected image, and the image is collected after the microscope, and a total of S multi-focus images are obtained. In order to fully extract the features of the to-be-tested article, different numbers of sequence images are collected at different depth of field ranges when the multi-focus images are collected, that is, the S multi-focus images are a plurality of sequence images, and each sequence image includes a plurality of images with continuously changing distances within a depth of field range; S2, group the sequence images, set the parameter N according to the input image number of the trained deep learning network model, divide the images with the same distance in each sequence image into a group, that is, divide the S multi-focus images in all sequence images into several groups according to that each group includes N images, and the maximum number of groups is M = ⌈S / N ⌉; In S2, S / N is not necessarily an integer, that is, there is a case that the number of images in the last group is less than N, at this time, the last group of images is expanded, which can be expanded by randomly copying the images in other sequence images, so that there are N images in the last group, participating in the grouping and fusion process this time; or the last group of images is retained and does not participate in the grouping and fusion process this time, and enters the next grouping and fusion process as the images that do not participate in the last image fusion process.
[0020] S3, input each group of images including N images into the trained deep learning network model for image fusion to generate the corresponding fusion generated image, that is, the clear image; In this embodiment, each group of images including N images is input into the trained deep learning network model, and each group of images is input into the trained deep learning network model for fusion generation, in each fusion generation process, N images in the group are input, and the fusion generated image, that is, the clear image, is output. The training process of the deep learning network model includes: S31, constructing a deep learning network model; the deep learning network model includes N feature extraction modules, a feature fusion module, a definition recognition module, and a Decoder image generation module; The structure diagram of the deep learning network model is as shown in Figure 3 Each feature extraction module is used for feature extraction of the input image, and the input is an image and the output is an image feature; in this embodiment, the N feature extraction modules are all realized by using a convolutional neural network, and the network structure in this part can use the common backbone structure in the current deep learning field, for example, Resnet50, EfficientnetV2, etc., which is not limited here; The feature fusion module is connected with the N feature extraction modules, and is used for fusion coding of the input image features, the input is N image features, and the output is coded features; in this embodiment, the image features output by the N feature extraction modules are combined together through a concat connection layer first, and then sent to the feature fusion module Encoder, that is, the features of all images are fused and coded, and the structure of this module can use the structure based on CNN or the structure based on vision transformer encoder, which is not specifically limited in this application; The sharpness recognition module is connected with the feature fusion module, and is used for generating a sharpness map with a size of H*W*N, wherein H, W and N respectively represent height, width and channel number; the sharpness map has the same width and height as the original image, and the probability distribution of each pixel point on the sharpness map on N represents that the maximum sharpness of the pixel point comes from which image, and N also corresponds to N images in a group of images, so as to facilitate backtracking of the maximum sharpness position. The decoder image generation module is connected with the feature fusion module, and is used for outputting a sharp image; in the embodiment, a deconvolution layer or a vision transformer decoder structure is adopted, and the deconvolution layer is a deconvolution layer, which is a common image generation method based on a CNN in previous years; The deep learning model fusion method is adopted to improve the image fusion effect, and the fusion failure problem of the traditional method in a non-texture area is solved.
[0021] S32, a sample set of a deep learning network model is constructed; the sample set includes a plurality of sequence images of different objects, each sequence image includes a plurality of images with continuously changed distances within a depth of field range; a target ground truth sharp image is obtained by manually removing noise points after image fusion according to a current traditional image processing method; and a sharpness map required for training is obtained by calculating the similarity between the fused sharp image and the original image; The sample set includes a plurality of sequence images of different objects, including collecting different numbers of sequence images of each object through a microscope, collecting sequence images of different positions of the object according to different ups and downs of different parts of the surface of the object, and collecting sequence images of different positions of the object according to different depth of field settings of the microscope. For different objects, such as circuit boards, screws, rubber, pens, steel plates, glass plates, metal pads, wrenches and other tools, different depth of field settings are set to collect different numbers of sequence images through the microscope. Since the ups and downs of different parts of each object are different, in this embodiment, 50 groups of sequence images of different positions of each object can be collected on average, so as to expand the entire data set. During training, N images are randomly selected from each group of sequence images as input images each time. The target ground truth clear image obtained by training is obtained by manually removing noise points after image fusion according to the current traditional image processing method, and the required clarity map for training is obtained by calculating the similarity between the fused clear image and the original image. For example, if we have obtained the fused clear image, then the clear image and the original N images are convolved at a certain pixel point with a size of k*k to obtain the region with the maximum value, which is the most similar region. Therefore, it can be considered that the most clear position of the pixel point is on the image. k is selected according to the actual situation, and is generally an integer less than or equal to 5.
[0022] S33, training the deep learning network model using the sample set to obtain a trained deep learning network model; As shown in FIG. 1, the sample set includes a plurality of sequence images of different objects, including collecting different numbers of sequence images of each object through a microscope, collecting sequence images of different positions of the object according to different ups and downs of different parts of the surface of the object, and collecting sequence images of different positions of the object according to different depth of field settings of the microscope. For different objects, such as circuit boards, screws, rubber, pens, steel plates, glass plates, metal pads, wrenches and other tools, different depth of field settings are set to collect different numbers of sequence images through the microscope. Since the ups and downs of different parts of each object are different, in this embodiment, 50 groups of sequence images of different positions of each object can be collected on average, so as to expand the entire data set. During training, N images are randomly selected from each group of sequence images as input images each time. The target ground truth clear image obtained by training is obtained by manually removing noise points after image fusion according to the current traditional image processing method, and the required clarity map for training is obtained by calculating the similarity between the fused clear image and the original image. For example, if we have obtained the fused clear image, then the clear image and the original N images are convolved at a certain pixel point with a size of k*k to obtain the region with the maximum value, which is the most similar region. Therefore, it can be considered that the most clear position of the pixel point is on the image. k is selected according to the actual situation, and is generally an integer less than or equal to 5. Figure 4As shown, during training, first-stage training is performed, that is, from each group of sequence images, N images are randomly selected as input images, and the network output is subtracted from the clarity map and the clear image prepared in the data set to calculate the loss function Loss for training. The process is the same as the common loss training of deep learning image generation, and the loss function formula can use the existing formula, which is not limited here. The loss function loss calculation of the clarity map is recorded as Loss_map, and the loss function Loss calculation of the clear image is recorded as Loss_img. These two Losses are combined by adding a weight during actual training, that is, the total loss function Loss = Loss_map * lambda + Loss_img, lambda is the weight, and is valued according to experience. After the total loss function Loss meets the preset threshold, the first-stage training process is completed, and the second-stage training is entered. In the second-stage training, the deep learning network model has a certain fusion ability, and the clear image generated by the deep learning network model trained in the first stage is numbered according to the sample set, mixed into the sample set, and trained again, that is, the second-stage training. From the sample set mixed with the clear image, N images are randomly selected as input images, input into the deep learning network model for training, and the second-stage training process is completed after the total Loss meets the preset threshold.
[0023] S4, the generated clear image is added to the original data set, and the clarity map corresponding to the clear image is recorded; S5, the images not participating in the last image fusion process are grouped again according to S2, and S6 is entered. The images not participating in the last image fusion process include the clear images generated by the deep learning network model after the last training, such as the images numbered from S+1 in S4, and if the last group of images in the last grouping is not expanded, all images in the last group in the last grouping are also included.
[0024] S6: each group of images including N images is input into the trained deep learning network model for fusion generation.
[0025] S7: whether the number of images not fused is 1 is judged, that is, all images except the image generated by the last fusion participate in the image fusion process, if yes, S8 is performed; if not, return to S4 and continue to execute; S8: the clear image obtained by the last image fusion is taken as the final clear image, the clarity map record in each fusion process is traced back according to the clarity map of the final clear image, and the number of each pixel in the final clear image is found.
[0026] The application can apply a deep learning model to an actual scene by grouping images collected by a microscope, construct an image fusion model with multi-image input, obtain a final clear image and a pixel-level clarity score map, and indicate the corresponding clearest image of each pixel of the final clear image in the initial sequence image.
[0027] The application further discloses a computer device including a processor and a memory, the memory storing a computer program, and the processor being configured to execute the computer program to implement the multi-focus image fusion method applied to a microscope according to any one of the above. The memory can be various types of memory, such as random access memory, read-only memory, flash memory, etc.
[0028] The application further provides a computer storage medium storing a computer program, and when the computer program is executed, a device executing the computer program implements the multi-focus image fusion method applied to a microscope according to the above. The computer storage medium can be a readable storage medium, a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium can include, but is not limited to, a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various storage medium capable of storing program codes.
[0029] The above description is only the preferred embodiments of the application, and it should be pointed out that those skilled in the art can make some improvements and refinements without departing from the principles of the application, and these improvements and refinements should also be considered as the protection scope of the application.
Claims
1. A multi-focus image fusion method applied to a microscope, characterized by, The method comprises: S1, collecting a plurality of sequence images of the object to be measured by a microscope, each sequence image comprising a plurality of images with continuously changing distances within a depth of field range; S2, grouping the sequence images, setting a parameter N according to the number of input images of the trained deep learning network model, and dividing the images with the same distance in each sequence image into a group, that is, dividing the S multi-focus images in all sequence images into a plurality of groups according to the condition that each group comprises N images; S3, inputting each group of images comprising N images into the trained deep learning network model for image fusion to generate a corresponding fusion generated image, that is, a clear image; S4, adding the generated clear image to the original data set and recording the clarity map corresponding to the clear image; S5, grouping the images not involved in the last image fusion process again according to S2, and entering S6; the images not involved in the last image fusion process include the clear image generated by the trained deep learning network model in the last time; S6, inputting each group of images comprising N images into the trained deep learning network model for fusion generation; S7, judging whether the number of unfused images is 1, that is, all images except the image generated by the last fusion have participated in the image fusion process, if yes, proceeding to S8; if no, returning to S4 and continuing to execute; S8, taking the clear image obtained by the last image fusion as the final clear image, backtracking the clarity map record in each fusion process according to the clarity map of the final clear image, and finding the image number of each pixel in the final clear image, that is, the image sequence number in S1.
2. The multi-focus image fusion method applied to a microscope according to claim 1, characterized in that: In S2, if S / N is not an integer, the last group of images is expanded so that there are N images in the last group, which participate in the grouping and fusion process this time; or the last group of images is retained and does not participate in the grouping and fusion process this time, and enters the next grouping and fusion process as the images not involved in the last image fusion process.
3. The multi-focus image fusion method applied to a microscope according to claim 1, characterized in that: The training process of the deep learning network model comprises: S31, constructing a deep learning network model; the deep learning network model comprises N connected feature extraction modules, a feature fusion module, a clarity recognition module and a Decoder image generation module; S32, constructing a sample set of the deep learning network model; the sample set comprises a plurality of sequence images of different objects, each sequence image comprising a plurality of images with continuously changing distances within a depth of field range; the target groundtruth clear image obtained by manual removal of noise points after image fusion according to the current traditional image processing method; the clarity map required for training is obtained by calculating the similarity between the clear image obtained by fusion and the original image; S33, training the deep learning network model using the sample set to obtain the trained deep learning network model.
4. The multi-focus image fusion method applied to a microscope according to claim 3, characterized in that: Each feature extraction module is used for feature extraction of an input image, and the input is an image and the output is an image feature; feature The fusion module is connected with the N feature extraction modules, and is configured to fuse and encode the input image features, and has N image features as the input and encoded features as the output. The sharpness recognition module is connected with the feature fusion module, and is configured to generate a sharpness map.
5. The multi-focus image fusion method applied to a microscope according to claim 4, characterized in that: The size of the sharpness map is H*W*N, and H, W and N represent height, width and channel number, respectively.
6. The multi-focus image fusion method applied to a microscope according to claim 3, characterized in that: The sharpness map has the same width and height as the original image, and the probability distribution of each pixel point on the sharpness map in N represents the maximum sharpness of the pixel point from which image.
7. The multi-focus image fusion method applied to a microscope according to claim 3, characterized in that: In S33, the total loss function Loss = Loss_map * lambda + Loss_img, lambda is a weight, Loss_map is a loss function of the sharpness map, and Loss_img is a loss function of the sharpness image.
8. The multi-focus image fusion method applied to a microscope according to claim 3, characterized in that: In S33, the sample set is used to train the deep learning network model, including two-stage training.
9. A computer device, comprising: In the first stage, the sample set is directly used for training until the total loss function Loss meets the preset threshold.
10. A computer storage medium, characterized in that, In the second stage, the sharpness image generated by the deep learning network model trained in the first stage according to the sample set is mixed into the sample set, and the deep learning network model is trained again until the total loss function Loss meets the preset threshold. The sample set includes a plurality of sequence images of different objects, which are collected by a microscope. The computer device includes a processor and a memory, and the memory stores a computer program. The computer program is executed on the processor to implement the multi-focus image fusion method applied to the microscope according to any one of claims 1-8. The computer program is executed on the processor to implement the multi-focus image fusion method applied to the microscope according to any one of claims 1-8.
Citation Information
Patent Citations
Multi-focus sequence image fusion method
CN104182952A
Multi-focus image fusion method based on zero sample learning
CN113313663A
Multi-focus image fusion method combining depth context and convolution conditional random field
CN113763300A
Multi-focus image end-to-end fusion system
CN117372832A