A multi-image automatic and rapid stitching method for corneal laser confocal microscopy
Through the multi-image automatic rapid stitching method for the corneal laser confocal microscope, the neural network model is used to automatically and quickly stitch multiple small field of view into large field of view images, solving the problem of small field of view imaging of laser confocal microscopes, and achieving high-precision and high-efficiency image stitching.
Patent Information
- Application Number
- CN202510178130.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-18
AI Technical Summary
While obtaining high signal-to-noise ratio images, laser confocal microscopes have the problem of small imaging field of view, which is difficult to solve through hardware optimization. In addition, traditional image stitching methods have low accuracy and cannot accelerate in corneal image processing.
A multi-image automatic rapid stitching method for corneal laser confocal microscopy is adopted to generate corneal images with large field of view through data acquisition, controlled data augmentation and neural network model construction and training. This method includes feature extraction, feature maximum correlation matching and regression subnet prediction heads, and uses mixed loss function to optimize the model to achieve automatic and rapid stitching of images.
It effectively solves the problem of small field of vision for laser confocal microscope imaging, improves the accuracy and speed of image stitching, can accelerate through GPU, significantly improves the stitching efficiency, and provides doctors and researchers with more powerful tools.
Smart Images

Figure CN119887518B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of image processing, biomedical engineering, and artificial intelligence, and particularly relates to a multi-image automatic and rapid stitching method for a corneal laser confocal microscope. Background Art
[0002] A laser confocal microscope uses the principle of optical conjugation for imaging. Since the illumination point, the detection point, and the object point are conjugate to each other, it has a series of advantages such as high signal-to-noise ratio, strong depth penetration ability, and three-dimensional imaging, and has been widely used in the medical field. As a non-invasive high-resolution imaging tool, it can provide detailed images of the ocular surface structure, greatly enhancing the doctor's understanding and diagnostic ability of eye diseases. More importantly, it also allows doctors and researchers to observe living tissues at the subcellular level, which is crucial for studying diseases such as dry eye, corneal dystrophy, and glaucoma.
[0003] However, due to the limitations of the imaging principle of the laser confocal microscope, while obtaining high signal-to-noise ratio images, it brings the problem of a small imaging field of view, which is difficult to solve through hardware optimization and design. However, accurately diagnosing diseases requires large-field images to bring more information. Therefore, the present invention can use an image stitching algorithm to fuse and stitch multiple small-field images with overlapping regions in similar perspectives into a large-field image, and retain as much information as possible from the original multiple small-field images. Traditional image stitching methods applied to laser confocal corneal images will result in insufficient detection of feature points or uneven distribution of feature points. This manually designed feature greatly reduces the stitching accuracy, and it cannot be accelerated using a GPU.
[0004] It should be noted that the information disclosed in the above background art section is only used for understanding the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The main purpose of the present invention is to solve the problems existing in the above background art, and provide a multi-image automatic and rapid stitching method for a corneal laser confocal microscope.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A multi-image automatic and rapid stitching method for a corneal laser confocal microscope, comprising the following steps:
[0008] S1. Data acquisition: Use a corneal laser confocal microscope to collect corneal images and construct an initial data set, where the images contain the structural information of corneal tissue;
[0009] S2. Controlled data augmentation: Screen the initial data set to remove out-of-focus or invalid images; perform brightness correction on the remaining images to make their brightness tend to be consistent; crop an image patch of a fixed size in the central region of the image as a reference image, and perform random displacement perturbation on the image patch to generate a target image and a corresponding offset label, and construct a synthetic data set with labels.
[0010] S3. Neural network model construction and training: Build a neural network model, which includes a feature extractor, a feature maximum correlation matching layer, and a regression sub-network prediction head; use the synthetic data set to train the model, extract image features through the feature extractor, calculate the correlation between features by the feature maximum correlation matching layer, and the regression sub-network prediction head predicts the offset between image pairs based on the correlation result, and use a hybrid loss function to optimize the model until the model can accurately predict the offset.
[0011] S4. Image stitching: Preprocess multiple collected corneal images, adjust the image size and perform brightness correction; sequentially select adjacent image pairs and input them into the trained neural network model to predict the offset between the image pairs; when the predicted offset exceeds a preset threshold, terminate the stitching of the current group of images, and use the previous images as a group of images to be stitched; for each group of images to be stitched, calculate the global coordinate system according to the predicted offset, and fuse and stitch the images to generate a large-field-of-view corneal image.
[0012] Further, step S2 specifically includes:
[0013] Screen the initial data set to remove out-of-focus images or images that do not contain corneal structure information.
[0014] Perform brightness correction on the screened images to make the image brightness tend to be consistent, so as to reduce boundary artifacts during stitching and highlight the main structure of the cornea.
[0015] Crop an image patch of a fixed size in the central region of the image as a reference image.
[0016] Perform random displacement perturbation on the cropped image patch to generate a target image and a corresponding offset label, and the displacement perturbation is consistent in the horizontal and vertical directions, and the displacement range is a preset value.
[0017] The image pairs and their offset labels generated through the above steps constitute a synthetic data set, and the synthetic data set is used for the training and testing of the neural network model.
[0018] Further, step S3 specifically includes:
[0019] The feature extractor is composed of multiple convolutional blocks stacked together. Each convolutional block contains a two-dimensional convolutional operation and a non-linear activation function. Except for the last convolutional block, the remaining convolutional blocks are downsampled through a max-pooling layer;
[0020] The feature extractor extracts features from the reference image and the target image to form feature maps of the same size and shares weights;
[0021] Feature maps of different scales are used for offset prediction to capture the offset information of the image;
[0022] Among them, by predicting the offset step by step at different scales and estimating the displacement of the image pair at different scales, accurate prediction is achieved.
[0023] Furthermore, the offset prediction includes:
[0024] At the first scale, deep feature maps are used for prediction. The deep feature maps contain rich semantic information;
[0025] The deep feature maps of the reference image and the target image are input into the feature maximum correlation matching layer to calculate the correlation at different positions of the feature maps and obtain the correlation result;
[0026] The correlation result is input into the first regression sub-network prediction head to predict the displacement of the image pair in the x and y directions;
[0027] At the second scale, shallow feature maps are used for prediction. The shallow feature maps contain more edge, texture, and structure information;
[0028] According to the prediction result of the first scale, the shallow feature map of the target image is deformed and input into the feature maximum correlation matching layer together with the shallow feature Figure 1 of the reference image to obtain the correlation result of the second scale;
[0029] The correlation result of the second scale is input into the second regression sub-network prediction head to predict the displacement of the image pair in the x and y directions.
[0030] Furthermore, the final offset prediction result is obtained by fusing the prediction results of the first scale and the second scale. Among them, the fusion method is to add the prediction results of the two scales.
[0031] Furthermore, in step S3, the training of the neural network model also includes:
[0032] Construct a loss function consisting of two parts. The first part is the L2 loss, which is used to make the network focus on the overall displacement information, and the second part is the deformation loss, which is used to make the network focus on the content difference in the overlapping area of the images;
[0033] The L2 loss consists of the mean of the L2 losses at two different scales, where the L2 loss at the first scale is based on the prediction results at a single scale, and the L2 loss at the second scale is based on the difference between the accumulated value of the prediction results at two scales and the actual displacement;
[0034] The deformation loss calculates the content difference between the deformed image and the reference image in the overlapping region by applying the same deformation matrix to the mask and the input image;
[0035] The total loss is the weighted sum of the L2 loss and the deformation loss, where the weights of the L2 loss and the deformation loss are set to balance the attention to the overall displacement information and the content difference in the overlapping region.
[0036] Further, step S4 specifically includes:
[0037] Testing the trained neural network model on the synthetic dataset to verify the accuracy and speed of the model;
[0038] Deploying the model and creating a Web service interface, including a health check interface and a prediction interface, where the prediction interface is used to receive input data and call the model for prediction;
[0039] Collecting multiple real images and performing size adjustment and brightness correction on them to ensure the consistency of the input data and reduce boundary artifacts;
[0040] Starting from the first image, sampling adjacent image pairs in sequence and calling the API interface to predict the offset of the image pair;
[0041] When the predicted offset exceeds the preset threshold, terminating the prediction of the current image group and marking the image group as an image group to be stitched;
[0042] Repeating the above process until the prediction of all images is completed, obtaining multiple groups of image groups to be stitched;
[0043] For each group of images to be stitched, calculating the global coordinate system according to the predicted local offset and performing image fusion to obtain the stitching result.
[0044] Further, the calculating the global coordinate system according to the predicted local offset includes: multiplying each group of predicted local offsets by a scaling factor and performing calculations to obtain the global coordinate system, and gradually performing image fusion.
[0045] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the multi-image automatic and fast stitching method for a corneal laser confocal microscope.
[0046] A computer program product, when the computer program is executed by a processor, implements the multi-image automatic and rapid stitching method for a corneal laser confocal microscope as described above.
[0047] The present invention has the following beneficial effects:
[0048] The present invention proposes an innovative multi-image automatic and rapid stitching method for a corneal laser confocal microscope. This method effectively solves the problem of small imaging field of view of the laser confocal microscope, makes up for the hardware defects, and can process the automatic and rapid stitching of multiple images. Compared with the traditional image stitching method, the method of the present invention has significant improvements in terms of accuracy, speed and stability. By designing a novel controlled data augmentation method, the present invention can construct a synthetic data set with labels from the original data, thus providing high-quality data for model training. At the same time, the present invention designs a new registration neural network model, which includes three modules: a deep feature extractor, a feature maximum correlation matching layer, and a regression sub-network prediction head, and can complete the prediction of the offset at different scales, thereby improving the prediction accuracy. This multi-scale neural network model can adapt to diverse data distributions, model complex non-linear features, and use the different rich information contained in different-level feature maps to perform prediction regression at multiple scales. The method of the present invention can not only improve the stitching accuracy, but also be accelerated by GPU hardware, greatly improving the stitching efficiency, providing a powerful tool for doctors and researchers to observe living tissues at the subcellular level, and is of great significance for studying diseases such as dry eye, corneal dystrophy, and glaucoma.
[0049] Other beneficial effects in the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of the multi-image automatic and rapid stitching method for corneal confocal images in the embodiments of the present invention.
[0051] Figure 2 is a schematic diagram of the controlled data augmentation process in the embodiments of the present invention.
[0052] Figure 3 is a neural network structure diagram in the embodiments of the present invention.
[0053] Figure 4 is a model loss diagram in the embodiments of the present invention.
[0054] Figure 5 is an example 1 of the stitching result in the embodiments of the present invention.
[0055] Figure 6 is an example 2 of the stitching result in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is merely exemplary and not intended to limit the scope and application of the present invention.
[0057] A multi-image automatic rapid stitching method for a corneal laser confocal microscope includes the following steps:
[0058] Step S1, data acquisition: Use a corneal laser confocal microscope to collect corneal images and construct an initial data set, and the images contain the structural information of corneal tissue.
[0059] Step S2, controlled data augmentation: Screen the initial data set to remove out-of-focus or invalid images; perform brightness correction on the remaining images to make their brightness tend to be consistent; crop an image block of a fixed size in the central area of the image as a reference image, and perform random displacement perturbation on the image block to generate a target image and its corresponding offset label, and construct a synthetic data set with labels.
[0060] In a preferred embodiment, step S2 specifically includes: screening the initial data set to remove out-of-focus images or images that do not contain corneal structural information; performing brightness correction on the screened images to make the image brightness tend to be consistent, so as to reduce boundary artifacts during stitching and highlight the main structure of the cornea; cropping an image block of a fixed size in the central area of the image as a reference image; performing random displacement perturbation on the cropped image block to generate a target image and its corresponding offset label, and the displacement perturbation is consistent in the horizontal and vertical directions, and the displacement range is a preset value. The image pairs and their offset labels generated by the above steps constitute a synthetic data set, and the synthetic data set is used for the training and testing of the neural network model.
[0061] Step S3, neural network model construction and training: Build a neural network model, and the model includes a feature extractor, a feature maximum correlation matching layer, and a regression sub-network prediction head; use the synthetic data set to train the model, extract image features through the feature extractor, calculate the correlation between features by the feature maximum correlation matching layer, and the regression sub-network prediction head predicts the offset between image pairs based on the correlation result, and use a hybrid loss function to optimize the model until the model can accurately predict the offset.
[0062] In a preferred embodiment, step S3 specifically includes: The feature extractor is composed of multiple convolutional blocks stacked together. Each convolutional block includes a two-dimensional convolutional operation and a non-linear activation function. Except for the last convolutional block, the remaining convolutional blocks are downsampled through a max pooling layer; The feature extractor extracts features from the reference image and the target image to form feature maps of the same size and shares weights; Feature maps of different scales are used to predict the offset to capture the offset information of the image; Among them, by predicting the offset progressively at different scales and estimating the displacement of the image pair at different scales, accurate prediction is achieved.
[0063] In a further preferred embodiment, the offset prediction includes: At the first scale, deep feature maps are used for prediction. The deep feature maps contain rich semantic information; The deep feature maps of the reference image and the target image are input into the feature maximum correlation matching layer to calculate the correlation at different positions of the feature maps and obtain the correlation result; The correlation result is input into the first regression sub-network prediction head to predict the displacement of the image pair in the x and y directions; At the second scale, shallow feature maps are used for prediction. The shallow feature maps contain more edge, texture and structure information; The shallow feature map of the target image is deformed according to the prediction result of the first scale and input into the feature maximum correlation matching layer together with the shallow feature Figure 1 map of the reference image to obtain the correlation result at the second scale; The correlation result at the second scale is input into the second regression sub-network prediction head to predict the displacement of the image pair in the x and y directions. Preferably, the final offset prediction result is the sum of the prediction results of the first scale and the second scale. The prediction results of the first scale and the second scale are fused to obtain the final offset prediction result, where the fusion method is to add the prediction results of the two scales.
[0064] In a preferred embodiment, in step S3, the training of the neural network model further includes: constructing a loss function consisting of two parts. The first part is the L2 loss, which is used to make the network focus on the overall displacement information, and the second part is the deformation loss, which is used to make the network focus on the content difference in the overlapping area of the image; The L2 loss is composed of the mean of the L2 losses at two different scales. Among them, the L2 loss at the first scale is based on the prediction result of a single scale, and the L2 loss at the second scale is based on the difference between the accumulated value of the prediction results of the two scales and the actual displacement; The deformation loss calculates the content difference between the deformed image and the reference image in the overlapping area by applying the same deformation matrix to the mask and the input image; Among them, the total loss is the weighted sum of the L2 loss and the deformation loss, and the weights of the L2 loss and the deformation loss are set to balance the attention degree of the overall displacement information and the content difference in the overlapping area. Through the design of the above loss function, the accuracy of the overall displacement and the consistency of the content in the overlapping area are optimized simultaneously during the model training process to improve the accuracy and quality of the stitched image.
[0065] Step S4, Image stitching: Preprocess the multiple corneal images collected, adjust the image size and perform brightness correction; sequentially select adjacent image pairs and input them into the trained neural network model to predict the offset between the image pairs; when the predicted offset exceeds the preset threshold, terminate the stitching of the current group of images, and use the previous images as a group of images to be stitched; for each group of images to be stitched, calculate the global coordinate system according to the predicted offset, and fuse and stitch the images to generate a large-field corneal image.
[0066] In a preferred embodiment, step S4 specifically includes: using the trained neural network model to test on the synthetic dataset to verify the accuracy and speed of the model; deploying the model and creating a Web service interface, including a health check interface and a prediction interface, where the prediction interface is used to receive input data and call the model for prediction; collecting multiple real images and performing size adjustment and brightness correction on them to ensure the consistency of the input data and reduce boundary artifacts; starting from the first image, sequentially sample adjacent image pairs and call the API interface to predict the offset of the image pairs; when the predicted offset exceeds the preset threshold, terminate the prediction of the current image group and mark the image group as a group of images to be stitched; repeat the above process until all image predictions are completed to obtain multiple groups of images to be stitched; for each group of images to be stitched, calculate the global coordinate system according to the predicted local offset and perform image fusion to obtain the stitching result. Preferably, the calculating the global coordinate system according to the predicted local offset includes: multiplying the predicted local offset of each group by a scaling factor and performing calculations to obtain the global coordinate system, and gradually performing image fusion.
[0067] The multi-image automatic and fast stitching method based on supervised learning designed by the present invention is data-driven. By training on a synthetic dataset with labels, it can better learn the rules in the data, and thus be applied to unknown data for prediction. The neural network model designed by the present invention adaptively extracts features in the image through a convolutional neural network. The shallow feature maps contain more basic information in the image, such as edges, textures, structures, etc., while the deep feature maps contain richer high-level semantic information in the image. The multi-scale neural network model of the present invention can adapt to diverse data distributions, model complex non-linear features, and the present invention uses the different rich information contained in different-level feature maps to perform prediction regression at multiple scales to achieve a high stitching accuracy. In addition, the method of the present invention can be accelerated by using a GPU hardware with more computing cores, greatly improving the stitching efficiency.
[0068] The following further describes the specific embodiments, algorithm examples and experimental verifications of the present invention.
[0069] Based on the supervised learning technology in deep learning, for the corneal images collected by a corneal laser confocal microscope, the present invention designs a method that can automatically and quickly stitch multiple images, effectively increasing the imaging field of view of the corneal laser confocal microscope. This method can be implemented through the following steps to achieve the automatic stitching of corneal confocal images: The first step is to collect corneal confocal images through a corneal laser confocal microscope to construct an initial data set; the second step is to perform controlled data augmentation, and construct a labeled synthetic data set for training the model from the initial data set through steps such as image brightness correction, random cropping, and displacement perturbation; the third step is to build a registration neural network model for estimating the affine matrix, which mainly consists of three modules: a deep feature extractor, a feature maximum correlation matching layer, and a regression subnetwork prediction head; the fourth step is to train the registration neural network model of the present invention with the synthetic data set, predict the affine matrix at different scales, and use a hybrid loss function for supervision; the fifth step is model testing and verification, and verify the performance parameters of the model of the present invention on the synthetic test set; the sixth step is model deployment and use, package and deploy the trained model locally through FastAPI, then scale the multiple real images collected to the required size and input them into the model for inference, predict the offset between adjacent images, and terminate when the predicted value exceeds the threshold. Take all the images before termination as a group of images to be stitched, calculate the global coordinate system based on the predicted value, and gradually fuse the images to be stitched to obtain the final multi-image stitching result.
[0070] I. Construction of Synthetic Data Set
[0071] First, take corneal images of the subject through a corneal laser confocal microscope to obtain an initial data set. Next, the present invention proposes a novel controlled data augmentation (CDA) method to construct a synthetic data set, and the specific steps are as follows: First, screen the initial data set, filter out out-of-focus and images that do not contain any information, and then perform brightness correction on the remaining images to make the brightness of the captured images relatively similar, which can avoid obvious boundary artifacts during stitching, and through image brightness correction, the main structure in the image can also be highlighted, such as the nerves in the corneal anterior elastic layer; secondly, perform local cropping on the image, and here select to crop an image block of 256×256 in size in the central area of the image as the reference image, that is Figure 2 the image in the red box; finally, randomly perturb the position of the image block. It should be noted that the offsets in the x and y directions of the four vertices of the rectangle are kept consistent respectively, and the range of the offset is set to 64, that is:
[0072]
[0073] Through this step, the present invention not only obtains the target image, that is, Figure 2 the image in the green box, but also obtains the true label of the input image pair's offset and . Here, factors such as rotation, shear, and scaling are ignored, and only the movement of the microscope lens during the shooting process is considered. Because the time taken to shoot each image is extremely short and the contact lens results in the local corneal area being close to a plane rather than a curved surface. Through this data augmentation method, the present invention can construct a synthetic dataset from the initial dataset. Each piece of data in the synthetic dataset contains a pair of input image pairs to be registered and the label offset. Then, the synthetic dataset is divided into a training set and a test set according to a certain proportion for training and testing the model respectively.
[0074] II. Construction and Training of Neural Network Model
[0075] The neural network of the present invention mainly consists of three modules: a feature extractor, a feature maximum correlation matching layer, and a regression sub-network prediction head, and predicts the offset at two scales. The specific structure is as Figure 3 shown.
[0076] The feature extractor module is stacked by 4 blocks, and each block contains two Conv2d + RELU. Here, Conv2d refers to the two-dimensional convolution operation. During the implementation process, the size of the convolution kernel is set to 3×3. Here, RELU refers to a non-linear activation function. Except for the last block, the other three blocks are downsampled by a max-pooling layer with a stride of 2. The feature extractor structures for extracting features from the reference image (input 1) and the target image (input 2) are the same and share weights. It is a siamese network. In other words, at each scale, the sizes of the feature maps they obtain are exactly the same. Then, the 1 / 4-sized feature map before the last max-pooling layer and the 1 / 8-sized feature map output by the last block are used for subsequent parameter prediction to capture the offset information of the image at different scales.
[0077] First, the 1 / 8-sized feature map is used for prediction at scale 1. This deep feature map contains richer high-dimensional semantic information, and after multiple downsamplings, the surrounding spatial information aggregated by each feature vector is more abundant. The 1 / 8-sized feature maps of input 1 and input 2 are input into the feature maximum correlation matching layer to establish the mapping of the correlation between different positions of the deep feature map, find the connection between the regions with the highest correlation, obtain the correlation result at scale 1, and then input it into the regression sub-network prediction head 1 to predict the offset parameters. Here, only the displacements in the x and y directions are predicted, while factors such as rotation, shear, and scaling are ignored. The deformation matrix can be simplified as:
[0078]
[0079] The predicted result on scale 1 is recorded as , Then, the 1 / 4-sized feature map is used to make predictions at the next scale, i.e., scale 2. This shallow feature map contains more information about edges, textures, structures, etc. on the image. First, the 1 / 4-sized feature map of input 2 is distorted according to the prediction result of scale 1 and the deformation matrix, and then it is combined with the 1 / 4-sized feature map of input 1. Figure 1 It is input into the feature maximum correlation matching layer to obtain the correlation result on scale 2, and then input into the regression subnetwork prediction head 2 for prediction. The result predicted on scale 2 is recorded as , , the final prediction result of the neural network is the sum of the two, that is, , .
[0080] The loss function consists of two parts: L2 loss and deformation loss. L2 loss hopes that the network can pay attention to the overall displacement information, while deformation loss hopes that the network can pay attention to the content differences in the overlapping areas of the image. L2 loss is composed of the mean of L2 losses at two scales:
[0081]
[0082] The deformation loss is as follows, filtering the overlapping regions by a mask deformed by the same deformation matrix:
[0083]
[0084] The total loss is composed of L2 loss and deformation loss weighted:
[0085]
[0086] By predicting the offset step by step at different scales, the method of the present invention estimates the displacement of image pairs at different scales and achieves accurate prediction. The Adam optimizer is used for training on two Nvidia RTX 2080Ti, with the initial learning rate set to 0.0001 and an exponential decay strategy. The loss function during training is as follows: Figure 4 As shown, blue, purple, orange, pink and green represent , , , and , it can be found that the loss converges to a smaller range.
[0087] 3. Model Testing and Deployment Application
[0088] After completing the training of the registration network, the Registration_Network.pth weight file was obtained, and then the model of the present invention was tested on the synthetic dataset, and its accuracy and speed both reached a relatively high level.
[0089] Next, the model was deployed and applied. A Web service interface was created using FastAPI, including a health check interface, a prediction interface, etc. The prediction interface was used to receive the input data passed by the client, call the model for prediction, and return the results. Here, the model of the present invention used GPU acceleration for inference during prediction. Then, multiple real images were collected. To ensure consistent dimensions, their dimensions were adjusted before calling the interface for prediction, and the image dimensions were adjusted to 256×256. And a brightness correction algorithm was adopted to avoid boundary artifacts. Then, starting from the first image, adjacent two images were sampled and the API interface was called to predict the offsets in the x and y directions of the image pair. The predicted offsets were denoted as 、 。Here, i represents the offset predicted for the i-th and the (i + 1)-th images. When the predicted value exceeds the threshold, it means that there is no corresponding offset relationship between the i-th and the (i + 1)-th images, and the prediction for this group is terminated. The images numbered 1 to i are used as a group of images to be stitched, and then the next group of predictions is started until all images have been predicted.
[0090] After all images have been predicted, k groups of images to be stitched are obtained. The local offsets predicted for each group are multiplied by the scaling factor and calculated to obtain the global coordinate system, and image fusion is gradually performed to obtain the multi-image stitching results of the confocal images of the ocular surface for each group.
[0091] Experimental Results
[0092] It was tested using the CCMID dataset. A multi-image stitching result obtained by stitching 10 confocal images of the ocular surface is as shown in Figure 5 . It can be found that the anterior elastic layer nerves on it maintained good continuity and there were no obvious breakpoints, which proved the effectiveness of the method of the present invention. Another multi-image stitching result obtained by stitching 8 confocal images of the ocular surface is as shown in Figure 6 . Its imaging field of view is approximately 3 to 4 times that of the original image.
[0093] Generally speaking, the important features and innovative points of the present invention include:
[0094] 1. A novel method for automatic and fast multi-image stitching for corneal laser confocal microscopy is proposed. It is no longer limited to the stitching of adjacent two images, but extended to the automatic and fast stitching of multiple images, which can make up for the defects in the hardware of the laser confocal microscope, effectively increase the imaging field of view, and its accuracy, speed and stability are all better than the existing methods.
[0095] 2. A novel controlled data augmentation method is proposed to construct a synthetic dataset with labels from the original data, mainly including steps such as brightness correction, cropping, and displacement perturbation.
[0096] 3. A new registration neural network model is designed, which consists of three modules: a deep feature extractor, a feature maximum correlation matching layer, and a regression sub-network prediction head, to predict the offset at different scales and improve the prediction accuracy.
[0097] In summary, the automatic and fast stitching method innovatively proposed in the present invention greatly increases the imaging field of view of the corneal laser confocal microscope, not only improves the accuracy of image analysis but also can obtain a very fast stitching speed with the acceleration of the GPU, providing a brand-new technical platform for further exploring eye diseases.
[0098] An embodiment of the present invention also provides a storage medium for storing a computer program, which when executed, at least executes the method described above.
[0099] An embodiment of the present invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein, the processor is used to execute the computer program to at least execute the method described above.
[0100] An embodiment of the present invention also provides a processor, which executes a computer program to at least execute the method described above.
[0101] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memories.
[0102] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0103] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0104] In addition, in each embodiment of the present invention, each functional unit can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.
[0105] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments. The aforementioned storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0106] Alternatively, if the above integrated units of the present invention are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0107] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0108] The features disclosed in the several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0109] The features disclosed in the several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0110] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention belongs, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and as long as the performance or use is the same, they should all be regarded as falling within the protection scope of the present invention.
Claims
1. A method for automatic rapid stitching of multiple images for corneal laser confocal microscopy, characterized in that: The following steps are involved: S1. Data acquisition: corneal images are collected using a corneal laser confocal microscope to construct an initial data set, wherein the images contain structural information of corneal tissue; S2, controlled data enhancement: the initial data set is screened to remove out-of-focus or invalid images; the remaining images are brightness corrected to make their brightness consistent; a fixed-size image block is cropped in the center area of the image as a reference image, and the image block is randomly displaced to generate a target image and a corresponding offset label, and a synthetic data set with labels is constructed; S3. Neural network model construction and training: Build a neural network model, which includes a feature extractor, a feature maximum correlation matching layer and a regression subnetwork prediction head; use the synthetic data set to train the model, extract image features through the feature extractor, calculate the correlation between features through the feature maximum correlation matching layer, and predict the offset between image pairs based on the correlation results by the regression subnetwork prediction head, and optimize the model using a mixed loss function until the model can accurately predict the offset; wherein, the feature extractor is composed of a plurality of convolution blocks stacked together, each of which contains a two-dimensional convolution operation and a nonlinear activation function, and except for the last convolution block, the remaining convolution blocks are downsampled through a maximum pooling layer; the feature extractor extracts features from the reference image and the target image to form feature maps of the same size and share weights; use feature maps of different scales to predict the offset to capture the offset information of the image; wherein, by predicting the offset step by step at different scales, the displacement of the image pair is estimated at different scales to achieve accurate prediction; S4, image stitching: pre-process the multiple corneal images collected, adjust the image size and perform brightness correction; select adjacent image pairs in turn and input them into the trained neural network model to predict the offset between the image pairs; when the predicted offset exceeds the preset threshold, terminate the stitching of the current group of images and use the previous images as a group of images to be stitched; for each group of images to be stitched, calculate the global coordinate system according to the predicted offset, and fuse and stitch the images to generate a corneal image with a large field of view.
2. The method for automatic rapid stitching of multiple images for corneal laser confocal microscopy according to claim 1, characterized in that: The specific process of step S2 includes: The initial dataset was screened to remove images that were out of focus or did not contain corneal structural information; Perform brightness correction on the screened images to make the image brightness consistent, so as to reduce boundary artifacts during stitching and highlight the main structure of the cornea; Crop a fixed-size image patch in the center area of the image as a reference image; Performing random displacement perturbation on the cropped image block to generate a target image and a corresponding offset label, wherein the displacement perturbation remains consistent in the horizontal and vertical directions and the displacement range is a preset value; Through the above specific process, image pairs of reference images and target images and offset labels between the image pairs are generated to form a synthetic data set for training and testing of neural network models.
3. The method for automatic rapid stitching of multiple images for corneal laser confocal microscopy according to claim 1 or 2, characterized in that: The offset prediction includes: At the first scale, deep feature maps are used for prediction, which contain rich semantic information; The deep feature maps of the reference image and the target image are input into the feature maximum correlation matching layer, and the correlation of different positions of the feature maps is calculated to obtain the correlation result; Inputting the correlation result into the prediction head of the first regression subnetwork to predict the displacement of the image pair in the x and y directions; At the second scale, shallow feature maps are used for prediction, which contain more edge, texture and structure information; The shallow feature map of the target image is deformed according to the prediction result of the first scale, and is input into the feature maximum correlation matching layer together with the shallow feature map of the reference image to obtain the correlation result of the second scale; The correlation result of the second scale is input into the prediction head of the second regression sub-network to predict the displacement of the image pair in the x and y directions.
4. The method for automatic rapid stitching of multiple images for corneal laser confocal microscopy according to claim 3, characterized in that: The final offset prediction result is obtained by fusing the prediction results of the first scale and the second scale, wherein the fusion method is to add the prediction results of the two scales.
5. The method for automatic rapid stitching of multiple images for corneal laser confocal microscopy according to claim 1 or 2, characterized in that: In step S3, the neural network model training further includes: Construct a loss function consisting of two parts. The first part is L2 loss, which is used to make the network focus on the overall displacement information. The second part is deformation loss, which is used to make the network focus on the content differences in the overlapping areas of the image. The L2 loss is composed of the average of the L2 losses at two different scales, where the L2 loss of the first scale is based on the prediction result of a single scale, and the L2 loss of the second scale is based on the difference between the accumulated value of the prediction results of the two scales and the actual displacement; The deformation loss calculates the content difference between the deformed image and the reference image in the overlapping area by applying the same deformation matrix to the mask and the input image; The total loss is a weighted sum of the L2 loss and the deformation loss, where the weights of the L2 loss and the deformation loss are set to balance the attention paid to the overall displacement information and the content difference of the overlapping regions.
6. The method for automatic rapid stitching of multiple images for corneal laser confocal microscopy according to claim 1 or 2, characterized in that: The specific process of step S4 includes: Use the trained neural network model to test on synthetic data sets to verify the accuracy and speed of the model; Deploy the model and create Web service interfaces, including the health check interface and the prediction interface. The prediction interface is used to receive input data and call the model for prediction. Collect multiple real images, resize them and correct their brightness to ensure the consistency of input data and reduce boundary artifacts; Starting from the first image, sample adjacent image pairs in sequence and call the API interface to predict the offset of the image pairs; When the predicted offset exceeds a preset threshold, the prediction of the current image group is terminated, and the image group is marked as an image group to be spliced; Repeat the above specific process until all image predictions are completed, and obtain multiple groups of image groups to be spliced; For each group of images to be stitched, the global coordinate system is calculated according to the predicted local offset, and the images are fused to obtain the stitching result.
7. The method for automatic rapid stitching of multiple images for corneal laser confocal microscopy according to claim 6, characterized in that: Calculating the global coordinate system according to the predicted local offsets includes: multiplying each group of predicted local offsets by a scaling factor and performing calculations to obtain a global coordinate system, and gradually performing image fusion.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for automatically and quickly stitching multiple images for corneal laser confocal microscopy as described in any one of claims 1 to 7 is implemented.
9. A computer program product, characterized in that When the computer program is executed by a processor, the method for automatically and quickly stitching multiple images for corneal laser confocal microscopy as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image splicing method based on pyramid structure super-resolution network
CN115841422A
Image splicing method and system for optimizing splicing quality based on deep learning
CN118644385A