Multi-focal-length embryo image fusion method based on multi-gradient collaborative fusion and structure global decoding
Through the method of multi-gradient collaborative fusion and structural global decoding, the subjectivity and stability problems of the image evaluation process in multi-focal length embryo image processing are solved, efficient and automated image fusion is achieved, clear and complete embryo images are generated, and the stability and consistency of image processing are improved.
Patent Information
- Application Number
- CN202510871761.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The existing technology in multi-focal embryo image processing has the following problems: the image evaluation process is highly subjective, and the stability and generalization ability are weak, which makes it difficult to meet the needs of high-throughput image analysis. In addition, traditional methods cannot effectively deal with the problem of inconsistent spatial focal length distribution in multi-focal length images.
A multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and global structural decoding is adopted. An end-to-end fusion network model is constructed through multi-scale gradient feature extraction, feature fusion and global decoding of embryonic structure, and optimized using self-supervised gradient feature alignment loss and unsupervised structural similarity loss.
It realizes automated and efficient multi-focal length embryo image fusion, generates clear and complete fused images, improves the stability and consistency of image processing, and enhances the image focusing quality and structural restoration capability.
Smart Images

Figure CN120689222A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image autofocus and image fusion, and in particular to a multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding. Background Art
[0002] In vitro fertilization (IVF), one of the most widely used and technologically mature assisted reproductive methods, comprises the core process of egg collection, IVF, embryo culture, and the selection and transfer of high-quality embryos. Embryo quality assessment and screening directly determine the ultimate pregnancy success rate and are crucial to the entire IVF treatment process. Clinically, physicians typically rely on embryo development images captured in delayed culture systems to systematically analyze embryo morphological characteristics to assess their developmental potential and implantation likelihood. This assessment process is not only the core basis for embryo selection but also a key factor influencing the success of an IVF cycle.
[0003] Embryo images are typically captured by a camera system within a time-lapse incubator. To obtain structurally intact and sharply focused images, clinical practice generally does not rely solely on fixed focal lengths for image acquisition. Due to factors such as focal length settings, embryo position offset, and uneven culture dish thickness, single images often appear blurry or lack information. Therefore, continuously acquiring multiple images at different focal lengths in a very short period of time has become standard operating procedure. Automatically and accurately selecting or fusing multi-focal length images of the same embryo at the same time point to create the clearest image has become a key challenge in current IVF image processing.
[0004] Traditional methods typically quantitatively analyze image quality based on a series of focus evaluation metrics, selecting the optimal focal length image for subsequent evaluation. These metrics include, but are not limited to, image clarity, edge sharpness, texture detail intensity, and frequency domain energy distribution. While these metrics can, to a certain extent, reflect the image's focus and help identify image samples with the clearest structures, these methods generally rely on predefined rules and manually extracted features, and are significantly affected by factors such as imaging conditions, parameter selection, and image complexity. The evaluation process is highly subjective, and stability and generalization capabilities are limited.
[0005] Furthermore, as the volume of embryonic image data continues to grow in clinical and scientific research, traditional methods are facing bottlenecks in processing efficiency and automation capabilities, making it difficult to meet the actual needs of high-throughput image analysis. Further complicating matters is the fact that in multi-focal images acquired in practice, some areas are often clear while others are blurred, indicating that the images are spatially unevenly focused. This inconsistent spatial focal length distribution makes it difficult for image selection strategies based on unified evaluation metrics to fully reflect the true interpretability and clinical value of the images. Summary of the Invention
[0006] In order to overcome the shortcomings of the existing technology in multi-focal embryo image processing, the present invention proposes a multi-focal embryo image fusion method based on multi-gradient collaborative fusion and global structural decoding, aiming to significantly improve the overall focusing quality and structural restoration ability of the fused image, thereby providing a more stable and reliable image basis for morphological evaluation and clinical interpretation.
[0007] In order to achieve the above-mentioned object, the present invention adopts the following technical solutions:
[0008] The multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding of the present invention is characterized in that it is performed according to the following steps:
[0009] Step 1: Obtain two embryo microscopic images taken at the same time with different focal lengths ,in, represents the first embryo microscopic image, represents the second embryo microscopic image, represents the length of the embryo microscopic image, represents the width of the embryo microscopic image; 3 represents the number of channels of the embryo microscopic image;
[0010] right Grayscale processing is performed to obtain the embryo microscopic grayscale image group ,in, represents the first embryo microscopic grayscale image, represents the second embryo microscopic grayscale image; 1 represents the number of channels of the embryo microscopic grayscale image;
[0011] According to the channel dimension of the grayscale image and Stack to form a multi-channel input feature map ;
[0012] take Different gradient operators are used to extract Structural gradient map group ,in, Indicates the The structural gradient map group extracted by the gradient operator, Indicates the Gradient operator extraction The structural gradient map, Indicates the Gradient operator extraction Structural gradient map of
[0013] right and Perform pixel-by-pixel maximum operation to obtain the Zhang embryo reference gradient map , thus obtaining Zhang embryo reference gradient map ;
[0014] Step 2: Construct a multi-gradient collaborative fusion network, including: a multi-scale feature extraction module guided by multiple gradient features, a feature fusion module, and Processing is performed to obtain multi-scale embryo gradient fusion features ;
[0015] Step 3: Construct the global decoding module of embryo structure, including: Patch embedding layer, rotation-invariant position encoding layer, multi-layer global decoding unit and Patch restoration layer; wherein the multi-layer global decoding unit is composed of multiple layers of Transformer decoding layers with the same structure, each Transformer decoding layer includes: multi-head attention unit, feedforward neural network, normalization and residual connection unit; and Processing to obtain focused embryo images ;
[0016] Step 4: Construct a joint loss function ;
[0017] Step 5: Use the AdamW optimizer to train the multi-focal length embryo image fusion network and minimize the joint loss function The network parameters are updated to obtain the optimal multi-focal length embryo image fusion model, which is used to realize the adaptive focus fusion of multi-focal length embryo images.
[0018] The multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding according to the present invention is also characterized in that step 2 is performed as follows:
[0019] Step 2.1: Construct a multi-scale feature extraction module, including: parallel branches, each containing Convolutional attention modules of different scales; each scale convolutional attention module includes: convolution layer, normalization layer, activation function layer and cross-channel attention mechanism layer; the cross-channel attention mechanism layer includes: pooling layer, fully connected layer, activation function layer and channel attention weighted layer; and Enter After the parallel branches, The convolution attention modules of different scales are processed to obtain the The intermediate gradient feature sequence output by each branch ,in, Indicates the The intermediate gradient features at the k-th scale output by the branch, , ;
[0020] Step 2.2, construct a feature fusion network, including: splicing unit, gating network and gradient fusion unit; among them, the gating network includes: two convolution layers, activation function layer and normalization processing unit; and the intermediate gradient feature sequence output by all branches Processing to obtain embryo fusion feature map .
[0021] Furthermore, step 2.2 is performed as follows:
[0022] Step 2.2.1, the splicing unit uses formula (1) to After the intermediate gradient features with the same scale in the branches are spliced by channel, we get Splicing features at different scales :
[0023] (1)
[0024] In formula (1), Indicates the Splicing features at different scales; Represents a splicing operation;
[0025] Step 2.2.2, Input into the gating network for processing and get the At the scale Weighted graph , thus obtaining Weight graphs at different scales ,in, Indicates the The first scale weights;
[0026] Step 2.2.3: Gradient fusion unit is obtained using formula (2) Fusion gradient features at different scales :
[0027] (2)
[0028] In formula (2), represents element-wise multiplication; Indicates the Fusion gradient features at different scales;
[0029] Step 2.2.4: Use formula (3) to fusion gradient features Perform upsampling and splicing by channel to obtain splicing features :
[0030] (3)
[0031] In formula (3), Express Perform upsampling so that its size is the same as Alignment;
[0032] Step 2.2.5, Input into the gating network for processing, and get Weighted graph The gradient fusion unit uses formula (4) to obtain the embryo fusion feature map :
[0033] (4)
[0034] In formula (4), represents the j-th weight graph.
[0035] Furthermore, step 3 is performed as follows:
[0036] Step 3.1: Embryo fusion feature map The input is divided into a set of two-dimensional image blocks in the Patch Embedding Layer, and each two-dimensional image block is mapped to a Token vector of uniform length;
[0037] Step 3.2: The rotation-invariant position encoding layer uses two-dimensional polar coordinate position encoding to add the rotation-invariant position vector to each token vector, and then fuses it with the corresponding token vector to obtain different global position-aware token vectors.
[0038] Step 3.3: Input each global location-aware token vector into a multi-layer global decoding unit for decoding to obtain the corresponding global location-aware token decoding vector.
[0039] Step 3.4, the Patch reduction layer uses linear projection to restore each global position-aware Token decoding vector to the corresponding image block vector, so as to The original spatial position of the two-dimensional image block in the image is reorganized to obtain the reconstructed focused embryo image. .
[0040] Furthermore, step 4 is performed as follows:
[0041] Step 4.1: Use formula (5) to establish the self-supervised gradient feature alignment loss:
[0042] (5)
[0043] In formula (5), Representative embryo reference gradient map The embryo reference gradient map after normalization, Indicates the A normalized embryo reference gradient map; represent The output of the branch Intermediate gradient feature sequences of different scales The intermediate gradient features after normalization, Indicates the The output of the branch The intermediate gradient features after normalization at each scale; represent The self-supervised gradient feature alignment loss of each branch, Indicates the Self-supervised gradient feature alignment loss for each branch;
[0044] Step 4.2: Use formula (6) to establish the structural similarity loss function :
[0045] (6)
[0046] In formula (6), represents the structural similarity index;
[0047] Step 4.3: Use formula (7) to construct the joint loss function :
[0048] (7)
[0049] In formula (7), represent The corresponding regularization parameter; and represents two regularization parameters.
[0050] The electronic device of the present invention includes a memory and a processor, and is characterized in that the memory is used to store a program that supports the processor to execute the multi-focal embryo image fusion method, and the processor is configured to execute the program stored in the memory.
[0051] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium. The computer-readable storage medium is characterized in that the computer program executes the steps of the multi-focal embryo image fusion method when the computer program is run by a processor.
[0052] Compared with the prior art, the beneficial effects of the present invention are embodied in:
[0053] 1. This paper proposes an end-to-end multi-focal length embryo image fusion network model that automatically processes multiple embryo images of different focal lengths and generates a clear, complete fused image. This method requires no human intervention and can effectively improve the efficiency and consistency of image processing, resolving the problems of traditional methods that rely on manual screening, are cumbersome to operate, and have poor repeatability.
[0054] 2. This paper designs a parallel multi-branch gradient feature extraction network, each branch of which contains a multi-scale convolutional attention module. This multi-scale convolutional module extracts image gradient features at different scales. Combined with a cross-channel attention mechanism, it adaptively enhances channel responses highly correlated with embryonic structure, providing multi-layered, complementary feature support for subsequent fusion.
[0055] 3. During the feature fusion stage, the present invention introduces a gated feature fusion mechanism. This mechanism first concatenates the gradient features extracted from each branch along the channel dimension. Then, a gating network dynamically generates multiple spatial weight maps, achieving pixel-by-pixel weighted integration of the features from each branch. This approach effectively enhances the structural response of clear areas while suppressing redundant or ambiguous information, thereby achieving efficient fusion of multi-scale gradient features.
[0056] 4. During the embryonic image reconstruction phase, the present invention constructs a global embryonic structural decoding network. This network employs a global modeling strategy based on rotationally invariant position encoding to achieve spatial semantic restoration of fused features within the decoding framework. By combining a self-supervised gradient feature alignment loss with an unsupervised structural similarity loss, a complete optimization objective function is constructed, effectively improving the focus and structural fidelity of the fused image. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a schematic diagram of the overall process of the method proposed in the present invention;
[0058] Figure 2 Schematic diagram of the structure of the multi-branch gradient feature extraction network in the present invention;
[0059] Figure 3 Schematic diagram of the process of multi-gradient feature fusion performed by the gated network in the present invention;
[0060] Figure 4 A visual comparison of the fusion effects of different methods on the same set of embryo images. DETAILED DESCRIPTION
[0061] In this embodiment, a multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding is firstly gray-scaled and stacked on the embryo image data; secondly, three convolutional attention modules with the same architecture are constructed; features of different scales and gradients are respectively fused using a gating method; the fused features are then fed into the embryonic structure global decoding network to obtain a multi-focal length fused embryo image; finally, a self-supervised gradient feature alignment loss and an unsupervised structural similarity loss function are established to optimize the overall network architecture to achieve fusion of the focal length embryo images; Figure 1 Specifically, the method is performed in the following steps:
[0062] Step 1: In this example, tens of thousands of embryo image data were collected from multiple reproductive centers. These data cover the entire process of embryo development from fertilized eggs to late blastocysts ready for transplantation. Embryo images are acquired based on the built-in camera in the time-lapse incubator. A set of images is taken continuously in a very short period of time, collecting two embryo images with different focal lengths. Obtain a set of two embryo microscopic images with different focal lengths taken at the same time. ,in, represents the first embryo microscopic image, represents the second embryo microscopic image, 512 represents the length of the embryo microscopic image, 612 represents the width of the embryo microscopic image; and 3 represents the number of channels of the embryo microscopic image.
[0063] right Grayscale processing is performed to obtain the embryo microscopic grayscale image group ,in, represents the first embryo microscopic grayscale image, represents the second embryo microscopic grayscale image; 1 represents the number of channels of the embryo microscopic grayscale image;
[0064] According to the channel dimension of the grayscale image and Stack to form a multi-channel input feature map .
[0065] Three different gradient operators (Canny, Tenengrad and Laplacian) are used to extract Structural gradient map group , among which express and The feature map after Canny gradient operator processing, express and The feature map after the Tenengrad gradient operator processing, express and The feature map after Laplacian gradient operator processing; then the structure gradient map group Perform pixel-by-pixel maximum operations on the graphs with the same gradient operator to obtain three unified embryo reference gradient maps. The embryonic image data used in this example comes from tens of thousands of images collected from multiple reproductive centers. This data covers the entire embryonic developmental process, from fertilized eggs to late-stage blastocysts ready for transplantation. Embryo images are captured using a built-in camera in a time-lapse incubator. A set of images is captured continuously over a very short period of time, capturing two embryo images at different focal lengths.
[0066] Step 2. In the present invention, in order to improve the perception of information of different scales and enhance the performance of fused embryo images in terms of structural integrity and local clarity, a three-branch parallel gradient feature extraction module is designed, which introduces three gradient algorithms with complementary characteristics: Canny, Tenengrad and Laplacian. In view of the actual needs of the significant scale differences in the structural features and focus of embryo images at different developmental stages, a multi-gradient collaborative fusion network is constructed, including: a multi-scale feature extraction module guided by multiple gradient features, a feature fusion module, and Processing is performed to obtain multi-scale embryo gradient fusion features , the network architecture is as follows Figure 1 As shown:
[0067] Step 2.1. Due to the actual needs of different structural scales that need to be paid attention to in embryonic images at different developmental stages, a unified multi-scale feature extraction module is constructed to cope with the characteristics of embryonic images at different developmental stages, including: three parallel branches, each branch contains three convolutional attention modules of different scales, guided by the three classic gradient algorithms of Canny, Tenengrad and Laplacian respectively; among them, the convolutional attention module of each scale includes: convolution layer, normalization layer, activation function layer and cross-channel attention mechanism layer; among them, the convolution layer sets appropriate padding parameters to ensure that the image space size remains unchanged after the convolution operation. First, the multi-channel input feature map Through the first convolution block, using The convolution kernel is set with padding of 1 to keep the spatial size unchanged, and the number of channels is increased from 2 to 16. Batch normalization is then performed to stabilize the feature distribution, and the ReLU activation function is used to enhance the nonlinear expression capability. The output feature size is Then, the feature map is used as input to the second convolution block, and the convolution layer uses The convolution kernel, padding is set to 2, the number of channels is further increased to 32, and after normalization and ReLU activation, the output size is Then enter the third convolution block and use Convolution kernel, padding is 3, the number of channels is expanded to 64, and the normalization and activation operations are repeated, and the final output is Through this structure, the model achieves the layer-by-layer expansion of the receptive field and the gradual enhancement of channel information while keeping the spatial size unchanged, thereby effectively integrating multi-level gradient information such as local edges, detailed textures and global structures in the embryo image, providing richer and more stable feature support for subsequent fusion processing. The cross-channel attention mechanism layer includes: pooling layer, fully connected layer, activation function layer and channel attention weighted layer. The overall structure is as follows Figure 2 As shown in the figure. First, the pooling method used by the pooling layer is global average pooling, which compresses the spatial dimension of each channel into a scalar, thereby forming a global statistical feature at the channel level; the features after the pooling layer are input to the first fully connected layer, reducing the feature channel to 1 / 8 of the original; then the ReLu activation function is used to introduce nonlinear mapping to enhance the discriminative ability of channel features, and then the number of channels is restored to the original input channels through the second fully connected layer. Finally, the Sigmoid activation function is used for normalization to obtain the attention weight corresponding to each channel; and The input is fed into three parallel branches and processed by three convolutional attention modules of different scales in turn, thereby obtaining the intermediate gradients of three different scales output by the first branch. , the intermediate gradients of three different scales output by the second branch And the intermediate gradients of three different scales output by the third branch .
[0068] Step 2.2: Construct a feature fusion network, including a splicing unit, a gating network, and a gradient fusion unit. The gating network includes two convolutional layers, an activation function layer, and a normalization processing unit. Figure 3 As shown, the convolution block 1 and convolution block 2 in the gated network have the same network architecture, where convolution block 1 uses a convolution kernel size of , the activation function is ReLU; the convolution block 2 uses a convolution kernel size of , the activation function is Softmax; the output dimension of the gating network is consistent with the spatial size of the input embryo feature map, and the three weight maps generated will act as channel weighting factors on the corresponding intermediate gradient features respectively, realizing the weighted fusion of different branch features in the spatial dimension, thereby enhancing the structural response of the effective area, suppressing redundant or interfering information, and providing a stable and reliable weight basis for the subsequent fusion embryo image generation; and the intermediate gradient feature sequences output by all branches are Processing to obtain embryo fusion feature map .
[0069] Step 2.2.1: To achieve unified expression of feature information from different gradient branches at the same scale, the splicing unit uses formula (1) to splice the intermediate gradient features of the same scale in the three branches by channel, and obtains three splicing features of different scales. :
[0070] (1)
[0071] In formula (1), Represents the splicing features at three scales respectively; Represents a splicing operation.
[0072] Step 2.2.2: After completing the channel-wise splicing of the intermediate gradient features of the three different branches, in order to further achieve the adaptive control of the contribution of each branch feature, the spliced features at the three scales are Input into the gating network for processing, and the features at each scale obtain three weight maps ,in Indicates the Three weights at different scales, thus obtaining weight graphs at three different scales ,like Figure 3 shown.
[0073] Step 2.2.3: The gradient fusion unit uses formula (2) to obtain the fusion gradient features at three different scales. :
[0074] (2)
[0075] In formula (2), represents element-wise multiplication; Represent the fusion features at three scales respectively; the weighted process can realize the dynamic allocation of importance for different areas in the image, so that the branch features with stronger responses obtain greater retention weights at the structural boundaries, detail areas or clear focus, while in redundant or fuzzy areas, the corresponding feature channel contributions are automatically weakened.
[0076] Step 2.2.4: Use formula (3) to fusion gradient features Perform upsampling and splicing by channel to obtain splicing features :
[0077] (3)
[0078] In formula (3), , Express Perform upsampling so that the size is the same as Alignment,The purpose of upsampling is to achieve channel alignment between,features of different scales.
[0079] Step 2.2.5, Input into the gating network for processing to obtain three weight maps , and then use the gradient fusion unit to obtain the embryo fusion feature map using formula (4) :
[0080] (4)
[0081] Step 3: Construct the global decoding module of embryo structure, including: Patch embedding layer, rotation-invariant position encoding layer, multi-layer global decoding unit and Patch restoration layer; the multi-layer global decoding unit is composed of multiple layers of Transformer decoding layers with the same structure, each Transformer decoding layer includes multi-head attention unit, feedforward neural network, normalization and residual connection unit; and Processing to obtain focused embryo images .
[0082] Step 3.1: Embryo fusion feature map The input is divided into a set of two-dimensional image blocks in the Patch Embedding Layer, and each two-dimensional image block is mapped to a Token vector of uniform length;
[0083] Step 3.2: The rotation-invariant position encoding layer uses two-dimensional polar coordinate position encoding to add the rotation-invariant position vector to be learned to each token vector, and then fuses it with the corresponding token vector to obtain different global position-aware token vectors. Polar coordinate encoding significantly improves the model's tolerance for random embryo orientation (such as a ±90° deviation in the placement of the culture dish), avoiding structural misalignment caused by traditional Cartesian coordinates. At the same time, the rotation-invariant position encoding vector can be initialized as a learnable parameter and optimized along with the network during training, thereby achieving long-term memory and adaptation of the relative and absolute position information of the embryo image block.
[0084] Step 3.3: Input each global location-aware token vector into a multi-layer global decoding unit for decoding to obtain the corresponding global location-aware token decoding vector.
[0085] Step 3.4, the Patch reduction layer uses linear projection to restore each global position-aware Token decoding vector to the corresponding image block vector, so as to The original spatial position of the two-dimensional image block in the image is reorganized to obtain the reconstructed focused embryo image. ; Through the two-dimensional embryo image block restoration and reconstruction process, the semantic features output by the decoding network are successfully restored to embryo image representations with real spatial structure and focus details, realizing end-to-end reconstruction from multi-focal length embryo image fusion features to a single clear embryo image.
[0086] Step 4: Construct a joint loss function :
[0087] Step 4.1: In order to guide multiple independent gradient branches in the fusion network to learn their corresponding features respectively, realize gradient semantic decoupling, and make the fusion process have clear feature sources and interpretability, the self-supervised gradient feature alignment loss is established using formula (5):
[0088] (5)
[0089] In formula (5), Representative embryo reference gradient map The normalized embryo reference gradient map corresponds to the feature maps processed by the three gradient operators Canny, Tenengrad and Laplacian in the embryo image; Represents the third scale of the three branches The intermediate gradient features after normalization correspond to the Canny, Tenengrad, and Laplacian gradient features in the embryo image learned by the three gradient branches respectively; the L1 norm is used to encourage the model to fit pixel by pixel, making it easy to train and converge stably; the normalization process maintains the consistency of the numerical scale. 、 and Represents the self-supervised gradient feature alignment loss in the three branches, which is used to balance the influence weights of feature alignment losses of different branches.
[0090] Step 4.2: Focused embryo image for model reconstruction Embryo images with multifocal length In terms of consistency at the structural level, the output image of the fusion network decoder is improved to retain the information of the original embryo image in terms of spatial structure and texture details. The final reconstructed embryo image is optimized under unsupervised conditions. A structural similarity loss function is introduced. This loss uses each input image in the multi-focal length embryo image as a reference to compare the fidelity of the final fusion image in terms of local structure, brightness and contrast. It aims to constrain the image reconstruction effect from the perceptual level. The structural similarity loss function is established using formula (6): :
[0091] (6)
[0092] In formula (6), represents the structural similarity index; The smaller the loss function value, the closer the fused embryo image is to the original image at the structural level. By minimizing this loss during training, the fused embryo image's ability to retain the original structural information is effectively enhanced, improving its clarity, edge fidelity, and visual consistency, thereby improving the overall quality of embryo image fusion.
[0093] Step 4.3: Use formula (7) to construct the joint loss function :
[0094] (7)
[0095] In formula (7), 、 and Represents the regularization parameter corresponding to the gradient feature alignment loss; and Represents the regularization parameters of the gradient feature alignment loss and structural similarity loss function in network optimization, which is used to balance the optimization objectives between self-supervised embryo image structural fidelity and local branch feature guidance, so that the network has the ability to learn multiple gradient features of embryo images while maintaining structural consistency.
[0096] Step 5: In the training phase of the present invention, the pre-processed embryo image data is input into the multi-focal length embryo image fusion network for training. The network realizes the joint perception and feature integration of structural information and clear areas in the multi-focal length embryo image by introducing a gradient-guided feature extraction module and a cross-channel attention mechanism. During the training process, the AdamW optimizer is used to train the multi-focal length embryo image fusion network and minimize the joint loss function. The network parameters are updated to obtain the optimal multi-focal length embryo image fusion model, which is used to realize the adaptive focus fusion of multi-focal length embryo images.
[0097] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0098] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are executed.
[0099] In the specific implementation, the network model under the joint optimization framework is used to evaluate the results of the test set. Experiments show that the regularization parameter of the hybrid loss , In addition, the present invention conducted a large number of comparative experiments on 50 sets of embryo image data sets to verify the effectiveness of the proposed method. The experimental results are shown in Table 1:
[0100] Table 1. Experimental results of different models on embryo image dataset
[0101]
[0102] The experimental results in Table 1 show that the proposed multi-focal embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding has high performance in six mainstream image fusion evaluation indicators: normalized mutual information , human visual perception indicators , gradient fidelity , phase consistency characteristic index , structural similarity index and visual information fidelity They are all superior to other methods (1. CNN method in the Medical image fusion based on convolutional neural networks and non-subsampled contourlet transform article, 2. U-Net method in the Fast multi-focus fusion based on deep learning for early-stage embryo image enhancement article, 3. Transformer method in the Image fusion transformer article, and 4. U2Fusion method in the Unified unsupervised image fusion network article). The index reaches 0.7244, which is about 16% higher than the best comparison method U2Fusion. Figure 4 The figure shows the visualization results of image fusion using different methods. It can be seen from the figure that the image fused by the method proposed in the present invention can achieve the best, showing the significant advantages of the method of the present invention (Ours) in structure restoration and clarity enhancement.
[0103] To further validate the effectiveness of the key components of the model architecture, the authors designed a series of ablation experiments, focusing on evaluating the contributions of three gradient-feature-guided branches (Canny, Tenengrad, and Laplacian) to overall fusion performance. The results showed that retaining only a single gradient branch generally degraded the fusion performance, particularly when using only the Laplacian branch, where all metrics were significantly inferior to the full model. However, by simultaneously introducing all three gradient information and synergistically fusing them through a gating mechanism, the model achieved optimal results across all metrics. This demonstrates the complementary nature of the three gradient operators in capturing different structural features, such as edges, details, and grayscale transitions. The synergistic fusion mechanism can significantly improve the overall focus quality and interpretability of multi-focal-length embryo images.
Claims
1. A multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding, characterized in that: The steps are as follows: Step 1: Obtain two embryo microscopic images taken at the same time with different focal lengths ,in, represents the first embryo microscopic image, represents the second embryo microscopic image, represents the length of the embryo microscopic image, represents the width of the embryo microscopic image; 3 represents the number of channels of the embryo microscopic image; right Grayscale processing is performed to obtain the embryo microscopic grayscale image group ,in, represents the first embryo microscopic grayscale image, represents the second embryo microscopic grayscale image; 1 represents the number of channels of the embryo microscopic grayscale image; According to the channel dimension of the grayscale image and Stack to form a multi-channel input feature map ; take Different gradient operators are used to extract Structural gradient map group ,in, Indicates the The structural gradient map group extracted by the gradient operator, Indicates the Gradient operator extraction The structural gradient map, Indicates the Gradient operator extraction Structural gradient map of right and Perform pixel-by-pixel maximum operation to obtain the Zhang embryo reference gradient map , thus obtaining Zhang embryo reference gradient map ; Step 2: Construct a multi-gradient collaborative fusion network, including: a multi-scale feature extraction module guided by multiple gradient features, a feature fusion module, and Processing is performed to obtain multi-scale embryo gradient fusion features ; Step 3: Construct the global decoding module of embryo structure, including: Patch embedding layer, rotation-invariant position encoding layer, multi-layer global decoding unit and Patch restoration layer; wherein the multi-layer global decoding unit is composed of multiple layers of Transformer decoding layers with the same structure, each Transformer decoding layer includes: multi-head attention unit, feedforward neural network, normalization and residual connection unit; and Processing to obtain focused embryo images ; Step 4: Construct a joint loss function ; Step 5: Use the AdamW optimizer to train the multi-focal length embryo image fusion network and minimize the joint loss function The network parameters are updated to obtain the optimal multi-focal length embryo image fusion model, which is used to realize the adaptive focus fusion of multi-focal length embryo images.
2. The multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding according to claim 1 is characterized in that: Step 2 is performed as follows: Step 2.1: Construct a multi-scale feature extraction module, including: parallel branches, each containing Convolutional attention modules of different scales; each scale convolutional attention module includes: convolution layer, normalization layer, activation function layer and cross-channel attention mechanism layer; the cross-channel attention mechanism layer includes: pooling layer, fully connected layer, activation function layer and channel attention weighted layer; and Enter After the parallel branches, The convolution attention modules of different scales are processed to obtain the The intermediate gradient feature sequence output by each branch ,in, Indicates the The intermediate gradient features at the k-th scale output by the branch, , ; Step 2.2, construct a feature fusion network, including: splicing unit, gating network and gradient fusion unit; among them, the gating network includes: two convolution layers, activation function layer and normalization processing unit; and the intermediate gradient feature sequence output by all branches Processing to obtain embryo fusion feature map .
3. The multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding according to claim 2 is characterized in that: Step 2.2 is performed as follows: Step 2.2.1, the splicing unit uses formula (1) to After the intermediate gradient features with the same scale in the branches are spliced by channel, we get Splicing features at different scales : (1) In formula (1), Indicates the Splicing features at different scales; Represents a splicing operation; Step 2.2.2, Input into the gating network for processing and get the At the scale Weighted graph , thus obtaining Weight graphs at different scales ,in, Indicates the The first scale weights; Step 2.2.3: Gradient fusion unit is obtained using formula (2) Fusion gradient features at different scales : (2) In formula (2), represents element-wise multiplication; Indicates the Fusion gradient features at different scales; Step 2.2.4: Use formula (3) to fusion gradient features Perform upsampling and splicing by channel to obtain splicing features : (3) In formula (3), Express Perform upsampling so that its size is the same as Alignment; Step 2.2.5, Input into the gating network for processing, and get Weighted graph The gradient fusion unit uses formula (4) to obtain the embryo fusion feature map : (4) In formula (4), represents the j-th weight graph.
4. The multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding according to claim 3 is characterized in that: Step 3 is performed as follows: Step 3.1: Embryo fusion feature map The input is divided into a set of two-dimensional image blocks in the Patch Embedding Layer, and each two-dimensional image block is mapped to a Token vector of uniform length; Step 3.2: The rotation-invariant position encoding layer uses two-dimensional polar coordinate position encoding to add the rotation-invariant position vector to each token vector, and then fuses it with the corresponding token vector to obtain different global position-aware token vectors. Step 3.3: Input each global location-aware token vector into a multi-layer global decoding unit for decoding to obtain the corresponding global location-aware token decoding vector. Step 3.4, the Patch reduction layer uses linear projection to restore each global position-aware Token decoding vector to the corresponding image block vector, so as to The original spatial position of the two-dimensional image block in the image is reorganized to obtain the reconstructed focused embryo image. .
5. The multi-focal length embryo image fusion method based on multi-gradient collaborative fusion and structural global decoding according to claim 4 is characterized in that: Step 4 is performed as follows: Step 4.1: Use formula (5) to establish the self-supervised gradient feature alignment loss: (5) In formula (5), Representative embryo reference gradient map The embryo reference gradient map after normalization, Indicates the A normalized embryo reference gradient map; represent The output of the branch Intermediate gradient feature sequences of different scales The intermediate gradient features after normalization, Indicates the The output of the branch The intermediate gradient features after normalization at each scale; represent The self-supervised gradient feature alignment loss of each branch, Indicates the Self-supervised gradient feature alignment loss for each branch; Step 4.2: Use formula (6) to establish the structural similarity loss function : (6) In formula (6), represents the structural similarity index; Step 4.3: Use formula (7) to construct the joint loss function : (7) In formula (7), represent The corresponding regularization parameter; and represents two regularization parameters.
6. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the multi-focal embryo image fusion method according to any one of claims 1 to 5, and the processor is configured to execute the program stored in the memory.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-focal embryo image fusion method according to any one of claims 1 to 5 are executed.
Citation Information
Patent Citations
Polarization image fusion method based on double attention mechanism generative adversarial network
CN117649349A
Skin lesion image segmentation method, system and device and storage medium
CN117893545A
Method for image motion deblurring, apparatus, electronic device and medium therefor
US20240404025A1
Cited By
A method, system, and medium for multi-focus view morphological analysis of embryo images
CN122391243A
A method, system, and medium for multi-focus view morphological analysis of embryo images
CN122391243B