An embryo implantation prediction device based on multimodal fusion multifocal plane images

By constructing a multimodal fusion embryo multifocal plane image implantation prediction model and utilizing the channel-space separation multi-head attention mechanism SMHA to calculate key features, the problems of fuzzy redundancy and incomplete information interaction in non-focused areas of blastocyst multifocal images are solved, thereby improving the accuracy and efficiency of embryo implantation prediction.

CN116433612BActive Publication Date: 2026-04-03ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for embryo implantation prediction using multifocal plane images have significant redundancy and blurring in the non-focused areas of blastocyst-stage multifocal images. This results in insufficient information exchange, making it impossible to fully extract information from blastocyst-stage multifocal images. Furthermore, the computational load is too large, leading to inaccurate embryo implantation prediction.

Method used

An embryo multifocal plane image implantation prediction device based on multimodal fusion is adopted. By constructing a multimodal fusion embryo multifocal plane image implantation prediction model, including a core image generation module, a deep feature extraction module, a feature fusion module, and an implantation prediction module, key features are calculated using the channel-space separation multi-head attention mechanism SMHA, and the model parameters are optimized by the cross-entropy function to reduce the amount of computation and improve information interaction and prediction accuracy.

Benefits of technology

It improves the accuracy of embryo implantation prediction, helps embryologists select high-quality embryos more accurately and efficiently, increases the success rate of embryo transfer, and reduces the time and computational load for training sample preparation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116433612B_ABST
    Figure CN116433612B_ABST
Patent Text Reader

Abstract

This invention discloses an embryo implantation prediction device based on multimodal fusion of multifocal plane images, belonging to the field of medical image processing technology. The device includes: extracting depth features from three focal plane images of the embryo and depth features from a core image; calculating key features of the three focal planes using a channel-space separation multi-head attention mechanism; then concatenating the key features of the three focal planes with the depth features of the core image to obtain fused features; and using these fused features for embryo implantation prediction, thereby improving prediction efficiency and accuracy. This helps embryologists more accurately and efficiently classify and select high-quality embryos for transplantation, thus increasing the embryo implantation success rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing technology, specifically, it relates to an implantation prediction device for embryo multifocal plane images based on multimodal fusion. Background Technology

[0002] In vitro fertilization (IVF) is one of the most common medical methods for treating infertility. Multiple embryo transfer is a common technique in IVF, but it increases the inherent risks of pregnancy. Therefore, selecting the highest quality embryos during embryo transfer is crucial. Typically, embryo transfer occurs on days 3 (D3) and days 5 (D5), also known as cleavage-stage embryo transfer and blastocyst-stage embryo transfer, respectively. Studies have shown that blastocyst-stage embryo transfer helps improve implantation success rates.

[0003] For D5 embryos, embryologists used incubators with built-in microscopes for culture. This method provides a stable recording environment while avoiding the impact of removing embryos from the incubator on embryonic development quality. At the end of day five, embryologists captured images at different focal planes using a microscope, recording the zona pellucida (ZP), inner cell mass (ICM), and trophoblast (TE) of the blastocyst. Embryos were scored by observing these three focal plane images of embryonic development. Studies show that high-quality embryos have a greater chance of implantation, while poor-quality embryos have correspondingly greater difficulty in implantation.

[0004] Blastocyst development is divided into stages 1-6, with higher stages indicating more complete development. Stage 3 and later blastocysts are commonly used for embryo transfer. For stage 3 and later blastocysts, the inner cell mass and trophoblast can be graded using three A, B, and C grades, representing quality from highest to lowest. Grade A indicates a large number of tightly packed cells; Grade B indicates a loosely structured epithelial cell layer with few cells; and Grade C indicates a very small number of sparsely packed cells. Traditional blastocyst grading methods rely on embryologists manually analyzing multiple focal plane images to arrive at a grade. However, this manual analysis process is very cumbersome, and the results obtained through human scoring—such as "many," "few," "tight," and "loose"—are subjective, leading to significant differences in results among different embryologists.

[0005] With the development of artificial intelligence technology and medical image analysis, machine learning and deep learning methods are being applied to assist in the diagnosis of medical images, helping experts make more accurate diagnoses. Deep learning technology is also being used in IVF embryo implantation prediction using multiple focal plane (MFP) images to assist embryologists in grading and classifying blastocysts to select high-quality embryos and improve the mother's pregnancy success rate.

[0006] To assist embryologists in grading and classifying embryos to select high-quality embryos, multifocal plane images of embryos at the blastocyst stage need to be recorded. However, existing methods for embryo implantation prediction using multifocal plane images still have the following problems: During the blastocyst implantation classification process, embryologists grade the focal plane images at ZP, ICM, and TE stages separately to score the overall developmental quality of the blastocyst and select well-developed embryos for multiple embryo transfer. However, multifocal images at the blastocyst stage have significant blurring and redundancy in non-focused areas, are difficult to fuse, have insufficient information exchange, cannot fully extract information from multifocal images at the blastocyst stage, and involve excessive data computation. These problems prevent accurate embryo grading and classification, thus hindering the selection of high-quality embryos for transfer. Summary of the Invention

[0007] In view of the above, the purpose of this invention is to provide an embryo implantation prediction device based on multimodal fusion of multifocal plane images, which improves the accuracy of embryo implantation prediction by combining multimodal focal plane images.

[0008] To achieve the above-mentioned objective, this invention provides an embryo multifocal plane image implantation prediction device based on multimodal fusion, comprising a memory and a processor. The memory stores a computer-executable program for performing embryo multifocal plane image implantation prediction based on multimodal fusion. The processor is communicatively connected to the memory and configured to execute the computer-executable program stored in the memory. When the processor executes the computer-executable program, it performs the following steps:

[0009] For the zona pellucida, inner cell mass, and trophoblast at day 5 of the blastocyst stage, implantation tags were labeled to obtain multifocal image samples of the embryo.

[0010] A multimodal fusion embryo multifocal plane image implantation prediction model is constructed, including a core image generation module, a depth feature extraction module, a feature fusion module, and an implantation prediction module. The core image generation module is used to generate a core image by adaptively weighting three focal plane images. The depth feature extraction module is used to extract the corresponding depth features of multiple focal plane images and the core image in their respective extraction networks. The feature fusion module is used to calculate key features by using the channel-space separation multi-head attention mechanism SMHA to calculate the depth features extracted from the three focal plane images and the depth features extracted from the core image. After dimensionality reduction of the three key features, they are concatenated with the corresponding depth features of the core image to obtain fused features, so as to enhance the information interaction between the various focal plane images. The fused features are used for the final prediction. The implantation prediction module is used to predict embryo implantation based on the input fused features.

[0011] All embryo multifocal plane image samples were fed into the embryo multifocal plane image implantation prediction model for training, and the model was continuously optimized by updating the model parameters.

[0012] Embryo implantation prediction was performed using a parameter-optimized multifocal plane image implantation prediction model.

[0013] Preferably, the core image generation module uses convolution operations to calculate the region weights of the three focal plane images, and then combines the three focal plane images in a weighted manner to generate the core image.

[0014] Preferably, the depth feature extraction module includes extraction networks for each of the three focal plane images and an extraction network for the core image. The extraction networks use residual networks to extract the depth features of the multiple focal plane images and the depth features of the core image.

[0015] Preferably, the channel-space separated multi-head attention mechanism (SMHA) calculates key features, including: the channel-space separated multi-head attention mechanism (SMHA) only exists after the last two feature extraction layers of the residual network, assuming the feature shape is C×H×W, where H represents the height of the vector feature, W represents the width, C represents the number of channels, and f q f represents the input query vector. kv Refers to key vectors and value vectors.

[0016] f q and f kv After the input space SMHA is transformed into a two-dimensional matrix through average pooling and reshape operations, it undergoes spatial SMHA as follows:

[0017] Spatial-SMHA(f q ,f kv ) = MHA(AvgPool(f q),f kv )

[0018] The AvgPool operation, as an average pooling operation, converts the input f... q Transform it into a 2D matrix of shape 1×C, and then use a reshape operation to transform the input f. kv Transform it into a two-dimensional matrix of shape (H×W)×C, and use these two matrices to calculate the MHA value;

[0019] f q and f kv The input is transformed into a two-dimensional matrix through convolution, specifically:

[0020] Channel-SMHA(f q ,f kv )=MHA(Conv(f q ),f kv )

[0021] The Conv operation is a convolution operation that transforms the input f... q Transform the output channel of a convolutional layer with a kernel size of 1×1 into a 1×(H×W) two-dimensional matrix, and then use a reshape operation to transform f kv Transform it into a two-dimensional matrix of C×(H×W), and use these two matrices to calculate the MHA value.

[0022] Combining the two operations above, the depth features f of the core image are... core Under supervision, the key features f' corresponding to the three focal planes i The extraction steps are as follows:

[0023] f` i =Channel-SMHA(f core Spatial-SMHA(f i ,f i ))

[0024] Where f i Used to refer to the depth features corresponding to the three focal plane images: zona pellucida, inner cell mass, and trophoblast.

[0025] Preferably, the feature fusion module concatenates the three key features with the corresponding depth features of the core image to obtain fused features, including: concatenating the three key features with the depth features of the core image after dimensionality reduction, and further fusion of the concatenated depth features to the original number of channels to obtain fused features.

[0026] Preferably, the implantation prediction module uses a residual network to input the fused features into a fully connected layer for calculation and outputs the embryo implantation prediction result.

[0027] Preferably, the embryo multi-focal plane image implantation prediction model further includes multiple sets of alternately connected deep feature extraction modules and feature fusion modules. The deep feature extraction module is used to extract corresponding high-order deep features from the three key features and fusion features input from the previous group in their respective extraction networks. The feature fusion module is used to calculate high-order key features by combining the three extracted high-order deep features with the high-order deep features extracted from the fusion features through a channel-space separation multi-head attention mechanism. After dimensionality reduction of the three high-order key features, they are concatenated with the high-order deep features corresponding to the fusion features to obtain high-order fusion features. The three high-order key features and the high-order fusion features serve as inputs to the next set of deep feature extraction modules.

[0028] Preferably, during training, the cross-entropy function is used to calculate the loss between the embryo implantation prediction result and the implantation label, and the model parameters are updated to continuously optimize the model.

[0029] To achieve the above-mentioned objectives, the present invention also provides an embryo multi-focal plane image implantation prediction device based on multimodal fusion, characterized in that it includes a data acquisition unit, a model construction unit, a training unit, and an application unit.

[0030] The data acquisition unit is used to label implantation tags on three focal plane images of the zona pellucida, inner cell mass, and trophoblast at the fifth day of the blastocyst stage to obtain embryo multifocal plane image samples.

[0031] The model building unit is used to construct a multimodal fusion embryo multifocal plane image prediction model, including a core image generation module, a depth feature extraction module, a feature fusion module, and an embryo multifocal plane image implantation prediction module. The core image generation module is used to generate a core image by performing adaptive weighting on batch-enhanced multifocal plane image samples. The depth feature extraction module is used to extract the corresponding depth feature vectors of multiple focal plane images and the core image in their respective extraction networks. The feature fusion module is used to fuse the depth feature vectors extracted from multiple focal plane images with the depth feature vectors extracted from the core image through an improved multi-head attention mechanism (SMHA) to enhance the information interaction between the various focal plane images. After multiple feature extraction layers and feature fusion layers, the final output core image fusion feature is used for the final prediction. The embryo multifocal plane image implantation prediction module is used to predict embryo implantation based on the core image fusion feature.

[0032] The training unit is used to send all embryo multifocal plane image samples into the multimodal fusion embryo multifocal plane image implantation prediction device for training, and continuously optimize it by updating parameters;

[0033] The application unit is used to predict embryo implantation using a parameter-optimized multimodal fusion embryo multifocal plane image implantation prediction device.

[0034] Compared with the prior art, the technical effects of the present invention include at least the following:

[0035] When annotating the multifocal plane images of the training sample D5 embryos, only the successful implantation status needs to be annotated, without the need to segment and annotate each multifocal plane image. This reduces the training sample production time. By using pairwise information interaction between the core image and multiple other focal plane images, the problem of information interaction in more than three equal modalities in general multimodal tasks is solved, thereby obtaining richer and more accurate core image fusion features and increasing prediction accuracy. The Channel-Spatial Separated Multi-Head Attention (SMHA) mechanism reduces the problem of excessive computation in convolutional neural networks, improves the running efficiency of the prediction model, and effectively improves the accuracy of embryo implantation prediction. This can help embryologists select high-quality embryos more accurately and efficiently, and improve the success rate of embryo transfer. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a flowchart of the embryo multi-focal plane image implantation prediction method based on multimodal fusion provided in this embodiment of the invention;

[0038] Figure 2 This is a flowchart of the training process for the embryo multi-focal plane image implantation prediction model based on multimodal fusion provided in this embodiment of the invention.

[0039] Figure 3 This is a structural diagram of the embryo multi-focal plane image implantation prediction model based on multimodal fusion provided in this embodiment of the invention;

[0040] Figure 4 This is a structural diagram of the channel-spatial separation multi-head attention mechanism (SMHA) in the embryo multifocal plane image implantation prediction model based on multimodal fusion provided in this embodiment of the invention.

[0041] Figure 5 This is a structural diagram of the embryo multifocal plane image implantation prediction device based on multimodal fusion provided in an embodiment of the present invention. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of this invention.

[0043] To address the problems in existing technologies, such as excessive blurring and redundancy in the non-focused areas of multifocal images during the blastocyst stage, difficulty in fusing multifocal images during the blastocyst stage, insufficient information exchange, inability to fully extract information from multifocal images during the blastocyst stage, and excessive data computation leading to inaccurate embryo implantation prediction, this embodiment provides a method and apparatus for embryo implantation prediction based on multimodal fusion of multifocal planar images.

[0044] like Figure 1 As shown in the embodiment, an embryo multifocal plane image implantation prediction method based on multimodal fusion is provided, including the following steps:

[0045] S110, for the zona pellucida, inner cell mass and trophoblast at day 5 of the blastocyst stage, three focal plane images are labeled with implantation tags to obtain embryo multifocal image samples.

[0046] In this embodiment, D5 blastocyst images are captured using a multifocal plane microscope, including three focal plane images: zona pellucida (ZP), inner cell mass (ICM), and trophoblast (TE). Each image is a 224*224 three-channel RGB image, containing the embryo image under the multifocal plane microscope and a small amount of patient information. The main body of the embryo is generally located in the center of the image, but some embryos in a few images may be offset.

[0047] Embryologists labeled each individual sample to indicate whether successful implantation was possible, with labels 0 and 1 corresponding to implantation failure and implantation success, respectively. Based on the labeled data, the two categories were divided into identical training, validation, and test sets in an 8:1:1 ratio, with a 52:48 ratio of blastocyst formation to non-blastocyst formation. Data was randomly selected during the partitioning process to avoid similar data being grouped together.

[0048] In this embodiment, since camera adjustments are required when capturing multi-focal-plane images at different focal lengths, and the embryo may also sway slightly during this process, it is necessary to first perform a uniform alignment operation on the three multi-focal-plane images in each group of images to ensure that their features are essentially in the same position. For the blastocyst stage image data in the training set, all multi-focal-plane images of a single embryo need to be uniformly scaled and cropped, randomly horizontally flipped, randomly vertically flipped, randomly varied in brightness, and randomly rotated at an angle. Because it is necessary to ensure that the data of the same sample is at the same angle, the above data augmentation needs to be uniformly processed for the three multi-focal-plane images of a single sample.

[0049] S120: Construct a multimodal fusion embryo multifocal plane image implantation prediction model, including a core image generation module, a deep feature extraction module, a feature fusion module, and an implantation prediction module.

[0050] In the embodiment, the overall training process is as follows: Figure 2 As shown, multiple sets (typically batch=8) of multifocal plane image datasets (each set containing 3 images) are input into the multimodal fusion embryo multifocal plane image implantation prediction model. The structure of the multimodal fusion embryo multifocal plane image implantation prediction model is as follows. Figure 3 As shown. Since it is very difficult to exchange information pairwise between the three modalities in a multimodal task, the region weights of each multifocal plane image are first calculated by using three convolutional layers with large kernels. Then, the core image is generated by weighting these three images.

[0051] like Figure 3 As shown, the depth feature extraction module includes extraction networks for each of the multiple focal plane images and the core image. A pre-trained ResNet-18 is used to extract depth features from both the multiple focal plane images and the core image. Specifically, the first five layers of ResNet-18 are used to extract depth features, comprising five convolutional blocks. The first convolutional block is a single-layer convolutional layer. The next four convolutional blocks form residual network blocks with increasingly larger channel numbers by concatenating multiple Bottleneck blocks. Each convolutional block is activated by the ReLU function. Since pairwise information exchange between the three modalities is very difficult in multimodal tasks, the region weights for each focal plane image are first calculated using three convolutional layers with large kernels. Then, a weighted sum of these three images is performed to obtain the generated core image.

[0052] like Figure 3 As shown, the depth feature extraction module includes extraction networks for each of the multiple focal plane images and the extraction network for the core image. It uses a pre-trained ResNet-18 to extract the depth features of the multiple focal plane images and the depth features of the core image.

[0053] like Figure 4As shown, the feature fusion module includes a channel-spatial separated multi-head attention mechanism (SMHA) to calculate the SMHA values ​​of depth features from multiple focal plane images and the core image. It also performs dimensionality reduction on each SMHA and concatenates it with the depth features of the core image. The concatenated depth features are then further fused to the original number of channels, and the fused features are output. After the last two feature extraction layers of ResNet-18, the multi-focal plane image features and the core image features are flattened and MHA is calculated for information interaction. To reduce computational cost, the MHA is replaced with channel-spatial separated MHA (SMHA). The MHA operation is as follows:

[0054]

[0055] Among them W Q ∈R dk×dk W K ∈R dk×dk W V ∈R dk×dk These represent the projection matrices of the query, key, and value, respectively. In MHA, f x It represents query and f y This represents the key and value, with d and k representing the dimensions of the query and key vectors. The Softmax operation transforms the output into a probability distribution on the order of 0 to 1, using and to represent the weight relationship between each feature value.

[0056] The specific operation of spatial SMHA is to first process the query vector f q The key-value vector f is transformed into a 1×C two-dimensional matrix through an average pooling layer, and then reshaped using a reshape operation. kv Transform it into a two-dimensional matrix of (H×W)×C, and then calculate MHA:

[0057] Spatial-SMHA(f q ,f kv ) = MHA(AvgPool(f q ),f kv )

[0058] The specific operation of channel SMHA is to first set f q The convolutional layer with 1 output channel and 1×1 kernel size is transformed into a 1×(H×W) two-dimensional matrix, and then the f is reshaped. kv Transform it into a two-dimensional matrix of C×(H×W), and then calculate MHA:

[0059] Channel-SMHA(f q ,fkv )=MHA(Conv(f q ),f kv )

[0060] Combining the two operations above, the depth features f of the core image are... core Under supervision, the key features f' corresponding to the three focal planes i The extraction steps are as follows:

[0061] f` i =Channel-SMHA(f core Spatial-SMHA(f i ,f i ))

[0062] Where f i Used to refer to the depth features corresponding to the three focal plane images: zona pellucida, inner cell mass, and trophoblast.

[0063] After the key feature extraction is completed, f` i The data will be input into the next feature extraction unit. In order to reduce redundancy in the key features, the key features of each multi-focal plane image are dimensionality reduced (compressed from the current channel to 4 channels) and concatenated with the core image features. Finally, the concatenated features are dimensionality reduced to the original number of channels for further fusion.

[0064] like Figure 3 As shown, the embryo multi-focal plane image implantation prediction module is used to predict embryo implantation based on core image fusion features.

[0065] S130, all embryo multifocal plane image samples are fed into the embryo multifocal plane image implantation prediction model for training, and the model is continuously optimized by updating the model parameters.

[0066] During training, training sample images are weighted by ResNet-18 to generate core images. These core images are then paired with features extracted from multiple focal plane images of the embryo. After SMHA dimensionality reduction and concatenation with the deep features of the core images, the concatenated deep features are further fused to the original number of channels and output as fused features. These fused features are then input into the embryo implantation prediction model to predict the final result. The classification loss function used in model training is CrossEntropyLoss, with 250 iterations. Each iteration uses a batch size of 8 as input data, calculates the aforementioned loss, backpropagates, and updates the model parameters until training is complete. The model with the best validation performance is saved during each training iteration. Furthermore, by modifying hyperparameters, including the learning rate (lr) and the lr descent rate (decreasing by K% every N iterations), the loss value, accuracy, and recall on the validation set are optimized, resulting in better generalization performance of the model.

[0067] S140, using a parameter-optimized embryo multifocal plane image implantation prediction model to predict embryo implantation.

[0068] In this embodiment, an embryo implantation prediction model with optimized parameters is used to predict embryo implantation, including: inputting multifocal plane image sample data of the D5 embryo to be predicted into the embryo implantation prediction model for prediction, and outputting two types of results: embryo implantation and embryo failure to implant.

[0069] To address the issues of significant blurring and redundancy in non-focused regions of blastocyst multifocal images, making image fusion difficult and information exchange incomplete, this embodiment utilizes ResNet-18 to construct a core image generation module. Based on the first five layers of ResNet-18, this model has five convolutional modules. The first convolutional module is a single-layer convolutional layer, followed by modules 2-4 which form residual modules with increasingly larger channel numbers through concatenated Bottleneck modules. Each convolutional module is activated by the ReLU function. Since pairwise information exchange between the three modalities is extremely difficult in multimodal tasks, the system first calculates the regional weights of each MFP image using three convolutional layers with large kernels. Then, a weighted summation operation is performed on these three images to generate the core image. By generating the core image and using pairwise interaction, the system solves the problem of information exchange among the three equal modalities in multimodal tasks, thereby improving the clarity of the embryonic multifocal images and the degree of interaction between images, ultimately enhancing the model's prediction accuracy.

[0070] To address the issue of excessive data computation, in this embodiment, features are extracted from the four images—the embryonic multifocal plane image and the core image—through each feature extraction layer of a pre-trained ResNet-18. After the last two feature extraction layers of ResNet-18, the MFP image features and core image features are flattened and the MHA is calculated for information interaction. To reduce computational cost, the MHA is replaced with Channel-Spatial Separated MHA (SMHA). To reduce redundancy in key features, the key features of each MFP image are dimensionality-reduced (compressed from the current channel to 4 channels) and concatenated with the core image features. Finally, the concatenated features are dimensionality-reduced back to the original number of channels for further fusion. This reduces data computation and improves the computational efficiency of the prediction model.

[0071] Based on the same inventive concept, the embodiment also provides an embryo multifocal plane image implantation prediction device based on multimodal fusion, including a memory and a processor. The memory is used to store a computer program, and when the processor executes the computer program, it implements the steps of the embryo multifocal plane image implantation prediction method based on multimodal fusion provided in the above embodiment, including the following steps:

[0072] S110, for the zona pellucida, inner cell mass and trophoblast of the fifth day of blastocyst, three focal plane images are labeled with implantation tags to obtain embryo multifocal image samples;

[0073] S120, Construct a multimodal fusion embryo multifocal plane image implantation prediction model, including a core image generation module, a deep feature extraction module, a feature fusion unit, and an embryo multifocal plane image implantation prediction module;

[0074] S130, all embryo multifocal plane image samples are fed into the embryo multifocal plane image implantation prediction model for training, and the model is continuously optimized by updating the model parameters;

[0075] S140, using a parameter-optimized embryo multifocal plane image implantation prediction model to predict embryo implantation.

[0076] In this embodiment, the memory can be a volatile memory located in the near end, such as RAM, or a non-volatile memory, such as ROM, FLASH, floppy disk, hard disk, etc., or a remote storage cloud. The processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP), or field-programmable gate array (FPGA), that is, the steps of the embryo multi-focal plane image implantation prediction method based on multimodal fusion can be implemented through these processors.

[0077] Based on the same inventive concept, the embodiments also provide an embryo multi-focal plane image implantation prediction device 500 based on multimodal fusion, including a data acquisition unit 510, a model building unit 520, a training unit 530, and an application unit 540.

[0078] The data acquisition unit 510 is used to label implantation tags on three focal plane images of the zona pellucida, inner cell mass, and trophoblast at the fifth day of the blastocyst stage to obtain multifocal plane image samples of the embryo.

[0079] The model building unit 520 is used to construct a multimodal fusion embryo multifocal plane image prediction model, including a core image generation module, a depth feature extraction module, a feature fusion module, and an embryo multifocal plane image implantation prediction module. The core image generation module is used to generate a core image by performing adaptive weighting operations on batch-enhanced multifocal plane image samples. The depth feature extraction module is used to extract the corresponding depth feature vectors of multiple focal plane images and the core image in their respective extraction networks. The feature fusion module is used to fuse the depth feature vectors extracted from multiple focal plane images with the depth feature vectors extracted from the core image through an improved multi-head attention mechanism (SMHA) to enhance the information interaction between the various focal plane images. After multiple feature extraction layers and feature fusion layers, the final output core image fusion feature is used for the final prediction. The embryo multifocal plane image implantation prediction module is used to predict embryo implantation based on the core image fusion feature.

[0080] The training unit 530 is used to feed all embryo multifocal plane image samples into the multimodal fusion embryo multifocal plane image implantation prediction device for training, and continuously optimizes it by updating parameters.

[0081] Application unit 540 is used to predict embryo implantation using a multimodal fusion embryo multifocal plane image implantation prediction device with optimized parameters.

[0082] It should be noted that the embryo implantation prediction device based on multimodal fusion of embryo multifocal plane images provided in the above embodiments should be illustrated by the division of the above functional units when performing embryo implantation prediction. The functions can be assigned to different functional units as needed, that is, the internal structure of the terminal or server can be divided into different functional units to complete all or part of the functions described above. Furthermore, the embryo implantation prediction device based on multimodal fusion of embryo multifocal plane images provided in the above embodiments and the embryo implantation prediction method based on multimodal fusion of embryo multifocal plane images belong to the same concept. For details of its implementation process, please refer to the embryo implantation prediction method based on multimodal fusion of embryo multifocal plane images, which will not be repeated here.

[0083] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A device for embryo implantation prediction based on multimodal fusion multifocal plane images, comprising a memory and a processor, wherein the memory stores a computer-executable program for performing embryo implantation prediction based on multimodal fusion multifocal plane images, and the processor is communicatively connected to the memory and configured to execute the computer-executable program stored in the memory, characterized in that: When the processor executes the computer-executable program, it performs the following steps: For the zona pellucida, inner cell mass, and trophoblast at day 5 of the blastocyst stage, implantation tags were labeled to obtain multifocal image samples of the embryo. A multimodal fusion embryo multifocal plane image implantation prediction model is constructed, including a core image generation module, a depth feature extraction module, a feature fusion module, and an implantation prediction module. The core image generation module is used to generate a core image by adaptively weighting three focal plane images. The depth feature extraction module is used to extract the corresponding depth features of multiple focal plane images and the core image in their respective extraction networks. The feature fusion module is used to calculate key features by using the channel-space separation multi-head attention mechanism SMHA to calculate the depth features extracted from the three focal plane images and the depth features extracted from the core image. After dimensionality reduction of the three key features, they are concatenated with the corresponding depth features of the core image to obtain fused features, so as to enhance the information interaction between the various focal plane images. The fused features are used for the final prediction. The implantation prediction module is used to predict embryo implantation based on the input fused features. The key features are calculated using the channel-space separation multi-head attention mechanism SMHA, which exists only after the last two feature extraction layers of the residual network. It is assumed that the feature shape is C×H×W, where H represents the height of the vector feature, W represents the width, C represents the number of channels, and f... q f represents the input query vector. kv Refers to key vectors and value vectors. f q and f kv The input space SMHA is transformed into a two-dimensional matrix through average pooling and reshape operations, specifically: Spatial-SMHA(f q ,f kv )=MHA(AvgPool(f q ),f kv ) The AvgPool operation, as an average pooling operation, converts the input f... q Transform it into a 2D matrix of shape 1×C, and then use a reshape operation to transform the input f. kv Transform it into a two-dimensional matrix of shape (H×W)×C, and use these two matrices to calculate the MHA value; f q and f kv The input is transformed into a two-dimensional matrix through convolution, specifically: Channel-SMHA(f q ,f kv )=MHA(Conv(f q ),f kv ) The Conv operation is a convolution operation that transforms the input f... q Transform the convolutional layer with 1 output channel and 1×1 kernel size into a 1×(H×W) two-dimensional matrix, and then use the reshape operation to transform f kv Transform it into a C×(H×W) two-dimensional matrix, and use these two matrices to calculate the MHA value. Combining the above two operations, the depth features f of the core image are obtained. core Under supervision, the key features f' corresponding to the three focal planes i The extraction steps are as follows: f` i =Channel-SMHA(f core ,Spatial-SMHA(f i ,f i )) Where f i Used to refer to the depth features corresponding to the three focal plane images: zona pellucida, inner cell mass, and trophoblast. All embryo multifocal plane image samples were fed into the embryo multifocal plane image implantation prediction model for training, and the model was continuously optimized by updating the model parameters. Embryo implantation prediction was performed using a parameter-optimized multifocal plane image implantation prediction model.

2. The embryo multifocal plane image implantation prediction device based on multimodal fusion according to claim 1, characterized in that, The core image generation module uses convolution operations to calculate the region weights of the three focal plane images, and then combines the three focal plane images in a weighted manner to generate the core image.

3. The embryo multi-focal plane image implantation prediction device based on multimodal fusion according to claim 1, characterized in that, The depth feature extraction module includes extraction networks for each of the three focal plane images and an extraction network for the core image. The extraction networks use residual networks to extract the depth features of the three focal plane images and the core image.

4. The embryo multifocal plane image implantation prediction device based on multimodal fusion according to claim 1, characterized in that, The feature fusion module concatenates three key features with the corresponding depth features of the core image to obtain fused features, including: concatenating the three key features with the depth features of the core image after dimensionality reduction, and further fusion of the concatenated depth features with the original number of channels to obtain fused features.

5. The embryo multifocal plane image implantation prediction device based on multimodal fusion according to claim 1, characterized in that, The implantation prediction module uses a residual network to input fused features into a fully connected layer for calculation and outputs embryo implantation prediction results.

6. The embryo multifocal plane image implantation prediction device based on multimodal fusion according to any one of claims 1-5, characterized in that, The embryo multifocal plane image implantation prediction model also includes multiple sets of alternating deep feature extraction modules and feature fusion modules. The deep feature extraction module is used to extract corresponding high-order deep features from the three key features and fusion features input from the previous group in their respective extraction networks. The feature fusion module is used to calculate high-order key features by combining the three extracted high-order deep features with the high-order deep features extracted from the fusion feature through a channel-space separation multi-head attention mechanism. After dimensionality reduction of the three high-order key features, they are concatenated with the high-order deep features corresponding to the fusion feature to obtain the high-order fusion feature. The three high-order key features and the high-order fusion feature serve as the input to the next set of deep feature extraction modules.

7. The embryo multifocal plane image implantation prediction device based on multimodal fusion according to any one of claims 1-5, characterized in that, During training, the cross-entropy function is used to calculate the loss between the embryo implantation prediction result and the implantation label, and the model parameters are updated to continuously optimize the model.

8. A device for predicting implantation of embryos based on multimodal fusion multifocal plane images, characterized in that, It includes a data acquisition unit, a model building unit, a training unit, and an application unit. The data acquisition unit is used to label implantation tags on three focal plane images of the zona pellucida, inner cell mass, and trophoblast at the fifth day of the blastocyst stage to obtain embryo multifocal plane image samples. The model building unit is used to construct a multimodal fusion embryo multifocal plane image prediction model, including a core image generation module, a depth feature extraction module, a feature fusion module, and an embryo multifocal plane image implantation prediction module. The core image generation module is used to generate a core image by performing adaptive weighting on batch-enhanced multifocal plane image samples. The depth feature extraction module is used to extract the corresponding depth feature vectors of multiple focal plane images and the core image in their respective extraction networks. The feature fusion module is used to fuse the depth feature vectors extracted from multiple focal plane images with the depth feature vectors extracted from the core image through an improved multi-head attention mechanism (SMHA) to enhance the information interaction between the various focal plane images. After multiple feature extraction layers and feature fusion layers, the final output core image fusion feature is used for the final prediction. The embryo multifocal plane image implantation prediction module is used to predict embryo implantation based on the core image fusion feature. The key features are calculated using the channel-space separation multi-head attention mechanism SMHA, which exists only after the last two feature extraction layers of the residual network. It is assumed that the feature shape is C×H×W, where H represents the height of the vector feature, W represents the width, C represents the number of channels, and f... q f represents the input query vector. kv Refers to key vectors and value vectors. f q and f kv The input space SMHA is transformed into a two-dimensional matrix through average pooling and reshape operations, specifically: Spatial-SMHA(f q ,f kv )=MHA(AvgPool(f q ),f kv ) The AvgPool operation, as an average pooling operation, converts the input f... q Transform it into a 2D matrix of shape 1×C, and then use a reshape operation to transform the input f. kv Transform it into a two-dimensional matrix of shape (H×W)×C, and use these two matrices to calculate the MHA value; f q and f kv The input is transformed into a two-dimensional matrix through convolution, specifically: Channel-SMHA(f q ,f kv )=MHA(Conv(f q ),f kv ) The Conv operation is a convolution operation that transforms the input f... q Transform the convolutional layer with 1 output channel and 1×1 kernel size into a 1×(H×W) two-dimensional matrix, and then use the reshape operation to transform f kv Transform it into a C×(H×W) two-dimensional matrix, and use these two matrices to calculate the MHA value. Combining the above two operations, the depth features f of the core image are obtained. core Under supervision, the key features f' corresponding to the three focal planes i The extraction steps are as follows: f` i =Channel-SMHA(f core ,Spatial-SMHA(f i ,f i )) Where f i Used to refer to the depth features corresponding to the three focal plane images: zona pellucida, inner cell mass, and trophoblast. The training unit is used to send all embryo multifocal plane image samples into the multimodal fusion embryo multifocal plane image implantation prediction device for training, and continuously optimize it by updating parameters; The application unit is used to predict embryo implantation using a parameter-optimized multimodal fusion embryo multifocal plane image implantation prediction device.

Citation Information

Patent Citations

  • Embryo development potential prediction method and system, equipment and storage medium

    CN113469958A

  • Embryo pregnancy prediction method and system based on space-time attention and cross-modal fusion

    CN114972167A