A method and system for predicting embryonic growth stages based on self-supervised learning
Through self-supervised learning and multi-core loss function optimization prediction model, the problem of insufficient generalization ability of embryo growth prediction model when facing new blastocyst images is solved, and higher prediction accuracy and robustness are achieved, reducing calculation complexity and labeling costs.
Patent Information
- Application Number
- CN202510432149.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Existing embryo growth prediction models lack generalization ability when facing brand new blastocyst images, and cannot accurately predict the growth stage of embryos, especially when dealing with outliers and noise.
Using a self-supervised learning method, a prediction model including an encoder and a decoder is constructed by segmenting the embryo image into several blocks and randomly masking it, and the model parameters are optimized using multi-core loss functions, combining VideoMamba encoder and multiple kernel functions to improve the robustness and generalization capabilities of the model.
It improves the prediction accuracy and robustness of the model for blastocyst images, can better handle outliers and noise in complex images, reduces dependence on labeled data, and enhances the generalization ability of the model.
Smart Images

Figure CN119942249B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method and system for predicting embryo growth stages based on self-supervised learning. Background Art
[0002] In the field of Assisted Reproductive Technology (ART), the assessment of blastocyst development quality is a crucial step in improving the success rate of in vitro fertilization (IVF). A blastocyst is an embryo stage that develops to the 5th to 6th day during in vitro fertilization (IVF), and its quality is of vital significance for embryo transfer decisions and the success rate of pregnancy. During the IVF process, medical experts use morphological assessment to determine whether a blastocyst has higher developmental potential. To improve the efficiency and accuracy of the assessment, using machine learning technology to automatically predict embryo growth has become an important means to increase the success rate of IVF.
[0003] In recent years, Self-Supervised Learning (SSL) has rapidly emerged in the field of machine learning. Especially in scenarios where data is scarce or the annotation cost is high, self-supervised learning can significantly reduce the dependence on labeled data by using the internal relationships within the data for training. In embryo growth prediction, self-supervised learning methods can learn more feature representations from unlabeled data under limited labeled data, thereby enhancing the prediction ability of the model.
[0004] Although self-supervised learning provides new possibilities for blastocyst growth prediction, the existing technologies still face some challenges. The development process of blastocysts is complex and variable, involving multiple aspects of characteristics such as morphological changes of embryos, cell division rates, and trophoblast development. These characteristics vary greatly among different individuals. As a result, traditional prediction models may perform well on training data, but have insufficient generalization ability when encountering new or unseen blastocyst image data and cannot accurately predict the embryo growth stage. Summary of the Invention
[0005] The present invention proposes a method and system for predicting embryo growth stages based on self-supervised learning, which solves the problem of insufficient generalization ability of existing prediction models when facing brand-new blastocyst images.
[0006] To solve the above technical problems, the present invention provides a method for predicting embryo growth stages based on self-supervised learning, including the following steps:
[0007] Step S1: Collect embryo images at different growth stages as a training set, and divide each embryo image in the training set into several blocks of the same size;
[0008] Step S2: Randomly select half of the blocks in each embryo image for masking to obtain a masked training set;
[0009] Step S3: Construct a prediction model including an encoder and a decoder. The encoder extracts a feature sequence from the masked training set, and the decoder reconstructs the embryo image according to the feature sequence;
[0010] Step S4: Use multiple weighted combined kernel functions as the loss function of the prediction model. Adjust the weights of each kernel function through cross-validation and gradient descent, and use backpropagation to optimize the parameters of the prediction model to obtain an optimal prediction model;
[0011] Step S5: Input the embryo image to be predicted into the optimal prediction model to obtain the prediction result of the growth stage of the embryo.
[0012] Preferably, the encoder in Step S3 extracting the feature sequence from the masked training set includes the following steps:
[0013] Step S31: Divide the masked training set into uniform spatio-temporal blocks according to the time dimension;
[0014] Step S32: Convert each spatio-temporal block into a high-dimensional vector through an embedding layer, and perform positional encoding on each high-dimensional vector according to time and space information;
[0015] Step S33: The encoder includes multiple Transformer blocks. Each Transformer block contains a self-attention layer and a multi-layer perceptron module. The self-attention layer captures the spatial and temporal correlations between high-dimensional vectors through multiple parallel attention heads, and the multi-layer perceptron module learns complex feature representations in high-dimensional vectors through two fully connected layers and a non-linear activation function;
[0016] Step S34: Use residual connections to add the output of each Transformer block to the input, and use the output of each Transformer block as the input of the next Transformer block to obtain the feature sequence of the masked training set.
[0017] Preferably, the expression of the loss function in Step S4 is:
[0018] ;
[0019] In the formula, is the loss function; N is the total number of embryo images in the masked training set; is the weight of the i th kernel function corresponding to the j th embryo image; is thej a kernel function; is the i th embryo image; is the i predicted value of the th embryo image; i is the j true value of the
[0020] Preferably, step S4 includes the following steps:
[0021] Step S41: Define the hyperparameter space of the prediction model and initialize the hyperparameter combination of the prediction model;
[0022] Step S42: Divide the masked training set into several subsets. For each subset: use all subsets except the current subset to train the prediction model and evaluate the performance of the prediction model on the current subset;
[0023] Step S43: Calculate the gradient of the multi-kernel loss function with respect to the weights of each kernel function, and use gradient descent to update the weights of each kernel function in the multi-kernel function;
[0024] Step S44: Calculate the gradient of the multi-kernel loss function with respect to the parameters of the prediction model, and use gradient descent to update the hyperparameter combination of the prediction model;
[0025] Step S44: Repeat steps S41 to S43 until the set first number of iterations is reached;
[0026] Step S45: Record the performance metrics of each set of hyperparameter combinations in cross-validation, adjust the search space of the hyperparameters according to the average performance of cross-validation, and repeat steps S42 to S44 until the set second number of iterations is reached;
[0027] Step S46: Calculate the average performance of all cross-validations, select the hyperparameter combination corresponding to the cross-validation with the best average performance as the optimal parameters, and obtain the optimal prediction model.
[0028] Preferably, after collecting embryo images at different growth stages as the training set in step S1, preprocess the training set, including the following steps: convert the collected embryo images into grayscale images, stack the grayscale images, and convert all embryo images into a high-dimensional vector.
[0029] The present invention also provides an embryo growth stage prediction system based on self-supervised learning, which is implemented based on the above-mentioned embryo growth stage prediction method based on self-supervised learning, and includes an image preprocessing module, an embryo feature extraction module, a masked self-supervised module, and a multi-kernel loss module;
[0030] The image preprocessing module: collects embryo images and preprocesses the collected embryo images;
[0031] The embryo feature extraction module: extracts features related to embryo growth from the preprocessed embryo images;
[0032] The masked self-supervised module: trains the prediction model through self-supervised learning by using the internal relationships in the embryo images;
[0033] The multi-kernel loss module: optimizes the prediction model through a multi-kernel loss function to enhance the ability of the prediction model to capture features related to embryo development.
[0034] Preferably, preprocessing the embryo images includes the following steps: adjusting the sizes of all embryo images to be consistent and converting the colored embryo images into grayscale images.
[0035] Preferably, the embryo feature extraction module uses a deep learning network to identify and extract the morphological features of the embryos in the embryo images.
[0036] Preferably, the self-supervised learning includes the following steps: dividing the feature maps extracted by the embryo feature extraction module into several image patches, masking a high proportion of the segments, and only inputting the unmasked parts into the encoder to extract latent features, and the decoder uses these features to reconstruct the masked parts.
[0037] Preferably, the multi-kernel loss module selects multiple kernel functions as the loss function of the prediction model, optimizes the weights of each kernel function, and optimizes the parameters of the prediction model through backpropagation.
[0038] The advantages of the present invention at least include:
[0039] 1. By dividing the embryo image into several image patches and randomly selecting half of the image patches for masking, the prediction model is forced to reconstruct the entire image from partial visible information, enhancing the generalization ability of the model to new data;
[0040] 2. The multi-kernel loss function enables the prediction model to better handle outliers and noise. Different kernel functions are suitable for learning different features in the embryo images and can provide a more robust feature representation for the reconstruction and prediction of the model, improving the robustness of the model when facing complex images. Description of the Drawings
[0041] Figure 1 is the embryo image collected in the embodiment of the present invention;
[0042] Figure 2 is the schematic flow chart of the method in the embodiment of the present invention;
[0043] Figure 3Schematic flowchart of masking the embryo image in the embodiment of the present invention;
[0044] Figure 4 Schematic diagram of the structure of the bidirectional Mamba block in the prediction model of the embodiment of the present invention;
[0045] Figure 5 Schematic diagram of the system structure of the embodiment of the present invention. Detailed implementation manners
[0046] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] In the existing automated embryo growth prediction model, the reconstruction loss function based on the Masked Autoencoder (MAE) model is often used to train the model. However, as Figure 1 shown, there are usually a large number of outliers and noises in the embryo image data, and the mean square error (MSE) calculates the average of the squares of the differences between the predicted value and the true value. Therefore, larger errors will be amplified, thus having a greater impact on the prediction ability of the model.
[0048] To solve this problem, the Multiple Kernel Learning (MKL) loss function is usually introduced in the prior art, and the nonlinear characteristics of the kernel function are used to capture the complex features in the data. Through the kernel function mapping, the data can be transformed into a high-dimensional space, so that the model can better distinguish normal and abnormal embryo development patterns. In addition, different kernel functions, such as Gaussian kernel and polynomial kernel, are suitable for learning different features and can provide a more robust feature representation for the reconstruction and prediction of the model.
[0049] Although self-supervised learning and multi-kernel loss functions have solved some problems in blastocyst growth prediction to a certain extent, they still face some technical challenges. For example, how to effectively combine multiple kernel functions without significantly increasing the computational complexity, how to balance the importance of different features, and how to select the optimal combination of kernel functions to improve the performance of the model. In addition, due to the diversity and complexity of embryo development, the model may still show a certain lack of generalization ability when facing new or unseen blastocyst data.
[0050] Therefore, one of the core goals of the embodiments of the present invention is to design an efficient prediction model that can adapt to different blastocyst data characteristics under the condition of limited computing resources and labeled data. This requires that the model should not only be able to handle outliers and noise, but also have good generalization ability to adapt to the diversity and complexity of embryonic development. Through these improvements, the prediction accuracy of the model can be improved, and better results can be achieved in the field of automated embryo growth prediction.
[0051] like Figure 2 As shown, an embodiment of the present invention provides an embryo growth stage prediction method based on self-supervised learning, comprising the following steps:
[0052] Step S1: Collect embryo images at different growth stages as a training set, and divide each embryo image in the training set into several blocks of the same size.
[0053] Specifically, the data collection work of the embodiment of the present invention covers a number of top reproductive medical centers, which provide a wealth of embryo image resources for this embodiment with their professional experience and advanced equipment in the field of assisted reproductive technology. Each embryo is photographed in the first five days of development, from the pronuclear stage to the blastocyst stage of embryonic development, a total of 500 pictures, and a total of nearly 100,000 images of embryonic development have been collected from multiple reproductive centers. For embryos with less than or more than 500 photos, the photos are copied, added or deleted on average from the existing photos until the number reaches 500.
[0054] These images were taken in a time-lapse incubator, which can continuously record the development of the embryo without affecting its development. The collected image dataset comprehensively covers all stages of embryo development from the pronuclear stage to the blastocyst stage, and includes images of empty dishes for comparison. The embryo images are sorted in sequence according to the embryonic development time, and standardized preprocessing operations are performed to ensure that the size of the embryo images is consistent, which is convenient for subsequent input into the model. This process ensures the quality and consistency of the data, providing a reliable basis for model training and prediction.
[0055] In the research of embryo growth prediction, the color information of images is usually not the core of the analysis. Grayscale images are sufficient to provide the key visual information required for autofocus. To simplify the processing flow and make full use of the advantages of grayscale images, the cv2.cvtColor() function in the OpenCV library is used to convert the color embryo images into grayscale format. This conversion process not only removes the interference of colors but also retains the texture and contour information of the images, which is crucial for the embryo growth and development prediction task. To meet the input requirements of the neural network model and optimize the efficiency of data processing, these grayscale images are specially processed. Specifically, to meet the input requirements of the neural network model and optimize the data processing efficiency, all embryo images are stacked together. By adding a new dimension, a set of data containing N embryo images is stacked along this dimension, and the shape of each tensor is ([[]] N , H , W ), where N represents the length of the image sequence, H and W are the height and width of the image respectively. This processing method enables the visual information of each group of images to be transmitted to the neural network along the new dimension. At the same time, the data structure is more compact, which is conducive to the network learning the image sequence and improving the processing efficiency.
[0056] Step S2: Randomly select half of the blocks in each embryo image for masking to obtain a masked training set.
[0057] Specifically, in the field of self-supervised learning research, the masking technique plays a crucial role. It introduces random uncertainty into the data, stimulating the model to mine more stable and generalized feature representations. As Figure 3 shown, in the embodiment of the present invention, 50% of the pixels in the embryo image are masked. This process first divides the embryo image into 36 small blocks. This division helps the model to more carefully understand and reconstruct the local features of the image. Especially in blastocyst images, the features of key regions such as the inner cell mass and the outer trophoblast are particularly important. Then, randomly select half of these small blocks for covering. This high proportion of masking operation simulates the partial occlusion situation that may be encountered in actual image acquisition, forcing the model to infer the features of the covered part from the limited visible part, so as to reconstruct the complete image.
[0058] After covering half of the image patches, the model needs to utilize the features in the other half of the un-covered image patches to predict the content of the covered part. This process requires the model to learn how to extract key information from the visible part and use this information to predict the features of the covered part. In the study of embryo growth prediction, this strategy not only improves the model's adaptability to occlusion and information loss, but also enhances the ability to extract various complex features of blastocysts. In this way, the model can better understand and process the complexity and noise in blastocyst images, especially the outliers and noise that may appear in key regions such as the inner cell mass and trophectoderm, improving the robustness of the model when processing blastocyst growth sequences, and providing a new perspective and strategy for the application of self-supervised learning in complex visual tasks.
[0059] Step S3: Construct a prediction model including an encoder and a decoder, where the encoder extracts a feature sequence from the masked training set, and the decoder reconstructs the embryo image according to the feature sequence.
[0060] Step S4: Use multiple weighted combined kernel functions as the loss function of the prediction model, adjust the weights of each kernel function through cross-validation and gradient descent, and use backpropagation to optimize the parameters of the prediction model to obtain the optimal prediction model.
[0061] Specifically, the core architecture of the prediction model in the embodiments of the present invention is a MAE (Masked Autoencoder) model. Since self-supervised learning does not require manual annotation, MAE greatly reduces the annotation cost. At the same time, the learning difficulty brought by masking improves the generalization ability of the model and reduces the risk of overfitting. In addition, MAE only processes part of the input segments, significantly reducing the computational cost, and is particularly suitable for data with significant spatio-temporal redundancy such as embryo development image sequences. And in the embodiments of the present invention, a multi-kernel loss function is introduced on the basis of the MAE model to enhance the model's ability to capture complex features of blastocyst images, reduce the sensitivity to outliers, and enhance the model's learning ability for different feature subspaces. The entire model is mainly composed of four modules: an input preprocessing module, a feature extraction module, a masked self-supervised module, and a multi-kernel loss module.
[0062] To further improve the performance of the model for the embryo growth prediction task, the embodiment of the present invention uses the state space model VideoMamba for efficient video understanding as the encoder. The VideoMamba encoder, as the core feature extraction component, is specifically designed for the spatio-temporal characteristics of video data. Its construction process first divides the video data into uniform spatio-temporal blocks along the time dimension. These spatio-temporal blocks represent the basic units of consecutive frames in the video and contain various dynamic changes in the scene at different time points. Each spatio-temporal block is transformed into a high-dimensional vector through the embedding layer. This transformation maps the blastocyst growth sequence into a format that the model can process, and at the same time, by adding temporal and spatial position encodings, the features in the video can be understood.
[0063] Structure of the VideoMamba encoder Figure 4 As shown, the SSM framework therein is the integration of the spring, spring MVC, and mybatis frameworks. The encoder adopts the self-attention mechanism based on Transformer, which can effectively handle the spatio-temporal dependence relationships in the blastocyst growth sequence. Through multiple parallel attention heads, the model can capture the spatial and temporal correlations between video frames in different subspaces. Each attention head independently analyzes the local and global features in the video frames. Such a setting enables VideoMamba to process multiple features simultaneously in multiple dimensions, thereby enhancing the model's ability to understand video content. After the self-attention layer, the encoder uses residual connections to add the output to the input, which helps the flow of information and alleviates the problem of gradient vanishing. Subsequently, the training process is further stabilized through the normalization layer to ensure that the output of each layer remains within a reasonable range. This design makes the encoder more stable during the training process and can quickly adapt to different video input patterns. To enhance the model's expressive power and non-linear fitting ability, the VideoMamba encoder adds a multi-layer perceptron (MLP) module after each self-attention layer. The MLP consists of two fully connected layers with a non-linear activation function ReLU in the middle, enabling the model to learn more complex feature representations.
[0064] The VideoMamba encoder is constructed by stacking multiple Transformer blocks. Each block contains a self-attention layer and an MLP layer, and the output of each block serves as the input of the next block. This layer-by-layer stacking structure enables the model to abstract and refine the important information in the video data at different levels, gradually forming a high-dimensional feature representation of the video sequence. Finally, the output of the VideoMamba encoder is a high-dimensional representation containing rich spatio-temporal features, which can comprehensively reflect the content of the blastocyst growth sequence.
[0065] The decoder trains the prediction ability of the model by recovering the masked regions from the feature representations extracted by the encoder. The decoder is designed to be lightweight and efficient, aiming to precisely reconstruct the morphological features of the blastocyst, enabling the model to fully utilize the powerful feature extraction ability of the VideoMamba encoder. Meanwhile, the performance of embryo growth prediction is optimized through a task-specific decoder and a multi-kernel loss function, thereby achieving higher accuracy and robustness when processing blastocyst growth sequences. The encoder uses a masking strategy to extract features related to the quality of blastocyst development from the input images, including the morphology of the inner cell mass, trophectoderm, etc., forcing the model to predict the masked part from the visible part. Then, the decoder reconstructs the feature representations extracted by the encoder into the form of the original data, i.e., the reconstructed image. In self-supervised learning, the decoder needs to recover the masked regions from the output of the encoder.
[0066] The use of a multi-kernel loss function effectively enhances the ability of the prediction model to capture complex features related to blastocyst development. Traditional reconstruction loss functions usually adopt the pixel-level mean squared error (MSE). However, due to the complexity and high noise of embryo image data, MSE will amplify the errors and it is difficult to fully capture key local features. To solve this problem, embodiments of the present invention select multiple kernel functions, such as combined Gaussian kernel, polynomial kernel, and linear kernel, to replace the loss function. These kernel functions can map different features to a high-dimensional space, capture various features such as edges, textures, and color distributions in the blastocyst image, and optimize the importance of each kernel function for the task through weight adjustment. The use of multi-kernel functions can not only balance the influence of outliers in complex data but also improve the reconstruction quality and the robustness of the model.
[0067] Specifically, embodiments of the present invention select four kernel functions for various features such as the basic morphological and distribution features of the image, the consistency of the size and morphology of blastomeres, and the stratification and migration patterns of cells. The weight coefficients are adjusted through automatic learning to determine the relative importance of each kernel function in the current task, thereby optimizing the quality of image reconstruction. The expression of the constructed multi-kernel loss function is:
[0068] ;
[0069] In the formula, is the loss function; N is the total number of embryo images in the masked training set; is the i th weight of the j rd kernel function corresponding to the th embryo image; j is the th kernel function, which is used to map the original features to a high-dimensional space to capture complex features in the blastocyst image; i is the i th embryo image; The predicted value of the i th embryo image; The i true value of the j th embryo image under the
[0070] By constructing the masked learning features and loss function, better prediction results of embryo growth can be achieved. In the specific implementation process of the method of the embodiment of the present invention, according to the characteristics of the algorithm for adaptively adjusting the learning rate, the Adam optimizer is selected to optimize the model. The initial learning rate is set to 0.005, the batch size is 16, and the number of training epochs is 60. And a learning rate decay strategy is adopted to ensure that the adjustment of model parameters is refined step by step during the training process. Specifically, after a certain number of epochs, the learning rate will be reduced according to a predetermined decay rate to promote the model to more closely approximate the optimal solution in the later stage of training.
[0071] Cross-validation is used to evaluate the performance of the model optimized by gradient descent. The dataset is divided into several subsets, and each time one subset is selected as the validation set, and the rest are used as the training set. On the training set, the gradient descent algorithm continues to adjust the weights to minimize the loss function. Then, the performance of the model is evaluated on the validation set, and the weight coefficients are further adjusted according to these performance results. This process is repeated multiple times, and each time a different subset is selected as the validation set to ensure the consistency of the model's performance on different data subsets. Finally, according to the average performance of the model during multiple validation processes, the optimal weight coefficients are selected.
[0072] By combining cross-validation and gradient descent, the weight coefficients of the multi-core loss function are optimized and the generalization ability is improved. Specifically, first, the weight coefficients are randomly initialized, and then the weights are iteratively updated by the gradient descent algorithm. Subsequently, the K-fold cross-validation method is used to evaluate the generalization ability of the model under different hyperparameter settings, including the learning rate and batch size, etc. By comparing the performances under different settings, the optimal combination of hyperparameters is selected for the formal training of the model. After the model training is completed, cross-validation is used again to evaluate the final performance of the model to ensure that the model has good prediction ability on unseen data. This process may require multiple iterations until satisfactory performance is achieved, so as to provide more accurate data for subsequent analysis and research. The Pytorch framework is used to implement the model of the present invention, and all models are trained and tested on an NVIDIA GeForce GTX 1660 GPU.
[0073] Through the mutual cooperation among the four modules, the model's comprehensive feature extraction ability for the blastocyst development process is jointly improved. By capturing the complementarity of the core functions of each module, the overfitting risk caused by the model's dependence on a single module is reduced, and the prediction accuracy and generalization ability of the system are further improved. In the model design, the feature extraction module enhances the spatial information perception through a bidirectional state space model and positional embedding. The masked self-supervised module uses a random masking mechanism to learn the global semantic relationships of the data. The input preprocessing module provides high-quality standardized data for the subsequent process. Finally, the multi-kernel loss function module dynamically adjusts the weight ratio of the features by integrating multiple kernel functions, thereby optimizing the reconstruction and prediction performance of the model.
[0074] In addition, the annotation of embryo image data is a cumbersome task that requires experienced professionals to comprehensively evaluate and determine the quality of various features of a set of blastocyst images, which is very time-consuming and laborious. Therefore, the embodiments of the present invention reasonably combine multi-kernel learning with self-supervised learning, effectively reducing the need for labeled embryo images. By constructing a self-supervised task network model through the feature information of the embryo images themselves, this not only reduces the need for professional participation but also reduces the annotation cost of blastocyst images. At the same time, through the close linkage of these four modules, the model can also effectively reduce the sensitivity to abnormal data and significantly improve the processing ability of complex biomedical data and the generalization performance for unseen data.
[0075] Step S5: Input the embryo image to be predicted into the optimal prediction model to obtain the prediction result of the growth stage of the embryo.
[0076] To verify the effectiveness of the method proposed in the embodiments of the present invention, a large number of parameter tuning experiments and comparative experiments with different methods were carried out. Since a set of embryo images has high redundancy and temporal correlation, a design with an extremely high tubular mask ratio is adopted. This higher mask ratio helps prevent information leakage during the mask modeling process, making the reconstruction of the gradient features of embryo images a meaningful self-supervised training task, and also increasing the difficulty of reconstruction to a certain extent. The embodiments of the present invention explored the accuracy when the mask ratio increased from 50% to 90%. The experimental results show that when the mask ratio is 50%, the growth prediction accuracy of embryo images is 82.7%; when the mask ratio is 70%, the auto-focus recognition accuracy of embryo images is 84.9%; when the mask ratio is 80%, the recognition accuracy is 85.2%; when the mask ratio is 90%, the recognition accuracy can reach the highest 87.8%, which is 5.1% higher than the result under the 50% mask ratio. This shows that at a higher mask ratio, the multi-task cooperation model of the present invention can obtain more excellent performance, further indicating that the design of the present invention can force the adopted network model to capture more useful spatio-temporal information in embryo images.
[0077] In addition, to further verify the effectiveness of the present method, embodiments of the present invention compared the effects of three different loss functions on the experimental results.
[0078] Table 1 Prediction accuracy obtained on embryo images using different loss functions
[0079]
[0080] As can be seen from Table 1, when the mean squared error is selected as the loss function, the obtained accuracy is 79.4%; when the cross-entropy is selected as the loss function, the obtained accuracy is 82.9%; when the multi-kernel function is selected as the loss function, the obtained accuracy is 87.8%. This indicates that since the growth of embryo images is jointly affected by multiple features, a single function as the loss function cannot fully calculate the loss, and the highest accuracy can be obtained when the four kernel functions in the multi-kernel function act together, and the multi-kernels cooperate with each other to achieve the best effect.
[0081] In summary, the method of the embodiments of the present invention makes full use of the relevant information contained in embryo images, adopts the VideoMamba model, combines with the self-supervised mask of embryo images for modeling, and studies the embryo growth prediction task. Experiments show that the method proposed by the present invention has more excellent performance.
[0082] As Figure 5 shown, embodiments of the present invention also provide an embryo growth stage prediction system based on self-supervised learning, which is implemented based on the above-mentioned embryo growth stage prediction method based on self-supervised learning, and includes an image preprocessing module, an embryo feature extraction module, a mask self-supervised module, and a multi-kernel loss module.
[0083] The image preprocessing module is used to collect embryo images, preprocess the collected embryo images, adjust the sizes of all embryo images to be consistent, and convert the color embryo images into grayscale images.
[0084] The embryo feature extraction module uses a deep learning network to extract features related to embryo growth from the preprocessed embryo images.
[0085] The mask self-supervised module is used to divide the feature maps extracted by the embryo feature extraction module into several non-overlapping image patches, mask a high proportion of the segments, and only input the unmasked part into the encoder to extract latent features. The decoder uses these features to reconstruct the masked part, and trains the prediction model using the internal relationships of the embryo images.
[0086] This module infers the semantics of the masked part through the context of visible fragments, thereby effectively learning the association between global and local features. Since self-supervised learning does not require manual annotation, MAE significantly reduces the annotation cost. At the same time, the learning difficulty brought about by masking improves the generalization ability of the model and reduces the risk of overfitting. In addition, MAE only processes partial input fragments, significantly reducing the computational cost, and is particularly suitable for data with significant spatio-temporal redundancy such as embryo development image sequences.
[0087] The multi-kernel loss module selects multiple kernel functions as the loss function of the prediction model, optimizes the weights of each kernel function, and optimizes the parameters of the prediction model through backpropagation to enhance the ability of the prediction model to capture embryo development-related features.
[0088] Traditional reconstruction loss functions usually adopt pixel-level mean squared error (MSE). However, due to the complexity and high noise of embryo image data, the single loss function MSE, which calculates the average of the squares of the differences between predicted values and true values, will amplify errors and is difficult to fully capture local features such as the cell density of the inner cell mass and the cell arrangement of the outer trophoblast, as well as their changes at different developmental stages, which are crucial for evaluating the quality and developmental potential of blastocysts. To solve this problem, the embodiments of the present invention adopt multiple kernel functions, such as combined Gaussian kernel, polynomial kernel, and linear kernel, to replace the loss function. These kernel functions can map different features to a high-dimensional space, capture various features such as edges, textures, and color distributions in blastocyst images, and at the same time optimize the importance of each kernel function for the task through weight adjustment. The use of multi-kernel functions can not only balance the influence of outliers in complex data but also improve the reconstruction quality and the robustness of the model. By dynamically adjusting the weight coefficients through cross-validation and combining optimization algorithms such as gradient descent, the model performance is further optimized, and finally, the model realizes efficient prediction of the quality of blastocyst development under the condition of limited labeled data.
[0089] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present invention. As long as the combinations of these technical features do not conflict, they should be considered as the scope described in this specification.
[0090] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A method for predicting embryo growth stages based on self-supervised learning, characterized in that, It includes the following steps: Step S1: Collect embryo images at different growth stages as the training set, and segment each embryo image in the training set into several blocks of the same size; Step S2: Randomly select half of the blocks in each embryo image for masking to obtain the masked training set; Step S3: Construct a prediction model including an encoder and a decoder. The encoder extracts a feature sequence from the masked training set, and the decoder reconstructs the embryo image according to the feature sequence; Step S4: Select four kernel functions for the basic morphology and distribution features of the image, the consistency of the size and morphology of blastomeres, and the stratification and migration patterns of cells. Use multiple weighted combined kernel functions as the loss function of the prediction model. The expression of the constructed multi-kernel loss function is: Where L is the loss function; N is the total number of embryo images in the masked training set; ω ij is the weight of the j-th kernel function corresponding to the i-th embryo image; K j is the j-th kernel function, which is used to map the original features to a high-dimensional space to capture the complex features in the blastocyst image; x i is the i-th embryo image; f(x i ) is the predicted value of the i-th embryo image; y ij is the true value of the i-th embryo image under the j-th feature; Adjust the weights of each kernel function through cross-validation and gradient descent, and optimize the parameters of the prediction model using backpropagation. The steps to obtain the optimal prediction model are as follows: Step S41: Define the hyperparameter space of the prediction model and initialize the hyperparameter combination of the prediction model; Step S42: Divide the masked training set into several subsets. For each subset: Train the prediction model using all subsets except the current subset, and evaluate the performance of the prediction model on the current subset; Step S43: Calculate the gradient of the multi-kernel loss function with respect to the weights of each kernel function, and use gradient descent to update the weights of each kernel function in the multi-kernel function; Step S44: Calculate the gradient of the multi-kernel loss function with respect to the parameters of the prediction model, and use gradient descent to update the hyperparameter combination of the prediction model; Step S44: Repeat Step S41 to Step S43 until the set first iteration number is reached; Step S45: Record the performance metrics of each hyperparameter combination in cross-validation, adjust the search space of the hyperparameters according to the average performance of cross-validation, and repeat Step S42 to Step S44 until the set second iteration number is reached; Step S46: Calculate the average performance of all cross-validations, and select the hyperparameter combination corresponding to the cross-validation with the best average performance as the optimal parameters to obtain the optimal prediction model; Step S5: Input the embryo image to be predicted into the optimal prediction model to obtain the prediction result of the growth stage of the embryo.
2. The method for predicting the embryonic growth stage based on self-supervised learning according to claim 1, wherein: The steps for the encoder in Step S3 to extract the feature sequence from the masked training set include the following steps: Step S31: Divide the masked training set into uniform spatio-temporal blocks according to the time dimension; Step S32: Convert each spatio-temporal block into a high-dimensional vector through an embedding layer, and perform positional encoding on each high-dimensional vector according to time and space information; Step S33: The encoder includes multiple Transformer blocks. Each Transformer block contains a self-attention layer and a multi-layer perceptron module. The self-attention layer captures the spatial and temporal correlations between high-dimensional vectors through multiple parallel attention heads. The multi-layer perceptron module learns complex feature representations in high-dimensional vectors through two fully connected layers and a non-linear activation function; Step S34: Add the output of each Transformer block to the input using residual connection, and use the output of each Transformer block as the input of the next Transformer block to obtain the feature sequence of the masked training set.
3. A method for predicting the embryonic growth stage based on self-supervised learning according to claim 1, characterized in that: After collecting embryo images at different growth stages as the training set in Step S1, preprocess the training set, including the following steps: convert the collected embryo images into grayscale images, stack the grayscale images, and convert all embryo images into a high-dimensional vector.
4. An embryo growth stage prediction system based on self-supervised learning, which is implemented based on an embryo growth stage prediction method based on self-supervised learning according to any one of claims 1 to 3, and is characterized in that: It includes an image preprocessing module, an embryo feature extraction module, a masked self-supervised module, and a multi-kernel loss module; The image preprocessing module: collect embryo images and preprocess the collected embryo images; The embryo feature extraction module: extract features related to embryo growth from the preprocessed embryo images; The masked self-supervised module: train the prediction model by self-supervised learning using the internal relationships within the embryo images; The multi-kernel loss module: optimize the prediction model through a multi-kernel loss function to enhance the ability of the prediction model to capture features related to embryo development.
5. The embryo growth stage prediction system based on self-supervised learning according to claim 4, characterized in that: Preprocessing the embryo images includes the following steps: adjust the sizes of all embryo images to be consistent, and convert the colored embryo images into grayscale images.
6. The embryo growth stage prediction system based on self-supervised learning according to claim 4, wherein: The embryo feature extraction module uses a deep learning network to identify and extract the morphological features of embryos in the embryo images.
7. The embryo growth stage prediction system based on self-supervised learning according to claim 4, wherein: The self-supervised learning includes the following steps: divide the feature map extracted by the embryo feature extraction module into several image patches, mask a high proportion of the segments, and only input the unmasked part into the encoder to extract latent features, and the decoder uses these features to reconstruct the masked part.
8. The embryo growth stage prediction system based on self-supervised learning according to claim 4, wherein: The multi-kernel loss module selects multiple kernel functions as the loss function of the prediction model, optimizes the weights of each kernel function, and optimizes the parameters of the prediction model through backpropagation.
Citation Information
Patent Citations
Unsupervised domain adaptation method for pulmonary tuberculosis diagnosis through chest X-ray image
CN118015358A
Embryo image automatic focusing method and device based on multi-task mask feature modeling
CN118351400A
Deep learning method and system for embryonic development multi-stage classification
CN119131507A