Embryo growth stage prediction method and system based on self-supervised learning

Through a self-supervised learning method, the embryo image is masked and the model parameters are optimized using multi-core loss function, which solves the problem of insufficient generalization ability of the existing model and achieves higher prediction accuracy and robustness in the embryo growth stage.

CN119942249AActive Publication Date: 2025-05-06WUHAN MUTUAL UNITED TECH CO LTD

Patent Information

Application Number
CN202510432149.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

The existing embryo growth stage prediction model lacks generalization ability when facing brand new blastocyst images and cannot accurately predict the embryo growth stage.

Method used

Using a self-supervised learning method, by segmenting the embryo image into several blocks and randomly selecting half of the blocks for masking processing, a prediction model including the encoder and the decoder is constructed, and the model parameters are optimized using the multi-core loss function to improve the generalization ability and robustness of the model.

Benefits of technology

It enhances the model's ability to generalize new data, improves the robustness when facing complex images, and improves the accuracy of prediction of embryo growth stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942249A_ABST
    Figure CN119942249A_ABST
Patent Text Reader

Abstract

The invention discloses an embryo growth stage prediction method and system based on self-supervised learning, and the method comprises the following steps: collecting embryo images at different growth stages as a training set, and segmenting each embryo image in the training set into a plurality of blocks with the same size; randomly selecting half of the blocks from each embryo image for mask processing to obtain a mask training set; a prediction model comprising an encoder and a decoder is constructed, the encoder extracts a feature sequence from the mask training set, and the decoder reconstructs an embryo image according to the feature sequence; the kernel functions of the multiple weighted combinations serve as loss functions of the prediction model, the weight of each kernel function is adjusted through cross validation and gradient descent, parameters of the prediction model are adjusted and optimized through back propagation, and an optimal prediction model is obtained; and inputting a to-be-predicted embryo image into the optimal prediction model to obtain a growth stage prediction result of the embryo.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for predicting embryo growth stages based on self-supervised learning. Background Art

[0002] In the field of Assisted Reproductive Technology (ART), the evaluation of the developmental quality of blastocysts is a key step in improving the success rate of in vitro fertilization (IVF). Blastocysts are embryos that develop to the 5th to 6th day during in vitro fertilization (IVF). Their quality is of vital importance to embryo transfer decisions and the success rate of pregnancy. During the IVF process, medical experts use morphological evaluation to determine whether the blastocyst has a higher developmental potential. In order to improve the efficiency and accuracy of the evaluation, the use of machine learning technology to automatically predict embryo growth has become an important means to improve the success rate of IVF.

[0003] In recent years, self-supervised learning (SSL) has rapidly emerged in the field of machine learning. Especially in scenarios where data is scarce or the cost of labeling is high, self-supervised learning can significantly reduce the reliance on labeled data by using the relationships within the data for training. In embryo growth prediction, self-supervised learning methods can learn more feature representations from unlabeled data with limited labeled data, thereby improving the model's predictive ability.

[0004] Although self-supervised learning provides new possibilities for blastocyst growth prediction, existing technologies still face some challenges. The development process of blastocysts is complex and changeable, involving many aspects of embryo morphology, cell division rate, trophoblast development, etc. These characteristics vary greatly between individuals. As a result, traditional prediction models may perform well on training data, but they lack generalization ability when encountering new or unseen blastocyst image data, and cannot accurately predict the growth stage of the embryo. Summary of the invention

[0005] The present invention proposes a method and system for predicting embryo growth stages based on self-supervised learning, which solves the problem that the existing prediction model has insufficient generalization ability when facing new blastocyst images.

[0006] In order to solve the above technical problems, the present invention provides a method for predicting embryo growth stages based on self-supervised learning, comprising the following steps: Step S1: collecting embryo images at different growth stages as a training set, and dividing each embryo image in the training set into a number of blocks of the same size; Step S2: randomly select half of the blocks in each embryo image for mask processing to obtain a mask training set; Step S3: constructing a prediction model including an encoder and a decoder, wherein the encoder extracts a feature sequence from the mask training set, and the decoder reconstructs an embryo image according to the feature sequence; Step S4: using multiple weighted combined kernel functions as the loss function of the prediction model, adjusting the weight of each kernel function through cross-validation and gradient descent, and using back propagation to tune the parameters of the prediction model to obtain the optimal prediction model; Step S5: Input the embryo image to be predicted into the optimal prediction model to obtain the embryo growth stage prediction result.

[0007] Preferably, in step S3, the encoder extracting the feature sequence from the mask training set comprises the following steps: Step S31: Divide the mask training set into uniform spatiotemporal blocks according to the time dimension; Step S32: convert each spatiotemporal block into a high-dimensional vector through an embedding layer, and perform position encoding on each high-dimensional vector according to the temporal and spatial information; Step S33: the encoder includes a plurality of Transformer blocks, each Transformer block includes a self-attention layer and a multi-layer perceptron module, the self-attention layer captures the spatial and temporal correlations between high-dimensional vectors through a plurality of parallel attention heads, and the multi-layer perceptron module learns complex feature representations in high-dimensional vectors through two fully connected layers and a nonlinear activation function; Step S34: Use residual connection to add the output of each Transformer block to the input, and use the output of each Transformer block as the input of the next Transformer block to obtain the feature sequence of the mask training set.

[0008] Preferably, the loss function in step S4 is expressed as: ; In the formula, is the loss function; N is the total number of embryo images in the mask training set; For the i The embryo image corresponds to j The weight of the kernel function; For the j Kernel function; For the i embryo image; For the i Prediction value of embryo images; For the i The embryo image is in j The true value of a feature.

[0009] Preferably, step S4 comprises the following steps: Step S41: define the hyperparameter space of the prediction model and initialize the hyperparameter combination of the prediction model; Step S42: Divide the mask training set into several subsets, and for each subset: use all subsets except the current subset to train the prediction model, and evaluate the performance of the prediction model on the current subset; Step S43: Calculate the gradient of the multi-core loss function with respect to the weight of each kernel function, and use gradient descent to update the weight of each kernel function in the multi-core function; Step S44: Calculate the gradient of the multi-core loss function with respect to the prediction model parameters, and use gradient descent to update the hyperparameter combination of the prediction model; Step S44: repeating steps S41 to S43 until the set first iteration number is reached; Step S45: Record the performance indicators of each set of hyperparameter combinations in cross-validation, adjust the search space of hyperparameters according to the average performance of cross-validation, and repeat steps S42 to S44 until the set second iteration number is reached; Step S46: Calculate the average performance of all cross-validations, select the hyperparameter combination corresponding to the cross-validation with the best average performance as the optimal parameter, and obtain the optimal prediction model.

[0010] Preferably, after collecting embryo images at different growth stages as a training set in step S1, the training set is preprocessed, including the following steps: converting the collected embryo images into grayscale images, stacking the grayscale images, and converting all embryo images into a high-dimensional vector.

[0011] The present invention also provides an embryo growth stage prediction system based on self-supervised learning, which is implemented based on the above-mentioned embryo growth stage prediction method based on self-supervised learning, and includes an image preprocessing module, an embryo feature extraction module, a mask self-supervision module and a multi-core loss module; The image preprocessing module is used to collect embryo images and preprocess the collected embryo images; The embryo feature extraction module is used to extract features related to embryo growth from the preprocessed embryo image; The mask self-supervision module trains the prediction model by using the internal relationship of the embryo image through self-supervision learning; The multi-core loss module optimizes the prediction model through a multi-core loss function to enhance the prediction model's ability to capture features related to embryonic development.

[0012] Preferably, preprocessing the embryo images comprises the following steps: adjusting the sizes of all embryo images to be consistent, and converting the color embryo images into grayscale images.

[0013] Preferably, the embryo feature extraction module uses a deep learning network to identify and extract the morphological features of the embryo in the embryo image.

[0014] Preferably, the self-supervised learning includes the following steps: dividing the feature map extracted by the embryo feature extraction module into several image blocks, shielding high-proportion fragments, and only inputting the unshielded part into the encoder to extract potential features, and the decoder uses these features to reconstruct the shielded part.

[0015] Preferably, the multi-kernel loss module selects multiple kernel functions as loss functions of the prediction model, optimizes the weight of each kernel function, and optimizes the parameters of the prediction model through back propagation.

[0016] The benefits of the present invention include at least: 1. By dividing the embryo image into several image blocks and randomly selecting half of the image blocks for masking, the prediction model is forced to reconstruct the entire image from partial visible information, which enhances the model's generalization ability to new data; 2. The multi-kernel loss function enables the prediction model to better handle outliers and noise. Different kernel functions are suitable for learning different features in embryo images, which can provide more robust feature representation for model reconstruction and prediction, and improve the robustness of the model when facing complex images. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Embryo images collected by the embodiment of the present invention; Figure 2 A schematic diagram of a method flow chart of an embodiment of the present invention; Figure 3 This is a schematic diagram of a process of performing mask processing on an embryo image in an embodiment of the present invention; Figure 4 It is a schematic diagram of the structure of a bidirectional Mamba block in a prediction model according to an embodiment of the present invention; Figure 5 Schematic diagram of the system structure of an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.

[0019] In existing automated embryo growth prediction models, the reconstruction loss function based on the Masked Autoencoder (MAE) model is often used to train the model. Figure 1 As shown in the figure, there are usually a large number of outliers and noise in embryonic image data, and the mean square error (MSE) calculates the average of the squares of the differences between the predicted values ​​and the true values. Therefore, larger errors will be amplified, which will have a greater impact on the predictive ability of the model.

[0020] To solve this problem, the existing technology usually introduces the kernel loss function (Multiple Kernel Learning, MKL), which uses the nonlinear characteristics of the kernel function to capture the complex features in the data. Through kernel function mapping, the data can be transformed into a high-dimensional space, so that the model can better distinguish between normal and abnormal embryonic development patterns. In addition, different kernel functions, such as Gaussian kernels and polynomial kernels, are suitable for learning different features and can provide more robust feature representations for model reconstruction and prediction.

[0021] Although self-supervised learning and multi-kernel loss functions have solved some problems in blastocyst growth prediction to a certain extent, they still face some technical challenges. For example, how to effectively combine multiple kernel functions without significantly increasing the computational complexity, how to weigh the importance of different features, and how to select the optimal combination of kernel functions to improve the performance of the model. In addition, due to the diversity and complexity of embryonic development, the model may still show a certain lack of generalization ability when facing new or unseen blastocyst data.

[0022] Therefore, one of the core goals of the embodiments of the present invention is to design an efficient prediction model that can adapt to different blastocyst data characteristics under the condition of limited computing resources and labeled data. This requires that the model should not only be able to handle outliers and noise, but also have good generalization ability to adapt to the diversity and complexity of embryonic development. Through these improvements, the prediction accuracy of the model can be improved, and better results can be achieved in the field of automated embryo growth prediction.

[0023] like Figure 2 As shown, an embodiment of the present invention provides an embryo growth stage prediction method based on self-supervised learning, comprising the following steps: Step S1: Collect embryo images at different growth stages as a training set, and divide each embryo image in the training set into several blocks of the same size.

[0024] Specifically, the data collection work of the embodiment of the present invention covers a number of top reproductive medical centers, which provide a wealth of embryo image resources for this embodiment with their professional experience and advanced equipment in the field of assisted reproductive technology. Each embryo is photographed in the first five days of development, from the pronuclear stage to the blastocyst stage of embryonic development, a total of 500 pictures, and a total of nearly 100,000 images of embryonic development have been collected from multiple reproductive centers. For embryos with less than or more than 500 photos, the photos are copied, added or deleted on average from the existing photos until the number reaches 500.

[0025] These images were taken in a time-lapse incubator, which can continuously record the development of the embryo without affecting its development. The collected image dataset comprehensively covers all stages of embryo development from the pronuclear stage to the blastocyst stage, and includes images of empty dishes for comparison. The embryo images are sorted in sequence according to the embryonic development time, and standardized preprocessing operations are performed to ensure that the size of the embryo images is consistent, which is convenient for subsequent input into the model. This process ensures the quality and consistency of the data, providing a reliable basis for model training and prediction.

[0026] In the study of embryo growth prediction, the color information of the image is usually not the core of the analysis, and the grayscale image is sufficient to provide the key visual information required for autofocus. In order to simplify the processing flow and make full use of the advantages of grayscale images, the cv2.cvtColor() function in the OpenCV library is used to convert the color embryo image into a grayscale format. This conversion process not only removes the interference of color, but also retains the texture and contour information of the image, which is crucial for the task of embryo growth and development prediction. In order to adapt to the input requirements of the neural network model and optimize the efficiency of data processing, these grayscale images are specially processed. Specifically, in order to adapt to the input requirements of the neural network model and optimize the efficiency of data processing, all embryo images are stacked together. By adding a new dimension, a set of N The data of the embryo images are stacked along this dimension, and the shape of each tensor is ( N , H , W ),in N represents the length of the image sequence, H and W are the height and width of the image respectively. This processing method allows the visual information of each group of images to be passed to the neural network along the new dimension. At the same time, the data structure is more compact, which is conducive to the network's learning of image sequences and improves processing efficiency.

[0027] Step S2: Randomly select half of the blocks in each embryo image for masking to obtain a mask training set.

[0028] Specifically, in the field of self-supervised learning, masking technology plays a vital role by introducing random uncertainty into the data, stimulating the model to mine more stable and generalized feature representations. Figure 3 As shown, in an embodiment of the present invention, 50% of the pixels in the embryo image are masked. This process first divides the embryo image into 36 small blocks. This segmentation helps the model to understand and reconstruct the local features of the image more carefully, especially in the blastocyst image, the features of key areas such as the inner cell mass and the outer trophoblast are particularly important. Then, half of these small blocks are randomly selected for masking. This high-proportion masking operation simulates the partial occlusion that may be encountered in actual image acquisition, forcing the model to infer the features of the masked part from the limited visible part, thereby reconstructing the complete image.

[0029] After masking half of the image block, the model needs to use the features in the other half of the image block that is not masked to predict the content of the masked part. This process requires the model to learn how to extract key information from the visible part and use this information to predict the features of the masked part. In the study of embryonic growth prediction, this strategy not only improves the model's adaptability to occlusion and information loss, but also enhances the ability to extract a variety of complex features of the blastocyst. In this way, the model can better understand and handle the complexity and noise in the blastocyst image, especially the outliers and noise that may appear in key areas such as the inner cell mass and the outer trophoblast, which improves the robustness of the model in processing blastocyst growth sequences and provides new perspectives and strategies for the application of self-supervised learning in complex visual tasks.

[0030] Step S3: construct a prediction model including an encoder and a decoder, wherein the encoder extracts a feature sequence from the mask training set, and the decoder reconstructs the embryo image according to the feature sequence.

[0031] Step S4: Use multiple weighted combined kernel functions as the loss function of the prediction model, adjust the weight of each kernel function through cross-validation and gradient descent, and use back propagation to tune the parameters of the prediction model to obtain the optimal prediction model.

[0032] Specifically, the core architecture of the prediction model of the embodiment of the present invention is the MAE (Masked Autoencoder) model. Since self-supervised learning does not require manual annotation, MAE greatly reduces the annotation cost. At the same time, it improves the generalization ability of the model by reducing the learning difficulty caused by masking and reduces the risk of overfitting. In addition, MAE only processes part of the input fragments, which significantly reduces the computational cost and is particularly suitable for data with significant spatiotemporal redundancy such as embryonic development image sequences. And the embodiment of the present invention introduces a multi-core loss function based on the MAE model to enhance the model's ability to capture complex features of blastocyst images, reduce sensitivity to outliers, and enhance the model's learning ability for different feature subspaces. The entire model is mainly composed of four modules: an input preprocessing module, a feature extraction module, a mask self-supervision module, and a multi-core loss module.

[0033] In order to further improve the performance of the model for the embryo growth prediction task, the embodiment of the present invention adopts the state space model VideoMamba for efficient video understanding as an encoder, and the VideoMamba encoder is a core feature extraction component, which is specially designed for the spatiotemporal characteristics of video data. Its construction process first divides the video data into uniform spatiotemporal blocks according to the time dimension. These spatiotemporal blocks represent the basic units of continuous frames in the video and include various dynamic changes in the scene at different time points. Each spatiotemporal block is converted into a high-dimensional vector by an embedding layer. This conversion maps the blastocyst growth sequence to a format that the model can handle, and at the same time, by adding the position encoding of time and space, the features in the video are understood.

[0034] VideoMamba encoder architecture Figure 4As shown in the figure, the SSM framework is an integration of spring, spring MVC and mybatis frameworks. The encoder adopts a Transformer-based self-attention mechanism, which can effectively handle the spatiotemporal dependencies in the blastocyst growth sequence. Through multiple parallel attention heads, the model can capture the spatial and temporal correlations between video frames in different subspaces, and each attention head independently analyzes the local and global features in the video frame. This setting enables VideoMamba to process multiple features simultaneously in multiple dimensions, thereby improving the model's ability to understand video content. After the self-attention layer, the encoder uses a residual connection to add the output to the input, which helps the flow of information and alleviates the problem of gradient disappearance. Subsequently, the training process is further stabilized by the normalization layer to ensure that the output of each layer remains within a reasonable range. This design makes the encoder more stable during training and can quickly adapt to different video input modes. In order to enhance the model's expressiveness and nonlinear fitting capabilities, the VideoMamba encoder adds a multi-layer perceptron (MLP) module after each self-attention layer. The MLP consists of two fully connected layers with a nonlinear activation function ReLU in the middle, allowing the model to learn more complex feature representations.

[0035] The VideoMamba encoder is constructed by stacking multiple Transformer blocks, each of which contains a self-attention layer and an MLP layer, and the output of each block serves as the input of the next block. This layer-by-layer stacking structure enables the model to abstract and refine important information in the video data at different levels, gradually forming a high-dimensional feature representation of the video sequence. Ultimately, the VideoMamba encoder outputs a high-dimensional representation containing rich spatiotemporal features that can fully reflect the content of the blastocyst growth sequence.

[0036] The decoder trains the model's prediction capabilities by recovering masked areas from the feature representation extracted by the encoder. The decoder is designed to be lightweight and efficient, aiming to accurately reconstruct the morphological features of the blastocyst, allowing the model to fully utilize the powerful feature extraction capabilities of the VideoMamba encoder, while optimizing the performance of embryo growth prediction through task-specific decoders and multi-core loss functions, thereby achieving higher accuracy and robustness when processing blastocyst growth sequences. The encoder uses a masking strategy to extract features related to the quality of blastocyst development from the input image, including the morphology of the inner cell mass and trophectoderm, forcing the model to predict the masked part from the visible part. The decoder then reconstructs the feature representation extracted by the encoder into the form of original data, that is, reconstructs the image. In self-supervised learning, the decoder needs to recover the masked area from the output of the encoder.

[0037] The use of multi-kernel loss functions effectively enhances the ability of the prediction model to capture complex features related to blastocyst development. Traditional reconstruction loss functions usually use pixel-level mean square error (MSE), but due to the complexity and noise of embryonic image data, MSE will amplify the error and it is difficult to fully capture key local features. In order to solve this problem, an embodiment of the present invention selects a variety of kernel functions, such as a combination of Gaussian kernels, polynomial kernels, and linear kernels to replace the loss function. These kernel functions can map different features to high-dimensional space, capture multiple features such as edges, textures, and color distributions in blastocyst images, and optimize the importance of each kernel function to the task through weight adjustment. The use of multi-kernel functions can not only balance the impact of outliers in complex data, but also improve the reconstruction quality and the robustness of the model.

[0038] Specifically, the embodiment of the present invention selects four kernel functions based on various features such as the basic morphology and distribution characteristics of the image, the size and morphology consistency of the blastomere, the stratification and migration pattern of the cells. The weight coefficient is adjusted through automatic learning to determine the relative importance of each kernel function in the current task, thereby optimizing the quality of image reconstruction. The expression of the constructed multi-kernel loss function is: ; In the formula, is the loss function; N is the total number of embryo images in the mask training set; For the i The embryo image corresponds to j The weight of the kernel function; For the j A kernel function is used to map the original features into a high-dimensional space in order to capture the complex features in the blastocyst image; For the i embryo image; For the i Prediction value of embryo images; For the i The embryo image is in j The true value of a feature.

[0039] By constructing masked learning features and loss functions, better embryo growth prediction effects can be achieved. In the specific implementation of the method of the embodiment of the present invention, the Adam optimizer is selected to optimize the model in view of the characteristics of the algorithm's adaptive learning rate adjustment. The initial learning rate is set to 0.005, the batch size is 16, the number of training rounds is 60, and a learning rate decay strategy is adopted to ensure that the adjustment of model parameters is gradually refined during the training process. Specifically, after a certain number of epochs, the learning rate will decrease according to a predetermined decay rate to promote the model to more carefully approach the optimal solution in the later stage of training.

[0040] Cross-validation is used to evaluate the performance of a model after gradient descent optimization. The dataset is divided into several subsets, and one subset is selected each time as the validation set, and the rest are used as the training set. On the training set, the gradient descent algorithm continues to adjust the weights to minimize the loss function. The performance of the model is then evaluated on the validation set, and the weight coefficients are further adjusted based on these performance results. This process is repeated multiple times, and a different subset is selected as the validation set each time to ensure that the model performs consistently on different data subsets. Finally, the optimal weight coefficient is selected based on the average performance of the model during multiple validation processes.

[0041] By combining cross-validation with gradient descent, the weight coefficients of the multi-core loss function are optimized and the generalization ability is improved. Specifically, the weight coefficients are first randomly initialized, and then the weights are iteratively updated by the gradient descent algorithm. The K-fold cross-validation method is then used to evaluate the generalization ability of the model under different hyperparameter settings, including learning rate and batch size. By comparing the performance under different settings, the optimal hyperparameter combination is selected for formal training of the model. After the model training is completed, cross-validation is used again to evaluate the final performance of the model to ensure that the model has good predictive ability on unseen data. This process may require multiple iterations until satisfactory performance is achieved, thereby providing more accurate data for subsequent analysis and research. The model of the present invention is implemented using the Pytorch framework, and all models are trained and tested on an NVIDIA GeForce GTX 1660 GPU.

[0042] Through the mutual collaboration among the four modules, the model's comprehensive feature extraction capability for the blastocyst development process is jointly improved. By capturing the complementarity of the core functions of each module, the risk of overfitting caused by the model's dependence on a single module is reduced, further improving the system's prediction accuracy and generalization ability. In the model design, the feature extraction module enhances spatial information perception through a bidirectional state space model and position embedding, the mask self-supervision module uses a random masking mechanism to learn the global semantic relationship of the data, and the input preprocessing module provides high-quality standardized data for subsequent processes. Finally, the multi-core loss function module dynamically adjusts the weight ratio of the features by integrating multiple kernel functions, thereby optimizing the reconstruction and prediction performance of the model.

[0043] In addition, the labeling of embryo image data is a cumbersome task, which requires experienced professionals to comprehensively evaluate and judge the advantages and disadvantages of a group of blastocyst images. This process is very time-consuming and labor-intensive. Therefore, the embodiment of the present invention reasonably combines multi-core learning with self-supervised learning, effectively reducing the demand for labeled embryo images. The self-supervised task network model is constructed through the characteristic information of the embryo image itself, which not only reduces the need for professional participation, but also reduces the labeling cost of blastocyst images. At the same time, through the close linkage of these four modules, the model can also effectively reduce the sensitivity to abnormal data, significantly improve the processing ability of complex biomedical data and the generalization performance of unseen data.

[0044] Step S5: Input the embryo image to be predicted into the optimal prediction model to obtain the embryo growth stage prediction result.

[0045] In order to verify the effectiveness of the method proposed in the embodiment of the present invention, a large number of parameter tuning experiments and comparative experiments of different methods were carried out. Since a group of embryo images have high redundancy and time correlation, a design with an extremely high tubular mask ratio is adopted. This higher mask ratio helps to prevent information leakage during mask modeling, making embryo image gradient feature reconstruction a meaningful self-supervised training task, and also increases the difficulty of reconstruction to a certain extent. The embodiment of the present invention explores the accuracy when the mask ratio increases from 50% to 90%. The experimental results show that when the mask ratio is 50%, the growth prediction accuracy of the embryo image is 82.7%; when the mask ratio is 70%, the automatic focusing recognition accuracy of the embryo image is 84.9%; when the mask ratio is 80%, the recognition accuracy is 85.2%; when the mask ratio is 90%, the recognition accuracy can reach the highest 87.8%, which is 5.1% higher than the result under the mask ratio of 50%. This shows that at a higher mask ratio, the multi-task collaborative model of the present invention can achieve more excellent performance, which can further illustrate that the design of the present invention can force the adopted network model to capture more useful spatiotemporal information in embryo images.

[0046] In addition, in order to further verify the effectiveness of the method, the embodiment of the present invention compares the effects of three different loss functions on the experimental results.

[0047] Table 1 Prediction accuracy obtained on embryo images using different loss functions

[0048] As can be seen from Table 1, when the loss function selects the mean square error, the accuracy is 79.4%; when the loss function selects the cross entropy, the accuracy is 82.9%; when the loss function selects the multi-kernel function, the accuracy is 87.8%. This shows that since the growth of embryo images is caused by the joint action of multiple features, a single function as a loss function cannot fully calculate the loss, while in the multi-kernel function, the highest accuracy can be obtained when the four kernel functions work together, and the multi-kernels work together to achieve the best effect.

[0049] In summary, the method of the embodiment of the present invention makes full use of the relevant information contained in the embryo image, adopts the VideoMamba model, and combines the embryo image self-supervised mask to model and study the embryo growth prediction task. Experiments show that the method proposed in the present invention has more excellent performance.

[0050] like Figure 5 As shown, an embodiment of the present invention also provides an embryo growth stage prediction system based on self-supervised learning, which is implemented based on the above-mentioned embryo growth stage prediction method based on self-supervised learning, including an image preprocessing module, an embryo feature extraction module, a mask self-supervision module and a multi-core loss module.

[0051] The image preprocessing module is used to collect embryo images and preprocess the collected embryo images, adjust the sizes of all embryo images to be consistent, and convert color embryo images into grayscale images.

[0052] The embryo feature extraction module uses a deep learning network to extract features related to embryo growth from preprocessed embryo images.

[0053] The masked self-supervision module is used to divide the feature map extracted by the embryo feature extraction module into several non-overlapping image blocks, masking high-proportion fragments, and only input the unmasked parts into the encoder to extract potential features. The decoder uses these features to reconstruct the masked parts and uses the internal relationships of the embryo image to train the prediction model.

[0054] This module infers the semantics of the masked part through the context of the visible fragment, thereby effectively learning the association between global and local features. Since self-supervised learning does not require manual annotation, MAE greatly reduces the annotation cost. At the same time, it improves the generalization ability of the model by reducing the learning difficulty caused by masking and reduces the risk of overfitting. In addition, MAE only processes part of the input fragment, significantly reducing the computational cost, which is particularly suitable for data with significant spatiotemporal redundancy such as embryonic development image sequences.

[0055] The multi-kernel loss module selects multiple kernel functions as the loss functions of the prediction model, optimizes the weights of each kernel function, optimizes the parameters of the prediction model through back propagation, and enhances the prediction model's ability to capture characteristics related to embryonic development.

[0056] The traditional reconstruction loss function usually uses the mean square error (MSE) at the pixel level. However, due to the complexity and noise of embryo image data, the single loss function MSE, which is the average of the square of the difference between the predicted value and the true value, will amplify the error and it is difficult to fully capture the cell density of the inner cell mass and the cell arrangement of the outer trophoblast, as well as their changes at different developmental stages, which are crucial for evaluating the quality and developmental potential of the blastocyst. In order to solve this problem, the embodiment of the present invention uses a variety of kernel functions, such as a combination of Gaussian kernels, polynomial kernels and linear kernels to replace the loss function. These kernel functions can map different features to high-dimensional space, capture multiple features such as edges, textures and color distribution in blastocyst images, and optimize the importance of each kernel function to the task through weight adjustment. The use of multi-kernel functions can not only balance the impact of outliers in complex data, but also improve the reconstruction quality and the robustness of the model. The weight coefficients are dynamically adjusted through cross-validation, and combined with optimization algorithms such as gradient descent, the model performance is further optimized, and finally the model achieves efficient prediction of the quality of blastocyst development with limited labeled data.

[0057] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is more specific and detailed, but it cannot be understood as limiting the scope of the present invention. As long as there is no contradiction in the combination of these technical features, they should be considered as within the scope of this specification.

[0058] It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these modifications and improvements all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.

Claims

1. A method for predicting embryo growth stages based on self-supervised learning, characterized in that: The following steps are involved: Step S1: collecting embryo images at different growth stages as a training set, and dividing each embryo image in the training set into a number of blocks of the same size; Step S2: randomly select half of the blocks in each embryo image for mask processing to obtain a mask training set; Step S3: constructing a prediction model including an encoder and a decoder, wherein the encoder extracts a feature sequence from the mask training set, and the decoder reconstructs an embryo image according to the feature sequence; Step S4: using multiple weighted combined kernel functions as the loss function of the prediction model, adjusting the weight of each kernel function through cross-validation and gradient descent, and using back propagation to tune the parameters of the prediction model to obtain the optimal prediction model; Step S5: Input the embryo image to be predicted into the optimal prediction model to obtain the embryo growth stage prediction result.

2. The method for predicting embryo growth stages based on self-supervised learning according to claim 1, characterized in that: Step S3, the encoder extracting a feature sequence from the mask training set, comprises the following steps: Step S31: Divide the mask training set into uniform spatiotemporal blocks according to the time dimension; Step S32: convert each spatiotemporal block into a high-dimensional vector through an embedding layer, and perform position encoding on each high-dimensional vector according to the temporal and spatial information; Step S33: the encoder includes a plurality of Transformer blocks, each Transformer block includes a self-attention layer and a multi-layer perceptron module, the self-attention layer captures the spatial and temporal correlations between high-dimensional vectors through a plurality of parallel attention heads, and the multi-layer perceptron module learns complex feature representations in high-dimensional vectors through two fully connected layers and a nonlinear activation function; Step S34: Use residual connection to add the output of each Transformer block to the input, and use the output of each Transformer block as the input of the next Transformer block to obtain the feature sequence of the mask training set.

3. The method for predicting embryo growth stages based on self-supervised learning according to claim 1, characterized in that: The expression of the loss function in step S4 is: ; In the formula, is the loss function; N is the total number of embryo images in the mask training set; For the i The embryo image corresponds to j The weight of the kernel function; For the j Kernel function; For the i embryo image; For the i Prediction value of embryo images; For the i The embryo image is in j The true value of a feature.

4. The method for predicting embryo growth stages based on self-supervised learning according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: define the hyperparameter space of the prediction model and initialize the hyperparameter combination of the prediction model; Step S42: Divide the mask training set into several subsets, and for each subset: use all subsets except the current subset to train the prediction model, and evaluate the performance of the prediction model on the current subset; Step S43: Calculate the gradient of the multi-core loss function with respect to the weight of each kernel function, and use gradient descent to update the weight of each kernel function in the multi-core function; Step S44: Calculate the gradient of the multi-core loss function with respect to the prediction model parameters, and use gradient descent to update the hyperparameter combination of the prediction model; Step S44: repeating steps S41 to S43 until the set first iteration number is reached; Step S45: Record the performance indicators of each set of hyperparameter combinations in cross-validation, adjust the search space of hyperparameters according to the average performance of cross-validation, and repeat steps S42 to S44 until the set second iteration number is reached; Step S46: Calculate the average performance of all cross-validations, select the hyperparameter combination corresponding to the cross-validation with the best average performance as the optimal parameter, and obtain the optimal prediction model.

5. The method for predicting embryo growth stages based on self-supervised learning according to claim 1, characterized in that: After collecting embryo images at different growth stages as a training set in step S1, the training set is preprocessed, including the following steps: converting the collected embryo images into grayscale images, stacking the grayscale images, and converting all embryo images into a high-dimensional vector.

6. An embryo growth stage prediction system based on self-supervised learning, implemented based on an embryo growth stage prediction method based on self-supervised learning as claimed in any one of claims 1 to 5, characterized in that: It includes image preprocessing module, embryo feature extraction module, mask self-supervision module and multi-core loss module; The image preprocessing module is used to collect embryo images and preprocess the collected embryo images; The embryo feature extraction module is used to extract features related to embryo growth from the preprocessed embryo image; The mask self-supervision module trains the prediction model by using the internal relationship of the embryo image through self-supervision learning; The multi-core loss module optimizes the prediction model through a multi-core loss function to enhance the prediction model's ability to capture features related to embryonic development.

7. The embryo growth stage prediction system based on self-supervised learning according to claim 6, characterized in that: The preprocessing of the embryo images includes the following steps: adjusting the sizes of all embryo images to be consistent, and converting the color embryo images into grayscale images.

8. The embryo growth stage prediction system based on self-supervised learning according to claim 6, characterized in that: The embryo feature extraction module uses a deep learning network to identify and extract the morphological features of the embryo in the embryo image.

9. The embryo growth stage prediction system based on self-supervised learning according to claim 6, characterized in that: The self-supervised learning includes the following steps: dividing the feature map extracted by the embryo feature extraction module into several image blocks, shielding high-proportion segments, inputting only the unshielded parts into the encoder to extract potential features, and the decoder uses these features to reconstruct the shielded parts.

10. The embryo growth stage prediction system based on self-supervised learning according to claim 6, characterized in that: The multi-kernel loss module selects multiple kernel functions as loss functions of the prediction model, optimizes the weight of each kernel function, and optimizes the parameters of the prediction model through back propagation.

Citation Information

Patent Citations

  • Embryo division process analysis and pregnancy rate intelligent prediction method and system

    CN111785375A

  • Cloud computing resource scheduling method and system for geographic big data

    CN112035264A

  • Embryo development stage prediction and quality evaluation system based on edge enhancement

    CN116844143A

  • Unsupervised domain adaptation method for pulmonary tuberculosis diagnosis through chest X-ray image

    CN118015358A

  • Embryo image automatic focusing method and device based on multi-task mask feature modeling

    CN118351400A

Cited By

  • IVF-ET embryo selection method based on time difference imaging multi-task learning

    CN120356013A

  • Time sequence self-supervised learning method and system for rail transit engineering video images

    CN120930713A

  • Medicine bottle defect unsupervised detection method and system based on Vision Mama and dynamic mask generation

    CN121527044A

  • A medicine bottle defect unsupervised detection method and system based on vision mamba and dynamic mask generation

    CN121527044B

  • Embryo development quality prediction method and system based on multi-modal fusion and adaptive adjustment

    CN122413159A