IVF-ET embryo selection method based on time difference imaging multi-task learning

Through a multi-task learning framework and deep learning network, combined with improved ResNet3D and Bi-LSTM, the problem of insufficient subjectivity and timing characteristics of embryo evaluation is solved, and more accurate and objective evaluation of embryo quality and development potential is achieved, and evaluation efficiency and pregnancy rate are improved.

CN120356013AActive Publication Date: 2025-07-22CHENGDU UNIVERSITY OF TECHNOLOGY +1

Patent Information

Application Number
CN202510820338.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-22
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing embryo evaluation methods rely on single modal data, resulting in strong subjectivity, inaccurate evaluation results, difficulty in objectively describing embryo quality and development potential, and insufficient utilization of timing characteristics.

Method used

The multi-task learning method based on time difference imaging is adopted, and the spatiotemporal characteristics of embryonic development are extracted through the improved ResNet3D and Bi-LSTM networks, combined with hard parameter sharing and virtual predictors, and multi-task learning is used to realize the comprehensive evaluation of Istanbul consensus and Gardner grading.

Benefits of technology

It improves the accuracy and objectivity of embryo evaluation, reduces the bias of predicted results, improves evaluation efficiency, and can more accurately select embryos with high implantation potential, and improves clinical pregnancy rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356013A_ABST
    Figure CN120356013A_ABST
Patent Text Reader

Abstract

The invention discloses an IVF-ET embryo selection method based on time difference imaging multi-task learning, and belongs to the field of image data processing and assisted reproduction, and the method comprises the steps: firstly, effectively extracting spatial-temporal characteristics in an embryo time difference imaging sequence through an expansion three-dimensional convolutional network and a bidirectional long-short-term memory network by using a coding module; then, under a hard parameter sharing multi-task learning framework, a plurality of key indexes such as quality grading based on lstanbul consensus and development grade based on Gardner grading are predicted at the same time through a plurality of task specific predictors by utilizing shared feature representation; in the training process, an AdaTask optimization algorithm and a Frank-Wolfe method are adopted to optimize model parameters, and a virtual predictor mechanism is introduced to enhance the universality of the model; the method can more accurately, objectively and efficiently evaluate the embryo quality and development potential, relieve the burden of embryo scholars and improve the success rate of IVF and ET, and can be applied to the fields of animal husbandry, endangered species protection and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image data processing and assisted reproductive technology, and particularly to an IVF-ET embryo selection method based on time-lapse imaging multitask learning. Background Art

[0002] In vitro fertilization (IVF) technology has become an important means to solve infertility problems. The success rate of IVF depends to a large extent on embryo selection and transplantation processes. One of the main challenges faced by embryologists is how to select the embryo most likely to successfully implant and develop from multiple available embryos.

[0003] In in vitro fertilization-embryo transfer (IVF-ET) technology, embryo quality grading and blastocyst formation prediction are key links in determining embryo selection.

[0004] Currently, the internationally common embryo scoring standard is the Istanbul Consensus, and the embryo quality grading at the blastocyst stage mainly adopts the Gardner blastocyst grading mechanism. However, traditional embryo evaluation methods mainly rely on embryologists' observation and judgment of embryo morphology and development rate at different time points with the aid of an optical microscope. Due to subjective differences and experience differences between different observers or the same observer at different times, the embryo grading results are often subjective and inconsistent. This subjective evaluation method is not only inefficient but also limited in accuracy, and the clinical pregnancy rate after implantation into the mother is only about 35%.

[0005] With the development of artificial intelligence technology, deep learning technology has been widely used in the field of embryo evaluation. In terms of embryo quality grading, existing technologies mainly include methods based on convolutional neural networks (CNNs), such as the pronuclear embryo classification method proposed by Dimitriadis et al. and the embryo automatic grading system proposed by Cher et al. These methods achieve automatic evaluation of morphological features such as cytoplasm, pronucleus, and zona pellucida by analyzing embryo images. In terms of blastocyst formation prediction, existing technologies mainly adopt time series analysis methods, such as the time flow model combining DenseNet201 and LSTM proposed by Liao et al., and the multi-focal segment feature aggregation network proposed by Xie et al. These methods predict the potential of embryos to form blastocysts by analyzing the temporal features during embryo development.

[0006] However, the existing technologies have the following problems: First, existing methods mainly use single-modality data to evaluate embryo quality and developmental potential, mainly focusing on analyzing embryo images using CNN, resulting in many data related to embryo evaluation not being effectively utilized; second, existing methods only use a single evaluation standard to evaluate embryos, such as the Istanbul consensus or Gardner grading, but there is still no evaluation standard that can fully guarantee an objective and standardized evaluation of embryos. This leads to biased evaluation results of embryos by existing methods, making it difficult to objectively describe the quality and developmental potential of embryos. In addition, existing methods do not make sufficient use of time-lapse imaging data, and fail to effectively capture the temporal characteristics and dynamic change information during embryonic development.

[0007] Therefore, how to further improve the accuracy of embryo evaluation and screening, reduce the bias of prediction results, and improve the efficiency of embryo evaluation and screening is a technical problem that needs to be urgently solved in this field. Summary of the invention

[0008] The purpose of the present invention is to provide an IVF-ET embryo selection method based on time-difference imaging multi-task learning to solve the above problems.

[0009] In order to achieve the above object, the technical solution adopted by the present invention is as follows: an IVF-ET embryo selection method based on time-difference imaging multi-task learning, comprising the following steps: S1: Data preparation and preprocessing: S11. Data collection: Collect time-lapse imaging videos or image sequences of embryonic development during IVF; S12. Data preprocessing: preprocessing the time-lapse imaging video or image sequence; S2: Extract spatiotemporal features through encoding module: S21. Input encoding module: input the image sequence preprocessed by S12 into the encoding module; S22. ResNet3D feature extraction: using the improved ResNet3D network to process the input sequence and extract the spatiotemporal features of the embryonic development process; specifically, the improvement of the "improved ResNet3D network" is to reduce the number of parameters of the standard 3D convolution and use 3D depthwise separable convolution (Depthwise Separable Conv), and improve the convolution kernel size to 13×3×3; S23. Regularization: Apply a spatial Dropout layer after the ResNet3D layer to maintain dimensionality; S24. Feature pooling: Apply spatial max pooling to obtain features with a dimension of 8×256, capturing significant spatial structure information; apply spatial average pooling to obtain features with a dimension of 8×256, retaining the overall spatial information; S25. Temporal feature fusion at S25: Feed the results of max pooling and average pooling into a bidirectional long short-term memory network layer; S3: Multi-task learning and optimization: S31. Shared representation: The feature vector output in step S2 is used as the shared representation and input into subsequent task-specific predictors, where is the input sample, is the shared encoder, is the encoded feature vector; S32. Task-specific predictors: Construct independent predictors for each evaluation task ; S33. Application of the multi-task learning framework: Adopt a hard parameter sharing structure to share the parameters of the shared encoder ; adopt the Frank-Wolfe-based method, regard the problem as multi-objective optimization, and iteratively update the shared parameters by solving sub-problems to find the Pareto optimal solution; S34. Model training and optimization: Loss function: For classification tasks, adopt the cross-entropy loss function as ; the total loss is the sum of the losses of each task; , where: is the cross-entropy loss of the k-th task; is the true label of class c, that is, the value in the one-hot encoding, 0 or 1; log is the natural logarithm function; is the probability of class c predicted by the model; Optimizer: Adopt the AdaTask optimization method to update the model parameters ; Training process: Use the labeled embryo dataset to train the model, and iteratively update the network parameters through backpropagation and the optimizer until the model converges or reaches the preset number of training epochs; The present invention preferably searches for the Pareto optimal solution based on Frank-Wolfe: Multi-objective optimization perspective: Multi-task learning can be regarded as a multi-objective optimization problem, that is, minimizing the loss functions of all tasks simultaneously . The goal is to find the Pareto optimal solution, that is, it is impossible to improve the performance of any one task without sacrificing the performance of at least one task.

[0010] Frank-Wolfe (Conditional Gradient) method: This method can be used to find Pareto stationary points. Its core idea is to find an effective descent direction by solving a linear programming subproblem in each iteration. Specifically, it calculates the gradients of all tasks , and searches for a direction (usually the negative gradient direction of a certain task or its convex combination) such that the losses of all tasks have the potential for improvement or reach a certain balance in this direction.

[0011] Solving the subproblem: In the update of the shared parameters , the FW algorithm needs to solve the following subproblem: Find the that minimizes , where G is a matrix whose columns are the gradients of each task with respect to the shared parameters , is the weight vector (belonging to the simplex). This usually means choosing the single task gradient direction that is most opposite to the current combined gradient direction.

[0012] Parameter update: Then update the shared parameters along the found direction . This method helps to balance between the objectives of different tasks and find a set of solutions that can better balance the performance of all tasks, rather than simply optimizing the weighted sum.

[0013] S4: Classification prediction and embryo selection: S41. Feature input: Input the shared feature vector output by the encoding module into the task-specific predictors of the classification module; S42. Multilayer perceptron processing: First hidden layer: 128 neurons, ReLU activation function, followed by Dropout(0.5), Second hidden layer: 64 neurons, ReLU activation function, followed by Dropout(0.5), S43. Output prediction: Istanbul consensus classification task: The output layer has 5 neurons, and the Softmax activation function is applied to output the probability distribution corresponding to 5 quality levels; Gardner grading task: The output layer has 6 neurons, and the Softmax activation function is applied to output the probability distribution corresponding to 6 development levels; S44. Embryo selection decision: Based on the multiple prediction results output by the model, combined with clinical experience and specific requirements, formulate an embryo selection strategy.

[0014] As a preferred technical solution, In step S11, each embryo sample contains consecutive image frames covering key developmental stages; each sample contains 32 grayscale images, and the resolution of each frame is 256×256 pixels; In step S12, the preprocessing includes image alignment, brightness normalization, and denoising operations.

[0015] As a preferred technical solution, In S22, the ResNet3D network contains three 3D residual blocks, and each residual block contains two 3D convolutional layers with a convolutional kernel size of 13×3×3; after being processed by multiple residual blocks, the dimension of the feature map gradually decreases, and the final output feature map dimension is: 8×32×32×256.

[0016] In S23, the dimension 8×32×32×256 is kept unchanged to prevent overfitting; In S24, applying spatial max pooling and applying spatial average pooling both result in features with a dimension of 8×256; In S25, each direction of the bidirectional long short-term memory network layer contains 128 units, processes the information in the time dimension, and captures the temporal dependence relationship of embryo development; the bidirectional long short-term memory network layer finally outputs a 256-dimensional feature vector that fuses spatio-temporal information.

[0017] Regarding the bidirectional long short-term memory network layer, namely Bi-LSTM: Function: Bi-LSTM is good at processing time series data and can learn the dependence relationship in the time dimension. The calculations of the forget gate, input gate, and output gate of Bi-LSTM can be expressed as: , , , Among them, 、 and are the activation values of the forget gate, input gate, and output gate respectively, and b are learnable parameters, is the sigmoid function. The update of the cell state and hidden state can be expressed as: , , , where: t represents the time step, is the forget gate, is the input gate, is the output gate, 、 , , is the weight matrix, , , , are the bias vectors, is the hidden state at the previous moment, is the input at the current moment, is the sigmoid activation function, is the candidate cell state, is the cell state at the current moment, is the cell state at the previous moment, is the hidden state at the current moment, is the element-wise multiplication operation, and tanh is the hyperbolic tangent activation function.

[0018] As a preferred technical solution, In S32, the predictor adopts a multi-layer perceptron structure, including: Input layer: Receives the feature representation from the shared encoder; Hidden layer: Contains multiple fully connected layers, each followed by a ReLU activation function and Dropout regularization, for non-linear transformation of features and preventing overfitting; Output layer: Uses the Softmax activation function according to the specific task to output the probability distribution of the corresponding level: , where: is the probability that the input x belongs to the category c, is the original score (logits) of the category c, is the total number of categories of the task k, and exp is the natural exponential function; The task includes Istanbul consensus grading and Gardner grading.

[0019] The classification module of the present invention adopts a multi-layer perceptron (MLP) structure for predicting and evaluating embryo quality.

[0020] As a preferred technical solution, In S33, the hard parameter sharing structure: contains a shared encoder and task-specific predictors , and the objective function is the sum of the losses of all tasks: , where: is the set of all model parameters, is the parameter of the shared encoder, is the parameter of the k-th task-specific predictor, is the shared encoder, is the k-th task-specific predictor, is the i-th input sample, is the true label of the i-th sample on the k-th task, is the loss function of the k-th task, is the number of samples, and K is the number of tasks, represents function composition.

[0021] As a preferred technical solution, In S33, a virtual predictor is introduced, whose structure is the same as that of the true predictor but the parameters are randomly initialized and fixed during the training process; by constraining the gradient norm of the virtual predictor, the shared encoder is forced to generate representations that are friendly to any predictor, rather than only adapting to the optimal predictor; A penalty term for the gradient norm of the virtual predictor is added to the training objective to improve the generality of the shared representation, where the loss function of a single task is the original loss function with a penalty term for the gradient norm of the virtual predictor added: , The specific form is: , where: is the predicted value of the k-th task for the i-th sample, is the predicted value of the virtual predictor for the i-th sample, is the shared representation of the i-th sample, is the penalty term coefficient, used to balance the task loss and the virtual gradient norm; The total loss function of multi-task is the sum of all task losses: , where the loss of each task consists of two parts: True predictor loss : The traditional task loss generated by the true predictor generating the prediction and calculated; Virtual predictor gradient norm penalty: By the virtual predictor generating the prediction and calculating the gradient norm of its loss with respect to its own parameters, which is used to improve the generality of the shared representation.

[0022] where: is the total loss of the k-th task, is the number of samples, is the true predictor loss function for the k-th task, is the true label of the i-th sample for the k-th task, is the predicted value of the true predictor for the i-th sample of the k-th task, is the virtual predictor loss function for the k-th task, is the predicted value of the virtual predictor for the i-th sample of the k-th task, are the parameters of the virtual predictor for the k-th task, is the gradient with respect to the parameters of the virtual predictor, is the Frobenius norm, is the penalty coefficient.

[0023] Generalization quantification: Definition: The generalization of the shared encoder is defined as the reciprocal of the difference between the "optimal predictor loss" and the "arbitrary predictor loss". If any randomly initialized "virtual predictor" can achieve a low loss based on the shared representation, then the generalization is high; Theorem: Under the assumption of convexity of the loss function, the generalization is inversely proportional to the gradient norm of the virtual predictor: , Note: is the generalization measure of the shared encoder, are the parameters of the virtual predictor for the k-th task, represents the gradient with respect to the parameters of the virtual predictor, is the Frobenius norm, and α represents a proportional relationship.

[0024] As a further preferred technical solution, In step S34, the classification task includes the Istanbul consensus and the Gardner grading; the total loss includes a penalty term for the gradient norm of the virtual predictor; The AdaTask optimization method is as follows: Initialization: Set the exponential decay factor v, the initial learning rate , the smoothing factor ε, and initialize the cumulative gradient variable for each task k; Gradient calculation: At training step t, calculate the gradient of the parameter θ for each task; Cumulative gradient update: For each task, update the cumulative gradient by exponential decay averaging: , Parameter update amount calculation: Calculate the parameter update amount for each task: , Model parameter update: Update the model parameters according to the parameter update amounts of all tasks: , where: t is the training step, k is the task index, i is the parameter index, is the cumulative gradient of task k for parameter i at step t, is the gradient of task k for parameter i at step, v is the exponential decay factor, is the initial learning rate, ε is the smoothing factor, is the update amount of task k for parameter i, is the value of parameter i at step t.

[0025] Adopt the AdaTask optimization method: Task-based cumulative gradient separation: When the traditional adaptive learning rate optimizer is used to optimize the shared parameters of the MTL model, the statistic will aggregate the gradients of all tasks, which is likely to cause some tasks to dominate the cumulative gradient and learning rate. AdaTask maintains a cumulative gradient variable for each shared parameter and task separately, avoiding task conflicts and improving the optimization efficiency. The present invention preferably realizes it by combining RMSProp and AdaTask.

[0026] Optimization effect: By maintaining independent cumulative gradient statistics for each task, AdaTask can avoid gradient interference between tasks, enabling each task to obtain an appropriate learning rate, thereby improving the overall performance of multi-task learning.

[0027] As a preferred technical solution, In S44, the final prediction result is obtained by taking the category with the highest probability: , where: is the predicted category of the k-th task, argmax means taking the category c that makes the probability the largest, is the predicted probability that the input x belongs to category c.

[0028] This probability-based prediction method not only provides the grade evaluation of the embryo, but also gives the confidence of the prediction, which is helpful for clinical decision-making.

[0029] The main innovation of the method of the present invention lies in adopting multi-task learning and using two optimization methods to solve various problems in multi-task learning to improve the training effect.

[0030] The innovation of the multi-task learning framework of the present invention lies in: 1. The present invention first uses this framework in the field to solve the embryo selection problem; 2. Optimize multi-task learning by using an adaptive learning rate and a method for enhancing generality: 2.1 The generality enhancement strategy mainly solves the problem that the feature extraction module extracts single-task features with bias; 2.2 The adaptive learning rate solves the problem that the adjustment of the learning rate of a single task is too large and affects the remaining tasks, and realizes the separation of the learning rates of each task.

[0031] The present invention can improve the accuracy of embryo evaluation and screening and reduce the bias of prediction results. Specifically: In the direction of improving accuracy: 1. Adaptive learning rate The adaptive learning rate mechanism adopted by the present invention can dynamically adjust the learning rate according to the convergence speed and gradient information of different tasks, avoiding the problems of insufficient learning or premature stopping of some tasks that may be caused by the traditional fixed learning rate in the multi-task scenario, enabling the model to more fully learn the deep features of each task, and thus improving the overall prediction accuracy. For example, in the early embryo development stage and the blastocyst formation stage, their dynamic characteristics and the indicative weights for the final developmental potential may be different, and the adaptive learning rate can better capture the contributions of these dynamic changes to different prediction tasks; 2. Virtual prediction mechanism By introducing the method of the virtual prediction mechanism, the present invention can reduce the overfitting of the model to a specific training data distribution and improve the generalization ability of the model on unseen new data or data from different clinics.

[0032] In terms of reducing bias: 1. Adaptive learning rate By dynamically adjusting the learning pace for different tasks or data subsets, the model can avoid premature convergence to a sub-optimal solution due to some dominant features or majority-group data, thus giving more sufficient learning opportunities to the minority group or insignificant features, enabling the model to show more consistent performance for different sources or types of embryo data, and reducing the prediction bias caused by data imbalance or feature differences; 2. Virtual prediction mechanism The virtual prediction mechanism forces the model to learn features that are robust to different data distributions, rather than simply relying on accidental associations or biased features that may exist in the training data. This enables the model to make more fair and less biased predictions when facing data with different potential biases.

[0033] Compared with existing artificial intelligence-based evaluation methods, for example, compared with the method based on multi-task deep learning and dynamic programming (MTDL-DP) proposed by Zihan Liu et al.: 1. Different evaluation criteria Different from the MTDL-DP method of Zihan Liu et al. which mainly focuses on the morphological classification at the embryonic development stage, through an innovative multi-task learning framework, the present invention not only integrates the macroscopic developmental potential judgment of the Istanbul Consensus and the fine morphological indicators of Gardner grading, but also endeavors to directly predict the evaluation indicators highly relevant to clinical pregnancy outcomes, achieving a leap from'stage identification' to 'potential prediction', and the evaluation results are more clinically instructive; 2. Different learning rate adjustment strategies The adaptive learning rate mechanism introduced in the present invention can dynamically adjust the learning weights according to the complexity and convergence of each sub-task in multi-task learning, avoiding the problems of insufficient learning or overfitting of some key tasks caused by improper learning rate setting in traditional MTDL methods, ensuring the full learning of the features of each evaluation dimension, and thus improving the accuracy of evaluation as a whole. In contrast, if existing methods such as MTDL-DP adopt a fixed learning rate, it may be difficult to optimally balance the learning of classification tasks at different developmental stages; 3. Different optimizations for data generality Aiming at the problems of performance degradation and prediction bias that may occur when existing AI methods (including MTDL-DP and most CNN-based methods) are applied to different data sources (such as different reproductive centers, different culture systems, different imaging devices), the present invention adopts a virtual prediction mechanism. This strategy forces the model to learn the robust essential features of embryos across domains rather than the biased information of specific data sets. Therefore, when dealing with data from different sources, the present invention can exhibit stronger generalization ability and lower prediction bias, making the evaluation results more objective and fair.

[0034] Adopting the method of the present invention, including using improved ResNet3D (3D depthwise separable convolution, optimizing the convolutional kernel) and Bi-LSTM networks, etc., can automatically and precisely capture the three-dimensional spatial morphological structure and its dynamic evolution information during the whole process of embryo development from early stage to blastocyst formation in time-lapse imaging sequences, as well as the long-term temporal dependence relationship between these features. This means that the model can identify subtle but crucial dynamic patterns and morphological cues related to embryo developmental potential (whether it can successfully implant and develop into a fetus eventually), which are difficult or easy to be overlooked by the human eye. For example, specific cell division times, minor changes in fragmentation degree, the expansion speed and uniformity of the blastocoel cavity, etc., which may be related to the chromosomal state and intrinsic vitality of the embryo. More precisely identifying these "high-quality signals" can more accurately screen out embryos with high implantation potential. Through hard parameter sharing, simultaneously learn and predict the embryo grades based on the Istanbul Consensus (macro developmental potential) and Gardner grading (fine morphological indicators). This enables the evaluation results to integrate the advantages of different evaluation systems and form a more comprehensive and three-dimensional understanding of embryo quality. For example, an embryo may perform mediocrely in a certain sub-item of Gardner grading, but perform excellently in the Istanbul Consensus assessment indicating the overall developmental potential (or vice versa). Multi-task learning can balance this information, avoid the one-sidedness that may be brought by a single evaluation criterion, and thus make a judgment closer to the true developmental potential of the embryo and select embryos with better comprehensive quality. The AI model performs well on a specific dataset, but its performance may decline when applied to data from different centers, different devices, or different culture systems, that is, the generalization ability is insufficient. By introducing a "virtual predictor" with fixed and randomly initialized parameters and adding its gradient norm as a penalty term to the loss function, it forces the feature representation learned by the shared encoder to not only adapt to the current real task predictor, but also be able to adapt to potential prediction tasks or data distribution changes not clearly defined in the training. This means that the model learns more essential and universal embryo development laws, rather than the "preferences" or noises of a specific dataset. Therefore, when applied to new embryo data from different sources (such as different reproductive centers, different imaging devices, different culture conditions), the model can still maintain a high evaluation accuracy. This improvement in robustness ensures that truly high-quality embryos can be screened out in a wider range of clinical practices, thereby steadily increasing the pregnancy rate. After applying the method of the present invention, it is reasonably expected that the clinical pregnancy rate can be increased by 10%-20%.

[0035] Compared with the prior art, the advantages of the present invention are as follows: The present invention constructs a new embryo evaluation model that takes into account both the Istanbul consensus and Gardner grading, achieving more objective and accurate embryo evaluation. At the same time, through an innovative multi-task learning framework and optimization strategy, the accuracy of embryo evaluation and screening is improved, and the bias of prediction results is reduced. The present invention can reduce the workload of embryologists, improve the efficiency of embryo evaluation and screening, avoid observation omissions and judgment errors caused by fatigue, and provide a technical solution that can be used for reference in fields such as animal husbandry, endangered species protection, and embryonic stem cell research, contributing to the realization of ecological balance and sustainable development. Specifically: (1) Improved the accuracy of embryo evaluation: Improvement in accuracy and objectivity brought by the deep spatio-temporal feature extraction and multi-dimensional collaborative evaluation of the present invention: The ResNet3D and bidirectional LSTM (i.e., Bi-LSTM) networks in the encoding module adopted by this method can automatically and precisely capture the spatio-temporal dynamic features of the entire development process of embryos from fertilization to the blastocyst stage. The reason is that ResNet3D can extract deep information on the evolution of spatial morphology over time from volumetric data, while Bi-LSTM further captures the long-term dependence relationships and dynamic patterns among these spatio-temporal features. The organic combination of the two overcomes the limitations of traditional manual evaluation that relies on the subjective visual judgment of embryologists, because the machine can identify and quantify subtle changes that are difficult for the human eye to detect or easily overlooked, thus making the depth and consistency of feature extraction far exceed traditional manual observation, laying a solid foundation for accurate evaluation.

[0036] (2) Improved the objectivity of embryo evaluation: More comprehensive objective evaluation: By combining a multi-task learning framework with hard parameter sharing, this method can simultaneously learn and predict the embryo grades based on multiple industry-recognized evaluation criteria such as the Istanbul consensus and Gardner grading. The logic is that the shared encoder is optimized to learn a general feature representation that is beneficial to all tasks, which means the model is forced to understand embryo quality from multiple complementary perspectives (i.e., different evaluation criteria). This enables the evaluation results to integrate the advantages of different evaluation systems, overcomes the bias or incomplete information problems that may be brought about by a single evaluation criterion, and thus the output evaluation conclusions are more comprehensive and objective; (3) Improved the robustness and reliability of embryo evaluation: Improvement in robustness and reliability brought by the dual optimization mechanism of learning quality and efficiency: The present invention fundamentally addresses the core challenges of feature generality and task specificity, as well as optimization efficiency and task balance in multi-task learning through the precise collaboration of the virtual predictor mechanism and the AdaTask optimization algorithm, and then achieves a significant improvement in learning quality and efficiency, enhancing the robustness of the model and the reliability of the evaluation results: The virtual predictor mechanism enhances the model's generality and robustness: This mechanism introduces a penalty on the gradient norm of a "virtual predictor" with fixed parameters and random initialization into the optimization objective. The principle is that this forces the feature representation learned by the shared encoder to be adaptable not only to the current real-task predictor but also to potential prediction tasks or data distribution changes not explicitly defined during training. Therefore, the model's adaptability to data variations, noise interference, and different clinical scenarios that may exist in the real world is enhanced, showing better generalization ability and robustness; The AdaTask optimization algorithm ensures the efficiency and balance of multi-task learning: When optimizing shared parameters in multi-task learning, the gradients of each task may conflict or have an imbalance in magnitude. The AdaTask algorithm solves this problem by independently maintaining cumulative gradient statistics for the gradient components contributed by each task to the shared parameters and adaptively adjusting their learning rates. This ensures that even while pursuing a general feature representation (guided by the virtual predictor mechanism), each specific evaluation task (such as lstanbu grading, Gardner grading) can be fully and specifically optimized, avoiding the situation where some tasks are under-learned due to their gradients being dominated or overwhelmed during the optimization process, thereby improving the overall optimization efficiency and the balance of the performance of each task; The synergistic effect improves the overall performance: The virtual predictor mechanism ensures the "breadth" of the learning direction and the "robustness" of the features, while AdaTask guarantees the "efficiency" of multi-task progress and the "fairness" of the learning process in this optimization direction. The logical advantage of their organic combination is that they jointly enable the model to stably and efficiently learn embryo features that are both widely applicable (insensitive to unknown data variations and not easily overfitted to specific task subsets) and perform excellently on various specific evaluation indicators (well-trained in each evaluation dimension). This carefully designed combination enables this method to output more reliable, accurate, and clinically valuable embryo evaluation results in complex multi-criteria evaluation scenarios, thus comprehensively improving the scientific nature and success rate of embryo selection.

[0037] (4) Greatly improve the evaluation efficiency, lower the application threshold, and expand the scope of technology application: Automation and high efficiency: The fully automated evaluation process, because it performs fast and deterministic calculations based on a pre-trained deep learning model, can significantly reduce the workload of embryologists, freeing them from the tedious, repetitive, and time-consuming film reading work, enabling them to focus more on clinical decision-making and difficult cases. The shortening of the evaluation cycle is directly due to the high efficiency of the calculation and also because it minimizes human intervention, reduces the risks of evaluation errors caused by subjective judgment, visual fatigue, and individual operation differences, and helps reduce related operating costs; (5) Wide applicability and promotion value: The core technical solution of the present invention, that is, combining deep spatio-temporal feature extraction with an optimized multi-task learning framework for dynamic evaluation of complex biological processes, not only has great significance for innovating embryo selection in the field of human assisted reproduction, but also has promotional value.

[0038] The logical basis for its promotion lies in: the inherent advantages demonstrated by this methodology in terms of accuracy, objectivity, robustness, and efficiency through the above mechanisms, which have considerable universality. Therefore, it can be conveniently applied to other fields that require fine evaluation of dynamic sequence data, such as the breeding of excellent varieties in animal husbandry, the formulation of breeding protection strategies for endangered species, and the research on the differentiation potential and developmental trajectory of embryonic stem cells. This is not only expected to bring significant social and economic benefits, but also will strongly promote the technological progress of related interdisciplinary fields. Brief Description of the Drawings

[0039] Figure 1 is a schematic diagram of the brief process of this method; Figure 2 is a schematic diagram of the structure of the residual 3D convolutional module in this method; Figure 3 is a schematic diagram of the structure of the bidirectional LSTM module in this method; Figure 4 is a schematic diagram of the specific algorithm process of the virtual prediction mechanism in this method; Figure 5 is a schematic diagram of the model framework adopted in this method. Detailed Embodiment

[0040] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0041] An IVF-ET embryo selection method based on time-lapse imaging multi-task learning. For the schematic diagram of the brief process of the entire method, please refer to Figure 1 , and its specific process includes the following steps: S1: Data preparation and preprocessing S11. Data collection: Collect time-lapse imaging videos or image sequences of embryo development during the IVF process. Each embryo sample contains consecutive image frames covering key developmental stages (in this embodiment, including from fertilization to the blastocyst stage). In this embodiment, each sample contains 32 grayscale images, and the resolution of each frame is 256×256 pixels; S12. Data preprocessing: Preprocess the original image sequence, including operations such as image alignment, brightness normalization, and denoising. These operations all adopt well-known existing methods to improve the stability and accuracy of subsequent model processing.

[0042] S2: Extract spatio-temporal features through the encoding module: S21. Input the encoding module: Input the preprocessed image sequence (dimension: 32×256×256×1) into the encoding module; S22. ResNet3D feature extraction: Use the improved ResNet3D network to process the input sequence and extract spatio-temporal features during embryonic development. The ResNet3D network contains three 3D residual blocks. Figure 2 The structural schematic diagram of the residual 3D convolution module in this method is shown. Each residual block contains two 3D convolution layers, and the convolution kernel size is 13×3×3. After being processed by multiple residual blocks, the dimension of the feature map gradually decreases, and the final output feature map dimension is: 8×32×32×256; S23. Regularization: Apply a SpatialDropout layer after the ResNet3D layer to keep the dimension 8×32×32×256 unchanged to prevent overfitting; S24. Feature pooling: Apply Spatialmax.pool to obtain features with a dimension of 8×256, capturing significant spatial structure information; apply Spatialavg.pool to obtain features with a dimension of 8×256, retaining the overall spatial information; S25. Temporal feature fusion: Input the results of max pooling and average pooling (both with a dimension of 8×256) into a bidirectional long short-term memory network (Bi-LSTM) layer. Figure 3 The structural schematic diagram of the Bi-LSTM module in this method is shown. Each direction of this Bi-LSTM layer contains 128 units, processes the information in the time dimension (8 time steps), and captures the temporal dependence of embryonic development. The Bi-LSTM layer finally outputs a 256-dimensional feature vector that fuses spatio-temporal information.

[0043] S3: Multi-task learning and optimization S31. Shared representation: The 256-dimensional feature vector output in step S2 is used as the shared representation , and is input into the subsequent task-specific predictors; S32. Task-specific predictors: Construct independent predictors for each evaluation task (such as lstanbul consensus grading and Gardner grading) ; In this embodiment, the predictor adopts a multi-layer perceptron (MLP) structure.

[0044] S33. Application of the multi-task learning framework: Adopt a hard parameter sharing structure to share the parameters of the encoder , and introduce a virtual predictor , which has the same structure as the real predictor but with randomly initialized and fixed parameters. Figure 4 Specifically shows the schematic diagram of the algorithm flow of the virtual prediction mechanism in this method. Add a penalty term for the virtual predictor gradient norm to the training objective , to improve the generality of the shared representation; adopt the Frank-Wolfe-based method, regard the problem as multi-objective optimization, and iteratively update the shared parameters by solving sub-problems to find the Pareto optimal solution; S34. Model training and optimization: Loss function: For classification tasks (Istanbul consensus and Gardner grading), use the cross-entropy loss function (Cross EntropyLoss) as . The total loss is the sum of the losses of each task (which may include the virtual gradient penalty term); Optimizer: Use the AdaTask optimization method (LAdaTask-RMSProp) to update the model parameters . AdaTask maintains independent cumulative gradient statistics for each task and shared parameters (or layers), avoiding task conflicts and improving the optimization efficiency. Set appropriate initial learning rates , momentum parameters (such as of Adam), regularization coefficients and other hyperparameters; Training process: Use the labeled embryo dataset to train the model, and iteratively update the network parameters through backpropagation and the optimizer until the model converges or reaches the preset number of training epochs.

[0045] S4: Classification prediction and embryo selection: S41. Feature input: Input the 256-dimensional shared feature vector output by the encoding module into each task-specific predictor (MLP) of the classification module; S42. MLP processing: First hidden layer: 128 neurons, ReLU activation function, followed by Dropout(0.5), Second hidden layer: 64 neurons, ReLU activation function, followed by Dropout(0.5), S43. Output prediction: Istanbul consensus classification task: The output layer has 5 neurons, apply the Softmax activation function, and output the probability distribution corresponding to 5 quality levels; Gardner grading task: The output layer has 6 neurons, apply the Softmax activation function, and output the probability distribution corresponding to 6 development levels; S44. Embryo selection decision: Based on multiple prediction results of the model and combined with clinical experience and specific requirements, this embodiment proposes a weighted scoring system to formulate an embryo selection strategy. The system quantifies the selection by assigning weights to two prediction results and calculating the comprehensive score of the embryo. Specifically, the comprehensive score (Score) of each embryo is calculated by the following formula: , wherein, and respectively represent the weights of the Istanbul grade probability and the Gardner grade probability, and the sum of the two is 1. According to the final score, the most suitable embryo for transplantation is selected. As shown in Figure 5 , a schematic diagram of the complete model framework adopted in this method is shown.

[0046] The above embodiments are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An IVF-ET embryo selection method based on multi-task learning of time-difference imaging, characterized in that, It includes the following steps: S1: Data preparation and preprocessing: S11. Data collection: Collect time-lapse imaging videos or image sequences of the embryo development process during the IVF process; S12. Data preprocessing: Preprocess the time-lapse imaging videos or image sequences; S2: Extract spatio-temporal features through an encoding module: S21. Input to the encoding module: Input the image sequence preprocessed in S12 into the encoding module; S22. ResNet3D feature extraction: Use an improved ResNet3D network to process the input sequence and extract spatio-temporal features during embryo development; S23. Regularization: Apply a spatial Dropout layer after the ResNet3D layer to maintain the dimension; S24. Feature pooling: Apply spatial max pooling to obtain features with a dimension of 8×256 to capture significant spatial structure information; apply spatial average pooling to obtain features with a dimension of 8×256 to retain the overall spatial information; S25. Temporal feature fusion: Send the results of max pooling and average pooling into a bidirectional long short-term memory network layer; S3: Multi-task learning and optimization: S31. Shared Representation: The feature vector output by step S2 is used as the shared representation , and is input into subsequent task-specific predictors, where is the input sample, is the shared encoder, is the encoded feature vector; S32. Task-Specific Predictor: Build an independent predictor for each evaluation task ; S33. Application of the multi-task learning framework: Adopt a hard parameter sharing structure to share the parameters of the encoder ; Adopt a Frank-Wolfe-based method, regard the problem as multi-objective optimization, and iteratively update the shared parameters by solving sub-problems to find the Pareto optimal solution; S34. Model training and optimization: Loss function: For classification tasks, the cross-entropy loss function is used as ; The total loss is the sum of the losses of each task; , Wherein: is the cross-entropy loss for the k-th task; is the true label of class c, i.e., the value in the one-hot encoding, 0 or 1; log is the natural logarithm function; is the probability of class c predicted by the model; Optimizer: The AdaTask optimization method is adopted to update the model parameters ; Training process: Use the labeled embryo dataset for model training, and iteratively update the network parameters through backpropagation and an optimizer until the model converges or reaches the preset number of training epochs; S4: Classification prediction and embryo selection: S41. Feature input: Input the shared feature vector output by the encoding module into each task-specific predictor in the classification module; S42. Processing by a multi-layer perceptron: First hidden layer: 128 neurons, ReLU activation function, followed by Dropout(0.5), Second hidden layer: 64 neurons, ReLU activation function, followed by Dropout(0.5), S43. Output prediction: Istanbul Consensus classification task: The output layer has 5 neurons, apply the Softmax activation function, and output the probability distribution corresponding to 5 quality grades; Gardner grading task: The output layer has 6 neurons, apply the Softmax activation function, and output the probability distribution corresponding to 6 development grades; S44. Embryo selection decision: According to the multiple prediction results output by the model, combined with clinical experience and specific requirements, formulate an embryo selection strategy.

2. The method according to claim 1, wherein in step S11, each embryo sample contains consecutive image frames covering key development stages; each sample contains 32 grayscale images, and the resolution of each frame is 256×256 pixels; in step S12, the preprocessing includes image alignment, brightness normalization, and denoising operations.

3. The method according to claim 1, wherein in S22, the ResNet3D network includes three 3D residual blocks, each residual block contains two 3D convolutional layers, and the convolutional kernel size is 13×3×3; after being processed by multiple residual blocks, the dimension of the feature map gradually decreases, and the final output feature map dimension is: 8×32×32×256; In S23, the dimension 8×32×32×256 is kept unchanged to prevent overfitting; In S24, applying spatial max pooling and applying spatial average pooling both result in features with a dimension of 8×256; In S25, each direction of the bidirectional long short-term memory network layer contains 128 units, processes the information in the time dimension, and captures the temporal dependencies of embryonic development; the bidirectional long short-term memory network layer finally outputs a 256-dimensional feature vector that fuses spatio-temporal information.

4. The method according to claim 1, wherein In S32, the predictor adopts a multi-layer perceptron structure, including: Input layer: Receives the feature representation from the shared encoder; Hidden layer: Contains multiple fully connected layers, each followed by a ReLU activation function and Dropout regularization for non-linear transformation of features and preventing overfitting; Output layer: Uses the Softmax activation function according to the specific task and outputs the probability distribution of the corresponding level: , Wherein: is the probability that the input x belongs to class c, is the original score (logits) of class c, is the total number of classes for task k, and exp is the natural exponential function; The tasks include Istanbul Consensus Grading and Gardner Grading.

5. The method according to claim 1, wherein In S33, the hard parameter sharing structure: includes a shared encoder and task-specific predictors , and the objective function is the sum of the losses of all tasks: , Wherein: is the set of all model parameters, are the parameters of the shared encoder, are the parameters of the k-th task-specific predictor, is the shared encoder, is the k-th task-specific predictor, is the i-th input sample, is the true label of the i-th sample on the k-th task, is the loss function of the k-th task, is the number of samples, K is the number of tasks, represents function composition.

6. The method according to claim 1, wherein In S33, a virtual predictor is introduced , whose structure is the same as that of the real predictor but with randomly initialized and fixed parameters; a penalty term for the gradient norm of the virtual predictor is added to the training objective to improve the generality of the shared representation Among them, the loss function of the single task is to add a penalty term of the virtual predictor gradient norm to the original loss function: , The specific form is: , Wherein: is the predicted value of the k-th task for the i-th sample, is the predicted value of the virtual predictor for the i-th sample, is the shared representation of the i-th sample, is the penalty term coefficient, used to balance the task loss and the virtual gradient norm; The total loss function of the multi-task is the sum of the losses of all tasks: , Among them, the loss of each task consists of two parts: True predictor loss : The traditional task loss generated by the true predictor to generate predictions and calculated; Virtual predictor gradient norm penalty: Generate predictions through the virtual predictor and calculate the gradient norm of its loss with respect to its own parameters for enhancing the generality of the shared representation; ​ Wherein: is the total loss of the k-th task, is the number of samples, is the true predictor loss function of the k-th task, is the true label of the i-th sample of the k-th task, is the predicted value of the true predictor for the i-th sample of the k-th task, is the virtual predictor loss function of the k-th task, is the predicted value of the virtual predictor for the i-th sample of the k-th task, is the parameter of the virtual predictor of the k-th task, is the gradient of the virtual predictor parameter, is the Frobenius norm, is the penalty term coefficient.

7. The method according to claim 6, wherein In step S34, the classification tasks include Istanbul Consensus and Gardner Grading; the total loss contains a penalty term of the virtual predictor gradient norm; The AdaTask optimization method is: Initialization: Set the exponential decay factor v, the initial learning rate , the smoothing factor ε, and initialize the cumulative gradient variable for each task k ; Gradient calculation: At training step t, compute the gradient of the parameter θ for each task ; Cumulative gradient update: For each task, update the cumulative gradient by exponential decay averaging: , Parameter update amount calculation: Calculate the parameter update amount for each task: , Model parameter update: Update the model parameters according to the parameter update amounts of all tasks: , Where: t is the training step, k is the task index, i is the parameter index, is the cumulative gradient of task k for parameter i at step t, is the gradient of task k for parameter i at step, v is the exponential decay factor, is the initial learning rate, ε is the smoothing factor, is the update amount of task k for parameter i, is the value of parameter i at step t.

8. The method according to claim 1, wherein In S44, the final prediction result is obtained by taking the category with the highest probability: , Wherein: is the predicted category of the k-th task, and argmax represents taking the category c that maximizes the probability, is the predicted probability that the input x belongs to the category c.

Citation Information

Patent Citations

  • Grade classification method for evaluating in-vitro fertilization treatment embryo based on cleavage behaviors

    CN104718298A

  • Embryo development prediction and morphology preferential selection system based on artificial intelligence technology

    CN117274706A

  • Embryo growth stage prediction method and system based on self-supervised learning

    CN119942249A

  • Devices and processes for machine learning prediction of in-vitro fertilization

    EP3951793A1

  • System and method for outcome evaluations on human IVF-derived embryos

    US20240185567A1

Cited By

  • Gradient guiding differential evolution method and system based on layer optimization difference

    CN121052320A