Embryo development quality evaluation method and system based on text image contrast learning
The text-image contrastive learning method addresses the lack of specificity in embryo quality assessment by integrating embryology knowledge into feature extraction, enhancing accuracy and consistency through a CLIP model alignment, thus improving embryo quality prediction.
Patent Information
- Application Number
- CN202510404135.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
The existing image processing methods are not targeted when processing embryo image features, resulting in insufficient accuracy and consistency of embryo quality assessment and lack of expertise in effective fusion embryology and reproductive medicine.
Using a method based on text image comparison learning, an embryo image feature description set is constructed, and the embryo image is combined with text feature description through a prediction model. The CLIP model is used for pre-training and fine-tuning, and the loss function is designed to maximize the matching similarity and minimize mismatch similarity to achieve the evaluation of embryo development quality.
It improves the professionalism and accuracy of embryonic development quality assessment, reduces the dependence of manual annotation, enhances the model's understanding and processing ability of embryonic development characteristics, and improves the consistency and accuracy of the assessment.
Smart Images

Figure CN120318185A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an embryo development quality assessment method and system based on text-image contrast learning. Background Art
[0002] In the field of assisted reproductive technology (ART), accurately assessing the quality and developmental potential of embryos is crucial for improving the success rate of in vitro fertilization (IVF). Methods for predicting embryo development quality not only assist in selecting the most suitable embryos for transplantation but also have important implications for studying key events in the early stages of embryo development, diagnosing, and treating developmental defects.
[0003] In ART, traditional embryo assessment mainly relies on morphological scoring, which is a method of evaluating the quality of embryos by assessing the morphological characteristics of embryos at different developmental stages. This method includes the assessment of stages such as the pronuclear stage, day 3 embryos, and blastocysts, and is currently the most common means of embryo quality assessment. It relies on artificial observation of the morphology and developmental changes of embryos under a microscope and scores accordingly. This method is rapid, non-invasive, and easy to operate. Although the traditional morphological scoring method is widely used, it has some significant problems. First, due to relying on subjective judgment, its accuracy may be limited. Second, the scoring criteria have not been unified, lacking precise quantitative indicators, resulting in a lack of repeatability of the results. Modern embryo assessment usually relies on artificial intelligence technology. These technologies reduce human variables through automated means, improving the consistency and accuracy of assessment.
[0004] Although artificial intelligence technology shows potential in embryo assessment, there are still some challenges. Existing technologies have failed to effectively integrate professional knowledge in fields such as embryology and reproductive medicine into the feature extraction process, resulting in the model being unable to fully utilize this knowledge for accurate feature recognition. Therefore, existing image processing methods may have insufficient pertinence when dealing with the features of embryo images. Summary of the Invention
[0005] The present invention proposes an embryo development quality assessment method and system based on text-image contrast learning, which solves the problem of insufficient pertinence of existing image processing methods when dealing with the features of embryo images.
[0006] To solve the above technical problems, the present invention provides an embryo development quality assessment method based on text-image contrast learning, including the following steps: Step S1: Construct an embryo image feature description set according to embryo development characteristics, and use the embryo image feature description set to label multiple continuously captured embryo images to obtain a training set; Step S2: Construct a prediction model including an image feature extraction module and a text feature extraction module, design a prompt template for the embryo development quality assessment task, and the text feature extraction module uses the prompt template to convert the descriptions in the embryo image feature description set into natural language descriptions, and the prediction model outputs the similarity between each embryo image in the training set and all feature descriptions; Step S3: Design a loss function with the goal of maximizing the similarity between matching image-text pairs and minimizing the similarity between non-matching image-text pairs, and update the parameters of the prediction model according to the loss function; Step S4: Input the embryo image to be evaluated into the updated prediction model to obtain the evaluation result of the embryo development quality.
[0007] Preferably, the embryo image feature description set in step S1 includes: pronucleus cells, uniform blastomeres, degree of cytoplasmic fragmentation, and trophoblast cells.
[0008] Preferably, in step S2, before inputting the training set and the embryo image feature description set into the prediction model, it further includes pre-training the prediction model, including the following steps: Step S201: Collect a large-scale image-text data set, and the image-text data set includes images in various fields and the corresponding description texts for each field; Step S202: For each image, use the text description that matches the image as a positive sample and the non-matching text description as a negative sample; Step S203: Input the image-text data set into the prediction model, and the image feature extraction module and the text feature extraction module respectively extract the feature vectors of the image and the text; Step S204: Calculate the similarity between the image feature vector and the text feature vector, calculate the contrast loss according to the similarity, and use the gradient descent method to update the parameters of the prediction model.
[0009] Preferably, before inputting the training set into the prediction model in step S2, preprocess the training set, including the following steps: add a new dimension, stack all the embryo images in the training set along the newly added dimension, and convert the training set into an embryo image sequence.
[0010] Preferably, the image feature extraction module in step S2 is used to extract the image features in the training set, including the following steps: Step S201: Divide each embryo image in the embryo image sequence into several image blocks of equal size, and expand each image block into a one-dimensional vector; Step S202: Convert the vector into an embedded vector of a fixed dimension through linear transformation, and add position encoding to each of the embedded vectors; Step S203: Input the sequence of image patches processed by embedding and position encoding into multiple stacked Transformer encoder layers, calculate the attention scores between all image patches through the multi-head self-attention mechanism, and input the output of the self-attention layer into the feed-forward network to extract the features in all image patches; Step S204: Use the attention scores to perform weighted summation on all image patch features to obtain a global feature vector as the image feature.
[0011] Preferably, the text feature extraction module in step S2 is used to extract the text features in the embryo image feature description set, including the following steps: Step S211: Segment the text in the embryo image feature description set into several words, and use the hint module to convert the words into natural language descriptions; Step S212: Convert the natural language description into a text feature vector through word embedding, and add position encoding to each text feature vector; Step S213: Input the text feature vectors with position encoding into multiple stacked Transformer encoder layers, calculate the attention scores between all text feature vectors through the multi-head self-attention mechanism, and input the output of the self-attention layer into the feed-forward network to extract the features of all text feature vectors; Step S214: Use the attention scores to perform weighted summation on all text feature vectors to obtain a global feature vector as the text feature.
[0012] Preferably, the prediction model in step S2 outputs the similarity between each embryo image in the training set and all feature descriptions, including the following steps: Step S221: Standardize the image features extracted by the image feature extraction module and the text features extracted by the text feature extraction module; Step S222: Calculate the similarity between the standardized image features and text features : ; where, is the image feature; is the text feature; represents the dot product of vectors; represents the norm of the vector; Step S223: Calculate the similarity between each image feature and all text features to form a similarity matrix. Each row of the similarity matrix represents an embryo image, and each column represents a feature description; Step S224: Perform normalization processing on each row of the similarity matrix, and output the similarity between each embryo image and all feature descriptions according to the normalized similarity matrix.
[0013] Preferably, the expression of the loss function in step S3 is: ; In the formula, is the loss function; is the mean absolute error MAE; is the mean square error MSE; is the image feature; is the text feature; is the adjustable parameter.
[0014] The present invention also provides an embryo development quality evaluation system based on text-image contrast learning, which is implemented based on the above-mentioned embryo development quality evaluation method based on text-image contrast learning, and includes: a model pre-training module, a data collection module, a feature extraction module, and a contrast learning module; The model pre-training module: Use a large-scale image-text data set including several fields to pre-train the prediction model; The data collection module: Collect an embryo image set and an embryo image feature description set for training and testing the model, and label the embryo image set based on the embryo image feature description set; The feature extraction module: Extract the image feature vectors in the embryo image set and the text feature vectors in the embryo image feature description set; The contrast learning module: Use the matching image-text pairs as positive samples and the non-matching image-text pairs as negative samples, calculate the contrast loss between the output result of the prediction model and the true label corresponding to the embryo image, and improve the discrimination ability of the prediction model by optimizing the contrast loss.
[0015] Preferably, the contrast learning module calculates the gradient of the parameters of the prediction model through the backpropagation algorithm, and updates the parameters of the prediction model using gradient descent.
[0016] The beneficial effects of the present invention at least include: 1. Construct a feature description set according to the embryo development characteristics to ensure that the annotation of the training set is closely related to the key characteristics of embryo development, thereby improving the professionalism and accuracy of the model in evaluating the embryo development quality; 2. The prompt template designed for the embryo development quality assessment task enables the model to better understand and process professional terms and descriptions related to embryo development, improving the model's ability to understand and process the language of embryo development characteristics. 3. By maximizing the similarity between the matching image-text pairs and minimizing the similarity between the non-matching pairs, it helps the prediction model to more accurately distinguish the embryo images from their characteristic descriptions related to the development quality, thereby improving the accuracy of the prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention; Figure 2 It is a schematic diagram of the embryo image collected according to an embodiment of the present invention; Figure 3 It is a schematic flowchart of the image feature extraction module extracting image features according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the structure of the decoder according to an embodiment of the present invention; Figure 5 It is a schematic diagram of training the prediction model by comparing similarities according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Next, with reference to the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0019] As Figure 1 shown, an embodiment of the present invention provides an embryo development quality assessment method based on text-image contrast learning, including the following steps: Step S1: Construct an embryo image feature description set according to embryo development characteristics, and use the embryo image feature description set to label multiple continuously captured embryo images to obtain a training set.
[0020] Specifically, the embodiments of the present invention collected 100,000 embryo images from multiple reproductive centers. These images were taken in a time-lapse incubator and covered all developmental stages of the embryo from the single-cell stage to the blastocyst stage, and also included empty dish control images. For example Figure 2 is a schematic diagram of the development process of a certain embryo cell. At the pronuclear stage, the number of embryo cells is two and their sizes are basically the same; at the cleavage stage, the number of blastomeres of the embryo cell is 7, which meets the standard range of 6-8, and the uniformity of the blastomere sizes is good.
[0021] However, due to the limited resources of embryologists with professional knowledge and experience, manually analyzing embryo images to predict whether the development is good is extremely time-consuming and laborious. This is not only a time-consuming and labor-intensive task, but also means that they cannot engage in other more valuable professional activities during this period. Therefore, there is an urgent need for a more efficient method to replace this mechanical image prediction work. Specifically, there are the following challenges in predicting embryo development through embryo images: 1. During the process of acquiring and processing embryo images, image quality is crucial. If the image is blurred or the embryo boundary is not clear, it will directly affect the accuracy of subsequent analysis and classification. And foreign objects may obscure the embryo or interfere with the detection of the embryo's position, further increasing the prediction difficulty.
[0022] 2. In the early stages of embryo development, due to the short development time and rapid changes, accurately identifying these stages is very challenging. Research shows that the identification of early embryo development stages is one of the key difficulties in predicting embryo development.
[0023] 3. Traditional embryo assessment methods rely on morphological scoring, which is highly dependent on experience and subjective. Different assessors may give different scores to the same embryo, resulting in a lack of reproducibility of the results. Although the application of artificial intelligence technology can reduce human variables and improve the consistency and accuracy of assessment, a large amount of training data is still required to ensure the accuracy and robustness of the algorithm.
[0024] Considering the above problems, the embodiment of the present invention adopts the multi-modal pre-training model CLIP as the prediction model. Through pre-training on a large-scale dataset, this model can learn effective feature representations in an unsupervised environment. The CLIP model uses contrastive learning technology to process image and text data simultaneously. During the training process, the model receives a pair of matching image-text pairs as positive samples and non-matching image-text pairs as negative samples. The model contains two independent encoders, one for processing images and the other for processing text, which convert them into high-dimensional feature vectors, and evaluate the correlation between images and text by calculating the similarity between these vectors. The CLIP model effectively solves the image annotation problem through contrastive learning and multi-modal pre-training, reduces the dependence on manually annotated data, and demonstrates excellent performance in various visual tasks. Its multi-modal contrastive learning method and powerful zero-shot learning ability enable the model to automatically generate accurate labels or descriptions for images, effectively reducing the workload of manual annotation.
[0025] In the process of selecting embryo image feature descriptions in the embodiments of the present invention, extensive literature research, professional consultations were carried out, and expert opinions were solicited to construct a text description set for embryo image features. By inviting senior experts in the field of reproductive medicine to participate in the evaluation work, and using the professional knowledge and experience of these experts, the professionalism and advancement of each text feature description were scored. According to the evaluation opinions of these experts, image tags including "pronucleus stage", "even blastomeres", and "degree of cytoplasmic fragmentation" were determined. For example Figure 2 the embryo cells in can be described as "the number of pronuclei is 2, the number of blastomeres is 6 - 8, the division is even, and the degree of cytoplasmic fragmentation < 15%". After constructing the embryo image feature description set, the collected embryo images are labeled using the embryo image feature description set to obtain a training set. The brightness, contrast, and saturation of all embryo images in the training set are adjusted, and all embryo images are scaled to a unified size and format to improve the generalization ability and accuracy of the model. This method not only ensures that the selected images meet the requirements of the best focal length at the technical level, but also is biologically representative and can accurately reflect the true developmental status of the embryo. Through this strict image screening and labeling process, a high-quality and representative embryo image data set is successfully constructed, laying a solid foundation for the pre-training of the CLIP model and providing a reliable reference standard for the evaluation and verification of the model prediction results.
[0026] Step S2: Construct a prediction model, design a prompt template for the embryo development quality assessment task, use the prompt template to convert the descriptions in the embryo image feature description set into natural language descriptions, and the prediction model outputs the similarity between each embryo image in the training set and all feature descriptions.
[0027] Specifically, before applying the CLIP model to embryo images, it is first pre-trained on a large-scale image-text data set, and then fine-tuned on a specific embryo image data set. Using the semantic understanding ability of the CLIP model, the recognition and classification performance of the model for embryo images is improved. In the pre-training stage, through the batch normalization layer and activation function, the local features of the image are extracted and enhanced. At the same time, in order to adapt to the problems of possible edge blurring and image distortion in embryo images, a lightweight multi-level visual feature adaptation and comparison framework is designed. This framework integrates multiple residual adapters into the pre-trained visual encoder and adopts a multi-level, pixel-level visual language feature alignment loss function to gradually enhance the visual features at different levels.
[0028] The CLIP model consists of an encoder and a decoder, where the encoder is further subdivided into a text encoder and an image encoder. The image encoder is used to process the input training set of embryo images. By adding a new dimension and stacking multiple embryo image data on this dimension, a multi-dimensional tensor containing all visual information is formed. This design not only makes the data structure compact but also facilitates batch training of the neural network model, speeds up the iteration of model weights, and improves learning efficiency. In addition, this multi-dimensional tensor simplifies the data processing flow, improves the speed and stability of model training, and provides a solid foundation for the training and prediction of the prediction model. By reducing the time for data conversion and loading, the CLIP model can invest more resources and time in model optimization and feature engineering, thereby enhancing the overall performance of the model.
[0029] The image encoder uses a Vision Transformer (ViT) as the backbone. Since embryo images may contain details of different sizes and resolutions, the CLIP model adopts a multi-scale feature extraction method to capture the details of embryo images at different scales, improving the understanding and generalization ability of the prediction model for images. As Figure 3 shown, the image encoder first divides the input image sequence into image patches of a fixed size, flattens these image patches into one-dimensional vectors, and converts them into embedding vectors of a fixed dimension through a linear transformation. Position encoding is added to all embedding vectors. The addition of position encoding preserves the spatial position information of the image patches, enabling the CLIP model to understand the relative positions of different parts in the image. The sequence of image patches processed by the multi-head self-attention mechanism and the feed-forward network is fed into the Transformer encoder layer to enhance the feature learning ability of the prediction model. Layer normalization and residual connections are used to stabilize the training process and improve the convergence of the model. A special classification token is added to all image patch sequences for the final classification task. After being processed by the Transformer encoder, the output of this token will be fed into a fully connected layer for classification. Through the above steps, ViT can effectively process image data and capture global features.
[0030] In the process of text feature extraction, first, the descriptions in the embryo image feature description set are converted into phrases through segmentation and augmentation operations. For example, "whether the blastomeres are uniform" is converted into "this blastomere is uniform". Then, the phrases are converted into natural language descriptions using a prompt template. For example, the phrase "this blastomere is uniform" is converted into "this blastomere has a uniform characteristic". Finally, the natural language descriptions are output to the text encoder to obtain tokens.
[0031] The text encoder converts tokens into word embeddings and also adds positional encodings to the word embeddings. The introduction of positional encodings enables the Transformer model to handle sequence order. The text encoder also includes a multi-head self-attention mechanism, a feed-forward network, layer normalization, and residual connections, which convert the input sequence into a high-dimensional vector for text classification.
[0032] Since the CLIP model sees complete sentences during pre-training, using phrases during inference will affect the model's prediction performance. Therefore, the embodiments of the present invention adopt a prompt template and optimize the prompt template according to the characteristics of embryonic development to guide the model to process embryo cell-related data and information more accurately, such as "pronucleus cell", "blastomere", "trophoblast cell", etc., and the weights of these words can also be increased by highlighting keywords.
[0033] The pre-trained text encoder is fine-tuned using the text encoder fine-tuning technique TextCraftor. The embryo detection process requires high-quality images to reveal the structure and characteristics of cells. TextCraftor enhances the pre-trained text encoder through a reward function, improving the semantic matching degree during the model prediction process by improving the image quality and the accuracy of text-image alignment, which is crucial for accurately identifying and classifying different types of cells in embryo cell detection.
[0034] In the CLIP model, there is no concept of a traditional decoder because the model is unconditional and does not need to predict the next word or sequence. During the test phase, after inputting the text description, the CLIP model directly calculates the vector related to the embryo image for this text description and finds the image that best matches this text description through similarity comparison. Figure 4As shown, the decoder takes image Tokens and data Tokens as the input of the Mamba Block. The Mamba Block removes the translational part in the calculation process through Root Mean Square Layer Normalization (RMS Norm) and only retains the scaling part. This simplification not only improves the training speed but also maintains the performance of the model. The input embedding is obtained through linear projection, and then the input embedding is convolved to prevent independent token calculation. The SiLU activation function enhances the model's ability to capture complex patterns and relationships in sequence data through its non-linearity and smoothness, improving the computational efficiency and model performance. The selective SSM framework including Spring, Spring MVC, and MyBatis is adopted, allowing the model to dynamically adjust its state propagation mechanism according to the input, enabling the model to filter out relevant information, thereby reducing unnecessary computational load. The output of the model is converted into a probability distribution through the activation functions linear and softmax, making the output more interpretable and more suitable for classification tasks. Through these steps, the CLIP model can effectively evaluate the correlation between images and texts, providing a powerful multi-modal learning ability for embryo image development prediction.
[0035] Step S3: Design a loss function with the goal of maximizing the similarity between the matched image-text pairs and minimizing the similarity between the unmatched image-text pairs, and update the parameters of the prediction model according to the loss function.
[0036] Specifically, in the embodiments of the present invention, a contrastive loss function is adopted to measure the similarity between the matched image-text pairs and maximize these similarities as much as possible, while minimizing the similarity between the unmatched image-text pairs.
[0037] The basic idea of the contrastive loss function is achieved by optimizing the inner product between two vectors. Specifically, the CLIP model hopes that the larger the inner product between the matched image and text, the more similar they are; conversely, the smaller the inner product between the unmatched image and text, the less similar they are. As Figure 5 shown in the training process of the model, the CLIP model is trained to predict N × N similarities of possible pairs, where N is the number of image and text pairs in the batch. The contrastive loss function is defined as: ; In the formula, is the loss function; is the Mean Absolute Error (MAE); is the Mean Squared Error (MSE); is the image feature; is the text feature.
[0038] The loss (Mean Absolute Error, MAE) and the loss (Mean Squared Error, MSE) are two common regression loss functions. The loss calculates the absolute difference between the predicted value and the true value, while the loss calculates the squared difference between the predicted value and the true value. In the embodiments of the present invention, an adjustable parameter is introduced on the basis of comparing the loss functions to make the model adapt to different data characteristics and error distributions. The comparative loss function in the embodiments of the present invention is defined as: .
[0039] Compared with the conventional comparative loss function, the introduction of the adjustable parameter has the following advantages: 1. If is larger, the loss will be more sensitive to larger errors. If is smaller, the loss will be smoother. Therefore, it helps the model to be more flexible in dealing with errors of different scales; 2. The error distributions of different data sets may be different. Through the adjustable parameter , the model can better adapt to these different data distributions, thereby improving the performance of the model.
[0040] In actual operation, the optimal weight ratio can be determined through experiments. For example, or the weight of the loss function can be gradually increased or decreased, and the performance of the model on the validation set can be observed. This method can help find a weight configuration that can maintain the model performance without being overly complicated.
[0041] Step S4: Input the embryo image to be evaluated into the updated prediction model to obtain the evaluation result of the embryo development quality.
[0042] In the embodiment of the present invention, a mixed-precision training method is adopted. Set the learning rate of the CLIP model to 3e-4, the batch size to 16, the number of training epochs to 300, and select the appropriate optimizer AdamW. Use mixed-precision training (use_amp=True) to improve the training efficiency, and enable multi-card training (use_dp=True) to accelerate the training process. Set gradient clipping (CLIP_GRAD=5.0) to prevent gradient explosion, and record the highest score (Best_ACC) to save the best model during the training process. During the training process, adopt the discriminative fine-tuning (DFT) method to dynamically allocate the learning rate to optimize the loss function and achieve fast convergence. Use the Pytorch framework to implement the model of the present invention, and all models are trained and tested on the NVIDIA GeForce GTX1660 Ti GPU.
[0043] To verify the prediction performance of the CLIP model in the embodiment of the present invention, use the confusion matrix to evaluate the performance of the model, and calculate the sensitivity and specificity of the model respectively. The formula for sensitivity is , and the formula for specificity is , is the true positive class, that is, the true class of the sample is the positive class, and the result recognized by the model is also the positive class; is the false negative class, that is, the true class of the sample is the positive class, but the model recognizes it as the negative class; is the true negative class, that is, the true class of the sample is the negative class, and the model recognizes it as the negative class; is the false positive class, that is, the true class of the sample is the negative class, but the model recognizes it as the positive class. A high sensitivity means that the model can better recognize all positive samples, and a high specificity means that the model can better distinguish negative samples for calculation.
[0044] Calculate the false positive rate (FPR) and false negative rate (FNR) to quantify the error prediction ability of the model. The false positive rate is the proportion of cases where negative examples are recognized as positive examples among all negative examples, and its calculation formula is , and the false negative rate is the proportion of cases where positive examples are recognized as negative examples among all positive examples, and its calculation formula is 。Mixed-precision training significantly accelerates the training cycle and overall training time when training the embryo image processing model by reducing memory requirements and allowing for larger batch sizes. Additionally, the discriminative fine-tuning method achieves fast convergence and high accuracy of the model by dynamically allocating the learning rate. The experimental results show that the model of the embodiment of the present invention reaches the state-of-the-art performance in 50 training cycles, and the performance of the model can be evaluated through the confusion matrix and sensitivity and specificity metrics.
[0045] The embodiment of the present invention also provides an embryo development quality assessment system based on text-image contrast learning, which is implemented based on the above-mentioned method for embryo development quality assessment based on text-image contrast learning, and includes: a model pre-training module, a data collection module, a feature extraction module, and a contrast learning module.
[0046] The model pre-training module pre-trains the prediction model using a large-scale image-text dataset including several domains.
[0047] Specifically, a network structure including a CLIP model is constructed. The CLIP model provides a joint feature representation of images and texts, while the Mamba prediction model is used to predict the development degree of embryos. The ResNet network is used as the basic network, and the Faster R-CNN network framework is combined to predict and locate the number of blastomeres. And a large-scale image-text dataset including several domains is used to pre-train the CLIP model.
[0048] The data collection module collects an embryo image set and an embryo image feature description set for training and testing the model, and annotates the embryo image set based on the embryo image feature description set.
[0049] Specifically, the data collection module first obtains embryo images at different developmental stages from a medical center and records the shooting time points for training and testing the model. The embryo images are pre-processed, including root mean square layer normalization operation, to reduce the influence of light source fluctuations on the images. The collected embryo images are divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0050] The feature extraction module: extracts the image feature vectors in the embryo image set and the text feature vectors in the embryo image feature description set.
[0051] The contrast learning module: uses the matching image-text pairs as positive samples and the non-matching image-text pairs as negative samples, calculates the contrast loss between the output result of the prediction model and the true label corresponding to the embryo image, and improves the discrimination ability of the prediction model by optimizing the contrast loss.
[0052] Specifically, the CLIP model is trained using the training set to obtain model parameters. The CLIP model is evaluated using the validation set, and the model parameters are adjusted to optimize performance. The trained CLIP model is used to predict the embryo image sequence of the test set, and the prediction results are statistically analyzed. The prediction results are compared with the actual development results of the embryo. If they are all normal or abnormal development, the prediction is considered to be accurate and successful, otherwise it is considered to be a failure and the prediction is inaccurate. The judgment results of multiple groups of embryo image sequences are output as a test report. If the prediction accuracy does not reach the predetermined target, return to the model training step to further optimize the model.
[0053] Binary cross entropy (BCE) is used as the loss function to measure the difference between the probability value predicted by the model and the actual label to evaluate the performance of the model.
[0054] ; In the formula, is the binary cross entropy; is the true label of the embryo image sample, with a value of 0 or 1. A value of 0 indicates a negative sample, and a value of 1 indicates a positive sample; The probability value of whether the embryo can develop well predicted by the CLIP model ranges from (0 to 1). The closer the probability value is to 1, the better the quality of embryo development. Conversely, the closer the probability value is to 0, the worse the quality of embryo development.
[0055] The present invention explores the different prediction accuracies of pre-trained image sets of different sizes. The experimental results show that when the image set only takes the first 100 images of the embryonic cell development process, the model prediction accuracy is 75.5%; when the image set only takes the first 200 images of the embryonic cell development process, the model prediction accuracy is 83.7%; when the image set only takes the first 300 images of the embryonic cell development process, the model prediction accuracy can reach the highest 88.1%, which is 12.6% higher than the result of using only the first 100 embryo images. This shows that when as many images of the development process as possible are selected, the multimodal CLIP model of the present invention can obtain more excellent performance, further illustrating that the design of the present invention can more accurately predict the results of whether embryonic cells can develop successfully through pre-training.
[0056] Table 1 shows the accuracy of different models in the task of embryonic image development prediction. From the experimental results in Table 1, it is found that for the ResNet model, ResNet101 can achieve the highest accuracy of 84.7%, which is 3.4% lower than the method proposed in the present invention; the highest accuracy of the present invention is 1.6% higher than the result of the BERT model in the task of embryonic image development prediction, which shows that the method proposed in the present invention has a higher recognition effect among the currently popular methods.
[0057] Table 1 Accuracy of Different Models in Predicting Embryo Image Development
[0058] In summary, the present invention makes full use of the image and text features contained in embryo images, adopts a text-image pre-trained contrastive learning model, combines the text features of embryo images for modeling, and studies the task of predicting embryo image development. A large number of experiments show that the method proposed by the present invention has more excellent performance.
[0059] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present invention. As long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0060] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. An embryo development quality assessment method based on text-image contrast learning, characterized in that, It includes the following steps: Step S1: Construct an embryonic image feature description set according to embryonic development characteristics, and use the embryonic image feature description set to label multiple continuously captured embryonic images to obtain a training set; Step S2: Construct a prediction model including an image feature extraction module and a text feature extraction module, design a prompt template for the embryonic development quality assessment task, and the text feature extraction module uses the prompt template to convert the descriptions in the embryonic image feature description set into natural language descriptions, and the prediction model outputs the similarity between each embryonic image in the training set and all feature descriptions; Step S3: Design a loss function with the goal of maximizing the similarity between matching image-text pairs and minimizing the similarity between non-matching image-text pairs, and update the parameters of the prediction model according to the loss function; Step S4: Input the embryonic image to be evaluated into the updated prediction model to obtain the evaluation result of the embryonic development quality.
2. The method for evaluating embryo development quality based on text-image contrast learning according to claim 1, wherein: The embryonic image feature description set in Step S1 includes: pronucleus cells, uniform blastomeres, degree of cytoplasmic fragments, and trophoblast cells.
3. The embryo development quality assessment method based on text-image contrast learning according to claim 1, characterized in that: In Step S2, before inputting the training set and the embryonic image feature description set into the prediction model, it also includes pre-training the prediction model, including the following steps: Step S201: Collect a large-scale image-text data set, and the image-text data set includes images in various fields and the corresponding description texts for each field; Step S202: For each image, use the text description that matches the image as a positive sample and the text description that does not match as a negative sample; Step S203: Input the image-text data set into the prediction model, and the image feature extraction module and the text feature extraction module respectively extract the feature vectors of the image and the text; Step S204: Calculate the similarity between the image feature vector and the text feature vector, calculate the contrast loss according to the similarity, and use the gradient descent method to update the parameters of the prediction model.
4. A method for evaluating the quality of embryonic development based on text-image contrast learning according to claim 1, characterized in that: In Step S2, preprocess the training set before inputting the training set into the prediction model, including the following steps: add a new dimension, stack all the embryonic images in the training set along the newly added dimension, and convert the training set into an embryonic image sequence.
5. The embryo development quality assessment method based on text-image contrast learning according to claim 4, wherein: The image feature extraction module in Step S2 is used to extract the image features in the training set, including the following steps: Step S201: Divide each embryonic image in the embryonic image sequence into several equally sized image patches, and expand each image patch into a one-dimensional vector; Step S202: Convert the one-dimensional vector into an embedding vector with a fixed dimension through a linear transformation, and add a position encoding to each embedding vector; Step S203: Input the sequence of image patches after embedding and position encoding processing into multiple stacked Transformer encoder layers, calculate the attention scores between all image patches through the multi-head self-attention mechanism, and input the output of the self-attention layer into the feed-forward network to extract the features in all image patches; Step S204: Use the attention scores to perform weighted summation on all image patch features to obtain a global feature vector as the image feature.
6. The embryonic development quality evaluation method based on text-image contrast learning according to claim 1, characterized in that: In step S2, the text feature extraction module is used to extract the text features in the embryo image feature description set, including the following steps: Step S211: Segment the text in the embryo image feature description set into several words, and use the hint module to convert the words into natural language descriptions; Step S212: Convert the natural language descriptions into text feature vectors through word embedding, and add position encoding to each text feature vector; Step S213: Input the text feature vectors with position encoding into multiple stacked Transformer encoder layers, calculate the attention scores between all text feature vectors through the multi-head self-attention mechanism, and input the output of the self-attention layer into the feed-forward network to extract the features of all text feature vectors; Step S214: Use the attention scores to perform weighted summation on all text feature vectors to obtain a global feature vector as the text feature.
7. The embryo development quality assessment method based on text-image contrast learning according to claim 1, wherein: In step S2, the prediction model outputs the similarity between each embryo image in the training set and all feature descriptions, including the following steps: Step S221: Normalize the image features extracted by the image feature extraction module and the text features extracted by the text feature extraction module; Step S222: Calculate the similarity between the standardized image features and text features : ; In the formula, is the image feature; is the text feature; represents the dot product of vectors; represents the norm of a vector; Step S223: Calculate the similarity between each image feature and all text features to form a similarity matrix, where each row of the similarity matrix represents an embryo image and each column represents a feature description; Step S224: Perform normalization processing on each row of the similarity matrix, and output the similarity between each embryo image and all feature descriptions according to the normalized similarity matrix.
8. A method for evaluating the quality of embryo development based on contrastive learning of text images according to claim 1, characterized in that: The expression of the loss function in step S3 is: ; Wherein, is the loss function; is the mean absolute error MAE; is the mean square error MSE; is the image feature; is the text feature; is the adjustable parameter.
9. An embryo development quality assessment system based on text-image contrast learning, which is implemented based on an embryo development quality assessment method based on text-image contrast learning according to any one of claims 1 to 8, characterized in that, Including: A model pre-training module, a data collection module, a feature extraction module, and a contrast learning module; The model pre-training module: Use a large-scale image-text dataset including several domains to pre-train the prediction model; The data collection module: Collect an embryo image set and an embryo image feature description set for training and testing the model, and label the embryo image set based on the embryo image feature description set; The feature extraction module: Extract the image feature vectors in the embryo image set and the text feature vectors in the embryo image feature description set; The contrast learning module: Use the matching image-text pairs as positive samples and the non-matching image-text pairs as negative samples, calculate the contrast loss between the output result of the prediction model and the true label corresponding to the embryo image, and improve the discrimination ability of the prediction model by optimizing the contrast loss.
10. The embryo development quality assessment system based on text-image contrast learning according to claim 9, characterized in that: The contrast learning module calculates the gradient of the parameters of the prediction model through the backpropagation algorithm and updates the parameters of the prediction model using gradient descent.
Citation Information
Patent Citations
Embryo evaluation method and device based on multi-modal data
CN116434841A
Contrast learning prediction method based on image-text fusion
CN117611576A
Image quality evaluation model training method, electronic equipment and storage medium
CN118799677A
Embryo image automatic focusing method and device based on paired comparison learning, and medium
CN119722658A