Embryo image automatic focusing method, device and medium based on paired comparison learning
Through the method based on pairwise comparison learning, embryo sample pairs and training neural network models, the problem of inconsistent number of embryo image samples is solved, automatic focus and feature extraction of embryo images are achieved, and research efficiency and analysis accuracy are improved.
Patent Information
- Application Number
- CN202510206274.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Deep learning models are difficult to deal with the problem of inconsistent number of embryonic image samples, resulting in poor autofocusing.
Using a pairwise comparison learning method, we use the method to collect embryo image sets, obtain the best focal image, construct embryo sample pairs, and train them using neural network models to achieve automatic focus of embryo images.
End-to-end automatic focus learning of variable-length embryo images is achieved, and useful features can be automatically extracted and learned and optimized without manual intervention, improving research efficiency and analysis accuracy.
Smart Images

Figure CN119722658B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of embryo image detection, and in particular to an embryo image automatic focusing method, device and medium based on paired comparison learning. Background Art
[0002] Through precise image analysis, embryologists can select the best embryos for transplantation, greatly improving the success rate of pregnancy while also reducing the risk of multiple pregnancies. In addition, clear embryo images can also help study key factors in embryonic development and provide valuable data resources for future clinical applications and research. Embryonic images captured by the built-in CCD in the time-lapse incubator can reflect the morphological changes of the embryo during development in real time. These image sequences are crucial for evaluating the quality and developmental potential of the embryo. However, because the embryo is in a dynamic growth environment, even a slight change in focus can cause image blur, making evaluation difficult.
[0003] In actual operation, due to various factors, the number of samples in each group of embryo images may not be consistent. First, the developmental conditions of embryonic cells in different culture dishes are different. This difference may be due to the physiological characteristics of the cells themselves, the different culture environments, and the different developmental stages. These subtle differences place higher demands on imaging equipment, especially the equipment mode and parameter settings must be suitable for embryo shooting. In the actual imaging process, in order to capture the clearest and most accurate embryonic images, operators are usually not satisfied with the results of a single shot. Therefore, multiple rounds of shooting have become a routine practice for obtaining high-quality embryonic images. Through this multiple rounds of shooting, researchers can ensure that the embryonic cells on the culture dish, regardless of their developmental status, can be recorded as clearly as possible. This not only improves the accuracy of subsequent analysis, but also makes long-term tracking and comparative studies of embryos possible.
[0004] However, multiple rounds of embryo images will lead to inconsistent numbers of embryo samples in each group. The requirements of different reproductive centers and the operating habits of different doctors will also lead to inconsistent numbers of embryo image samples in each group. Different reproductive centers may have different operating standards and procedures. Some reproductive centers may require more frequent shooting to ensure that every key stage of embryo development is recorded, while some centers may reduce the number of shots based on experience to save time and resources. In addition, different doctors have their own preferences and habits during the operation. For example, some doctors may prefer to adjust the focal length and shooting angle multiple times to capture the most ideal image, while some doctors may rely more on rapid continuous shooting to obtain more samples. These differences in operation will undoubtedly further lead to inconsistent numbers of embryo image samples in each group, thereby affecting the subsequent realization of automatic focusing of embryo images.
[0005] In recent years, deep learning has been widely recognized as a powerful tool in the field of artificial intelligence, and its potential in image recognition and analysis has been widely recognized. However, the input of general deep learning models requires fixed-length input data for effective feature extraction and learning. Traditional deep learning methods are not applicable to variable-length data. In addition, deep learning models with a single-mode architecture cannot accurately capture the subtle global and local features of embryo images, resulting in poor final results. Summary of the invention
[0006] The present invention proposes an embryo image automatic focusing method, device and medium based on paired comparison learning to solve the technical problem that deep learning models are difficult to train using inconsistent numbers of embryo image samples.
[0007] To solve the above technical problems, the present invention provides an embryo image automatic focusing method based on paired comparison learning, comprising the following steps:
[0008] Step S1: collecting a number of sets of embryo image sets, each set of embryo image sets is a number of embryo images taken during the focusing process, and presents a pattern of blurring, focusing, and then blurring again;
[0009] Step S2: Obtain the sequence number of the embryo image with the best focal length in each set of embryo images ; The embryo image with the best focal length is paired with the remaining embryo images in the group to construct embryo sample pairs and ,Will The label of is reconstructed to 0, The label is reconstructed to 1; the adjacent images in each group of embryo images are paired to construct embryo sample pairs and , ,like , then The label of is reconstructed to 0, The label of is reconstructed to 1. , then The label of is reconstructed to 1, The label of is reconstructed to 0;
[0010] Step S3: inputting the embryo sample pair into a neural network model for training;
[0011] Step S4: Input the embryo image set to be automatically focused into the trained neural network model to obtain the embryo image with the best focal length in the embryo image set to be automatically focused.
[0012] Preferably, the embryo sample pairs obtained in step S2 are stacked to form embryo sample pairs with three-dimensional tensors; in step S3, the embryo sample pairs with three-dimensional tensors are input into the neural network model for training.
[0013] Preferably, each embryo image in the plurality of embryo image sets collected in step S1 is converted into a grayscale image.
[0014] Preferably, during the training process, the neural network model is trained using a mixed loss function consisting of a ranking loss calculated based on the labels of embryo sample pairs and a cross entropy loss function.
[0015] Preferably, the method for calculating the sorting loss based on the labels of the embryo sample pairs comprises the following steps:
[0016] Step S21: Calculate the difference between the scores of two embryo image samples:
[0017] ;
[0018] In the formula, and Represents embryo sample pairs Embryo images i and embryo images j The relative focus score output after passing through the neural network model;
[0019] Step S22: Convert the difference of embryo images into probability through Sigmoid function :
[0020] ;
[0021] Step S23: Calculate the sorting loss of the embryo image sample pair according to the relative position of the embryo image sample pair :
[0022] ;
[0023] In the formula, represents the true label;
[0024] Step S24: Sum all embryo image samples to get the total ranking loss :
[0025] ;
[0026] In the formula, Represents the set of all sample pairs in the embryo image training set.
[0027] Preferably, the hybrid loss function The expression is:
[0028] ;
[0029] In the formula, represents the regularization parameter; represents the cross entropy loss function.
[0030] Preferably, the neural network model extracts features through a feature extraction module and then performs a Flatten operation to obtain an expanded feature vector; the feature vector is input into a KAN module to locate the embryo image with the best focal length.
[0031] Preferably, the feature extraction module sequentially passes through four Conv-Trans modules to perform feature extraction;
[0032] The Conv-Trans module first undergoes a downsampling operation, and then inputs the downsampled features into the CNN module and the Transformer module simultaneously for feature extraction and splicing.
[0033] The present invention also provides an electronic device, comprising: a memory, a processor and a computer program, wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the above method.
[0034] The present invention also provides a computer-readable storage medium, in which a computer program is stored. The computer program is executed by a processor to implement the above method.
[0035] The beneficial effects of the present invention include at least: the present invention realizes the end-to-end automatic focusing learning task of variable-length embryo images and trains the deep learning network. This method enables the algorithm to automatically and accurately extract useful features from embryo image training sample data of different lengths, and learn and optimize without manual intervention. Through the application of deep learning algorithms, researchers can be freed from the tedious work of feature extraction and focus more on core tasks such as research design and result analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 A schematic diagram of a method flow chart of an embodiment of the present invention;
[0037] Figure 2 A schematic diagram of the focusing state of an embryo image according to an embodiment of the present invention;
[0038] Figure 3 A schematic diagram of constructing an embryo sample pair according to an embodiment of the present invention;
[0039] Figure 4Schematic diagram of the CTK-Net network model structure of an embodiment of the present invention;
[0040] Figure 5 This is a schematic diagram of the model training and testing process after the solutions of Embodiments 1 to 3 are integrated in an embodiment of the present invention;
[0041] Figure 6 A schematic diagram of a detailed testing process of a set of embryo images according to an embodiment of the present invention;
[0042] Figure 7 Schematic diagram of the effects of different methods on focusing on actual embryo images. DETAILED DESCRIPTION
[0043] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention.
[0044] In today's assisted reproductive technology, the collection and analysis of embryo images plays a vital role. Through the time-difference incubator, scientists can observe the entire development process of the embryo from the fertilized egg to the blastocyst stage, which is of great significance for evaluating the quality of the embryo and predicting its implantation potential. The embryo image data adopted in the present invention not only comes from multiple different reproductive centers, but also covers the whole process of embryonic cell development in the time-delay incubator, as well as the image data in the empty culture dish, which provides a very rich and diverse source of information for research. However, in the actual embryo image acquisition process, each group of embryo images has differences in sample quantity due to a variety of different factors. This has brought great challenges to subsequent research work.
[0045] There are many reasons for the inconsistency in the number of embryo image samples in different groups. The main reason is that the developmental states of embryonic cells in different culture dishes show diversity. This difference puts higher standards on embryo imaging technology, especially the imaging equipment needs to adjust its mode and parameters to adapt to the shooting needs of the embryo. In actual operation, technicians usually use multiple rounds of imaging to obtain clear and accurate embryo images. It is this method based on multiple imaging that leads to the problem of inconsistent number of sample groups. Secondly, the inconsistency in the number of embryo image samples may be due to differences in requirements of different reproductive centers and different operating habits of doctors. These differences in factors will undoubtedly further lead to inconsistencies in the number of embryo image samples in each group, thereby affecting the subsequent research on the automatic focusing task of embryo images.
[0046] In addition, building a high-quality embryo image dataset is also a complex and tedious process. In actual operation, screening out images with the best focal length is a very challenging task. Experienced embryologists need to select the most representative image from a set of embryo images based on the characteristics of different stages of embryonic development. This requires not only a deep understanding of embryology, but also rich practical experience and keen observation. And because of the huge number of embryo images, data screening and optimization becomes a time-consuming and labor-intensive task. However, this step is crucial for the subsequent morphological analysis of embryos, because only images with the best focal length can provide more accurate morphological features, thereby helping embryologists make correct judgments and decisions.
[0047] Faced with such a huge amount of data, it is particularly important to choose the right research method. Traditional manual analysis methods require professional personnel to perform analysis, which is not only inefficient, but also different professionals may have different judgment standards and subjective biases, which will lead to inconsistent analysis results. And manual analysis faces the dual challenges of time and cost when processing large-scale data sets, which makes it difficult to meet the needs of quickly and accurately assessing embryo quality. In addition, due to the multiple rounds of shooting methods used in the embryo image acquisition process and the different operating habits of reproductive centers and doctors, it is impossible to ensure that the number of images in each group is unified.
[0048] In order to overcome these limitations, the present invention adopts embodiment 1 to solve the problem.
[0049] Example 1
[0050] like Figure 1 As shown, an embodiment of the present invention provides an embryo image automatic focusing method based on paired comparison learning, comprising the following steps:
[0051] Step S1: Collect several groups of embryo image sets, each group of embryo image sets is a number of embryo images taken during the focusing process, and presents a pattern of blurring, focusing, and then blurring again.
[0052] The present invention has collected tens of thousands of embryo images from multiple reproductive centers. These images record the whole process of embryonic cell development from fertilized eggs to blastocysts in detail, providing a rich data basis for research. In order to obtain clear layers in the embryonic development process, the embryonic development state is usually recorded by a multi-round shooting method, as well as the requirements of different reproductive institutions and the habits of doctors. These factors lead to inconsistent numbers of each group of embryo images, which are generally between 5 and 16. Five professionals analyze each group of embryo images in detail. During the analysis process, the professionals use the principle of "the minority obeys the majority" to jointly decide the image of the best focal length as the label of the group. This principle ensures the fairness and accuracy of the decision-making process and avoids the influence of personal subjective bias on the results. According to the different stages of embryonic development, more than 50,000 embryo images are selected, totaling 5000 groups. The data set is divided into a training set and a test set using a ten-fold cross validation method.
[0053] Exemplarily, in the research of automated embryo image focusing technology, color information is often not a decisive factor, and grayscale images are sufficient to provide the main visual clues required for automatic focusing. In order to reduce the complexity of processing and give play to the characteristics of grayscale images, the present embodiment adopts the cv2.cvtColor() function in the OpenCV library to achieve the conversion of colored embryo images to grayscale images. This step not only effectively eliminates the unnecessary interference brought by color, but also can greatly improve the accuracy and efficiency of the network model to the best focus picture judgment, and provide a reliable basis for subsequent analysis and research.
[0054] Step S2: Obtain the sequence number of the embryo image with the best focal length in each set of embryo images ; The embryo image with the best focal length is paired with the remaining embryo images in the group to construct embryo sample pairs and ,Will The label of is reconstructed to 0, The label is reconstructed to 1; the adjacent images in each group of embryo images are paired to construct embryo sample pairs and , ,like , then The label of is reconstructed to 0, The label of is reconstructed to 1. , then The label of is reconstructed to 1, The label of is reconstructed to 0.
[0055] Specifically, according to the prior knowledge that there is a blur-focus-blur rule in embryo images, in each embryo image set, a series of images before the embryo image with the best focal length often show a trend of gradually changing from blur to clarity, and after the best focal length point, the embryo image will gradually become blurred. The specific situation is as follows: Figure 2 shown.
[0056] like Figure 3 The following is a rule for constructing sample pairs based on prior knowledge of embryo images. First, the embryo image with the best focal length in each group of embryo images is extracted based on the label information. Then, in the first step, the embryo image with the best focal length is paired with the remaining embryo images in the group to construct embryo sample pairs. and , where the sample pair The label of is reconstructed to 0, and the sample pair The label is reconstructed to 1; the second step is to construct sample pairs by pairing adjacent images in sequence according to the prior knowledge of embryo images. and , , if both samples are within the blur-focus position range, then the sample pair The label of is reconstructed to 0, and the sample pair The label of is reconstructed to 1. If both samples are in the focus-blur position range, then the sample pair The label of is reconstructed to 1, and the sample pair The label is reconstructed to 0.
[0057] In order to further improve the efficiency of the network model to sample pair processing, the present embodiment stacks the sample pairs in advance. Specifically, by adding a new dimension and stacking the sample pairs along the dimension, a three-dimensional tensor embryo sample pair data is formed, which can make the data of a sample pair be completely retained. This design not only makes the data structure more compact, but also facilitates the neural network model to be trained in batch mode, speeds up the iteration speed of the model weight, and thus accelerates the learning process. At the same time, stacking images into multidimensional tensors simplifies the data processing process and improves the speed and stability of model training. In practical applications, this is conducive to the preprocessing, enhancement and loading of data, and lays a more solid foundation for the training and prediction of the model. The time for data conversion and loading is reduced, so that more resources and time can be used for the optimization and feature engineering of the model, further improving the overall performance of the model.
[0058] Through the above method, the embryo image training set used in this embodiment is reconstructed, while the test set does not need to construct sample pairs in advance because the image with the best focal length cannot be known in advance.
[0059] Step S3: inputting the embryo sample pairs into the neural network model for training;
[0060] Step S4: Input the embryo image set to be automatically focused into the trained neural network model to obtain the embryo image with the best focal length in the embryo image set to be automatically focused.
[0061] Traditional algorithms rely on professionals to manually extract different features for research, which not only increases the professional requirements for researchers, but also limits research efficiency. Researchers need to have in-depth field knowledge and rich experience to accurately select and extract key features. This process is often time-consuming and labor-intensive, and is easily affected by subjective factors, thereby affecting the reliability of the research. However, using the method of the present embodiment, end-to-end automatic focusing learning of variable length embryo images is achieved. The algorithm is enabled to automatically extract useful features from raw data, and learn and optimize without manual intervention. Through the application of deep learning algorithms, researchers can be liberated from cumbersome feature extraction work and focus more on core tasks such as research design and result analysis.
[0062] Example 2
[0063] Based on the problem that each group of embryo image samples is inconsistent, this embodiment forms sample pairs of embryo images and compares the focus quality of a pair of embryo images, which can more effectively train the network to distinguish different focus levels. To achieve this goal, this embodiment optimizes the network model by calculating the sorting loss function obtained by the discounted cumulative gain value of the embryo sample pair based on Example 1.
[0064] like Figure 3 As shown, suppose there is a set of embryo images , among which the group has embryo image samples, the index number of the embryo image with the best focal length is The training sample pairs are formed by sequentially constructing sample pairs from the embryo image with the best focal length and other embryo images in the group. The sample pair set constructed from this group of embryo images can be expressed as , whose label is For a training sample pair , input it into the neural network model, and obtain the network model for the embryo image sample pairs Relative focus score in and .
[0065] Step 1: Calculate the difference between the scores of two embryo image samples:
[0066] (1)
[0067] Step 2: Convert the difference of embryo images into probability through Sigmoid function:
[0068] (2)
[0069] Step 3: Calculate the ranking loss of the embryo image sample pairs according to their relative positions:
[0070] (3)
[0071] in, .
[0072] Step 4: Sum all embryo image samples to get the total ranking loss:
[0073] (4)
[0074] in, It refers to the set of all sample pairs in the embryo image training set. The advantage of the ranking loss is that it is suitable for dealing with the relative order problem between embryo image sample pairs, and can effectively optimize the performance of the model in the task of ranking the focus of embryo images. By comparing the score differences between embryo image sample pairs, the ranking relationship between them can be learned.
[0075] The cross entropy loss function is widely used in classification tasks and has many advantages. In the present invention, it maximizes the log-likelihood of the embryo image by minimizing the loss, and improves the prediction accuracy of the model for the best focal length image in the embryo image. It has high sensitivity, can quickly reflect the performance changes of the model, and promotes rapid convergence. The cross entropy loss is easy to implement and has numerical stability and probabilistic interpretability, making it a standard choice in the task of the present invention. The cross entropy loss function is calculated as follows:
[0076] (5)
[0077] in, Refers to the set of all samples in the embryo image training set; It is i The label corresponding to each embryo image sample is converted into a binary value by a 1-hot encoder. It is i The classification probability that is relatively more focused in a sample pair is output by the classifier model after each embryo image sample passes through the classifier model; It is j The label corresponding to each embryo image sample is converted into a binary value by a 1-hot encoder. It is j After the embryo image samples pass through the classifier model, the classification probability that is relatively more focused in a sample pair is output.
[0078] According to the specific embryo image automatic focusing task based on paired comparison learning in the present invention, a hybrid loss function consisting of a ranking loss and a cross entropy loss function is used to jointly optimize the network model. The hybrid loss function is calculated as follows:
[0079] (6)
[0080] in, refers to the ranking loss function; refers to the cross entropy loss function; Refers to the regularization parameter, which represents the regularization weight of the two loss functions. The value of can adjust the actual contribution of the ranking loss function to the entire network during the model training process; since the role of the relative ranking module is to strengthen the learning of the relative ranking of the focusing degree of the embryo images in the sample pairs in order to strengthen the feature extraction model in the model training stage, it is deleted in the testing stage. Therefore, the model adopted in this embodiment is still considered to be an end-to-end framework, and no additional cost is added due to the relative ranking module.
[0081] Example 3
[0082] In order to further extract more representative focus features in embryo images, this embodiment designs an innovative hybrid network architecture based on Embodiment 1 to replace the neural network model in step S3, referred to as CTK-Net. CTK-Net organically combines three advanced technologies such as Convolutional Neural Network (CNN), Transformer and KAN (Kolmogorov-Arnold Networks). Figure 4 As shown in the figure, the network's ability to capture fine-grained features in embryo images is significantly improved, while the modeling performance of global features is optimized.
[0083] First, the embryo image data undergoes a series of preprocessing operations before being input into the network to ensure that the data can maintain consistency and higher effectiveness during the feature extraction process. Subsequently, the preprocessed embryo image data is input into a feature extraction network consisting of four consecutive Conv-Trans modules. Each Conv-Trans module fully combines the advantages of CNN and Transformer to model the local details and global patterns of the embryo image respectively. The CNN module focuses on extracting the texture features between cells, the clarity of cell boundaries, and local significant features such as particle distribution and cell shape; the Transformer module is responsible for capturing the symmetry and uniformity of the overall structure in the embryo image, as well as the long-range dependencies of spatial distribution. This local and global collaborative modeling method enables this stage to fully capture the important information in the embryo image, providing a solid foundation for subsequent feature processing and evaluation. Next, the extracted multidimensional features are flattened into a one-dimensional feature vector through the Flatten operation. This operation not only retains the high-dimensional feature information extracted in the previous stage, but also converts it into a unified representation so that subsequent modules can process these features more efficiently. At the same time, the Flatten operation can ensure that the information of local features and global features does not interfere with each other, laying the foundation for feature importance analysis and subsequent weight optimization. Finally, the expanded feature vector is input into the KAN module. The KAN module uses its unique learnable edge activation function and B-spline nonlinear transformation mechanism to perform fine-grained weight optimization on the input features. In this process, the KAN module dynamically evaluates and adjusts the importance of each feature node, focusing on key features in the embryo image that are closely related to developmental assessment and focal length optimization. For example, it can more accurately capture the fissure clarity of embryonic cells, the morphological characteristics of the nucleus, and the overall symmetry distribution through node weight allocation, thereby accurately locating and generating embryo images with optimal focal length. This process not only significantly improves the accuracy of focal length optimization, but also provides more comprehensive and reliable data support for embryo quality assessment, laying a solid foundation for subsequent medical decision-making.
[0084] The design of the Conv-Trans module fully combines the characteristics of embryo images, aiming to provide accurate and efficient feature representation for embryo images in focus analysis tasks through the collaborative modeling of local and global information. The module mainly consists of three parts: downsampling, parallel structure of CNN and Transformer, and feature splicing. The first step of the Conv-Trans module is the downsampling operation, which reduces the spatial dimension of the input embryo image features through average pooling. In embryo images, edge information and texture information are key features for distinguishing the focus state, and downsampling can effectively preserve the integrity of this information while providing support for multi-scale receptive fields for subsequent processing. The downsampled features are simultaneously input into the CNN and Transformer branches. This parallel structure gives full play to the advantages of both and collaboratively models the local details and global context of the input features. In the CNN branch, the feature extraction process consists of three consecutive small-size convolution operations, with the sizes of the convolution kernels being 3×3, 1×1, and 1×1, respectively. This design gradually reduces the receptive field, thereby more accurately capturing the local texture information in the embryo image. After each convolution operation, the ReLU activation function is connected to introduce nonlinear transformation and enhance the model's ability to express detailed information in the embryonic image. The design of the CNN branch structure can not only effectively extract the multi-scale features of the embryonic image in terms of details, but also compress irrelevant information while retaining key details, thereby significantly improving the model's ability to express embryonic image features, and ultimately providing a richer and more accurate feature representation for the following feature extraction network. In the Transformer branch, it is mainly composed of normalization layers, multi-head attention mechanisms, feedforward neural networks, and residual connections, focusing on capturing the global pattern of embryonic images. By normalizing the input features of the embryonic image in the channel dimension, ensuring the stability of the feature value range, it helps to balance the feature strength of different regions in the embryonic image, such as comparing the difference between the subtle structure inside the embryo and the background brightness. The main purpose of the multi-head attention mechanism is to capture the global dependency of the input features of the embryonic image, that is, to allow each feature point to perceive the information of other feature points. The feedforward neural network is responsible for nonlinear transformation and further feature enhancement of the output of the multi-head attention mechanism. It contains two fully connected layers and an activation function layer, which can further enhance the modeling ability of global morphological information in embryo images, such as the consistency assessment of the overall structure of the embryo. The residual connection retains the feature information of the original embryo image input and adds it directly to the output. This mechanism not only avoids the gradient vanishing problem, but also improves the robustness of the network to key features by retaining embryo detail information. For example, it can help the model retain weak but important local texture features in embryo images at a deep level.At the same time, the residual path can also promote the information flow of deep networks, allowing the model to effectively transfer features in deeper structures, and ultimately improve the global modeling capabilities and feature expression performance of the Transformer branch. The final stage of the Conv-Trans module is feature splicing, which integrates the features extracted by the CNN branch and the Transformer branch to achieve effective fusion of local information and global information, providing rich and unified feature representation for subsequent networks. Through this design, the Conv-Trans module can accurately capture the detailed information and global patterns in embryonic images, providing powerful feature expression capabilities for embryonic images in focusing tasks, while providing more efficient input for downstream network modules, thereby further improving the overall performance and application effects of the model.
[0085] The core of the KAN module is to achieve nonlinear transformation of input features through learnable edge activation functions, breaking through the limitations of fixed node activation functions in traditional multi-layer perceptrons (MLPs), thereby significantly improving the model's adaptability to complex data patterns. Specifically, KAN uses the B-Spline function as the parameterized form of the activation function, and flexibly adjusts the shape of the activation function by learning control points and node positions to meet the needs of different tasks. In the task of analyzing the focus state of embryonic images, the embryonic image features extracted by the Conv-Trans module are flattened and then input into the KAN module. The KAN module can use its powerful nonlinear modeling capabilities to deeply analyze the embryonic image features from both local and global dimensions. Through the flexibility of the B-Spline activation function, the KAN module can dynamically capture the subtle differences between the focus state features of the two embryonic images. This high-precision capture of feature differences provides a reliable basis for the selection and decision-making of focus features. In addition, the KAN module dynamically adjusts the shape of the activation function to amplify the significant focus discriminant features in the embryo image, while suppressing background noise or irrelevant features. This significantly improves the assessment and classification accuracy of the entire model on the embryo focus situation, providing highly reliable support for subsequent medical decision-making.
[0086] Example 4
[0087] In order to further illustrate the effectiveness of the network model proposed in Example 3, this example integrates Examples 1 to 3 and illustrates the model training, testing and experimental results. The training and testing process is as follows: Figure 5 shown.
[0088] The present invention explores an automatic focusing technology for embryo images based on paired comparison learning, aiming to solve the problem of inconsistent sample numbers in each group of embryo images. By constructing a sample pair of embryo images for pairwise comparison learning, relatively focused samples in the sample pair are selected, and then a group of multiple embryo image samples are compared to improve the performance of the automatic focusing task of the embryo image. In the experimental setting, the Adam optimizer was selected for parameter adjustment, the learning rate was set to 0.005, the batch size was 16, and 60 training cycles were performed. The optimizer combines weight decay and L2 regularization strategies, effectively solving the challenges of slow network convergence and parameter overfitting. In detail, the present invention uses the Adam optimizer to optimize all network model parameters involved, with the purpose of minimizing the overall loss function L, and then optimizing the parameters of the entire network structure. This optimization method can generate an excellent network model specifically for performing the automatic focusing task of embryo images. This model can accurately identify the best focused pictures in a group of embryo images, thereby providing a more accurate data basis for further analysis and research of embryo images.
[0089] The sample pairs constructed in advance during the training phase are based on the best focal length images in a known set of embryo images. This operation is to stack the embryo images into a multi-dimensional tensor, thereby simplifying the model's data processing flow and improving the speed and stability of model training. However, in the test phase for embryo images, the number of samples and labels of a set of embryo images cannot be known in advance to the network model. Therefore, the test data processing flow needs to be further explained in detail during the test phase.
[0090] like Figure 6 The following is a test process of a set of embryo images. First, the images are grayed and resized in the same way as in the training set to ensure consistency and comparability of the test set. Second, the first two samples are taken in the order of embryo image retention to construct a sample pair. , this is for the sake of orderliness and rationality of comparison. Input into the trained network model for sorting. If the model predicts that the label of the sample pair is 0, it means that the latter is more focused in the sample pair; if the model predicts that the label of the sample pair is 1, it means that the former is more focused in the sample pair. Record the relatively focused sample in the first sample pair Next, construct the second sample pair , the processing steps are the same as the first sample pair, and the relatively focused samples in the sample pair are recorded. And so on, until the last sample pair is constructed. , continue to compare the focusing degree of the two samples according to the above method, and finally obtain the best focal length image The model returns the index number of the best focal length image of the group of embryo images .
[0091] This testing process not only improves the accuracy of the selection, but also enhances the reliability of the results, thereby ensuring the quality of subsequent embryo image analysis and research. In addition, the efficiency of this method is reflected in its gradual elimination process. Instead of repeated comparisons of the entire test set, the embryo image with the best focal length can be determined in one traversal, which greatly saves computing resources and time.
[0092] In order to verify the validity of the proposed network model, the present embodiment carries out a large number of experiments on model parameter tuning and contrast experiments of different methods. Because the number of data samples of different groups in the embryo image data set taken by the present embodiment is variable, this causes the present invention to be unable to adopt other traditional deep learning models to carry out contrast experiments and comparisons between different methods. Therefore, the contrast experiment of the present invention is mainly compared with the traditional algorithm. The traditional algorithm used in the contrast experiment mainly adopts three different gradient features of Canny, Tenengrad and Laplacian extracted manually to carry out experimental verification. Figure 7 The results of the best focal length images obtained by different methods on the same set of embryo images are shown. Figure 7 It can be seen that different gradient features have different advantages in autofocusing of embryo images at different developmental stages. The Canny gradient feature is better for identifying embryo images at the cell stage; the Tenengrad gradient feature is more suitable for embryo images at the blastocyst stage; and the Laplacian gradient feature can achieve better performance on empty dish images. These forms of manually extracting features based on traditional algorithms can perform better on embryo images at a certain developmental stage, but worse on embryo images at all developmental stages. Through experiments, it was found that the CTK-Net network model proposed in the present invention achieves excellent performance on embryo images at different developmental stages.
[0093] The present invention also provides an electronic device, comprising: a memory, a processor and a computer program, wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the above method.
[0094] The present invention also provides a computer-readable storage medium, in which a computer program is stored. The computer program is executed by a processor to implement the above method.
[0095] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. Only the preferred embodiments of the present invention are expressed. The description is more specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. As long as there is no contradiction in the combination of these technical features, they should be considered as the scope recorded in this specification.
[0096] It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these modifications and improvements all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. An embryo image automatic focusing method based on paired comparison learning, characterized in that: The following steps are involved: Step S1: collecting a number of sets of embryo image sets, each set of embryo image sets is a number of embryo images taken during the focusing process, and presents a pattern of blurring, focusing, and then blurring again; Step S2: Obtain the sequence number of the embryo image with the best focal length in each set of embryo images ; The embryo image with the best focal length is paired with the remaining embryo images in the group to construct embryo sample pairs and ,Will The label of is reconstructed to 0, The label is reconstructed to 1; the adjacent images in each group of embryo images are paired to construct embryo sample pairs and , ,like , then The label of is reconstructed to 0, The label of is reconstructed to 1. , then The label of is reconstructed to 1, The label of is reconstructed to 0; Step S3: inputting the embryo sample pair into a neural network model for training; Step S4: Input the embryo image set to be automatically focused into the trained neural network model to obtain the embryo image with the best focal length in the embryo image set to be automatically focused.
2. The method for automatically focusing embryo images based on paired comparison learning according to claim 1, characterized in that: The embryo sample pairs obtained in step S2 are stacked to form embryo sample pairs with three-dimensional tensors; in step S3, the embryo sample pairs with three-dimensional tensors are input into the neural network model for training.
3. The embryo image automatic focusing method based on paired comparison learning according to claim 1, characterized in that: Each embryo image in the plurality of embryo image sets collected in step S1 is converted into a grayscale image.
4. The embryo image automatic focusing method based on paired comparison learning according to claim 1, characterized in that: During the training process, the neural network model is trained using a mixed loss function consisting of a ranking loss calculated based on the labels of embryo sample pairs and a cross entropy loss function.
5. The embryo image automatic focusing method based on paired comparison learning according to claim 4, characterized in that: The method for calculating the sorting loss based on the labels of embryo sample pairs includes the following steps: Step S21: Calculate the difference between the scores of two embryo image samples: ; In the formula, and Represents embryo sample pairs Embryo images i and embryo images j The relative focus score output after passing through the neural network model; Step S22: Convert the difference of embryo images into probability through Sigmoid function : ; Step S23: Calculate the sorting loss of the embryo image sample pair according to the relative position of the embryo image sample pair : ; In the formula, represents the true label; Step S24: Sum all embryo image samples to get the total ranking loss : ; In the formula, Represents the set of all sample pairs in the embryo image training set.
6. The embryo image automatic focusing method based on paired comparison learning according to claim 5, characterized in that: The hybrid loss function The expression is: ; In the formula, represents the regularization parameter; represents the cross entropy loss function.
7. The embryo image automatic focusing method based on paired comparison learning according to claim 1, characterized in that: The neural network model extracts features through a feature extraction module and then performs a Flatten operation to obtain an expanded feature vector; the feature vector is input into a KAN module to locate an embryo image with an optimal focal length.
8. The embryo image automatic focusing method based on paired comparison learning according to claim 7, characterized in that: The feature extraction module sequentially passes through four Conv-Trans modules to perform feature extraction; The Conv-Trans module first undergoes a downsampling operation, and then inputs the downsampled features into the CNN module and the Transformer module simultaneously for feature extraction and splicing.
9. An electronic device, comprising: A memory, a processor and a computer program, characterized in that the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Pathological image imaging quality assessment method
CN119048446A
Automatic focusing method based on embryo image sequence, electronic equipment and storage medium
CN119151911A