Optimal focusing embryo image screening method of ConvNeXt network based on fusion position information

Through the optimal focused embryo image screening method based on the ConvNeXt network based on the fusion location information, the problem of manual selection in the embryo image acquisition process in the prior art is solved, and the efficiency and accuracy of automatic screening of the best focused embryo image is achieved.

CN120014399AActive Publication Date: 2025-05-16HEFEI UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510137869.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-16
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

In the process of embryo image acquisition in assisted reproductive technology, doctors need to manually select the best focused embryo image, which is time-consuming and labor-intensive. Especially when facing large batches of data, the workload is increasing exponentially.

Method used

The optimal focus embryo image screening method based on the ConvNeXt network based on the fusion position information is adopted. Through technical means such as image graying, downsampling and image differential, the focus degree of the embryo image is automatically analyzed to screen out the optimal focus embryo image.

Benefits of technology

It effectively reduces the computational overhead of the model, improves the performance of image data processing and analysis, significantly improves the screening efficiency and accuracy of embryonic image data, reduces manual intervention, and improves the overall work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014399A_ABST
    Figure CN120014399A_ABST
Patent Text Reader

Abstract

The invention discloses an optimal focusing embryo image screening method based on a ConvNeXt network fusing position information. The optimal focusing embryo image screening method comprises the following steps: 1, combining two gray-scale embryo images to manufacture two-channel combined data conforming to a model input pattern; 2, subtracting data on two channels of the dual-channel combined data to obtain a difference image, and performing downsampling; 3, extracting overall features of the difference image by using a convolutional neural network to obtain a corresponding feature matrix; 4, fusing the position index prior information of the original embryo image corresponding to the dual-channel combined data into the feature matrix; 5, constructing a loss function, and training an optimal focusing image dichotomy model; and 7, processing a group of embryo images by using the trained dichotomy model so as to realize screening of optimal focusing images. According to the method, the focusing degree of the embryo image can be effectively analyzed, and the method is excellent in the optimal focusing embryo image screening task and has certain interpretability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and specifically refers to a method for screening best-focused embryo images based on a supervised learning algorithm, which aims to locate the best-focused embryo image in a group of embryo images. Background Art

[0002] In the rapid development of assisted reproductive technology, the field of in vitro fertilization and embryo culture is particularly hot. However, in this process, embryo quality assessment is always a critical and complex task. The accuracy of the assessment directly affects the success rate of embryo transplantation. Therefore, how to accurately judge the embryo development stage through scientific and objective methods and screen out high-quality embryos to achieve the best transplantation effect is one of the core issues in current assisted reproductive technology research.

[0003] In the practice of assisted reproduction, high-quality embryo images are the basis of embryo evaluation. Doctors need to judge the development of the embryo by observing the morphological characteristics of the embryo image, including the cell division, morphological integrity and internal structure of the embryo. However, it is not easy to obtain high-quality images. Factors affecting image quality include the performance of the microscopic imaging equipment, the stability of sample preparation, and the complexity of the imaging environment. Clear embryo images with appropriate contrast can provide doctors with a reliable basis for judgment, while blurred or highly deviated images will lead to uncertainty in the evaluation results and potential misjudgment. Therefore, improving the efficiency and accuracy of embryo image acquisition and screening has become a key link in improving the quality of embryo evaluation.

[0004] In the existing embryo image acquisition process, the time-series incubator has become a commonly used tool. The microscope in the incubator usually automatically shoots the embryo sample multiple times at a set time interval, and obtains images of different focal planes in each shot. This multi-focal plane shooting method avoids the shortcomings of manual focusing to a certain extent, especially reducing the problem of incomplete images caused by microscope focusing errors. However, this technology also brings new challenges: each embryo sample folder usually contains a large number of unfocused images, which cannot be used directly for evaluation. Doctors need to select the embryo images in the best focal plane, that is, the best focused embryo images, from the folder one by one for subsequent analysis and decision-making. This manual selection process is time-consuming and labor-intensive, especially when faced with large quantities of embryo data, the workload increases exponentially. Summary of the invention

[0005] In order to solve the difficulties faced by the above-mentioned scenarios, the present invention proposes a method for screening the best focused embryo images based on a ConvNeXt network that integrates position information, in order to automatically analyze the focusing degree of the embryo image, thereby effectively screening out the best focused embryo image from a group of embryo images and making the result more interpretable.

[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme:

[0007] The present invention provides a method for screening the best focused embryo image based on a ConvNeXt network integrating position information, characterized in that the method is performed according to the following steps:

[0008] Step 1: Obtain an RGB three-channel embryo image dataset ,make Middle The embryo image data of the group is recorded as ,and ,in, express The Embryo image, express The number of embryo images in is the height of an embryo image, is the width of an embryo image; for The embryo images and the best focused images, ;

[0009] Step 2: After graying, we get Grayscale embryo image data ,in, express Grayscale embryo image, let express Grayscale embryo image;

[0010] Step 3: Respectively with the previous Grayscale embryo images Combine them to get the first channel category Two-channel combined data of embryo images ,in, For the first channel category Group Embryo image dual-channel combined data, , and , Before In the grayscale embryo image Grayscale embryo images, where The data of the first channel consists of a grayscale image constitute; The second channel data consists of a grayscale image constitute; The true channel category label is recorded as ;

[0011] Step 4: Respectively and later Grayscale embryo images Combine them to get the second channel category Two-channel combined data of embryo images ,in, For the second channel category Group Embryo image dual-channel combined data, where , and , After indicating In the grayscale embryo image Grayscale embryo images, where The data of the first channel consists of a grayscale image constitute; The second channel data consists of a grayscale image constitute; The true channel category label is recorded as ;

[0012] Step 5: Construct a ConvNeXt network that integrates position information and and Process them separately and get The binary probability of and The binary probability of ,in, and They are The probability of belonging to the first channel category and the second channel category, and and They are The probability of belonging to the first channel category and the second channel category;

[0013] Step 6: Use formula (1) to construct the total loss function of the ConvNeXt network :

[0014] (1)

[0015] In formula (1), represents the first loss function, represents the second loss function;

[0016] Step 7: Use the backpropagation algorithm and optimizer to minimize the total loss function , to train the ConvNeXt network that integrates position information and update network parameters. When the preset maximum number of training rounds or the total loss function is reached When convergence occurs, the training is stopped and a trained best focus image binary classification model is obtained to realize the screening of the best focus image.

[0017] The best focused embryo image screening method based on the ConvNeXt network fused with position information of the present invention is characterized in that the ConvNeXt network fused with position information in step 5 comprises: an image difference module, ConvNeXt modules and fully connected layers;

[0018] Step 5.1: The image difference module will After subtracting the data on the two channels, we get the first channel category Group Difference image ;

[0019] Step 5.2: Difference Image via ConvNeXt modules are used to process the first channel category. Group Passed After processing by the ConvNeXt module Feature Matrix ;

[0020] Step 5.3: Perform a flattening operation to obtain a flattened one-dimensional vector ,in, Represents a one-dimensional vector The length of ;Will Send it to the fully connected layer for processing to obtain dual-channel combined data The binary probability of ,in, and They are The probability of belonging to the first channel category and the second channel category;

[0021] Step 5.4: Follow the process of steps 5.1-5.3 to combine the two-channel data Input the fused position information into the ConvNeXt network for processing, and get The binary probability of ,in and They are The probability of belonging to the first channel category and the second channel category.

[0022] Further, in step 5.2, The ConvNeXt module includes: Downsampling module, Position information fusion module, The depth convolution module, The layer normalization module, Layer scaling module, The path discarding module, A residual connection module; ;

[0023] when When the difference image Enter the In the ConvNeXt module, Downsampling modules are scaled according to the right Downsampling is performed to obtain the first channel category Group After the downsampling module Downsampled difference image ,in, represents the height of the downsampled difference image, , represents the width of the downsampled difference image, , ;

[0024] No. A pair of deep convolution modules Perform channel expansion processing to obtain the first channel category Group After being processed by the deep convolution module A deep convolution feature matrix ,in, is the number of channels of the feature matrix, ;

[0025] No. Position information fusion module calculation Location information , and Expand the dimension and get Expanded location information ; then Expand to Have the same number of channels , get the The location information after further expansion , thus and After adding, we get the first channel category Group The first fusion position information Position information feature matrix ;

[0026] No. Layer normalization module calculation Middle The mean of the channels and variance , and get Specifically, calculate the first channel category Group After layer normalization, The first Feature Matrix , so that the first channel category can be obtained Group After layer normalization, Layer normalized feature matrix ,in, is a constant that prevents the denominator from being zero, ;

[0027] No. The layer scaling module performs a feature matrix Perform channel scaling and change the number of channels from Scaling to 1, thus obtaining the first channel category Group After the layer scaling module Layer scaled feature matrix Specifically, It is through The weighted sum along the channel dimension is: .in It is The learnable weights corresponding to the channels are ;

[0028] No. The path discarding module uses Bernoulli distribution to generate a A random tensor of the same shape and dimension , and Each element in is divided by the set path drop probability Then, the first channel category is obtained. Group After the above operation, Path discard tensors , so as to obtain the first channel category Group After the path discard module processes Path discard feature matrix ,in , ; It means multiplying the corresponding elements of two feature matrices of the same shape;

[0029] No. The residual connection module discards the path feature matrix and downsampled difference image Add the corresponding elements to get the first channel category Group After the residual connection module is processed The residual feature matrix ,in, ;

[0030] when When the first channel category Group After the After processing by the ConvNeXt module Feature Matrix Enter to The first channel category is finally obtained. Group After the After processing by the ConvNeXt module Feature Matrix ,in, express Height, express Width.

[0031] The electronic device of the present invention includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the optimal focus embryo image screening method, and the processor is configured to execute the program stored in the memory.

[0032] The present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the optimal focus embryo image screening method when the computer program is executed by a processor.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] 1. The present invention adopts technical means such as image graying, downsampling and image difference, thereby effectively reducing the computational overhead of the model. Specifically, image graying converts color images into grayscale images, reduces the dimension and complexity of image data, reduces the amount of calculation for subsequent processing, and retains the basic structure and texture information of the image, laying the foundation for subsequent feature extraction. Downsampling further reduces the scale of image data by reducing the resolution of the image, speeds up the processing speed of the model, improves the operating efficiency of the model, enables the model to process more embryo image data in a shorter time, and improves the overall screening efficiency.

[0035] 2. The present invention cleverly uses image difference technology to highlight the difference in the degree of focus between the two embryo images. By performing a difference operation on the two embryo images, the change in the degree of focus of the image can be clearly displayed. This difference provides an important feature for the training of the model, thereby improving the model's ability to extract effective features of dual-channel combined data. This improvement in feature extraction capability enables the model to more accurately screen out the best focused embryo image when processing complex multi-focal plane embryo image data.

[0036] 3. The present invention utilizes ConvNeXt network to extract data features, which significantly improves the performance of embryo image data processing and analysis. ConvNeXt network performs layer-by-layer feature extraction on the input embryo image through deep convolution operation. Each layer of convolution operation can capture the different features in the embryo image data. This multi-level feature extraction method enables the ConvNeXt network to comprehensively analyze the embryo image data. In addition, the network also introduces technologies such as layer normalization, residual connection, and path discarding, which further enhances the training effect and stability of the network.

[0037] 4. The present invention enhances the full extraction of the model's prior knowledge of embryo image data by incorporating embryo image position prior information into the model, improves the final focusing accuracy, and thus makes the present invention perform well in the optimal focus embryo image screening task, and has strong interpretability. Specifically, when processing embryo image data, the method not only considers the characteristics of a single embryo image, but also comprehensively considers the relationship between images and position index information, so as to more accurately judge the optimal focal plane. This method improves the focusing accuracy, makes the screening result more reliable, and also enhances the interpretability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 Schematic diagram constructed for the dataset;

[0039] Figure 2 is an overall flow chart of the method of the present invention;

[0040] Figure 3 A detailed diagram of the testing process for a set of embryo images. DETAILED DESCRIPTION

[0041] In this example, a method for selecting the best focused embryo image based on a ConvNeXt network with fused position information is based on a ConvNeXt model. On this basis, the position information of the embryo image is integrated into the training process of the model. At the same time, based on the idea of ​​comparing adjacent images, a group of embryo images are automatically screened to determine the best focused embryo image. Specifically, first, for the preparation of the training data set, all embryo images are grayed out, and then Figure 1 As shown in FIG. 1 , the present invention first manually determines the best focused embryo image in a group of embryo images in the training set, and combines it with other images in the same group. If the combined image is located in front of the best focused image, it is called the first channel category, i.e., category 1; otherwise, it is called the second channel category, i.e., category 2. Figure 2 As shown, in order to enhance the model's full extraction of data features, the present invention incorporates the position prior information of the two combined embryo images into the model during the model training phase. Figure 3 As shown, for the test phase, the present invention combines adjacent images in a group of embryo images, processes them using a trained model, and automatically screens them round by round and in sequence, discarding embryo images that do not meet the expectations of the model until the last embryo image is output, which is the best focused embryo image, thereby achieving the screening of the best focused embryo image. Specifically, the method is carried out in the following steps:

[0042] Step 1: Obtain an RGB three-channel embryo image dataset ,make Middle The embryo image data of the group is recorded as ,and ,in, express The Embryo image, express The number of embryo images in is the height of an embryo image, is the width of an embryo image; for The embryo images and the best focused images, .

[0043] Step 2: After graying, we get Grayscale embryo image data ,in, express Grayscale embryo image, let express Grayscale embryo image; graying the image can simplify the calculation, reduce the amount of calculation, thereby improving processing efficiency and speeding up the algorithm operation.

[0044] Step 3: Respectively with the previous Grayscale embryo images Combine them to get the first channel category Two-channel combined data of embryo images ,in, For the first channel category Group Embryo image dual-channel combined data, , and , Before In the grayscale embryo image Grayscale embryo images, where The data of the first channel consists of a grayscale image constitute; The second channel data consists of a grayscale image constitute; The true channel category label is recorded as .

[0045] Step 4: Respectively and later Grayscale embryo images Combine them to get the second channel category Two-channel combined data of embryo images ,in, For the second channel category Group Embryo image dual-channel combined data, where , and , After indicating In the grayscale embryo image Grayscale embryo images, where The data of the first channel consists of a grayscale image constitute; The second channel data consists of a grayscale image constitute; The true channel category label is recorded as .

[0046] Step 5: Construct a ConvNeXt network that integrates position information, including: image difference module, ConvNeXt modules and fully connected layers;

[0047] Step 5.1: Image difference module will After subtracting the data on the two channels, we get the first channel category Group Difference image Image difference can suppress some noise that is not related to the spatial position, and at the same time enhance the contrast between different regions in the image, making the target area of ​​interest more obvious. In addition, the dual-channel data is converted into single-channel data through image difference, reducing the computational complexity of the model.

[0048] Step 5.2: Difference Image via ConvNeXt modules are used to process the first channel category. Group Passed After processing by the ConvNeXt module Feature Matrix Among them, The ConvNeXt module includes: Downsampling module, Position information fusion module, The depth convolution module, The layer normalization module, Layer scaling module, The path discarding module, A residual connection module; ;

[0049] when When the difference image Enter the In the ConvNeXt module, Downsampling modules are scaled according to the right Downsampling is performed to obtain the first channel category Group After the downsampling module Downsampled difference image ,in, represents the height of the downsampled difference image, , represents the width of the downsampled difference image, ,and ; Downsampling can reduce the computational overhead of the network;

[0050] No. A pair of deep convolution modules Perform channel expansion processing to obtain the first channel category Group After being processed by the deep convolution module The depth convolution matrix ,in, is the number of channels of the feature matrix, ; Deep convolution can expand the receptive field and extract multi-level features, thereby enhancing the expressive power of the model.

[0051] No. Position information fusion module calculation Location information , and Expand the dimension and get Expanded location information ; then Expand to Have the same number of channels , get the The location information after further expansion , thus and After adding, we get the first channel category Group The first fusion position information Position information feature matrix ; The fusion of position information helps the model to better extract the correlation between two embryo images by using the sequential correlation information of embryo images, thereby improving the model's feature extraction ability for embryo image data.

[0052] No. Layer normalization module calculation Middle The mean of the channels and variance , and get Specifically, calculate the first channel category Group After layer normalization, The first Feature Matrix , so that the first channel category can be obtained Group After layer normalization, Layer normalized feature matrix ,in, is a constant that prevents the denominator from being zero, ; The layer normalization module normalizes the input data of each layer, which can speed up the training and enhance the stability of the model, so that the network can learn feature representation more effectively and avoid the problem of gradient disappearance or explosion.

[0053] No. The layer scaling module performs a feature matrix Perform channel scaling and change the number of channels from Scaling to 1, thus obtaining the first channel category Group After the layer scaling module Layer scaled feature matrix Specifically, It is through The weighted sum along the channel dimension is: .in, It is The learnable weights corresponding to the channels are The layer scaling module introduces learnable weights for the residual connection, which can appropriately scale and adjust the gradient, making the gradient more stable during the back propagation process, effectively alleviating the gradient vanishing problem, so that the network can be trained more smoothly.

[0054] No. The path discarding module uses Bernoulli distribution to generate a A random tensor of the same shape and dimension , and Each element in is divided by the set path drop probability Then, we get the first channel category Group After the above operation, Path discard tensors , so as to obtain the first channel category Group After the path discard module processes Path discard feature matrix ,in, ,and ; It represents the multiplication of the corresponding elements of two feature matrices of the same shape. The path discarding module can choose to discard some paths according to the actual situation, which can reduce the overfitting of the model to the training data.

[0055] No. The residual connection module discards the feature matrix and downsampled difference image Add the corresponding elements to get the first channel category Group After the residual connection module is processed The residual feature matrix ,in, ; Residual connections can fuse features at different levels, allowing the network to simultaneously utilize shallow low-level features and deep high-level features, which can improve the feature extraction capability of the model.

[0056] when When the first channel category Group After the After processing by the ConvNeXt module Feature Matrix Enter to The first channel category is finally obtained. Group After the After processing by the ConvNeXt module Feature Matrix ,in, express Height, express Width.

[0057] Step 5.3: Perform a flattening operation to obtain a flattened one-dimensional vector ,in, Represents a one-dimensional vector The length of .Will Send it to the fully connected layer for processing to obtain dual-channel combined data The binary probability of ,in, and They are The probability of belonging to the first channel category and the second channel category;

[0058] Step 6: Follow the process in step 5 to Input the fused position information into the ConvNeXt network for processing The binary probability of ,in and They are The probability of belonging to the first channel category and the second channel category.

[0059] Step 7: Use formula (1) to construct the total loss function of the ConvNeXt network :

[0060] (1)

[0061] In formula (1), represents the first loss function, represents the second loss function.

[0062] Step 8: Use the backpropagation algorithm and optimizer to minimize the total loss function , to train the ConvNeXt network that integrates position information and update network parameters. When the preset maximum number of training rounds or the total loss function is reached When convergence occurs, the training is stopped and a trained best focus image binary classification model is obtained to realize the screening of the best focus image.

[0063] Step 9: Model testing:

[0064] For Group embryo images , which is grayed out to ,The practice of selecting the best focused embryo images is carried out in the following steps.

[0065] Step 9.1: Set variables as well as Used to record the embryo images and corresponding indexes screened out in each round of the model from step 5 to step 6.

[0066] Step 9.2: Initialize variables as well as .

[0067] Step 9.3: Set the variable The two adjacent embryo images in are combined in order to obtain dual-channel data suitable for the model: . Process the dual-channel data using the model from step 5 to step 6 to obtain the corresponding category 1 or 2. If it is category 1, retain the form In (data of the second channel) to the variable , and the index Record to variable If it is category 2, keep In (data of the first channel) to the variable , and at the same time Record to variable .

[0068] Step 9.4: Remove variables When the same image is saved twice, delete the extra image. The image index in is also deleted.

[0069] Step 9.5: Assume variables Updated to ,variable Updated to .

[0070] Step 9.6: Update the variables The two adjacent embryo images in the sequence are combined to obtain the dual-channel combined data suitable for the model: The dual-channel combined data is processed using the models of step 5 to step 6 in sequence to obtain the corresponding category 1 or 2, and then the variables are updated again based on this. and .

[0071] Step 9.7: Repeat steps 9.4 to 9.6.

[0072] Step 9.8: When the variable When there is only one embryo image in the image, stop step 9.7 and you can use the variable The corresponding image index is obtained from Group embryo images The index of the best focused embryo image in .

[0073] In this embodiment, an electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0074] In this embodiment, a computer-readable storage medium stores a computer program on the computer-readable storage medium, and the computer program executes the steps of the above method when executed by a processor.

[0075] The following is the implementation process of this embodiment:

[0076] 1. Data collection and processing stage:

[0077] (1) Preparation of embryo images:

[0078] The embryo images used in this example were collected by a time-difference incubator (TLS301) developed by Wuhan Huchuang United Technology Co., Ltd. The 2000 sets of embryo images used in the experiment were from this device, each set containing 7 embryo images.

[0079] In the research of automatic embryo image focusing technology, color information is not a decisive factor, and grayscale images are sufficient to provide the main visual clues required for automatic focusing. In order to reduce the complexity of data processing, the present invention converts the color embryo image into a grayscale image.

[0080] (2) Preparation of training data:

[0081] In a group of embryo images, there are different degrees of difference between the best focused embryo image and other embryo images in the same group. This difference is the basis for screening the best focused embryo images. Figure 1 As shown, in this group of embryo images, the position index of the best focused embryo image is 4. Combining it with embryo images with position indexes 1, 2, and 3, data belonging to category 1 is obtained; combining it with images with position indexes 5, 6, and 7, data belonging to category 2 is obtained.

[0082] 2. Model training phase:

[0083] (1) Selection of network model:

[0084] In order to screen out the best focused embryo image from a group of embryo images, the present invention designs a ConvNeXt network that incorporates embryo image position information. The original ConvNeXt modified the convolutional network and referred to some key designs of Transformer, such as: deeper network layer design, using LayerNorm instead of traditional BatchNorm. In addition, some new designs are introduced, such as Depthwise Convolution, which can reduce the number of parameters and computational complexity. At the same time, the GELU activation function is used to replace the traditional ReLU activation function, providing a smoother gradient. ConvNeXt also uses larger convolution kernels (such as 7×7) to allow the network to capture a larger receptive field and enhance the ability to model global information.

[0085] However, for a group of embryo images, since the feature relationship between the data of the two channels of the dual-channel combined data is strongly correlated with the position index within the group of the corresponding original embryo image, the lack of explicit position information may cause the model to misjudge some details for ConvNeXt that does not incorporate position index information. Therefore, incorporating the position index information of embryo images into the ConvNeXt network has certain potential for improving the accuracy of classification tasks. Making full use of this position information can help the model better understand the feature distribution and category discrimination of the data.

[0086] like Figure 2 As shown, two embryo images (including a best focus image) are first grayed and then combined. After the combination, it can be regarded as data with a channel number of 2. Then, the difference image of the two images is calculated, that is, the two channel data are subtracted. The purpose of doing this is to highlight the morphological differences between the best focused image and other images, which is conducive to the model to better extract the features of the constructed data and also reduce the amount of calculation. Afterwards, the difference image is sent to the ConvNeXt that fuses the position information for training. In each ConvNeXt module, the present invention will use the position index information of the two original embryo images (including a best focus image) corresponding to the difference image as prior knowledge and fuse it into the network, so that the model can better extract the features between the two embryo images.

[0087] (3) Model training settings:

[0088] In the experiment, the present invention used an NVIDIA GeForce RTX 4090 GPU with a video memory of 24GB. During the model training process, the epoch was set to 100, the batch size was set to 1, and the learning rate was set to 0.001. The training was carried out on a Linux platform, using a virtual Python environment set up by Anaconda 24.1.2 and PyCharm 2024.1, and the Python version was 3.8.

[0089] 3. Evaluation indicators:

[0090] The present invention takes "focusing accuracy ( )" is used as the evaluation index, and the calculation of this index is based on the following formula:

[0091] (2)

[0092] (3)

[0093] in, It refers to the best focused embryo image screening method based on the ConvNeXt network that integrates position information and combines the idea of ​​combining adjacent embryo images. It is The first images, refers to The inference prediction results obtained after processing this group of embryo images, is the position index of the best focused embryo image in the group of embryo images.

[0094] when When the prediction is correct, the present invention calls it a correct inference prediction. Based on the above ideas, "focus accuracy ( )" is calculated as follows: In the test samples of the group, the number of correct inference predictions of the method proposed in the present invention is ,but ;

[0095] 4. Model testing phase:

[0096] In the testing phase, the present invention designs a method based on the combination of adjacent embryo images, and then classifies and compares them by the model to obtain the best focused embryo image. This method uses the ConvNeXt binary classification model that integrates the position information after training as the selector, and selects the images that meet the model expectations in the adjacent embryo images in turn, thereby achieving the final screening of the best focused embryo image. Figure 3As shown in the figure, a group of embryo images, two adjacent images are combined after grayscale processing to form a data format that can be processed by the binary classification model. After the model processing, the image that meets the model's expectations will be selected. After the deduplication operation, these images will be combined again in a new round of adjacent images, and then sent to the model for processing. After multiple rounds of comparison, the best focused embryo image that meets the model's expectations can be obtained.

[0097] In addition, the present invention also adopts a 5-fold cross validation method. The embryo image data set is divided into five non-overlapping subsets. Each subset is used as a test set in turn, and the remaining four subsets are used for training. This method allows the present invention to conduct five independent experiments, thereby comprehensively evaluating the performance of the present invention. Table 2 shows the accuracy obtained using 5-fold cross validation. The method proposed by the present invention achieved the highest accuracy of 97.25% in the second fold, and the overall average accuracy was 96.25%.

[0098] 5. Experimental results:

[0099] The best focused embryo image screening method based on the ConvNeXt network with fused position information is based on ConvNeXt as the basic model, and the main improvement is to integrate the position information of the embryo image into the feature extraction of the ConvNeXt block. To this end, two groups of comparative experiments are set up in the present invention. One group is trained and tested with the original ConvNeXt network, and the other group is trained and tested with the ConvNeXt network with fused position information. In this way, the excellent effect of the improved network is demonstrated. The present invention uses 2000 groups of samples for comparative experiments and verifies them using the 5-fold cross validation method. The results of the comparative experiments are shown in Tables 1 and 2:

[0100] Table 1 shows the experimental results of the original ConvNeXt;

[0101]

[0102] Table 2 shows the experimental results of the ConvNeXt network that integrates position information;

[0103]

[0104] In general, compared with the original method, the method proposed in the present invention has a certain degree of improvement after verification by a 5-fold crossover experiment.

Claims

1. A method for screening the best focused embryo image based on a ConvNeXt network integrating position information, characterized in that: The steps are as follows: Step 1: Obtain an RGB three-channel embryo image dataset ,make Middle The embryo image data of the group is recorded as ,and ,in, express The Embryo image, express The number of embryo images in is the height of an embryo image, is the width of an embryo image; for The embryo images and the best focused images, ; Step 2: After graying, we get Grayscale embryo image data ,in, express Grayscale embryo image, let express Grayscale embryo image; Step 3: Respectively with the previous Grayscale embryo images Combine them to get the first channel category Two-channel combined data of embryo images ,in, For the first channel category Group Embryo image dual-channel combined data, , and , Before In the grayscale embryo image Grayscale embryo images, where The data of the first channel consists of a grayscale image constitute; The second channel data consists of a grayscale image constitute; The true channel category label is recorded as ; Step 4: Respectively and later Grayscale embryo images Combine them to get the second channel category Two-channel combined data of embryo images ,in, For the second channel category Group Embryo image dual-channel combined data, where , and , After indicating In the grayscale embryo image Grayscale embryo images, where The data of the first channel consists of a grayscale image constitute; The second channel data consists of a grayscale image constitute; The true channel category label is recorded as ; Step 5: Construct a ConvNeXt network that integrates position information and and Process them separately and get The binary probability of and The binary probability of ,in, and They are The probability of belonging to the first channel category and the second channel category, and and They are The probability of belonging to the first channel category and the second channel category; Step 6: Use formula (1) to construct the total loss function of the ConvNeXt network : (1) In formula (1), represents the first loss function, represents the second loss function; Step 7: Use the backpropagation algorithm and optimizer to minimize the total loss function , to train the ConvNeXt network that integrates position information and update network parameters. When the preset maximum number of training rounds or the total loss function is reached When convergence occurs, the training is stopped and a trained best focus image binary classification model is obtained to realize the screening of the best focus image.

2. The method for screening the best focused embryo image based on the ConvNeXt network integrating position information according to claim 1, characterized in that: The ConvNeXt network for integrating position information in step 5 includes: an image difference module, ConvNeXt modules and fully connected layers; Step 5.1: The image difference module will After subtracting the data on the two channels, we get the first channel category Group Difference image ; Step 5.2: Difference Image through ConvNeXt modules are used to process the first channel category. Group After processing Feature Matrix ; Step 5.3: Perform a flattening operation to obtain a flattened one-dimensional vector ,in, Represents a one-dimensional vector The length of ;Will Send it to the fully connected layer for processing to obtain dual-channel combined data The binary probability of ,in, and They are The probability of belonging to the first channel category and the second channel category; Step 5.4: Follow the process of steps 5.1-5.3 to combine the two-channel data Input the fused position information into the ConvNeXt network for processing, and get The binary probability of ,in and They are The probability of belonging to the first channel category and the second channel category.

3. The method for screening the best focused embryo image based on the ConvNeXt network integrating position information according to claim 2, characterized in that: In step 5.2 The ConvNeXt module includes: Downsampling module, The depth convolution module, Position information fusion module, The layer normalization module, Layer scaling module, The path discarding module, A residual connection module; ; when When the difference image Enter the In the ConvNeXt module, Downsampling modules are scaled according to the right Downsampling is performed to obtain the first channel category Group After the downsampling module Downsampled difference image ,in, represents the height of the downsampled difference image, , represents the width of the downsampled difference image, , ; No. A pair of deep convolution modules Perform channel expansion processing to obtain the first channel category Group After being processed by the deep convolution module A deep convolution feature matrix ,in, is the number of channels of the feature matrix, ; No. Position information fusion module calculation Location information , and Expand the dimension and get Expanded location information ; then Expand to Have the same number of channels , get the The location information after further expansion , thus and After adding, we get the first channel category Group The first fusion position information Position information feature matrix ; No. Layer normalization module calculation Middle The mean of the channels and variance , and calculate the first channel category Group After layer normalization, The first Feature Matrix , thus obtaining the first channel category Group After layer normalization, Layer normalized feature matrix ,in, is a constant that prevents the denominator from being zero, ; No. The layer scaling module performs a feature matrix Perform channel scaling and change the number of channels from Scaling to 1, thus obtaining the first channel category Group After the layer scaling module Layer scaled feature matrix ,in, It is The weights to be learned corresponding to the channels are ;and ; No. The path discarding module uses Bernoulli distribution to generate a A random tensor of the same shape and dimension , and Each element in is divided by the set path drop probability Then, the first channel category is obtained. Group After the above operation, Path discard tensors , and then get the first channel category Group After the path discard module processes Path discard feature matrix ,in, ,and ; No. The residual connection module discards the feature matrix and downsampled difference image Add the corresponding elements to get the first channel category Group After the residual connection module is processed The residual feature matrix ,in, ; when When the first channel category Group After the After processing by the ConvNeXt module Feature Matrix Enter to The first channel category is finally obtained. Group After the After processing by the ConvNeXt module Feature Matrix ,in, express Height, express Width.

4. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the best focused embryo image screening method according to any one of claims 1 to 3, and the processor is configured to execute the program stored in the memory.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the best focused embryo image screening method according to any one of claims 1 to 3 are executed.

Citation Information

Patent Citations

  • Automatic focusing method and device, electronic equipment and shooting system

    CN116405775A

  • Embryo image automatic focusing method and device based on multi-task mask feature modeling

    CN118351400A

  • Embryo pronucleus target counting method based on fuzzy elimination and multi-focus image fusion

    CN118379288A

  • Automatic focusing method based on embryo image sequence, electronic equipment and storage medium

    CN119151911A

  • Image generation device, image generation method, recording medium, and processing method

    US20170270662A1