Embryo development stage hierarchical classification method based on convolutional neural network
Through the combination of convolutional neural network and cross-attention mechanism, high-precision automated classification of embryonic development stage is achieved, and the inefficiency and low accuracy of traditional manual evaluation is solved, and it is suitable for assisted reproductive technology and embryology research.
Patent Information
- Application Number
- CN202510587553.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional embryonic developmental stage evaluation relies on manual examination, which is time-consuming and susceptible to subjective factors, resulting in low accuracy and efficiency, making it difficult to achieve high-precision embryo selection.
Convolutional neural network (CNN) combined with cross attention mechanism is used to automatically capture subtle changes and complex patterns in embryonic development through multi-stage feature extraction and fine-grained classification, dynamically adjust the attention to key areas, and achieve accurate classification.
It significantly improves the classification accuracy of the embryonic development stage, reduces the possibility of misclassification, improves the model's ability to handle abnormal situations and atypical samples, and is suitable for assisted reproductive technology and embryological research.
Smart Images

Figure CN120495760A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing and analysis, and in particular relates to a hierarchical classification method for embryonic development stages based on convolutional neural networks. Background Art
[0002] Biological development follows a strict sequence, particularly in mammalian embryonic development, where each stage progresses irreversibly, from single-cell cleavage to morula and blastocyst stages. This monotonic constraint ensures the stability and consistency of developmental progression. However, in IVF practice, selecting the most viable embryos for transfer is crucial for achieving a successful pregnancy. Traditional embryo assessment relies on manual examination by clinicians using a microscope, a time-consuming and subjectivist approach that limits its accuracy and efficiency.
[0003] Therefore, the development of an automated, high-precision embryonic development stage classification model is urgently needed. This invention is designed to address this problem. It combines the convolutional neural network (CNN) and the cross-attention mechanism in deep learning to achieve accurate classification of embryonic development stages. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a hierarchical classification method for embryonic development stages based on convolutional neural networks, which combines the convolutional neural network (CNN) and cross-attention mechanism in deep learning to achieve accurate classification of embryonic development stages.
[0005] The present invention provides a method for hierarchical classification of embryonic development stages based on a convolutional neural network, comprising:
[0006] Acquire embryo images;
[0007] acquiring a feature map according to the embryo image;
[0008] The feature maps are fused and coarse-grained classified to obtain the main stage probability;
[0009] Fusing the feature maps after the coarse-grained classification and performing fine-grained classification to obtain sub-stage probabilities;
[0010] Obtaining category probabilities based on the main stage probabilities and the sub-stage probabilities;
[0011] The embryonic development stages are hierarchically classified according to the category probabilities.
[0012] Optionally, acquiring the embryo image includes:
[0013] Acquire initial embryo images;
[0014] The initial embryo image is preprocessed to obtain the embryo image.
[0015] Optionally, obtaining a feature map according to the embryo image includes:
[0016] Inputting the embryo image into a pre-trained ResNet50 model to obtain the feature map, wherein the ResNet50 model includes: an initial convolutional layer, a batch normalization layer, a maximum pooling layer, and a residual block;
[0017] The initial convolutional layer is used to extract low-level features;
[0018] The batch normalization layer is used to normalize the low-level features;
[0019] The maximum pooling layer is used to retain the most important feature information by selecting the maximum value in the local area;
[0020] The residual block is used for deep feature extraction.
[0021] Optionally, the residual block includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a batch normalization layer and a ReLU activation layer;
[0022] The first convolutional layer is used to reduce the dimension through a 1×1 convolution operation, thereby reducing the number of channels of the feature map and reducing computational complexity;
[0023] The second convolutional layer is used to extract local features through a 3×3 convolution operation while preserving the spatial information of the feature map;
[0024] The third convolutional layer is used to increase the dimension through a 1×1 convolution operation, restore the number of channels of the feature map, and enhance the expressiveness of the features.
[0025] Optionally, the first convolutional layer and the third convolutional layer are jump-connected, and each convolutional layer is connected to a batch normalization layer and a ReLU activation layer.
[0026] Optionally, fusing the feature maps includes:
[0027] Rearranging the dimensions of the feature graph and calculating a similarity score matrix of the feature graph using matrix multiplication;
[0028] Based on the similarity score matrix, the Softmax function is used to obtain the attention weight;
[0029] The feature maps are weighted combined using the attention weights to obtain a fused feature map. Optionally, coarse-grained classification is performed on the fused feature map to obtain the main stage probability, including:
[0030] P m =Softmax(GlobalAveragePooling(F))
[0031] Among them, F is the fused feature map, P m is the main stage probability, GlobalAveragePooling (F) represents the global average pooling operation on the feature map, and the Softmax function is used to convert the output into a probability distribution.
[0032] Optionally, fusing the feature maps after the coarse-grained classification includes:
[0033] Rearranging the dimensions of the feature graph after the coarse-grained classification, and calculating the similarity score matrix of the feature graph using matrix multiplication;
[0034] Based on the similarity score matrix, the Softmax function is used to obtain the attention weight;
[0035] The feature maps are weighted combined using the attention weights to obtain a fused feature map. Optionally, fine-grained classification is performed on the fused feature map to obtain sub-stage probabilities, including:
[0036] P sub =Softmax(Classifier(F sub ))
[0037] Among them, P sub is the sub-stage probability, Classifier(F sub ) represents a fine-grained classification operation on the feature map, and the Softmax function is used to convert the output into a probability distribution.
[0038] Compared with the prior art, the present invention has the following advantages and technical effects:
[0039] 1. By introducing convolutional neural networks (CNNs) for multi-stage feature extraction, the present invention can capture subtle changes and complex patterns during embryonic development, thereby significantly improving classification accuracy. Compared with traditional methods that rely on limited manually defined features, deep learning models can automatically learn more discriminative features from large amounts of data.
[0040] 2. This invention utilizes a cross-attention mechanism to dynamically adjust the focus on key areas at different levels, further enhancing the ability to capture important features. This not only reduces the possibility of misclassification but also improves the model's ability to handle abnormal or atypical samples. Cross-attention not only helps capture correlations between focal length data but also promotes information exchange between different features, thereby improving the overall model's effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0042] Figure 1 This is a flow chart of a method for hierarchical classification of embryonic development stages based on a convolutional neural network according to an embodiment of the present invention;
[0043] Figure 2 This is a schematic diagram of coarse-grained division based on embryonic development stages according to an embodiment of the present invention;
[0044] Figure 3 is an overall structural diagram of the classification method according to an embodiment of the present invention;
[0045] Figure 4 This is a diagram of the ResNet50 module feature extraction network structure of an embodiment of the present invention;
[0046] Figure 5 This is the overall structure diagram of the multiplication of coarse and fine granularity category probabilities in an embodiment of the present invention;
[0047] Figure 6 This is a structural diagram of the cross-attention feature fusion module of an embodiment of the present invention. DETAILED DESCRIPTION
[0048] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0049] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0050] This embodiment proposes a hierarchical classification method for embryonic development stages based on convolutional neural networks, such as Figure 1 As shown, the specific steps include:
[0051] Acquire embryo images;
[0052] Obtain a feature map based on the embryo image;
[0053] Fuse the feature maps and perform coarse-grained classification to obtain the main stage probability;
[0054] The feature maps after coarse-grained classification are fused and fine-grained classification is performed to obtain sub-stage probabilities;
[0055] Obtain category probabilities based on the main stage probabilities and sub-stage probabilities;
[0056] Embryonic development stages are hierarchically classified based on class probabilities.
[0057] Specifically, the following steps are included:
[0058] (1) Acquire and preprocess embryo images at three different focal lengths;
[0059] (2) Input the preprocessed image into the constructed ResNet50 model and perform feature extraction using the first three stages;
[0060] (3) Applying a cross-attention mechanism to the feature maps of the same scale obtained at different focal lengths to dynamically adjust the attention to the key areas;
[0061] (4) Use global average pooling and fully connected layers to perform main stage classification on the fused feature map to obtain the main stage probability;
[0062] (5) Continue to extract features from the fused feature map and apply the cross attention mechanism again;
[0063] (6) Use global average pooling and fully connected layers to perform sub-stage classification on the further extracted feature maps to obtain sub-stage probabilities;
[0064] (7) According to the result of multiplying the main stage probability and the sub-stage probability, the maximum category probability of embryonic development is determined and output.
[0065] This method uses CNN to automatically learn and capture subtle changes and complex patterns in embryonic development by introducing multi-stage feature extraction and a cross-attention mechanism, thereby achieving accurate classification of embryonic development stages. Specifically, this embodiment first receives or generates three embryo images of different focal lengths, then uses the ResNet50 model for preliminary feature extraction, and applies these feature maps to the cross-attention module to enhance feature representation. Subsequently, the fused feature maps are used for classification tasks of the main stage and sub-stage, and ultimately outputs the maximum category probability of embryonic development. This method not only improves classification accuracy but also reduces the possibility of misclassification, making it suitable for assisted reproductive technology and embryology research.
[0066] Furthermore, obtaining the embryo image includes:
[0067] Acquire initial embryo images;
[0068] The initial embryo image is preprocessed to obtain an embryo image.
[0069] Specifically, in step (1), preprocessing includes but is not limited to adjusting the image size to a grayscale image of 224x224.
[0070] Furthermore, obtaining a feature map based on the embryo image includes:
[0071] The embryo image is input into the pre-trained ResNet50 model to obtain the feature map, where the ResNet50 model includes: initial convolution layer, batch normalization layer, maximum pooling layer and residual block;
[0072] Initial convolutional layer, used to extract low-level features;
[0073] Batch normalization layer, used to normalize low-level features;
[0074] The maximum pooling layer is used to retain the most important feature information by selecting the maximum value in the local area;
[0075] Residual block, used for deep feature extraction.
[0076] Specifically, in step (2), the ResNet50 model includes four stages, each stage consisting of multiple residual blocks, but only the first three stages are used in this embodiment.
[0077] Furthermore, the residual block includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a batch normalization layer, and a ReLU activation layer;
[0078] The first convolutional layer is used to reduce the dimension through 1×1 convolution operation, reducing the number of channels of the feature map, thereby reducing the computational complexity;
[0079] The second convolutional layer is used to extract local features through 3×3 convolution operations while preserving the spatial information of the feature map;
[0080] The third convolutional layer is used to increase the dimension through 1×1 convolution operation, restore the number of channels of the feature map, and enhance the expressiveness of the features.
[0081] Furthermore, the first convolutional layer and the third convolutional layer are skip-connected, and each convolutional layer is connected to a batch normalization layer and a ReLU activation layer.
[0082] Furthermore, the feature map fusion includes:
[0083] Rearrange the dimensions of the feature map and use matrix multiplication to calculate the similarity score matrix of the feature map;
[0084] Based on the similarity score matrix, the Softmax function is used to obtain the attention weight;
[0085] The feature maps are weightedly combined using the attention weights to obtain the fused feature maps.
[0086] Furthermore, the fused feature map is subjected to coarse-grained classification to obtain the main stage probability, including:
[0087] The fused feature map is globally average pooled to obtain the main stage probability.
[0088] Specifically, the coarse-grained classification of the fused feature map and the acquisition of the main stage probability include:
[0089] P m =Softmax(GlobalAveragePooling(F))
[0090] Among them, F is the fused feature map, P m is the main stage probability, GlobalAveragePooling (F) represents the global average pooling operation on the feature map, and the Softmax function is used to convert the output into a probability distribution.
[0091] More specifically, in step (3), the cross-attention mechanism calculates the similarity score matrix between each feature map through matrix multiplication, applies the Softmax function to the score matrix to obtain the attention weights, and then applies these weights to the original feature map to generate the weighted feature representation;
[0092] In steps (4) and (6), global average pooling is used to reduce the spatial dimension of the feature map, and the fully connected layer is used to output the probability distribution of the classification results.
[0093] Furthermore, the fusion of the feature maps after coarse-grained classification includes:
[0094] The dimensions of the feature maps after coarse-grained classification are rearranged, and the similarity score matrix of the feature maps is calculated using matrix multiplication;
[0095] Based on the similarity score matrix, the Softmax function is used to obtain the attention weight;
[0096] The feature maps are weightedly combined using the attention weights to obtain the fused feature maps.
[0097] Specifically, 1. Rearrange the dimension of the feature map: rearrange the dimension of the feature map F after coarse-grained classification and express it in the matrix form M.
[0098] 2. Calculate the similarity score matrix: Calculate the similarity score matrix S of the feature graph through matrix multiplication, and its formula is:
[0099] S=M·M T
[0100] Among them, M T is the transpose of the matrix M.
[0101] Obtaining attention weight: Based on the similarity score matrix S, the Softmax function is used to calculate the attention weight A, and its formula is:
[0102] A=Softmax(S)
[0103] Weighted combination feature map: Use the attention weight A to perform weighted combination on the feature map F to obtain the fused feature map F sub , the formula is:
[0104] F sub =A·F
[0105] Furthermore, fine-grained classification is performed on the fused feature map to obtain sub-stage probabilities, including:
[0106] P sub =Softmax(Classifier(F sub ))
[0107] Among them, P sub is the sub-stage probability, Classifier(F sub ) represents a fine-grained classification operation on the feature map, and the Softmax function is used to convert the output into a probability distribution.
[0108] The present embodiment will be described in detail below with reference to the accompanying drawings:
[0109] This embodiment proposes a convolutional neural network hierarchical classification method for embryonic development stages, such as Figure 3 As shown, the following steps are included:
[0110] (1) Image input and preprocessing:
[0111] (2) The specific calculation process of step (1) is as follows: obtain and preprocess three embryo images I1, I2, and I3 with different focal lengths, and uniformly adjust the size of each image to a grayscale image of 224×224 pixels to ensure the consistency of subsequent processing. Obtain and preprocess three embryo images I1, I2, and I3 with different focal lengths, and uniformly adjust the size of each image to a grayscale image of 224×224 pixels to ensure the consistency of subsequent processing.
[0112] (2) Multi-focal image reception and storage:
[0113] The specific calculation process of step (2) is as follows: the system receives or generates three embryo images I1, I2, and I3 corresponding to different focal lengths, and stores the data of these images and their associated information in a multidimensional array A, while assigning corresponding main classification labels and sub-classification labels to each image.
[0114] like Figure 4 As shown, step (3) uses the pre-trained ResNet50 model to extract features from the multidimensional array A, but only uses the first three stages (Stage1-3). The specific parameters are as follows;
[0115] The specific calculation process of step (3) is as follows:
[0116] The preprocessed image data A is input into the constructed ResNet50 model. ResNet50 is a deep convolutional neural network that includes four stages, each consisting of multiple residual blocks. However, in the implementation process of this embodiment, only the first three stages of ResNet50 are used for feature extraction.
[0117] First, the image data for each focal length is fed as a separate channel into an initial convolutional layer (Conv2d) for preliminary feature extraction. The parameters for this convolutional layer are: kernel size 7x7, stride 2, padding 3, 3 input channels (one channel per focal length), and 64 output channels. The primary purpose of this convolutional layer is to extract low-level features, such as edges and texture, from the original image and reduce the spatial size of the image, paving the way for more complex feature extraction. The resulting feature map then passes through a batch normalization layer and a Reluctant Unit (ReLU) activation layer. The batch normalization layer normalizes the input of each layer, ensuring that each batch of data has zero mean and unit variance, accelerating training and improving model stability. The Reluctant Unit (ReLU) activation function introduces nonlinearity, enabling the network to learn more complex mappings. These operations help stabilize and accelerate the overall training process while ensuring model effectiveness and generalization. Next, the feature map enters a max pooling layer with the following parameters: kernel size 3x3, stride 2, and padding 1. The max pooling layer retains the most important feature information by selecting the maximum value within a local area, while significantly reducing the spatial dimension of the feature map, reducing computational complexity, and enhancing the model's robustness to translation and small deformations. The feature map is then fed into the residual blocks in the first three stages of ResNet50 for deep feature extraction. Each stage contains a different number of residual blocks. Each residual block includes the following components:
[0118] The first convolutional layer: The convolution kernel size is 1x1, the stride is determined by the location, usually 1 or 2, used to adjust the spatial size, padded with 0, and the number of output channels is determined by the network design.
[0119] The second convolutional layer: The convolution kernel size is 3x3, the stride is 1, the padding is 1, and the number of output channels is determined according to the network design.
[0120] The third convolutional layer: The convolution kernel size is 1x1, the stride is 1, the padding is 0, and the number of output channels is the same as the number of input channels.
[0121] Each convolutional layer is followed by a batch normalization layer and a ReLU activation layer to stabilize the training process and introduce nonlinearity. When necessary, a 1x1 convolution is used to adjust the size difference between the input and the output of the third convolutional layer to ensure that they can be added together. This forms part of the skip connection. The skip connection adds the input directly to the output of the third convolutional layer. If the size does not match, the number of channels and spatial size are adjusted through a 1x1 convolution to ensure the effectiveness of the skip connection.
[0122] like Figure 2 As shown, step (4) applies the obtained feature maps of the same scale at different focal lengths to cross attention;
[0123] The specific calculation process of step (4) is as follows: First, the feature maps of the three focal lengths obtained in step (3) are applied to the first layer of cross attention, as shown in Figure (4). The cross attention is as follows: Figure 6 As shown in Figure 2, cross attention is a mechanism for capturing the relationship between different feature maps. For the I1 feature map, we first rearrange its dimensions so that the channel dimension is moved to the end to facilitate matrix multiplication. Then, the similarity score matrix A between each feature map is calculated by matrix multiplication. i , and apply the Softmax function to the score matrix to obtain the attention weights. After that, these weights are applied to the original feature map I again using matrix multiplication i On the top, generate weighted feature representation Finally, the three weighted feature maps are linearly combined according to the predefined weight vector to obtain the final fused feature map X1.
[0124] Step (5) performs main stage classification on the acquired fused feature map to obtain the main stage probability;
[0125] The specific calculation process of step (5) is as follows: first, the fused feature map X1 obtained in step (4) is globally average pooled, and then the full connection layer is used to output the main stage classification probability.
[0126] Step (6) continues to extract and apply the fusion to cross attention to continue extracting features;
[0127] The specific calculation process of step (6) is the same as that of step (4).
[0128] Step (7) performs sub-stage classification on the acquired fused feature map to obtain the sub-stage probability;
[0129] The specific calculation process of step (7) is the same as that of step (5).
[0130] Step (8) obtains the maximum class probability and outputs the embryonic stage;
[0131] like Figure 5 As shown, step (8) multiplies the main stage probability obtained in step (6) and the sub-stage probability obtained in step (7) to obtain the embryonic stage by taking the maximum probability of the category;
[0132] This example provides an efficient and accurate hierarchical classification method for embryonic development stages. It not only overcomes the limitations of traditional manual assessment methods but also demonstrates significant potential for improving IVF success rates. By introducing deep learning techniques and a cross-attention mechanism, this example represents a new breakthrough in assisted reproductive technology and is expected to play a significant role in future research and clinical applications.
[0133] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A hierarchical classification method for embryonic development stages based on convolutional neural networks, characterized in that: include: Acquire embryo images; acquiring a feature map according to the embryo image; The feature maps are fused and coarse-grained classified to obtain the main stage probability; Fusing the feature maps after the coarse-grained classification and performing fine-grained classification to obtain sub-stage probabilities; Obtaining category probabilities based on the main stage probabilities and the sub-stage probabilities; The embryonic development stages are hierarchically classified according to the category probabilities.
2. The method for hierarchical classification of embryonic development stages based on convolutional neural networks according to claim 1, wherein: Acquiring the embryo image includes: Acquire initial embryo images; The initial embryo image is preprocessed to obtain the embryo image.
3. The method for hierarchical classification of embryonic development stages based on convolutional neural networks according to claim 1, wherein: Acquiring a feature map according to the embryo image includes: Inputting the embryo image into a pre-trained ResNet50 model to obtain the feature map, wherein the ResNet50 model includes: an initial convolutional layer, a batch normalization layer, a maximum pooling layer, and a residual block; The initial convolutional layer is used to extract low-level features; The batch normalization layer is used to normalize the low-level features; The maximum pooling layer is used to retain the most important feature information by selecting the maximum value in the local area; The residual block is used for deep feature extraction.
4. The method for hierarchical classification of embryonic development stages based on convolutional neural networks according to claim 3, wherein: The residual block includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a batch normalization layer and a ReLU activation layer; The first convolutional layer is used to reduce the dimension through a 1×1 convolution operation, thereby reducing the number of channels of the feature map and reducing computational complexity; The second convolutional layer is used to extract local features through a 3×3 convolution operation while preserving the spatial information of the feature map; The third convolutional layer is used to increase the dimension through a 1×1 convolution operation, restore the number of channels of the feature map, and enhance the expressiveness of the features.
5. The method for hierarchical classification of embryonic development stages based on convolutional neural networks according to claim 4, wherein: The first convolutional layer and the third convolutional layer are jump-connected, and each convolutional layer is connected to a batch normalization layer and a ReLU activation layer.
6. The method for hierarchical classification of embryonic development stages based on convolutional neural networks according to claim 1, wherein: Fusing the feature maps includes: Rearranging the dimensions of the feature graph and calculating a similarity score matrix of the feature graph using matrix multiplication; Based on the similarity score matrix, the Softmax function is used to obtain the attention weight; The feature maps are weightedly combined using the attention weights to obtain a fused feature map.
7. The method for hierarchical classification of embryonic development stages based on convolutional neural networks according to claim 6, wherein: Perform coarse-grained classification on the fused feature map to obtain the main stage probability, including: P m =Softmax(GlobalAveragePooling(F)) Among them, F is the fused feature map, P m is the main stage probability, GlobalAveragePooling (F) represents the global average pooling operation on the feature map, and the Softmax function is used to convert the output into a probability distribution.
8. The method for hierarchical classification of embryonic development stages based on convolutional neural networks according to claim 7, wherein: Fusing the feature maps after the coarse-grained classification includes: Rearranging the dimensions of the feature graph after the coarse-grained classification, and calculating the similarity score matrix of the feature graph using matrix multiplication; Based on the similarity score matrix, the Softmax function is used to obtain the attention weight; The feature maps are weightedly combined using the attention weights to obtain a fused feature map.
9. The method for hierarchical classification of embryonic development stages based on convolutional neural networks according to claim 8, wherein: Perform fine-grained classification on the fused feature map to obtain sub-stage probabilities, including: P sub =Softmax(Classifier(F sub )) Among them, P sub is the sub-stage probability, Classifier(F sub ) represents a fine-grained classification operation on the feature map, and the Softmax function is used to convert the output into a probability distribution.
Citation Information
Patent Citations
Bone mass recognition method and system based on three-dimensional convolutional neural network and CT image
CN116416428A
Embryo development stage prediction and quality evaluation system based on rotation equivariant network
CN116883996A
Attention-based double-branch fine-grained network alien invasive plant identification method
CN117636151A
Deep learning method and system for embryonic development multi-stage classification
CN119131507A
Global and local feature reconstruction network-based medical image segmentation method
US20230274531A1
Cited By
Improved multi-network fusion embryo image classification method
CN121564382A