Biological Tissue Processing Method Based on Semantic Consistency Network
By adopting the I2C2Net framework in blastocyst tissue image segmentation and combining intra-class and inter-class context modules, the problem of insufficient image segmentation performance of blastocyst tissue in the prior art is solved, and higher segmentation accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510218521.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The prior art has problems with insufficient recognition performance in blastocyst tissue image segmentation, especially in the treatment of inconspicuous inter-class structure and edge areas.
The I2C2Net framework based on semantic consistency network is adopted to capture the intra-class and inter-class features of blastocyst tissue through the combination of intra-class context module (IACCM) and inter-class context module (IRCCM), and to improve the learning and expression capabilities of the model using consistency module.
The performance of blastocyst tissue image segmentation is significantly improved, the aggregation of intra-class features and the capture of inter-class relationships is enhanced, and the accuracy and robustness of segmentation are improved.
Smart Images

Figure CN119723576B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for biological tissue recognition and processing based on a context semantic consistency network, specifically to an intra-class and inter-class context and consistency network for supervised and semi-supervised blastocyst image segmentation, which is based on the I2C2Net framework and belongs to the technical field of image pattern recognition based on artificial neural networks and the field of artificial intelligence technology applied to medicine. Background Art
[0002] In recent years, the progress of deep learning technology has continuously improved the performance of medical image segmentation. The deep learning network commonly used for medical image segmentation is U-Net [3] and its derivative variants. Some studies have attempted to segment blastocyst tissue images, but the segmentation performance still needs to be further improved. In the blastocyst tissue images obtained by photographing, there are significant differences in the intra-class features and adhesiveness of various tissue categories, such as Figure 1 shown in a), b) and c) below, where a) shows the original photo image of the blastocyst tissue, b) shows the details of the inner cell mass (ICM) region of the original photo image of the blastocyst tissue, and c) shows the details of the inter-class region between the inner cell mass (ICM) and other category tissues of the original photo image of the blastocyst tissue. Due to the transparency characteristics of the blastocyst itself and the fact that blastocyst tissue images are generally collected under low light conditions, the intra-class features of the blastocyst tissue are blurred and the edge features are difficult to distinguish; as Figure 1 shown in c) below, the differences in the inter-class structures of different tissue types in the blastocyst tissue image are not obvious, especially in the edge regions of the tissue. In the results of traditional blastocyst segmentation of the blastocyst using U-Net, no inter-class clustering was observed, such as Figure 1 shown in d), e) and f) below. In particular, some regions of the inner cell mass (Inner Cell Mass, ICM) were not recognized, some regions of the blastocoel were misidentified as trophoblast, and the inter-class structure was not prominent, such as Figure 1 shown in f) below, and the errors were mainly concentrated on the boundaries and some difficult-to-process pixels.
[0003] In recent years, some works have attempted to segment blastocyst images, mainly focusing on identifying the inner cell mass (ICM) and trophoblast (trophectoderm, TE). For example, using the level set method, which defines an initial contour and applies gradient information to segment the inner cell mass and trophoblast; however, in clinical practice, the zona pellucida (ZP) of the blastocyst may rupture, and at this time the inner cell mass is outside the zona pellucida, so the recognition effect is not significant. Rad et al. proposed using a stacked dilated U-Net model for accurate segmentation of the inner cell mass, and proposed a U-Net variant by combining dilated convolution, which only increased the receptive field of a single pixel but did not change the attribute relationship between pixels, and the Jaccard index was lower than 82%.
[0004] The methods in the prior art are either used for single tissue segmentation or the recognition of several tissues. Therefore, methods for improving the performance of blastocyst segmentation still need to be further developed, and the overall recognition of single tissues and the structural associations of all tissues should be concerned. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a method for processing biological tissues based on a semantic consistency network, including the steps of:
[0006] Construct a network model for modeling intra-class and inter-class features of tissues based on semantic consistency. The network model includes N encoding blocks, N - 1 I2C modules, N compression blocks, and N logic blocks; wherein, each encoding block, compression block, and logic block is a convolutional module based on a convolutional neural network, and each I2C module includes an intra-class context module for aggregating pixel representations within a specific category region and an inter-class context module capable of capturing the morphological structures between tissues; and, the N encoding blocks are connected in series in sequence, the output of the Nth encoding block is input to the corresponding Nth compression block for processing, and the result of the Nth compression block is input to the Nth logic block for processing and then upsampled and input to the (N - 1)th I2C module. The (N - 1)th I2C module receives the output of the (N - 1)th encoding block and the upsampled output processed by the above Nth logic block, outputs the processed result to the (N - 1)th logic block for processing and then upsampled and input to the (N - 2)th functional block, and so on, until the output after the processing of the 1st logic block is output as the prediction result; N is an integer greater than 2;
[0007] Moreover, the network model for modeling intra-class and inter-class features of tissues based on semantic consistency further includes a consistency module for training the network model;
[0008] Obtain training sample data, which is biological tissue image data. A part of the biological tissue image data is labeled data for supervised training, and another part is unlabeled data for semi-supervised training;
[0009] Use the training sample data to train the network model for modeling intra-class and inter-class features of tissues based on semantic consistency to obtain a trained model for biological tissue recognition and processing;
[0010] Use the trained model for biological tissue recognition and processing to recognize and process biological tissue image data.
[0011] In the above technical solution, N is 5.
[0012] In the above technical solution, each encoding block includes two convolutional network layers, and each convolutional network layer includes two batch normalization layers, two activation layers, and one max pooling layer.
[0013] In the above technical solution, the encoding block is defined as , and the encoding block E projects the pixels in each input biological tissue image X into a non-linear embedding space to obtain a pixel representation matrix :
[0014]
[0015] where has a size of , C is the number of channels of ; X is the input biological tissue image; W and H are the width and height of the biological tissue image X respectively.
[0016] In the above technical solution, the compression block is defined as , and the compression block F compresses the pixel representation matrix into :
[0017]
[0018] where represents the features obtained by aggregating intra-class and inter-class information from different scales of the entire image using the I2C module, has a size of , is the number of categories.
[0019] In the above technical solution, the logic block is defined as , and the logic block L is used to obtain the supervised logic graph of the biological tissue image X:
[0020]
[0021] where the logic block L consists of a convolutional layer, a batch normalization layer, and an activation layer.
[0022] In the above technical solution, the upsampling function for obtaining the probability map from the output of the logic block is .
[0023] In the above technical solution, the consistency module includes a first convolutional block, a second convolutional block, a first self-attention block, a third convolutional block, a second self-attention block, and a convolutional output layer connected in sequence; among them, the first convolutional block, the second convolutional block, and the third convolutional block have the same structure, and each includes a convolutional layer, a spectral normalization layer, and an activation function layer.
[0024] The present invention also provides a biological tissue processing device based on a semantic consistency network, including:
[0025] A model construction module, configured to construct a network model for modeling intra-class and inter-class features of tissues based on semantic consistency. The network model includes N encoding blocks, N - 1 I2C modules, N compression blocks, and N logic blocks; wherein, each encoding block, compression block, and logic block is a convolutional module based on a convolutional neural network, and each I2C module includes an intra-class context module for aggregating pixel representations within a specific category region and an inter-class context module capable of capturing the morphological structure between tissues; and, the N encoding blocks are connected in series in sequence, the output of the Nth encoding block is input to the corresponding Nth compression block for processing, and the result of the Nth compression block is input to the Nth logic block. After being processed, it is upsampled and input to the (N - 1)th I2C module. The (N - 1)th I2C module receives the output of the (N - 1)th encoding block and the upsampled output processed by the above-mentioned Nth logic block, outputs the processed result to the (N - 1)th logic block, and after being processed, it is upsampled and input to the (N - 2)th functional block, and so on, until the output after being processed by the 1st logic block is output as a prediction result; N is an integer greater than 2;
[0026] And, the network model for modeling intra-class and inter-class features of tissues based on semantic consistency further includes a consistency module for training the network model;
[0027] A training sample acquisition module, configured to acquire training sample data, where the training sample data is biological tissue image data, and a part of the biological tissue image data is labeled data for supervised training, and another part is unlabeled data for semi-supervised training;
[0028] A model training module, using the training sample data to train the network model for modeling intra-class and inter-class features of tissues based on semantic consistency, to obtain a biological tissue recognition and processing training model;
[0029] A data processing module, using the biological tissue recognition and processing training model to recognize and process biological tissue image data.
[0030] The present invention also provides a biological tissue processing system based on a semantic consistency network, including a processor and a storage medium; program code is stored on the storage medium; the processor is configured to call the program code stored in the storage medium to execute the biological tissue processing method based on the semantic consistency network in the above technical solution.
[0031] The present invention has achieved the following technical effects: The present invention proposes a general framework I2C2Net, which continuously improves the performance of biological tissue image segmentation such as blastocysts by utilizing intra-class and inter-class information; it also proposes a semi-supervised version of I2C2Net based on semantic consistency hypothesis and clustering hypothesis, which utilizes labeled and unlabeled data in two different domains without structural changes; the present invention designs a simple and effective IACCM to enhance the aggregation within the same category, and a novel IRCCM to capture the relationship information between different categories; in addition, the present invention also uses a consistency module to improve the learning and expression ability of the basic model without increasing additional computational costs; the present invention adopts a top-down feature enhancement path, coupling the backbone encoder to improve the representation ability of the feature map in a coarse-to-fine manner through intra-class enhancement and inter-class enhancement. Qualitative and quantitative experimental results prove the effectiveness of the method proposed by us. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a diagram of a blastocyst tissue image and its related processing information; where a) is the original photo image of the blastocyst tissue, b) is the detail of the inner cell mass (ICM) region of the original photo image of the blastocyst tissue, c) is the detail of the inter-class region between the inner cell mass (ICM) and other category tissues of the original photo image of the blastocyst tissue, d) is the true label of different tissues of the blastocyst tissue obtained by manual processing, e) is the prediction map obtained using U-Net, and f) is the difference between e) and d);
[0033] Figure 2 It is a schematic diagram of the I2C2Net network structure;
[0034] Figure 3 It is a schematic diagram of the network structure of the intra-class context module (IACCM);
[0035] Figure 4 It is a schematic diagram of the network structure of the inter-class context module (IRCCM);
[0036] Figure 5 It is a schematic diagram of the network structure of the consistency module (CM);
[0037] Figure 6 It is a schematic diagram of the network structure of the semi-supervised version of I2C2Net. DETAILED DESCRIPTION OF THE INVENTION
[0038] In order to facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0039] To solve the above problems, the present invention proposes an I2C2Net network architecture to effectively model intra-class and inter-class features for the segmentation task of blastocyst tissue images. The I2C2Net network architecture consists of an Intra-Class Context Module (IACCM) and an Inter-Class Context Module (IRCCM); optionally, the I2C2Net network architecture further includes a Consistency Module (CM) for training.
[0040] The Intra-Class Context Module (IACCM) focuses on aggregating pixel representations within specific category regions, strengthening the learning of homogeneous features, and enhancing the model's ability to recognize classification regions related to true labels. By aggregating intra-class features, the Intra-Class Context Module (IACCM) can simplify the multi-class recognition task into multiple binary classification tasks while retaining the ability to learn intra-class features, thereby being able to adaptively enhance intra-class features and suppress other types of features, enhancing the degree of aggregation within the class, making the segmentation results more consistent, and improving the accuracy and efficiency of segmentation.
[0041] In addition, the Intra-Class Context Module (IACCM) also simplifies a K-class recognition task into K binary classification tasks, reducing the complexity of model learning. However, since the Intra-Class Context Module (IACCM) is more likely to learn simple samples, a large number of simple samples dominate the gradient, resulting in challenges for the model in recognizing difficult pixels and weight values. In particular, pixels in the edge region are difficult to recognize due to their small proportion and unclear features. In addition, although the decomposition task reduces the complexity of model learning, the Intra-Class Context Module (IACCM) does not consider the structural correlation between inter-class samples. Therefore, the present invention proposes an Inter-Class Context Module (IRCCM) to solve the problems of the Intra-Class Context Module (IACCM) in boundary conditions and difficult pixel processing through a weighted mapping function. The Inter-Class Context Module (IRCCM) is designed for the morphological structure reflected in the blastocyst tissue image, and can capture the information changes between various cell tissues during the development of the blastocyst tissue from the inner layer to the outer layer, further enhancing the model's comprehensive understanding and segmentation ability of the blastocyst tissue. The Inter-Class Context Module (IRCCM) also makes up for the deficiencies of the Intra-Class Context Module (IACCM) through the interaction between different classes.
[0042] The Intra-class Context Module (IACCM) and the Inter-class Context Module (IRCCM) together constitute a network architecture (abbreviated as I2C2Net) that can effectively model intra-class and inter-class features of tissues, and is responsible for generating segmentation results of biological tissue images (especially blastocyst tissue images). These two modules are applied at each level of the network, and through a top-down feature enhancement path, continuously provide coarse-to-fine supervision information. During this process, sufficient interaction occurs between multi-scale feature information. At the high level of the network, the topological structures of different classes are integrated based on semantic information; while at the low level of the network, the image edges are refined according to low-level statistical features such as color, edge, and texture. The intra-class and inter-class interactions enable the model based on the I2C2Net network to learn the potential features of the embryo, thereby improving the accuracy and robustness of blastocyst segmentation.
[0043] In order to improve the accuracy of embryo tissue recognition without increasing additional parameters and the computational burden of inference, the present invention focuses on optimizing the model's ability to handle data distribution differences. These differences, especially class boundaries and distribution characteristics, are crucial for the quality assessment of image segmentation. Therefore, the present invention designs an innovative Consistency Module (CM) for supervised training, which organically integrates the original embryo image and the segmentation result, and improves the model's ability to learn the original data distribution, thereby improving the recognition accuracy without increasing the computational burden. During the training process of I2C2Net, CM plays a crucial role, which promotes the improvement of the distribution consistency between the model prediction result and the true label. More importantly, in order to address the challenge of scarce labeled data and better serve clinical needs, the present invention proposes a semi-supervised version (Semi-supervised learning, SSL) of I2C2Net based on the semantic consistency hypothesis and the clustering hypothesis across heterogeneous domains, providing an effective way to solve the above problems. The above method can more efficiently capture the data distribution differences between embryo tissue classes and improve the accuracy of embryo segmentation.
[0044] The network (I2C2Net) framework provided by the present invention that can effectively model intra-class and inter-class features of tissues is as Figure 2 shown. The I2C2Net network includes N encoding blocks (as Figure 2 shown in ), the I2C module (such as IACCM, IACCM shown in Figure 2 ), the compression block (as Figure 2 shown in ), the logic block (such as Figure 2 shown in as shown); each I2C module includes an I2C framework composed of an intra-class context module IACCM and an inter-class context module IRCCM. The blastocyst image data X is sequentially input into N encoding blocks in a serial manner , where i ranges from 1 to N; among them, the output of the Nth encoding block is input into the corresponding Nth compression block for processing, and the result of the Nth compression block is input into the Nth logic block After processing, it is upsampled and input into the (N - 1)th I2C module. The (N - 1)th I2C module receives the output of the (N - 1)th encoding block and the upsampled output after processing of the Nth logic block After processing, it is output to the (N - 1)th logic block After processing, it is upsampled and input into the (N - 2)th functional block, and so on, until the 1st logic block After processing, it is output as the prediction result. Among them, the encoding block , the compression block , and the logic block are all convolutional modules.
[0045] During the training and working (testing) process, the I2C2Net framework uses the intra-class and inter-class context information extracted by the intra-class context module (IACCM) and the inter-class context module (IRCCM) to segment biological tissues (such as blastocyst images). The consistency module (CM) only participates in the training phase, so no additional computational workload is added during the working (testing) phase.
[0046] For the training and working (testing) of I2C2Net, a set of images S can be considered. Let , where represents the original input embryo image, and represents the corresponding ground truth obtained through manual processing. Since each input image is processed separately, for simplicity, the subscript m is omitted in the specific processing flow of each embryo image in S.
[0047] Define the encoding block as , where n represents the nth encoding block. Each consists of two convolutional layers, and each convolutional layer has two batch normalization (BN) layers, two ReLU activation layers, and one max pooling layer. I2C2Net first projects the pixels in each input embryo tissue image X into a non-linear embedding space using the encoding block E to obtain the pixel representation matrix , and then projects it N - 1 times in nested order:
[0048] (1)
[0049] Among them, The size of is , C is the number of channels of ; n is the number of nested projections; X is the input embryonic tissue image; W and H are the width and height of the embryonic tissue image X respectively; N is the number of encoding blocks E, that is, the total number of times of nested projection into the non-linear embedding space. In the present invention, preferably N = 5. The convolution module can be composed of convolution kernels with different kernel sizes, which plays a role in feature mapping.
[0050] The present invention uses a compression block to compress the pixel representation matrix into , which has K channels (K is the number of categories). The process is as follows:
[0051] (2)
[0052] Among them, represents the feature obtained by aggregating intra-class and inter-class information from different scales of the entire image using the I2C framework, The size of is
[0053] The I2C framework is composed of an intra-class context model and an inter-class context model . The intra-class context model and the inter-class context model simultaneously process the high-level features output by the encoding block with the corresponding serial number and the low-level features output by the logic block with the previous serial number.
[0054] After compressing the channels of the representation through the compression block , a supervised logic map of the image X is obtained using the logic block , and its size is .
[0055] (3)
[0056] Among them, the logic block is composed of a convolutional layer, a BN layer, and a ReLU layer. Therefore, before obtaining the aggregated information of the intra-class and inter-class modules, the probability map can be determined from the low-level logic map , as shown in Figure 2 .
[0057] To match the high-level feature scale, the present invention uses an upsampling function to obtain a probability map of size from the output of the logic block . .
[0058] (4)
[0059] where is the sigmoid function.
[0060] The upsampling function is specifically .
[0061] Then, the probability map is used for the I2C module to enhance features :
[0062] (5)
[0063] (6)
[0064] (7)
[0065] Finally, the prediction output maps of different scales are represented as:
[0066] (8)
[0067] For the intra-class context module (IACCM):
[0068] In the I2C module, the IACCM aims to capture the context information of the same class in the image. As Figure 3 shown, feature enhancement can be achieved between the same scale and different scales. Before the intra-class scale enhancement, each layer of the logic map passes through a supervised module before fusion. Then, these logic maps are activated together with adjacent high-level features that focus on the context information of the same class.
[0069] To match the adjacent high-level features with the probability map organized by K classes, the present invention orderly divides into K groups. is a matrix of size where represents the number of channels of class k. Accordingly, the number of parameters can be effectively reduced.
[0070] (9)
[0071] where, Denotes the floor operation.
[0072] For the convenience of representation, define as the logical graph under category k. To aggregate the context information of each same category, calculate the representation area of each category k as follows:
[0073] (10)
[0074] Finally, orderly connect all and use to obtain features, which contain the context information of the same category in the image. Through grouping, the present invention decomposes the multi-category recognition task into multiple two-category recognition tasks, thereby reducing the learning difficulty of the model and allowing the model to focus on learning the features of the same category without being interfered by other categories. And, to a certain extent, the number of parameters is reduced.
[0075] For the inter-category context module (IRCCM):
[0076] Although IACCM effectively solves the problems of feature representation and learning within the same category, it does not consider the inter-category relationship. As Figure 1 shown in f), the proportion of prediction errors between different categories in the image is relatively large, including some pixels with unclear structural features. To solve this problem, the present invention designs IRCCM , which can process the relationship information between different categories and some difficult pixels that are difficult to identify, as Figure 4 shown.
[0077] IACCM and IRCCM have the same high-level features and follow the same grouping principle. Intra-category grouping considers the number of organized categories, while inter-category grouping considers the correlation between organizations, which is represented by K - 1 groups. Taking a typical blastocyst tissue image as an example, generally speaking, from the inside to the outside, it is: inner cell mass, blastocoel, trophoblast, zona pellucida, and background. Therefore, generally only consider four tissue relationships, namely: 1, 1 + 2, 1 + 2 + 3, and 1 + 2 + 3 + 4, where 1, 2, 3, 4 represent the inner cell mass, blastocoel, trophoblast, and zona pellucida respectively. Comparing Figure 3 and Figure 4 shown, the feature information included in the upsampled logical graph processed in IACCM is: 0 (no feature or background feature), 1 (inner cell mass feature), 2 (blastocoel feature), 3 (trophoblast feature), 4 (zona pellucida feature); the feature information included in the upsampled logical graph processed in IRCCM is: 1 (inner cell mass feature), 1 + 2 (inner cell mass + blastocoel feature), 1 + 2 + 3 (inner cell mass + blastocoel + trophoblast feature), 1 + 2 + 3 + 4 (inner cell mass + blastocoel + trophoblast + zona pellucida feature).
[0078] To aggregate context information of different categories, the calculation representation area is as follows:
[0079] (11)
[0080] Where is a weighted mapping function that can be used to enhance the representation of boundaries and difficult pixels. According to formula (11), the structural feature map between tissues from the inside to the outside (inner cell mass to zona pellucida) can be obtained.
[0081] According to the characteristics of cross-entropy loss, when the probability in IRCCM approaches 0.5, the greater the loss, the more difficult the sample is to identify. For embryo tissue samples, the boundaries between different tissues are very small. In addition, since the blastocyst is transparent and needs to be photographed under low light conditions, its boundary features are not clear. Therefore, the present invention designs a weighted mapping function to reassign the probability weights of each category, that is, pixels that are not easily recognized (such as boundary regions) usually have a probability value of about 0.5, while pixels that are easily recognized have values close to 1 or 0. In this way, the weight values of sample pixels that are not easily recognized are close to , and the weight values of sample pixels that are easily recognized are close to 1. Therefore, the model has ideal learning and reasoning capabilities for edges and difficult pixels. It can be expressed as:
[0082] (12)
[0083] Where α and β can be selected as activation factors fixed at 1 and 2 respectively. When taking such values, formula (12) is formally the same as the upsampling function identical.
[0084] Finally, are connected in an orderly manner, and is used for convolution to obtain features that contain context information of different categories in the image.
[0085] For the semantic consistency module (CM):
[0086] Based on the assumption of semantic consistency, the present invention proposes a semantic consistency module (CM) for training, which helps to approximate the data distribution of the results output by I2C2Net to the distribution of real labels, and improves the generalization ability of the model.
[0087] The present invention takes the final results of I2C2Net under different conditions as generated samples and sends them to the Consistency Module (CM), and obtains a score for evaluating the quality. This scoring result can positively supervise the segmentation network, fit the segmentation results from the perspective of distribution, and particularly play a role in distinguishing the distributions of different categories. At the same time, the segmentation result is a multi-category value map with relatively less information. The real sample (original RGB image) and the segmentation result are combined and input into the discriminator for scoring to improve the semantic understanding ability and segmentation performance, as Figure 5 shown.
[0088] The Consistency Module (CM) includes a first convolutional block, a second convolutional block, a first self-attention block, a third convolutional block, a second self-attention block, and a convolutional output layer connected in sequence; among them, the first convolutional block, the second convolutional block, and the third convolutional block have the same structure, and each includes a convolutional layer, a spectral normalization layer, and an activation function layer.
[0089] In actual operation, the manually indexed real label or the segmentation result predicted by I2C2Net is combined with the original image as the input of the Consistency Module (CM), where the combination of the real label and the original image is the real result, and the combination of the segmentation result and the original image is the false result. The combined input data is processed sequentially through the first convolutional block, the second convolutional block, and the third convolutional block, and a spectral normalization (SN for short) operation and a LeakyReLU activation function are added to each convolutional block (as Figure 5 shown). SN is widely used in the training stage of GAN to ensure stability. The present invention makes the Consistency Module meet the 1-Lipschitz condition by applying SN, limits the degree of function change, and makes the model more stable. After passing through the three convolutional blocks, it also passes through two self-attention blocks and a convolutional block in the middle, and then obtains the final result through a single-layer convolution.
[0090] The present invention uses a consistency loss function to train the Consistency Module (CM), and the present invention uses a cross-entropy loss function to supervise the logical graph. The loss function can be expressed as:
[0091] (13)
[0092] where represents a one-hot label scaled to size. Obtained by formula (14):
[0093] (14)
[0094] is the predicted output map of different scales.
[0095] represents the cross - entropy loss, represents the sum of the values at all positions in the input image X.
[0096] The total cross - entropy loss is expressed as:
[0097] (15)
[0098] To improve the learning and expression ability of the basic model without increasing the computational cost, the present invention proposes an adversarial loss function to compensate for the cross - entropy loss. The present invention will use the result obtained by I2C2Net combined with the original embryo image X as "fake data" and input it into the consistency module, and input the combination of the real label Y and the original image X as "real data" into the consistency module. The consistency module extracts the combined data features from shallow to deep and outputs true / false results. This process can be expressed by Equation 16:
[0099] (16)
[0100] where Score represents the discrimination score, which is the output of the consistency module, . represents the concatenation operation. represents the consistency module. is the segmented image, which can be represented by the real label or the predicted map of the same input X.
[0101] To make the training stable, the LSGAN loss is adopted in our consistency module. By alternately updating, the LSGAN loss is introduced into the training, which helps the consistency loss :
[0102] (17)
[0103] where X represents the original image, is the data distribution of the original image. Y represents the real label, is the prediction result of X, obtained by obtained. represents the segmentation network, i.e., I2C2Net in the present invention.
[0104] Finally, the total loss function is expressed as:
[0105] (18)
[0106] where represents the hyperparameter used to balance and losses, preferably . Using the joint loss function, the model parameters are jointly learned through backpropagation, and the proposed I2C2Net can be trained in a coarse-to-fine manner.
[0107] Through the method of the present invention, it can be found that although the textures of blastocyst images from two domains are diverse, it can still be found that they contain the same tissues (e.g., blastocoel, inner cell mass, trophoblast, and zona pellucida). That is to say, these heterogeneous domains share the same semantic context information. Based on this discovery, the present invention proposes a semantic consistency hypothesis: when processing images from another heterogeneous domain, a well-trained network may still maintain partial (not all) semantic consistency. In fact, after being trained on a public dataset, I2C2Net can still distinguish the rough regions of these tissues in self-collected images to a certain extent (as Figure 6 shown). On the other hand, the clustering hypothesis, that is, the decision boundary should be in a low-density region, is a basic hypothesis in semi-supervised learning. That is, the partial semantic consistency obtained from unlabeled images helps to increase the density of the labeled dataset, thereby improving the accuracy. Based on these two hypotheses, the design of the entire training strategy is as Figure 6 shown. It mainly includes the following five steps:
[0108] Step S100: I2CNet is first trained on the labeled data using the consistency module and the total loss function in Equation (18).
[0109] Step S200: Unlabeled blastocyst images from another heterogeneous domain are segmented by the pre-trained I2CNet in the first step, and these segmentation maps are then used as inputs for the corresponding pseudo-labels . According to the semantic consistency hypothesis, the semantic information contained in the pseudo-labels is consistent in physiological structure.
[0110] Step S300: For the labeled dataset ( and the corresponding ), the supervised consistency loss is obtained by the following formula
[0111] (19)
[0112] where is the segmentation result of I2CNet corresponding to the input . The total loss is:
[0113] (20)
[0114] where is calculated by Equation (15). Preferably, is equal to 2.0.
[0115] Step S400: For the unlabeled images and their pseudo-labels , an unsupervised consistency loss is designed:
[0116] (21)
[0117] where is the segmentation map. The consistency module attempts to utilize semantic consistency and encourage the density of different segmentation classes.
[0118] Step S400: As Figure 6 shown, I2C2Net considers both labeled and unlabeled data simultaneously. Therefore, the total loss function for network training is finally:
[0119] (22)
[0120] where is a parameter, preferably set to 0.1, equals 2.0.
[0121] As for the consistency module, it is:
[0122] (23)
[0123] I2C2Net uses I2CNet as the segmentation backbone and achieves semantic consistency in the semi-supervised learning scenario.
[0124] The supervised dataset used in the test of the present invention is the human blastocyst public dataset, which contains 249 blastocyst images and their ground truth (GT) provided by the Pacific Center for Reproductive Medicine (PCRM). The training and test sets contain 199 (80%) and 50 images (20%). The dataset collected in the test data includes 84 human blastocyst images, which are from the Reproductive Medicine Center of the First Affiliated Hospital of Anhui Medical University. The blastocyst images were taken on an inverted microscope (IX-71; Olympus, Japan) and a heated micromanipulator (Nikon, eclipse Ti2, Japan). The relevant work has been approved by the Medical Ethics Committee of Anhui Medical University and follows the Declaration of Helsinki.
[0125] The backbone of I2C2Net of the present invention and the two integrated context modules are randomly initialized by Kaiming initialization. The "poly" learning rate policy with a factor of is used for training. The optimizer is Adam, and the initial learning rate and weight decay are set to The number of training epochs is 2000. Synchronized batch normalization implemented using PyTorch is used during training. For data augmentation, random scaling, random rotation, random cropping, and horizontal flipping are performed on each sample during the training phase. The input image size is It is implemented using PyTorch (version ≥ 1.3) and trained using four NVIDIA Tesla V100 GPUs (each with 32 GB of memory). All test programs are executed on a single NVIDIA Tesla V100 GPU.
[0126] To evaluate the performance of the network, five commonly used metrics are used: Accuracy, Precision, Recall, Dice coefficient, and Jaccard index. The calculation of these metrics is as follows:
[0127] Accuracy = (24)
[0128] Precision = , (25)
[0129] Recall = (26)
[0130] Dice coefficient = , (27)
[0131] Jaccard index = , (28)
[0132] Where TP represents the number of pixels correctly identifying the tissue area; TN is the number of pixels correctly identified as the background (unsegmented area); FP represents the number of pixels misclassifying the background as blastocoel, inner cell mass, trophoblast, and zona pellucida regions; FN represents the number of pixels misclassified as the background. Accuracy indicates the effectiveness of the method in correctly distinguishing the background and blastocyst tissue. Precision represents the proportion of all pixels correctly predicted as blastocyst tissue. Recall indicates the effectiveness of the method in detecting blastocyst tissue. The Dice coefficient and Jaccard index are crucial for evaluating the overall segmentation performance and are applicable to handling class imbalance; in addition, the impact of misidentification and missed regions is also considered. The range of the Dice coefficient and Jaccard index is from 0 to 1, and the closer the metric is to 1, the better the segmentation performance.
[0133] The ablation experiment results verify the effectiveness of IACCM, IRCCM and CM proposed in the present invention. Compared with other supervised methods, the I2C2Net of the present invention achieves the best performance in accuracy, precision, recall, Dice coefficient and Jaccard index, which are 93.79%, 91.69%, 92.23%, 91.95% and 85.33% respectively. Moreover, the accuracy, Dice and Jaccard index of I2C2Net on the supervised data set are improved by 0.49%, 0.31% and 0.09% respectively; the semi-supervised version of I2C2Net also achieves the best performance compared with other mainstream semi-supervised methods of the same period; in particular, the model accuracy index of the present invention is improved by at least 0.53%, the precision is improved by at least 0.54%, the Dice coefficient is improved by at least 0.67%, and the Jaccard index is improved by at least 1.14%.
[0134] Through the detailed description of the specific implementation methods and embodiments of the present invention above, ordinary technicians in the field can understand that various changes, modifications, substitutions and variations can be made to these implementation methods and embodiments without departing from the principles and basic concepts of the present invention, and the scope of the present invention is defined by the attached claims and their equivalents. Contents not described in detail in this specification, such as the details of the specific structure and implementation of a typical neural network and its functional layer, are within the scope of what can be achieved by professional and technical personnel in the field based on the known prior art and the technical teachings of the present invention.
Claims
1. A biological tissue processing method based on semantic consistency network, characterized in that Includes steps: A network model for modeling intra-class and inter-class features of tissues based on semantic consistency is constructed, and the network model includes N encoding blocks, N-1 I2C modules, N compression blocks, and N logic blocks; wherein each encoding block, compression block, and logic block is a convolutional module based on a convolutional neural network, and each I2C module includes an intra-class context module representing the aggregated pixels in the category area and an inter-class context module capable of capturing the morphological structure between tissues; and the N encoding blocks are connected in series in sequence, and the output of the Nth encoding block is input to the corresponding Nth compression block for processing, and the Nth compression block The result is input to the Nth logic block for processing and then input to the N-1th I2C module after upsampling. The N-1th I2C module receives the output of the N-1th coding block and the upsampled output after processing by the Nth logic block, outputs the processed result to the N-1th logic block for processing and then input to the N-2th functional block after upsampling, and so on, until it is processed by the 1st logic block and output as the prediction result; N is an integer greater than 2; Furthermore, the network model for organizing intra-class and inter-class features based on semantic consistency modeling also includes a consistency module for training the network model; Acquire training sample data, where the training sample data is biological tissue image data, a portion of which is labeled data for supervised training, and another portion of which is unlabeled data for semi-supervised training; Using the training sample data, the network model for modeling tissue intra-class and inter-class features based on semantic consistency is trained to obtain a biological tissue recognition processing training model; Use biological tissue recognition and processing training models to recognize and process biological tissue image data; The consistency module includes a first convolution block, a second convolution block, a first self-attention block, a third convolution block, a second self-attention block, and a convolution output layer connected in sequence; the first convolution block, the second convolution block, and the third convolution block have the same structure, and all include a convolution layer, a spectrum normalization layer, and an activation function layer; The consistency loss function is used to train the consistency module, and the cross entropy loss function is used to supervise the logical graph. The loss function is expressed as: ; in, Represents a unique hot label Zoom to The size of; K is the number of categories, W and H are the width and height of the biological tissue image X respectively; Obtained using the following formula: ; Output graphs of predictions at different scales; represents the cross entropy loss, Represents the input biological tissue image Sum the values of all positions in ; The total cross entropy loss is expressed as: ; The results obtained using this network model Original biological tissue images The combination is input into the consistency module as "fake data" to convert the real label and original biological tissue images The combination of is input into the consistency module as "real data"; the consistency module extracts the combined data features from shallow to deep and outputs true / false results; it is expressed by the following formula: ; Among them, Score represents the discrimination score, which is the output of the consistency module. ; Indicates a connection operation; Represents the consistency module; is a segmented image, represented by the true label or the predicted map of the same input biological tissue image X; In order to make the training stable, LSGAN loss is used in the consistency module, and LSGAN loss is introduced into the training through alternating updates: ; in, It is the original biological tissue image Data distribution; represents the true label, is the prediction result, get; Represents a segmentation network.
2. The biological tissue processing method based on semantic consistency network according to claim 1, characterized in that: The N is an integer of 5 or greater.
3. The biological tissue processing method based on semantic consistency network according to claim 2, characterized in that: Each encoding block consists of two convolutional network layers, and each convolutional network layer consists of two batch normalization layers, two activation layers and a maximum pooling layer.
4. The biological tissue processing method based on semantic consistency network according to claim 3, characterized in that: Define the encoding block as , coding block Each input biological tissue image The pixels in are projected into the nonlinear embedding space to obtain the pixel representation matrix : ; in, The size is , yes The number of channels.
5. The biological tissue processing method based on semantic consistency network according to claim 4, characterized in that: Define the compression block as , compressed block Representing pixels as a matrix Compress to : ; in, represents the features obtained by aggregating intra-class and inter-class information from different scales of the entire image using the I2C module, The size is .
6. The biological tissue processing method based on semantic consistency network according to claim 5, characterized in that: Define the logic block as , logical block Used to obtain images of biological tissues Supervised Logical Graph : ; Among them, the logic block It consists of a convolutional layer, a batch normalization layer, and an activation layer.
7. The biological tissue processing method based on semantic consistency network according to claim 6, characterized in that: Get the probability map from the output of the logic block The upsampling function is .