Method for identifying and processing biological tissue based on task consistency
By adopting a semi-supervised method of a unified task-aware consistent interaction framework in image recognition and segmentation, combined with a skeleton network and a multi-branch convolutional neural network, the problem of unlabeled data utilization is solved, efficient identification and segmentation of the internal structure of blastocyst tissue is achieved, and the performance of medical image segmentation is improved.
Patent Information
- Application Number
- CN202510229399.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The prior art is difficult to effectively utilize unlabeled data and combine embryonic skeleton and boundary features for automatic identification and evaluation, especially in the identification and segmentation of internal structures of blastocyst tissue.
The semi-supervised image recognition and segmentation method based on a unified task-aware and consistent interaction framework is adopted to extract multi-scale image features through the skeleton network, and parallel convolutional neural network training is used to train the main branch and additional branches (including segmentation branch, horizontal set branch and point set branch) to achieve the integration of output results at pixel level, geometric level and point level.
It improves the performance of semi-supervised medical image segmentation, can effectively utilize unlabeled data, enhances learning dynamics, and improves the accuracy and robustness of blastocyst image recognition and segmentation.
Smart Images

Figure CN119723252B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for biological tissue recognition and processing based on task consistency, specifically to a semi-supervised image recognition and segmentation method based on a unified task-aware consistency interaction framework, belonging to the technical field of image pattern recognition based on artificial neural networks and the field of artificial intelligence technology applied to medicine. Background Art
[0002] In in vitro fertilization technology, it is necessary to culture fertilized cells until the seventh day to form embryos. At this time, it is necessary to evaluate their implantation potential according to the morphology of various tissues in the embryos. Currently, this tissue morphology analysis mainly relies on the subjective judgment of operators.
[0003] In the prior art, attempts have been made to explore the automatic recognition and evaluation of different embryo tissue morphologies to provide a more objective evaluation. For example, Chinese patent application CN202210978474.8 proposes a multi-stage intelligent analysis system based on embryo development images, which applies a YOLO network including 24 convolutional layers, 4 pooling layers, and 2 fully connected layers to output class labels of developmental stages for blastocyst tissue images to achieve multi-stage recognition and segmentation. However, it can only segment the entire blastocyst image from the culture dish background and cannot achieve the recognition and segmentation of the internal structure image of the blastocyst tissue. Another example is that in the prior art, traditional methods, such as the level set method, have been used to segment the inner cell mass and trophoblast of the blastocyst image. With the progress of deep learning, some studies have attempted to combine traditional methods with deep learning for the recognition of the internal structure of blastocyst tissue in blastocyst images. For example, traditional features are extracted and input into a two-layer neural network to recognize the zona pellucida (ZP), trophoblast (TE), and inner cell mass (ICM) in blastocyst images, and a fully convolutional network (FCN) is used to further segment the inner cell mass (ICM). However, these data-driven networks usually require a large amount of labeled data for network training, but obtaining labeled data in clinical practice takes time and a large amount of resources, which hinders the widespread application of deep learning technology.
[0004] In the prior art, methods that exploit semi-supervised learning by minimizing the entropy of predicted probabilities have also begun to be explored to be applicable to unlabeled data. For example, adversarial training is used to regularize the segmentation model, aiming to minimize the difference between the segmentation distributions of labeled and unlabeled images; interpolation consistency (IC) is used to apply consistency between different versions of the input data; a regularized discard network (RD) is used to enforce consistent segmentation between the discarded information of the network; a teacher-student framework is used to enforce prediction consistency between labeled and unlabeled samples during the SSL process, and Monte Carlo discard is combined to estimate the uncertainty in the teacher model prediction; an aggregation and decoupling framework is used to capture distribution-invariant features; a method that uses global information in segmentation masks to estimate segmentation uncertainty, etc. However, these methods focus on pixel-level consistency without fully exploring the underlying representation consistency between different feature spaces.
[0005] According to existing experience, how to effectively use unlabeled data to identify different problems by combining embryo skeleton and boundary features is a technical problem that needs to be solved urgently. Summary of the invention
[0006] In order to solve the above technical problems, the present invention provides a biological tissue recognition and processing method based on task consistency, comprising the steps of: constructing a network model based on a unified task-aware consistency interaction framework, the network model comprising a skeleton network for extracting multi-scale image features, a main branch and at least two additional branches; wherein the skeleton network is based on an encoder-decoder architecture of a deep learning model, and the output result of the skeleton network is input into the main branch and at least two additional branches; the main branch and the at least two additional branches are convolutional neural networks in a parallel relationship, wherein the main branch is used to perform a main task, the task performed by at least one additional branch has a strong correlation with the main task, and the task performed by at least another additional branch has a weak correlation with the main task; obtaining training sample data, the training sample data being biological tissue image data, a part of the biological tissue image data being labeled data for supervised training, and the other part being unlabeled data for semi-supervised training; using the training sample data to train the network model based on the unified task-aware consistency interaction framework to obtain a biological tissue recognition and processing training model; and using the biological tissue recognition and processing training model to recognize and process the biological tissue image data.
[0007] In the above technical scheme, the main branch is a segmentation branch, which is used to generate a pixel-level probability map and generate a pixel-level output result; an additional branch is a level set branch, which is used to generate a target level set representation and generate a geometric level output result, which has a strong correlation with the main task; another additional branch is a point set branch, which is used to synthesize a point representation related to the target structure and generate a point-level output result, which has a weak correlation with the main task.
[0008] In the above technical solution, the segmentation branch is supervised and trained, and the level set branch and the point set branch are simultaneously supervised and semi-supervised trained.
[0009] In the above technical solution, the biological tissue image data has a dimensional structure of, where C represents the number of color channels of the biological tissue image, H and W respectively represent the height and width of the biological tissue image, and R represents the real number space; the biological tissue image data is processed by the skeleton network to obtain feature sets of different scales denoted by, where the dimension , and respectively correspond to processing C, H, and W with the downsampling factor of ; the integer D represents the degree of downsampling performed at each stage.
[0010] In the above technical solution, the loss function of the network model based on the unified task-aware consistency interaction framework when trained with labeled data is:
[0011]
[0012] Among them, is the loss function for the supervised training of the segmentation branch, is the loss function for the supervised training of the point set branch, is the loss function for the supervised training of the level set branch, is the strong and weak task-aware consistency function, , , , are weight coefficients.
[0013] In the above technical solution, the loss function of the network model based on the unified task-aware consistency interaction framework when trained with unlabeled data is:
[0014]
[0015] Among them, is the loss function for the unsupervised training of the point set branch, is the loss function for the unsupervised training of the level set branch, is the strong and weak task-aware consistency function, , , are weight coefficients.
[0016] In the above technical solution, the loss function for semi-supervised learning of the network model based on the unified task-aware consistency interaction framework has the overall expression as follows:
[0017]
[0018] Among them, and are weight coefficients, is the loss function when the network model based on the unified task-aware consistency interaction framework is trained with labeled data, is the loss function when the network model based on the unified task-aware consistency interaction framework is trained with unlabeled data.
[0019] In the above technical solution, training the network model based on the unified task-aware consistency interaction framework using the training sample data includes:
[0020] Step S100: Randomly initialize the segmentation branch parameters , the point set branch parameters , and the level set branch parameters ;
[0021] Step S200: When the predetermined loss expectation or the number of training times is not completed, execute:
[0022] Step S210: Obtain each batch of training samples ; Among them, are the input image and the label map corresponding to the labeled data set respectively; is the input image corresponding to the unlabeled data set ;
[0023] Step S220: Calculate the prediction results of the segmentation branch, the point set branch, and the level set branch respectively;
[0024] Step S230: Use to generate through the conversion formula of the point set branch , and generate through the conversion formula of the level set branch ;
[0025] Step S240: Generate the point set representation feature from the prediction result through the formula ;
[0026] Step S250: Generate from the prediction result Generate the level set representation feature ;
[0027] Step S260: For the prediction result Generate the pixel representation feature through the formula ; ;
[0028] Step S270: For the prediction result Generate the point set representation feature through the formula ; ;
[0029] Step S280: Obtain the overall loss ;
[0030] Step S290: Update the model parameters through backpropagation ;
[0031] Step S300: When the predetermined loss expectation or the number of training times is completed, end the loop;
[0032] Step S400: Return the obtained segmentation branch parameters , the point set branch parameters , the level set branch parameters .
[0033] The present invention also provides a biological tissue recognition and processing device based on task consistency, comprising: a model construction module, configured to construct a network model based on a unified task-aware consistency interaction framework, the network model including a backbone network for extracting multi-scale image features, a main branch, and at least two additional branches; wherein, the backbone network is based on an encoder-decoder architecture of a deep learning model, and the output result of the backbone network is input into the main branch and at least two additional branches; the main branch and at least two additional branches are convolutional neural networks in a parallel relationship, wherein the main branch is used to execute a main task, the task executed by at least one additional branch has a strong correlation with the main task, and the task executed by at least another additional branch has a weak correlation with the main task; a training sample acquisition module, configured to acquire training sample data, the training sample data being biological tissue image data, a part of the biological tissue image data being labeled data for supervised training, and another part being unlabeled data for semi-supervised training; a model training module, using the training sample data to train the network model based on the unified task-aware consistency interaction framework to obtain a biological tissue recognition and processing training model; wherein, the biological tissue image data in the training sample data is input into the backbone network to obtain feature sets of different scales, and the feature sets are respectively input into the main branch and at least two additional branches for training, the main branch and at least two additional branches have loss functions corresponding to the tasks they execute, and these loss functions are compared and guided by information between related tasks to utilize the synergistic effect between these diverse tasks; a data processing module, using the biological tissue recognition and processing training model to recognize and process biological tissue image data.
[0034] The present invention has achieved the following technical effects:
[0035] Based on the consistent inference results with different task awareness, a new unified task-aware consistency interaction (UniTask+) framework is developed, which fully utilizes the synergistic effect between these diverse tasks, promotes the effective utilization of task-aware knowledge, and improves the performance of semi-supervised medical image segmentation. The UniTask+ framework method promotes the exploration of inherent segmentation perturbations at three key levels (pixel level, geometric level, and point level), enhances the learning dynamics, and benefits both supervised and semi-supervised scenarios. The experimental evaluation on the ICM, blastocyst, and left atrium datasets has reached the state-of-the-art level, and it is effective and robust in dealing with the complex challenges of medical image segmentation. Description of the Drawings
[0036] Figure 1 A correlation diagram for processing three tasks on a blastocyst image;
[0037] Figure 2For the structural diagram of the unified task perception consistency interaction (UniTask+) framework;
[0038] Figure 3 For the network structure diagram of the segmentation branch;
[0039] Figure 4 For the network structure diagrams of the point set branch and the level set branch. Specific implementation manners
[0040] To facilitate the understanding and implementation of the present invention by those of ordinary skill in the art, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0041] As Figure 1 shown, for determining whether an embryo can continue to develop better, factors such as the filling degree and proportion of the inner cell mass (ICM) characterized by the blastocyst image are key evaluation indicators; at the same time, other structural morphologies at the blastocyst stage, such as the blastocoel, trophoblast, and zona pellucida, are also of great significance. Therefore, it is necessary to effectively combine the tasks of identifying structural information and boundary information with the traditional segmentation task to achieve the automatic segmentation of different blastocyst tissues; and, in this case, it is also necessary to make full use of unlabeled data. Based on this requirement, the segmentation task of the blastocyst image needs to include at least the following main contents: (1) The point set representation task, which is used to represent the structural information of the inner cell mass (ICM). In Figure 1 , the skeleton point set depicted by the point set representation task effectively captures and represents the overall distribution trend of the inner cell mass (ICM) at the point level; however, this representation does not include boundary information. (2) The semantic segmentation task; the result of which shows a significant irregular pattern at the pixel level in the clustering in the blastocyst image. (3) The distance transformation task, which is used to represent the boundary information of the inner cell mass (ICM). There is an inherent connection between the output results obtained from these three tasks. Specifically, as Figure 1 shown, the clustering of the inner cell mass (ICM) in the blastocyst image shows a significant irregular pattern at the pixel level; there is a weak correlation between the output result of the point set representation task and the output result of the semantic segmentation task, and there is a strong correlation between the output result of the semantic segmentation task and the output result of the distance transformation task.
[0042] Therefore, fusing and superimposing the output results of the point set representation task, the semantic segmentation task, and the distance transformation task can take into account the above weak correlation and strong correlation. As Figure 1 shown, through the semantic segmentation task as a mediator, the geometric-level information extracted from the distance transformation task can be converted into a point representation and connected to the point-level task obtained in the point set representation task. Thus, it is necessary to consider the shared basic features and similar processing pipelines between these tasks within a network framework to achieve Figure 1Task-aware consistency among the three tasks shown
[0043] Based on the task-aware smoothness (TAS) assumption, that is, assuming that the mutual information between the output results of the above three tasks remains unchanged during task transformation, including strong correlation (i.e., strong task-aware consistency) and strong-to-weak correlation (i.e., strong-to-weak task-aware consistency). This stable exchanged information promotes feature consistency and effectively minimizes the difference between the learned features and their intrinsic attributes, especially suitable for processing unlabeled data.
[0044] The mathematical definition of the task-aware smoothness assumption is: given an input , and two outputs , For two different tasks A and B, after applying the task transformation to , the corresponding results and should be the same.
[0045] When the task transformation is bidirectional in the task-aware smoothness assumption, task B is defined as strongly correlated with task A, and vice versa. When is unidirectional, task B is called a weakly correlated task. Finally, the task-aware consistency is obtained by minimizing:
[0046] (1)
[0047] where is the metric function; and represent executing task A and task B and sharing the same backbone network; represents the conditional probability of the model. When B is a weakly correlated task, the corresponding weak task-aware consistency is the same as formula (1). This equation can effectively cover many previous studies using the dual-anchored two-stream architecture in semi-supervised learning. These studies utilized two independent branches in their methodology to handle SSL tasks, usually one branch for handling labeled data and the other branch for leveraging unlabeled data or combining labeled and unlabeled information in a collaborative manner. The equation under discussion may reflect the core principles or mathematical formulas of these methods for handling consistency specifications, feature extraction, or joint learning between different branches to improve model performance under limited supervision.
[0048] In addition, when B is a strongly correlated task, the corresponding strong task consistency is:
[0049] (2)
[0050] Among them, is 's inverse transformation. This method helps to integrate and interact the task-aware consistency at multiple levels in a unified semi-supervised learning framework.
[0051] For the above formula (2), if a strongly related task A and a weakly related task C can be found in the main segmentation task B, this kind of consistency can be achieved:
[0052] (3)
[0053] Among them, is the composite transformation from the strongly related task to the weakly related task, and can also be expressed as the transformation to , followed by a continuous transformation of a one-way transformation ; represents the execution of task C. By utilizing this sufficient task-aware consistency in formula (3), more rich and discriminative feature representations can be learned from locally and globally related auxiliary tasks, thereby potentially improving the generalization ability and accuracy of the main segmentation task B.
[0054] To solve the above technical problems, the present invention proposes a method for identifying and segmenting blastocyst images through semi-supervised learning from a task perspective, focusing on the task invariance and task-specific dependence of segmentation. Based on the derivation results with consistent task awareness of different tasks, the present invention develops a new unified task-aware consistency interaction (UniTask+) framework, which includes a main branch for medical segmentation (MS) based on a backbone network, and two additional branches based on the same backbone network to perform strongly / weakly related tasks; among them, a contour / level set (LS) branch focuses on extracting and processing geometric information to promote strong consistency; while a point set (PS) branch is used to generate point set representations to stimulate weak consistency under underlying task perturbations.
[0055] More specifically, the present invention provides a method for biological tissue recognition and processing based on task consistency, including the steps of: constructing a network model based on a unified task-aware consistency interaction framework, the network model including a backbone network for extracting multi-scale image features, a main branch, and at least two additional branches; wherein, the backbone network is based on the encoder-decoder architecture of a deep learning model, and the output result of the backbone network is input into the main branch and at least two additional branches; the main branch and at least two additional branches are convolutional neural networks in a parallel relationship, wherein the main branch is used to execute the main task, at least one additional branch executes a task that has a strong correlation with the main task, and at least another additional branch executes a task that has a weak correlation with the main task; obtaining training sample data, the training sample data being biological tissue image data, a part of the biological tissue image data being labeled data for supervised training, and another part being unlabeled data for semi-supervised training; using the training sample data to train the network model based on the unified task-aware consistency interaction framework to obtain a biological tissue recognition and processing training model; using the biological tissue recognition and processing training model to recognize and process biological tissue image data.
[0056] Preferably, after the biological tissue image data in the training sample data is input into the backbone network, feature sets of different scales are obtained, and the feature sets are respectively input into the main branch and at least two additional branches for training. The main branch and at least two additional branches have loss functions corresponding to the tasks they execute, and these loss functions are compared and guided by the information between related tasks to utilize the synergistic effects between these diverse tasks.
[0057] Preferably, the unified task-aware consistency interaction (UniTask+) framework provided by the present invention, as Figure 2 shown, includes a backbone network and three branches based on the backbone network, namely: (i) a medical segmentation branch (MS), mainly used to generate a pixel-level probability map; (ii) a level set branch (LS), mainly used to generate a target level set representation; (iii) a point set branch (PS), mainly used to synthesize a point representation related to the target structure. Input the blastocyst image X into the backbone network, and after passing through the backbone network, feature maps of different scales are obtained , and enter the MS, LS, and PS branches respectively to generate pixel-level, geometric-level, and point-level output results. The loss function is compared and guided by the information between the main task and the strong / weak related tasks. This coherent framework makes full use of the synergistic effects between these diverse tasks, promotes the effective utilization of task-aware knowledge, and improves the performance of semi-supervised medical image segmentation.
[0058] The backbone network in the unified task-aware consistency interaction (UniTask+) framework provided by the present invention is based on a typical encoder-decoder architecture. Or rather, the unified task-aware consistency interaction (UniTask+) framework can be seamlessly integrated into any encoder-decoder architecture. As Figure 2 shown, it includes an input layer, an encoding layer / encoder, several hidden layers, a decoding layer / decoder, and one or more output layers. The encoder-decoder architecture is a deep learning model structure widely used in fields such as natural language processing, image processing, and speech recognition. This structure can handle sequence-to-sequence tasks such as machine translation, text summarization, dialogue systems, and voice conversion. Among them, the role of the encoder is to receive the input sequence and convert it into a fixed-length context vector. The encoder can be any type of deep learning model. The goal of the decoder is to convert the context vector generated by the encoder into an output sequence. The decoder can also be various types of deep learning models, but usually the same type of model as the encoder is used to maintain consistency. When training a model based on the encoder-decoder architecture, the goal is to minimize the difference between the output sequence predicted by the model and the actual output sequence, usually achieved by calculating a loss function and using optimization algorithms such as backpropagation and gradient descent for parameter update.
[0059] As Figure 2 shown, for the main task, the MS branch obtains the segmentation result ; for the strongly related task, the LS branch obtains the geometric representation under the level set function ; for the weakly related task, the PS branch obtains the point set of the skeleton representation of the object . When processing labeled data, the outputs of these three tasks - segmentation prediction ( ), level set representation ( ), and point set representation ( ) - are respectively compared with the corresponding ground truth to evaluate their respective supervised losses. In the case of processing unlabeled data, the unsupervised losses are calculated using the level set reconstruction from the LS branch ( ) and the point set representation from the PS branch ( ), and the segmentation prediction from the MS branch ( )As a consistency anchor. This cross-task participation enhances the ability of the UniTask+ framework network to extract meaningful representations from semi-supervised learning scenarios, regardless of whether the data is labeled or unlabeled. By minimizing the differences between pixel and pixel, pixel and contour, pixel and skeleton point, and contour and point, the UniTask+ framework network can learn consistent features of pixel-level, geometric-level, and point-level data in three tasks, maintaining strong, weak, and strong-to-weak task-aware consistency, thereby updating the entire UniTask+ framework network, making full use of unlabeled data and improving the performance of the UniTask+ framework network in identifying and segmenting blastocyst images.
[0060] For the blastocyst image input into the backbone network has a dimensional structure, where C represents the number of channels. For RGB images, the number of color channels is usually C = 3, while H and W represent the height and width of the image X, respectively. The blastocyst image undergoes a series of computational processes in the backbone network and sequentially processes to generate feature maps of different scales, denoted by , where the dimensions and are reduced according to the downsampling coefficients of , that is
[0061]
[0062] The integer D represents the degree of downsampling performed at each stage. Each such feature map has a corresponding number of channels, denoted as . After the multi-scale feature extraction of the backbone network, the output result of the backbone network provides a final feature representation, called .
[0063] A multi-scale branch is introduced after the backbone network to form the backbone network architecture. The segmentation branch network utilizes the feature set output by the backbone network and derives the final recognition feature space through a series of convolutional operations. The output has a dimensional structure, where K represents the number of categories, usually K = 2 for binary segmentation tasks. The network structure and loss function of the segmentation branch are as Figure 3 shown. Its typical structure includes an input layer, several convolutional network layers, and an output layer; optionally, the input layer of the segmentation branch can be merged / shared with an output layer of the backbone network.
[0064] Meanwhile, a point set branch network is also introduced after the backbone network to obtain the blastocyst image The point - set representation of skeleton information explores potential weak - task - aware structures to capture the complex structural details of the skeleton. Specifically, as shown in Figure 4 , the input of the point - set branch is represented as , which is the feature map in the skeleton network. It contains two learning heads: a class head and a location head. It reflects that each point in the data should carry both its class identity and relative spatial coordinates simultaneously. The initial mapping of the input feature produces an estimated point representation , which consists of an object confidence score and a relative position coordinate map . The class head mainly identifies whether a given point belongs to the skeleton structure, while the location head is specifically used to estimate the positional relationship between each point and the center point of the image. To align with the ground truth, is introduced in the ground truth for transformation, resulting in the actual point - set representation , including the target score and the relative position coordinate map .
[0065] The transformation function is used to convert the pixel - level segmentation mask into a skeleton representation suitable for the point - set representation . The mathematical expression of this transformation is as follows:
[0066] (4)
[0067] where \(U\) represents the union operation, and the symbols and represent morphological erosion and opening operations respectively. is a pre - defined structuring element, and \(N\) is determined as the maximum value such that the erosion operation does not result in an empty set, that is . This iterative process essentially extracts a skeleton - like structure by repeatedly eroding and subtracting the eroded image from the original image, thus maintaining the core shape and connectivity of the segmented object, in the form of the point - set representation . Then the target score , which can also be represented as a matrix with binary values from . And we convert to the point - set , where \(a\) and \(b\) represent the horizontal and vertical relative distances between the current point and the center of the image, and represents the data point in .
[0068] In addition, by leveraging the estimation from the target score and relative position coordinates The calculated combined loss introduces the overall optimization strategy of the PS branch. This method plays an important role in preventing misidentification between a single point and its spatial position.
[0069] For some predicted points, they may have high confidence scores , but their position estimates may deviate significantly. Therefore, effectively combining these two types of information becomes crucial for reducing ambiguity and disambiguating between different categories, which will ensure that the predicted points can not only be confidently classified but also correctly located, thus contributing to improving the performance in identifying and partitioning complex structures in the skeleton. However, the results predicted by the MS branch cannot be reconstructed solely from the information generated by the PS branch. The PS branch follows formula (1), so the PS branch is called a weakly related task.
[0070] The present invention associates the assignment of strongly related tasks with a contour distance representation branch (i.e., the level set branch) to capture the geometric active contour and distance information of the internal structure of the blastocyst tissue image. From this, geometric-level and pixel-level information can be obtained, and these consistencies are minimized to further learn the boundary, thereby increasing the consistency of strongly related tasks. Here, the feature set output by the skeleton network is used as the input of the level set branch, and the role of the level set task is simulated through the level set branch. Specifically, this is achieved through a 1×1 convolutional layer and normalized using the tanh function to obtain the final level set output , as shown in the level set branch of Figure 4 .
[0071] As strongly related tasks and should be bi-directionally transformed. The transformation from to can be set as the forward transformation , which is defined as:
[0072] (5)
[0073] where, Inf represents the infimum; represents the L2 norm; , are two different pixel points from , respectively; is the zero level set, which also represents the contour of the target object; and represent the internal region and the external region of the target object respectively.
[0074] Due to the non-differentiability of formula (5), it is difficult to solve its inverse mapping problem, and the following smoothing function can be used as the inverse mapping:
[0075] (6)
[0076] After the input is magnified by k times, it is differentiable and invertible, which can increase the task perception consistency.
[0077] Based on the above, given a pair of supervised data , after applying the backbone network to obtain features of different scales, the MS branch obtains the segmentation result , the LS branch obtains the geometric representation under the level set , and the PS branch obtains the skeleton point representation of the object .
[0078] For the loss function of the main task, as shown in the loss function of the MS branch in Figure 3 , under the supervised data, directly use and to calculate the cross-entropy to obtain
[0079] (7)
[0080] Among them, represents calculating the cross-entropy of and .
[0081] For the loss of weakly related tasks, as shown in the loss function of the PS branch in Figure 4 , first convert the ground truth label to (including the ground truth target score and the ground truth location map ) through the skeleton conversion method, and use and (including the predicted target score and the predicted location map ), and the PS branch calculates the loss including the score loss and the location information loss
[0082] (8)
[0083] Therefore, we can obtain the loss function of the point set representation
[0084] (9)
[0085] Among them, μ 1 and μ 2 respectively represent weight coefficients. Preferably, μ 1 = 0.1 and μ 2 = 1.0.
[0086] Considering the weak task loss consistency in formula (1), use and ( is obtained through formula ) to obtain the weak task consistency loss .
[0087] (10)
[0088] The training objective is to ensure that the predicted Y is very close to the true value Y. Therefore, the weight assigned to each loss is set to 1. The overall loss of the PS branch is
[0089] (11)
[0090] For the loss function of strongly correlated tasks, as shown by the loss function of the LS branch in Figure 4 , first generate through formula , then use and generated by the LS branch to calculate the mean square error (MSE) as the distance loss function:
[0091] (12)
[0092] Secondly, the distance loss function can also be calculated according to predicted by the LS branch and generated through :
[0093] (13)
[0094] In addition, generated by the LS branch can be reversed to obtain , so the consistency loss of this part can be calculated :
[0095] (14)
[0096] Among them, is function.
[0097] The sum of equations (13) and (14) conforms to formula (2), reflecting the strong consistency goal of the task to be achieved by the present invention. Therefore, the overall level set branch (LS) loss can be defined as :
[0098] (15)
[0099] wherein 、 、 are weight coefficients; preferably = 0.01, = 0.01 and = 1.
[0100] As Figure 4 shown, the strong and weak task-aware consistency in the UniTask+ framework is also introduced.
[0101] (16)
[0102] In summary, the supervision loss of the UniTask+ framework is defined as:
[0103] (17)
[0104] wherein 、 、 、 are weight coefficients, preferably = 1.0, = 0.006, = 1.0 and = 0.001.
[0105] It is different from the previous UniTask framework:
[0106] (18)
[0107] The semi-supervised learning process of the UniTask+ framework of the present invention is as follows:
[0108] Perform semi-supervised training of task-aware consistency through UniTask+:
[0109] Input:
[0110] wherein are the input image and label map corresponding to the labeled dataset respectively; is the input image corresponding to the unlabeled dataset .
[0111] Output: MS branch parameter , PS branch parameter , LS branch parameter (Note that they share the same backbone)
[0112] 1: Random initialization ;
[0113] 2: When the predetermined loss expectation or the number of training times is not completed, execute:
[0114] 1): Obtain each batch of training samples ;
[0115] 2): Calculate the results of the three branches , and the generated formulas are respectively , and ;
[0116] 3): Use to generate through the point set transformation formula , and generate through the level set transformation formula ;
[0117] 4) Generate the point set representation feature for the prediction result through the formula ;
[0118] 5) Generate the level set representation feature for the prediction result through the formula ;
[0119] 6) Generate the pixel representation feature for the prediction result through the formula ;
[0120] 7) Generate the point set representation feature for the prediction result through the formula ;
[0121] 8) Obtain the overall loss through formulas (17), (21), and (22);
[0122] 9) Update the model parameters through backpropagation ;
[0123] 3: End the loop;
[0124] 4: Return the parameters ;
[0125] For unlabeled data, although no label information can be used, the feature differences between different tasks can be minimized, and task consistency between tasks can be found, so as to effectively utilize unlabeled data.
[0126] Specifically, for the loss function of the PS branch, it can also be obtained as in formula (10) :
[0127] (19)
[0128] The loss function of the LS branch can be expressed as:
[0129] (20)
[0130] where the weights and are the same as in the supervised method (as shown in formula (15)). Similarly, the strong and weak task consistency in UniTask+ (as shown in Figure 4 and given by formula (16)) is used for unlabeled data.
[0131] Therefore, the total loss of unlabeled data is expressed as follows:
[0132] (21)
[0133] where, , , are weight coefficients, preferably = 0.06, = 1.0, = 0.01.
[0134] Finally, the overall loss of semi-supervised learning is expressed as:
[0135] (22)
[0136] where, , are weight coefficients, preferably = 1.0.
[0137] To verify the effectiveness of the method of the present invention in medical image analysis, a publicly available human blastocyst dataset was selected to evaluate the performance of the UniTask+ framework of the present invention. At the same time, the generality of the model was also tested on the left atrium (LA) segmentation dataset provided by Zhaohan et al. Among them, this human blastocyst dataset contains 249 microscopic images of human blastocysts and their corresponding ground truth (GT) labels, including the annotations of the main tissues of the zona pellucida (ZP), trophectoderm (TE), inner cell mass (ICM), and blastocoel provided by the Pacific Reproductive Medicine Center. For the binary segmentation task, the main goal is to identify the inner cell mass because the ICM is the basis of embryonic development. For the multi-class segmentation task, the targets include ZP, TE, ICM, and blastocoel, and the size of the input image was adjusted to 256×256 during the test. The LA dataset contains 100 sets of gadolinium-enhanced magnetic resonance imaging (GE-MRIs) images and labels, which is the most commonly used dataset for evaluating semi-supervised performance, and the size of the input image was adjusted to 112×112×80 during the test.
[0138] All tests were performed on an NVIDIA Tesla V100 GPU using PyTorch 1.10.0. To avoid overfitting, data augmentation methods were adopted to increase the diversity of training samples, including random scaling, random rotation, brightness, and contrast enhancement. The maximum number of training epochs was set to 2000, the initial learning rate was set to 0.0001, and the weight decay was set to 0.00005. The Adam optimizer was selected for the experiment. To evaluate the performance of the network, four evaluation criteria for the ICM are usually used: accuracy, recall, Dice coefficient (Dice), and Jaccard index (Jaccard). For ICM segmentation, the original UNet was selected as the backbone network. For LA segmentation, the same four evaluation metrics were used: Dice, Jaccard index, average surface distance (ASD), and 95% Hausdorff distance (95HD). For the experiment on the LA dataset, VNet was selected as the backbone network.
[0139] In the test validation conducted on the human blastocyst dataset, among 190 training images, 10% (19 scans) were used as labeled data, while the remaining 171 images were used as unlabeled data. The UniTask+ framework method of the present invention achieved accuracy, recall, Dice, and Jaccard metrics of 97.67%, 92.76%, 92.09%, and 86.75% respectively. Compared with other existing methods, the UniTask+ framework method of the present invention improved by at least 0.92% and 1.32% in Dice and Jaccard respectively. When 50% of the labeled data was used for training, the UniTask+ framework method of the present invention achieved better results, reaching accuracy, recall, Dice, and Jaccard metrics of 98.37%, 95.28%, 94.76%, and 90.68% respectively; and was superior to other methods by at least 1.42% and 1.85% in Dice and Jaccard metrics.
[0140] In the left atrium (LA) segmentation dataset test experiment, among 80 training scans, 10% (8 scans) were used as labeled data, while the remaining 72 scans were used as unlabeled data. The UniTask+ framework method of the present invention achieved Dice, Jaccard, ASD (voxel), and 95HD (voxel) score metrics of 90.61%, 83.03%, 6.80%, and 2.25% respectively. Its results were superior to other comparison methods by at least 1.78% and 2.72% in Dice and Jaccard metrics. When 20% of the labeled data was used for training, the UniTask+ framework method of the present invention achieved Dice, Jaccard, ASD (voxel), and 95HD (voxel) scores of 92.74%, 86.55%, 4.87%, and 1.58% respectively. Its results were superior to other methods by at least 1.43% and 2.47% in Dice and Jaccard metrics respectively.
[0141] The above test experiments show that there is an inherent nature because fine pixel-level information needs to be abstracted during the segmentation process. Through these dual-task certification efforts, the UniTask+ framework method promotes the exploration of inherent segmentation perturbations at three key levels: pixel level, geometric level, and point level. This enhances the learning dynamics and benefits both supervised and semi-supervised scenarios. The experimental evaluation of the present invention on the ICM, blastocyst, and left atrium datasets has reached the state-of-the-art level. These excellent performances prove the effectiveness and robustness of the UniTask+ framework in dealing with the complex challenges of medical image segmentation.
[0142] Through the above detailed description of the specific embodiments and examples of the present invention, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments and examples without departing from the principles and basic concepts of the present invention. The scope of the present invention is defined by the appended claims and their equivalents. The content not described in detail in this specification, such as the specific structure and implementation details of typical neural networks and their functional layers, can be achieved by those skilled in the art based on the well-known prior art and the technical teachings of the present invention.
Claims
1. A biological tissue recognition and processing method based on task consistency, characterized in that Includes steps: A network model based on a unified task-aware consistency interaction framework is constructed, the network model comprising a skeleton network for extracting multi-scale image features, a main branch, and at least two additional branches; wherein the skeleton network is based on an encoder-decoder architecture of a deep learning model, and the output result of the skeleton network is input to the main branch and the at least two additional branches; the main branch and the at least two additional branches are a neural network in a parallel relationship, wherein the main branch is used to perform a main task, the task performed by at least one additional branch has a strong correlation with the main task, and the task performed by at least another additional branch has a weak correlation with the main task; Acquire training sample data, where the training sample data is biological tissue image data, and a portion of the biological tissue image data is labeled data for supervised training; Using the training sample data to train the network model based on the unified task perception consistency interaction framework, to obtain a biological tissue recognition processing training model; Use biological tissue recognition and processing training models to recognize and process biological tissue image data; The main branch is a segmentation branch, which is used to generate a pixel-level probability map and generate a pixel-level output result; an additional branch is a level set branch, which is used to generate a target level set representation and generate a geometric output result, which has a strong correlation with the main task; another additional branch is a point set branch, which is used to synthesize a point representation related to the target structure and generate a point-level output result, which has a weak correlation with the main task; The biological tissue image data has dimensional structure, wherein C represents the number of color channels of the biological tissue image, H and W represent the height and width of the biological tissue image respectively, and R represents the real number space; the biological tissue image data is processed by the skeleton network to obtain feature sets of different scales Indicates that, the dimension , and Corresponding to the use of The downsampling factor of C, H, and W is processed; the integer D represents the degree of downsampling performed at each stage; Loss function when training with labeled data for: ; in, is the loss function for supervised training of the segmentation branch, is the loss function for supervised training of the point set branch, is the loss function for supervised training of the level set branch, is the strong and weak task-aware consistency function, , , , is the weight coefficient; Loss function when training with unlabeled data for: ; in, is the loss function for unsupervised training of the point set branch, is the loss function for unsupervised training of the level set branch, is the strong and weak task-aware consistency function, , , is the weight coefficient; in, , where is the loss function, Obtaining the skeleton point representation of the object for the point set branch; For the prediction results By formula The generated point set represents the feature; For the prediction results By formula Generate pixel representation features; It is the geometric representation obtained by the level set branch under the level set.
2. The biological tissue identification and processing method based on task consistency as claimed in claim 1, characterized in that: Another part of the biological tissue image data is unlabeled data for semi-supervised training; the segmentation branch performs supervised training, and the level set branch and the point set branch perform supervised training and semi-supervised training simultaneously.
3. The biological tissue identification and processing method based on task consistency as claimed in claim 2, characterized in that: The loss function for semi-supervised learning of the network model based on the unified task-aware consistency interaction framework The overall expression is: ; in, , is the weight coefficient, The loss function used for training with labeled data is used for the network model based on the unified task-aware consistency interaction framework. The loss function used when training with unlabeled data is adopted for network models based on a unified task-aware consistency interaction framework.
4. The biological tissue identification and processing method based on task consistency as claimed in claim 3, characterized in that: Using the training sample data to train the network model based on the unified task-aware consistency interaction framework includes: Step S100: Randomly initialize segmentation branch parameters , point set branch parameter , level set branch parameter ; Step S200: When the predetermined loss expectation or training times are not completed, execute: Step S210: Obtain each batch of training samples ;in, They are labeled datasets The corresponding input image and label map; For unlabeled dataset The corresponding input image; Step S220: Calculate the prediction results of the segmentation branch, point set branch, and level set branch respectively ; Step S230: Utilize Transformation formula through point set branching generate , through the transformation formula of the level set branch generate ; Step S240: Prediction results By formula Generate point set representation features ; Step S250: Prediction results By formula Generate level set representation features ; Step S260: Prediction results By formula Generate pixel representation features ; Step S270: Prediction results By formula Generate point set representation features ; Step S280: Obtaining the overall loss ; Step S290: Update model parameters by back propagation ; Step S300: When the predetermined loss expectation or training times are completed, the loop ends; Step S400: Return the obtained segmentation branch parameters , point set branch parameter , level set branch parameter .
Citation Information
Patent Citations
A Multi-Stage Intelligent Analysis Method and System Based on Embryonic Development Images
CN115049908B