Transform-based unsupervised cell segmentation method

By applying the unsupervised cell segmentation framework based on Transformer in the field of cell segmentation, combining multimodal alignment and MMX module, the problem of traditional methods dependence on annotation data is solved, and efficient segmentation and computational efficiency is improved in complex cell environments.

CN120047460AActive Publication Date: 2025-05-27HANGZHOU DIANZI UNIV

Patent Information

Application Number
CN202411672434.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-05-27
Estimated Expiration
2044-11-21

AI Technical Summary

Technical Problem

The effectiveness of traditional supervised learning methods is limited when dealing with rare or highly heterogeneous cell types, mainly because these methods rely heavily on a large number of excellent annotated data.

Method used

An unsupervised cell segmentation framework based on Transformer is adopted, combining multimodal alignment and open domain applicability of the Owl-vit model, and an adaptive zero-shot segmentation mechanism is realized through a low-rank attention mechanism and a matching matrix feature optimization module (MMX).

Benefits of technology

This method can efficiently handle complex cell environments in an unsupervised environment, improves computational efficiency and feature recognition processes, and significantly improves segmentation performance on different cell types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047460A_ABST
    Figure CN120047460A_ABST
Patent Text Reader

Abstract

The invention relates to an unsupervised cell segmentation method based on Transform, and the method is characterized in that a multi-modal text image alignment module aims at effectively fusing text and image data, and achieves the high alignment of multi-modal information through a Transform architecture; the mutual relevance of the data is enhanced through a low-rank attention mechanism, so that the multi-modal features can be extracted and aligned more accurately in an unsupervised environment. The matching matrix feature optimization module further processes the aligned feature data. According to the method, a unique matching matrix optimization algorithm is utilized, the precision of feature matching is remarkably improved, parameters of segmented cells are extracted and adjusted through the optimized matching matrix, and a more accurate initial prompt is provided for the subsequent segmentation process. The optimized features are input to an SAM segmentation module. And the SAM realizes high-precision cell segmentation by utilizing the strong segmentation capability of the SAM. The module gives full play to the advantages of a low-rank attention mechanism and matching matrix optimization, and ensures the accuracy and robustness of a segmentation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an unsupervised cell segmentation method based on Transformer, belonging to the technical field of medical image segmentation. Background Art

[0002] Traditional supervised learning has made some progress in the field of cell segmentation, but its effectiveness is limited when dealing with rare or highly heterogeneous cell types. This is mainly because these methods rely heavily on a large amount of excellent annotated data. Our solution to this problem is an innovative unsupervised cell segmentation framework based on Transformer, which combines the benefits of multimodal alignment and the open-domain applicability of the Owl-vit model. The pre-trained encoder directly inputs the high-dimensional features of the image into the image encoder. In addition to improving the model's ability to handle complex cell environments, the MMX module also enables the framework to segment cell types that have not been directly trained due to its adaptive zero-shot segmentation mechanism. In addition, a low-rank attention mechanism is implemented within the framework to improve computational efficiency and optimize the feature recognition process.

[0003] In summary, how to obtain high-quality complex inference datasets at low cost is a topic worthy of in-depth study. This special topic starts from the direction of the chain of thought and model self-enhancement to explore, solve the difficulties and key points of the current methods, and form a complete unsupervised cell segmentation method based on Transformer. Summary of the Invention

[0004] In order to overcome the deficiency of the existing optimized model in unsupervised cell segmentation ability, the present invention provides an unsupervised cell segmentation method based on Transformer, which utilizes the low-rank attention mechanism in a new way and embeds an MMX module to optimize cell parameters after zero-shot feature detection. Comprehensive experimental verification shows that the performance of the framework is better and can be generalized to different cell types in the segmentation task. Especially when dealing with cell categories, compared with the traditional excellent zero-shot segmentation model, significant improvements have been achieved in the segmentation performance on three public datasets.

[0005] An unsupervised cell segmentation method based on Transformer includes a multimodal text-image alignment module, a matching matrix feature optimization module MMX (MatchMatrix), and a SAM segmentation module. The method includes the following steps:

[0006] Step 1: Prepare the dataset, including three cell tissue datasets using hematoxylin and eosin (H&E) and immunohistochemical staining (IHC);

[0007] Step 2: Perform data preprocessing. The data preprocessing includes the preprocessing of images and text content. The main purpose is to enhance the features of cell images, increase data diversity, and obtain text word vectors;

[0008] Step 3: After preprocessing, perform multimodal alignment of image-text pairs. The initialized data is passed through a Transformer encoder to obtain high-dimensional features of text and vision. A projection layer is used to project the text and vision features into a shared multimodal representation space. A contrastive loss function is used to move unrelated images and texts far away from each other in the representation space and bring semantically related images and texts closer;

[0009] Step 4: After data alignment, map the feature vectors to the detection head, including the classification head and the regression head. For images, the classification head is responsible for obtaining the confidence scores of cell categories, and the regression head is responsible for linearly regressing the initial object detection box vector. For text, the text embedding information of the text is obtained;

[0010] Step 5: Use the obtained regression object detection position as the input to the Matching Matrix Feature Optimization Module (MMX). Extract local image features according to the detection box, mainly dealing with areas with a large number of cells. Use a vision Transformer encoder to obtain a new round of object positions, combine with the initial object detection box, and use the Non-Maximum Suppression (NMS) algorithm to eliminate overlapping bounding boxes in the object detection task, obtaining the final object boundary to be sent to the SAM model for segmentation.

[0011] Step 6: After obtaining the final object boundary, use it as the prompt box of SAM and use the powerful segmentation ability of SAM to segment the cells constrained by the text.

[0012] The data preparation steps in Step 1 are as follows:

[0013] 1.1: Prepare three publicly available cell datasets, the Kaggle dataset, the Kumar dataset, and the Lizard dataset. The three datasets contain 665, 226, and 30 stained images respectively.

[0014] 1.2: For the photos in each dataset, unify them to the PNG format for subsequent image feature extraction.

[0015] 1.3: The label formats corresponding to the original images are RLE, mat, and csv respectively. Unify them and output them in the csv format.

[0016] The data preprocessing method in Step 2 is as follows:

[0017] 2.1: The image is scaled to a unified size of 1024×1024, and the RGB three-channel images are weighted and averaged to obtain a single-channel grayscale image, which can reduce the computational complexity while removing color redundancy and improving robustness.

[0018] 2.2: Gaussian filtering is performed on the grayscale image to remove image noise and enhance the edge features of the cell image. Then, the pixel values of the filtered image are normalized to the standard range [0,1] to improve the model efficiency.

[0019] 2.3: The stained Kaggle (H&E), Kumar (H&E), and Lizard (IHC) can be observed closely and recognized by machine in terms of their morphology and color, and word segmentation and word vectors of the text during segmentation are added.

[0020] The multi-modal alignment method in step three is as follows:

[0021] Dataset D N has n images I N,n , and these images will be divided into M×M patches of a fixed size of P×P (Patch N,ni ∈R M×M×d , where d is the dimension of the image feature vector). Each patch is flattened into a one-dimensional vector, and after normalization, the corresponding position encoding (Position Encoding i ) is added to retain the spatial information of the image. For the convenience of subsequent processing by the Transformer, these image patches form a batch of data P N,n , and the i-th image patch vector is denoted as P N,n,i .

[0022] P N,n,i = Flatten(Patch N,ni ) + PositionEncoding i

[0023] P N,n = {P N,n,i | i ∈ {1, 2,..., M×M}

[0024] The corresponding text E of the image N,n is embedded as word vectors, and the corresponding position encoding is also added to the word vectors. The j-th word Word N,n,j ∈R l×d (where l is the number of words in the n-th text and d is the dimension of the feature vector), and E N,n,j is specifically expressed as:

[0025] E N,n,j = WordEmbedding(WordN,n,j ) + PositionEncoding j

[0026] The image patch vectors and word vectors are used to generate the subsequent image feature sequence I n,M×M and text feature sequence T n,l .

[0027] I n,M×M = [P N,n,1 , P N,n,2 , …, P N,n,M×M

[0028] T n,l = [E N,n,1 , E N,n,2 , …, E N,n,l .

[0029] This method mainly extracts and optimizes the specified features to be processed in each picture.

[0030] The method for obtaining the initial coordinate box and its confidence in step 4 is as follows:

[0031] After the input image and text feature sequence are processed by the self-attention and feed-forward neural network of the Transformer encoder layer, the image feature output F N,n,i and text feature output Text N,n,i are obtained. The feature F N,n,i is mapped to the detection head, and the detection head includes a classification head and a regression head. The classification head predicts the category of the target and generates the confidence C N,n,i , and the regression head is used to predict the bounding box B N,n,i of the target.

[0032] F N,n,i = ViT(I n,M×M )

[0033] Text N,n,i = TTE(T n,l )

[0034]

[0035] c is the number of categories of each image patch, and W c is the weight matrix of the classification branch. (F N,n,i · W c ) k represents the k-th element of the category score vector F N,n,i · W c , that is, the score of the k-th category.

[0036] ​Predict the bounding box parameters BP (center coordinates x, y and width and height w, h) for each image patch. The output of the regression branch is B N,n,i , and map it back to the coordinate system of the original image.

[0037] BP = F N,n,i ·W R

[0038]

[0039] W R is the weight matrix of the regression branch.

[0040] The method of the matching matrix feature optimization module (MMX) in the fifth step is as follows:

[0041] The MMX module crops the local image feature I N,n using B Tailor N,n,i .

[0042]

[0043] where i represents the i-th block in the set of bounding boxes, and B N,n, is the set of bounding boxes output by the regression branch.

[0044] After scaling and padding the cropped local image feature, the local image R N,n,i is obtained. It is divided into W×W P×P Patch blocks (here 16×16 pixel blocks are used). The Patch vector of each image patch t is flattened into a one-dimensional vector and added with position encoding.

[0045] R N,N,i = Pad(Resize(I Tailor N,n,i ))

[0046] Patch N,n,i,j = Patchify(R N,N,i )

[0047] P (t) N,n,i,j = Flatten(Patch N,n,i,j ) + PositionRncoding j

[0048] P (t) N,n,i = {P (t) N,n,i,j |j ∈ {1, 2,..., W×W}

[0049] where PatchN,n,i,j is the vector of the j-th image patch, is the set of patch vectors of the flattened image, D Patch is the vector dimension after flattening the image patch.

[0050] Project P (t) N,n,i,j linearly project the high-dimensional feature vector to a fixed feature space as the input feature of the Transformer to obtain the output feature F of the second round (t) N,n,i .

[0051] Proj N,n,i = W p ·P (t) N,n,i + b p

[0052] F (t) N,n,i = ViT(Proj N,n,i )

[0053] is the projected high-dimensional feature vector, is the weight matrix of the linear projection, is the bias vector, and D is the dimension after projection.

[0054] Map F (t) N,n,i to the detection head to obtain the corresponding class score C (t) N,n,i and the bounding box parameters B (t) N,n,i .

[0055]

[0056] c is the number of classes for each image patch, W c is the weight matrix of the classification branch, W R is the weight matrix of the regression branch, (F (t) N,n,i ·W c ) k represents the k-th element of the class score vector F (t) N,n,i ·W c i.e., the score of the k-th class.

[0057] The SAM segmentation method in step six:

[0058] By using hint clues in the image segmentation task, SAM utilizes the hint strategy in the field of natural language processing (NLP) to quickly segment any object and adapt to different downstream tasks.

[0059] The preprocessed image passes through a convolutional base, which divides the image into fixed-size patches. After normalization, each patch is flattened into a one-dimensional vector and the corresponding position encoding (PositionEncoding i ) is added to obtain the above vector set P N,n , and then its feature vectors are input into the Transformer to obtain the high-dimensional feature output F SAM N,n,i .

[0060] P N,n ={P N,n,i |i∈{1,2,...,M×M}

[0061] F SAM N,n,i =ViT(P N,n,i )

[0062] The Prompt Encoder converts each bounding box into a feature vector BB N,n,j .

[0063] BB N,n,j =PromptEncoder(B (t) N,n,i )

[0064] The image embedding and the prompt embedding are combined, and the segmentation mask M is generated through the Mask Decoder N,n,j .

[0065] M N,n,j =MaskDecoder(F SAM N,n,i ,BB N,n,j )

[0066] To obtain a more accurate prediction mask, since the generated mask contains connection values, further improvements can be made subsequently. Compared with the prior art, the beneficial effects of the present invention are as follows:

[0067] The present invention includes three core modules: a multimodal text-image alignment module, a matching matrix feature optimization module MMX (MatchMatrix), and a SAM segmentation module. First, the multimodal text-image alignment module aims to effectively fuse text and image data, achieving a high degree of alignment of multimodal information through the Transformer architecture. This process enhances the mutual correlation of data through the low-rank attention mechanism, enabling more accurate extraction and alignment of multimodal features in an unsupervised environment. Second, the matching matrix feature optimization module further processes the aligned feature data. This module uses a unique matching matrix optimization algorithm to significantly improve the accuracy of feature matching. Through the optimized matching matrix, the parameters of the segmented cells are extracted and adjusted, providing more accurate initial cues for the subsequent segmentation process. Finally, the optimized features are input into the SAM segmentation module. SAM (Segmentation-Aware Module) utilizes its powerful segmentation ability to achieve high-precision cell segmentation based on the input bounding box cues. This module fully exploits the advantages of the low-rank attention mechanism and matching matrix optimization to ensure the accuracy and robustness of the segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0069] Figure 1 It is a schematic structural diagram of an unsupervised cell segmentation method based on Transformer of the present invention;

[0070] Figure 2 It is a comparison diagram of the input sample (left side) of an unsupervised cell segmentation method based on Transformer of the present invention and the cell segmentation result mask and the ground truth label mask (right side). DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0072] An unsupervised cell segmentation method based on Transformer, including three core modules: a multimodal text-image alignment module, a matching matrix feature optimization module MMX (MatchMatrix), and a SAM segmentation module. The method includes the following steps:

[0073] Step 1: Prepare the dataset, including three cell tissue datasets stained with hematoxylin and eosin (H&E) and immunohistochemistry (IHC);

[0074] Step 2: Perform data preprocessing, which includes preprocessing of images and text content. The main purpose is to enhance the features of cell images, increase data diversity, and obtain text word vectors;

[0075] Step 3: After preprocessing, perform multimodal alignment of image-text pairs. The initialized data is passed through a transformer encoder to obtain high-dimensional features of text and vision. A projection layer is used to project the text and vision features into a shared multimodal representation space. A contrast loss function is used to move unrelated images and text away from each other in the representation space and bring semantically related images and text closer;

[0076] Step 4: After data alignment, map the feature vectors to the detection head, including a classification head and a regression head. For images, the classification head is responsible for obtaining the confidence scores of cell categories, and the regression head is responsible for linearly regressing the initial object detection box vector. For text, the text embedding information of the text is obtained;

[0077] Step 5: Use the obtained regression object detection positions as the input to the matching matrix feature optimization module (MMX). Extract local image features based on the detection boxes, mainly dealing with areas with a large number of cells. Use a vision transformer encoder to obtain a new round of object positions, combine with the initial object detection boxes, and use the non-maximum suppression (NMS) algorithm to eliminate overlapping bounding boxes in the object detection task, obtaining the final object boundaries to be fed into the SAM model for segmentation.

[0078] Step 6: After obtaining the final object boundaries, use them as the prompt boxes for SAM, and use the powerful segmentation ability of SAM to segment text-constrained cells.

[0079] The data preparation steps in Step 1 are as follows:

[0080] 1.1: Prepare three publicly available cell datasets, the Kaggle dataset, the Kumar dataset, and the Lizard dataset. The three datasets contain 665, 226, and 30 stained pictures respectively.

[0081] 1.2: For the photos in each dataset, unify them into the PNG format for subsequent image feature extraction.

[0082] 1.3: The label formats corresponding to the original images are RLE, mat, and csv respectively, and they are uniformly output in csv format.

[0083] The data preprocessing method in the second step is as follows:

[0084] 2.1: The image is scaled to a uniform size of 1024×1024, and the RGB three-channel images are weighted and averaged to obtain a single-channel grayscale image, which can reduce the color redundancy and improve the robustness while reducing the computational complexity.

[0085] 2.2: Gaussian filtering is performed on the grayscale image to remove image noise and enhance the edge features of the cell image. Then, normalization is used to scale the pixel values of the filtered image to the standard range [0,1] to improve the model efficiency.

[0086] 2.3: The stained Kaggle (H&E), Kumar (H&E), and Lizard (IHC) can be observed closely and recognized by machine in terms of their morphology and color, and word segmentation and word vectors of the text during segmentation are added.

[0087] The multi-modal alignment method in the third step is as follows:

[0088] Dataset D N There are n images I N,n , and these images will be divided into M×M patches of a fixed size of P×P (Patch N,ni ∈R M×M×d , where d is the dimension of the image feature vector). Each patch is flattened into a one-dimensional vector, and after normalization, the corresponding position encoding (Position Encoding i ) is added to retain the spatial information of the image. For the convenience of subsequent processing by the Transformer, these image patches form a batch of data P N,n , and the i-th image patch vector is denoted as P N,n,i .

[0089] P N,n,i = Flatten(Patch N,ni ) + PositionEncoding i

[0090] P N,n ={P N,n,i |i∈{1,2,...,M×M}

[0091] The corresponding text E of the image N,n is embedded as a word vector, and the corresponding position encoding is also added to the word vector. The j-th word Word in the n-th textN,n,j ∈R l×d (where l is the number of words in the nth text and d is the dimension of the feature vector), E N,n,j Specifically expressed as:

[0092] E N,n,j = WordEmbedding(Word N,n,j ) + PositionEncoding j

[0093] The image patch vector and the word vector are used to generate the subsequent image feature sequence I n,M×M and the text feature sequence T n,l .

[0094] I n,M×M = [P N,n,1 , P N,n,2 , …, P N,n,M×M

[0095] T n,l = [E N,n,1 , E N,n,2 , …, E N,n,l .

[0096] This method mainly extracts and optimizes the specified features to be processed in each picture.

[0097] The method for obtaining the initial coordinate box and its confidence in step 4 is as follows:

[0098] After the input image and the text feature sequence are processed by the self-attention and feed-forward neural network of the Transformer encoder layer, the image feature output F N,n,i and the text feature output Text N,n,i are obtained. Map the feature F N,n,i to the detection head, and the detection head includes a classification head and a regression head. The classification head predicts the category of the target and then generates the confidence C N,n,i , and the regression head is used to predict the bounding box B N,n,i of the target.

[0099] F N,n,i = ViT(I n,M×M )

[0100] Text N,n,i = TTE(T n,l )

[0101]

[0102] c is the number of categories of each image patch, W c is the weight matrix of the classification branch. (F N,n,i · W​c ) k represents the category score vector F N,n,i ·W c The k-th element of, i.e., the score of the k-th category.

[0103] Predict the bounding box parameters BP (center coordinates x, y and width and height w, h) of each image patch, and the output of the regression branch is B N,n,i , mapped back to the coordinate system of the original image.

[0104] BP = F N,n,i ·W R

[0105]

[0106] W R is the weight matrix of the regression branch.

[0107] The method of the matching matrix feature optimization module (MMX) in the fifth step is as follows:

[0108] The MMX module crops the local image feature I by using B N,n as the bounding box of the preliminary mapping Tailor N,n,i .

[0109]

[0110] where i represents the i-th block in the set of bounding boxes, and B N,n, is the set of bounding boxes output by the regression branch.

[0111] After scaling and filling the cropped local image feature, the local image R is obtained N,n,i , which is unfolded into W×W P×P Patch blocks (here 16×16 pixel blocks are used), and the Patch vector of each image patch t is flattened into a one-dimensional vector and added with position encoding.

[0112] R N,N,i = Pad(Resize(I Tailor N,n,i ))

[0113] Patch N,n,i,j = Patchify(R N,N,i )

[0114] P (t) N,n,i,j = Flatten(Patch N,n,i,j ) + PositionRncoding j

[0115] P(t) N,n,i = {P (t) N,n,i,j | j ∈ {1, 2, ..., W × W}

[0116] where Patch N,n,i,j is the j-th image patch vector, is the set of patch vectors of the flattened image, D Patch is the vector dimension after flattening the image patches.

[0117] Project P (t) N,n,i,j linearly into a high-dimensional feature vector and map it to a fixed feature space as the input feature of the Transformer to obtain the output feature F of the second round (t) N,n,i .

[0118] Proj N,n,i = W p · P (t) N,n,i + b p

[0119] F (t) N,n,i = ViT(Proj N,n,i )

[0120] is the high-dimensional feature vector after projection, is the weight matrix of the linear projection, is the bias vector, D is the dimension after projection.

[0121] Map F (t) N,n,i to the detection head to obtain the corresponding class scores C (t) N,n,i and bounding box parameters B (t) N,n,i .

[0122]

[0123] c is the number of classes for each image patch, W c is the weight matrix of the classification branch, W R is the weight matrix of the regression branch, (F (t) N,n,i · W c ) k represents the k-th element of the class score vector F (t) N,n,i · W c i.e., the score of the k-th class.

[0124] The SAM segmentation method in step six:

[0125] By using hint clues in the image segmentation task, SAM utilizes hint strategies in the field of natural language processing (NLP) to quickly segment any object and adjust to different downstream tasks.

[0126] The preprocessed image passes through a convolutional base, divides the image into fixed-size patches, flattens each patch into a one-dimensional vector after normalization, and adds the corresponding position encoding (Position Encoding i ), obtaining the above vector set P N,n , and then inputs its feature vector into the Transformer to obtain the high-dimensional feature output F SAM N,n,i .

[0127] P N,n ={P N,n,i |i∈{1,2,...,M×M}

[0128] F SAM N,n,i =ViT(P N,n,i )

[0129] The Prompt Encoder converts each bounding box into a feature vector BB N,n,j .

[0130] BB N,n,j =PromptEncoder(B (t) N,n,i )

[0131] Combining the image embedding and the prompt embedding, generate the segmentation mask M through the Mask Decoder N,n,j .

[0132] M N,n,j =MaskDecoder(F SAM N,n,i ,BB N,n,j )

[0133] To obtain a more accurate prediction mask, since the generated mask contains connection values, further improvement can be made subsequently.

[0134] The above has described the embodiments of the present invention in detail in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations made to these embodiments still fall within the protection scope of the present invention.

Claims

1. A Transformer-based unsupervised cell segmentation method, characterized in that: The following steps are involved: Step 1: Prepare the dataset, including three cell tissue datasets stained with hematoxylin and eosin and immunohistochemistry; Step 2: Perform data preprocessing, including preprocessing of images and text content; Step 3: After preprocessing, perform multimodal alignment of image-text pairs. Pass the initialization data through the transformer encoder to obtain high-dimensional features of text and vision. Use the projection layer to project the text and visual features into a shared multimodal representation space. Use the contrast loss function to keep unrelated images and texts away from each other in the representation space and bring semantically related images and texts closer together. Step 4: After data alignment, the feature vector is mapped to the detection head, including the classification head and the regression head. For images, the classification head is responsible for obtaining the confidence score of the cell category, and the regression head is responsible for linear regression of the initial target detection box vector. For text, the text embedding information of the text is obtained; Step 5: Get the regressed target detection position as the input of the matching matrix feature optimization module, extract local image features according to the detection frame, and use the visual transformer encoder to get a new round of target positions for areas with more cells. Combine the initial target detection frame and use the non-maximum suppression algorithm to eliminate overlapping bounding boxes in the object detection task, and get the final target boundary sent to the SAM model for segmentation; Step 6: After obtaining the final target boundary, use SAM’s powerful segmentation capability as the prompt box to segment out the cells constrained by the text.

2. The unsupervised cell segmentation method based on Transformer according to claim 1, characterized in that: The steps for preparing the data set in step 1 are as follows: 1.1: Prepare three public cell datasets: Kaggle dataset, Kumar dataset and Lizard dataset; 1.2: Unify the photos of each dataset into PNG format; 1.3: The label formats corresponding to the original images are RLE, mat, and csv, which are uniformly output in csv format.

3. The unsupervised cell segmentation method based on Transformer according to claim 1, characterized in that: The data preprocessing method in step 2 is as follows: 2.1: Scale the image to a uniform size of 1024×1024, and perform weighted averaging of the three RGB channel images to obtain a single-channel grayscale image, removing color redundancy to improve robustness while reducing computational complexity; 2.2: Perform Gaussian filtering on the grayscale image to remove image noise and enhance the edge features of the cell image. Post-normalization is then performed to scale the pixel values ​​of the filtered image to the standard range [0,1] to improve model efficiency. 2.3: After dyeing, Kaggle, Kumar and Lizard can be observed closely and recognized by machine in terms of shape and color, and the word segmentation and word vector of the text are added during segmentation.

4. The unsupervised cell segmentation method based on Transformer according to claim 1, characterized in that: The multimodal alignment method in step 3 is as follows: Dataset D N There are n images I N,n , the image will be divided into M×M patches of fixed size P×P. N,ni ∈R M×M×d , where d is the dimension of the image feature vector, each block is flattened into a one-dimensional vector, and the corresponding position code is added after standardization to retain the spatial information of the image. After processing the subsequent Transformer, the image blocks form a batch data P N,n , the i-th image block vector is represented as P N,n,i , P N,n,i =Flatten(Patch N,ni )+PositionEncoding i P N,n ={P N,n,i |i∈{1,2,...,M×M} The text E corresponding to the image N,n Embedded as word vector, the word vector also adds the corresponding position encoding, the jth word in the nth text Word N,n,j ∈R l×d , where l is the number of words in the nth text, d is the dimension of the feature vector, and E N,n,j Specifically expressed as: E N,n,j =WordEmbedding(Word N,n,j )+PositionEncoding j Image block vectors and word vectors are used to generate subsequent image feature sequences I n,M×M and text feature sequence T n,l , I n,M×M =[P N,n,1 ,P N,n,2 ,…,P N,n,M×M ] T n,l =[And N,n,1 ,AND N,n,2 ,…,AND N,n,l ]。 5. The unsupervised cell segmentation method based on Transformer according to claim 1, characterized in that: The step 4 specifically includes: After the input image and text feature sequence are processed by the self-attention and feedforward neural network of the Transformer encoder layer, the image feature output F is obtained. N,n,i And text feature output Text N,n,i , the feature F N,n,i Mapped to the detection head, the detection head includes the classification head and the regression head. The classification head predicts the category of the target and generates the confidence C N,n,i , the regression head is used to predict the bounding box B of the target N,n,i , F N,n,i =ViT(I n,M×M ) Text N,n,i =TTE(T n,l ) c is the number of categories for each image block, W c The weight matrix of the classification branch, (F N,n,i ·W c ) k Represents the category score vector F N,n,i ·W c The k-th element of is the score of the k-th category, Predict the bounding box parameters BP of each image block, and the output of the regression branch is B N,n,i , mapped back to the coordinate system of the original image, BP = F N,n,i ·W R W R is the weight matrix of the regression branch.

6. The unsupervised cell segmentation method based on Transformer according to claim 1, characterized in that: The matching matrix feature optimization module method in step 5 is as follows: MMX modules use B N,n As a preliminary mapping bounding box to crop the local image features I Tailor N,n,i , Where i represents the i-th block in the bounding box set, B N,n, is the set of bounding boxes output by the regression branch, The cropped local image features are scaled and filled to obtain the local image R N,n,i , expand it into W×W P×P patches, and the patch vector of each image block t is flattened into a one-dimensional vector and added with position encoding. R N,N,i =Pad(Resize(I Tailor N,n,i )) Patch N,n,i,j =Patchify(R N,N,i ) P (t) N,n,i =Flatten(Patch N,n,i,j )+PositionRncoding j P (t) N,n,i ={P (t) N,n,i,j |j∈{1,2,...,W×W} Patch N,n,i,j is the j-th image block vector, is the set of block vectors of the flattened image, D Patch is the vector dimension after flattening the image block, P (t) N,n,i,j The high-dimensional feature vector of linear projection is mapped to a fixed feature space as the input feature of Transformer to obtain the output feature F of the second round (t) N,n,i , Project N,n,i =W p ·P (t) N,n,i +b p F (t) N,n,i =ViT(Proj N,n,i ) is the high-dimensional feature vector after projection, is the weight matrix of the linear projection, is the bias vector, D is the dimension after projection, F ’ N,n,i Mapped to the detection head to get the corresponding category score C (t) N,n,i and the bounding box parameter B (t) N,n,i , c is the number of categories for each image block, W c The weight matrix of the classification branch, W R is the weight matrix of the regression branch, (F (t) N,n,i ·W c ) k Represents the category score vector F (t) N,n,i ·W c The k-th element of is the score of the k-th category.

7. The unsupervised cell segmentation method based on Transformer according to claim 1, characterized in that: The SAM segmentation method in step 6 specifically includes: By using hints in image segmentation tasks, SAM leverages hint strategies from the field of natural language processing to quickly segment arbitrary objects and adjust them for different downstream tasks. The preprocessed image passes through a convolution base, which divides the image into patches of fixed size. After normalization, each patch is flattened into a one-dimensional vector and the corresponding position encoding is added. i , we get the above vector set P N,n , input its feature vector into Transformer to obtain high-dimensional feature output F SAM N,n,i , P N,n ={P N,n,i |i∈{1,2,...,M×M} F SAM N,n,i =ViT(P N,n,i ) PromptEncoder converts each bounding box into a feature vector BB N,n,j , BB N,n,j =PromptEncoder(B (t) N,n,i ) Combine the image embedding and the hint embedding, and generate the segmentation mask M through MaskDecoder N,n,j , M N,n,j =MaskDecoder(F SAM N,n,i ,BB N,n,j )。

Citation Information

Patent Citations

  • Multi-modal microscopic image cell segmentation method based on convolutional neural network

    CN116229457A

  • Image segmentation method and system based on multi-modal dialogue language model

    CN117036706A

  • Method for identifying cross-modal features from spatially resolved datasets

    CN118176527A

  • Transform weak supervision semantic segmentation method combined with context attention

    CN118411522A

  • Bone marrow cell recognition system based on multi-modal enhancement

    CN118865373A

Cited By

  • RT-Deet intelligent traffic target detection optimization method and system fused with Prompt prompt mechanism

    CN120388248A