Cross-graph context aware cell nucleus instance segmentation method and system

Through a cross-image context-aware cell nucleus instance segmentation method, image features are integrated using feature extraction and context injection modules, which solves the problem of the inability to effectively utilize image context in existing technologies and achieves more efficient cell nucleus instance segmentation.

CN120823602AActive Publication Date: 2025-10-21SUZHOU UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511341115.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-10-21
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Existing technologies cannot effectively utilize the contextual information of images in cell nucleus instance segmentation, resulting in insufficient segmentation performance, especially when cell nuclei are densely distributed in cell images, which easily leads to missed detections.

Method used

A cross-image context-aware cell nucleus instance segmentation method is adopted. The feature extraction module extracts the prior features and domain features of the cell image, and uses a multi-level feature refinement block to refine the features. The context information of the surrounding slices is integrated with the context injection module to generate image features. The hint encoding module is used to generate point embeddings of the cell nucleus. Finally, the mask decoder and texture encoder are used to generate the final segmentation result.

Benefits of technology

The accuracy and efficiency of cell nucleus instance segmentation are improved, missed detections are reduced, the method adapts to the complex structure of cell images, and the segmentation performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823602A_ABST
    Figure CN120823602A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a cross-graph context-aware cell nucleus instance segmentation method and system, and the method comprises the steps: obtaining a cell image, segmenting the cell image into slice groups, and constructing a cell nucleus instance segmentation model which comprises a feature extraction module, a context injection module, a prompt coding module, a mask decoder and a texture encoder; a feature extraction module extracts prior features and domain features of the slices, refines the prior features through a multi-stage feature refinement block, and extracts and fuses missing features from the domain features to obtain enhanced features; a context injection module integrates context information according to the enhanced features to generate image features, a prompt coding module generates point embedding, a mask decoder obtains a mask of each cell nucleus according to the image features and the point embedding, and the masks are combined to obtain an instance segmentation result; a texture encoder encodes the enhancement feature and the instance segmentation result into a texture feature, which is stored in a context injection module together with the enhancement feature. The embodiment segmentation effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a cross-image context-aware cell nucleus instance segmentation method and system. Background Art

[0002] Nucleus instance segmentation is a fundamental step in extracting meaningful biological information. Identifying the relative topology, size, and shape of nuclei is crucial for downstream tasks such as pathological image diagnosis and biological cell research. Consequently, this area of ​​research has attracted increasing attention. Nucleus instance segmentation is typically performed on digitized cell images, such as whole-slide images (WSIs). However, these digitized cell images have extremely high resolution and contain numerous nuclei, making instance segmentation challenging.

[0003] Existing techniques use deep learning for instance segmentation of cell nuclei, such as automated cell nucleus segmentation models like HoVer-Net, U-Net, and nn-UNet. These models use encoders to extract features from images and decoders to process these features into the final segmentation results. However, these models only extract features from the current input image and fail to consider contextual information from surrounding images (including the spatial distribution and tissue structure of cell nuclei). Consequently, it is difficult to achieve optimal segmentation results by relying solely on local features in the input image.

[0004] SAM 2 is a foundational model for video segmentation that includes a streaming memory for real-time video processing and instantaneous propagation. It has demonstrated excellent performance and generalization on large-scale datasets, such as Segment Anything Video (SA-V). Therefore, SAM 2 has potential for instance segmentation. However, applying SAM 2 to cell images presents several challenges. First, SAM 2 is not designed for cell image datasets and cannot extract domain-specific features. Directly applying SAM 2 can degrade segmentation performance. Second, although SAM 2 supports instantaneous propagation and can segment the entire video using only cues from the first frame, cell nuclei are relatively densely distributed in cell images. Capturing all nuclei instances using instantaneous propagation alone can hinder segmentation performance and lead to missed detections. Summary of the Invention

[0005] To this end, the technical problem to be solved by the present invention is to overcome the deficiencies in the existing technology and provide a cross-graph context-aware cell nucleus instance segmentation method and system, which can effectively extract domain features, integrate historical context information, and improve the effect of cell nucleus instance segmentation.

[0006] To solve the above technical problems, the present invention provides a cross-image context-aware cell nucleus instance segmentation method, comprising: Acquire a cell image and scan and segment it into a slice group including a plurality of slices, and construct a cell nucleus instance segmentation model, wherein the cell nucleus instance segmentation model includes a feature extraction module, a context injection module, a hint encoding module, a mask decoder, and a texture encoder; The feature extraction module extracts the prior features and domain features of the slice through a multi-layer encoder, refines the prior features of each layer through a multi-level feature refinement block, extracts and fuses the missing features from the domain features of each layer, and obtains enhanced features; The context injection module integrates context information from surrounding segmented slices according to the enhanced features to generate image features, the hint encoding module generates a point embedding of each cell nucleus according to the domain features, and the mask decoder obtains a mask of each cell nucleus according to the image features and the point embedding; The mask of each cell nucleus is combined to obtain the instance segmentation result of the slice, and the instance segmentation results of all slices are combined to obtain the instance segmentation result of the cell image; The instance segmentation result and the enhanced features of the slice are input into the texture encoder to generate texture features, and the texture features and the enhanced features are stored in the context injection module to provide context information for the next slice.

[0007] Furthermore, the feature extraction module includes a CNN encoder and a ViT encoder, the CNN encoder is a multi-layer convolutional neural network encoder, and the ViT encoder includes a Patch Embedded module and a multi-layer feature refinement layer. The number of layers of the feature refinement layer is the same as the number of layers of the convolutional neural network encoder. Each layer of the feature refinement layer includes multiple feature refinement blocks, and each feature refinement block includes the multi-level feature refinement block and the ViT block.

[0008] Furthermore, the feature extraction module extracts the prior features and domain features of the slice through a multi-layer encoder, refines the prior features of each layer through a multi-level feature refinement block, extracts and fuses the missing features from the domain features of each layer, and obtains enhanced features, specifically: When each slice is input into the feature extraction module, the i-th layer of the CNN encoder outputs the i-th layer of domain features; The ViT encoder first extracts the first layer of prior features through the Patch Embedded module, fuses the i-th layer of prior features and the i-th layer of domain features through the i-th layer of feature refinement layer to obtain the i+1-th layer of prior features, and uses the output of the last layer of feature refinement layer as the enhanced feature.

[0009] Furthermore, the calculation method of the i+1th layer prior features is: , in, is the i+1th layer prior feature, ViT( ) is the operation of the ViT block, is the i-th layer prior feature, MFRB( ) is the operation of the multi-level feature refinement block, [ ] n Indicates that n ViT(MFRB()) operations are performed, where n is the number of feature refinement blocks in the feature refinement layer; In the i-th feature refinement layer, the feature of the j-th feature refinement block is recorded as , j=1,2…,n, That is , the features obtained after the multi-level feature refinement block in the jth feature refinement block are recorded as , the operations of the multi-level feature refinement block are as follows: , in, is the cross attention operation, LN( ) is the layer normalization operation, is the domain feature of the i-th layer, After refining domain feature representation , The result obtained after the operation of the ViT block is .

[0010] Furthermore, the The calculation method is: , Among them, Adapter( ) is the Adapter fine-tuning operation.

[0011] Furthermore, the context injection module generates image features by integrating context information from surrounding segmented slices according to the enhanced features, specifically: The context injection module includes an environment memory library and a texture memory library. For the i-th slice of the current input, the environment memory library stores the enhanced features corresponding to the previous i-1 slices, which are recorded as , j=1,2,…,i-1; the enhanced feature corresponding to the i-th slice is ,when When entering the context injection module, the closest The M1 enhanced features construct the shortest distance memory library, recorded as ;application Injecting surrounding environment information Get the features with environmental information, recorded as ,Will Input texture memory; The texture memory bank stores the texture features of Mt slices of completed instance segmentation with the largest similarity difference, and selects the texture memory bank and The closest M2 texture features build the maximum similarity memory bank, recorded as ;application Injecting foreground information Get the image features, record as .

[0012] Furthermore, the shortest distance memory bank is: , For the closest The kth feature among the M1 enhanced features; described The calculation method is: ,in, attention operations for memory; The maximum similarity memory bank is: , For the closest The qth feature among the M2 texture features; described The calculation method is: .

[0013] Furthermore, when the texture memory stores the texture features of Mt slices that have completed instance segmentation and have the maximum similarity difference, the storage method is: The texture feature corresponding to the current i-th slice is recorded as ,Will The confidence level is recorded as ; The texture features of the Mt slices of completed instance segmentation stored in the texture memory are recorded as ,j=1,2,…,Mt;calculate The sum of the similarities between any texture feature and all other texture features, select the texture feature corresponding to the maximum sum, and record it as ,Will The confidence level is recorded as ; like Greater than Then delete it from the texture memory and store , otherwise the texture memory remains unchanged.

[0014] Furthermore, the hint encoding module generates a point embedding for each cell nucleus based on the domain features, specifically: The hint encoding module includes a prediction head and a hint encoder, and the prediction head includes a regression head and a classification head; The prediction head predicts the points of each cell nucleus based on the domain features, and the hint encoder encodes the points into point embeddings.

[0015] The present invention also provides a cross-graph context-aware cell nucleus instance segmentation system, comprising: An image acquisition module, which acquires cell images and scans and segments the images into a slice group including a plurality of slices; A model construction module constructs a cell nucleus instance segmentation model, which includes a feature extraction module, a context injection module, a hint encoding module, a mask decoder, and a texture encoder. The feature extraction module extracts the prior features and domain features of the slice through a multi-layer encoder, refines the prior features of each layer through a multi-level feature refinement block, extracts and fuses missing features from the domain features of each layer, and obtains enhanced features. The context injection module integrates context information from surrounding segmented slices based on the enhanced features to generate image features. The hint encoding module generates a point embedding for each cell nucleus based on the domain features. The mask decoder obtains a mask for each cell nucleus based on the image features and the point embedding. The instance segmentation module combines the mask of each cell nucleus to obtain the instance segmentation result of the slice, and combines the instance segmentation results of all slices to obtain the instance segmentation result of the cell image; the instance segmentation result and the enhanced features of the slice are input into the texture encoder to generate texture features, and the texture features and enhanced features are stored in the context injection module to provide context information for the next slice.

[0016] The above technical solution of the present invention has the following beneficial effects compared with the prior art: The present invention extracts the prior features and domain features of cell images through a feature extraction module, and uses a multi-level feature refinement block to refine the prior features and integrate the domain features to achieve effective extraction of domain features; on this basis, the historical context information around each cell nucleus in the cell image is integrated through a context injection module, making the model more adaptable to the cell image and improving the effect of cell nucleus instance segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein:

[0018] Figure 1 Flowchart of the method in the preferred embodiment of the present invention.

[0019] Figure 2 Schematic diagram of the structure of the cell nucleus instance segmentation model in a preferred embodiment of the present invention.

[0020] Figure 3 Schematic diagram of the structure of the feature refinement layer in a preferred embodiment of the present invention.

[0021] Figure 4 Schematic diagram of the structure of a multi-level feature refinement block in a preferred embodiment of the present invention.

[0022] Figure 5 Schematic diagram of the structure of the context injection module in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0023] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.

[0024] Reference Figure 1 As shown, the present invention discloses a cross-graph context-aware cell nucleus instance segmentation method, comprising the following steps:

[0025] S1: Acquire a cell image and scan and segment it into a slice group consisting of multiple slices. The entire cell image obtained by scanning is regarded as a video, and each slice in the slice group can be regarded as a video frame. The cell image is used as the input image, denoted as I. In this embodiment, a sliding window is used to scan I in a counterclockwise direction starting from the center to generate a slice group, denoted as P, where P={P1,…,P i ,…,P N}, P i represents the slice of the i-th input, and N represents the number of slices in the slice group.

[0026] S2: Build Figure 2 The cell nucleus instance segmentation model shown in the figure is an improvement based on the existing SAM 2 model. The cell nucleus instance segmentation model includes a feature extraction module, a context injection module, a hint encoding module, a mask decoder and a texture encoder.

[0027] S3: The feature extraction module extracts the prior features and domain features of the slice through a multi-layer encoder, refines the prior features of each layer through a multi-level feature refinement block (MFRB), extracts and fuses the missing features from the domain features of each layer, and obtains enhanced features, which are recorded as ; This enables effective extraction of specific domain features in slices.

[0028] The feature extraction module includes a CNN encoder and a ViT encoder. The CNN encoder is a multi-layer convolutional neural network encoder. The CNN encoder in this embodiment uses a common ConvNeXt encoder with a 4-layer structure. The ViT encoder is based on Hiera and includes a Patch Embedded module and a multi-layer feature refinement layer. The number of layers of the feature refinement layer is the same as the number of layers of the convolutional neural network encoder. In this embodiment, the feature refinement layer also has 4 layers, and each layer of the feature refinement layer includes multiple feature refinement blocks, such as Figure 3 As shown, each feature refinement block includes the multi-level feature refinement block and the ViT block. The multi-level feature refinement block is used to facilitate information flow between the CNN encoder and the ViT encoder (i.e., the ViT branch and the CNN branch). In this embodiment, the four feature refinement layers have K1, K2, K3, and K4 feature refinement blocks, respectively. The values ​​of K1, K2, K3, and K4 are adjusted based on actual needs. The Patch Embedded module is the Patch Embedded module in the ViT model, and the ViT block is the single-layer Transformer Encoder in the ViT model.

[0029] When each slice is input into the feature extraction module, the CNN encoder layer i outputs the i-th domain feature, denoted as , i=L1, L2, L3, L4, representing the 1st layer, 2nd layer, 3rd layer, and 4th layer respectively.

[0030] The ViT encoder first extracts the first layer of prior features through the Patch Embedded module, which is recorded as , the i+1th layer of prior features is obtained by fusing the i-th layer of prior features and the i-th layer of domain features through the i-th layer of feature refinement layer, and the output of the last layer of feature refinement layer is used as the enhanced feature, that is, .

[0031] The calculation method of the i+1th layer prior features is: , in, is the i+1th layer prior feature, ViT( ) is the operation of the ViT block, is the i-th layer prior feature, MFRB( ) is the operation of the multi-level feature refinement block, [ ] n Indicates that n ViT(MFRB()) operations are performed, where n is the number of feature refinement blocks in the feature refinement layer. In this embodiment, the four feature refinement layers perform K1, K2, K3, and K4 ViT(MFRB()) operations respectively. In the i-th feature refinement layer, the feature of the j-th feature refinement block is recorded as , j=1,2…,n, That is ,like Figure 4 As shown in Figure 2, the operations of the multi-level feature refinement block are as follows: , in, is the Cross-Attention operation, LN( ) is the layer normalization operation, After refining domain feature representation , The result obtained after the operation of the ViT block is ; The calculation method is: , Adapter( ) is an adapter fine-tuning operation, used for initial fine-tuning to include domain knowledge.

[0032] S4: The context injection module integrates the context information from the surrounding segmented slices according to the enhanced features to generate image features, which are recorded as .

[0033] In order to help the model obtain contextual information from surrounding slices when segmenting the current slice, the present invention designs a context injection module (CIM). The context injection module extends the memory module of SAM 2 (memory bank and memory attention in SAM 2) and can combine contextual information from surrounding slices during model prediction, thereby achieving more accurate mask prediction. Figure 5 As shown, the context injection module includes an environment memory bank (EMB) and a texture memory bank (TMB). The environment memory bank is used to supplement the surrounding tissue microenvironment information (i.e., background information), and the texture memory bank provides additional foreground information (texture and size of the cell nucleus, etc.). Therefore, CIM can effectively extract and utilize the surrounding context information.

[0034] For the i-th slice of the current input, the environment memory bank stores the enhanced features corresponding to the previous i-1 slices, recorded as , j=1,2,…,i-1; the enhanced feature corresponding to the i-th slice is ,when When entering the context injection module, the closest The M1 enhanced features construct the shortest distance memory library, recorded as , , For the closest The kth feature among the M1 enhanced features; in this embodiment, the closest When using the M1 enhanced features, the Euclidean distance is used to judge the distance, that is, to calculate and The Euclidean distance between ,pass The value of determines the distance. Application Injecting surrounding environment information Get the features with environmental information, recorded as ,Will Enter the texture memory library. The calculation method is: ,in, ( ) is the memory attention operation. Since the environment memory is used to store and provide enhanced features from the same image, it is cleared every time the input image is changed.

[0035] In order to better provide The texture memory stores the texture features of Mt slices with the largest similarity difference that have completed instance segmentation. The closest M2 texture features build the maximum similarity memory bank, recorded as , , For the closest The qth feature among the M2 texture features; in this embodiment, the closest When using M2 texture features, cosine similarity is used to judge the distance. Injecting foreground information Get the image features, record as , The calculation method is: .

[0036] When the texture memory stores the texture features of Mt slices that have completed instance segmentation and have the largest similarity difference, the storage method is: the texture feature corresponding to the current i-th slice is recorded as ,Will The confidence level is recorded as The texture features of the Mt slices of completed instance segmentation stored in the texture memory are recorded as , j=1,2,…, Mt, Mt is the length of the texture memory bank; calculate The sum of the similarities between any texture feature and all other texture features, select the texture feature corresponding to the maximum sum, and record it as ,Will The confidence level is recorded as ;like Greater than Then delete it from the texture memory and store Otherwise, the texture memory remains unchanged. By storing the texture features of Mt slices with the largest similarity difference that have completed instance segmentation in the texture memory, the reliability and diversity of the features are achieved.

[0037] The existing SAM 2 model's memory module utilizes foreground information from previously segmented slices to aid segmentation of the current block. However, its first-in, first-out access strategy fails to effectively provide the necessary foreground information. Furthermore, background information (such as the tissue microenvironment) plays a crucial role in identifying cell nucleus instances, which SAM 2 also fails to provide. To address these issues, the present invention designs a CIM to supplement foreground information (e.g., cell nucleus texture and size) with background information.

[0038] S5: The hint encoding module generates a point embedding for each cell nucleus based on the domain features to reduce the number of missed detections.

[0039] In this embodiment, the hint coding module includes a prediction head and a hint encoder. In this embodiment, the hint encoder used is the hint encoder in SAM 2. The prediction head includes a regression head and a classification head.

[0040] The prediction head predicts the points of each cell nucleus according to the domain features, and the prompt encoder encodes the points into point embeddings, recorded as The point of the cell nucleus refers to the cell nucleus that can be identified on the current slice. Once the cell nucleus is identified, a point is generated at the position of the cell nucleus on the current slice. In this embodiment, some points are first pre-scattered on the slice, and then the regression head predicts the offset of the points based on the characteristics of the slice, thereby changing the position of all points; then, the classification head determines whether these points belong to the cell nucleus or the background. The combination of the two can finally locate each cell nucleus, that is, generate a point on each cell nucleus.

[0041] S6: Mask decoder based on image features and point embedding The mask of the cell nucleus corresponding to each point, that is, the mask of each cell nucleus, is predicted.

[0042] S7: Combine the masks of each cell nucleus to obtain the instance segmentation result of the slice, and combine the instance segmentation results of all slices to obtain the instance segmentation result of the cell image.

[0043] S8: Inputting the instance segmentation result and the enhanced features of the slice into the texture encoder to generate texture features, and storing the texture features and the enhanced features together into the context injection module to provide context information for the next slice.

[0044] In this embodiment, the instance segmentation result and the enhanced features of the slice are passed through a texture encoder to obtain texture features, which are recorded as , the texture features are stored in the texture memory bank in the context injection module, and the enhanced features (denoted as ) is stored in the environment memory in the context injection module to provide context information for the next slice.

[0045] The present invention also discloses a cross-graph context-aware cell nucleus instance segmentation system, comprising: An image acquisition module, which acquires cell images and scans and segments the images into a slice group including a plurality of slices; A model construction module constructs a cell nucleus instance segmentation model, which includes a feature extraction module, a context injection module, a hint encoding module, a mask decoder, and a texture encoder. The feature extraction module extracts the prior features and domain features of the slice through a multi-layer encoder, refines the prior features of each layer through a multi-level feature refinement block, extracts and fuses missing features from the domain features of each layer, and obtains enhanced features. The context injection module integrates context information from surrounding segmented slices based on the enhanced features to generate image features. The hint encoding module generates a point embedding for each cell nucleus based on the domain features. The mask decoder obtains a mask for each cell nucleus based on the image features and the point embedding. The instance segmentation module combines the mask of each cell nucleus to obtain the instance segmentation result of the slice, and combines the instance segmentation results of all slices to obtain the instance segmentation result of the cell image; the instance segmentation result and the enhanced features of the slice are input into the texture encoder to generate texture features, and the texture features and enhanced features are stored in the context injection module to provide context information for the next slice.

[0046] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, which implements a cross-graph context-aware cell nucleus instance segmentation method when the computer program is executed by a processor.

[0047] The present invention also discloses a device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a cross-graph context-aware cell nucleus instance segmentation method is implemented.

[0048] Compared with the prior art, the advantages of the present invention are: 1. The feature extraction module is used to extract the prior features and domain features of the cell image, and the multi-level feature refinement block is used to refine the prior features and integrate the domain features to achieve effective extraction of specific domain features.

[0049] 2. The context injection module integrates the historical context information around each cell nucleus in the cell image, making the model more adaptable to the cell image and improving the effect of cell nucleus instance segmentation.

[0050] 3. By adding a prediction head including a regression head and a classification head after the convolution branch to automatically generate cell nucleus points, compared with the manual point annotation process in SAM 2, efficiency is improved and labor costs are reduced.

[0051] To further demonstrate the beneficial effects of the present invention, in this embodiment, the present invention is used to perform the prior art training of U-Net, U-Net, Mask R-CNN, HoVer-Net, SMILE (see the paper "Pan, **peng, et al. "SMILE: Cost-sensitive multi-task learning for nuclear segmentation and classification with imbalanced annotations." Medical Image Analysis 88 (2023): 102867."), PointNu-Net (see the paper "Yao, Kai, et al. "Pointnu-net: Keypoint-assisted convolutional neural network for simultaneousmulti-tissue histology nuclei segmentation and classification." IEEE Transactions on Emerging Topics in Computational Intelligence 8.1 (2023): 802-813."), SAM (Zero-Shot) (see the paper "Kirillov, Alexander, et al. "Segmentanything." Proceedings of the IEEE / CVF international conference on computervision. 2023.”), MedSA (for details, see the paper “Wu J, Ji W, Liu Y, et al. Medical SAMAdapter: Adapting Segment Anything Model for Medical Image Segmentation[J]. 2023.”), HQ-SAM (for details, see the paper “Ke, Lei, et al. "Segment anything in high quality." Advances in Neural Information Processing Systems 36 (2023): 29914-29934.”), CellViT (for details, see the paper “Hrst F, Rempe M, Heine L, et al.CellViT: VisionTransformers for precise cell segmentation and classification[J].MedicalImage Analysis, 2024, 94(000).DOI:10.1016 / j.media.2024.103143.”), SAM2, MedSAM2 (for details, see the paper “Zhu J, Hamdi A, Qi Y, et al.Medical SAM 2: Segment medical imagesas video via Segment Anything Model 2[J]. 2024.”) for cell nucleus instance segmentation simulation experiments. Among these methods, U-Net, U-Net, Mask R-CNN, HoVer-Net, SMILE, PointNu-Net, SAM (Zero-Shot), MedSA, HQ-SAM, and CellViT do not generate point embeddings. SAM2 and MedSAM2 generate point embeddings by manually annotating points and then encoding them through the hint encoder.

[0052] The Dice coefficient (Dice), Aggregate Jaccard Index (AJI), Detection Quality (DQ), Segmentation Quality (SQ), and Panoramic Quality (PQ) were used as evaluation metrics to assess the performance of cell nucleus instance segmentation. The experimental results are shown in Tables 1 and 2.

[0053] Table 1. Cell nucleus instance segmentation results of different methods on the CPM-17 dataset

[0054]

[0055] Table 2. Cell nucleus instance segmentation results of different methods on the MoNuSeg dataset

[0056]

[0057] As can be seen from Table 1 and Table 2, all indicators of the present invention are superior to those of the prior art methods, which proves that the present invention can significantly improve the performance of cell nucleus instance segmentation.

[0058] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0059] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0060] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0061] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0062] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A cross-graph context-aware cell nucleus instance segmentation method, characterized in that: include: Acquire a cell image and scan and segment it into a slice group including a plurality of slices, and construct a cell nucleus instance segmentation model, wherein the cell nucleus instance segmentation model includes a feature extraction module, a context injection module, a hint encoding module, a mask decoder, and a texture encoder; The feature extraction module extracts the prior features and domain features of the slice through a multi-layer encoder, refines the prior features of each layer through a multi-level feature refinement block, extracts and fuses the missing features from the domain features of each layer, and obtains enhanced features; The context injection module integrates context information from surrounding segmented slices according to the enhanced features to generate image features, the hint encoding module generates a point embedding of each cell nucleus according to the domain features, and the mask decoder obtains a mask of each cell nucleus according to the image features and the point embedding; The mask of each cell nucleus is combined to obtain the instance segmentation result of the slice, and the instance segmentation results of all slices are combined to obtain the instance segmentation result of the cell image; The instance segmentation result and the enhanced features of the slice are input into the texture encoder to generate texture features, and the texture features and the enhanced features are stored in the context injection module to provide context information for the next slice.

2. The cross-graph context-aware cell nucleus instance segmentation method according to claim 1, characterized in that: The feature extraction module includes a CNN encoder and a ViT encoder. The CNN encoder is a multi-layer convolutional neural network encoder. The ViT encoder includes a Patch Embedded module and a multi-layer feature refinement layer. The number of layers of the feature refinement layer is the same as the number of layers of the convolutional neural network encoder. Each layer of the feature refinement layer includes multiple feature refinement blocks, and each feature refinement block includes the multi-level feature refinement block and the ViT block.

3. The cross-graph context-aware cell nucleus instance segmentation method according to claim 2, characterized in that: The feature extraction module extracts the prior features and domain features of the slice through a multi-layer encoder, refines the prior features of each layer through a multi-level feature refinement block, extracts and fuses the missing features from the domain features of each layer, and obtains enhanced features, specifically: When each slice is input into the feature extraction module, the i-th layer of the CNN encoder outputs the i-th layer of domain features; The ViT encoder first extracts the first layer of prior features through the Patch Embedded module, fuses the i-th layer of prior features and the i-th layer of domain features through the i-th layer of feature refinement layer to obtain the i+1-th layer of prior features, and uses the output of the last layer of feature refinement layer as the enhanced feature.

4. The cross-graph context-aware cell nucleus instance segmentation method according to claim 3, characterized in that: The calculation method of the i+1th layer prior features is: , in, is the i+1th layer prior feature, ViT( ) is the operation of the ViT block, is the i-th layer prior feature, MFRB( ) is the operation of the multi-level feature refinement block, [ ] n Indicates that n ViT(MFRB()) operations are performed, where n is the number of feature refinement blocks in the feature refinement layer; In the i-th feature refinement layer, the feature of the j-th feature refinement block is recorded as , j=1,2…,n, That is , the features obtained after the multi-level feature refinement block in the jth feature refinement block are recorded as , the operations of the multi-level feature refinement block are as follows: , in, is the cross attention operation, LN( ) is the layer normalization operation, is the domain feature of the i-th layer, After refining domain feature representation , The result obtained after the operation of the ViT block is .

5. The cross-graph context-aware cell nucleus instance segmentation method according to claim 4, characterized in that: described The calculation method is: , Among them, Adapter( ) is the Adapter fine-tuning operation.

6. The cross-graph context-aware cell nucleus instance segmentation method according to claim 1, characterized in that: The context injection module generates image features by integrating context information from surrounding segmented slices based on the enhanced features, specifically: The context injection module includes an environment memory library and a texture memory library. For the i-th slice of the current input, the environment memory library stores the enhanced features corresponding to the previous i-1 slices, which are recorded as , j=1,2,…,i-1; the enhanced feature corresponding to the i-th slice is ,when When entering the context injection module, the closest The M1 enhanced features construct the shortest distance memory library, recorded as ;application Injecting surrounding environment information Get the features with environmental information, recorded as ,Will Input texture memory; The texture memory bank stores the texture features of Mt slices of completed instance segmentation with the largest similarity difference, and selects the texture memory bank and The closest M2 texture features build the maximum similarity memory bank, recorded as ;application Injecting foreground information Get the image features, record as .

7. The cross-graph context-aware cell nucleus instance segmentation method according to claim 6, characterized in that: The shortest distance memory bank is: , For the closest The kth feature among the M1 enhanced features; described The calculation method is: ,in, attention operations for memory; The maximum similarity memory bank is: , For the closest The qth feature among the M2 texture features; described The calculation method is: .

8. The cross-graph context-aware cell nucleus instance segmentation method according to claim 6, characterized in that: When the texture memory stores the texture features of Mt slices with the maximum similarity difference that have completed instance segmentation, the storage method is: The texture feature corresponding to the current i-th slice is recorded as ,Will The confidence level is recorded as ; The texture features of the Mt slices of completed instance segmentation stored in the texture memory are recorded as ,j=1,2,…,Mt;calculate The sum of the similarities between any texture feature and all other texture features, select the texture feature corresponding to the maximum sum, and record it as ,Will The confidence level is recorded as ; like Greater than Then delete it from the texture memory and store , otherwise the texture memory remains unchanged.

9. The cross-graph context-aware cell nucleus instance segmentation method according to claim 1, characterized in that: The hint encoding module generates a point embedding for each cell nucleus based on the domain features, specifically: The hint encoding module includes a prediction head and a hint encoder, and the prediction head includes a regression head and a classification head; The prediction head predicts the points of each cell nucleus based on the domain features, and the hint encoder encodes the points into point embeddings.

10. A cross-graph context-aware cell nucleus instance segmentation system, characterized in that: include: An image acquisition module, which acquires cell images and scans and segments the images into a slice group including a plurality of slices; A model construction module constructs a cell nucleus instance segmentation model, which includes a feature extraction module, a context injection module, a hint encoding module, a mask decoder, and a texture encoder. The feature extraction module extracts slice prior features and domain features through a multi-layer encoder, refines the prior features of each layer through a multi-level feature refinement block, extracts and fuses missing features from the domain features of each layer, and obtains enhanced features. The context injection module integrates context information from surrounding segmented slices according to the enhanced features to generate image features, the hint encoding module generates a point embedding of each cell nucleus according to the domain features, and the mask decoder obtains a mask of each cell nucleus according to the image features and the point embedding; The instance segmentation module combines the masks of each cell nucleus to obtain the instance segmentation results of the slice, and combines the instance segmentation results of all slices to obtain the instance segmentation results of the cell image; The instance segmentation result and the enhanced features of the slice are input into the texture encoder to generate texture features, and the texture features and the enhanced features are stored in the context injection module to provide context information for the next slice.

Citation Information

Patent Citations

  • Glioma image segmentation method and system based on multi-feature fusion

    CN117710670A

  • Tumor tissue cell image segmentation method, device, equipment, medium and product

    CN118447041A

  • Cell nucleus segmentation model construction method, cell nucleus segmentation method and construction device

    CN119478938A

  • Systems and methods for particle classification using machine learning

    US20250182506A1

  • Systems and methods for image-based cell segmentation, cell division detection, and cell tracking

    WO2025029788A1