Organoid Image Segmentation Method and System

Through visual image neural network and multidimensional interframe information fusion technology, the problems of interframe continuity and high computing resources in organoid OCT image segmentation are solved, and efficient and accurate three-dimensional segmentation effect is achieved, adapting to organoid segmentation tasks under different imaging conditions.

CN119887809BActive Publication Date: 2025-07-18HANGZHOU REGENOVO BIOTECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510377833.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The prior art has problems such as discontinuity between frames, boundary fracture, serious noise interference and high computing resource requirements in organoid OCT image segmentation, especially in organoids with complex morphology and multi-scale dispersion.

Method used

The visual graph neural network and multidimensional inter-frame information fusion technology are used to aggregate the three-dimensional neighborhood information of neighboring frames through a three-dimensional convolution kernel, and graph node connections with consistent physical structures are dynamically generated. Combined with a lightweight two-dimensional segmentation network, Dice Loss and cross-entropy loss function are used to optimize the segmentation effect.

Benefits of technology

It improves the intra-frame integrity and inter-frame consistency of organoid segmentation, reduces the computing resource requirements, enhances the model's robustness to noise and morphological heterogeneity, and realizes efficient and accurate three-dimensional OCT image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119887809B_ABST
    Figure CN119887809B_ABST
Patent Text Reader

Abstract

The present invention discloses an organoid image segmentation method and system, which relates to the field of medical image processing. The method includes: generating graph nodes and neighboring graph nodes based on the high-dimensional feature maps of the current frame and neighboring frames; dynamically generating the connection relationships between the graph nodes and the neighboring graph nodes by using the spatial position constraint principle of the original three-dimensional image to construct a graph structure with physical structure consistency; passing the multi-dimensional information of the neighboring graph nodes to the target graph node through graph convolution operations to compensate for the feature loss of the current frame, and aggregating the three-dimensional neighborhood information of the neighboring frames through a three-dimensional convolution kernel to achieve cross-dimensional inter-frame feature complementary fusion. According to the technical solution of the present application, it is possible to achieve a more accurate and robust segmentation effect of the two-dimensional segmentation network in complex organoid OCT scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image processing, and relates to an image segmentation method based on optical coherence tomography, and particularly to an organoid image segmentation method and system. Background Art

[0002] Patient-derived cancer organoids are a three-dimensional tissue culture system that can faithfully reproduce the molecular, histopathological, and phenotypic characteristics of patient tumors, and are widely used in tumor pathology modeling, drug screening, and precision medicine. However, organoids self-organize in a dispersed manner during culturing, showing multi-scale complex tissue morphology, which is closely related to tumor heterogeneity.

[0003] As a non-invasive three-dimensional imaging technology, optical coherence tomography is widely used in three-dimensional morphological analysis and dynamic monitoring of patient-derived cancer organoids due to its high resolution, deep penetration ability, and large field of view. However, organoids show multi-scale dispersed self-organization characteristics during culturing, and their complex morphology is highly related to tumor heterogeneity, which poses a severe challenge to the segmentation accuracy of OCT images. Existing methods have the following significant deficiencies in dealing with such tasks: Although the two-dimensional segmentation network based on single-frame images has the advantages of high computational efficiency and flexible design, it completely ignores the context information of adjacent frames in the three-dimensional volume data, resulting in discontinuity and boundary breakage in the segmentation results across frames. Especially in areas with low contrast between organoids and culture media and severe speckle noise interference, the limitations of single-frame information further exacerbate the segmentation error; The inherent speckle noise, stripe artifacts in OCT imaging, and the significant diversity of organoid morphology pose strict requirements on segmentation algorithms. Traditional methods rely on manually designed features or local context modeling and are difficult to dynamically adapt to changes in different imaging conditions and organoid states, resulting in large fluctuations in segmentation accuracy and a high background mis-segmentation rate. Therefore, the present invention aims to break through the above technical bottlenecks and achieve efficient, accurate, and scalable three-dimensional organoid segmentation by introducing a visual graph neural network and a multi-dimensional inter-frame information fusion strategy. Summary of the Invention

[0004] Technical Objectives

[0005] To solve the above problems, the purpose of the present invention is to provide an organoid image segmentation method and system, aiming to solve the performance bottleneck in the segmentation of three-dimensional organoid OCT images caused by the neglect of frame continuity in two-dimensional segmentation methods, the excessive computational resource requirements of three-dimensional segmentation methods, and the scarcity of labeled data through visual graph neural networks and multi-dimensional inter-frame information fusion technology. Specifically, it includes: improving the intra-frame integrity and inter-frame consistency of organoid segmentation; optimizing the cross-frame feature fusion ability of two-dimensional networks without the condition of continuous frame annotation; reducing video memory occupancy and computational complexity to achieve efficient processing of large-scale three-dimensional OCT data; enhancing the robustness of the model to speckle noise, morphological heterogeneity, and image quality degradation under drug intervention.

[0006] Technical solution

[0007] To achieve the above purpose, the present invention provides an organoid image segmentation method and system. The method and system extract multi-level features of the current frame and adjacent frames, and dynamically construct a graph structure through three-dimensional convolution to aggregate the feature information of multi-dimensional adjacent nodes, replacing the traditional k-nearest neighbor algorithm to reduce computational overhead, and restoring high-resolution segmentation results by fusing encoder features through upsampling and skip connections. This network is cascaded with a two-dimensional segmentation network, combined with a composite loss function and data augmentation strategy, and can achieve high-precision segmentation across patients and drug scenarios under the condition of only sparse annotation.

[0008] In the first aspect, the present invention provides an organoid image segmentation method, including:

[0009] Generating graph nodes and adjacent graph nodes based on the high-dimensional feature maps of the current frame and adjacent frames, where the current frame is encoded as the target graph node, and the context information of its three-dimensional adjacent region is encoded as the adjacent graph node;

[0010] Dynamically generating the connection relationship between graph nodes and adjacent graph nodes by using the spatial position constraint principle of the original three-dimensional image, and constructing a graph structure with physical structure consistency;

[0011] Transferring the multi-dimensional information of adjacent graph nodes to the target graph node through graph convolution operations, compensating for the feature loss of the current frame, and aggregating the three-dimensional neighborhood information of adjacent frames through three-dimensional convolution kernels to achieve cross-dimensional inter-frame feature complementary fusion.

[0012] Furthermore, the multi-dimensional inter-frame information fusion network is sequentially connected with the two-dimensional segmentation network as a preprocessing module, and the enhanced feature map output by the fusion network is input into the two-dimensional segmentation network to improve the intra-frame integrity and inter-frame continuity of organoid segmentation and significantly reduce computational resource requirements.

[0013] Furthermore, multi-label annotation is performed on the three-dimensional OCT images to distinguish organoids, Matrigel, and background regions, and the three-dimensional continuity of the annotation is verified frame by frame. Data augmentation is performed on the training set through random affine transformation, elastic deformation, and combined intensity transformation to enhance the robustness of the model to imaging noise and morphological heterogeneity.

[0014] Furthermore, the three-dimensional OCT image is represented as:

[0015]

[0016] To enhance the segmentation performance of the current frame , the d slice stacks of its neighboring frames are selected as:

[0017]

[0018] Furthermore, the process of the multi-label annotation includes:

[0019] Import the three-dimensional OCT data into the annotation tool frame by frame and draw multi-label masks frame by frame;

[0020] Ensure the continuity and consistency of the annotation during the annotation process, especially maintaining the coherence of the organoid and Matrigel boundaries when crossing frames;

[0021] In areas where the organoids are large or morphologically complex, a finer annotation granularity is used for annotation;

[0022] In the Matrigel region, the annotation covers the direct support area around the organoids and avoids extending into the background area.

[0023] Furthermore, multi-level feature maps of the current frame and neighboring frames are extracted through a lightweight two-dimensional convolutional network and stacked into a three-dimensional tensor. Neighborhood information aggregation is performed on the three-dimensional tensor using zero-padding three-dimensional convolution operations to generate neighboring graph nodes consistent with the physical space distribution.

[0024] Furthermore, the multi-level feature maps are extracted through the following formula:

[0025]

[0026] where is the k-th level feature map of ; is a two-dimensional convolution; contains and with a convolution kernel size of 1; has a convolution kernel size of 3; is the ReLU activation function; It is a max pooling downsampling operator;

[0027] Furthermore, the multi-frame feature maps are processed by the following formula and stacked into a three-dimensional tensor:

[0028]

[0029] Further, based on the three-dimensional spatial position constraint, neighboring nodes that are closely related to the target graph nodes are automatically screened, redundant connections are removed, the feature similarity between nodes is characterized by an adjacency matrix, and the connection weights are dynamically adjusted to avoid the computational complexity brought by global search.

[0030] Further, in the graph convolution operation, only the connections that are closely related to the current node in the physical space are retained for the connection relationship of neighboring nodes, and the adjacency matrix is determined by the feature similarity of the neighboring nodes.

[0031] Further, the three-dimensional neighborhood information of neighboring frames is aggregated through three-dimensional convolution operation, and the complementary fusion of multi-dimensional feature information is realized based on the dynamically generated graph node connection relationship, where the dynamic connection relationship is automatically generated by the three-dimensional spatial position constraint.

[0032] Further, the three-dimensional neighborhood information of neighboring frames is aggregated by the following formula:

[0033]

[0034] In the formula, , that is , is a three-dimensional convolution with zero padding; is the neighboring graph node of.

[0035] Further, a multi-layer perceptron is used to fuse the features of the target graph node and its neighboring nodes, and the update formula is:

[0036]

[0037] In the formula, is the feature of neighboring nodes aggregated by three-dimensional convolution; is the multi-layer perceptron.

[0038] This method significantly reduces the computational complexity of graph construction, and at the same time improves the efficiency of feature fusion. Through spatial position constraints, the model can automatically generate node connection relationships consistent with the physical structure, avoiding the waste of computing resources caused by global search in traditional methods. In addition, it also enhances the generalization ability of the model for cross-patient and cross-drug scenarios, enabling it to adapt to the organoid segmentation task under different imaging conditions.

[0039] Furthermore, through progressive upsampling and skip connection operations, the fused feature map is concatenated with the multi-level features of the encoder at the channel level, and the high-resolution segmentation result is restored based on the U-Net architecture to enhance the integrity of the organoid boundary and the continuity of the internal structure.

[0040] Furthermore, the graph nodes with updated features are merged into the decoder to improve the clarity and integrity of the organoid features in the feature map. The decoding formula is as follows:

[0041]

[0042] In the formula, is the upsampling operation; is a two-dimensional convolution operator; contains GroupNorm and ReLU. The convolution kernel size of is (1, 1); The convolution kernel size of

[0043] Furthermore, the size of the image that finally outputs the current frame fused with the neighboring frame information is the same as the input B-scan.

[0044] Furthermore, the cascaded network is applicable to the OCT image segmentation task of patient-derived organoids, can dynamically adapt to the decreased image contrast and increased noise caused by drug intervention, and balances the segmentation weights of organoids and the background through a composite loss function.

[0045] Furthermore, the model training uses a composite loss function combining Dice Loss and cross-entropy loss, conducts end-to-end training through the Adam optimizer, quantifies the segmentation accuracy based on the Dice coefficient, background mis-segmentation rate, and inter-frame continuity index, and optimizes the three-dimensional convolution kernel size and the number of neighboring frames.

[0046] Furthermore, the initial learning rate of the model training is 0.00005, and a random seed is set to ensure the reproducibility of the experiment.

[0047] Furthermore, the dataset is divided according to the ratio of 20% training set and 80% test set to ensure that the model learns the core features under limited training samples. The test set is used to evaluate the model performance, and focuses on detecting its generalization ability on large-scale data.

[0048] Furthermore, no data augmentation is performed on the test set, and only basic operations such as normalization are applied.

[0049] Further, the Dice coefficient is used to evaluate the segmentation accuracy of the model for the organoid region, the background mis-segmentation coefficient is used to evaluate the mis-segmentation of the background region by the model, and the inter-frame continuity index is used to ensure the continuity of cross-frame segmentation.

[0050] While ensuring segmentation accuracy, this architecture significantly reduces the computational resource requirements, enabling it to efficiently process large-scale three-dimensional OCT data; through the composite loss function, the model can better balance the segmentation weights of small and large targets, improving the overall quality of the segmentation results. Its lightweight design makes it possible to apply the model in clinical real-time scenarios, providing technical support for the efficient analysis and precise treatment of organoids.

[0051] In a second aspect, the present invention also provides an organoid image segmentation system. The system is based on the method described in the first aspect above and includes:

[0052] An encoder for extracting multi-level feature maps of the current frame and adjacent frames and converting them into three-dimensional tensors. Its input is the current frame and adjacent frames, and the output is the high-dimensional feature map of each frame;

[0053] A multi-dimensional inter-frame information fusion module for aggregating neighborhood information of the three-dimensional tensors to generate adjacent graph nodes, fusing adjacent node information in a graph convolution manner, transferring the features of adjacent frames to the current frame nodes, and automatically generating node connection relationships based on three-dimensional spatial position constraints, thereby avoiding the high computational complexity brought by applying the k-nearest neighbor algorithm. The feature map of each frame is represented as a set of graph nodes , where represents the number of channels;

[0054] A decoder module for outputting an enhanced current frame feature map through upsampling and feature fusion, and inputting it into a two-dimensional segmentation network to complete organoid segmentation.

[0055] In a third aspect, the present invention also provides a computer device, including a management platform and a memory. The management platform is connected to the memory. The memory is used to store a computer program, and the management platform is used to execute the computer program stored in the memory so that the computer device executes at least one step of the organoid image segmentation method described above.

[0056] In a fourth aspect, the present invention also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by the management platform, it realizes at least one step of the organoid image segmentation method described above.

[0057] The present invention utilizes a three-dimensional convolutional kernel to aggregate the three-dimensional neighborhood information of adjacent frames, dynamically generates node connection relationships based on physical space positions, and preserves boundary information through zero-padding three-dimensional convolution, avoiding the complexity of manually defining connection rules; cascades MIFFNet with a two-dimensional segmentation network, balances performance and resource consumption through lightweight design, and combines Dice Loss and cross-entropy loss to dynamically adjust the weight allocation between the organoids and the background. The method and system have a wide range of application scenarios and can enable the two-dimensional segmentation network to have more accurate and robust segmentation effects in complex organoid OCT scenarios.

[0058] Advantageous Effects

[0059] By implementing the organoid image segmentation method and system provided by the present invention, the following technical effects are achieved:

[0060] (1) Utilizes a three-dimensional convolutional kernel to aggregate the three-dimensional neighborhood information of adjacent frames, dynamically generates node connection relationships based on physical space positions, and preserves boundary information through zero-padding three-dimensional convolution, avoiding the complexity of manually defining connection rules. This method significantly reduces the computational complexity of graph construction, while improving the efficiency of feature fusion. Through spatial position constraints, the model can automatically generate node connection relationships consistent with the physical structure, avoiding the waste of computational resources caused by global search in traditional methods. In addition, it enhances the generalization ability of the model for cross-patient and cross-drug scenarios, enabling it to adapt to organoid segmentation tasks under different imaging conditions.

[0061] (2) Cascades MIFFNet with a two-dimensional segmentation network, balances performance and resource consumption through lightweight design, and combines Dice Loss and cross-entropy loss to dynamically adjust the weight allocation between the organoids and the background. This architecture significantly reduces the computational resource requirements while ensuring segmentation accuracy, enabling it to efficiently process large-scale three-dimensional OCT data; through the composite loss function, the model can better balance the segmentation weights of small and large targets, improving the overall quality of the segmentation results. Its lightweight design makes it possible to apply the model in clinical real-time scenarios, providing technical support for the efficient analysis and precise treatment of organoids. Description of the Drawings

[0062] To make the above organoid image segmentation method and system of the present invention more clearly understandable, the following will briefly introduce the drawings required for the specific implementation of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0063] Figure 1 Represents the flowchart of the organoid image segmentation method;

[0064] Figure 2 It shows a schematic diagram of the overall architecture of the multi-dimensional inter-frame information fusion module;

[0065] Figure 3 It shows a schematic diagram of the overall architecture of the cascading of the multi-dimensional inter-frame information fusion module and the two-dimensional network. Specific implementation manners

[0066] To facilitate the understanding of the embodiments of the present invention, first, the abbreviations and key terms that may be involved in the embodiments of the present invention are explained or defined. For the undefined abbreviations or key terms, they are all commonly understood by those skilled in the art.

[0067] OCT: Optical Coherence Tomography;

[0068] Regenovo OCT: Optical Coherence Tomography Imaging Device;

[0069] ITK-SNAP: Medical Image Segmentation Software;

[0070] torchio: A Python library for efficient processing of 3D medical images in deep learning;

[0071] RandomAffine: A class in the torchio module for performing random affine transformations on images;

[0072] RandomElasticDeformation: A class in the torchio module for random elastic deformation transformation in medical image enhancement;

[0073] B-Scan: A two-dimensional cross-sectional view of optical coherence tomography, which shows the depth information (longitudinal) and transverse information (horizontal) of the organoid;

[0074] ReLU: Activation function, Rectified Linear Unit function;

[0075] Dice Loss: A loss function applied to image segmentation tasks;

[0076] Cross-Entropy Loss: Cross-entropy loss;

[0077] Dice coefficient: A statistical metric for measuring the similarity between two sample sets;

[0078] MIFFNet: A classification network for feature fusion;

[0079] TransUnet: A deep learning model that combines Transformer and U-Net structures;

[0080] GDNet: A deep learning segmentation model for OCT retinal images

[0081] nnUNet 3D: An Adaptive Medical Image Segmentation Framework

[0082] FRNet: A Deep Learning Model for OCT Retinal Vascular Images

[0083] Infer-time: Inference Time

[0084] Memory: Memory

[0085] Example 1:

[0086] An organoid image segmentation method and system are provided as described below.

[0087] The organoid image segmentation method is as follows Figure 1 shown and includes:

[0088] Step 1: Downsample the input image and extract high-dimensional feature information.

[0089] Step 2: Generate graph nodes and adjacent graph nodes from the high-dimensional feature maps of the current frame and adjacent frames, use the spatial position constraints of the original three-dimensional image to generate the connection relationships between graph nodes and adjacent nodes, construct a graph structure, and transfer the information in the adjacent graph nodes to the target graph node based on the convolution update of the graph structure to compensate for the feature information of the current frame.

[0090] Step 3: Restore the high-dimensional feature information and the current frame feature map after fusion update to an output image, and input the output image into a two-dimensional segmentation network to segment the organoids.

[0091] The input image is a patient-derived organoid OCT image collected by a Regenovo OCT device and is manually multi-label annotated, specifically including:

[0092] 1. Annotation tool:

[0093] a) Use the ITK-SNAP open-source software to perform multi-label annotation on the OCT image, supporting fine pixel-level classification.

[0094] 2. Annotation object:

[0095] a) Organoids: Manually delineate the organoids in the OCT image and label them as foreground (label value is 1);

[0096] b) Matrigel: Label the Matrigel area around the organoids in the OCT image as a sub-important area in the background (label value is 2);

[0097] c) Background: Label the area other than the organoids and Matrigel as the general background (label value is 0).

[0098] 3. Annotation process:

[0099] a) Import the 3D OCT data frame by frame into the annotation tool and draw multi-label masks frame by frame;

[0100] b) Ensure the continuity and consistency of the annotation, especially the coherence of the boundaries of the organoids and the Matrigel across frames;

[0101] c) In areas where the organoids are large or have complex morphologies, adopt a finer annotation granularity;

[0102] d) In the Matrigel area, the annotation should cover the direct support area around the organoids and avoid extending into the background area.

[0103] 4. Annotation verification:

[0104] a) Have the annotation results reviewed by experts to ensure the accuracy of the boundaries and classification of the organoids and the Matrigel;

[0105] b) Use a 3D visualization inspection tool to verify the 3D continuity of the annotation.

[0106] 5. Annotation output:

[0107] a) Output a 3D mask file with multi-label annotations, where each pixel corresponds to a class label (0: background, 1: organoid, 2: Matrigel);

[0108] b) Each file contains the original OCT image and the corresponding multi-label annotation.

[0109] Divide the annotated dataset into a training set and a test set at a ratio of 20% and 80% respectively, ensuring that the model learns the core features with limited training samples. The test set is used to evaluate the model performance, focusing on detecting its generalization ability on large-scale data, specifically including:

[0110] 1. Data sorting standards:

[0111] a) The 3D organoid image data collected from the Regenovo OCT device, including OCT image samples from multiple patients (such as Patient 1, Patient 2, etc.);

[0112] b) Keep the dataset complete without excluding any images to ensure the richness of the data volume and the generalization ability of the model;

[0113] c) Each data sample includes:

[0114] i. A 3D OCT image (resolution 512 × 800 × 800, length × width × number of cross-sections);

[0115] ii. The corresponding multi-label annotation file (including organoids, Matrigel, and background regions).

[0116] 2. Basis for data division:

[0117] a) The data of Patient 1 and Patient 2 are divided in a ratio of 20%:80%:

[0118] i. Training set: 20% of the data of Patient 1 and Patient 2;

[0119] ii. Test set: 80% of the data of Patient 1 and Patient 2.

[0120] b) After division, the training set is mainly used for model training, and the test set is used for model performance evaluation, focusing on detecting its generalization ability on large-scale data.

[0121] Before training the OCT images, data preprocessing is performed on the training set. The data preprocessing includes at least one of normalization processing, spatial transformation, random flipping, and intensity transformation, specifically including:

[0122] 1. Spatial transformation: Use the torchio data augmentation toolkit to apply spatial transformation operations to the OCT images, including:

[0123] a) Random affine transformation: Apply rotation and scaling transformations to the image with a probability of 0.5, and the parameter range is:

[0124] i. The rotation angle range is [-10°, 10°];

[0125] ii. The scaling factor range is [0.9, 1.1].

[0126] b) Random elastic deformation: Apply elastic deformation to the image through control point interpolation, and the control parameters are:

[0127] i. Lock the boundary to limit the pixels in the edge region of the image from being affected by elastic deformation;

[0128] ii. The number of control points is 4 to generate a moderate elastic distortion.

[0129] c) Comprehensive spatial transformation: Randomly select one from RandomAffine and RandomElasticDeformation with a probability of 70% and apply it to the image.

[0130] 2. Random flipping:

[0131] a) Randomly flip the image along the horizontal, vertical, or depth direction with a flipping probability of 0.5;

[0132] b) This operation can enhance the robustness of the model to different morphologies of organoids.

[0133] 3. Intensity transformation:

[0134] a) Apply a random transformation of the intensity level to simulate variations in imaging and enhance the model's adaptability to different image qualities;

[0135] b) Random gamma transformation: Non-linearly adjust the intensity of the image pixel values;

[0136] c) Random blur: Apply Gaussian blur to the image to simulate blur artifacts in OCT imaging;

[0137] d) Random bias field: Add a bias field to the image to simulate the non-uniformity that may be introduced by the imaging device;

[0138] e) Combined intensity transformation: Randomly select one of the above three intensity transformations with a probability of 70% and apply it to the image.

[0139] 4. Applied to image data:

[0140] a) Apply the above preprocessing pipeline to each sample of the divided training set data;

[0141] b) No data augmentation is performed on the test set, and only basic operations such as normalization are applied.

[0142] The input image is specifically a three-dimensional OCT image containing multiple B-Scans , expressed as:

[0143]

[0144] To enhance the segmentation performance of the current frame , select the d slice stacks of its neighboring frames as:

[0145]

[0146] The spatial transformation is to perform random affine transformation, random elastic deformation, and comprehensive spatial transformation operations on the OCT image through the torchio data augmentation toolkit.

[0147] The random flip is to randomly flip the image along the horizontal, vertical, or depth direction.

[0148] The intensity transformation is to apply a random transformation of the intensity level to simulate variations in imaging, including random gamma transformation, random blur, random bias field, and combined intensity transformation.

[0149] No data augmentation is performed on the test set, and only basic operations such as normalization are applied.

[0150] In step 1, features are extracted through a lightweight two-dimensional convolutional network, and multi-scale three-dimensional context information is captured based on the slice stack. The multi-level feature maps of each frame are extracted by the following formula:

[0151]

[0152] wherein, is the k-th level feature map of ; is a two-dimensional convolution; contains and with a convolution kernel size of 1; has a convolution kernel size of 3; is the ReLU activation function; is the max pooling downsampling operator;

[0153] The multi-frame feature maps are processed and stacked into a three-dimensional tensor by the following formula:

[0154]

[0155] In step 2, the generation of graph nodes and neighboring graph nodes from the high-dimensional feature maps of the current frame and neighboring frames is specifically as follows: Given the current frame as input, the current frame is encoded as a graph node, and the information of its three-dimensional adjacent region is encoded as neighboring graph nodes.

[0156] The three-dimensional neighborhood information of neighboring frames is aggregated through three-dimensional convolution operations, and complementary fusion of multi-dimensional feature information is achieved based on the dynamically generated graph node connection relationships, where the dynamic connection relationships are automatically generated through three-dimensional spatial position constraints.

[0157] The three-dimensional neighborhood information of neighboring frames is aggregated by the following formula:

[0158]

[0159] wherein, , that is , is a three-dimensional convolution with zero padding; is the neighboring graph node of

[0160] In the graph convolution operation, the connection relationships of neighboring nodes only retain the connections that are closely related to the current node in physical space. The connection relationships of neighboring nodes can be characterized by an adjacency matrix, and the adjacency matrix is determined by the feature similarity of the neighboring nodes.

[0161] For each graph node , fuse adjacent nodes , fuse multi-dimensional information of adjacent nodes, and perform feature update through graph convolution operation:

[0162]

[0163] In the formula, is a multi-layer perceptron for updating the feature representation of graph nodes.

[0164] Through upsampling and skip connection operations, perform channel concatenation based on feature maps at different levels and the fused feature map to enhance the boundary integrity and internal structure continuity of the segmentation result.

[0165] Gradually restore the low-resolution feature map to the same resolution as the input image through step-by-step upsampling, and merge the graph nodes with updated features into the decoder to improve the clarity and integrity of the organoid features in the feature map. The decoding formula is as follows:

[0166]

[0167] In the formula, is the upsampling operation; is a two-dimensional convolution operator; contains GroupNorm and ReLU. The convolution kernel size of is (1, 1); The convolution kernel size of

[0168] The size of the image of the current frame that finally fuses the information of adjacent frames is the same as the input B-scan.

[0169] The network is trained end-to-end through a composite loss function of Dice Loss and Cross-Entropy Loss, and the weight balance between the organoid and the matrix gel is achieved based on the Adam optimizer and a batch size of 4. Among them, the initial learning rate is 0.00005, and a random seed is set to ensure the repeatability of the experiment.

[0170] The evaluation metrics of the model include the Dice similarity coefficient, the background mis-segmentation coefficient, the inter-frame continuity index, etc.; the Dice coefficient is used to evaluate the segmentation accuracy of the model for the organoid region, the background mis-segmentation coefficient is used to evaluate the mis-segmentation of the background region by the model, and the inter-frame continuity index is used to ensure the continuity of cross-frame segmentation.

[0171] An organoid image segmentation system, the system is based on the foregoing method, and includes:

[0172] An encoder, which is used to extract multi-level feature maps of the current frame and adjacent frames and convert them into three-dimensional tensors, and its input is the current frame and adjacent frames, and the output is the high-dimensional feature map of each frame;

[0173] A multi-dimensional inter-frame information fusion module, whose overall architecture is as Figure 2 shown, and the overall architecture cascaded with the two-dimensional network is as Figure 3 shown, which is used to aggregate neighborhood information of the three-dimensional tensors to generate adjacent graph nodes, fuse adjacent node information in the way of graph convolution, transfer the features of adjacent frames to the current frame nodes, and automatically generate node connection relationships based on three-dimensional spatial position constraints, so as to avoid the high computational complexity brought by applying the k-nearest neighbor algorithm. The feature map of each frame is represented as a set of graph nodes , where represents the number of channels;

[0174] A decoder, which is used to output an enhanced current frame feature map through upsampling and feature fusion, and input it into a two-dimensional segmentation network to complete organoid segmentation.

[0175] Example 2:

[0176] On the basis of the foregoing embodiment, the MIFFNet module is cascaded sequentially with two-dimensional segmentation networks such as TransUnet and GDNet.

[0177] The connection method of the sequential cascade includes:

[0178] Connect the MIFFNet as a preprocessing module sequentially with the two-dimensional segmentation network;

[0179] The MIFFNet receives the OCT images of the current frame and its adjacent frames, and generates an enhanced current frame feature map through multi-dimensional inter-frame information fusion;

[0180] The enhanced current frame feature map is input into a two-dimensional segmentation network such as TransUnet or GDNet for subsequent segmentation tasks.

[0181] While keeping the network structure simple, this method significantly improves the accuracy and consistency of the segmentation results. The MIFFNet module design follows the lightweight principle, replaces the graph construction of the computationally complex k-nearest neighbor algorithm with a single 3D convolution; controls the number of adjacent frames and the size of the 3D convolution kernel, so that the system achieves a balance between performance and computational overhead; improves the accuracy of organoid boundary and detail segmentation by aggregating adjacent frame information, and reduces the inconsistency between adjacent frames by using multi-dimensional inter-frame fusion. In addition, it can also enhance the robustness of segmentation in the case of image contrast degradation and noise enhancement caused by drug action.

[0182] Example 3:

[0183] Based on the foregoing embodiments, the performance of the visual graph neural network is verified.

[0184] Pre-read the labeled three-dimensional OCT image dataset of organoid groups from the storage path. The dataset contains organoid OCT images of multiple patients, multiple cancer types, and multiple drugs. Each sample consists of multiple B-scans and has a size of .

[0185] Suppose the patient numbers are P1 and P2, the cancer types include pancreatic cancer and cholangiocarcinoma, and the experimental conditions are the states of organoids under normal culture and after drug action. By normalizing each frame of B-scan image, the image contrast is enhanced, and data augmentation operations are performed on the images, while ensuring that the enhanced images retain the boundary features of the organoids.

[0186] Table 1. Comparison of experimental results of OCT images of organoid groups in the case of multiple patients, multiple cancer types, and multiple drugs

[0187]

[0188] As shown in Table 1, the OCT dataset of pancreatic cancer organoids of patient 1 was cultured under normal conditions, and the OCT dataset of intrahepatic cholangiocarcinoma of patient 2 included two drug conditions. In the OCT data of pancreatic cancer organoids of patient 1 without drug action, the two-dimensional segmentation network outperformed the pure three-dimensional network nnUNet 3D trained on the sparse annotation dataset in terms of segmentation accuracy, and the background mis-segmentation probability of all two-dimensional networks was lower than 10%. When TransUnet was cascaded with MIFFNet, the segmentation accuracy of the x-z cross-section increased by 1%. When FRNet was cascaded with MIFFNet, the in-frame segmentation accuracy decreased by 2%, and the background mis-segmentation probability decreased by 4%. When GDNet was cascaded with MIFFNet, both the segmentation accuracy and the background mis-segmentation probability were improved. In addition, in the ICC-2 organoid OCT data mixed with normal culture and drug action, the two-dimensional baseline method was lower than the pure three-dimensional network nnUNet 3D trained on the sparse annotation dataset in terms of segmentation accuracy. When the two-dimensional baseline network was cascaded with MIFFNet, the segmentation accuracy was close to or even exceeded that of nnUNet 3D, and the background mis-segmentation probability decreased significantly. All three networks could be controlled at about 10%. It can be seen that cascading the two-dimensional baseline network with MIFFNet can improve the consistency and stability of organoid segmentation before and after drug addition.

[0189] The segmentation accuracy and stability of the two-dimensional baseline method cascaded with MIFFNet are close to or even exceed those of the three-dimensional network, while the Infer-time and Memory metrics show that its calculation time and video memory overhead are significantly reduced. Compared with the three-dimensional network, the video memory decreases by about 50%, and the calculation time is shortened by more than 90%.

[0190] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable non-transitory storage media containing computer-usable program code.

[0191] The present invention can provide computer program instructions to the management platform of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed through the management platform of the computer or other programmable data processing devices generate a device for implementing the system.

[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions of the system.

[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions of the system.

Claims

1. An organoid image segmentation method, characterized in that, include: A lightweight two-dimensional convolutional network is used to extract multi-level feature maps of the current frame and neighboring frames, and the maps are stacked into a three-dimensional tensor. The three-dimensional tensor is aggregated with neighborhood information using a zero-filled three-dimensional convolution operation to generate a target graph node and neighboring graph nodes that are consistent with the physical space distribution; the current frame is encoded as a target graph node, and the context information of its three-dimensional adjacent area is encoded as a neighboring graph node; The spatial position constraint principle of the original three-dimensional image is used to dynamically generate the connection relationship between the graph nodes and the adjacent graph nodes, and a graph structure with physical structure consistency is constructed; The multi-dimensional information of neighboring graph nodes is transferred to the target graph node through graph convolution operation to compensate for the feature loss of the current frame, and the three-dimensional neighborhood information of neighboring frames is aggregated through the three-dimensional convolution kernel to achieve cross-dimensional inter-frame feature complementary fusion; The fused feature map is input into the two-dimensional segmentation network to complete the organoid segmentation.

2. The method according to claim 1, characterized in that: The MIFFNet network is used as a preprocessing module and sequentially connected to the two-dimensional segmentation network. The MIFFNet network receives images of the current frame and its adjacent frames, generates an enhanced feature map of the current frame through multi-dimensional inter-frame information fusion, and inputs it into the two-dimensional segmentation network.

3. The method according to claim 1, characterized in that: Multi-label annotation is performed on 3D OCT images to distinguish between organs, matrix gel and background areas, and the 3D continuity of the annotations is verified frame by frame. The training set is enhanced through random affine transformation, elastic deformation and combined intensity transformation.

4. The method according to claim 1, characterized in that: Based on the three-dimensional spatial position constraints, the adjacent nodes closely related to the target graph nodes are automatically screened, redundant connections are eliminated, the feature similarity between nodes is represented by the adjacency matrix, and the connection weights are dynamically adjusted.

5. The method according to claim 4, characterized in that: A multi-layer perceptron is used to fuse the features of the target graph node and its neighboring nodes. The update formula is: In the formula, is the neighboring node feature aggregated by 3D convolution; is the multi-layer perceptron.

6. The method according to claim 1, characterized in that: Through step-by-step upsampling and skip connection operations, the fused feature map and the encoder multi-level features are channel-concatenated to restore the high-resolution segmentation result based on the U-Net architecture.

7. The method according to claim 1, characterized in that: The cascade network is suitable for the OCT image segmentation task of patient-derived organoids. It can dynamically adapt to the image contrast loss and noise enhancement caused by drug intervention, and balance the segmentation weights of organoids and background through a composite loss function. The cascade network is the result of sequential cascading of the MIFFNet network and the two-dimensional segmentation network.

8. The method according to claim 1, characterized in that: The model training adopts a composite loss function that combines Dice Loss and cross entropy loss, and performs end-to-end training through the Adam optimizer. The segmentation accuracy is quantified based on the Dice coefficient, background mis-segmentation rate, and inter-frame continuity indicators.

9. An organoid image segmentation system, characterized in that: The implementation of the system is based on the method according to any one of claims 1 to 8; The system comprises: An encoder, which is used to extract multi-level feature maps of the current frame and adjacent frames and convert them into three-dimensional tensors; A multi-dimensional inter-frame information fusion module, which is used to aggregate adjacent node information and update features through the MIFFNet network; A decoder, which is used to output an enhanced current frame feature map through upsampling and feature fusion and input it into a two-dimensional segmentation network to complete organoid segmentation.

Citation Information

Patent Citations

  • Skeleton behavior recognition method based on 3D space-time diagram convolution

    CN111814719A