Pig organ automatic segmentation method and segmentation system based on CT anatomical structure relation
Through the automatic segmentation method of pig organs based on the relationship between CT anatomical structure, using the encoder-decoder architecture of visual state space blocks and the spatial link GRU module, the problems of unclear organ boundaries, large size differences and large shape changes in pig CT images are solved, and efficient and accurate automatic segmentation of pig multi-organs is achieved, improving segmentation accuracy and efficiency.
Patent Information
- Application Number
- CN202510631991.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-29
AI Technical Summary
The existing multi-organ segmentation method of pig CT images faces problems such as unclear organ boundaries, large size differences, large shape changes, uneven segmentation accuracy and time-consuming and labor-consuming segmentation, resulting in low accuracy and robustness of the segmentation model.
The automatic segmentation method of pig organs based on CT anatomical structure relationship is adopted, and the organ mask is manually marked by 3D Slicer, combined with random data augmentation technology and the automatic segmentation model of pig multi-organs, and semi-supervised training is carried out through the encoder-decoder architecture of visual state space blocks, spatial link GRU module, global organ category coding module and global organ category guidance module, and semi-supervised training is carried out to optimize segmentation losses, auxiliary losses and consistency losses to realize automatic segmentation of organs.
The batch processing and accurate segmentation of pig CT images is realized, the segmentation consistency of different organs is improved, the segmentation accuracy and efficiency are significantly improved, and the current problem of multi-organ segmentation of pigs is solved, providing efficient and accurate solutions for medical image analysis and animal husbandry management.
Smart Images

Figure CN120563533A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision and deep learning technology, and specifically to an automatic pig organ segmentation method and segmentation system based on CT anatomical structure relationships. Background Art
[0002] Pigs are not only an important food source for humans, but their anatomical characteristics and physiological properties also share certain similarities with humans, making them widely used in medical research and drug trials. In recent years, with the increasing maturity of computed tomography (CT) technology, the use of CT images to analyze animal body composition has gradually become a key technology and is widely used in the livestock production industry. By analyzing whole-body CT images of pigs using multi-organ automatic segmentation methods, information about their internal organ structure can be obtained non-invasively. By qualitatively and quantitatively analyzing this information, we can obtain data such as organ shape and size, drug uptake, biomarker signals, or metastasis distribution in imaging data, providing key support for medical research and breeding research.
[0003] However, existing multi-organ segmentation methods face many challenges when processing pig CT images. First, the imaging contrast of pig internal tissues is low, resulting in unclear organ boundaries and increasing the difficulty of segmentation. Second, the size of pig internal organs varies significantly. For example, there is a large size difference between large organs such as the heart and liver and small organs such as the gallbladder and bladder, which can easily lead to uneven segmentation accuracy. In addition, the shapes of organs such as the gallbladder and bladder vary greatly, making it difficult for traditional segmentation methods to accurately capture their morphological features. In addition, manual segmentation of CT images is time-consuming and labor-intensive, which limits the number of segmentation samples, resulting in low accuracy and robustness of existing segmentation models.
[0004] Therefore, an automatic pig organ segmentation method and segmentation system based on CT anatomical structure relationship is proposed. Summary of the Invention
[0005] The purpose of the present invention is to provide a method and system for automatic pig organ segmentation based on CT anatomical structure relationships, so as to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: a method for automatically segmenting pig organs based on CT anatomical structure relationships, comprising the following steps;
[0007] Step 1: Obtain a full-body CT scan of a pig, and have professional veterinarians and radiologists manually annotate masks of multiple organs using 3D Slicer;
[0008] Step 2: Using random data enhancement technology on the whole-body CT scan image of the pig and its mask, obtain a strongly enhanced whole-body CT scan image of the pig and its mask;
[0009] Step 3: Use the pig multi-organ automatic segmentation model to train the pig whole-body CT scan image and its mask and the strongly enhanced pig whole-body CT scan image and its mask to obtain the pig multi-organ prediction mask and complete the automatic segmentation of the pig multiple organs;
[0010] The training process of the pig multi-organ automatic segmentation model is as follows:
[0011] S1. A whole-body CT scan image sample of a pig and a corresponding strongly enhanced whole-body CT scan image sample of a pig are simultaneously input into the model, and global and local features are extracted through an encoder-decoder architecture based on the visual state space;
[0012] S2, input the global features and local features extracted in S1 into the spatial link GRU module, dynamically model the changes between slices, and obtain the relative position information and detailed features of different organs;
[0013] S3. For the whole-body CT scan image samples of live pigs, the global organ category features are extracted through the global organ category encoding module, and a category-aware 3D segmentation map M is generated. The segmentation loss and auxiliary loss are calculated with the ground-truth mask GT.
[0014] S4. Sample of a strongly enhanced whole-body CT scan image of a pig, using the global organ category guidance module to generate a guided 3D segmentation map G s , the unguided 3D segmentation map M is generated by the global organ category encoding module s , calculate G s and M s consistency loss.
[0015] Preferably, the random data enhancement technology in step 2 includes random flipping, rotation, contrast adjustment and Gaussian noise addition.
[0016] Preferably, the segmentation loss, auxiliary loss and consistency loss in the training of the pig multi-organ automatic segmentation model are training loss functions, the segmentation loss is used to guide network segmentation learning, the auxiliary loss is used to optimize the high-level semantic information and edge detail information of the extracted features, and the consistency loss is used to constrain the consistency of the segmentation results of the strongly enhanced pig whole-body CT scan image under different conditions;
[0017] The calculation formula of the segmentation loss is:
[0018]
[0019] in: is the cross entropy loss, for dice losses;
[0020] The auxiliary loss is calculated as follows:
[0021]
[0022] in: is the Softmax operation, M λ is the segmentation prediction result of the original whole-body CT scan image of the live pig or the strongly enhanced whole-body CT scan image of the live pig, It is a bilinear interpolation operation used to upsample the segmentation prediction result to the same size as the mask GT;
[0023] The calculation formula of the consistency loss is:
[0024]
[0025] in, is the Softmax operation, G s is the segmentation result of the strongly enhanced pig whole-body CT scan image guided by the global organ category expression feature, M s is the segmentation result of the strongly enhanced whole-body swine CT scan image without the guidance of the global organ category expression feature, n is the batch size, and i is the slice index;
[0026] Therefore, the optimized training loss function is:
[0027]
[0028] Among them, α, β, and γ are weight factors to balance the relationship between losses, and are set to 0.5, 0.5, and 1 respectively based on our experimental results on the selection of loss function weights.
[0029] Preferably, the encoder-decoder architecture based on the visual state space block is a U-shaped network architecture designed based on the visual state space block. The encoder has a five-layer structure, the first layer consists of a convolutional layer and a patch coding layer, and the following four layers are composed of 1, 1, 2, and 1 visual state space blocks and patch merging layers respectively. The decoder and the encoder are designed in a symmetrical structure, and jump connections are used to splice low-level features and high-level features to reduce the loss of detail information in the information flow.
[0030] Preferably, the spatial link GRU module workflow includes:
[0031] After the global features and local features are input into the spatial link GRU module, they are input into the forward GRU subunit and the backward GRU subunit in sequence to dynamically model the anatomical structure relationship, and the forward spatial features and the backward spatial features are obtained. Each GRU subunit is fed with the current slice information x sand the hidden state information h of the adjacent slices s-1 To calculate the reset gate unit r s and update the gate control unit u s , r s Determines the retained information in adjacent slices, u s Controls the flow of input and output information, both of which help update the memory information of the current slice.
[0032] Specifically, the input three-dimensional CT image is divided into a series of two-dimensional slices along the axial direction (i.e., the Z axis) to form a one-dimensional slice sequence. The SLG module processes this sequence as a "time series". The forward GRU processes the slices in order from the 1st slice to the Nth slice, and the backward GRU models them in reverse order to capture the spatial context information across slices. Each time step processes only one 2D slice, and uses the hidden state of the previous (or next) slice to model the spatial association with the adjacent slices, strengthening the coherence modeling between anatomical structures. The relevant calculation formula is:
[0033] r s =σ(W r x s +K r h s-1 +b r )
[0034] u s =σ(W u x s +K u h s-1 +b u )
[0035] Among them, σ(*) is the Sigmoid activation function, W r , K r 、W u , K u is the learnable weight matrix, b r 、b u is the bias term;
[0036] h s =(1-u s )h s-1 +u s (τ(W h x s +K h r s h s-1 +b h ))
[0037] Among them, τ(*) is the hyperbolic tangent activation function, W h , K hare the learnable weight matrices, b h is the bias term.
[0038] In the GRU subunit, we replace the traditional weight multiplication operation with a two-dimensional convolution with a kernel size of 3 to better preserve and utilize the spatial structure information of the input data. In the gated unit, the output is constrained to be between 0 and 1 through the Sigmoid activation function, thereby accurately controlling the flow and retention of information. Finally, we transform the forward feature of each slice into and backward features The images are then concatenated sequentially and fed into the decoder. The spatially linked GRU module is designed to share features between adjacent slices, effectively utilizing the spatial anatomical information in the pig scans. It also extracts features from each 2D slice, ultimately achieving segmentation results comparable to or even superior to those achieved using 3D convolution.
[0039] Preferably, the global organ category encoding module extracts global organ category features and generates a 3D segmentation map through a cross-attention mechanism, specifically comprising:
[0040] Input the multi-scale feature map generated in the decoding stage, perform feature alignment and feature fusion to obtain feature f;
[0041] Initialize the query vector Q. The query vector Q is used to learn organ category features. Its shape is 1, N, and C. Since Q is a globally shared learnable parameter, its first dimension is set to 1, indicating that the same set of category feature representations is shared between different samples. N refers to the number of categories, and C refers to the number of channels.
[0042] After aligning the feature maps from the three scales, we concatenate them into a single feature map f. The concatenated feature map f is then converted into tokens suitable for the attention mechanism, where each pixel position is considered as a token. Specifically, the channel information of each position in the feature map f is flattened into a token, resulting in the tokenized feature T f , which is a matrix containing multiple tokens. The tokenized feature T f The dimension is (HW,C * ), where HW is the spatial size of the concatenated feature map, C * is the number of channels per token;
[0043] Then T f As key and value input to the Cross Attention Module (CA) for update, the specific steps are as follows:
[0044] Query, key, and value calculations: The query vector q, key vector k, and value vector v are obtained by the following linear transformations:
[0045] q=QWQ k=T f W K v=T f W V
[0046] Where W Q 、W K 、W V is the parameter matrix of linear projection;
[0047] Attention score calculation and normalization: The similarity between the query and the key is calculated by dot product, and the score is normalized using Softmax to obtain the attention weight α:
[0048]
[0049] Weighted sum value: Use the attention weight to perform weighted summation on the value to obtain the updated feature representation:
[0050] CA(Q,T f )=α·vW
[0051] Among them, W is a learnable matrix used to perform the final linear transformation on the weighted summation result.
[0052] The above process is used to extract the global semantic feature representation of each organ category It is then used to obtain the updated global organ category expression vector Q1, as follows:
[0053]
[0054] Where Conv represents 1x1 convolution, LN refers to LayerNormalization;
[0055] Cross-attention module output features and the global organ class expression vector Q1, Generate a 3D segmentation map through the convolution layer;
[0056] The global organ category expression vector Q1 guides the generation of category-aware 3D segmentation maps of the strongly enhanced images.
[0057] A segmentation system for an automatic pig organ segmentation method based on CT anatomical structure relationships, comprising:
[0058] A pig CT image acquisition module, used to acquire the pig CT image to be segmented;
[0059] A multi-organ automatic segmentation module is used to process the pig CT image using the pig multi-organ automatic segmentation model to obtain a pig multi-organ prediction mask and complete the automatic segmentation of the pig multiple organs;
[0060] The pig multi-organ automatic segmentation module includes: an encoder-decoder architecture based on visual state space blocks, a spatial link GRU module, a global organ category encoding module, a global organ category guidance module and a semi-supervised training framework;
[0061] The encoder-decoder architecture based on the visual state space block uses linear complexity to capture long-range dependencies and extracts global and local features by filtering redundant information through a selective mechanism;
[0062] The spatial link GRU module dynamically models the changes between slices to mine the spatial anatomical structure relationship of the whole-body CT scan image of the pig, thereby obtaining the relative position information and detailed features of different organs;
[0063] The global organ category encoding module uses the attention mechanism to encode organ categories, which is used to help the network learn the category expression information of different organs and obtain global organ category expression features;
[0064] The global organ category guidance module guides the automatic segmentation of multiple organs in the strongly enhanced whole-body CT scan image of the pig by obtaining the global organ category expression features;
[0065] The semi-supervised training framework refers to a consistency regularization training framework. The model improves the automatic segmentation performance and generalization ability of the model by aligning the segmentation results of the strongly enhanced whole-body CT scan image of the pig guided by the global organ category expression features and the segmentation results without guidance.
[0066] An electronic device comprising a memory and a processor;
[0067] The memory is used to store computer programs;
[0068] The processor is configured to implement a method for automatically segmenting pig organs when executing the computer program.
[0069] The storage medium stores a computer program, and when the computer program is executed by the processor, the automatic pig organ segmentation method is implemented.
[0070] Compared with the prior art, the present invention has the following beneficial effects:
[0071] The present invention can batch process pig CT images, accurately segment them in the original image, and efficiently segment multiple organs in the CT image; the present invention designs an encoder-decoder architecture based on visual state space blocks, and efficiently extracts context information with linear complexity; constructs a spatial link GRU module to extract anatomical structure information and capture the dynamic spatial relationship between organs, thereby reducing segmentation deviations caused by changes in organ shape; develops a global organ category encoding module and a global organ category guidance module, and combines consistency regularization and attention mechanism to enable the network to accurately extract global organ category expression in the decoding stage. Even if the number of segmentation samples is limited, the segmentation consistency of organs of different sizes can be significantly improved, effectively solving the current problem of multi-organ segmentation in pigs, and providing a more efficient and accurate solution for medical image analysis and animal husbandry management of pigs. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 A schematic diagram of the basic flow of a method for automatic pig organ segmentation based on CT anatomical structure relationships provided by one embodiment of the present invention.
[0073] Figure 2 A schematic diagram of the visual state space module structure of a method for automatic pig organ segmentation based on CT anatomical structure relationships provided by one embodiment of the present invention.
[0074] Figure 3 A schematic diagram of the spatial link GRU module structure of an automatic pig organ segmentation method based on CT anatomical structure relationships provided by one embodiment of the present invention.
[0075] Figure 4 A schematic diagram of the global organ category encoding and guidance module structure of an automatic pig organ segmentation method based on CT anatomical structure relationships provided by one embodiment of the present invention.
[0076] Figure 5 This is the original segmentation image of the automatic pig organ segmentation method based on CT anatomical structure relationship provided by one embodiment of the present invention.
[0077] Figure 6 Annotated multi-organ image of a method for automatic pig organ segmentation based on CT anatomical structure relationships provided by one embodiment of the present invention.
[0078] Figure 7 This is a predicted mask image of a method for automatic pig organ segmentation based on CT anatomical structure relationships provided by one embodiment of the present invention.
[0079] Figure 8 This figure shows the comparison results of pig organ segmentation performance between the automatic pig organ segmentation method based on CT anatomical structure relationships provided by one embodiment of the present invention and the manual cutting method and traditional segmentation method. DETAILED DESCRIPTION
[0080] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0081] Example 1
[0082] A segmentation system for an automatic pig organ segmentation method based on CT anatomical structure relationships, comprising:
[0083] A pig CT image acquisition module, used to acquire the pig CT image to be segmented;
[0084] A multi-organ automatic segmentation module is used to process the pig CT image using the pig multi-organ automatic segmentation model to obtain a pig multi-organ prediction mask and complete the automatic segmentation of the pig multiple organs;
[0085] The pig multi-organ automatic segmentation module includes: an encoder-decoder architecture based on visual state space blocks, a spatial link GRU module, a global organ category encoding module, a global organ category guidance module and a semi-supervised training framework;
[0086] The encoder-decoder architecture based on the visual state space block uses linear complexity to capture long-range dependencies and extracts global and local features by filtering redundant information through a selective mechanism;
[0087] The spatial link GRU module dynamically models the changes between slices to mine the spatial anatomical structure relationship of the whole-body CT scan image of the pig, thereby obtaining the relative position information and detailed features of different organs;
[0088] The global organ category encoding module uses the attention mechanism to encode organ categories, which is used to help the network learn the category expression information of different organs and obtain global organ category expression features;
[0089] The global organ category guidance module guides the automatic segmentation of multiple organs in the strongly enhanced whole-body CT scan image of the pig by obtaining the global organ category expression features;
[0090] The semi-supervised training framework refers to a consistency regularization training framework. The model improves the automatic segmentation performance and generalization ability of the model by aligning the segmentation results of the strongly enhanced whole-body CT scan image of the pig guided by the global organ category expression features and the segmentation results without guidance.
[0091] Example 2
[0092] See also Figure 1-Figure 5 , a method for automatic segmentation of pig organs based on CT anatomical structure relationship, comprising the following steps:
[0093] Step 1: Full-body CT scans of live pigs were obtained. Veterinarians and radiologists manually annotated masks of multiple organs using 3D Slicer. Training and testing sample sets were then constructed through data partitioning. It should be noted that each sample refers to a 3D matrix generated from all CT slices of a single pig.
[0094] Step 2: Using random data enhancement technology on the whole-body CT scan image of the pig and its mask, obtain a strongly enhanced whole-body CT scan image of the pig and its mask;
[0095] Step 3: Use the pig multi-organ automatic segmentation model to train the pig whole-body CT scan image and its mask and the strongly enhanced pig whole-body CT scan image and its mask to obtain the pig multi-organ prediction mask and complete the automatic segmentation of the pig multiple organs;
[0096] The training process of the pig multi-organ automatic segmentation model is as follows:
[0097] S1. Input a whole-body CT scan image sample of a pig and a corresponding strongly enhanced whole-body CT scan image sample of a pig at the same time, and extract global features and local features through an encoder-decoder architecture based on the visual state space;
[0098] S2, input the global features and local features extracted in S1 into the spatial link GRU module, dynamically model the changes between slices, and obtain the relative position information and detailed features of different organs;
[0099] S3. For the whole-body CT scan image samples of live pigs, the global organ category features are extracted through the global organ category encoding module, and a category-aware 3D segmentation map M is generated. The segmentation loss and auxiliary loss are calculated with the ground-truth mask GT.
[0100] S4. Sample of a strongly enhanced whole-body CT scan image of a pig, using the global organ category guidance module to generate a guided 3D segmentation map G s , the unguided 3D segmentation map M is generated by the global organ category encoding module s , calculate G s and M s consistency loss.
[0101] The random data enhancement techniques in step 2 include random flipping, rotation, contrast adjustment, and Gaussian noise addition;
[0102] The segmentation loss, auxiliary loss and consistency loss in the training of the pig multi-organ automatic segmentation model are training loss functions, the segmentation loss is used to guide network segmentation learning, the auxiliary loss is used to optimize the high-level semantic information and edge detail information of the extracted features, and the consistency loss is used to constrain the consistency of the segmentation results of the strongly enhanced pig whole-body CT scan image under different conditions;
[0103] The calculation formula of the segmentation loss is:
[0104]
[0105] in: is the cross entropy loss, for dice losses;
[0106] The auxiliary loss is calculated as follows:
[0107]
[0108] in: is the Softmax operation, M λ is the segmentation prediction result of the original whole-body CT scan image of the live pig or the strongly enhanced whole-body CT scan image of the live pig, It is a bilinear interpolation operation used to upsample the segmentation prediction result to the same size as the mask GT;
[0109] The calculation formula of the consistency loss is:
[0110]
[0111] in, is the Softmax operation, G s is the segmentation result of the strongly enhanced pig whole-body CT scan image guided by the global organ category expression feature, M s is the segmentation result of the strongly enhanced whole-body swine CT scan image without the guidance of the global organ category expression feature, n is the batch size, and i is the slice index;
[0112] Therefore, the optimized training loss function is:
[0113]
[0114] Among them, α, β, and γ are weight factors to balance the relationship between losses. Based on our experimental results on the selection of loss function weights, they are set to 0.5, 0.5, and 1 respectively;
[0115] The encoder-decoder architecture based on the visual state space block is a U-shaped network architecture designed based on the visual state space block. The encoder has a five-layer structure. The first layer consists of a convolutional layer and a patch coding layer. The following four layers are composed of 1, 1, 2, and 1 visual state space blocks and a patch merging layer respectively. The decoder and the encoder are designed in a symmetrical structure. Skip connections are used to splice low-level features and high-level features to reduce the loss of detail information in the information flow. The structure of the visual state space block is as follows: Figure 3 As shown in (a), the design of the 2D-Selective-Scan (SS2D) module is adopted, and its processing flow includes the following three parts:
[0116] S1. Four-way cross sequence modeling: The input feature map is first scanned and expanded along four directions (horizontally from tile 1 to the left, horizontally from tile 3 to the right, vertically from tile 9 to the top, and vertically from tile 1 to the bottom). In the specific implementation, the feature map is divided by channel through row and column-level convolution operations and then serialized in sequence to obtain four groups of one-dimensional sequence representations. These sequences retain structural information in different directions and help model spatial context. The input feature map is first expanded by selective scanning in four directions: horizontal (left, right) and vertical (up, down). In the implementation, the SS2D module divides the input channels into several subgroups, and applies column-based (Conv-H) and row-based (Conv-V) convolution operations to generate directional embeddings (DirectionalEmbedding), which are then expanded into four groups of one-dimensional sequences in sequence. This design explicitly encodes spatial structures in different directions, which helps capture long-distance dependencies and cross-regional contextual information.
[0117] S2 and S6 selective fusion mechanism (see Figure 3 (b)):
[0118] To achieve dynamic suppression of redundant information and enhancement of important regions, the SS2D module introduces a gated selection mechanism, the S6 module. This module includes the following key steps:
[0119] Content Attention: Contextual modeling of sequences via lightweight MLP to generate position-dependent response strengths;
[0120] Position bias fusion: Fusion of learnable position biases to guide attention to focus on spatially salient areas;
[0121] Weighted fusion strategy: Softmax normalization is applied to the sequences in four directions to generate a weight distribution, and then weighted summation is performed to effectively retain the key structure and weaken redundant information.
[0122] S3. Sequence restoration and two-dimensional reconstruction: The fused four-way sequence results are aggregated into a unified feature expression through element-by-element summation, and then restored to a two-dimensional feature map of the original input resolution through deserialization mapping and channel integration operations, realizing the restoration of spatial structure and the injection of global context.
[0123] The spatial link GRU module workflow includes:
[0124] After the global features and local features are input into the spatial link GRU module, they are input into the forward GRU subunit and the backward GRU subunit in sequence to dynamically model the anatomical structure relationship, and the forward spatial features and the backward spatial features are obtained. Each GRU subunit is fed with the current slice information x s and the hidden state information h of the adjacent slices s-1 To calculate the reset gate unit r s and update the gate control unit u s , r s Determines the retained information in adjacent slices, u s Controls the flow of input and output information, both of which help update the memory information of the current slice.
[0125] Specifically, the input three-dimensional CT image is divided into a series of two-dimensional slices in the axial direction (i.e., the Z axis) to form a one-dimensional slice sequence. The SLG module treats the sequence as a "time series" for processing. The forward GRU processes the slices in order from the 1st slice to the Nth slice, and the backward GRU is modeled in reverse order to capture the spatial context information across slices. Only one 2D slice is processed at each time step, and the hidden state of the previous (or next) slice is used to model the spatial association with the adjacent slices, thereby strengthening the coherence modeling between anatomical structures. The relevant calculation formula is:
[0126] r s =σ(W r x s +K r h s-1 +b r )
[0127] u s =σ(W u x s +K u h s-1 +b u )
[0128] Among them, σ(*) is the Sigmoid activation function, W r , K r 、W u , K u is the learnable weight matrix, b r 、bu is the bias term;
[0129] h s =(1-u s )h s-1 +u s (τ(W h x s +K h r s h s-1 +b h ))
[0130] Among them, τ(*) is the hyperbolic tangent activation function, W h , K h are the learnable weight matrices, b h is the bias term;
[0131] In the GRU subunit, we replace the traditional weight multiplication operation with a two-dimensional convolution with a kernel size of 3 to better preserve and utilize the spatial structure information of the input data. In the gated unit, the output is constrained to be between 0 and 1 through the Sigmoid activation function, thereby accurately controlling the flow and retention of information. Finally, we transform the forward feature of each slice into and backward features The images are then concatenated sequentially and fed into the decoder. The spatially linked GRU module is designed to share features between adjacent slices, effectively utilizing the spatial anatomical information in the pig scans. It also extracts features from each 2D slice, ultimately achieving segmentation results comparable to or even superior to those achieved using 3D convolution.
[0132] The global organ category encoding module extracts global organ category features and generates a 3D segmentation map through a cross-attention mechanism, specifically including:
[0133] Input the multi-scale feature map generated in the decoding stage, perform feature alignment and feature fusion to obtain feature f;
[0134] Initialize the query vector Q. The query vector Q is used to learn organ category features. Its shape is 1, N, and C. Since Q is a globally shared learnable parameter, its first dimension is set to 1, indicating that the same set of category feature representations is shared between different samples. N refers to the number of categories, and C refers to the number of channels.
[0135] After aligning the feature maps from the three scales, we concatenate them into a single feature map f. The concatenated feature map f is then converted into tokens suitable for the attention mechanism, where each pixel position is considered as a token. Specifically, the channel information of each position in the feature map f is flattened into a token, resulting in the tokenized feature T f, which is a matrix containing multiple tokens. The tokenized feature T f The dimension is (HW,C * ), where HW is the spatial size of the concatenated feature map, C * is the number of channels per token;
[0136] Then T f As key and value input to the Cross Attention Module (CA) for update, the specific steps are as follows:
[0137] Query, key, and value calculations: The query vector q, key vector k, and value vector v are obtained by the following linear transformations:
[0138] q=QW Q k=T f W K v=T f W V
[0139] Where W Q 、W K 、W V is the parameter matrix of linear projection;
[0140] Attention score calculation and normalization: The similarity between the query and the key is calculated by dot product, and the score is normalized using Softmax to obtain the attention weight α:
[0141]
[0142] Weighted sum value: Use the attention weight to perform weighted summation on the value to obtain the updated feature representation:
[0143] CA(Q,T f )=α·vW
[0144] Among them, W is a learnable matrix used to perform the final linear transformation on the weighted summation result.
[0145] The above process is used to extract the global semantic feature representation of each organ category It is then used to obtain the updated global organ category expression vector Q1, as follows:
[0146]
[0147] Where Conv represents 1x1 convolution, LN refers to LayerNormalization;
[0148] Cross-attention module output features and the global organ class expression vector Q1, Generate a 3D segmentation map through the convolution layer;
[0149] The global organ category expression vector Q1 guides the generation of category-aware 3D segmentation maps of the strongly enhanced images.
[0150] The trained pig multi-organ automatic segmentation model is used to segment a whole-body CT image of a pig in the test set, and a multi-organ segmentation mask is obtained to complete the automatic and accurate segmentation of the pig's multiple organs.
[0151] Specifically, the unsegmented original CT images of pigs in the test set are converted into array form; these arrays are normalized; the arrays are input into the trained pig multi-organ automatic segmentation model for prediction, and the multi-organ image array predicted by the model is output to obtain the pig multi-organ segmentation results.
[0152] Example 3
[0153] See also Figure 6-Figure 8 The traditional technical solutions of pig multi-organ segmentation method and manual segmentation method were selected for comparison with the method of the present invention to verify the real effect of the method.
[0154] Experimental pigs were selected. First, the pigs were anesthetized by intramuscular injection of 0.1 mg / kg azaperidone, 0.2 mg / kg ketamine, and 0.22 mg / kg propofol. The pigs were then placed in a prone position on a CT bed and scanned using a CT device to obtain pig scan images. The CT images were saved in dicom format, and the size of each dicom image was set to 512 pixels × 512 pixels with a thickness of 2.5 mm. Professional veterinarians and radiologists used a semi-automated annotation tool called 3D slicer to perform manual annotations. Mask annotations were performed for each organ category based on the physiological structure of the pig and the morphological characteristics of each internal organ. A total of 60 pig scans were manually annotated, of which 40 samples were used as training data sets to participate in the training phase of the model, and the remaining 20 samples were used as test sets to evaluate the segmentation performance of the model.
[0155] A porcine organ automatic segmentation model based on CT anatomical structural relationships was constructed, comprising an encoder-decoder architecture based on visual state space blocks, a spatially linked GRU module, a global organ category encoding module, a global organ category guidance module, and a semi-supervised training framework. Model training was optimized using a training loss function consisting of a segmentation loss, an auxiliary loss, and a consistency loss. The segmentation loss guided network segmentation learning, the auxiliary loss optimized the high-level semantic information and edge detail of the extracted features, and the consistency loss constrained the consistency of the segmentation results of the strongly enhanced images under different conditions.
[0156] The comparison results of the traditional solution and the traditional manual segmentation solution are shown in Table 1 below:
[0157] Table 1: Comparison results between the present invention and the traditional solution
[0158] Comparison Object Manual cutting method Traditional segmentation methods Method of the present invention Accuracy / % 85 80 86 Segmentation accuracy / % 83 83 87 Time / s 3600-7200 25 10 efficiency Low Higher high
[0159] From the above comparison data, it can be seen that the method of the present invention is superior to the traditional scheme in accuracy, segmentation precision and efficiency, and can achieve rapid removal in a shorter time than the traditional scheme. The comparison results show that the method of the present invention can quickly, efficiently, accurately and automatically segment multiple organs in pig CT images.
[0160] Some steps in the embodiments of the present invention may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.
[0161] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for automatic pig organ segmentation based on CT anatomical structure relationships, characterized by: The following steps are included: Step 1: Obtain a full-body CT scan of a pig, and have veterinarians and radiologists manually annotate masks of multiple organs using 3D Slicer; Step 2: Using random data enhancement technology on the whole-body CT scan image of the pig and its mask, obtain a strongly enhanced whole-body CT scan image of the pig and its mask; Step 3: Use the pig multi-organ automatic segmentation model to train the pig whole-body CT scan image and its mask and the strongly enhanced pig whole-body CT scan image and its mask to obtain the pig multi-organ prediction mask and complete the automatic segmentation of the pig multiple organs; The training process of the pig multi-organ automatic segmentation model is as follows: S1. A whole-body CT scan image sample of a pig and a corresponding strongly enhanced whole-body CT scan image sample of a pig are simultaneously input into the model, and global and local features are extracted through an encoder-decoder architecture based on the visual state space; S2, input the global features and local features extracted in S1 into the spatial link GRU module, dynamically model the changes between slices, and obtain the relative position information and detailed features of different organs; S3. For the whole-body CT scan image samples of live pigs, the global organ category features are extracted through the global organ category encoding module, and a category-aware 3D segmentation map M is generated. The segmentation loss and auxiliary loss are calculated with the ground-truth mask GT. S4. Sample of a strongly enhanced whole-body CT scan image of a pig, using the global organ category guidance module to generate a guided 3D segmentation map G s , the unguided 3D segmentation map M is generated by the global organ category encoding module s , calculate G s and M s consistency loss.
2. The method for automatic pig organ segmentation based on CT anatomical structure relationship according to claim 1, characterized in that: The random data enhancement technology in step 2 includes random flipping, rotation, contrast adjustment and Gaussian noise addition.
3. The method for automatic pig organ segmentation based on CT anatomical structure relationship according to claim 1, characterized in that: The segmentation loss, auxiliary loss and consistency loss in the training of the pig multi-organ automatic segmentation model are training loss functions, the segmentation loss is used to guide network segmentation learning, the auxiliary loss is used to optimize the high-level semantic information and edge detail information of the extracted features, and the consistency loss is used to constrain the consistency of the segmentation results of the strongly enhanced pig whole-body CT scan image under different conditions; The calculation formula of the segmentation loss is: in: is the cross entropy loss, for dice losses; The auxiliary loss is calculated as follows: in: is the Softmax operation, M λ is the segmentation prediction result of the original whole-body CT scan image of the live pig or the strongly enhanced whole-body CT scan image of the live pig, It is a bilinear interpolation operation used to upsample the segmentation prediction result to the same size as the mask GT; The calculation formula of the consistency loss is: in, is the Softmax operation, G s is the segmentation result of the strongly enhanced pig whole-body CT scan image guided by the global organ category expression feature, M s is the segmentation result of the strongly enhanced whole-body swine CT scan image without the guidance of the global organ category expression feature, n is the batch size, and i is the slice index; Therefore, the optimized training loss function is: Among them, α, β, and γ are weight factors to balance the relationship between losses. Based on our experimental results on the selection of loss function weights, they are set to 0.5, 0.5, and 1 respectively.
4. The method for automatic pig organ segmentation based on CT anatomical structure relationship according to claim 1, characterized in that: The encoder-decoder architecture based on visual state space blocks is a U-shaped network architecture designed based on visual state space blocks. The encoder has a five-layer structure. The first layer consists of a convolutional layer and a patch encoding layer. The following four layers consist of 1, 1, 2, and 1 visual state space blocks and a patch merging layer, respectively. The decoder and the encoder are designed in a symmetrical structure, and skip connections are used to splice low-level features and high-level features to reduce the loss of detail information in the information flow.
5. The method for automatic pig organ segmentation based on CT anatomical structure relationship according to claim 1, characterized in that: The spatial link GRU module workflow includes: After the global features and local features are input into the spatial link GRU module, they are input into the forward GRU subunit and the backward GRU subunit in sequence to dynamically model the anatomical structure relationship, and the forward spatial features and the backward spatial features are obtained. Each GRU subunit is fed with the current slice information x s and the hidden state information h of the adjacent slices s-1 To calculate the reset gate unit r s and update the gate control unit u s , r s Determines the retained information in adjacent slices, u s Controls the flow of input and output information, both of which help update the memory information of the current slice. Specifically, the input three-dimensional CT image is divided into a series of two-dimensional slices in the axial direction to form a one-dimensional slice sequence. The SLG module treats this sequence as a "time series" for processing. The forward GRU processes the slices in order from the 1st to the Nth slice, and the backward GRU models them in reverse order to capture the spatial context information across slices. Each time step only processes one 2D slice, and uses the hidden state of the previous slice to model the spatial association with the adjacent slices, thereby strengthening the coherence modeling between anatomical structures. The relevant calculation formula is: r s =σ(W r x s +K r h s-1 +b r ) u s =σ(W u x s +K u h s-1 +b u ) Among them, σ(*) is the Sigmoid activation function, W r , K r 、W u , K u is the learnable weight matrix, b r 、b u is the bias term; h s =(1-u s )h s-1 +u s (τ(W h x s +K h r s h s-1 +b h )) Among them, τ(*) is the hyperbolic tangent activation function, W h , K h are the learnable weight matrices, b h is the bias term; In the GRU subunit, the traditional weight multiplication operation is replaced by a two-dimensional convolution with a kernel size of 3 to better preserve and utilize the spatial structure information of the input data. In the gated unit, the output is constrained to between 0 and 1 through the Sigmoid activation function, thereby accurately controlling the flow and retention of information. Finally, we transform the forward feature of each slice into and backward features After being spliced sequentially, they are input into the decoder. The spatial link GRU module is designed to share features between adjacent slices, effectively utilizing the spatial anatomical information in the pig body scan image, and simultaneously extracting features from each two-dimensional slice, ultimately achieving a segmentation effect that is comparable to or even better than using three-dimensional convolution.
6. The method for automatic pig organ segmentation based on CT anatomical structure relationship according to claim 1, characterized in that: The global organ category encoding module extracts global organ category features and generates a 3D segmentation map through a cross-attention mechanism, specifically including: Input the multi-scale feature map generated in the decoding stage, perform feature alignment and feature fusion to obtain feature f; Initialize the query vector Q. The query vector Q is used to learn organ category features. Its shape is 1, N, and C. Since Q is a globally shared learnable parameter, its first dimension is set to 1, indicating that the same set of category feature representations is shared between different samples. N refers to the number of categories, and C refers to the number of channels. After aligning the feature maps from the three scales, they are concatenated into a single feature map f. Then, the concatenated feature map f is converted into a token suitable for the attention mechanism, where each pixel position is regarded as a token. Specifically, the channel information of each position in the feature map f is flattened into a token to obtain the tokenized feature T. f , this is a matrix containing multiple tokens, the tokenized feature T f The dimension is (HW,C * ), where HW is the spatial size of the concatenated feature map, C * is the number of channels per token; Then T f As key and value input to the Cross Attention Module (CA) for update, the specific steps are as follows: Query, key, and value calculations: The query vector q, key vector k, and value vector v are obtained by the following linear transformations: q=QW Q k=T f W K v=T f W V Where W Q 、W K 、W V is the parameter matrix of linear projection; Attention score calculation and normalization: The similarity between the query and the key is calculated by dot product, and the score is normalized using Softmax to obtain the attention weight α: Weighted sum value: Use the attention weight to perform weighted summation on the value to obtain the updated feature representation: CA(Q,T f )=α vW Among them, W is a learnable matrix used to perform the final linear transformation on the weighted summation result; The above process is used to extract the global semantic feature representation of each organ category It is then used to obtain the updated global organ category expression vector Q1, as follows: Where Conv represents 1x1 convolution, LN refers to LayerNormalization; Cross-attention module output features and the global organ class expression vector Q1, Generate a 3D segmentation map through the convolution layer; The global organ category expression vector Q1 guides the generation of category-aware 3D segmentation maps of the strongly enhanced images.
7. The segmentation system of any one of claims 1 to 6, wherein: include: A pig CT image acquisition module, used to acquire the pig CT image to be segmented; A multi-organ automatic segmentation module is used to process the pig CT image using the pig multi-organ automatic segmentation model to obtain a pig multi-organ prediction mask and complete the automatic segmentation of the pig multiple organs; The pig multi-organ automatic segmentation module includes: an encoder-decoder architecture based on visual state space blocks, a spatial link GRU module, a global organ category encoding module, a global organ category guidance module and a semi-supervised training framework; The encoder-decoder architecture based on the visual state space block uses linear complexity to capture long-range dependencies and extracts global and local features by filtering redundant information through a selective mechanism; The spatial link GRU module dynamically models the changes between slices to mine the spatial anatomical structure relationship of the whole-body CT scan image of the pig, thereby obtaining the relative position information and detailed features of different organs; The global organ category encoding module uses the attention mechanism to encode organ categories, which is used to help the network learn the category expression information of different organs and obtain global organ category expression features; The global organ category guidance module guides the automatic segmentation of multiple organs in the strongly enhanced whole-body CT scan image of the pig by obtaining the global organ category expression features; The semi-supervised training framework refers to a consistency regularization training framework. The model improves the automatic segmentation performance and generalization ability of the model by aligning the segmentation results of the strongly enhanced whole-body CT scan image of the pig guided by the global organ category expression features and the segmentation results without guidance.
8. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by the processor, the method for automatic segmentation of pig organs according to any one of claims 1 to 6 is implemented.
9. An electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to execute the computer program to implement the automatic pig organ segmentation method according to any one of claims 1 to 6.
Citation Information
Cited By
Low-label pig CT image segmentation method based on multi-task self-supervised learning
CN121437893A