Method for segmentation of ct soil porosity based on liquid state neural network and edge enhancement

CN121685580BActive Publication Date: 2026-09-18HAINAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511915205.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-09-18
Estimated Expiration
2045-12-18

AI Technical Summary

Technical Problem

[0003]然而,现有基于UNet架构的CT土壤孔隙分割方法仍存在明显局限:首先,在解码阶段的特征融合过程中,常采用静态融合策略,难以充分捕捉多分辨率特征之间的动态依赖关系,导致细节信息丢失,尤其是在孔隙边缘区域;其次,由于CT图像中土壤孔隙与周围背景在灰度与纹理上高度相似,传统融合方式在边缘敏感特征提取与细粒度细节恢复方面表现不足,影响分割精度;再次,虽然液体神经网络在时序数据处理中表现出色,但其在图像分割领域的应用仍较为有限,尤其是在CT土壤图像等静态图像分析任务中,其动态建模潜力尚未得到有效发掘与利用

Benefits of technology

本发明的编码器采用了液态频率特征提取模块(Liquid-Frequency FeatureExtraction Module, LFFEM)。该模块基于液态神经网络原理,通过内部状态的迭代演化对输入特征进行动态处理,克服了传统卷积操作固有的静态局限性,能够更有效地捕捉CT土壤图像中孔隙结构的复杂多尺度上下文信息,实现了动态、自适应的特征提取,显著增强了模型的特征表示能力,同时大幅降低了模型复杂度和计算开销。此外,该模块特有的结构设计(如频率分解与多核提取)在提升特征质量的同时,实现了模型参数量和计算量的显著减少,达成了精度与效率的统一。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685580B_ABST
    Figure CN121685580B_ABST
Patent Text Reader

Abstract

The application discloses a CT soil pore segmentation method based on a liquid neural network and edge enhancement, and belongs to the technical field of image segmentation. The method comprises the following steps: acquiring and preprocessing a CT soil image to construct a data set; constructing a segmentation model based on an encoder-decoder architecture, wherein the encoder adopts a liquid frequency feature extraction module to realize dynamic multi-scale feature extraction, and the decoder adopts an edge enhancement fusion module to explicitly enhance edge information in the feature fusion process; training and optimizing the model by using the data set; and finally applying the trained model to realize high-precision pore segmentation. The liquid frequency feature extraction module improves the feature representation capability of the model and reduces the calculation overhead, and the edge enhancement fusion module significantly improves the segmentation accuracy of the pore edge, so that more accurate and efficient segmentation of the CT soil pore is realized as a whole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image segmentation technology, and particularly relates to a CT soil pore segmentation method based on liquid neural network and edge enhancement. Background Technology

[0002] Soil, as a typical porous medium, has a pore structure that significantly impacts soil aeration, water retention, and the biogeochemical processes occurring within it. Therefore, accurate segmentation of soil pores in CT images is crucial for quantitatively characterizing pore structure, studying water and air transport mechanisms, and guiding agricultural production and ecological restoration. In image segmentation, UNet and its improved architectures, due to their encoder-decoder structure and skip connection design, have become fundamental models for various image segmentation tasks, effectively extracting multi-scale contextual features. Existing methods typically employ feature fusion strategies in the decoding stage, such as element-wise addition, channel stitching, or weighted fusion, to integrate feature information at different resolutions. While maintaining computational efficiency, these methods achieve a certain degree of information fusion, providing a feasible path for pore segmentation in CT images. Furthermore, in recent years, Liquid Neural Networks (LNNs) have demonstrated excellent nonlinear fitting capabilities in time series and dynamic system modeling. Their dynamic time-varying state adjustment mechanism allows them to adapt to complex data patterns, and they have been explored for application in some medical image analysis tasks.

[0003] However, existing CT soil pore segmentation methods based on the UNet architecture still have significant limitations: First, during feature fusion in the decoding stage, static fusion strategies are often used, making it difficult to fully capture the dynamic dependencies between multi-resolution features, leading to the loss of detailed information, especially in the pore edge region. Second, because soil pores in CT images are highly similar to the surrounding background in grayscale and texture, traditional fusion methods are insufficient in edge-sensitive feature extraction and fine-grained detail recovery, affecting segmentation accuracy. Third, although liquid neural networks perform well in temporal data processing, their application in image segmentation is still relatively limited, especially in static image analysis tasks such as CT soil images, where their dynamic modeling potential has not been effectively explored and utilized. Therefore, existing technologies still face challenges in achieving high-precision and high-efficiency CT soil pore segmentation. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention proposes a CT soil pore segmentation method based on liquid neural networks and edge enhancement, thereby resolving the issues present in the prior art.

[0005] To achieve the above objectives, this invention provides a CT soil pore segmentation method based on liquid neural networks and edge enhancement, comprising: S1. Acquire and preprocess CT soil slice images to construct a dataset; S2. Construct a segmentation model based on an encoder-decoder architecture; the encoder includes a liquid frequency feature extraction module for extracting dynamic multi-scale features; the decoder includes an edge enhancement fusion module for fusing features of different resolutions and enhancing edge information. S3. Use the dataset to train and optimize the segmentation model to obtain a trained segmentation model; S4. Use the trained segmentation model to segment the target CT soil image and output the pore segmentation result.

[0006] Preferably, the execution process of the liquid frequency feature extraction module includes: performing frequency decomposition on the input features to obtain external input features; performing multi-kernel feature extraction on the internal hidden state; concatenating the external input features with the internal hidden state features, and solving the problem based on the iterative process of the liquid neural network to output a deep feature map.

[0007] Preferably, the frequency decomposition process of the input features includes: splitting the input feature map into a first sub-feature map and a second sub-feature map; extracting low-frequency features from the first sub-feature map through pooling, convolution, and upsampling operations; extracting high-frequency features from the second sub-feature map by subtracting the mean and then using the Sobel operator and convolution operations; and fusing the low-frequency features with the high-frequency features.

[0008] Preferably, the multi-kernel feature extraction process includes: splitting the internal hidden state into multiple sub-states; processing the multiple sub-states respectively using multiple depthwise separable convolutions with different kernel sizes; and fusing the processed sub-states.

[0009] Preferably, the iterative process of the liquid neural network includes: calculating the time constant and bias vector for the concatenated features; and iteratively updating the hidden state according to the set iteration parameters using the following formula: ; in, Hide the current state. For time step, It is a time constant. For bias vectors, These are the features after splicing.

[0010] Preferably, the execution process of the edge enhancement fusion module includes: performing convolutional preprocessing on the input low-resolution feature map and high-resolution feature map respectively; generating an edge information mask for the high-resolution feature map; calculating the dynamic fusion coefficients of the low-resolution feature map and the high-resolution feature map respectively; extracting fine-grained edge information from both; fusing the edge information based on the dynamic fusion coefficients; and concatenating and further fusing the fused edge information with the upsampled low-resolution feature map.

[0011] Preferably, the dynamic fusion coefficient is calculated by a pooling convolution activation module; the fine-grained edge information is extracted by the Sobel operator.

[0012] Preferably, in step S3, the cross-entropy loss function is used for model training, and the average crossover ratio is calculated on the validation set to evaluate the model performance. Training is stopped when the metric no longer improves within a preset number of iterations.

[0013] Preferably, in step S1, the preprocessing includes: performing interval sampling on the original CT slice image, cropping with the maximum inscribed square, and dividing it into image blocks of uniform size.

[0014] Preferably, in step S4, the process of outputting the aperture segmentation result is as follows: the multi-channel feature map output by the model is compared with the channel values ​​at each pixel point, and the channel category corresponding to the maximum value is used as the segmentation label of the pixel point to generate the final binary or semantic segmentation map.

[0015] Compared with the prior art, the present invention has the following advantages and technical effects: The encoder of this invention employs a Liquid-Frequency Feature Extraction Module (LFFEM). Based on the principles of liquid neural networks, this module dynamically processes input features through the iterative evolution of its internal states, overcoming the inherent static limitations of traditional convolution operations. It can more effectively capture the complex multi-scale contextual information of pore structures in CT soil images, achieving dynamic and adaptive feature extraction, significantly enhancing the model's feature representation capabilities while substantially reducing model complexity and computational overhead. Furthermore, the module's unique structural design (such as frequency decomposition and multi-kernel extraction) improves feature quality while significantly reducing the number of model parameters and computational cost, achieving a balance between accuracy and efficiency.

[0016] The decoder of this invention employs an Edge Enhanced Fusion Module (EEFM). This module explicitly extracts and enhances edge information when fusing deep semantic features from the encoder and shallow detail features from the decoder. By dynamically calculating fusion coefficients and weighting them with the edge map, this module enables the model to more accurately perceive and recover low-contrast, blurred boundaries between apertures and the background, effectively overcoming the problem of unclear edge segmentation caused by feature similarity.

[0017] The LFFEM module provides high-quality features rich in dynamic contextual information during the encoding stage, laying the foundation for subsequent fine-grained fusion. The EEFM module utilizes this information during the decoding stage to achieve accurate feature reconstruction and restoration with edge sensitivity. Through the synergistic effect of these two core modules, this invention significantly improves the overall segmentation performance of the final model on publicly available CT soil image datasets, outperforming traditional segmentation models such as UNet. Attached Figure Description

[0018] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of the CT soil pore segmentation method based on liquid neural network and edge enhancement module according to an embodiment of the present invention; Figure 2 The segmentation model (Liquid Frequency Extraction and Edge Enhancement Network, LFE) of this invention is shown in the embodiment. 3 -Net) architecture diagram; Figure 3 This is a structural diagram of the LFFEM model according to an embodiment of the present invention; Figure 4 This is a structural diagram of the edge enhancement module (EEFM) block according to an embodiment of the present invention; Figure 5 The images show a comparison of the experimental segmentation results of this invention; where (a) is a CT slice image; (b) is a label image; (c) is a ViT segmentation result image; (d) is a DDRNet segmentation result image; (e) is a UNet segmentation result image; and (f) is an LFE segmentation result image. 3 -Net segmentation result image. Detailed Implementation

[0019] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0020] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0021] Example 1 like Figure 1 As shown, this embodiment provides a CT soil pore segmentation method based on liquid neural networks and edge enhancement. It employs the ordinary differential equations of liquid neural networks to design a liquid frequency feature extraction module (LFFEM), which effectively improves the model's dynamic feature extraction capability and significantly reduces the number of model parameters. Edge information from feature maps of different resolutions is fused to enhance the model's ability to perceive and recognize soil pore edges. The method specifically includes: S1. Acquire and preprocess CT soil slice images to construct a dataset; Further, in step S1, the preprocessing includes: performing interval sampling on the original CT slice image, cropping with the maximum inscribed square, and dividing it into image blocks of uniform size.

[0022] Specifically, soil samples with a diameter of 10 cm and a height of 10 cm were collected at the Meiting Agro-Forestry Ecosystem Observation and Research Station in Chengmai County, Hainan Province. A total of 6450 CT soil slice images were obtained by scanning these samples along the Z-axis (vertical direction) using a CT scanner. To avoid interference from similar information in adjacent CT soil images, CT soil slices were selected at intervals of 50 to create the dataset. Each slice contains 3000 images. The image is a 3000-pixel grayscale image with noise information such as container walls and damaged soil at the edges. Therefore, the image is cropped to a resolution of 1536×1536 using the largest inscribed square, and then further cropped into nine 512×512 patches. The final cropped images are then divided into training, validation, and test sets in a 7:2:1 ratio.

[0023] S2. Construct a segmentation model based on an encoder-decoder architecture; the encoder includes a liquid frequency feature extraction module for extracting dynamic multi-scale features; the decoder includes an edge enhancement fusion module for fusing features of different resolutions and enhancing edge information. Furthermore, the execution process of the liquid frequency feature extraction module includes: performing frequency decomposition on the input features to obtain external input features; performing multi-kernel feature extraction on the internal hidden state; concatenating the external input features with the internal hidden state features, solving the problem based on the iterative process of the liquid neural network, and outputting a deep feature map.

[0024] Furthermore, the frequency decomposition process for the input features includes: splitting the input feature map into a first sub-feature map and a second sub-feature map; extracting low-frequency features from the first sub-feature map through pooling, convolution, and upsampling operations; extracting high-frequency features from the second sub-feature map by subtracting the mean and then using the Sobel operator and convolution operations; and fusing the low-frequency features with the high-frequency features.

[0025] Furthermore, the multi-kernel feature extraction process includes: splitting the internal hidden state into multiple sub-states; processing the multiple sub-states separately using multiple depthwise separable convolutions with different kernel sizes; and fusing the processed sub-states.

[0026] Furthermore, the iterative process of the liquid neural network includes: solving for the time constant and bias vector for the spliced ​​features; Furthermore, the execution process of the edge enhancement fusion module includes: performing convolutional preprocessing on the input low-resolution feature map and high-resolution feature map respectively; generating an edge information mask for the high-resolution feature map; calculating the dynamic fusion coefficients of the low-resolution feature map and the high-resolution feature map respectively; extracting fine-grained edge information from both respectively; fusing the edge information based on the dynamic fusion coefficients; and concatenating and further fusing the fused edge information with the upsampled low-resolution feature map.

[0027] Furthermore, the dynamic fusion coefficients are calculated using a pooling convolution activation module; the fine-grained edge information is extracted using the Sobel operator.

[0028] Specifically, the segmentation model is defined as LFE. 3 The -Net model is trained using data from the training set. LFE is one such model. 3 -Net is designed based on the UNet architecture, such as Figure 2 As shown. The model structure consists of two parts: an encoder and a decoder. For the encoder, the DoubleConv module is defined first. This module contains two consecutive convolutional modules Conv (the Conv module contains three basic modules: standard convolution Conv2d + batch normalization BN + activation function ReLU). It performs preliminary feature extraction on the input CT soil slice image and increases the number of channels from the original 3 channels to 64 channels, as shown in the following formula: (1) (2) in, Given the input CT soil slice image, Conv2d, BN, and ReLU represent the standard convolution operation, batch normalization operation, and activation function, respectively. This is the output feature map of the DoubleConv module.

[0029] Then, the MaxPool-DoubleConv module is defined. This module first uses the standard max-pooling MaxPool2d operation to downsample the feature map to 256×256, and then uses DoubleConv to extract features from the downsampled feature map and double the number of channels to 128, as shown in the following formula: (3) (4) in, This represents the standard max pooling operation. This is the output feature map of the MaxPool-DoubleConv module.

[0030] Repeat the above process, and define another MaxPool-DoubleConv module to process the feature map. Further downsampling was used to obtain a feature map with a resolution of 128×128 and 256 channels. As shown in the following formula: (5) (6) Next, define the MaxPool-LFFEM module. The LFFEM module is as follows: Figure 3 As shown. This module first uses the standard max-pooling (MaxPool2d) operation to downsample the feature map to 64×64, then uses LFFEM to extract features from the downsampled feature map and increases the number of channels to 512, as shown in the following formula: (7) (8) Repeat the above process, and define another MaxPool-LFFEM module to convert the feature map... Further downsampling was used to obtain a feature map with a resolution of 32×32 and the number of channels remaining at 512. As shown in the following formula: (9) (10) For the LFFEM module, this module uses FDM to perform frequency decomposition and fusion on the input feature map. This process first splits the input feature map into two sub-feature maps by channel, where the sub-feature maps... Low-frequency features are extracted using average pooling, convolution, and upsampling modules, with the following formula: (11) Another sub-feature map After obtaining high-frequency information by subtracting its mean, high-frequency features are extracted using the Sobel operator and convolution module.

[0031] (12) Then, the two sub-feature maps are concatenated by channels and fused using a convolutional module to obtain a new external input feature map. As shown in the following formula: (13) Meanwhile, the internal state of the module uses the multi-kernel feature extraction module MkFEM to extract features from the internal hidden state. This process first uses a convolution module for preprocessing, and then the processed hidden state is split into 4 sub-states by channel. , , and The four sub-states were processed using depthwise separable convolutions with kernels of 3, 5, 7, and 9, respectively. Then, the four sub-states were concatenated by channel and fused using a convolutional module to obtain the features of the hidden state. As shown in the following formula: (14) Then and Assembled according to channels Then, the iterative process of the liquid neural network is performed. The iterative process of the liquid neural network first uses the DPCM module to... Solving for the time constant and bias vector As shown in the following formula: (15) (16) Then, set the liquid time. T and number of iterations step And obtain the time step. The hidden state at the next time step is solved iteratively using the following formula: (17) in, Hide the current state. For time step, It is a time constant. For bias vectors, These are the features after splicing.

[0032] Finally, the convolution module is used to process the state at the last time step. The output feature map of LFFEM is obtained through calculation.

[0033] In the decoder stage, an EEFM module is used to fuse different features. The EEFM module is as follows: Figure 4 As shown, the formula is: (18) (19) (20) (twenty one) in, , , , These are the outputs of the four EEFM modules. The resolution is 64×64, and the number of channels is 256; The resolution is 128×128, and the number of channels is 128; The resolution is 256×256, and the number of channels is 64; The resolution is 512×512, and the number of channels is 64. For the EEFM module, the inputs are low-resolution feature maps. and high-resolution feature maps First, standard convolution operations are used to preprocess both to obtain... and Then to Perform upsampling, and then use a standard convolutional block to... Perform calculations to obtain the edge information mask of the high-resolution feature map. Further, two pooling convolutional activation modules (ACS) are used to respectively... and The mutual fusion coefficient was calculated. and Simultaneously using the Sobel operator pair and The fine-grained edge information of the two feature maps is calculated. and Then, the edge information is fused according to the following formula: (twenty two) (twenty three) Finally, and The channels are concatenated, and then fused using a DoubleConv module and a CBAM module to obtain the output feature map of the EEFM module.

[0034] After the last EEFM module, a standard convolution with a 1×1 kernel is used to convert the feature map... Convert to the final output of the model ,Right now: (twenty four) S3. Use the dataset to train and optimize the segmentation model to obtain a trained segmentation model; Furthermore, in step S3, the cross-entropy loss function is used for model training, and the average intersection-union ratio is calculated on the validation set to evaluate the model performance. Training stops when the metric no longer improves within a preset number of iterations.

[0035] Specifically, the model training parameters were set, the AdamW optimizer was used during training, the learning rate was set to 0.0001, the input image resolution was 512×512 pixels, the number of training iterations was 200, and the batch size was set to 4.

[0036] Based on the labels, the loss value is calculated using the cross-entropy loss function, as shown in the following formula: (25) In the formula, It's a real label. The model predicts the output. Gradients are calculated based on the calculated loss value, and the model parameters are updated.

[0037] The evaluation metrics are calculated using the model with updated parameters from each iteration on the validation set. and the optimal evaluation value When comparing, Update And save the current model parameters. The calculation method is as follows: (26) During training, if Training will stop if there are no updates for more than 40 times or if the number of iterations reaches 200.

[0038] S4. Use the trained segmentation model to segment the target CT soil image and output the pore segmentation result.

[0039] Further, in step S4, the process of outputting the aperture segmentation result is as follows: the multi-channel feature map output by the model is compared with the channel values ​​at each pixel point, and the channel category corresponding to the maximum value is used as the segmentation label of the pixel point to generate the final binary or semantic segmentation map.

[0040] Specifically, CT soil porosity segmentation is performed using the saved optimal model parameters. The CT soil image data to be segmented is input into the LFE. 3 The -Net model, calculated according to the process in step S2, yields a resolution of... N Prediction results of ×512×512 y (where N is the segmentation category), the classification label is obtained based on the index of the maximum value of N values ​​in each pixel, and the final segmentation result is obtained.

[0041] In this embodiment, training is stopped after 200 iterations, or when the optimal metric has not been updated for more than 40 iterations. The saved optimal model is used to predict the CT soil image to be segmented, and the final output of the model is obtained. This output is a 6×512×512 tensor (i.e., a feature map with a resolution of 512×512 and 6 channels). The values ​​of the 6 channels in each pixel are compared, and the index of the channel with the largest value is used as the classification result of that pixel. Finally, a 1×512×512 class mask, i.e., the segmentation result, can be obtained.

[0042] The experimental segmentation comparison results are shown in Table 1. The segmentation accuracy in the CT soil image dataset is significantly improved compared with the traditional UNet model, while the required model parameters and computational cost are greatly reduced.

[0043] Table 1 Comparison of experimental segmentation results, such as Figure 5 As shown, the proposed LFE 3 The -Net model can more accurately identify and segment soil pores.

[0044] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A CT soil pore segmentation method based on liquid state neural network and edge enhancement, characterized in that, Includes the following steps: S1. Acquire and preprocess CT soil slice images to construct a dataset; S2. Construct a segmentation model, which is based on an encoder-decoder architecture; the encoder includes a liquid frequency feature extraction module for extracting dynamic multi-scale features; The decoder includes an edge enhancement fusion module for fusing features of different resolutions and enhancing edge information; The execution process of the liquid frequency feature extraction module includes: performing frequency decomposition on the input features to obtain external input features; performing multi-kernel feature extraction on the internal hidden state; concatenating the external input features with the internal hidden state features, solving the problem based on the iterative process of the liquid neural network, and outputting a deep feature map. The multi-kernel feature extraction process includes: splitting the internal hidden state into multiple sub-states; processing the multiple sub-states separately using multiple depthwise separable convolutions with different kernel sizes; and fusing the processed sub-states. The iterative process of the liquid neural network includes: calculating the time constant and bias vector for the concatenated features; and iteratively updating the hidden state according to the set iteration parameters using the following formula: ; wherein, is the hidden state at the current time instant, is the time step, is the time constant, is the bias vector, is the concatenated features; The execution process of the edge enhancement fusion module includes: performing convolutional preprocessing on the input low-resolution feature map and high-resolution feature map respectively; generating an edge information mask for the high-resolution feature map; calculating the dynamic fusion coefficients of the low-resolution feature map and the high-resolution feature map respectively; extracting fine-grained edge information from both; fusing the edge information based on the dynamic fusion coefficients; and concatenating and further fusing the fused edge information with the upsampled low-resolution feature map. S3. Use the dataset to train and optimize the segmentation model to obtain a trained segmentation model; S4. Use the trained segmentation model to segment the target CT soil image and output the pore segmentation result.

2. The method according to claim 1, characterized in that, The frequency decomposition process for the input features includes: splitting the input feature map into a first sub-feature map and a second sub-feature map; extracting low-frequency features from the first sub-feature map through pooling, convolution, and upsampling operations; extracting high-frequency features from the second sub-feature map by subtracting the mean and then using the Sobel operator and convolution operations; and fusing the low-frequency features with the high-frequency features.

3. The method according to claim 1, characterized in that, The dynamic fusion coefficients are calculated using a pooling convolution activation module; the fine-grained edge information is extracted using the Sobel operator.

4. The method according to claim 1, characterized in that, In step S3, the cross-entropy loss function is used to train the model, and the average crossover ratio is calculated on the validation set to evaluate the model performance. Training stops when the metric no longer improves within a preset number of iterations.

5. The method according to claim 1, characterized in that, In step S1, the preprocessing includes: performing interval sampling on the original CT slice image, cropping the image to the maximum inscribed square, and dividing it into image blocks of uniform size.

6. The method according to claim 1, characterized in that, In step S4, the process of outputting the pore segmentation result is as follows: the multi-channel feature map output by the model is compared with the channel values ​​at each pixel point, and the channel category corresponding to the maximum value is used as the segmentation label of the pixel point to generate the final binary or semantic segmentation map.