Sarcopenia ct image segmentation method based on deep learning
By constructing a deep learning-based segmentation model for sarcopenia CT images, the problems of uneven gray-level distribution, blurred boundaries, and weak texture response in sarcopenia CT image segmentation were solved, achieving more stable and accurate segmentation results, especially efficient segmentation under complex boundary conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-31
AI Technical Summary
Existing CT image segmentation methods for sarcopenia suffer from unstable segmentation results and insufficient boundary correction capabilities when there are uneven gray-level distributions, blurred tissue boundaries, weak local texture responses, and unclear target region attribution relationships.
A deep learning-based CT image segmentation model for sarcopenia is constructed, including a salient tissue response enhancement module, a target region expression generation module, and a decoding correction mask generation module. The model generates an initial tissue response map through a local neighborhood averaging operator, constructs multi-scale heterogeneous response features, performs adaptive weighted aggregation to generate stable enhancement features, and constructs semantic traction features through convolution and spatial position ranking to finally generate a segmentation mask.
It improves the accuracy and stability of CT image segmentation for sarcopenia, enhances segmentation precision and consistency under complex boundary conditions, and improves the ability to express target regions and fit boundaries.
Smart Images

Figure CN122492891A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image segmentation, specifically relating to a deep learning-based method for segmenting sarcopenia CT images. Background Technology
[0002] Existing sarcopenia assessment techniques typically require the identification and quantitative analysis of skeletal muscle regions based on CT images to help determine the degree of muscle loss and related functional status in subjects. Accurate segmentation of skeletal muscle regions is a crucial prerequisite for subsequent area measurement, muscle mass assessment, and clinical grading analysis. However, many existing techniques still rely on manual delineation, semi-automatic segmentation, or processing methods based on fixed thresholds, region growth, and simple texture rules. These methods are not only cumbersome to operate, highly dependent on operator experience, and have poor repeatability and consistency, but also prone to problems such as inaccurate target region identification, boundary adhesion, or missed segmentation when sarcopenia CT images have uneven grayscale distribution, blurred tissue boundaries, weakened local textures, and significant individual anatomical differences. This affects the stability and reliability of sarcopenia CT image segmentation results.
[0003] With the gradual application of deep learning technology in the field of medical image segmentation, automatic segmentation methods based on convolutional neural networks have improved the processing efficiency of sarcopenia CT images to some extent. However, there is still room for further improvement in existing methods. On the one hand, some methods focus more on directly using the network to perform end-to-end segmentation prediction of the input image. They are insufficient in mining the gray-level perturbation relationship, local tissue initial response, multi-scale heterogeneous response differences, and the traction relationship between deep and shallow semantics in sarcopenia CT images. This results in the model's limited ability to represent the attribution relationship of weak boundary regions, complex texture regions, and target regions. On the other hand, existing methods often lack collaborative modeling of the intrinsic correlation between features at different levels in the processes of boundary detail restoration, target region interpretation enhancement, and segmentation mask generation. As a result, they are prone to problems such as insufficient boundary correction ability of segmentation results, insufficient expression of local regions, and decreased segmentation accuracy in complex scenes.
[0004] This deep learning-based CT image segmentation method for sarcopenia first constructs a gray-level perturbation potential difference map and an initial tissue response map for sarcopenia CT images, and further forms multi-scale heterogeneous response features. Then, based on stable enhancement features, it constructs a response ranking difference map, semantic traction features, and target region representation. Finally, it combines a balanced reference tensor and boundary correction mapping to output the final segmentation mask. Thus, it achieves continuous processing from low-level gray-level perturbation modeling to mid-level semantic interpretation enhancement, and then to high-level boundary correction and mask generation within the same technical framework. Compared with existing technologies, it can more effectively improve the target region representation ability, enhance the segmentation stability under complex boundary conditions, and improve the accuracy and consistency of the final segmentation mask. Therefore, it is more suitable for automatic segmentation and quantitative analysis of sarcopenia CT images. Summary of the Invention
[0005] This invention provides a deep learning-based CT image segmentation method for sarcopenia. Addressing the issues of uneven grayscale distribution, blurred tissue boundaries, weak local texture response, unclear target region attribution, and insufficient boundary correction capability of the final segmentation mask in existing sarcopenia CT image segmentation processes, this invention proposes a deep learning-based sarcopenia CT image segmentation model. This model consists of a significant tissue response enhancement module, a target region expression generation module, and a decoding correction mask generation module. It solves the problems of insufficient mining of underlying grayscale perturbation relationships in sarcopenia CT images, weak deep semantic traction capability, inadequate target region expression, and low segmentation accuracy and stability under complex boundary conditions in existing technologies.
[0006] The technical solution adopted by the present invention to achieve the above objectives specifically includes the following steps: S1. Collect raw CT scan data of sarcopenia and construct a CT image dataset of sarcopenia. S2. Based on sarcopenia CT images, the local neighborhood averaging operator is used to generate the initial tissue response map. Multi-scale heterogeneous response features are constructed based on the initial tissue response map, and the multi-scale heterogeneous response features are adaptively weighted and aggregated to generate stable enhancement features. S3. Based on stable enhancement features, a response ranking difference map is obtained through convolution and spatial location sorting. Nonlinear activation is performed on the local tissue texture response, and semantic traction features are constructed through traction difference responses at different scales. The stable enhancement features and semantic traction features are weighted and labeled using the attribution tendency map, and the target region expression is output. S4. Construct a balanced reference tensor based on the target region representation, and generate decoding expansion features in combination with the target region representation. Generate boundary correction features based on the decoding expansion features, and output a segmentation mask. Steps S5, S2, S3, and S4 are combined to construct a deep learning-based CT image segmentation model for sarcopenia. S6. Input sarcopenia CT image: Deep learning-based sarcopenia CT image segmentation model; Output image: Segmentation mask.
[0007] Preferably, in step S1, the method for constructing the sarcopenia CT image dataset involves collecting raw sarcopenia CT scan data, performing quality screening to remove data with severe artifacts, missing slices, and substandard imaging quality, and then performing format standardization, target area annotation, and size normalization on the remaining raw sarcopenia CT scan data to generate the sarcopenia CT image dataset.
[0008] Preferably, in step S2, based on the sarcopenia CT image, a tissue initial response map is generated using a local neighborhood averaging operator, multi-scale heterogeneous response features are constructed based on the tissue initial response map, and adaptive weighted aggregation is performed on the multi-scale heterogeneous response features to generate stable enhancement features. Based on sarcopenia CT images, the gray values within the local neighborhood set surrounding the current pixel are aggregated using a local neighborhood averaging operator to construct a local gray-level baseline relationship between the target tissue region and the surrounding background region, generating a gray-level perturbation potential difference map. Based on the gray-level perturbation potential difference map, spatial constraints are constructed by combining the positional distance and spatial attenuation coefficient within the local neighborhood. The local gray-level perturbation information is enhanced by combining the potential difference enhancement index. Then, the local neighborhood averaging operator is used for aggregation processing to generate an initial tissue response map, thereby enhancing the identification ability of the target tissue region.
[0009] Preferably, based on the initial response map of the tissue, the initial response map of the tissue is aggregated at multiple scales using local neighborhood averaging operators at different scales to construct the local deviation relationship between the initial response map of the tissue at each scale and the corresponding local neighborhood averaging result. The local deviation relationship is then enhanced by combining the heterogeneous enhancement index and combined according to the feature dimension to generate multi-scale heterogeneous response features to characterize the response differences of the target tissue region at different scales.
[0010] Preferably, based on the adaptive folding and aggregation relationship of multi-scale heterogeneous response features, the response components corresponding to different scales in the multi-scale heterogeneous response features are adaptively weighted and aggregated to generate stable enhanced features, so as to improve the collaborative stability of multi-scale heterogeneous response features at different scales and suppress the interference of local abnormal fluctuations on feature expression.
[0011] Furthermore, the significant tissue response enhancement module first constructs a gray-level perturbation potential difference map and an initial tissue response map based on sarcopenia CT images to highlight the local gray-level perturbation information of the target tissue region relative to the surrounding background region. Then, it constructs multi-scale heterogeneous response features based on the initial tissue response map to characterize the local deviation characteristics of the target tissue region from different scales and enhance the ability to represent complex tissue structure changes and subtle gray-level differences. Finally, it generates stable enhancement features based on the multi-scale heterogeneous response features to improve the cooperative stability between response components at different scales and reduce the interference of local abnormal fluctuations on feature expression, thereby providing more stable and discriminative feature support for subsequent target region recognition and boundary segmentation.
[0012] Preferably, in step S3, based on the stable enhancement features, a response ranking difference map is obtained through convolution and spatial location ranking, and the local tissue texture response is nonlinearly activated and semantic traction features are constructed through the differential response under different receptive ranges. The stable enhancement features and semantic traction features are weighted and labeled using the attribution tendency map, and the target region expression is output. Based on the stable enhancement features, the channel response difference between the first and second ranked positions at each spatial location is extracted by one-dimensional convolution, and a response ranking difference map is constructed. Under the action of the response ranking difference map, the local tissue texture response in the stable enhancement features is nonlinearly activated to generate candidate explanatory features.
[0013] Preferably, the candidate paraphrasing features are convolved at different scales, and differential responses are constructed between them and the convolution kernel with a size of 1, to obtain traction differential responses at different scales. The traction differential responses characterize the degree of deviation of the candidate paraphrasing features from the responses at different scales, thereby enhancing the ability of semantic traction features to focus on the target region in sarcopenia CT images.
[0014] Preferably, the traction differential responses at different scales are combined according to the feature dimension to generate semantic traction features. Based on the semantic traction features, channel mapping and normalization constraint processing are performed to compress the attribution tendency information of each position in the semantic traction features into an attribution tendency map in the range of 0 to 1, which is used to characterize the degree of attribution allocation of the target region expression at different positions. The attribution tendency map is used to perform position-related weighting of the stable enhancement features and semantic traction features, so that the stable enhancement feature expression at the higher position of the attribution tendency map is strengthened and the semantic traction feature expression at the lower position of the attribution tendency map is preserved, thus completing the recalibration of the stable enhancement feature expression and generating the target region expression.
[0015] Furthermore, the target region expression generation module first constructs a response ranking difference map based on stable enhancement features and generates candidate interpretive features, which can improve the initial recognition ability of target region related textures in sarcopenia CT images. Secondly, it constructs semantic traction features through traction differential responses at different scales, which can enhance the focusing ability on regions with blurred boundaries and similar gray levels. Finally, it uses the attribution tendency map to weight and label the stable enhancement features and semantic traction features to generate target region expressions with clear regional attribution tendencies. In summary, the target region expression generation module can improve the accuracy and stability of target region expression in sarcopenia CT image segmentation.
[0016] Preferably, in step S4, a balanced reference tensor is constructed based on the target region representation, and a decoding expansion feature is generated in combination with the target region representation. Boundary correction features are generated based on the decoding expansion feature, and a segmentation mask is output. Based on the overall reference state of the target region expression in the channel dimension, the two-dimensional responses of the target region expression in each channel are averaged and converged. A balanced reference tensor with the same channel dimension as the target region expression is constructed through replication expansion processing to characterize the balanced reference state of the target region expression. Decoding control information is generated based on the deviation relationship between the target region expression and the balanced reference tensor. Learnable decoding is performed in combination with the target region expression to generate decoding unfolded features, thereby improving the ability of the decoding unfolded features to express the target main region and key transition regions in sarcopenia CT images.
[0017] Preferably, based on the boundary offset relationship between the target region representation and the decoded unfolded features, the target region representation and the decoded unfolded features are subjected to difference mapping processing, and a boundary correction mapping is constructed by combining the boundary reattachment strength parameter to generate boundary correction features.
[0018] Preferably, channel compression and normalization constraint processing are performed based on boundary correction features to construct a mask output mapping, and the output segmentation mask is used as the corresponding segmentation result of the sarcopenia CT image.
[0019] Furthermore, the decoding correction mask generation module first constructs a balanced reference tensor based on the target region representation and generates decoding unfolded features, which can improve the representation of the target main region and key transition regions in sarcopenia CT images. Based on the boundary offset relationship between the target region representation and the decoding unfolded features, it generates boundary correction features, which can enhance the ability of the segmentation results to fit the boundary of the target region. Based on the boundary correction features, it constructs a mask output mapping and outputs the final segmentation mask. In summary, the decoding correction mask generation module can improve the accuracy, boundary stability and overall reliability of sarcopenia CT image segmentation results.
[0020] Preferably, in step S5, steps S2, S3 and S4 are combined to construct a deep learning-based CT image segmentation model for sarcopenia. Step S2 corresponds to the significant tissue response enhancement module, step S3 corresponds to the target region expression generation module, and step S4 corresponds to the decoding correction mask generation module, which together construct a deep learning-based sarcopenia CT image segmentation model.
[0021] Furthermore, this invention addresses the characteristics of sarcopenia CT image datasets, such as similar gray levels between the target region and surrounding tissues, subtle local texture differences, unclear boundary transitions, and fluctuating image quality. It constructs a deep learning-based sarcopenia CT image segmentation model, consisting of a salient tissue response enhancement module, a target region expression generation module, and a decoding correction mask generation module. The salient tissue response enhancement module performs local gray-level perturbation analysis, multi-scale heterogeneous response construction, and stable enhancement processing on the sarcopenia CT images, outputting stable enhancement features. This provides more stable and discriminative features for subsequent target region identification. Based on the feature set, the target region expression generation module generates a target region expression based on stable enhanced features, which can enhance the model's ability to express the texture, region affiliation, and boundary neighboring regions of the target region, providing more focused semantic support for subsequent segmentation and decoding. The decoding correction mask generation module generates the final segmentation mask based on the target region expression, which can improve the ability of the segmentation result to fit the target main region and boundary region. Through the step-by-step connection of the above modules, the output of the previous module can provide a more stable, more focused, and more suitable input for segmentation and decoding for the next module, thereby improving the accuracy, boundary stability, and overall reliability of sarcopenia CT image segmentation. Attached Figure Description
[0022] Figure 1 This diagram illustrates the steps of a deep learning-based CT image segmentation method for sarcopenia.
[0023] Figure 2 This is a structural diagram of a deep learning-based CT image segmentation model for sarcopenia.
[0024] Figure 3 A structural diagram of the module for significantly enhancing organizational response.
[0025] Figure 4 Generate a module structure diagram to represent the target region.
[0026] Figure 5 This is a structural diagram of the decoding correction mask generation module.
[0027] Figure 6 Input the deep learning-based sarcopenia CT image segmentation model and the corresponding segmentation mask image into the sarcopenia CT image. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Please see the appendix Figure 1 - Appendix Figure 6 To address the issues of insufficient target region recognition and inadequate boundary fitting in sarcopenia CT images, where the target region and surrounding tissues exhibit similar gray levels, subtle local texture differences, and unclear boundary transitions, this invention provides a deep learning-based sarcopenia CT image segmentation method. First, a significant tissue response enhancement module enhances the local gray-level perturbation information and multi-scale heterogeneous response information of the sarcopenia CT image dataset, generating stable enhancement features. Second, a target region expression generation module performs semantic guidance and attribution recalibration on the stable enhancement features to generate a target region expression. Finally, a decoding and correction mask generation module decodes, unfolds, corrects boundaries, and outputs the mask to produce the final segmentation mask. The specific steps are as follows: Figure 1 As shown.
[0030] A deep learning-based CT image segmentation model for sarcopenia is jointly constructed by a significant tissue response enhancement module, a target region expression generation module, and a decoding correction mask generation module, as shown in the following structure. Figure 2 As shown.
[0031] S1. Collect raw CT scan data of sarcopenia and construct a CT image dataset of sarcopenia.
[0032] Furthermore, regarding the construction method of the sarcopenia CT image dataset, a total of 1248 raw sarcopenia CT scans were collected. After quality screening, 110 images with severe artifacts, missing slices, and substandard imaging quality were removed. The remaining 1138 raw sarcopenia CT scans were then processed for format standardization, target area annotation, and size normalization to construct a sarcopenia CT image dataset in PNG format with a uniform size of 428×322.
[0033] S2. Based on sarcopenia CT images, the local neighborhood averaging operator is used to generate the initial tissue response map. Multi-scale heterogeneous response features are constructed based on the initial tissue response map, and adaptive weighted aggregation is performed on the multi-scale heterogeneous response features to generate stable enhancement features.
[0034] Furthermore, in step S2, a significant tissue response enhancement module is constructed, and the structure diagram of the significant tissue response enhancement module is shown below. Figure 3 As shown, the specific steps are as follows: Based on sarcopenia CT images, a local neighborhood averaging operator is used to aggregate the gray values within the local neighborhood set surrounding the current pixel, generating a local neighborhood averaging result. The gray-level perturbation difference between the sarcopenia CT image and the local neighborhood averaging result is calculated to construct a gray-level shift relationship. This is then constrained using a stability constant to generate a gray-level perturbation potential difference map. The mathematical model is as follows: ; in, This represents a CT image of sarcopenia. This represents the grayscale perturbation potential difference diagram. Represents the stability constant, in this embodiment Values , The mathematical model for the local neighborhood averaging operator is as follows: ; in, This represents the local neighborhood averaging result calculated using the local neighborhood averaging operator on sarcopenia CT images. This represents a 3×3 local neighborhood set centered at the current pixel. This represents the number of pixels in the local neighborhood set. This indicates that the corresponding local neighborhood set of CT images of sarcopenia The grayscale value of the location; Based on the grayscale perturbation potential difference map, spatial constraints are applied to the map by combining the location distance and spatial attenuation coefficient within the local neighborhood. The map is then enhanced using a potential difference enhancement index, and aggregated through a local neighborhood averaging operator to generate the tissue initial response map. The mathematical model is as follows: ; in, This represents the organization's initial response diagram. This represents the positional distance within a local neighborhood, used to characterize the spatial proximity of each position within the local neighborhood relative to the current pixel. This represents the spatial attenuation coefficient, used to control the degree to which different locations within a local neighborhood affect the initial tissue response map. In this embodiment, it is set to a value of 2.0. This represents the potential enhancement index, used to increase the contribution of high-potential regions in the grayscale perturbation potential difference map to the initial response map of the tissue. In this embodiment, the value is set to 1.5. This represents the result after applying a potential enhancement index to the absolute value of the gray-scale perturbation potential difference, used to improve the response contribution in high-potential-difference regions. This indicates element-wise multiplication. This represents an exponential function.
[0035] Furthermore, based on the local deviation unfolding relationship of the initial response map at different scales, a multi-scale heterogeneous response feature is constructed, and the mathematical model is as follows: ; in, This represents the characteristics of multi-scale heterogeneous responses. This represents the organization's initial response diagram. This represents the heterogeneous enhancement index, designed to provide moderate enhancement in significantly deviated regions while maintaining the stability of the overall heterogeneous response. In this embodiment, the value is 0.5. This indicates combination based on feature dimensions. , and Let represent the local neighborhood averaging operators at the first, second, and third scales, respectively. Indicates the first The local neighborhood averaging operator at each scale has the following mathematical model: ; in, The organization's initial response map is shown in the first... The local neighborhood averaging result calculated using the local neighborhood averaging operator at each scale. , and These represent the local neighborhood averages calculated using the local neighborhood averaging operator at the first, second, and third scales, respectively, representing the initial response map of the organization. Indicates the first A local neighborhood set of scales. Indicates the first The number of pixels in a local neighborhood set at each scale. Indicates the neighborhood at this scale The initial response value of the organization at the location, The scale number is represented as 1, 2, or 3 in this embodiment.
[0036] Furthermore, based on the adaptive folding and aggregation relationship of multi-scale heterogeneous response features, stable enhancement features are generated. The mathematical model is as follows: ; in, Indicates stable enhancement features, In representing the multi-scale heterogeneous response characteristics, the first... The response components corresponding to each scale Indicates the scale traversal number. This indicates the first step in the scale accumulation process. The response components corresponding to each scale This represents the summation of the response components at each scale. Represents the stability constant, in this embodiment Values .
[0037] S3. Based on stable enhancement features, a response ranking difference map is obtained through convolution and spatial location sorting. The local tissue texture response is nonlinearly activated and semantic traction features are constructed through traction difference responses at different scales. The stable enhancement features and semantic traction features are weighted and labeled using the attribution tendency map to output the target region expression.
[0038] Furthermore, in step S3, a target region representation generation module is constructed, and the structure diagram of the target region representation generation module is shown below. Figure 4 As shown, the specific steps are as follows: Based on stable enhancement features, through Convolution is used to extract the channel response difference between the first and second ranked channels at each spatial location, constructing a response ranking difference map. The mathematical model is as follows: ; in, This represents the response ranking difference plot. Indicates stable enhancement features, Indicates the kernel size as The convolution is used to perform channel semantic projection on the stable enhancement features. and These represent the channel response values of the 1st and 2nd ranked positions at each spatial location, respectively. Under the influence of the response ranking difference map, nonlinear activation is applied to the local tissue texture response in the stable enhancement features to generate candidate interpretive features. The mathematical model is as follows: ; in, Indicates the characteristics of candidate definitions. Indicates the kernel size as The convolution is used to extract local tissue texture responses from stable enhancement features. Represents a non-linear activation function. This indicates that corresponding elements are multiplied.
[0039] Furthermore, convolutions are performed on candidate paraphrasing features at different scales, and differential responses are constructed between these convolutions and kernels of size 1. This yields traction differential responses at different scales. The traction differential responses characterize the degree of deviation of candidate paraphrasing features from their responses at different scales, thereby enhancing the semantic traction features' ability to focus on target regions in sarcopenia CT images. The mathematical model is as follows: ; in, Indicates the first Traction differential response at each scale Indicates the first Under each scale In this embodiment, convolution corresponds to three convolution extraction branches with different receptive ranges. The scale number is represented as 1, 2, or 3 in this embodiment. The traction differential responses at different scales are combined according to the feature dimension to generate semantic traction features. The mathematical model is as follows: ; in, Indicates semantic traction features, This indicates combination based on feature dimensions. These are represented as the traction differential responses at the first, second, and third scales, respectively. Indicates the kernel size as The convolution.
[0040] Furthermore, based on semantic traction features, channel mapping and normalization constraint processing are performed to compress the attribution tendency information at each position in the semantic traction features into an attribution tendency map within the range of 0 to 1. This map is used to characterize the degree of attribution assignment to the target region at different positions. The mathematical model is as follows: ; in, This represents a belonging tendency graph. The Sigmoid activation function is used to constrain the attribution trend map to the range of 0 to 1. Position-related weighting of stable enhancement features and semantic traction features is applied using the attribution trend map, strengthening the stable enhancement features at higher positions on the attribution trend map and preserving the semantic traction features at lower positions. This completes the relabeling of the stable enhancement features and generates the target region representation. The mathematical model is as follows: ; in, Indicates the target region representation. This indicates a stable enhancement feature.
[0041] S4. Construct a balanced reference tensor based on the target region representation, and generate decoding expansion features in combination with the target region representation. Generate boundary correction features based on the decoding expansion features, and output a segmentation mask.
[0042] Furthermore, in step S4, a decoding correction mask generation module is constructed. The structure diagram of the decoding correction mask generation module is shown below. Figure 5 As shown, the specific steps are as follows: Based on the overall reference state of the target region representation in the channel dimension, the two-dimensional responses of the target region representation in each channel are averaged and aggregated. A uniform reference tensor with the same channel dimension as the target region representation is constructed through replication and expansion. This tensor is used to characterize the uniform reference state of the target region representation, thus providing a unified reference for subsequently characterizing the deviation relationship of the target region representation. The mathematical model is as follows: ; in, Represents the equilibrium reference tensor. Indicates target region expression In the Two-dimensional response on each channel This indicates the number of channels representing the target region. Indicates the channel index. This represents the mean plot of the target region along the channel dimension. The copy expansion operation is used to expand the mean map to the same channel dimension as the target region representation. Decoding control information is generated based on the deviation between the target region representation and the equilibrium reference tensor. This information is then combined with the target region representation for learnable decoding to generate decoded unfolded features. The decoding control information adjusts the degree of decoded unfolding of the target region representation at different locations, enhancing locations with larger deviations from the equilibrium reference tensor during learnable decoding. This improves the expressive power of the decoded unfolded features for the target main region and key transition regions in sarcopenia CT images. The mathematical model is as follows: ; in, Indicates the decoding expansion features, express Activation function Represents a non-linear decoding mapping function. This indicates that the basic decoding result is output after passing through the non-linear decoding mapping function. Used to output decoding control information, when a certain position deviates significantly from the equalization reference tensor, that position is more likely to correspond to the target main body region or key transition region, thereby enhancing the unfolding strength of that position during the decoding process.
[0043] Furthermore, based on the boundary offset relationship between the target region representation and the decoded unfolded features, a difference mapping process is performed on the target region representation and the decoded unfolded features. A boundary correction mapping is then constructed by combining the boundary reattachment strength parameter to generate boundary correction features. The mathematical model is as follows: ; in, Indicates boundary correction features, Indicates the decoding expansion features, Indicates the target region representation. The boundary reattachment strength parameter represents the overall offset between the target region representation and the decoded unfolded features at the boundary location, thus adjusting the correction magnitude of the boundary correction mapping on the decoded unfolded features. Its mathematical model is as follows: ; in, and These represent the height and width of the feature map, respectively. and Indicates the position index. The target region is expressed in location. The gradient vector at that point, This indicates the salience intensity of the boundary at that location, used to characterize whether the location is near the boundary of the target area. The target region is expressed in location. The channel response vector at that location, This indicates the position of the decoded expanded features after channel mapping. The channel response vector at that location, This indicates the degree of shift in their attribution at that location. Represents the stability constant, with a value of This is used to avoid the denominator being zero. The boundary reattachment strength parameter is obtained by weighted aggregation of the salient strength of the boundary in the target region expression and the degree of attribution offset between the target region expression and the decoded unfolded features. This makes the degree of attribution offset at the location with higher salient strength of the boundary occupy a higher weight in the boundary correction mapping, thereby enhancing the boundary correction effect when the decoded unfolded features show boundary expansion, boundary contraction or local blurring, and reducing the interference of non-boundary locations on the boundary correction mapping, thereby improving the edge fitting ability and regional stability of the boundary correction features.
[0044] Furthermore, based on boundary correction features, channel compression and normalization constraint processing are performed to construct a mask output mapping, outputting the final segmentation mask. The mathematical model is as follows: ; in, This represents the segmentation mask, used as the segmentation result for sarcopenia CT images. express An activation function is used to constrain the output to the range of 0-1.
[0045] Steps S5, S2, S3, and S4 are combined to construct a deep learning-based CT image segmentation model for sarcopenia.
[0046] Furthermore, in step S5, step S2 corresponds to the significant tissue response enhancement module, step S3 corresponds to the target region expression generation module, and step S4 corresponds to the decoding correction mask generation module. These three modules jointly construct a deep learning-based CT image segmentation model for sarcopenia. The structure diagram of the deep learning-based CT image segmentation model for sarcopenia is shown below. Figure 2 As shown, the specific steps are as follows: Sarcopenia CT images output stable enhancement features through a significant tissue response enhancement module. The mathematical model is as follows: ; in, Indicates stable enhancement features, This indicates a significant organizational response enhancement module. This represents a CT image of sarcopenia.
[0047] Furthermore, the stable enhancement features output the target region representation through the target region representation generation module. The mathematical model is as follows: ; in, Indicates the target region representation. This indicates the target region representation generation module.
[0048] Furthermore, the target region representation is used to output a segmentation mask through a decoding correction mask generation module. The mathematical model is as follows: ; in, Indicates a segmentation mask. This indicates the decoding correction mask generation module.
[0049] S6. Input sarcopenia CT image: Deep learning-based sarcopenia CT image segmentation model; Output image: Segmentation mask.
[0050] Furthermore, as shown in the appendix Figure 1 The deep learning-based CT image segmentation model for sarcopenia described in S6 takes a sarcopenia CT image as input, as shown in the attached image. Figure 6 As shown in (a), the output is the segmentation mask for the corresponding image, as shown in the appendix. Figure 6 As shown in (b).
[0051] Furthermore, the deep learning-based sarcopenia CT image segmentation model uses the CentOS operating system, Python 3.9.2 as the language, Jetson Xavier as the processor, OpenCV as the image processing library, 64G of physical memory on the server, SGD as the model training optimizer, and an initial learning rate of 0.001.
Claims
1. A deep learning-based sarcopenia CT image segmentation method, characterized in that, Includes the following steps: S1. Collect raw CT scan data of sarcopenia and construct a CT image dataset of sarcopenia. S2. Based on sarcopenia CT images, the local neighborhood averaging operator is used to generate the initial tissue response map. Multi-scale heterogeneous response features are constructed based on the initial tissue response map, and the multi-scale heterogeneous response features are adaptively weighted and aggregated to generate stable enhancement features. S3. Based on stable enhancement features, a response ranking difference map is obtained through convolution and spatial location sorting. Nonlinear activation is performed on the local tissue texture response, and semantic traction features are constructed through traction difference responses at different scales. The stable enhancement features and semantic traction features are weighted and labeled using the attribution tendency map, and the target region expression is output. S4. Construct a balanced reference tensor based on the target region representation, and generate decoding expansion features in combination with the target region representation. Generate boundary correction features based on the decoding expansion features, and output a segmentation mask. Steps S5, S2, S3, and S4 are combined to construct a deep learning-based CT image segmentation model for sarcopenia. S6. Input sarcopenia CT image: Deep learning-based sarcopenia CT image segmentation model; Output image: Segmentation mask. 2.The deep learning-based sarcopenia CT image segmentation method of claim 1, wherein, Based on sarcopenia CT images, the gray values within the local neighborhood set surrounding the current pixel are aggregated using a local neighborhood averaging operator to generate a gray-level perturbation potential difference map. Based on the gray-level perturbation potential difference map, spatial constraint relationships are constructed by combining the positional distance and spatial attenuation coefficient within the local neighborhood. The potential response in the gray-level perturbation potential difference map is enhanced by combining the potential difference enhancement index. Finally, the local neighborhood averaging operator is used to aggregate the data to generate an initial tissue response map. 3.The deep learning-based sarcopenia CT image segmentation method of claim 2, wherein, Based on the initial response map of the tissue, the initial response map of the tissue is aggregated at multiple scales using local neighborhood averaging operators at different scales. The local deviation relationship between the initial response map of the tissue at each scale and the corresponding local neighborhood averaging result is constructed. The local deviation relationship is then enhanced by combining the heterogeneous enhancement index and combined according to the feature dimension to generate multi-scale heterogeneous response features. 4.The deep learning-based sarcopenia CT image segmentation method of claim 3, wherein, Based on the adaptive folding and aggregation relationship of multi-scale heterogeneous response features, the response components corresponding to different scales in the multi-scale heterogeneous response features are adaptively weighted and aggregated to generate stable enhanced features. 5.The deep learning-based sarcopenia CT image segmentation method of claim 4, wherein, Based on the stable enhancement features, the channel response difference between the first and second ranked positions at each spatial location is extracted by one-dimensional convolution, and a response ranking difference map is constructed. Under the action of the response ranking difference map, the local tissue texture response in the stable enhancement features is nonlinearly activated to generate candidate explanatory features.
6. The deep learning-based CT image segmentation method for sarcopenia according to claim 5, characterized in that, The candidate paraphrasing features are convolved at different scales, and differential responses are constructed between them and the convolution kernel with a size of 1, resulting in traction differential responses at different scales.
7. The deep learning-based CT image segmentation method for sarcopenia according to claim 6, characterized in that, The traction differential responses at different scales are combined according to the feature dimension to generate semantic traction features. Based on the semantic traction features, channel mapping and normalization constraint processing are performed. The attribution tendency information of each position in the semantic traction features is compressed into an attribution tendency map in the range of 0 to 1. The attribution tendency map is used to perform position-related weighting and labeling of stable enhancement features and semantic traction features to generate target region representation.
8. The deep learning-based CT image segmentation method for sarcopenia according to claim 7, characterized in that, Based on the overall reference state of the target region representation in the channel dimension, the two-dimensional responses of the target region representation in each channel are averaged and aggregated. A balanced reference tensor with the same channel dimension as the target region representation is constructed through replication expansion processing. Decoding control information is generated based on the deviation relationship between the target region representation and the balanced reference tensor. Learnable decoding is performed in combination with the target region representation to generate decoding expansion features.
9. The deep learning-based CT image segmentation method for sarcopenia according to claim 8, characterized in that, Based on the boundary offset relationship between the target region representation and the decoded unfolded features, a difference mapping process is performed on the target region representation and the decoded unfolded features. Then, a boundary correction mapping is constructed by combining the boundary reattachment strength parameter to generate boundary correction features.
10. The deep learning-based CT image segmentation method for sarcopenia according to claim 9, characterized in that, Channel compression and normalization constraint processing are performed based on boundary correction features to construct a mask output mapping and output a segmentation mask.