Calcaneus fracture image segmentation method based on edge enhancement guidance and CNN-Transform network

The hybrid CNN-Transformer network with edge-enhanced guidance addresses the challenges of calcaneus fracture segmentation by improving edge detection, resulting in enhanced precision and accuracy.

CN120318507APending Publication Date: 2025-07-15NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308487.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, when processing images of calcaneal fractures, there are problems with poor image quality and low segmentation accuracy caused by irregularities in the fracture edges, especially in complex cases or emergency situations, with high risk of misdiagnosis or misdiagnosis.

Method used

Using a hybrid architecture based on edge enhancement boot and CNN-Transformer network, multi-scale features are extracted through convolution blocks, CF modules and KAN modules, and edge detection branches are introduced, combined with multi-strategy learning models to fusion images to generate the final segmentation result.

Benefits of technology

The segmentation accuracy of the calcaneal fracture image is significantly improved, the limitations of the existing methods are overcome, and the segmentation accuracy is improved under complex conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318507A_ABST
    Figure CN120318507A_ABST
Patent Text Reader

Abstract

The invention discloses a calcaneal fracture image segmentation method based on edge enhancement guidance and a CNN-Transform network, and the method comprises the steps: obtaining a first image set based on a collected calcaneal fracture image; constructing a first encoder of a CNN-Transform hybrid architecture, and extracting multi-scale features of images in the first image set; introducing an edge detection branch into the first encoder, and extracting fracture edge information of images in the first image set; performing layer-by-layer fusion on the multi-scale features and the fracture edge information to obtain image fusion features; and restoring the resolution of the images in the first image set and generating a final segmentation result. According to the method, the image segmentation task and the edge detection task are combined, and the limitation of an existing algorithm based on an encoder-decoder structure during calcaneal fracture image processing is effectively overcome. According to the invention, by introducing the edge detection task, the capability of capturing the fracture edge features is enhanced, so that the segmentation precision is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of medical image processing, and specifically relates to a calcaneal fracture image segmentation method based on edge enhancement guidance and CNN-Transformer network. Background Technique

[0002] Calcaneal fracture, as a common type of injury in orthopedic clinics, has seen a significant increase in its incidence in recent years with the rise of accidental injuries such as traffic accidents and falls from heights. The symptoms of calcaneal fracture are manifested as persistent pain and swelling in the affected heel, and severe cases may show subcutaneous ecchymosis, and some severe cases are unable to walk. Therefore, quickly and accurately determining whether there is a calcaneal fracture is crucial for timely formulating treatment plans and avoiding complications such as calcaneal deformity and walking pain. Traditional fracture diagnosis methods mainly rely on doctors' clinical experience and imaging examinations, such as X-ray films, CT scans, etc. Conventional X-ray photography is usually used for the initial assessment of calcaneal injuries, but it has the typical disadvantages of two-dimensional imaging. The modern assessment of calcaneal fractures largely relies on multi-detector computed tomography, which can better visualize and characterize fracture lines and fracture fragment displacement. However, these methods may have the risks of low judgment efficiency, misdiagnosis, or missed diagnosis when facing complex cases or emergencies.

[0003] With the development of deep learning, calcaneal fracture image segmentation faces many challenges. On the one hand, calcaneal fracture images usually contain a large amount of biometric information, and the image quality is affected by various factors. For example, low resolution or noise interference may lead to the loss of key details, and the differences in bone morphology among different patients may also affect the generalization ability of the model. These factors increase the complexity and diversity of calcaneal fracture images, bringing considerable difficulties to the segmentation task. On the other hand, calcaneal fractures are often accompanied by the displacement and comminution of fracture fragments, making the fracture edges irregular and difficult to define on the image. This irregularity not only increases the complexity of the segmentation task but may also cause errors in traditional segmentation methods when dealing with the edge regions. In addition, the texture and density changes in the fracture area may be similar to those of the surrounding healthy tissues, further increasing the difficulty of accurate segmentation.

[0004] To address these challenges, researchers have been continuously exploring new segmentation methods. Deep learning models based on the encoder-decoder structure have attracted much attention due to their powerful feature extraction and image reconstruction capabilities. Among them, CNN-based methods have made some progress in encoding more refined context representations, but still face problems such as limited receptive fields and connection localization. In contrast, Transformer-based models can capture global context information in images through the self-attention mechanism, but are prone to attention collapse in the medical image field due to data annotation quality issues and high image redundancy. The combination of the two also has this problem and cannot fully unleash the potential of the Transformer. Summary of the Invention

[0005] This application provides a calcaneal fracture image segmentation method based on edge enhancement guidance and a CNN-Transformer network to solve the above technical problems.

[0006] To solve the above technical problems, a technical solution adopted in this application is: a calcaneal fracture image segmentation method based on edge enhancement guidance and a CNN-Transformer network, including:

[0007] Based on the collected calcaneal fracture images, obtain the first image set;

[0008] Based on convolutional blocks, CF modules, and KAN modules, construct a first encoder with a CNN-Transformer hybrid architecture and extract multi-scale features of the images in the first image set;

[0009] Based on a gated convolutional layer, residual blocks, and a gating mechanism, introduce an edge detection branch in the first encoder to extract the fracture edge information of the images in the first image set;

[0010] Based on a multi-strategy learning model, layer-by-layer fuse the multi-scale features and the fracture edge information to obtain image fusion features;

[0011] Based on the image fusion features, restore the resolution of the images in the first image set and generate the final segmentation result.

[0012] Further, the method for constructing the first encoder with a CNN-Transformer hybrid architecture includes:

[0013] Based on a convolutional layer, a batch normalization layer, and a ReLU activation function, construct convolutional blocks;

[0014] Based on a pooling module, an attention module, and a convolutional feed-forward network, construct a CF module to obtain the CF module.

[0015] Reshape the output features into a patch sequence, introduce a CBR block and a residual, and construct a KAN module.

[0016] Furthermore, the method for constructing the CF module includes:

[0017] Based on formulas (1)-(4), construct the CF module; wherein, formulas (1)-(4) are:

[0018]

[0019] wherein, x i,j is a pixel; E q and E k are projection matrices; Q i,j and K i,j are intermediate variables; I i,j is the initial convolution kernel; learnable Gaussian distance map M and parameters θ, α; M i,j is the Gaussian distance map.

[0020] Furthermore, based on formula (5), obtain the KAN module, wherein, formula (5) is:

[0021] Z k = LN(Z k-1 + CBR(KAN(Z k-1 ))) (5);

[0022] wherein, Z k is the output feature map of the k-th layer, and K is set to 2; LN is layer normalization, and CBR is a convolutional layer, a batch normalization layer, and a ReLU activation function.

[0023] Furthermore, the edge detection branch includes a gated convolutional layer, a residual block, and a gating mechanism.

[0024] Furthermore, the method for constructing the edge detection branch includes:

[0025] Based on formula (6), obtain the attention map; wherein, formula (6) is:

[0026] α t = σ(C 1×1 (s t || r t )) (6);

[0027] wherein, α t is the attention map, r t and s t represent the corresponding intermediate representations of the segmentation task and the edge detection task processed by the gated convolutional layer; || represents the concatenation of feature maps;

[0028] Based on formula (7), obtain the gated convolutional layer; wherein, formula (7) is:

[0029]

[0030] Among them, is the gated convolutional layer at each pixel (i, j); is the attention map at each pixel (i, j); is the edge detection value at each pixel (i, j).

[0031] Furthermore, the method for layer-by-layer fusion of multi-scale features and fracture edge information includes:

[0032] Based on formula (8), obtain the first loss function for the edge detection task; where formula (8) is:

[0033]

[0034] Among them, BCE is the binary cross-entropy loss, N is the total number of samples, y i is the category to which the i-th sample belongs. Calcaneal fracture image segmentation is a binary classification, and p i is the predicted value of the i-th sample. This formula directly measures the similarity between two samples;

[0035] Based on formulas (9)-(10), obtain the second loss function for the segmentation task; where formulas (9)-(10) are:

[0036] BCE DiceLoss(p,t)=0.5·BCELoss(p,t)+Dice(p,t) (9);

[0037]

[0038] Among them, the weight of the BCE loss is 0.5, the weight of the Dice loss is 1, p i and t i are the predicted value and the target value of the i-th sample respectively. ∈ is a smoothing constant used to avoid the case of a zero denominator;

[0039] Perform weighted summation based on the first loss function and the second loss function to obtain the total loss function;

[0040] Based on the total loss function, obtain the image fusion features.

[0041] The beneficial effects of this application are as follows: This application combines the image segmentation task and the edge detection task, effectively overcoming the limitations of existing algorithms based on the encoder-decoder structure when processing calcaneal fracture images. Specifically, existing methods are easily affected by poor image quality (such as noise and low resolution) and the irregularity of fracture edges (such as fracture fragment displacement and comminuted fracture), resulting in relatively low segmentation accuracy. By introducing the edge detection task, this application enhances the ability to capture fracture edge features, thereby significantly improving the segmentation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 FIG. is a schematic flowchart of an embodiment of the calcaneal fracture image segmentation method based on edge enhancement guidance and CNN-Transformer network of this application;

[0043] Figure 2 FIG. is a schematic diagram of the network structure of an embodiment of the calcaneal fracture image segmentation method based on edge enhancement guidance and CNN-Transformer network of this application;

[0044] Figure 3 is Figure 1 a schematic flowchart of an embodiment of step S2 in

[0045] Figure 4 is Figure 2 a schematic diagram of the network structure of the CF module in an embodiment of step S2 in

[0046] Figure 5 is Figure 1 a schematic diagram of the network structure of the KAN module in an embodiment of step S2 in

[0047] Figure 6 FIG. is a bar chart of IoU and Dice scores in an embodiment of the calcaneal fracture image segmentation method based on edge enhancement guidance and CNN-Transformer network of this application;

[0048] Figure 7 FIG. is a comparison experiment diagram in an embodiment of the calcaneal fracture image segmentation method based on edge enhancement guidance and CNN-Transformer network of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the following provides a more detailed description of the present invention in conjunction with specific embodiments.

[0050] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the limitations of the specific embodiments disclosed in the following specification.

[0051] Refer to Figure 1-2 , Figure 1 which is a schematic flowchart of an embodiment of the calcaneal fracture image segmentation method based on edge enhancement guidance and CNN-Transformer network of the present application. The method includes:

[0052] Step S1. Based on the collected calcaneal fracture images, obtain the first image set.

[0053] Specifically, collect all fracture images in the medical system to form the first image set for image training.

[0054] Step S2. Based on the convolutional block, CF module, and KAN module, construct the first encoder with a CNN-Transformer hybrid architecture and extract the multi-scale features of the images in the first image set.

[0055] Specifically, the first encoder part uses a CNN-Transformer hybrid architecture, including a convolutional block, CF module, and KAN module. The CF module is a deep improved Transformer similar to CNN, which combines convolutional operations with a CNN-style self-attention mechanism to extract the global context features and local detail features of calcaneal fracture images. The KAN module is used to further enhance the feature extraction ability of the model.

[0056] Refer to Figure 3 , step S2 specifically includes:

[0057] Step S21. Based on the convolutional layer, batch normalization layer, and ReLU activation function, construct the convolutional block.

[0058] Specifically, each convolutional block of the first encoder consists of a convolutional layer, batch normalization layer, and ReLU activation function. The adopted kernel size is 3×3, the stride is 1, and the padding is 1. The convolutional block also integrates a max pooling layer with a size of 2×2.

[0059] Step S22. Based on the pooling module, attention module, and convolutional feed-forward network, construct the CF module to obtain the CF module.

[0060] Specifically, as Figure 4 shown, the CF module includes a pooling module, a CNN-style attention module, and a convolutional feed-forward network. The pooling module first captures local features with CBR, and then downsamples through max pooling with a kernel size of 2×2 and another group of CBR to maintain the resolution. The CF module processes long-range pixel dependencies through an adaptive convolutional kernel. For each pixel x i,j , use the learnable projection matrix E q and E kFuse the features of the 3×3 neighborhood to construct the basis of the convolutional kernel. Obtain two intermediate variables Q i,j and K i,j , as shown in formulas (1) and (2). Calculate the initial convolutional kernel I i,j through cosine similarity, as shown in formula (3). Introduce the learnable Gaussian distance map M and parameters θ, α to dynamically determine the size of the convolutional kernel. Multiply the convolutional kernel I i,j and the Gaussian distance map M i,j to obtain the final convolutional kernel A i,j . Establish adaptive long-range dependencies by multiplying A i,j with V, and finally use 1×1 convolution, batch normalization, and ReLU to integrate the features.

[0061]

[0062] Among them, x i,j is a pixel; E q and E k are projection matrices; Q i,j and K i,j are intermediate variables; I i,j is the initial convolutional kernel; the learnable Gaussian distance map M and parameters θ, α; M i,j is the Gaussian distance map.

[0063] Among them, the convolutional feed-forward network refines the features generated by the CNN-like attention module and consists of two CBR blocks. CFFN makes the CF module fully based on CNN, avoiding the confrontation between CNN and Transformer in the CNN-Transformer hybrid method during training.

[0064] Step S23. Reshape the output features into a patch sequence, introduce the CBR block and the residual, and construct the KAN module.

[0065] Specifically, the network structure of the KAN module is as Figure 5 shown. Reshape the output features of the CF module into a flattened 2D patch sequence and pass them into a series of KAN layers. After each KAN layer, the features pass through a CBR block. Additionally, a residual connection is used and the patch sequence is added as the residual. Then, layer normalization is applied and the output features are passed to the next block. The output of the k-th KAN block can be described as:

[0066] Z k = LN(Z k-1 + CBR(KAN(Z k-1 ))) (5).

[0067] Among them, Z kis the output feature map of the k-th layer, where K = 2 is set; LN is layer normalization, and CBR is a convolutional layer, a batch normalization layer, and a ReLU activation function.

[0068] Step S3. Based on the gated convolutional layer, the residual block, and the gating mechanism, an edge detection branch is introduced into the first encoder to extract the fracture edge information of the images in the first image set.

[0069] Specifically, an edge detection branch is introduced into the first encoder. The edge detection branch includes a gated convolutional layer, a residual block, and a gating mechanism. The edge information of the fracture region is explicitly learned through the gated convolutional layer and the residual block, where the gating mechanism is used to control the interaction between the edge information and the main segmentation feature.

[0070] The edge detection branch takes the image gradient and the output of the second convolutional layer of the segmentation branch as inputs and produces the semantic boundary as the output. The edge detection branch is composed of a few residual blocks interleaved with gated convolutional layers. Three gated convolutional layers are used between the two branches and are connected to the third, fourth, and last layers of the segmentation branch. Let m denote the number of positions, and t ∈ {0, 1,..., m} be a running index, where r t and s t represent the intermediate representations of the corresponding segmentation task and edge detection task processed using the gated convolutional layer. By concatenating r t and s t an attention map α t is obtained, and then through a normalized 1×1 convolutional layer and a Sigmoid function σ, we get:

[0071] α t = σ(C 1×1 (s t ||r t )) (6);

[0072] where || denotes the concatenation of the feature maps. Given the attention map α t , the gated convolutional layer is applied to s t , ⊙ is the element-wise product of the attention map α, and then a residual connection and channel weighting are performed using the kernel w t . At each pixel (i, j), the gated convolutional layer is calculated as:

[0073]

[0074] where, is the gated convolutional layer at each pixel (i, j); is the attention map at each pixel (i, j); is the edge detection value at each pixel (i, j).

[0075] Step S4. Based on the multi-strategy learning model, layer-by-layer fusion of the multi-scale features and the fracture edge information is performed to obtain image fusion features.

[0076] Specifically, a first loss function for the edge detection task is defined. The first loss function is the BCE loss, and the specific formula is as follows:

[0077]

[0078] where N is the total number of samples, y i is the category to which the i-th sample belongs. The calcaneus fracture image segmentation is a binary classification, and p i is the predicted value of the i-th sample. This formula directly measures the similarity between two samples.

[0079] A second loss function for the segmentation task is defined. The second loss function is a weighted combination of the BCE loss and the Dice loss, where the weight of the BCE loss is 0.5 and the weight of the Dice loss is 1. The specific formula is as follows:

[0080] BCE DiceLoss(p,t) = 0.5·BCELoss(p,t) + Dice(p,t) (9);

[0081]

[0082] where the weight of the BCE loss is 0.5 and the weight of the Dice loss is 1, p i and t i are respectively the predicted value and the target value of the i-th sample, and ∈ is a smoothing constant used to avoid the case of a zero denominator;

[0083] where p i and t i are respectively the predicted value and the target value of the i-th sample, and ∈ is a smoothing constant used to avoid the case of a zero denominator.

[0084] Finally, the first loss value function of the segmentation task and the second loss function of the edge detection task are weighted and summed as the final total loss function to optimize the network parameters and obtain the image fusion features.

[0085] Step S5. Based on the image fusion features, the resolution of the images in the first image set is restored and the final segmentation result is generated.

[0086] Specifically, a multi-task learning strategy is adopted to layer-by-layer fuse the multi-scale features extracted by the first encoder and the fracture edge information extracted by the edge detection branch, and the fused features are transmitted to the decoder part through skip connections for gradually restoring the image resolution and generating the final segmentation result, so as to realize the calcaneal fracture image segmentation based on edge enhancement guidance and a CNN-Transformer hybrid network.

[0087] See Figure 6-7 , Figure 6 shows the bar charts of the IoU and Dice scores of the method provided by the present invention and seven other methods. It can be seen from the scores of IoU and Dice in the figure that the method proposed in this application has the best performance. Figure 7 shows the schematic diagram of the comparison results of the method provided by this application and seven other methods. Through comparison, it can be known that the predicted values obtained by the calcaneal fracture image segmentation method proposed in this application are closer to the true values, and the prediction accuracy is higher.

[0088] The above are only the embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied to other related technical fields, shall be similarly included in the patent protection scope of the present application.

Claims

1. A method for calcaneal fracture image segmentation based on edge enhancement guidance and CNN-Transformer network, characterized in that, The method includes the following steps: Based on the collected calcaneal fracture images, obtain a first image set; Based on a convolutional block, a CF module, and a KAN module, construct a first encoder with a CNN-Transformer hybrid architecture, and extract multi-scale features of the images in the first image set; Based on a gated convolutional layer, a residual block, and a gating mechanism, introduce an edge detection branch into the first encoder to extract fracture edge information of the images in the first image set; Based on a multi-strategy learning model, layer-by-layer fuse the multi-scale features and the fracture edge information to obtain image fusion features; Based on the image fusion features, restore the resolution of the images in the first image set and generate a final segmentation result.

2. The method according to claim 1, wherein The method for constructing the first encoder with a CNN-Transformer hybrid architecture includes: Based on a convolutional layer, a batch normalization layer, and a ReLU activation function, construct the convolutional block; Based on a pooling module, an attention module, and a convolutional feed-forward network, construct the CF module to obtain the CF module. Reshape the output features into a patch sequence, introduce a CBR block and a residual, and construct the KAN module.

3. The method according to claim 2, wherein The method for constructing the CF module includes: Based on formulas (1)-(4), construct the CF module; where the formulas (1)-(4) are: where x i,j is a pixel; E q and E k are projection matrices; Q i,j and K i,j are intermediate variables; I i,j is the initial convolution kernel; learnable Gaussian distance map M and parameters θ, α; M i,j is the Gaussian distance map.

4. The method according to claim 2, wherein Based on formula (5), obtain the KAN module, where the formula (5) is: Z k = LN(Z k-1 + CBR(KAN(Z k-1 ))) (5); Among them, Z k is the output feature map of the k-th layer, where K = 2 is set; LN is layer normalization, and CBR is a convolutional layer, a batch normalization layer, and a ReLU activation function.

5. The method according to claim 1, wherein The edge detection branch includes a gated convolutional layer, a residual block, and a gating mechanism.

6. The method according to claim 1, wherein The method for constructing the edge detection branch includes: Based on formula (6), obtain an attention map; where the formula (6) is: α t = σ(C 1×1 (s t ||r t )) (6); Among them, α t is the attention map, r t and s t represent the intermediate representations of the corresponding segmentation task and edge detection task processed by the gated convolutional layer; || represents the concatenation of feature maps; Based on formula (7), obtain the gated convolutional layer; where the formula (7) is: wherein, is the gated convolutional layer at each pixel (i, j); is the attention map for each pixel (i, j); is the edge detection value for each pixel (i, j).

7. The method according to claim 6, wherein The method for layer-by-layer fusing the multi-scale features and the fracture edge information includes: Based on formula (8), obtain the first loss function for the edge detection task; where the formula (8) is: Among them, BCE is the binary cross-entropy loss, N is the total number of samples, and y i is the category to which the i-th sample belongs. The calcaneal fracture image segmentation is a binary classification, and p i is the predicted value of the i-th sample. This formula directly measures the similarity between two samples; Based on formulas (9)-(10), obtain the second loss function for the segmentation task; where the formulas (9)-(10) are: BCE DiceLoss(p,t) = 0.5·BCELoss(p,t) + Dice(p,t) (9); Among them, the weight of the BCE loss is 0.5, the weight of the Dice loss is 1, p i and t i are the predicted value and the target value of the i-th sample respectively, and ∈ is a smoothing constant used to avoid the case of a zero denominator; Based on the first loss function and the second loss function, perform weighted summation to obtain a total loss function; Based on the total loss function, obtain the image fusion features.