Method for detecting spine-shaped structure of vertebra in X-ray image

By introducing curve-perceived deformable convolution and shuffled channel attention mechanisms in the YOLOv8 framework, the existing model's insufficient accuracy in detecting vertebral spur-like structures is solved, and high-precision vertebral spur-like structure detection is achieved, suitable for edge platforms.

CN120339168APending Publication Date: 2025-07-18BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510257457.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When detecting vertebrae spur-like structures in X-ray images, existing deep learning models are prone to ignore local subtle features and are difficult to distinguish between targets and backgrounds, resulting in low detection accuracy.

Method used

The YOLOv8 framework is used to combine the CSP Darknet 53 backbone network, and the curve-sensing deformable convolution and shuffled channel attention mechanism is used to enhance feature extraction capabilities, realize multi-scale feature fusion through PAN structure, and build a vertebra spur-like structure detection network.

Benefits of technology

It improves the detection accuracy of vertebrae and spike-like structures, the model is lightweight and suitable for edge platform deployment, with high detection accuracy and low complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339168A_ABST
    Figure CN120339168A_ABST
Patent Text Reader

Abstract

The invention discloses a spine structure detection method for vertebrae in an X-ray image, and belongs to the field of computer vision and medical image processing. At present, a spinous structure of a model is generally not designed in a targeted manner, and local fine features of the spinous structure are very easy to ignore. And on the other hand, the target is influenced by surrounding tissues and artifacts, and the model needs to have enough capability to distinguish the target from the background. The invention provides a convolution operator used for capturing thorn-shaped structure features and a more effective channel attention mechanism, the convolution operator and the channel attention mechanism are embedded into a popular target detection framework, and the spine-shaped structure detection method for the vertebrae in the X-ray image is constructed. The method provided by the invention can accurately detect the vertebrae in the X-ray image and the spinous structure generated by the vertebrae, and provides powerful technical support for downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and medical image processing, and particularly relates to technologies such as computer image processing, deep learning, and object detection. Background Art

[0002] As people age, the human skeletal system gradually undergoes degenerative changes, which sometimes lead to small bone spurs at the edge of the vertebrae. In X-ray imaging examinations of the spine, it often appears as spiky structures around the vertebrae. Precise detection of these spiky structures can help people better understand their physical conditions and provide technical support for downstream tasks.

[0003] In recent years, deep learning has developed rapidly and made remarkable progress in various vision tasks. However, existing deep learning models still face many challenges in detecting spiky structures. Firstly, the models generally do not have a targeted design for spiky structures, and their local fine features are very easy to be ignored. On the other hand, the target is affected by surrounding tissues and artifacts, and the model needs to have sufficient ability to distinguish the target from the background.

[0004] To this end, the present invention proposes a convolution operator for capturing the features of spiky structures and a more effective channel attention mechanism, and embeds them into a popular object detection framework to construct a method for detecting spiky structures of vertebrae in X-ray images. The proposed method can accurately detect vertebrae and their generated spiky structures in X-ray images, providing strong technical support for downstream tasks. Summary of the Invention

[0005] The present invention proposes a method for detecting spiky structures of vertebrae based on X-ray images, which can automatically locate vertebrae one by one and detect whether they contain spiky structures.

[0006] To achieve the above objectives, the present invention proposes the following technical solutions:

[0007] Firstly, YOLOv8 is used as the basic framework, CSP Darknet 53 is used as the backbone network, and the PAN structure is utilized to combine image feature maps from different levels to achieve multi-scale feature fusion. Then, a curve-aware deformable convolution is proposed. By adding prior constraints to the deformable convolution, the sampling points of the convolution kernel can form a suitable curve to perceive spiky structures in the image, effectively enhancing the feature extraction ability. Finally, a shuffle channel attention mechanism is proposed to selectively weight the channels in the feature map, enhancing important channels and suppressing unimportant channels. Through the above methods, the feature extraction and expression capabilities of the model are improved, and accurate detection of vertebrae and spiky structures is achieved.

[0008] This solution includes three steps: the construction of a vertebral X-ray image dataset, the design of a vertebral spinous process detection network model, and the training of the network model. The implementation details of these steps are introduced in detail below.

[0009] Step 1: Construction of the vertebral X-ray image dataset

[0010] Currently, there is no publicly available vertebral spinous process detection dataset based on X-ray images. Therefore, the present invention cooperates with a domestic hospital to collect X-ray images of vertebral spinous processes, and then doctors annotate each image one by one. The present invention has constructed two datasets for the cervical vertebrae and lumbar vertebrae, and a number of target boxes containing vertebrae and target boxes containing spinous processes are annotated in all the X-ray images.

[0011] Step 2: Design of the vertebral spinous process detection network model

[0012] This model uses YOLOv8 as the basic framework, and then embeds the designed curve-aware deformable convolution and shuffle channel attention mechanism into it to construct a vertebral spinous process detection network model.

[0013] (1) Curve-aware deformable convolution

[0014] For the spinous processes in vertebrae, the standard rectangular convolution kernel cannot capture their features well. Therefore, the present invention first proposes a curve-aware deformable convolution. By adding prior constraints to the deformable convolution, the sampling points of the convolution kernel can form a suitable curve to perceive the spinous processes of vertebrae in X-ray images. The present invention embeds it into the C2f module of the model to better extract the features of spinous processes.

[0015] (2) Shuffle channel attention mechanism

[0016] The features extracted by deep models often contain multiple channels, and different channels contain different information, and not every channel is equally important. Therefore, channel attention mechanisms are often used in deep models to selectively weight the channels in the feature map. However, the classic channel attention uses a fully connected layer to generate attention weights, and the channel dimensionality reduction operation therein often causes information loss. Therefore, the present invention proposes a shuffle channel attention mechanism. By means of one-dimensional convolution and shuffling operations, the data loss caused by compressing channels is avoided, and the relative positions between different channels are shuffled, helping the model learn the correlations between features from a broader perspective. The present invention embeds it into the C2f module of the model to guide the model to focus on the channels containing important information, suppress the influence of irrelevant or redundant channels, and improve the model performance.

[0017] Step 3: Training of the network model

[0018] The vertebral X-ray image dataset constructed in Step 1 is divided into a training set and a test set, with a division ratio of 8:2. The loss function adopts a combination of BCE (Binary Cross-Entropy), CIOU (Complete Intersection Over Union), and DFL (Distribution Focal Loss) to optimize and supervise the training process of the network.

[0019] Through training, a model is obtained that can automatically locate the vertebrae in the X-ray image and detect the spiculate structures therein.

[0020] Compared with the prior art, the present invention has the following obvious advantages and practical application values:

[0021] The detection accuracy of the model is high

[0022] Compared with the mainstream object detection models, the proposed model can better learn the features of the vertebrae and the spiculate structures, and has a high detection accuracy.

[0023] The complexity of the model is low

[0024] The proposed convolution operator and channel attention do not introduce excessive parameters and computational amounts, ensuring the lightweight of the model and facilitating deployment on edge platforms. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 (a) is a partial image example of the cervical vertebra dataset.

[0026] Figure 1 (b) is a partial image example of the lumbar vertebra dataset.

[0027] Figure 2 is the structural diagram of the vertebral spiculate structure detection network model.

[0028] Figure 3 is the structural diagram of the shuffle channel attention mechanism. DETAILED DESCRIPTION OF THE INVENTION

[0029] Combined with the above method description and the drawings, the process of the present invention is further described in detail.

[0030] Step 1: Construction of the vertebral X-ray image dataset

[0031] The present invention cooperates with a domestic hospital to collect data of vertebral spiculate structure X-ray images, and finally constructs two datasets of cervical vertebrae and lumbar vertebrae. Then, the doctors of this hospital annotate the images one by one, and the annotation contents include the target boxes for locating the vertebrae and the target boxes for locating the spiculate structures. Partial image examples are as Figure 1 shown.

[0032] Step 2: Design of the vertebral spur-like structure detection network model

[0033] The structure diagram of the model is as shown Figure 2 in the figure. YOLOv8 is used as the basic framework. In sequence, the feature map will pass through the following modules: ConvModule_1, ConvModule_2, C2f_1, ConvModule_3, C2f_2, ConvModule_4, C2f_3, ConvModule_5, C2f_CDC1, SPPF, C2f_4, C2f_5, ConvModule_6, C2f_6, ConvModule_7, C2f_CDC2. The output head contains three groups, and each group contains two branches. The first group takes the feature map of a larger size as input. After passing through the first branch: ConvModule_8, ConvModule_9, Conv2d, the output of the target box information is obtained. After passing through the second branch: ConvModule_10, ConvModule_11, Conv2d, the output of the target class information is obtained. The second group takes the feature map of a medium size as input. After passing through the first branch: ConvModule_12, ConvModule_13, Conv2d, the output of the target box information is obtained. After passing through the second branch: ConvModule_14, ConvModule_15, Conv2d, the output of the target class information is obtained. The third group takes the feature map of a smaller size as input. After passing through the first branch: ConvModule_16, ConvModule_17, Conv2d, the output of the target box information is obtained. After passing through the second branch: ConvModule_18, ConvModule_19, Conv2d, the output of the target class information is obtained. The connection method of the model is introduced in detail below.

[0034] For the backbone network, it first uses a ConvModule for downsampling, then stacks 3 groups of combinations of ConvModule and C2f, 1 group of combinations of ConvModule and C2f_CDC, and finally uses an SPPF for spatial pyramid pooling. Subsequently, the PAN structure is used to combine the image feature maps from different levels to achieve multi-scale feature fusion. First, the output features of the last layer C2f_CDC1 in the backbone network are upsampled, then concatenated with the output features of the penultimate layer C2f_3. Next, the concatenated features pass through C2f_4, are upsampled and then concatenated with the output features of the second layer C2f_2, and features are extracted through the C2f_5 module. Next, it is downsampled through the ConvModule, then concatenated with the output features of C2f_4, and features are extracted through the C2f_6 module. Subsequently, it is downsampled through the ConvModule, concatenated with the output of the SPPF, and finally features are extracted through the C2f_CDC2 module. The output head consists of two groups, each group including two ConvModules and a single convolutional layer, which are connected to C2f_5, C2f_6, and C2f_CDC2 respectively. Thus, the construction of the entire network model is completed. The main module components of the model are introduced in detail below.

[0035] The main module components of the model include ConvModule, C2f, C2f_CDC, and SPPF. Among them, ConvModule is a basic convolutional module, consisting of 2D convolution, batch normalization (BN), and SiLU activation function. C2f is a feature extraction module. First, it is processed by a ConvModule, then by several Darknet Bottlenecks, and finally, the features of each module are concatenated and passed through another ConvModule to obtain the output. The Darknet Bottleneck in C2f contains two ConvModules, and whether to perform residual connection is determined according to the configuration. C2f_CDC is an improved feature extraction module. Each layer contains two CDC Darknet Bottlenecks. It replaces the standard convolution in the Darknet Bottleneck of C2f with the proposed Curve-aware Deformable Convolution (CDC) and adds the proposed Shuffle Channel Attention (SCA) at the end of the Darknet Bottleneck. SPPF is a spatial pyramid pooling module. First, it is processed by a ConvModule, then by three max pooling layers, and finally, the features of each module are concatenated and passed through another ConvModule to obtain the output. The parameters and configurations of each module are introduced in detail below.

[0036] Table 1 Sizes of input and output feature maps of each module of the model

[0037] Module Name Input Dimension Output Dimension ConvModule_1 640×640×3 320×320×64 ConvModule_2 320×320×64 160×160×128 C2f_1 160×160×128 160×160×128 ConvModule_3 160×160×128 80×80×256 C2f_2 80×80×256 80×80×256 ConvModule_4 80×80×256 40×40×512 C2f_3 40×40×512 40×40×512 ConvModule_5 40×40×512 20×20×512 C2f_CDC1 20×20×512 20×20×512 SPPF 20×20×512 20×20×512 C2f_4 40×40×1024 40×40×512 C2f_5 80×80×512 80×80×256 ConvModule_6 80×80×256 40×40×256 C2f_6 40×40×512 40×40×512 ConvModule_7 40×40×512 20×20×512 C2f_CDC2 20×20×1024 20×20×512 ConvModule_8 80×80×256 80×80×64 ConvModule_9 80×80×64 80×80×64 ConvModule_10 80×80×256 80×80×64 ConvModule_11 80×80×64 80×80×64 ConvModule_12 40×40×512 40×40×64 ConvModule_13 40×40×64 40×40×64 ConvModule_14 40×40×512 40×40×64 ConvModule_15 40×40×64 40×40×64 ConvModule_16 20×20×512 20×20×64 ConvModule_17 20×20×64 20×20×64 ConvModule_18 20×20×512 20×20×64 ConvModule_19 20×20×64 20×20×64

[0038] The input and output feature map sizes of each module are shown in Table 1. For the ConvModule module, numbers 1 to 7 are used for downsampling, with a convolutional kernel size of 3×3, a stride of 2, and a padding value of 1; numbers 8 to 19 are used as the detection head outputs, with a convolutional kernel size of 3×3, a stride of 1, and a padding value of 1. The last layer of the model detection head also includes a separate 2D convolution with a convolutional kernel size of 1×1, a stride of 1, and a padding value of 0; in the first and last layers of C2f and C2f_CDC, the convolutional kernel size is 1×1, the stride is 1, and the padding value is 0; in the Darknet Bottleneck, the convolutional kernel size is 3×3, the stride is 1, and the padding value is 1; in the SPPF, the convolutional kernel size is 1×1, the stride is 1, and the padding value is 0. For the C2f module, each layer contains several Darknet Bottlenecks, where numbers 1, 4, 5, and 6 contain 2 layers of Darknet Bottlenecks; numbers 2 and 3 contain 4 layers of Darknet Bottlenecks; numbers 1 to 3 use residual connections, and numbers 4 to 6 do not use residual connections. C2f_CDC is an improved feature extraction module, and each layer contains 2 layers of CDC Darknet Bottlenecks. The CDC Darknet Bottleneck in it replaces the standard convolution with the proposed curve-aware deformable convolution, with a convolutional kernel sampling point number of 9, a stride of 1, and a padding value of 0, and the proposed shuffle channel attention mechanism is added at the end. The proposed curve-aware deformable convolution and shuffle channel attention mechanism are introduced in detail below.

[0039] (1) Curve-aware deformable convolution

[0040] For a standard 2D convolution, the value at a certain position p0 on the output feature map y can be expressed as

[0041]

[0042] where x is the input feature map, k is the number of sampling points, w i represents the weight value of the i-th sampling point, which is continuously updated during the training iteration, and p i represents the predefined offset of the i-th sampling point relative to the center point p0, which is fixed in the standard 2D convolution. For example, for a standard 3×3 convolution, we have

[0043] p i ∈{(-1,-1),(-1,0),…,(0,1),(1,1)}#(2)

[0044] The deformable convolution introduces a learnable offset Δp i, where the positions of the sampling points are dynamic and can adapt to objects of different scales or shapes. At this time, the value at a certain position p0 on the output feature map can be expressed as

[0045]

[0046] where the offset Δp i =(Δp xi , Δp yi ), Δp xi is the abscissa of the offset, and Δp yi is the ordinate of the offset. In deformable convolution, the learning of these two components of Δp xi and Δp yi is independent and completely free without any constraints. This results in the loss of correlation between each sampling point and is not well applicable to the spiky structures of vertebrae in X-ray images.

[0047] The Taylor series can use the sum of several polynomials to approximate any smooth curve. Inspired by this, if the spiky structure in the image is regarded as a function curve, then the curve can be approximated by the weighted sum of polynomials. At this time, instead of directly learning the offset of each point, only the coefficients of each polynomial need to be learned to determine the shape of the curve. Subsequently, sampling at discrete points can quickly determine the offset of each point. At this time, the coordinates of the offset can be determined by formula (4):

[0048]

[0049] where Δp xi represents the abscissa of the offset Δp i of the i-th sampling point, and Δp yi is the corresponding ordinate, which can be expressed as

[0050]

[0051] where k n represents the coefficient of the n-th term in the Taylor polynomial. In the present invention, n = 4 is taken. Since the sum of the weighted polynomials can only roughly fit the curve, there is still a high-order error between the learned offset and the actual offset. Therefore, a free offset k i is added in the vertical direction to learn the error between the approximate solution and the actual solution and improve the accuracy.

[0052] On the other hand, different sampling points often have different degrees of importance. It is necessary to distinguish which sampling points the model focuses on and assign higher weights to them, otherwise assign lower weights. Therefore, the present invention further introduces a learnable coordinate weight m i, to distinguish the importance of different sampling points. At this time, y(p0) can be expressed as

[0053]

[0054] (2) Shuffle channel attention mechanism

[0055] The classic channel attention uses a fully connected layer to generate attention weights, and the channel dimension reduction operation often causes information loss. Therefore, the present invention uses a one-dimensional convolution to replace the fully connected layer to avoid data compression caused by the channel dimension reduction operation. However, the one-dimensional convolution only focuses on local channel information and cannot focus on global channel information. Therefore, a shuffle operation is further added to better enhance the interaction between channels. Its structure is as follows Figure 3 The algorithm is described in detail below.

[0056] First, the input features are shuffled to disrupt the relative positions of the features in the channel dimension. Then, a global average pooling operation is performed on each feature channel to obtain the global average of each channel. A one-dimensional convolution operation is then performed. The one-dimensional convolution operation is performed on the channel dimension to capture the interaction between channels. The step size is 1, and an adaptive convolution kernel size is used. The calculation formula can be expressed as:

[0057]

[0058] Among them, k is the convolution kernel size, C is the number of channels, γ and b are adjustment parameters, and the present invention takes γ = 2, b = 1. Finally, its output is passed through the Sigmoid activation function to generate attention weights, and these weights are used to weight the input feature map to obtain the output features. The proposed attention mechanism effectively improves the model's ability to represent features without introducing significant computational overhead. By introducing adaptive weights for each channel, the model pays more attention to important features and can suppress the influence of noise or irrelevant features, which helps to improve the performance and generalization ability of the model.

[0059] Step 3: Network model training

[0060] The present invention adopts a combination of BCE, CIOU and DFL as the loss function.

[0061] The BCE loss function is used to measure the difference between the category probability distribution predicted by the model and the true label. Its expression is:

[0062]

[0063] Where N is the total number of categories, y i is the true label of the i-th category, is the predicted probability of the i-th class.

[0064] The CIOU loss function is an improved IoU (Intersection over Union) loss function used to optimize the positioning accuracy of bounding boxes. By comprehensively considering the overlapping area, the distance between the center points, and the aspect ratio difference of the bounding boxes, it effectively improves the regression performance of object detection. Its expression is:

[0065]

[0066] where b represents the predicted target box by the model, b gt represents the true target box, IoU represents the intersection over union between the two target boxes, ρ represents the Euclidean distance between the center points of the predicted box and the true box, c represents the diagonal of the smallest enclosing rectangle containing the predicted box and the true box, α is the weight coefficient, v represents the shape difference between the predicted box and the true box, w represents the width of the predicted box, h represents the height of the predicted box, w gt represents the width of the true box, h gt represents the height of the true box.

[0067] The DFL loss function is a distribution loss function for bounding box regression. By modeling the distribution of discretized bounding box coordinates, it optimizes the prediction distribution, enabling the model to more accurately fit the true continuous target bounding box. Its expression is:

[0068]

[0069] where S i is the predicted distribution value at the i-th discrete position, S i+1 is the predicted distribution value at the (i + 1)-th discrete position, y is the value of the true continuous bounding box coordinate, y i is the discrete coordinate value at the i-th position after discretization, y i+1 is the discrete coordinate value at the (i + 1)-th position after discretization.

[0070] The total loss function can be expressed as:

[0071] Loss = L BCE + L CIoU + DFL#(13)

[0072] Using the above loss function to measure the difference between the model output and the label, AdamW is used as the optimizer to update the model parameters. The initial learning rate is set to 0.00125, and the momentum is 0.9. The training process lasts for 300 rounds, and the training result of the round with the highest mAP (mean Average Precision) value is taken as the final training result of the model. After the model is trained, when an X-ray image containing human vertebrae is input, the model can automatically locate the vertebrae and detect the spiky structures in them.

Claims

1. A method for detecting the spinous processes of vertebrae in X-ray images, characterized in that The following steps are involved: Step 1: Construction of vertebral X-ray image dataset Two datasets, cervical and lumbar spine, were constructed; Step 2: Design of the network model for vertebral spur detection The design of the vertebral spur detection network model incorporates the proposed curve-aware deformable convolution and shuffled channel attention mechanism; The proposed curve-aware deformable convolution and shuffled channel attention mechanism are introduced in detail below; (1) Curve-aware Deformable Convolution For a standard two-dimensional convolution, the value at a position p0 on the output feature map y is represented as where x is the input feature map, k is the number of sampling points, and w i represents the weight value of the i-th sampling point, which is continuously updated during the training iteration, and p i represents the predefined offset of the i-th sampling point relative to the center point p0; For a standard 3×3 convolution, we have p i ∈{(-1,-1),(-1,0),…,(0,1),(1,1)}#(2) Deformable convolution introduces a learnable offset Δp i , and the positions of its sampling points are dynamic; at this time, the value at a certain position p0 on the output feature map is represented as where the offset Δp i =(Δp xi , Δp yi ), Δp xi is the abscissa of the offset, and Δp yi is the ordinate of the offset; In deformable convolution, Δp xi and Δp yi The learning of these two components is independent and completely free; The thorny structure in the image is regarded as a function curve, and the weighted sum of polynomials is used to approximate this curve. At this time, there is no need to directly learn the offset of each point, and only the coefficients of each polynomial need to be learned to determine the shape of the curve. Then, the offset of each point can be determined by sampling at discrete points. At this time, the coordinates of the offset are determined by formula (4): where Δp xi represents the offset Δp at the i-th sampling point i of the abscissa, and Δp yi is the corresponding ordinate, expressed as where k n represents the coefficient of the n-th term in the Taylor polynomial, and n = 4. A free offset k is added in the vertical direction i to learn the error between the approximate solution and the actual solution; A learnable coordinate weight m is introduced into the proposed operator i to distinguish the importance of different sampling points; at this time, y(p0) is expressed as (2) Shuffle channel attention mechanism First, the input features are shuffled to disrupt the relative positions of the features in the channel dimension. Then, a global average pooling operation is performed on each feature channel to obtain the global average value of each channel. Then, a one-dimensional convolution operation is performed. The one-dimensional convolution operation is performed on the channel dimension to capture the interaction between channels. The step size is 1, and an adaptive convolution kernel size is used. The calculation formula is expressed as: Among them, k is the convolution kernel size, C is the number of channels, γ and b are adjustment parameters, where γ = 2, b = 1; finally, its output is passed through the Sigmoid activation function to generate attention weights, and these weights are used to weight the input feature map to obtain the output features; Step 3: Network model training A combination of BCE, CIOU and DFL is used as the loss function; The BCE loss function is used to measure the difference between the category probability distribution predicted by the model and the true label; its expression is: where N is the total number of categories, and y i is the true label of the i-th category, is the predicted probability of the i-th category; The CIOU loss function expression is: where b represents the predicted bounding box of the model, b gt represents the true bounding box, IoU represents the intersection over union between the two bounding boxes, ρ represents the Euclidean distance between the center points of the predicted box and the true box, c represents the diagonal of the smallest bounding rectangle containing the predicted box and the true box, α is the weight coefficient, v represents the shape difference between the predicted box and the true box, w represents the width of the predicted box, h represents the height of the predicted box, w gt represents the width of the true box, h gt represents the height of the true box; The DFL loss function expression is: where S i is the predicted distribution value at the i-th discrete position, and S i+1 is the predicted distribution value at the (i + 1)-th discrete position, y is the true continuous bounding box coordinate value, and y i is the discrete coordinate value at the i-th position after discretization, and y i+1 is the discrete coordinate value at the (i + 1)-th position after discretization; The total loss function is expressed as: Loss=L BCE +L CIoU + DFL#(13) The above loss function is used to measure the difference between the model output and the label. AdamW is used as the optimizer to update the model parameters. The initial learning rate is set to 0.00125 and the momentum size is 0.

9. The training process lasts for 300 rounds, and the training result with the highest mAP (meanAverage Precision) value is taken as the final training result of the model. After the model is trained, an X-ray image containing human vertebrae is input, and the model can automatically locate the vertebrae and detect the thorny structure therein.

2. The method according to claim 1, characterized in that, The design of the vertebral spur-like structure detection network model is as follows: The feature map will sequentially pass through the following modules: ConvModule_1, ConvModule_2, C2f_1, ConvModule_3, C2f_2, ConvModule_4, C2f_3, ConvModule_5, C2f_CDC1, SPPF, C2f_4, C2f_5, ConvModule_6, C2f_6, ConvModule_7, C2f_CDC2; the output head contains three groups, and each group contains two branches; the first group accepts the feature map with a larger size as input, and through the first branch: ConvModule_8, ConvModule_9, Conv2d, it gets the output of the target box information, and through the second branch: ConvModule_10, ConvModule_11, Conv2d, it gets the output of the target class information; the second group accepts the feature map with a medium size as input, and through the first branch: ConvModule_12, ConvModule_13, Conv2d, it gets the output of the target box information, and through the second branch: ConvModule_14, ConvModule_15, Conv2d, it gets the output of the target class information; the third group accepts the feature map with a smaller size as input, and through the first branch: ConvModule_16, ConvModule_17, Conv2d, it gets the output of the target box information, and through the second branch: ConvModule_18, ConvModule_19, Conv2d, it gets the output of the target class information.

Citation Information

Cited By

  • Vertebra staggered seam X-ray image intelligent identification method based on similar triangular structure labeling

    CN121962100A