A two-step method for processing spinal X-ray films based on single vertebral body segmentation
By using the YOLOv8 object detection model and the two-step processing method of the U-Net network in spinal X-ray images, the problems of low accuracy and insufficient accuracy of spinal segmentation in traditional methods are solved, and higher segmentation accuracy and recognition accuracy are achieved, supporting more effective diagnosis of spinal diseases.
Patent Information
- Application Number
- CN202410959702.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-17
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-07-17
AI Technical Summary
Traditional image segmentation methods have problems with low recognition accuracy and insufficient segmentation accuracy when segmenting spinal X-ray images. Especially because the vertebral blocks are similar to the background, and the same structure is covered by other tissues or are blurred, which increases the difficulty of segmentation and recognition.
A two-step spinal X-ray processing method based on single vertebral body segmentation was adopted. First, the YOLOv8 target detection model was used to identify the position of the spinal vertebral body and generate the target box. Then, the U-Net network was used for detailed segmentation in each target box, combining target detection and semantic segmentation to improve the accuracy of spinal segmentation and the accuracy of recognition.
This method significantly improves the accuracy and accuracy of spinal segmentation, can more accurately outline the spinal contour and identify key feature points of the vertebrae, and provide more effective information to support the intelligent diagnosis of spinal diseases.
Smart Images

Figure CN118918118B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical image processing, and particularly relates to a two-step spinal X-ray image processing method based on single vertebral body segmentation. Background Art
[0002] Accurately segmenting vertebral body blocks in full-spine X-ray images is crucial for the intelligent diagnosis of spinal diseases. Traditional image segmentation methods mainly include threshold-based segmentation methods, edge detection-based segmentation methods, region-based segmentation methods, etc. The vertebral body blocks in spinal X-ray images have a similar appearance to the background, and factors such as parts of the same structure in the image being covered by other tissues or being blurred make it more difficult for segmentation and recognition, posing a huge challenge to segmentation performance.
[0003] For spinal X-ray images with complex structures, traditional image segmentation methods often have problems such as low recognition accuracy and insufficient segmentation precision. Summary of the Invention
[0004] To solve the problems of low accuracy and insufficient segmentation precision in segmenting spinal X-ray images by traditional image segmentation methods, the present invention proposes a two-step spinal X-ray image processing method based on single vertebral body segmentation. The present invention mainly imports the target bounding boxes obtained by the YOLOv8 POSE model into the U-Net network for segmentation, combining object detection and semantic segmentation to improve the segmentation precision and recognition accuracy of the spine.
[0005] To achieve the above object, the present invention discloses a two-step spinal X-ray image processing method based on single vertebral body segmentation, including the following steps:
[0006] Step S1: Use the YOLOv8 object detection model to identify the position of each single spinal vertebral body in the X-ray image and generate single vertebral body target bounding boxes;
[0007] Step S2: Perform detailed segmentation within each detected single vertebral body target bounding box to accurately outline the spinal contour;
[0008] Step S3: Perform feature point recognition within the single vertebral body target bounding box, and use the YOLOv8 POSE model to identify the key feature points of the vertebrae to further improve the accuracy of feature point recognition.
[0009] As a further preferred technical solution of the above technical solution, step S1 is specifically implemented as the following steps:
[0010] Step S1.1: Perform image preprocessing, including operations such as normalization, rotation, and cropping on the X-ray image before inputting it into the YOLOv8 object detection model to meet the input requirements of the model;
[0011] Step S1.2: Feature extraction. Use the backbone network (CSPDarknet53) of the YOLOv8 object detection model to extract multi-scale features from the X-ray images. Use residual blocks to deepen the network and improve the effect of feature extraction, and introduce skip connections to solve the problems of gradient vanishing and gradient explosion;
[0012] Step S1.3: Feature fusion. Use FPN (Feature Pyramid Network) and PANet (Path Aggregation Network) to fuse features of different scales, improving the accuracy and robustness of object detection;
[0013] Step S1.4: Object bounding box prediction. Use the YOLO Head module of the YOLOv8 object detection model to predict object bounding boxes, class confidences, and location offsets on the feature map;
[0014] Step S1.5: Perform non-maximum suppression (NMS). Remove redundant bounding boxes through non-maximum suppression and only retain high-confidence bounding boxes to ensure that each object bounding box uniquely corresponds to one vertebra class.
[0015] As a further preferred technical solution of the above technical solution, for step S1.3:
[0016] Construct a top-down feature pyramid structure through FPN to extract multi-scale semantic information from feature maps at different levels to solve the problem of insufficient multi-scale information in object detection;
[0017] Introduce lateral connections and adaptive feature pooling mechanisms through PANet for information transfer and fusion between feature maps of different resolutions, thereby making better use of the multi-scale features generated by FPN.
[0018] As a further preferred technical solution of the above technical solution, for step S1.4:
[0019] Classification head: Predict the class label by applying a convolutional layer and an activation function (such as Softmax) on the feature map. The convolutional layer converts the information in the feature map into scores for each class, and the class with the highest score is the class label of the object within the object bounding box;
[0020] The Softmax activation function converts the raw scores (logits) into numerical values representing a probability distribution, making the probability value of each class between 0 and 1, and the sum of the probabilities of all classes equal to 1. The formula is as follows:
[0021]
[0022] Among them, \(e\) represents the base of the natural logarithm (Euler's number), \(n\) represents the number of categories, and \(z\) i is the original score of the \(i\)-th category;
[0023] Regression head: It is used to predict the coordinate parameters and confidence scores of each target box to determine the position and size of the target box. The coordinate parameters include the center coordinates, width, and height of the target box. The model uses a regression algorithm to predict these coordinate parameters to accurately locate the position of the target; the confidence score represents the confidence level of the model regarding whether the target is contained within the target box and is used to screen the detection results;
[0024] Through the combination of the above classification head and regression head, the category, position, and confidence information of the target are accurately predicted on the feature map.
[0025] As a further preferred technical solution of the above technical solution, step S1.5 is specifically implemented as the following steps:
[0026] Step S1.5.1: Perform confidence screening. Among the target boxes predicted by the model, first sort all the target boxes according to the confidence scores and retain the target box with the highest confidence.
[0027] Step S1.5.2: Calculate the overlap degree. For the remaining target boxes, calculate the overlap degree with the target box with the highest confidence through the Intersection over Union (IoU).
[0028] Step S1.5.3: Perform NMS operation. Remove the target boxes with an IoU higher than the set threshold from the remaining target boxes to ensure that redundant target boxes that overlap significantly with the target box with the highest confidence are not retained. Repeat this process until all target boxes are processed.
[0029] Step S1.5.4: Ensure uniqueness. After the NMS operation is completed, each retained target box should correspond to a unique target. If multiple target boxes highly overlap with the same target, NMS will select the target box with the highest confidence, and the other target boxes will be suppressed.
[0030] As a further preferred technical solution of the above technical solution, step S2 is specifically implemented as the following steps:
[0031] Step S2.1: Prepare data. Prepare (3000 cases) of single spine vertebral X-ray image data that has been classified, and standardize the data into images suitable for model input through operations including rotation and cropping.
[0032] Step S2.2: Perform data preprocessing on the extracted target boxes, including operations such as scaling, cropping, and normalization to meet the input requirements of the U-Net network;
[0033] Step S2.3: Construct the U-Net network architecture. This U-Net network takes an image of 1024×512×1 as input. Each encoder layer includes two 3×3 convolutional layers, followed by instance normalization, ReLU activation function, and 2×2 max pooling. Dropout is applied to each encoder-decoder stage of the network. Side outputs are generated at 128, 256, and 512 resolutions in the decoder, and then the final output is generated at the original resolution;
[0034] Step S2.4: Train the network. Use the extracted target boxes as training data to train the U-Net network. During the training process, optimize the network parameters through the backpropagation algorithm to enable the network to accurately learn the contour and structure of the vertebral body;
[0035] Step S2.5: Perform segmentation prediction. Input the extracted target boxes into the U-Net network to obtain pixel-level labels of the vertebral body contour;
[0036] Step S2.6: Post-process the segmentation results. Apply the region growing algorithm to fill the holes and apply Gaussian filtering to remove noise and smooth the edges to obtain a more accurate and smooth segmentation result.
[0037] As a further preferred technical solution of the above technical solution, Step S3 is specifically implemented as the following steps:
[0038] Step S3.1: Prepare data. Prepare (1000 cases) of single spinal vertebral X-ray image data with feature point markings, and standardize the data into images suitable for model input through operations including rotation, cropping, etc.;
[0039] Step S3.2: Perform data augmentation. Perform operations such as random rotation, scaling, and flipping on the training data to increase the diversity of the data, which helps the model better generalize to different vertebral postures and angles;
[0040] Step S3.3: Calculate the loss function. Ensure that the loss function effectively guides the model to learn the accurate feature point positions through the mean squared error (MSE) and the Euclidean distance of the key points;
[0041]
[0042] The mean squared error (MSE) is the sum of the squares of the distances between the target variable and the predicted values;
[0043]
[0044] Among them, ρ is the Euclidean distance between the point (x2, y2) and the point (x1, y1);
[0045] Step S3.4: Conduct model training. Use the pre-trained YOLOv8 POSE model as the basic model, perform machine learning to adapt to the vertebral key point detection task, and optimize the model parameters through the backpropagation algorithm to accurately identify the key feature points of the vertebrae;
[0046] Step S3.5: Apply the model. Apply the trained feature point recognition network optimized for a specific vertebral type to the spine segmentation data that has been classified.
[0047] To achieve the above object, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the two-step spine X-ray image processing method based on single vertebral body segmentation are implemented.
[0048] To achieve the above object, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the two-step spine X-ray image processing method based on single vertebral body segmentation are implemented.
[0049] Deep learning technology has enabled the rapid development of image segmentation methods. Compared with traditional image segmentation methods, the present invention has improved in both segmentation accuracy and segmentation precision, and can provide more effective information for clinicians. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of a two-step spine X-ray image processing method based on single vertebral body segmentation of the present invention.
[0051] Figure 2 is a schematic diagram of detecting the vertebral target box by the YOLOv8 object detection model of a two-step spine X-ray image processing method based on single vertebral body segmentation of the present invention.
[0052] Figure 3 is a schematic diagram of identifying the spine contour in the segmentation step of a two-step spine X-ray image processing method based on single vertebral body segmentation of the present invention.
[0053] Figure 4 is the actual operation result of feature point recognition of a two-step spine X-ray image processing method based on single vertebral body segmentation of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0054] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments in the following description are only examples, and those skilled in the art can think of other obvious variants. The basic principles of the present invention defined in the following description can be applied to other implementation schemes, variant schemes, improvement schemes, equivalent schemes, and other technical schemes without departing from the spirit and scope of the present invention.
[0055] In the preferred embodiment of the present invention, those skilled in the art should note that X-ray images and the like involved in the present invention can be regarded as prior art.
[0056] Preferred embodiment.
[0057] As Figure 1 shown, the present invention discloses a two-step spine X-ray image processing method based on single vertebral body segmentation, including the following steps:
[0058] Step S1: Use the YOLOv8 object detection model to identify the position of each single spine vertebral body in the X-ray image and generate a single vertebral body target box;
[0059] Step S2: Perform detailed segmentation within each detected single vertebral body target box to accurately outline the spine contour;
[0060] Step S3: Perform feature point recognition within the single vertebral body target box. Use the YOLOv8 POSE model to identify the key feature points of the vertebrae to further improve the accuracy of feature point recognition (the YOLOv8 POSE model is a variant based on the YOLOv8 model, specifically for pose estimation. The YOLOv8 model detects the position and category of an object, while the YOLOv8 POSE model further detects the pose of an object or the key points of a human body).
[0061] Specifically, step S1 is specifically implemented as the following steps:
[0062] Step S1.1: Perform image preprocessing. Before inputting the X-ray image into the YOLOv8 object detection model, perform preprocessing operations including normalization, rotation, and cropping to meet the input requirements of the model;
[0063] Step S1.2: Perform feature extraction. Use the backbone network (CSPDarknet53) of the YOLOv8 object detection model to extract multi-scale features from the X-ray image. Use residual blocks to deepen the network and improve the effect of feature extraction, and solve the problems of gradient disappearance and gradient explosion by introducing skip connections;
[0064] Step S1.3: Perform feature fusion. Use FPN (Feature Pyramid Network) and PANet (Path Aggregation Network) to fuse features of different scales, improving the accuracy and robustness of object detection;
[0065] Step S1.4: Perform object bounding box prediction. Use the YOLO Head module of the YOLOv8 object detection model to predict object bounding boxes, class confidences, and location offsets on the feature map;
[0066] Step S1.5: Perform non-maximum suppression (NMS). Remove redundant boxes through non-maximum suppression and only retain high-confidence boxes to ensure that each object bounding box uniquely corresponds to one vertebra class.
[0067] More specifically, for Step S1.3:
[0068] Construct a top-down feature pyramid structure through FPN, extract multi-scale semantic information from feature maps at different levels to solve the problem of insufficient multi-scale information in object detection;
[0069] Introduce lateral connections and adaptive feature pooling mechanisms through PANet for information transfer and fusion between feature maps of different resolutions, thereby (better) utilizing the multi-scale features generated by FPN.
[0070] Furthermore, for Step S1.4:
[0071] Classification head: Predict class labels by applying convolutional layers and activation functions (such as Softmax) on the feature map. The convolutional layer converts the information in the feature map into scores for each class, and the class with the highest score is the class label of the object within the object bounding box;
[0072] The Softmax activation function converts the raw scores (logits) into numerical values representing a probability distribution, such that the probability value of each class is between 0 and 1, and the sum of the probabilities of all classes is equal to 1. The formula is as follows:
[0073]
[0074] where e represents the base of the natural logarithm (Euler's number), n represents the number of classes, and z i is the raw score of the i-th class;
[0075] Regression head: It is used to predict the coordinate parameters and confidence scores of each target box to determine the position and size of the target box. The coordinate parameters include the center coordinates, width, and height of the target box. The model uses a regression algorithm to predict these coordinate parameters to accurately locate the position of the target. The confidence score represents the confidence level of the model regarding whether the target is contained within the target box and is used to filter the detection results.
[0076] Through the combination of the above classification head and regression head, the category, position, and confidence information of the target are accurately predicted on the feature map.
[0077] Furthermore, step S1.5 is specifically implemented as the following steps:
[0078] Step S1.5.1: Perform confidence filtering. Among the target boxes predicted by the model, first sort all the target boxes according to the confidence scores and retain the target box with the highest confidence.
[0079] Step S1.5.2: Calculate the overlap degree. For the remaining target boxes, calculate the overlap degree with the target box with the highest confidence through the Intersection over Union (IoU).
[0080] Step S1.5.3: Perform NMS operation. Remove the target boxes with IoU higher than the set threshold from the remaining target boxes to ensure that redundant target boxes with large overlaps with the target box with the highest confidence are not retained. Repeat this process until all target boxes are processed.
[0081] Step S1.5.4: Ensure uniqueness. After the NMS operation is completed, each retained target box should correspond to a unique target. If there are multiple target boxes with a high overlap with the same target, NMS will select the target box with the highest confidence, and the other target boxes will be suppressed.
[0082] Preferably, step S2 is specifically implemented as the following steps:
[0083] Step S2.1: Prepare data. Prepare (3000 cases) of classified single spine vertebral X-ray image data and standardize the data into images suitable for model input through operations including rotation and cropping.
[0084] Step S2.2: Perform data preprocessing. Perform data preprocessing on the extracted target boxes, including operations such as scaling, cropping, and normalization, to meet the input requirements of the U-Net network.
[0085] Step S2.3: Construct a U-Net network architecture. This U-Net network takes an image of 1024×512×1 as input. Each encoder layer includes two 3×3 convolutional layers, followed by instance normalization, ReLU activation function, and 2×2 max pooling. Dropout is applied to each encoder-decoder stage of the network. Side outputs are generated at the 128, 256, and 512 resolutions in the decoder, and then the final output is generated at the original resolution;
[0086] Step S2.4: Train the network. Use the extracted target boxes as training data to train the U-Net network. During the training process, optimize the network parameters through the backpropagation algorithm to enable the network to accurately learn the contour and structure of the vertebral body;
[0087] Step S2.5: Perform segmentation prediction. Input the extracted target boxes into the U-Net network to obtain pixel-level labels of the vertebral body contour;
[0088] Step S2.6: Post-process the segmentation results. Apply the region growing algorithm to fill the holes, and apply Gaussian filtering to remove noise and smooth the edges to obtain a more accurate and smooth segmentation result.
[0089] Preferably, Step S3 is specifically implemented as the following steps:
[0090] Step S3.1: Prepare data. Prepare (1000 cases) of single spinal vertebral X-ray image data with feature point markings, and standardize the data into images suitable for model input through operations including rotation and cropping;
[0091] Step S3.2: Perform data augmentation. Perform operations including random rotation, scaling, and flipping on the training data to increase the diversity of the data, which helps the model better generalize to different vertebral postures and angles;
[0092] Step S3.3: Calculate the loss function. Through the mean squared error (MSE) and the Euclidean distance of key points, ensure that the loss function effectively guides the model to learn the accurate positions of the feature points;
[0093]
[0094] The mean squared error (MSE) is the sum of the squares of the distances between the target variable and the predicted values;
[0095]
[0096] where ρ is the Euclidean distance between the point (x2, y2) and the point (x1, y1);
[0097] Step S3.4: Conduct model training. Use the pre-trained YOLOv8 POSE model as the base model, and perform machine learning to adapt to the vertebral key point detection task. Optimize the model parameters through the backpropagation algorithm to accurately identify the key feature points of the vertebrae;
[0098] Step S3.5: Apply the model. Apply the trained feature point recognition network optimized for a specific vertebral type to the spine segmentation data that has been classified.
[0099] Regarding the algorithm explanation in the present invention:
[0100] U-Net: U-Net is a classic convolutional neural network structure commonly used in medical image segmentation tasks. It has an encoder-decoder structure and can effectively capture the context information of the image. However, U-Net has difficulties in dealing with blurred boundaries and small object segmentation.
[0101] FCN (Fully Convolutional Network): FCN is a fully convolutional network that can perform pixel-level classification on the entire image. It can generate dense prediction results, but it will face problems of memory consumption and computational complexity when dealing with large-size images.
[0102] DeepLab: DeepLab is a segmentation algorithm based on dilated convolution, which captures broader context information by increasing the receptive field of the convolutional kernel. However, DeepLab has difficulties in dealing with fine structures and blurred boundaries.
[0103] Mask R-CNN: Mask R-CNN is an object detection and segmentation algorithm based on the Region Proposal Network. It can simultaneously achieve object detection and segmentation, but it will face problems of slow training and inference speed when dealing with large-scale datasets.
[0104] YOLO (You Only Look Once) is a deep learning-based object detection algorithm that simultaneously predicts the bounding boxes and classes of multiple objects in an image through a single forward pass. Its main characteristics are fast and accurate. However, since YOLO divides the image into grids for prediction, for small-size objects, the detection accuracy will be inaccurate.
[0105] The present invention also discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the two-step spine X-ray image processing method based on single vertebral body segmentation.
[0106] The present invention also discloses a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the double-step spinal X-ray image processing method based on single vertebral body segmentation are implemented.
[0107] It is worth mentioning that technical features such as X-ray images involved in this invention patent application should be regarded as prior art. For the specific structures, working principles, possible control methods, and spatial arrangement methods of these technical features, conventional selections in the art can be adopted, and they should not be regarded as the invention points of this invention patent. This invention patent will not be further specifically elaborated.
[0108] For those skilled in the art, it is still possible to modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A two-step spinal X-ray processing method based on single vertebra segmentation, characterized in that: The following steps are involved: Step S1: Use the YOLOv8 target detection model to identify the position of each single spinal vertebra in the X-ray image and generate a single vertebra target box; Step S1 is specifically implemented as follows: Step S1.1: Perform image preprocessing before inputting the X-ray image into the YOLOv8 target detection model, including normalization, rotation and cropping operations to adapt to the input requirements of the model; Step S1.2: Perform feature extraction, use the backbone network of the YOLOv8 target detection model to extract multi-scale features from the X-ray image, use residual blocks to deepen the network and improve the effect of feature extraction, and solve the gradient vanishing and gradient exploding problems by introducing jump connections; Step S1.3: Perform feature fusion, using FPN and PANet to fuse features of different scales to improve the accuracy and robustness of target detection; Step S1.4: Perform target box prediction, using the YOLO Head module of the YOLOv8 target detection model to predict the target box, category confidence, and position offset on the feature map; Step S1.5: Perform non-maximum suppression to remove redundant boxes and retain only high-confidence boxes to ensure that each target box uniquely corresponds to one vertebral category; For step S1.4: Classification head: The class label is predicted by applying convolutional layers and activation functions on the feature map. The convolutional layer converts the information in the feature map into a score for each class. The class with the highest score is the class label of the object in the target box. The Softmax activation function converts the original score into a numerical value representing the probability distribution, so that the probability value of each category is between 0 and 1, and the sum of the probabilities of all categories is equal to 1. The formula is as follows: Where e is the base of the natural logarithm, n is the number of categories, and z is i is the raw score of the ith category; Regression head: used to predict the coordinate parameters and confidence score of each target box to determine the position and size of the target box. The coordinate parameters include the center coordinates, width and height of the target box. The model predicts the coordinate parameters through the regression algorithm to accurately locate the position of the target. The confidence score indicates the model's confidence in whether the target box contains the target and is used to filter the detection results. The combination of the above classification head and regression head accurately predicts the category, location and confidence information of the target on the feature map; Step S2: Perform detailed segmentation within each detected single vertebral target frame to accurately outline the spine; Step S2 is specifically implemented as the following steps: Step S2.1: Perform data preparation to prepare the classified single spinal cone X-ray image data, and standardize the data into an image suitable for model input by including rotation and cropping operations; Step S2.2: Perform data preprocessing on the extracted target frame, including scaling, cropping and normalization operations to meet the input requirements of the U-Net network; Step S2.3: Construct a U-Net network architecture. This U-Net network takes a 1024×512×1 image as input. Each encoder layer consists of two 3×3 convolutional layers followed by instance normalization, ReLU activation function, and 2×2 maximum pooling. Dropout is applied to each encoder-decoder stage of the network. Side outputs are generated at 128, 256, and 512 resolutions of the decoder, and then the final output is generated at the original resolution. Step S2.4: Train the network, use the extracted target box as training data, and train the U-Net network. During the training process, optimize the network parameters through the back propagation algorithm so that the network can accurately learn the contour and structure of the vertebra; Step S2.5: Perform segmentation prediction and input the extracted target box into the U-Net network to obtain the pixel-level labeling of the vertebral contour; Step S2.6: Post-process the segmentation results, apply the region growing algorithm to fill the holes, and apply Gaussian filtering to remove noise and smooth the edges to obtain a more accurate and smooth segmentation result; Step S3: identify feature points within the single vertebral target frame, and use the YOLOv8 POSE model to identify key feature points of the vertebrae to further improve the accuracy of feature point recognition; Step S3 is specifically implemented as the following steps: Step S3.1: Perform data preparation, prepare single spinal cone X-ray image data with feature point labels, and standardize the data into an image suitable for model input by including rotation and cropping operations; Step S3.2: Perform data augmentation, including random rotation, scaling, and flipping of the training data to increase data diversity, which is used to help the model better generalize to different vertebral postures and angles; Step S3.3 calculates the loss function, and ensures that the loss function effectively guides the model to learn the accurate feature point position through the mean square error and the Euclidean distance of the key points; The mean square error is the sum of the squares of the distances between the target variable and the predicted value; Where ρ is the Euclidean distance between point (x2, y2) and point (x1, y1); Step S3.4: Perform model training, use the pre-trained YOLOv8 POSE model as the basic model, perform machine learning to adapt to the vertebral key point detection task, and optimize the model parameters through the back propagation algorithm to accurately identify the key feature points of the vertebrae; Step S3.5: Apply the model and apply the trained feature point recognition network optimized for a specific vertebral type to the classified spinal segmentation data.
2. A two-step spinal X-ray film processing method based on single vertebra segmentation according to claim 1, characterized in that: For step S1.3: A top-down feature pyramid structure is constructed through FPN to extract multi-scale semantic information from feature maps at different levels to solve the problem of insufficient multi-scale information in target detection. The lateral connection and adaptive feature pooling mechanism are introduced through PANet to transfer and fuse information between feature maps of different resolutions, thereby utilizing the multi-scale features generated by FPN.
3. A two-step spinal X-ray film processing method based on single vertebra segmentation according to claim 2, characterized in that: Step S1.5 is specifically implemented as follows: Step S1.5.1: Perform confidence screening. In the target box predicted by the model, first sort all the target boxes according to the confidence score and retain the target box with the highest confidence; Step S1.5.2: Calculate the overlap degree. For the remaining target frames, calculate the overlap degree with the target frame with the highest confidence by using the intersection-over-union ratio. Step S1.5.3: Perform NMS operation to remove target boxes with IoU higher than the set threshold from the remaining target boxes, ensuring that no redundant target boxes with large overlap with the target box with the highest confidence are retained, and repeat this process until all target boxes are processed; Step S1.5.4: Ensure uniqueness. After the NMS operation is completed, each retained target box corresponds to a unique target. If multiple target boxes highly overlap with the same target, NMS will select the target box with the highest confidence, and other target boxes will be suppressed.
Citation Information
Patent Citations
Lightweight YOLOv4 construction site safety helmet wearing detection method and device
CN113989852A
X-ray spine image key point detection and identification method
CN114463298A
Method for detecting target object grabbed by mechanical arm based on improved YOLOv5 algorithm
CN116630602A
Neural network and method for micronucleus recognition
CN118297119A