Method for Automatic Recognition of Thoracolumbar Key Points and Cobb Angle Calculation Based on Long- and Short-Range Features
The neural network model with short-range and long-range feature connections improves thoracic and lumbar vertebrae keypoint detection, addressing accuracy issues and enhancing Cobb angle estimation precision.
Patent Information
- Application Number
- CN202411508812.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-10-28
AI Technical Summary
The existing Cobb angle automatic estimation method based on key point detection. In the anterior and posterior spine X-ray image, the detection accuracy of vertebral key point is affected by the occlusion of tissues and external objects in the image, resulting in insufficient detection accuracy.
The automatic identification method of thoracolumbar key points based on long-term and short-range features is adopted. By constructing a network model, combining short-range and long-range feature connection modules, the connection between the corresponding position feature points is enhanced, and the central positioning, center comparison learning, center offset and angular positioning modules are used to distinguish between target vertebrae and non-target vertebrae, and reduce the impact of occlusion of foreign objects.
It improves the accuracy of identifying key points of the thoracic and lumbar spine, reduces the impact of external objects on detection, and provides more accurate results for vertebral position recognition.
Smart Images

Figure CN119445304B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method for automatically identifying thoracolumbar key points and calculating Cobb angles based on long-range and short-range features. Background Art
[0002] The judgment of scoliosis mainly measures the Cobb angle in the thoracolumbar segment of the anteroposterior X-ray image of the spine to quantify and diagnose spinal deformities. The Cobb angle is the angle between the upper endplate of the uppermost vertebra where the scoliosis begins and the lower endplate of the lowermost vertebra where the scoliosis ends in the anteroposterior X-ray image of the spine. Usually, a complete spine has three Cobb angles.
[0003] In recent years, compared with traditional machine learning algorithms, many automatic estimations of Cobb angles based on convolutional neural networks have shown great potential. The existing methods are mainly divided into two categories: 1) the regression-based method, 2) the method based on vertebral key point detection; the regression-based method maps the image to the Cobb angle by omitting the intermediate complex steps. Although the regression-based method can directly obtain three Cobb angles, it lacks the display of the basis for the results, so that it cannot be popularized in the clinical environment; compared with the regression-based method, the method based on vertebral key point detection calculates the Cobb angle by finding four key points of the vertebra or segmenting the vertebra in the image, so as to provide an intuitive basis for clinical decision-making.
[0004] However, the accuracy of the Cobb angle estimated by the method based on key point detection depends on the accuracy of vertebral key point detection. And vertebral key point detection is usually affected by the occlusion of tissues, organs and other foreign objects in the image. Therefore, the accurate identification of vertebrae is still a major problem in the automatic estimation of Cobb angles. For this reason, further improvement of the existing technology is needed. Summary of the Invention
[0005] The first technical problem to be solved by the present invention is directed to the above-mentioned existing technology, and provides an automatic identification method for thoracolumbar key points based on long-range and short-range features that can improve the detection accuracy.
[0006] The second technical problem to be solved by the present invention is to provide a method for calculating Cobb angles applying the above-mentioned automatic identification method for thoracolumbar key points based on long-range and short-range features.
[0007] The technical solution adopted by the present invention to solve the above first technical problem is: an automatic identification method for thoracolumbar key points based on long-range and short-range features, which is characterized by including the following steps:
[0008] Step 1: Obtain a certain number of anteroposterior X-ray images of the spine, and label the upper left, upper right, lower left, and lower right positions for each thoracic and lumbar vertebra in each X-ray image to form a sample set;
[0009] Step 2: Divide the sample set into a training set, a validation set, and a test set;
[0010] Step 3: Construct a network model. The constructed network model includes an encoder with d encoding layers, a decoder with d decoding layers, d - 1 vertebral position capture modules, and a processing unit, where d is a positive integer greater than 1;
[0011] The d encoding layers are connected in sequence, and the d decoding layers are also connected in sequence. The output end of the d-th encoding layer is connected to the input end of the 1st decoding layer; The 1st vertebral position capture module is arranged between the 2nd encoding layer and the 3rd encoding layer, the 2nd vertebral position capture module is arranged between the 3rd encoding layer and the 4th encoding layer,..., the d - 2nd vertebral position capture module is arranged between the d - 1st encoding layer and the d-th encoding layer, the d - 1st vertebral position capture module is arranged between the d-th encoding layer and the 1st decoding layer. The 1st vertebral position capture module is skip-connected to the d-th decoding layer, the 2nd vertebral position capture module is skip-connected to the d - 1st decoding layer,..., the d - 1st vertebral position capture module is skip-connected to the 2nd decoding layer;
[0012] The vertebral position capture module includes a short-range feature association module and a long-range feature association module. The specific processing process of the vertebral position capture module is as follows: The original feature map input to the vertebral position capture module is divided into tiles twice. One tile division is to evenly divide the original feature map into q first feature tiles, and the other tile division is to evenly divide the original feature map into q - 1 second feature tiles; The image sizes of the second feature tiles and the first feature tiles are the same, and q is a positive integer greater than 1;
[0013] Then, the short-range feature association module performs self-attention processing on each first feature tile and each second feature tile respectively to obtain the first feature map corresponding to each first feature tile and the second feature map corresponding to each second feature tile, and splice all the first feature maps in the order of the first feature tiles to obtain a first output feature map; In addition, splice all the second feature maps in the order of the second feature tiles to obtain a second output feature map. The image sizes of the first output feature map and the second output feature map are the same as that of the original feature map. Finally, fuse the first output feature map and the second output feature map to obtain the feature map output by the short-range feature association module;
[0014] In addition, the long-range feature connection module performs global average pooling on each first feature patch and each second feature patch respectively to obtain the third feature map corresponding to each first feature patch and the fourth feature map corresponding to each second feature patch; then, self-attention processing is performed on all the third feature maps to obtain the fifth feature map, and the fifth feature map is added to the original feature map to obtain the third output feature map; similarly, self-attention processing is performed on all the fourth feature maps to obtain the sixth feature map, and the obtained sixth feature map is added to the original feature patch to obtain the fourth output feature map; finally, the third output feature map and the fourth output feature map are fused to obtain the feature map output by the long-range feature connection module;
[0015] Adding the feature map output by the short-range feature connection module, the feature map output by the long-range feature connection module and the original feature map, the output of the vertebral position capture module is obtained;
[0016] The processing unit includes a center positioning module, a center contrast learning module, a center offset module and an angle positioning module, and the center positioning module, the center contrast learning module, the center offset module and the angle positioning module are connected to the output end of the d-th decoding layer;
[0017] Step 4: Input all the sample images in the training set into the network model constructed in Step 3 in batches for training, and use all the sample images in the validation set to verify the performance of the trained network model; after multiple trainings and validations, the optimal network model is selected;
[0018] Step 5: Input the test image in the test set into the optimal network model, and the key point positioning result of the test image is obtained; the key point positioning result of the test image includes the coordinates of the center point of each vertebral body predicted by the center offset module in the test image and the coordinates of the end points of each vertebral body predicted by the angle positioning module in the test image.
[0019] Preferably, before dividing the second feature patches, first leave a distance of h / 2q empty at the top and bottom of the original feature map, and then evenly divide the image in the original feature map except for the empty positions, so as to obtain q-1 second feature patches, where h is the height of the original feature map.
[0020] Specifically, the j-th feature patch in the feature map output by the short-range feature connection module The calculation formula is:
[0021]
[0022] Among them, is the j-th first feature map in the first output feature map, is the j-th first feature map in the second output feature map, and Mean(.) is mean processing.
[0023] Specifically, for the j-th feature map block in the feature map output by the long-range feature connection module The calculation formula is:
[0024]
[0025] Wherein, Is the j-th feature map in the third output feature map, Is the j-th feature map in the fourth output feature map.
[0026] Specifically, each time the network model is trained in step 4, the total loss function L all Is calculated, and the network parameters in the network model are updated backward according to the total loss function L all To obtain the network model after one training is completed;
[0027] L all The calculation formula is:
[0028] L all = L hm + L hc + L cenoff + L coroff
[0029] Wherein, L hm Is the first loss function calculated according to the output result of the center localization module, L hc Is the second loss function calculated according to the output result of the center contrast learning module, L cenoff Is the third loss function calculated according to the output result of the center offset module, L coroff Is the fourth loss function calculated according to the output result of the corner localization module.
[0030] Preferably, the specific calculation process of the center localization module is: by calculating the average value of the coordinates of the upper left, upper right, lower left, and lower right endpoints of each vertebral body in the original image, and using it as the center point coordinate of each vertebral body, then converting the center point of each vertebral body into a Gaussian heat map through Gaussian transformation, and finally using the Gaussian heat map to predict the center point coordinate of each vertebral body; the above original image is a training sample in the training set;
[0031] The calculation formula of the first loss function L hm Is:
[0032]
[0033] Wherein, N is the total number of pixel points on the Gaussian heat map, i is the pixel point position index on the Gaussian heat map, y i Is the pixel value at the i-th position on the original image, pi is the pixel value at the i-th position on the Gaussian heat map, and both α and β are hyperparameters.
[0034] Preferably, the specific calculation process of the central contrast learning module is as follows:
[0035] The output of the d-th decoding layer is passed through a feature extraction layer to obtain an output feature map, which has the same size as the Gaussian heat map. The positions of the pixel points in the Gaussian heat map with pixel values greater than T are selected and recorded as the first position set. The features at each position in the first position set in the output feature map are used as positive samples, and the features in the output feature map other than the positive samples are used as negative samples;
[0036] The second loss function L hc is calculated as follows:
[0037]
[0038] where P i is the set of positions of positive samples in the output feature map, A is the total number of positive samples in the output feature map, i is the position of all sample positions in the output feature map, E i is the output feature map, i + is the position of the positive sample in the output feature map, is the feature map corresponding to the positive sample in the output feature map, τ is a hyperparameter; N i is the set of positions of negative samples in the output feature map, i - is the position of the negative sample in the output feature map, is the feature map corresponding to the negative sample in the output feature map, and sim(.) is the cosine similarity.
[0039] Preferably, the central offset module predicts the coordinates of the center point of each vertebral body in the original image on the basis of the center point coordinates of each vertebral body predicted by the center positioning module;
[0040] The third loss function L cenoff is calculated as follows:
[0041]
[0042] where M is the total number of vertebral body center points, x′ k and y′ k are the abscissa and ordinate of the predicted center point of the k-th vertebral body in the original image, x k and y k are the theoretical abscissa and ordinate of the center point of the k-th vertebral body in the original image, n is the ratio of the size of the original image to the feature map restored by the decoder, is the floor function.
[0043] Preferably, the angle positioning module predicts the coordinates of the endpoints of each vertebral body in the original image on the basis of the center offset module predicting the coordinates of the center point of each vertebral body.
[0044] The fourth loss function L coroff has the following calculation formula:
[0045]
[0046] where P is the set of four endpoints of each vertebral body, x i ′ and y i ′ are the abscissa and ordinate of the predicted center point of the i-th vertebral body in the original image, x i and y i are the theoretical abscissa and ordinate of the center point of the i-th vertebral body in the original image, and x i,p and y i,p are the abscissa and ordinate of the predicted endpoint p of the i-th vertebral body in the original image.
[0047] The technical solution adopted by the present invention to solve the above second technical problem is: A Cobb angle calculation method, characterized in that: the automatic recognition method of thoracolumbar key points based on long and short range features as described above is applied.
[0048] Compared with the prior art, the advantages of the present invention are: by setting a short-range feature connection module and a long-range feature connection module in the vertebral body position capture module, the short-range feature connection module and the long-range feature connection module are used to strengthen the long and short range feature dependencies of the thoracolumbar key points, thereby strengthening the connection of the corresponding position feature points and improving the network recognition accuracy, reducing the influence of foreign object occlusion on the detection of thoracolumbar key points, and making the recognition of thoracolumbar key points more accurate; in addition, the target vertebral body and non-target vertebral body are further distinguished through the center positioning module, the center contrast learning module, the center offset module and the angle positioning module, thereby removing the interference to the detection of thoracic key points and improving the detection accuracy of thoracolumbar key points. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a schematic diagram of the architecture of the network model in the embodiment of the present invention;
[0050] Figure 2 is Figure 1 a schematic diagram of the architecture of the vertebral body position capture module in DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0052] The automatic recognition method of thoracolumbar key points based on long and short range features in this embodiment includes the following steps:
[0053] Step 1: Obtain a certain number of anteroposterior X-ray images of the spine, and label the upper left, upper right, lower left, and lower right positions of each thoracic and lumbar vertebra in each X-ray image to form a sample set;
[0054] In this embodiment, the obtained X-ray images are provided by two radiologists with at least 5 years of clinical experience, focusing on 17 vertebrae composed of thoracic and lumbar vertebrae. The Labelme tool is used to label the upper left, upper right, lower left, and lower right positions of each vertebra. Each X-ray image has 68 position markings and is not cropped, forming a sample set;
[0055] Step 2: Divide the sample set into a training set, a validation set, and a test set;
[0056] Step 3: Construct a network model. The constructed network model includes an encoder with d encoding layers, a decoder with d decoding layers, d - 1 vertebral position capture modules, and a processing unit, where d is a positive integer greater than 1; as Figure 1 shown, in this embodiment, d = 4, that is, the encoder includes 4 encoding layers, the decoder includes 4 decoding layers, and there are 3 vertebral position capture modules; the above encoding layers include a first feature extraction module and an upsampling module, and the decoding layers include a second feature extraction module and a downsampling module; the above encoding layers and decoding layers are all prior arts and will not be elaborated here;
[0057] The d encoding layers are connected in sequence, the d decoding layers are also connected in sequence, and the output end of the dth encoding layer is connected to the input end of the first decoding layer; the first vertebral position capture module is arranged between the second encoding layer and the third encoding layer, the second vertebral position capture module is arranged between the third encoding layer and the fourth encoding layer,..., the d - 2th vertebral position capture module is arranged between the d - 1th encoding layer and the dth encoding layer, the d - 1th vertebral position capture module is arranged between the dth encoding layer and the first decoding layer, the first vertebral position capture module is jump-connected to the dth decoding layer, the second vertebral position capture module is jump-connected to the d - 1th decoding layer,..., the d - 1th vertebral position capture module is jump-connected to the second decoding layer;
[0058] The vertebral position capture module includes a short-range feature connection module and a long-range feature connection module, as Figure 2As shown, the specific processing process of the vertebral body position capture module is as follows: The original feature map input to the vertebral body position capture module is divided into tiles twice. One tile division is to evenly divide the original feature map into q first feature tiles, and the other tile division is to evenly divide the original feature map into q - 1 second feature tiles; the image sizes of the second feature tiles and the first feature tiles are the same, and q is a positive integer greater than 1. In this embodiment, before dividing the second feature tiles, first leave a distance of h / 2q at the top and bottom of the original feature map, and then evenly divide the image of the original feature map except for the vacant positions, that is, q - 1 second feature tiles are obtained, where h is the height of the original feature map.
[0059] Then, the short-range feature connection module performs self-attention processing on each first feature tile and each second feature tile respectively to obtain the first feature map corresponding to each first feature tile and the second feature map corresponding to each second feature tile, and splices all the first feature maps in the order of the first feature tiles to obtain the first output feature map; in addition, splice all the second feature maps in the order of the second feature tiles to obtain the second output feature map. The first output feature map and the second output feature map have the same image size as the original feature map. Finally, fuse the first output feature map and the second output feature map to obtain the feature map output by the short-range feature connection module.
[0060] In addition, the long-range feature connection module performs global average pooling processing on each first feature tile and each second feature tile respectively to obtain the third feature map corresponding to each first feature tile and the fourth feature map corresponding to each second feature tile; then perform self-attention processing on all the third feature maps to obtain the fifth feature map, and add the fifth feature map to the original feature map to obtain the third output feature map; similarly, perform self-attention processing on all the fourth feature maps to obtain the sixth feature map, and add the obtained sixth feature map to the original feature tile to obtain the fourth output feature map; finally, fuse the third output feature map and the fourth output feature map to obtain the feature map output by the long-range feature connection module.
[0061] Add the feature map output by the short-range feature connection module, the feature map output by the long-range feature connection module, and the original feature map, that is, the output of the vertebral body position capture module is obtained.
[0062] In this embodiment, q = 8. As Figure 2 shown, the 8 first feature tiles are S1, S2, S3, S4, S5, S6, S7, and S8 respectively; the 7 second feature tiles are M1, M2, M3, M4, M5, M6, and M7 respectively; successively make k = 1, 2, 3... 8, and calculate according to to obtain is the first feature map corresponding to S1, is the first feature map corresponding to S2, is the first feature map corresponding to S8; Selfatt(.) is self-attention processing, which is a prior art and will not be elaborated here; similarly, let k = 1, 2, 3... 7 in sequence, and calculate according to to obtain is the second feature map corresponding to M1, is the second feature map corresponding to M2, is the second feature map corresponding to M7;
[0063] The first output feature map Cat(.) means concatenating all feature map blocks head to tail, and the second output feature map
[0064] The j-th feature map block in the feature map output by the short-range feature connection module The calculation formula is:
[0065]
[0066] where is the j-th first feature map in the first output feature map, is the j-th first feature map in the second output feature map, and Mean(.) is mean processing;
[0067] As Figure 2 shown, the fifth feature map S ia = Selfatt(S′1, S′2,,…, S′8), where S′1, S′2,,…, S′8 are the third feature maps obtained by performing global average pooling on S1, S2…S8 respectively; are the 8 feature images in the fifth feature map respectively; let k = 1, 2, 3…8 in sequence, and calculate according to
[0068] to obtain each feature image in the third output feature map
[0069] Similarly, the sixth feature map M ia = Selfatt(M′1, M′2,…, M′7), and M′1, M′2,…, M′7 are the fourth feature maps obtained by performing global average pooling on M1, M2, M3, M4, M5, M6 and M7 respectively, are the 7 feature images in the sixth feature map respectively, and let k = 1, 2, 3…7 in sequence, and calculate according to Calculate each feature image in the fourth output feature map
[0070] The j-th feature map block in the feature map output by the long-range feature connection module The calculation formula is as follows:
[0071]
[0072] Wherein, is the j-th feature map in the third output feature map, is the j-th feature map in the fourth output feature map;
[0073] The processing unit includes a center positioning module, a center contrast learning module, a center offset module, and a corner positioning module. The center positioning module, the center contrast learning module, the center offset module, and the corner positioning module are connected to the output end of the d-th decoding layer;
[0074] The specific calculation process of the center positioning module in this embodiment is as follows: By calculating the average value of the coordinates of the upper left, upper right, lower left, and lower right four endpoints of each vertebral body in the original image, and using it as the center point coordinate of each vertebral body, then transforming the center point of each vertebral body into a Gaussian heat map through Gaussian transformation, and finally using the Gaussian heat map to predict the center point coordinate of each vertebral body; The above original image is a training sample in the training set;
[0075] The calculation formula of the Gaussian heat map generated according to the center of the N-th vertebral body in this embodiment is:
[0076]
[0077] Wherein, (X,Y) is the pixel coordinate on the Gaussian heat map generated according to the center of the N-th vertebral body, (X N ,Y N ) is the center point coordinate of the N-th vertebral body, and a is the radius of the disk formed on the Gaussian heat map;
[0078] The specific process of using the Gaussian heat map to predict the center point coordinate of each vertebral body is as follows: Use max pooling of size b*b to perform non-maximum suppression on the Gaussian heat map, and select the point with the largest pixel value as the center prediction point of the current vertebral body; b is a positive integer, and in this embodiment, b = 3;
[0079] The specific calculation process of the center contrast learning module is:
[0080] The output of the d-th decoding layer passes through a feature extraction layer to obtain an output feature map, which has the same size as the Gaussian heat map. Filter out the pixel positions in the Gaussian heat map where the pixel values are greater than T, denoted as the first position set, and use the features at the positions in the first position set in the output feature map as positive samples, and use the features other than the positive samples in the output feature map as negative samples;
[0081] Based on the center coordinates of each vertebral body predicted by the center localization module, the center offset module predicts the coordinates of the center point of each vertebral body in the original image;
[0082] The specific calculation process for the center offset module in this embodiment to predict the coordinates of the center point of each vertebral body in the original image is as follows:
[0083] Assume that the coordinates of a certain pixel point in the original image are (x, y). After n times of downsampling, the mapped coordinates on the output feature map are Assume that the center offset of each vertebral body is (△x, △y). The calculation formulas for △x and △y are Adding the calculated center offset of each vertebral body to the coordinates (x′, y′) on the output feature map gives x″ = x′ + △x, y″ = y′ + △y. Scaling (x″, y″) by n times can obtain the position of the predicted center point of the vertebral body on the original image;
[0084] Based on the center coordinates of each vertebral body predicted by the center offset module, the corner localization module predicts the coordinates of the endpoints of each vertebral body in the original image;
[0085] The specific calculation process for the corner localization module in this embodiment to predict the coordinates of the endpoints of each vertebral body in the original image is as follows:
[0086] Calculate the offsets from the predicted center point of each vertebral body to the upper left, upper right, lower left, and lower right endpoints of the same vertebral body. (a k , b k ) represents the offset from the predicted center point of a certain vertebral body to the k-th endpoint of the vertebral body. Add the offsets at the four positions to the center coordinates of the vertebral body respectively to obtain x′ k = a k + x′, y′ k = b k + y′, (x′ k , y′ k ) is the coordinate of the k-th endpoint of the vertebral body on the feature map. Scaling by n times can obtain the positions of the predicted endpoints of each vertebral body on the original image;
[0087] Step 4: Input all the sample images in the training set into the network model constructed in Step 3 in batches for training, and use all the sample images in the validation set to verify the performance of the trained network model; after multiple trainings and validations, select the optimal network model;
[0088] When training the network model each time, calculate the total loss function L all , and update the network parameters in the network model backward according to the total loss function L all to obtain the network model after one training is completed;
[0089] L all The calculation formula of is:
[0090] L all = L hm + L hc + L cenoff + L coroff
[0091] Among them, L hm is the first loss function calculated according to the output result of the center localization module, L hC is the second loss function calculated according to the output result of the center contrast learning module, L cenoff is the third loss function calculated according to the output result of the center offset module, L coroff为 is the fourth loss function calculated according to the output result of the corner localization module;
[0092] In this embodiment, the calculation formula of the first loss function L hm is:
[0093]
[0094] Among them, N is the total number of pixel points on the Gaussian heatmap, i is the pixel point position index on the Gaussian heatmap, y i is the pixel value at the i-th position on the original image, p i is the pixel value at the i-th position on the Gaussian heatmap, and both α and β are hyperparameters;
[0095] The calculation formula of the second loss function L hc is:
[0096]
[0097] Among them, P i is the set of positions of positive samples in the output feature map, A is the total number of positive samples in the output feature map, i is all sample positions in the output feature map, E i is the output feature map, i + is the position of the positive sample in the output feature map, is the feature map corresponding to the positive samples in the output feature map, and τ is a hyperparameter; N i is the set of positions of the negative samples in the output feature map, and i - is the position of the negative sample in the output feature map, is the feature map corresponding to the negative samples in the output feature map, and sim(.) is the cosine similarity;
[0098] The third loss function L cenoff is calculated as follows:
[0099]
[0100] Among them, M is the total number of vertebral body center points, x′ k and y′ k are the abscissa and ordinate of the center point of the k-th vertebral body predicted in the original image, x k and y k are the theoretical abscissa and ordinate of the center point of the k-th vertebral body in the original image, and n is the ratio of the size of the original image to the feature map restored by the decoder, is the floor function;
[0101] The fourth loss function L coroff is calculated as follows:
[0102]
[0103] Among them, P is the set of four endpoints of each vertebral body, x i ′ and y i ′ are the abscissa and ordinate of the center point of the i-th vertebral body predicted in the original image, x i and y i are the theoretical abscissa and ordinate of the center point of the i-th vertebral body in the original image, x i,p and y i,p are the abscissa and ordinate of the endpoint P of the i-th vertebral body predicted in the original image;
[0104] Step 5: Input the test image in the test set into the optimal network model, and the key point localization result of the test image can be obtained; the key point localization result of the test image includes the coordinates of the center point of each vertebral body predicted by the center offset module in the test image and the coordinates of the endpoints of each vertebral body predicted by the angle localization module in the test image.
[0105] This embodiment also relates to a Cobb angle calculation method, which applies the above-mentioned automatic identification method of thoracolumbar key points based on long and short range features. Among them, the Cobb angle calculation is a prior art and will not be elaborated here.
[0106] This embodiment also relates to a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the method for automatically identifying thoracolumbar key points and calculating Cobb angle based on long- and short-range features as described above.
[0107] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An automatic recognition method for thoracolumbar key points based on long- and short-range features, characterized in that The method includes the following steps: Step 1: Obtain a certain number of anteroposterior X-ray images of the spine. Mark four positions, namely the upper left, upper right, lower left, and lower right positions, for each thoracic and lumbar vertebra in each X-ray image to form a sample set; Step 2: Divide the sample set into a training set, a validation set, and a test set; Step 3: Construct a network model. The constructed network model includes an encoder with d encoding layers, a decoder with d decoding layers, d - 1 vertebral position capture modules, and a processing unit, where d is a positive integer greater than 1; The d encoding layers are connected in sequence, and the d decoding layers are also connected in sequence. The output end of the d-th encoding layer is connected to the input end of the first decoding layer; the first vertebral position capture module is disposed between the second encoding layer and the third encoding layer, the second vertebral position capture module is disposed between the third encoding layer and the fourth encoding layer,..., the d - 2-th vertebral position capture module is disposed between the d - 1-th encoding layer and the d-th encoding layer, the d - 1-th vertebral position capture module is disposed between the d-th encoding layer and the first decoding layer. The first vertebral position capture module is jump-connected to the d-th decoding layer, the second vertebral position capture module is jump-connected to the d - 1-th decoding layer,..., the d - 1-th vertebral position capture module is jump-connected to the second decoding layer; The vertebral position capture module includes a short-range feature association module and a long-range feature association module. The specific processing process of the vertebral position capture module is as follows: The original feature map input to the vertebral position capture module is divided into tiles twice. One tile division is to evenly divide the original feature map into q first feature tiles, and the other tile division is to evenly divide the original feature map into q - 1 second feature tiles; the image sizes of the second feature tiles and the first feature tiles are the same, and q is a positive integer greater than 1; Before the second feature tile division, first leave a distance of h / 2q at the top and bottom of the original feature map, and then evenly divide the image of the original feature map except for the vacant positions to obtain q - 1 second feature tiles, where h is the height of the original feature map; Then, use the short-range feature association module to perform self-attention processing on each first feature tile and each second feature tile respectively to obtain the first feature map corresponding to each first feature tile and the second feature map corresponding to each second feature tile. Concatenate all the first feature maps in the order of the first feature tiles to obtain a first output feature map; in addition, concatenate all the second feature maps in the order of the second feature tiles to obtain a second output feature map. The image sizes of the first output feature map and the second output feature map are the same as that of the original feature map. Finally, fuse the first output feature map and the second output feature map to obtain the feature map output by the short-range feature association module; In addition, a long-range feature connection module is used to perform global average pooling on each first feature map block and each second feature map block respectively, obtaining a third feature map corresponding to each first feature map block and a fourth feature map corresponding to each second feature map block; then, self-attention processing is performed on all the third feature maps to obtain a fifth feature map, and the fifth feature map is added to the original feature map to obtain a third output feature map; similarly, self-attention processing is performed on all the fourth feature maps to obtain a sixth feature map, and the sixth feature map is added to the original feature map block to obtain a fourth output feature map; finally, the third output feature map and the fourth output feature map are fused to obtain the feature map output by the long-range feature connection module; Adding the feature map output by the short-range feature connection module, the feature map output by the long-range feature connection module, and the original feature map gives the output of the vertebral position capture module; The processing unit includes a center positioning module, a center contrast learning module, a center offset module, and a corner positioning module, and the center positioning module, the center contrast learning module, the center offset module, and the corner positioning module are all connected to the output end of the d-th decoding layer; Step 4: Input all the sample images in the training set into the network model constructed in Step 3 in batches for training, and use all the sample images in the validation set to verify the performance of the trained network model; after multiple trainings and validations, the optimal network model is selected; Step 5: Input the image to be tested in the test set into the optimal network model, and the key point positioning result of the image to be tested is obtained; the key point positioning result of the image to be tested includes the coordinates of the center point of each vertebral body predicted by the center offset module in the image to be tested and the coordinates of the end points of each vertebral body predicted by the corner positioning module in the image to be tested.
2. The automatic thoracic and lumbar key point recognition method according to claim 1, wherein: The j-th feature map block in the feature map output by the short-range feature connection module is calculated as follows: Among them, is the j-th first feature map in the first output feature map, is the j-th first feature map in the second output feature map, and Mean(.) is the mean processing.
3. The automatic recognition method of thoracolumbar key points according to claim 1, wherein: The j-th feature map block in the feature map output by the long-range feature connection module is calculated as follows: Among them, is the j-th feature map in the third output feature map, is the j-th feature map in the fourth output feature map.
4. The automatic identification method for thoracolumbar key points according to any one of claims 1 to 3, characterized in that: In each training of the network model in step 4, calculate the total loss function L all , and based on the total loss function L all backward update the network parameters in the network model to obtain the network model after one training is completed; L all The calculation formula is as follows: L all = L hm + L hc + L cenoff + L coroff Among them, L hm is the first loss function calculated according to the output result of the center positioning module, L hc is the second loss function calculated according to the output result of the center contrast learning module, L cenoff is the third loss function calculated according to the output result of the center offset module, L coroff is the fourth loss function calculated according to the output result of the corner positioning module.
5. The automatic recognition method of thoracolumbar key points according to claim 4, characterized in that: The specific calculation process of the center positioning module is as follows: by calculating the average value of the coordinates of the upper left, upper right, lower left, and lower right endpoints of each vertebral body in the original image and taking it as the center point coordinate of each vertebral body, then transforming the center point of each vertebral body through Gaussian transformation into a Gaussian heat map, and finally using the Gaussian heat map to predict the center point coordinate of each vertebral body; the above original image is a training sample in the training set; The first loss function L hm has the following calculation formula: Where N is the total number of pixels on the Gaussian heatmap, i is the pixel position index on the Gaussian heatmap, and y i is the pixel value at the i-th position on the original image, and p i is the pixel value at the i-th position on the Gaussian heatmap. Both α and β are hyperparameters.
6. The automatic thoracic and lumbar key point recognition method according to claim 5, wherein: The specific calculation process of the center contrast learning module is as follows: The output end of the d-th decoding layer is passed through a feature extraction layer to obtain an output feature map, which has the same size as the Gaussian heat map. The positions of the pixel points in the Gaussian heat map with pixel values greater than T are selected and recorded as the first position set, and the features at each position in the first position set in the output feature map are used as positive samples, and the features in the output feature map other than the positive samples are used as negative samples; The second loss function L hc is calculated as follows: Among them, P i is the set of positions of positive samples in the output feature map, A is the total number of positive samples in the output feature map, i is the position of all samples in the output feature map, E i is the output feature map, i + is the position of the positive sample in the output feature map, E i is the feature map corresponding to the positive sample in the output feature map, τ is a hyperparameter; N i is the set of positions of negative samples in the output feature map, i - is the position of the negative sample in the output feature map, is the feature map corresponding to the negative sample in the output feature map, sim(.) is the cosine similarity.
7. The automatic recognition method for thoracolumbar key points according to claim 5, wherein: The center offset module predicts the coordinates of the center point of each vertebral body in the original image based on the center point coordinates of each vertebral body predicted by the center positioning module; The third loss function L cenoff is calculated as follows: Among them, M is the total number of the center points of the vertebral bodies, x′ k and y′ k are the abscissa and ordinate of the center point of the k-th vertebral body predicted in the original image, x k and y k are the theoretical abscissa and theoretical ordinate of the center point of the k-th vertebral body in the original image, n is the ratio of the size of the original image to the feature map restored by the decoder, is the floor function.
8. The automatic recognition method of thoracolumbar key points according to claim 7, characterized in that: The corner positioning module predicts the coordinates of the end points of each vertebral body in the original image based on the center point coordinates of each vertebral body predicted by the center offset module; The fourth loss function L coroff is calculated by the following formula: Among them, P is the set of four endpoints of each vertebral body, x' i and y' i are the abscissa and ordinate of the center point of the i-th vertebral body obtained by prediction in the original image, x i and y i are the theoretical abscissa and ordinate of the center point of the i-th vertebral body in the original image, x i,p and y i,p are the abscissa and ordinate of the endpoint p of the i-th vertebral body obtained by prediction in the original image.
9. A Cobb angle calculation method, characterized in that: Apply the method for automatically identifying thoracolumbar key points based on long and short range features according to any one of claims 1 to 8 above.
Citation Information
Patent Citations
X-ray spine image key point detection and identification method
CN114463298A
Medical image segmentation method based on long and short distance features
CN114463341A