A bone age assessment method based on fine-grained classification
Through the fine-grained classification method, the region of interest of left-hand wrist bone X-rays was selected using ResNet50 and feature pyramid network, which solved the problem of time-consuming and low interpretability of traditional bone age assessment and achieved high-precision bone age assessment.
Patent Information
- Application Number
- CN202210039032.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-01-13
AI Technical Summary
Traditional bone age assessment methods require expertise and are time-consuming, and deep learning-based methods cannot determine which parts of the X-ray play a key role in bone age assessment and are less interpretable.
The bone age evaluation method based on fine-grained classification is adopted, and ResNet50 is used as a feature extractor, combining the feature pyramid network and the path aggregation network to select the region of interest, and optimizing the selection of the region of interest through the guidance subnet, and using the evaluation subnet for accurate bone age evaluation.
A high-precision bone age assessment of left-hand wrist bone X-rays is achieved, and the areas of interest with the maximum amount of information can be accurately selected, and the evaluation results are accurate to months, suitable for adolescents aged 0 to 18.
Smart Images

Figure CN114742745B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a bone age assessment method based on fine-grained classification. Background Art
[0002] Bone age assessment usually analyzes the X-ray images of the left wrist to determine the individual's bone age. Traditional bone age assessment methods are all based on manual work. The reader compares the size and morphology of the wrist bones in the X-ray to be evaluated with the standard X-ray to assess bone age. There are three traditional bone age assessment methods: counting method, atlas method, and scoring method. These assessment methods require experts with certain professional knowledge to perform, and the assessment of an X-ray often takes a long time. Thanks to the development of neural networks, more and more bone age assessment methods based on deep learning have been proposed. These methods can complete the bone age assessment of an X-ray in a few seconds, and the accuracy of the assessment results is comparable to that of relevant experts in the field of bone age. Among these methods, some methods directly extract features from the entire X-ray image to perform bone age assessment. Although such methods utilize the features of the entire X-ray, they cannot determine which parts of the X-ray play a key role in bone age assessment, and the interpretability of the methods is low. Another part of the methods first segments the X-ray film, extracts the regions of interest, and then uses convolutional neural networks to extract the image features of these regions of interest to assess bone age. These methods are highly interpretable, and the selection of regions of interest is mostly based on the evaluation criteria of the scoring method, which requires certain prior knowledge and additional annotation of the target regions of interest. Summary of the invention
[0003] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a bone age assessment method based on fine-grained classification.
[0004] The present invention solves the technical problem by adopting the following technical solution:
[0005] A bone age assessment method based on fine-grained classification comprises the following steps:
[0006] Step 1: Acquisition of X-ray images of the left wrist bones;
[0007] Step 2: input the bone age assessment network based on fine-grained classification designed in the present invention to perform bone age assessment;
[0008] Step 3: Obtain bone age assessment results.
[0009] Step 1 specifically includes:
[0010] 1) Use the Python library function scipy.misc.imread() to read the X-ray image to be evaluated and save it as an array data type.
[0011] 2) Use the python library function Image.fromarray() to convert the array data type to the Image data type.
[0012] 3) Use the python library function transforms.Resize() to convert the image data into a size of 448*448, and use the same size for all X-ray images to be evaluated.
[0013] 4) Use the Python library function transforms.ToTensor() to convert the image data into a Tensor vector for use in the subsequent bone age assessment network.
[0014] The bone age assessment network used in step 2 is divided into four parts, namely feature extractor, region of interest selection subnet, guidance subnet and evaluation subnet:
[0015] 1) Feature extractor;
[0016] Use ResNet50 as a feature extractor to extract image features of the left wrist X-ray. Remove the last fully connected layer and Softmax layer of the ResNet50 network and only use the feature extraction part of the network. After the feature extractor, the image features in the X-ray are extracted to generate a feature map.
[0017] 2) Select subnets by region of interest;
[0018] The main function of the region of interest selection subnet is to select the region of interest with the largest amount of information in the X-ray film, that is, the region of interest with the most image features, based on the image features extracted by the feature extractor.
[0019] The input of the region of interest selection subnet is the feature map generated by the feature extractor. For these feature maps, the structure of Feature Pyramid Networks (FPN) + Path Aggregation Network (PAN) is used to fuse different features. Its structure is shown in the figure. After feature fusion, the language information and spatial information in the feature map are enhanced. Then, according to the response value in the feature map, the region of interest with the largest amount of information is selected, denoted as R1, R2, ..., RK. Then the amount of information in the K regions of interest is calculated, and the amount of information is defined as I(R). The purpose of the region of interest selection subnet is to select K ROIs, R1, R2, ..., RK in order so that they meet the condition I(R1)>I(R2)>...>I(Rk).
[0020] 3) Guidance subnet;
[0021] The role of the guidance subnet is to optimize the interest region selection subnet so that the interest region selection subnet can select the most informative ROI. For the K interest regions selected by the ROI selection subnet, R1, R2, ..., RK, the guidance subnet classifies these interests respectively, calculates their confidence C(R1), C(R2), ..., C(RK), and feeds them back to the interest region selection subnet. Confidence refers to the probability of classifying the interest region as its corresponding true label. The higher the confidence, the more helpful it is for classifying the entire wrist bone X-ray. The interest region selection subnet is continuously optimized during the network training process based on the feedback of the guidance subnet, so that the final selected interest region satisfies formula (1):
[0022]
[0023] That is, for the ROI selected in the ROI selection subnet, if its information volume is larger, its confidence is also higher.
[0024] 4) Evaluate subnet;
[0025] The role of the assessment subnet is to perform bone age assessment. For the input wrist bone X-ray, the feature extractor extracts the features of the entire image, and the region of interest selection subnet selects K regions of interest with the largest amount of information, and the feature extractor extracts the image features. In the bone age assessment subnet, these K+1 features are connected to form new image feature features. After two fully connected layers, the final bone age assessment result is obtained.
[0026] Step 3: Obtain bone age assessment results
[0027] When assessing bone age, 0 to 18 years old are divided into months, with each month being a category, for a total of 228 categories.
[0028] After the X-ray film of the left wrist bone to be evaluated is input into the bone age evaluation network, the corresponding bone age value will be output, ranging from 0 to 228. The accuracy of the present invention in bone age evaluation is accurate to months.
[0029] The present invention has the following beneficial effects:
[0030] (1) Accurately assess bone age from left wrist bone X-rays.
[0031] (2) The representative region of interest in the X-ray film, that is, the feature region with the maximum amount of information, can be selected.
[0032] (3) The invention can accurately assess the bone age of the left wrist bones of adolescents aged 0 to 18 years old based on X-ray films of the left wrist bones, and has great application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is the overall flow chart of the present invention.
[0034] Figure 2 It is a structural diagram of FPN+PAN in the interest area selection subnet in the present invention. DETAILED DESCRIPTION
[0035] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0036] A bone age assessment method based on fine-grained classification includes the following steps:
[0037] Step 1: Acquisition of X-ray images of the left wrist bones;
[0038] Step 2: input the bone age assessment network based on fine-grained classification designed in the present invention to perform bone age assessment;
[0039] Step 3: Obtain bone age assessment results.
[0040] Step 1 specifically includes:
[0041] 1) Use the Python library function scipy.misc.imread() to read the X-ray image to be evaluated and save it as an array data type.
[0042] 2) Use the python library function Image.fromarray() to convert the array data type to the Image data type.
[0043] 3) Use the python library function transforms.Resize() to convert the image data into a size of 448*448, and use the same size for all X-ray images to be evaluated.
[0044] 4) Use the Python library function transforms.ToTensor() to convert the image data into a Tensor vector for use in the subsequent bone age assessment network.
[0045] The bone age assessment network used in step 2 is divided into four parts, namely feature extractor, region of interest selection subnet, guidance subnet and evaluation subnet:
[0046] 1) Feature extractor;
[0047] Use ResNet50 as a feature extractor to extract image features of the left wrist X-ray. Remove the last fully connected layer and Softmax layer of the ResNet50 network and only use the feature extraction part of the network. After the feature extractor, the image features in the X-ray are extracted to generate a feature map.
[0048] 2) Select subnets by region of interest;
[0049] The main function of the region of interest selection subnet is to select the region of interest with the largest amount of information in the X-ray film, that is, the region of interest with the most image features, based on the image features extracted by the feature extractor.
[0050] The input of the region of interest selection subnet is the feature map generated by the feature extractor. For these feature maps, the structure of Feature Pyramid Networks (FPN) + Path Aggregation Network (PAN) is used to fuse different features. Its structure is shown in the figure. After feature fusion, the language information and spatial information in the feature map are enhanced. Then, according to the response value in the feature map, the region of interest with the largest amount of information is selected, denoted as R1, R2, ..., RK. Then the amount of information in the K regions of interest is calculated, and the amount of information is defined as I(R). The purpose of the region of interest selection subnet is to select K ROIs, R1, R2, ..., RK in order so that they meet the condition I(R1)>I(R2)>...>I(Rk).
[0051] 3) Guidance subnet;
[0052] The role of the guidance subnet is to optimize the interest region selection subnet so that the interest region selection subnet can select the most informative ROI. For the K interest regions selected by the ROI selection subnet, R1, R2, ..., RK, the guidance subnet classifies these interests respectively, calculates their confidence C(R1), C(R2), ..., C(RK), and feeds them back to the interest region selection subnet. Confidence refers to the probability of classifying the interest region as its corresponding true label. The higher the confidence, the more helpful it is for classifying the entire wrist bone X-ray. The interest region selection subnet is continuously optimized during the network training process based on the feedback of the guidance subnet, so that the final selected interest region satisfies formula (1):
[0053]
[0054] That is, for the ROI selected in the ROI selection subnet, if its information volume is larger, its confidence is also higher.
[0055] 4) Evaluate subnet;
[0056] The role of the assessment subnet is to perform bone age assessment. For the input wrist bone X-ray, the feature extractor extracts the features of the entire image, and the region of interest selection subnet selects K regions of interest with the largest amount of information, and the feature extractor extracts the image features. In the bone age assessment subnet, these K+1 features are connected to form new image feature features. After two fully connected layers, the final bone age assessment result is obtained.
[0057] Step 3: Obtain bone age assessment results
[0058] When assessing bone age, 0 to 18 years old are divided into months, with each month being a category, for a total of 228 categories.
[0059] After the X-ray film of the left wrist bone to be evaluated is input into the bone age evaluation network, the corresponding bone age value will be output, ranging from 0 to 228. The accuracy of the present invention in bone age evaluation is accurate to months.
[0060] The present invention designs an end-to-end bone age assessment model based on a fine-grained image classification model. The model consists of four parts: a feature extractor, a region of interest selection subnet, a guidance subnet and an evaluation subnet. The feature extractor is implemented based on a convolutional neural network, and uses ResNet50 to extract image features in a left hand wrist bone X-ray. The region of interest selection subnet is used to select multiple regions of interest in the X-ray film. These feature regions contain representative image features and can help classification. The guidance subnet can optimize the region of interest selection subnet through feedback, thereby guiding the region of interest selection subnet to select regions of interest more accurately. The evaluation subnet uses the extracted image features to perform bone age assessment. The bone age assessment model proposed in the invention can extract the regions of interest with the largest amount of information in the left hand wrist bone X-ray, and use these regions of interest to improve the accuracy of bone age assessment. The present invention uses a fine-grained image classification method to evaluate the bone age of the left wrist bone X-ray film, designs and implements a bone age evaluation network based on fine-grained image classification. When performing bone age evaluation, only the X-ray film to be evaluated and the gender corresponding to the X-ray film need to be input to obtain the evaluated bone age value. For the input X-ray film, the bone age evaluation network can adaptively extract several interest regions containing the most feature information, and use the image features of these interest regions to perform accurate bone age evaluation. The invention can accurately evaluate the bone age of the left wrist bone X-ray films of adolescents aged 0 to 18, and has great application value.
Claims
1. A bone age assessment method based on fine-grained classification, comprising the following steps: Step 1: Acquisition of X-ray images of the left wrist bones; specifically including: 1) Use the Python library function scipy.misc.imread() to read the X-ray image to be evaluated and save it as an array data type; 2) Use the python library function Image.fromarray() to convert the array data type to the Image data type; 3) Use the Python library function transforms.Resize() to convert the image data into a size of 448*448, and use the same size for all X-ray images to be evaluated; 4) Use the Python library function transforms.ToTensor() to convert the image data into a Tensor vector for subsequent use in the bone age assessment network; Step 2: Input a bone age assessment network based on fine-grained classification to perform bone age assessment; the bone age assessment network includes a feature extractor, a region of interest selection subnet, a guidance subnet, and an evaluation subnet: The feature extractor uses ResNet50 as a feature extractor to extract image features of the left wrist X-ray film; removes the last fully connected layer and Softmax layer of the ResNet50 network, and only uses the feature extraction part of the network; extracts the image features in the X-ray film through the feature extractor to generate a feature map; The region of interest selection subnet selects the region of interest with the largest amount of information in the X-ray film, that is, the region of interest with the most image features, according to the image features extracted by the feature extractor; The input of the region of interest selection subnet is the feature map generated by the feature extractor. For these feature maps, the structure of Feature Pyramid Networks (FPN) + Path Aggregation Network (PAN) is used to fuse different features. After feature fusion, the language information and spatial information in the feature map are enhanced. Then, according to the response value in the feature map, the region of interest with the largest amount of information is selected, which is recorded as R1, R2, ..., RK. Then the amount of information in the K regions of interest is calculated, and the amount of information is defined as I(R). The purpose of the region of interest selection subnet is to select K ROIs, R1, R2, ..., RK in order so that they meet the condition I(R1)>I(R2)>...>I(Rk). The guiding subnet optimizes the interest region selection subnet so that the interest region selection subnet can select the ROI with the most information. For the K interest regions, R1, R2, ..., RK, selected by the ROI selection subnet, the guiding subnet classifies these interests respectively, calculates their confidences C(R1), C(R2), ..., C(RK), and feeds them back to the interest region selection subnet. The confidence refers to the probability of classifying the interest region as its corresponding true label. The higher the confidence, the more helpful it is for classifying the entire wrist bone X-ray. The interest region selection subnet is continuously optimized during the network training process according to the feedback of the guiding subnet, so that the finally selected interest region satisfies formula (1): That is, for the ROI selected in the ROI selection subnet, if its information content is larger, its confidence is also higher; The assessment subnet performs bone age assessment. For the input wrist bone X-ray, the feature extractor extracts the features of the entire image, and the region of interest selection subnet selects K regions of interest with the largest amount of information. The feature extractor extracts the image features. In the bone age assessment subnet, these K+1 features are connected to form new image features. After two fully connected layers, the final bone age assessment result is obtained. Step 3: Obtain bone age assessment results; When assessing bone age, the age of 0 to 18 years is divided into months, with each month as a category, for a total of 228 categories; After the left wrist bone X-ray to be evaluated is input into the bone age assessment network, its corresponding bone age value will be output, ranging from 0 to 228.
Citation Information
Patent Citations
A bone age evaluation method based on two-stage neural network
CN109345508A
Hand bone X-ray film bone age evaluation method based on heterogeneous data fusion network
CN110503635A