An automatic bone age prediction method based on style transfer
By adopting style transfer technology in bone age prediction technology, the target domain image is converted into source domain style images, which solves the problems of lack of data and poor generalization ability of model, and improves the accuracy of bone age prediction and heat map positioning ability.
Patent Information
- Application Number
- CN202311240749.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-09-25
AI Technical Summary
The existing bone age prediction technology has problems such as lack of data and poor generalization ability of model, resulting in low accuracy of bone age prediction.
The automatic bone age prediction method based on style transfer is adopted, and the image set of the source domain and the target domain is obtained, and the target domain image is converted into the source domain style image is used to use the style migration network, and the accuracy of bone age prediction is improved through attention mechanism and feature extraction technology.
Through style transfer technology, unifying data style and pixel distribution without adding additional tags, improving the accuracy of bone age prediction and heat map positioning ability, reducing the time cost of repeated training.
Smart Images

Figure CN117314852B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of adaptation and bone age prediction, and particularly relates to an automatic bone age prediction method based on style transfer. Background Art
[0002] Bone age assessment plays an important role in understanding the growth and development of children. It is a medical examination conducted by pediatricians and pediatric endocrinologists. This assessment method is used to diagnose and treat growth and endocrine disorders in children and adolescents by comparing the skeletal bone age of children with their actual age, and to predict their final adult height. In addition, it can also be applied to the diagnosis and treatment of surgical operations involving spinal correction, lower limb balance, etc., as well as the fields of sports and forensic identification. However, traditional manual bone age assessment has some disadvantages. First, it has strong subjectivity and inconsistency. The bone age results of the same X-ray film evaluated by different doctors are often different, and even the bone age results of the same X-ray film evaluated by the same doctor at different times may also be different. Second, bone age assessment requires professional knowledge and long-term strict training, and the assessment process takes a long time.
[0003] Deep neural networks have been widely used in the medical field because automatically estimating bone age using deep learning methods is much faster than manual bone age judgment, and the accuracy is also far higher than traditional methods.
[0004] Hyunkwang Lee et al. used GoogLeNet as the backbone and determined the final bone age by showing three to five reference images of the G&P atlas. Toan Duc Bui et al. used Faster-RCNN and Inception-v4 networks respectively for the detection and classification of regions of interest (ROIs). Chuanbin Liu et al. introduced an attention agent and an identification agent for proposing distinguishable skeletal parts and feature learning and age assessment respectively.
[0005] Currently, there are the following problems in bone age prediction: 1. Lack of data. Since the hand bone image dataset is relatively small, the accuracy of model training is poor, resulting in low accuracy of final bone age prediction; 2. The generalization ability of the model between different individuals is poor, that is, the performance of the model on new hand bone images may be inferior to its performance on training data. Summary of the Invention
[0006] To solve the above problems existing in the prior art, the present invention proposes an automatic bone age prediction method based on style transfer, including: obtaining a source domain image set and a target domain image set; processing the source domain images using a bone age prediction model; inputting the target domain images into the bone age prediction model according to the processing results of the source domain images to obtain bone age prediction results;
[0007] Processing the source domain image using the bone age prediction model includes: calculating the pixel histogram of the source domain image and performing enhancement processing on the image; using the attention mechanism to extract the region of interest features from the enhanced image to obtain the source domain attention heat map and attention weights; cropping the source domain attention heat map to obtain the hand bone feature map, finger bone feature map, and metacarpal bone feature map; fusing the hand bone feature map, finger bone feature map, and metacarpal bone feature map and adding gender features, and then inputting them into the bone age prediction network to obtain the source domain bone age prediction result; determining the weights of each feature map according to the source domain bone age prediction result;
[0008] Processing the target domain image using the bone age prediction model includes: performing enhancement processing on the target domain image according to the pixel histogram of the source domain image; using the style transfer network to convert the target domain image into a source domain style image and performing smoothing filtering processing on the converted image; extracting the region of interest features from the filtered image according to the attention weights to obtain the target domain attention heat map; cropping the target domain attention heat map to obtain the hand bone feature map, finger bone feature map, and metacarpal bone feature map; fusing all the feature maps according to the weights of the source domain feature maps, and inputting the fused feature map into the bone age prediction network to obtain the final bone age prediction result.
[0009] Preferably, obtaining the source domain image in the target domain style includes: performing binaryzation processing on the source domain image and the target domain image to obtain the border and the contour extracted by the Sobel operator; cropping the source domain image and the target domain image according to the border and the contour extracted by the Sobel operator to obtain the hand bone image; processing the source domain image using AHE according to the pixel histogram of the target domain, and inputting the processed image into the Cyclegan network for filtering processing to obtain the source domain image in the target domain style.
[0010] The beneficial effects of the present invention:
[0011] Through the style transfer of the source domain to the target domain, the present invention realizes the unification of data style and pixel distribution without adding additional labels, improving the ability of heat map localization and the accuracy of bone age prediction. Through the transfer learning of the common model, the present invention reduces the time cost of repeated training. The preprocessing and feature extraction methods in the source domain of the present invention depend on the image distribution and learning model of the target domain, without the need for additional preprocessing, and the operation is simpler, greatly reducing the workload of doctors. Description of the Drawings
[0012] Figure 1 It is the overall flowchart of the present invention;
[0013] Figure 2 It is the process diagram of the heat map network annotation process of the present invention;
[0014] Figure 3Schematic diagram of the feature area cutting of the present invention;
[0015] Figure 4 Prediction process diagram of the bone age regression network of the present invention;
[0016] Figure 5 Schematic diagram of the CBAM attention process of the present invention;
[0017] Figure 6 Schematic diagram of the process of style transfer of the present invention;
[0018] Figure 7 Schematic diagram of the losses of the source domain and the target domain based on transfer learning of the present invention. Detailed implementation manners
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] The present invention discloses a fully automatic bone age prediction method based on domain adaptation: related to the field of bone age prediction, the style of a private hand bone dataset with better overall prediction effect is transferred to a public hand bone dataset with poorer overall effect to improve the prediction accuracy. The bone age prediction method includes three groups of networks, a style transfer network, a heatmap localization network, and a bone age prediction network. The specific steps are as follows. The private dataset passes through the heatmap localization module to locate the region of interest (ROI) and obtain a heatmap. Through multiple trainings and croppings of the heatmap and the original image, the hand bones, metacarpal bones, and phalanges are intercepted. The cut dataset is connected to the bone age prediction module to obtain the final predicted bone age, and the network is saved. The private dataset passes through the domain adaptation module to make the styles of the source domain and the target domain unified. Then, through the trained localization model and prediction model, the shallow layer of feature extraction is locked, and the attention mechanism is adjusted and then trained to obtain a highly accurate predicted bone age.
[0021] An automatic bone age prediction method based on style transfer, as Figure 1 shown, the method includes:
[0022] S1. Obtain a source domain image set and a target domain image set;
[0023] S2. Calculate the pixel histogram of the source domain image and perform enhancement processing on the image;
[0024] S3. Perform enhancement processing on the target domain image according to the pixel histogram of the source domain image;
[0025] S4. Use a style transfer network to convert the target domain image into a source domain style image, and perform smoothing filtering on the converted image;
[0026] S5. Use an attention mechanism to extract region-of-interest features from the enhanced image to obtain a source domain attention heatmap and attention weights;
[0027] S6. Crop the source domain attention heatmap to obtain a hand bone feature map, a finger bone feature map, and a metacarpal bone feature map;
[0028] S7. Fuse the hand bone feature map, the finger bone feature map, and the metacarpal bone feature map, add gender features, and then input them into the bone age prediction network to obtain a source domain bone age prediction result; determine the weights of each feature map according to the source domain bone age prediction result
[0029] S8. Extract region-of-interest features from the filtered image according to the attention weights to obtain a target domain attention heatmap;
[0030] S9. Crop the target domain attention heatmap to obtain a hand bone feature map, a finger bone feature map, and a metacarpal bone feature map;
[0031] S10. Fuse all the feature maps according to the weights of the source domain feature maps, and input the fused feature maps into the bone age prediction network to obtain the final bone age prediction result;
[0032] S11. Repeat steps S2 - S10 until the bone age prediction network converges to obtain a high-accuracy bone age prediction model;
[0033] S7. Use the high-accuracy bone age prediction model to predict the bone age of the image to be detected.
[0034] Converting the pixel histogram into a source domain image with the target domain style according to the target domain image includes:
[0035] Step 1: The hand bone localization network outputs features output through Inception_v3. After two poolings and a fully connected layer, the final prediction value is output. Here, output is F ∈ R H*W*C The feature map after feature extraction, where C represents the number of channels of the feature, and H, W represent the length and width of the feature map.
[0036] Step 2: The feature map At is the sum of the generated features F in the channel dimension. That is Then f represents the i-th feature map, and C represents the number of channels of the feature.
[0037] Step 3: Use the Inception_v3 network to train the source domain and target domain of the dataset, and use the mean squared error between the output and the true label. The label is smoothed to the range [0, 240].
[0038] Adopting an attention mechanism to extract region-of-interest features from source domain images in the target domain style includes:
[0039] Step 1: Set As the threshold for measuring whether the aggregated feature map At can be used as an effective information region, then the mask Mt. Set the size of the mask, overlap the source domain picture with Mt and crop the feature region, and fill the cut region of the original cut map Mask_out with random values and save it.
[0040] Step 2: Re-enter Mask_out into the localization network and repeat the processes of S103 and S105. Obtain the available feature region of the phalanx.
[0041] In this embodiment, an automatic bone age prediction method based on style transfer includes:
[0042] S1: Divide the target domain Xt (private dataset) and source domain Xs (RSNA public dataset) datasets into two groups of male and female according to gender.
[0043] S2: Calculate the contrast and image histogram of the source domain Xt. And preprocess the image through PM filtering and AHE (Adaptive Histogram Equalization).
[0044] S3: After quickly locating the region of interest of Xt through the Inception_V3 network and adding gender information, pass through the GAP (Global Average Pooling) and then through the FC layer. And extract the global feature or local feature At of the whole image, and save the model modell, where the size of Xt is 560*560.
[0045] S4: Input the local feature or the global feature into the network, and the output result of the network is F∈R H*W*C . Perform a GAP (Global Average Pooling) operation on this output, and then connect a fully connected layer (FC) behind. Denote the average value of the k-th feature map after global average pooling as S k . Denote the weight matrix of the fully connected layer as W∈R C*T , T is the number of categories of the classification model).
[0046]
[0047] Among them, Y t is the output of the model, W ktIt represents the weight of the k-th input neuron and the t-th output neuron in the last fully connected layer, thereby locating the heatmap At with label t.
[0048] S5: After the fully connected layer (FC), dense(240). That is, Y t is a vector of 240 elements. And the bone age label of this method adopts a softened label distribution method. The formula is as follows, where t represents the true bone age value and i is the one-hot label distribution
[0049]
[0050] Among them, l takes 240, mapping the age label to [0, 240].
[0051] S6: The loss function selects the MAE (mean absolute error) function.
[0052] S7: As Figure 2 shown, set the area with larger values in the heatmap as the area concerned by the network. After scaling the heatmap At to the size of Xt, crop out Xt-hand (hand bone) and Xt-R1 (metacarpal bone).
[0053] S71: Set the threshold parameter S. Keep the pixel values of At greater than the threshold S unchanged, and define the pixel values less than the threshold as 0. After cropping the region of interest through the heatmap, apply a mask to the cropped region of Xt and save the cropped image Xt-mask.
[0054] S72: Put Xt-mask back into the heatmap localization module. After obtaining a new heatmap, set the threshold again and crop out Xt-R2 (finger bone).
[0055] S9: As Figure 3 shown, input the above cropped Xt-hand into the Xception network. Input Xt-R1 and Xt-R2 into the Resnet50 network.
[0056] S10: The model output of Xception is x1, and the model outputs of Resnet50 are x2 and x3. As Figure 5 shown, after passing through the CBAM attention mechanism to obtain the output, X1, X2, and X3 pass through the GAP layer with different parameters to uniformly output features of dimension (1, 1, N). Among them, CBAM contains a channel attention module (M c ) and a spatial attention module, two sub-modules (M s ), and the formula is as follows:
[0057] The expression of the channel attention mechanism is:
[0058]
[0059] The spatial attention mechanism includes:
[0060]
[0061] Among them, σ represents the Sigmoid activation function, F represents the hand bone feature map, MLP represents the shared fully connected layer, AvgPool represents average pooling, Max Pool represents max pooling, and the weights of MLP are shared by W 1 and W 0 The channel information of one Feature Map is aggregated using two poolings to generate two feature maps. f 7×7 represents a convolutional kernel of size 7×7.
[0062] As Figure 4 shown, the spatial attention module performs GAP and GMP on the input feature F in the channel direction to obtain a feature map of size H*W*1. Then, these two feature maps are concatenated in the depth direction, and after convolution, the dimension is reduced to H*W*1. Finally, it is input into the activation function to obtain the spatial attention feature. After multiplying the spatial attention feature by the initial input feature, the final output result feature is obtained.
[0063] S11: Map the male and female gender labels to [-1, 1] and fuse the features. That is, Concatenate([X1, X2, X3, gender]). Finally, obtain the predicted bone age value through the FC layer. And save the model model2.
[0064] S12: The cross-entropy loss function can calculate the probability distribution distance between the model's predicted bone age and the label.
[0065]
[0066] Among them, c is the total number of categories, p(y j ) represents the probability value of the true label, y j is the true label, and is the predicted bone age value.
[0067] L hand = -log(P H (c))
[0068] L raw = -log(P R (c))
[0069]
[0070] Among them, c represents the true bone age label of the input hand bone image, and P RP and PH are the bone age category probabilities output after the fully connected layer for the original image branch and the hand image branch respectively. P P(n) represents the category output by the nth region of the local attention area branch, and N represents the number of selected attention areas. In order to enable the final converged model to perform bone age assessment based on the hand region features of the object or the local fine-grained features of the attention.
[0071] The total loss is:
[0072] L = αL raw + βL hand + λL parts
[0073] Since the contributions of the three branches to the overall model training are different, it may occur that the weight distribution for different branches is uneven during the model training process, resulting in a reduction in the overall performance of the model. Therefore, different branch weight coefficients are added to optimize the training of the overall model. α represents the weight of the original image branch, β represents the weight of the hand region branch, and λ represents the weight of the local attention area branch.
[0074] S13: As Figure 6 shown, input the source domains Xs and Xt into the Cyclegan network to obtain the dataset in the target domain style According to the contrast of Xt and the image histogram, obtain a dataset unified with Xt through PM filtering and AHE (Adaptive Histogram Equalization).
[0075] S14: According to S3 and S11, save the heatmap localization and prediction model. After locking the shallow layer of the feature extraction model, input lower the learning rate to obtain the region of interest and finally obtain the predicted bone age with high accuracy.
[0076] In this embodiment, the RSNA and CQJTJ datasets are used to complete the training respectively. During the training process, first perform overall training on the two datasets without distinguishing genders, and then train the hand bone images of male and female children separately by gender. The main evaluation criteria adopted are MAE, RMSE, and the correct rate within ±1 year; the loss schematic diagrams of the source domain and the target domain based on transfer learning when using the RSNA and CQJTJ datasets to complete the training respectively are as Figure 7 shown.
[0077] The preset training set used in the source domain is the RSNA2017 bone age prediction competition dataset, which includes 12,000 png-format X-ray images of the palm bones of European and American teenagers and their corresponding bone age labels. The target domain uses the dataset of 3,898 left wrist bone X-ray images collected by the Radiology Department of Chongqing Jintongjia Children's Hospital and the bone age annotation reports of these images by professional detecting physicians. This dataset is named the CQJTJ dataset, and the bone age is in months. 1,000 images are selected as the validation set, 200 images as the test set, and the rest as the training set.
[0078] The backbone network for bone age prediction is selected as the Inception_v3 and ResNet-50 networks, which can better complete the extraction of detailed feature information. The Adam_betal parameter in the training process is set to 0.9, the Adam_beta2 parameter is set to 0.999, the basic learning rate is set to 0.001, and the exponential decay learning rate is used as the training progresses to improve the model fitting effect. The decay rate is set to 0.98, and the decay step is set to every 10 epochs. The size of the original input image is set to 448×448, and the size of the cropped hand region image is also set to 448×448. The number of epochs for the RSNA dataset is set to 60, and the number of epochs for the CQJTJ dataset is set to 50. The batchsize of the number of images iterated each time is set to 32.
[0079] The present invention uses the tensor flow framework to implement the construction and training of various neural network models. The language used is Python 3.6 programming, and the graphics card model used is 1 GeForce RTX 3090, with a graphics card video memory of 24G.
[0080] In the bone age assessment model, when iterating to the 40th epoch, the model begins to converge, and the loss value at this time is approximately 1.2. In the CQJTJ dataset, when the number of iterations is 29, the model begins to converge, and the loss value at this time is about 0.85.
[0081] The above-mentioned embodiments further elaborate on the purpose, technical solution, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An automatic bone age prediction method based on style transfer, characterized in that, it includes: Obtain the source domain image set and the target domain image set; process the source domain images using a bone age prediction model; According to the processing results of the source domain images, input the target domain images into the bone age prediction model to obtain bone age prediction results; Processing the source domain images using a bone age prediction model includes: calculating the pixel histogram of the source domain images and performing enhancement processing on the images; using an attention mechanism to extract region-of-interest features from the enhanced images to obtain a source domain attention heat map and source domain attention weights; cropping the source domain attention heat map to obtain hand bone feature maps, finger bone feature maps, and metacarpal bone feature maps; fusing the hand bone feature maps, finger bone feature maps, and metacarpal bone feature maps and adding gender features, then inputting them into a bone age prediction network to obtain source domain bone age prediction results; determining the weights of each feature map according to the source domain bone age prediction results; Processing the target domain images using a bone age prediction model includes: performing enhancement processing on the target domain images according to the pixel histogram of the source domain images; using a style transfer network to convert the target domain images into source domain style images and performing smoothing filtering on the converted images; performing region-of-interest feature extraction on the filtered images according to the source domain attention weights to obtain a target domain attention heat map; cropping the target domain attention heat map to obtain hand bone feature maps, finger bone feature maps, and metacarpal bone feature maps; fusing all the feature maps according to the weights of the source domain feature maps, and inputting the fused feature maps into a bone age prediction network to obtain the final bone age prediction results; Using a style transfer network to convert the target domain images into source domain style images includes: performing binarization processing on the source domain images and the target domain images to obtain borders and contours extracted by the Sobel operator; cropping the source domain images and the target domain images according to the borders and contours extracted by the Sobel operator to obtain hand bone images; processing the source domain images using AHE according to the pixel histogram of the target domain, and inputting the processed images into a Cyclegan network for filtering processing to obtain source domain style images; The bone age prediction network uses the XceptionNet network and the Resnet50 network as the backbone networks of the bone age prediction network; input the metacarpal bone feature maps into the Xception network to obtain feature map x1; input the finger bone feature maps and the hand bone feature maps into the Resnet50 network to output feature map x2 and feature map x3; pass feature map x1, feature map x2, and feature map x3 through the CBAM attention mechanism to obtain a metacarpal bone attention feature map X1, a finger bone attention feature map X2, and a hand bone attention feature map X3; pass X1, X2, X3 through GAP layers with different parameters to uniformly output features in the (1, 1, N) dimension; map the male and female gender labels to [-1, 1], and fuse the features, that is, Concatenate([X1, X2, X3, gender]); pass the fused feature map through the FC layer to obtain the bone age prediction value.
2. The automatic bone age prediction method based on style transfer according to claim 1, characterized in that, Processing the source domain image using AHE includes: setting a threshold; correcting the source domain image according to the set threshold; and performing brightness enhancement and contrast enhancement processing on the corrected source domain image using the adaptive histogram equalization algorithm.
3. A method for automatic bone age prediction based on style transfer according to claim 1, characterized in that The process of filtering the processed image by the Cyclegan network includes: Among them, I is the grayscale image to be filtered, t is the positioning label, ΔI represents the Laplacian operator, is the local image gradient operator, c(x, y) represents the diffusion coefficient of an image with size x*y, and K represents the heat conduction coefficient.
4. A method for automatic bone age prediction based on style transfer according to claim 1, characterized in that Using the attention mechanism to extract the region of interest features from the source domain image with the target domain style includes: The image passes through the localization network Inception_v3, and the attention mechanism is used to extract features from the image, and the feature output is output; the feature output is passed through two pooling layers and a fully connected layer, and finally the predicted value is output; where output is F∈R H*W*C The feature map after feature extraction, where C represents the number of channels of the feature, and H, W represent the length and width of the feature map.
5. A method for automatic bone age prediction based on style transfer according to claim 4, characterized in that The attention mechanism includes a channel attention mechanism and a spatial attention mechanism; the expression of the channel attention mechanism is: The spatial attention mechanism includes: Among them, σ represents the Sigmoid activation function, F represents the hand bone feature map, MLP represents the shared fully connected layer, AvgPool represents average pooling, and Max Pool represents max pooling; the weights of the MLP are shared by W 1 and W 0 to aggregate the channel information of one Feature Map using two poolings, generating two feature maps; f 7×7 represents a convolutional kernel of size 7×7.
Citation Information
Patent Citations
Hand bone X-ray film bone age evaluation method based on heterogeneous data fusion network
CN110503635A
Children hand bone X-ray image bone age evaluation method and system based on model integration
CN116342516A