A method for predicting bone age in children

By constructing the TENet model, using hand topology maps, edge feature enhancement modules and deep learning network improvement models, the problem of difficulty in accurately predicting bone age in contemporary Chinese children in the existing technology is solved, and higher prediction accuracy and reliability are achieved.

CN115423779BActive Publication Date: 2025-05-06CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211076541.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-05-06
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the bone age of contemporary Chinese children, especially under conditions based on the Chinese-05 standard.

Method used

A TENet model is constructed using a method that includes hand topology maps, edge feature enhancement modules and deep learning network improvement models. The model extracts and enhances the characteristic information of the hand X-ray through the improved Canny edge detection algorithm and Otsu algorithm, combining bilateral filtering and denoising and automatic color equalization algorithm.

Benefits of technology

It improves the accuracy and reliability of children's bone age prediction, can better reflect the development trends of contemporary Chinese children, and achieves a high bone age prediction performance on the public data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115423779B_ABST
    Figure CN115423779B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for predicting bone age of children, comprising the following steps: randomly selecting a part of hand X-rays from a public data set of hand X-rays with bone age labels to form a data set, and adjusting all pictures to a specified size; establishing a TENet model and constructing a training set at the same time; using the training set as the input of the TENet model, using the Adam optimizer to train the TENet model, and obtaining a trained TENet model when the maximum number of iterations is reached. The model of the present invention can accurately predict the bone age of children.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer image vision and big data medical image applications, and in particular to a method for predicting bone age in children. Background Art

[0002] Traditional bone age assessment usually involves filming the subject's hands and wrists, and then the doctor interprets the results based on the film. Interpretation methods include simple counting, atlas, scoring, and computer bone age scoring systems. The most commonly used methods are the GP atlas and TW3 scoring methods. Affected by racial differences and long-term trends in growth and development, the most suitable bone age standard for contemporary Chinese children is CHN-05. In 2006, the China-05 bone age standard was approved as an industry standard of the People's Republic of China, becoming the only industry bone age standard in China.

[0003] Recently, deep neural networks have been widely used in the medical field because the automatic estimation of bone age using deep learning methods is much faster than manual judgment of bone age, and the accuracy is far higher than traditional methods. HyunkwangLee et al. used GoogLeNet as the backbone to determine the final bone age by displaying three to five reference images of the G&P atlas. ToanDucBui et al. used Faster-RCNN and Inception-v4 networks for region of interest (ROI) detection and classification, respectively. ChuanbinLiu et al. introduced an attention agent and a recognition agent for proposing discriminative skeletal parts and feature learning and age assessment, respectively.

[0004] Although these methods can achieve good performance in BAA, in the actual detection process, due to the development and changes of the times, these methods cannot objectively and accurately evaluate the bone age of contemporary Chinese children. Based on the above problems, it is urgent to find a method to conduct research based on the standardized sample of Zhonghua-05. The standard of Zhonghua-05 is composed of normal children of urban Han nationality in China in 2005, which can truly reflect the development trend of contemporary Chinese children. The edge of the metacarpal phalanges and the length of the metacarpal phalanges of Chinese children are used as reference standards to obtain the correct bone age. Summary of the invention

[0005] In view of the above problems existing in the prior art, the technical problem to be solved by the present invention is: how to accurately predict the bone age of children.

[0006] To solve the above technical problems, the present invention adopts the following technical solution: a method for predicting bone age in children, comprising the following steps:

[0007] S100: Select a public dataset, which includes hand X-rays with bone age labels; randomly select some hand X-rays to form dataset A, and resize all images in A to 560×560;

[0008] S200: Establishing a TENet model, the TENet model includes a hand topology module, an edge feature enhancement module and a deep learning network improved model D; wherein the edge feature enhancement module is an improved Canny edge detection algorithm, and the improved Canny edge detection algorithm uses a bilateral filtering denoising algorithm and an Otsu algorithm;

[0009] S300: Build the training set of TENet model:

[0010] S310: Preprocess the images in A using an automatic color balancing algorithm to obtain a preprocessed image corresponding to each image in A;

[0011] S320: Using a hand topology module to perform feature extraction processing on the images in A to obtain a hand topology map corresponding to each image in A;

[0012] S330: Use an edge feature enhancement module to perform edge enhancement on the preprocessed image corresponding to the image in A to obtain an edge feature enhanced image corresponding to the image in A;

[0013] S340: Repeat S310-S330, traverse all hand X-rays in A, and obtain a preprocessing map corresponding to each image in A, a hand topology map corresponding to each image in A, and an edge feature enhancement map corresponding to each image in A;

[0014] S350: All pre-processed images, hand topology images and edge feature enhancement images obtained in S340 are combined into a training set of the TENet model, wherein the pre-processed images, hand topology images and edge feature enhancement images corresponding to each image form a training sample; S400: Preset the maximum number of iterations, preset the initial learning rate of D as μ, and train D:

[0015] S410: taking all training samples in the training set as input of D, training D using the Adam optimizer to update the parameters of D, wherein the MAE target loss function is used in the Adam optimizer;

[0016] When the maximum number of iterations is reached, the training is stopped and the trained deep learning network improved model D' is obtained;

[0017] S420: Using the TENet model of D' as the trained TENet model;

[0018] S500: Input a child's hand X-ray film whose unknown bone age needs to be predicted into the trained TENet model, and the output is the predicted value of the child's bone age.

[0019] Preferably, the structure of the deep learning network improved model D in S200 is to use a deep learning network as the backbone network, remove the topmost neural network layer in order from top to bottom, and then add a 3x3 downsampling layer, a 3x3 maximum pooling layer and a fully connected layer with 32 neurons in sequence after the bottom layer of the original backbone network, and integrate the gender information corresponding to the hand X-ray in A into the image features through the fully connected layer.

[0020] Preferably, the specific steps of obtaining the pre-processed image corresponding to the image in A in S310 are as follows:

[0021] S311: Select the cth picture from A:

[0022] S312: Calculate the adaptive filtering intermediate variable R of any pixel point p in the cth image c (p);

[0023]

[0024] Among them, p and q represent two pixels on the cth image, I c (p) and I c (q) represents the grayscale difference between p and q, d(p,q) represents the distance metric function between p and q, S α (·) represents the brightness expression function;

[0025] S313: Calculate the L(p) value of the cth image. The calculation expression is as follows:

[0026]

[0027] Among them, L(p) represents the preprocessing result of pixel p, and the mapping interval of pixel p is [0,255], [minR c (p), maxR c (p)] represents the full domain of L(p);

[0028] S314: repeat S312-S313, traverse all the pixels on the c-th image, pre-process all the pixels in the c-th image, and obtain a pre-processed image corresponding to the c-th image;

[0029] Preprocessing the images can solve the problem of low contrast between the background area and ROI in hand X-rays.

[0030] Preferably, the specific steps of obtaining the hand topology map corresponding to the image in A in S320 are as follows:

[0031] S321: Select the cth picture from A:

[0032] S322: Use the horizontal matrix to perform a plane convolution with the t-th pixel on c to obtain an approximate brightness difference G of the t-th pixel in the horizontal direction xt , the specific expression is as follows:

[0033]

[0034] Among them, G xt represents the approximate brightness difference of the t-th pixel in the horizontal direction, c t represents the tth pixel;

[0035] Use the vertical matrix to perform a plane convolution with the t-th pixel on c to obtain the approximate brightness difference G of the t-th pixel in the vertical direction yt , the specific expression is as follows:

[0036]

[0037] Among them, G yt Represents the approximate brightness difference of the t-th pixel in the vertical direction;

[0038] S323: Calculate the gradient approximation G of the t-th pixel on c t , the expression is as follows:

[0039]

[0040] Calculate the direction of the t-th pixel on c. The expression is as follows:

[0041]

[0042] Among them, α(x t ,y t ) represents the direction of the t-th pixel on c;

[0043] S324: Repeat S322-S323, traverse all the pixels on c, obtain the convolution value of each pixel, and the convolution values ​​of all the pixels form a matrix B, which is the gradient map c';

[0044] S325: Perform threshold binarization processing on c' to obtain a rough foreground edge map of the hand skeleton of c.

[0045] S326: Use the MediaPipe application framework to predict the positions of the joint points on the rough foreground edge map of the hand backbone corresponding to c to obtain a hand topology map of c.

[0046] The hand topology map obtained by using the rough foreground edge map of the hand skeleton can obtain structured ROI information, which can well express the corresponding semantic features in CHN-05, that is, the length of the metacarpophalangeal bones of the hand; it can be more conducive to the later assessment of bone age.

[0047] Preferably, the specific steps of obtaining the edge feature enhancement map corresponding to the image in A in S330 are as follows:

[0048] S331: Select the cth picture from A:

[0049] S332: Calculate the spatial proximity and pixel value similarity of each pixel point on the c-th image and perform bilateral filtering denoising on the c-th image. The expression of spatial proximity is as follows:

[0050]

[0051] Among them, i, j, k, l represent the coordinate values ​​of pixel points a and b, namely a(i, j) and b(k, l), δ d is the Gaussian standard deviation;

[0052] The expression of pixel value similarity is as follows:

[0053]

[0054] Among them, f(i,j) represents the pixel value of the image at point a(i,j), f(k,l) represents the pixel value of the image at point b(k,l), δ r is the Gaussian standard deviation, W r To calculate the difference between the near point a and the center point b;

[0055] Suppose the image c is obtained after bilateral filtering and denoising.

[0056] S333: Setting a threshold T to divide c" into background and target;

[0057] The optimal pixel threshold of c" is calculated using the Otsu algorithm. The calculation formula is as follows:

[0058] p0+p1=1; (9)

[0059]

[0060]

[0061] Among them, σ 2 represents the between-class variance, represents the average gray value of c", p0 represents the proportion of the background in the picture, p1 represents the proportion of the target in the picture, and the expression is as follows:

[0062]

[0063]

[0064] Where k = {1,…,T}, n represents the category of grayscale value, p n Represents the frequency of occurrence of each type of gray value, p0 represents the sum of the probability of occurrence of each type of gray value between 0 and k, L represents the pixel level of the image, and p1 represents the sum of the probability of occurrence of each type of gray value between k+1 and L;

[0065] m0 represents the average gray value of the background, m1 represents the average gray value of the target, and the expression is as follows:

[0066]

[0067]

[0068] S334: Substitute formula (10) into formula (11) and simplify to obtain formula (16), which is as follows:

[0069] σ 2 =p0p1(m0-m1) 2 ; (16)

[0070] S335: Traverse all pixel points on c'', and calculate the pixel points with the same gray value on c'' in turn through formula (9) - formula (16) to obtain N σ 2 ;

[0071] S336: The obtained N σ 2 Sort in descending order and select the largest value of σ 2 As the pixel optimal threshold of c”;

[0072] S337: Use c" processed by 336 as the edge feature enhancement image corresponding to the c-th image.

[0073] In the traditional Canny edge detection algorithm, linear Gaussian filtering is used for denoising, in which convolution with weighted coefficients is used to denoise the image. In essence, it is a mean blur, but Gaussian blur is based on weighted average. The closer the distance, the greater the weight, and the farther the distance, the smaller the weight. Therefore, calculating the mean will "blur" the edge information and feature information in the image, and many features will be lost. Bilateral filtering is a nonlinear filtering method. It is a compromise between the spatial proximity and pixel value similarity of the image, while considering spatial information and grayscale similarity. That is, it achieves the purpose by introducing spatial domain kernels and value domain kernels. This method not only considers the influence of the distance of the pixel in space, but also considers the influence of the similarity of pixel brightness. Therefore, the bilateral filter can well retain the edge features of the image and filter out the noise of low-frequency components.

[0074] In the traditional Canny edge detection algorithm, a dual threshold algorithm is used to classify pixels. However, the fixed threshold cannot be universal for each different image, so there will be an inherent problem that the TC, Laplacian, Prewitt and Roberts operators cannot identify the ossification center and the boundary of the metacarpal end. The Otsu algorithm, which calculates the best threshold based on different images instead of manually selecting the threshold, minimizes the uncertainty of the algorithm performance and makes the final result more consistent with the semantics of CHN-05, that is, the ossification center is clearly visible, disc-shaped, with a smooth and continuous edge, and the degree of fusion of the epiphysis and the diaphysis can be observed.

[0075] Compared with the prior art, the present invention has at least the following advantages:

[0076] 1. The model of the present invention uses a hand topology map, an edge feature enhancement module and a deep learning network improved model to construct a bone age prediction and evaluation model, namely the TENet model. The hand topology map is used in the TENet model to strengthen the information feature extraction of hand X-rays, thereby increasing the available information. The improved Canny edge detection algorithm, namely the edge feature enhancement module, is used to improve the amount and accuracy of information feature extraction of the edge part of the hand X-ray. At the same time, the Otsu algorithm is used instead of the manually selected threshold to minimize the uncertainty of the algorithm performance, thereby ensuring that the final result is more in line with the semantics of CHN-05.

[0077] 2. Pulse noise often exists in hand X-rays. Traditional denoising methods such as Gaussian filtering, median filtering, and box filtering cannot protect high-frequency information well, thereby blurring the edge details of the image. The bilateral filtering denoising algorithm used in the present invention can achieve the effect of edge-preserving denoising while retaining the edge features of the image by compromising the spatial proximity and pixel similarity of the image, and considering the spatial information and grayscale similarity.

[0078] 3. Use the Otsu algorithm to calculate the optimal threshold based on different images instead of manually selecting the threshold to minimize the uncertainty of the algorithm performance.

[0079] 4. In data preprocessing, an automatic color balancing algorithm is used to solve the problem of low contrast between the background area and ROI in hand X-rays. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 This is the overall framework diagram of the TENet model.

[0081] Figure 2 This is the network structure used in the present invention.

[0082] Figure 3 This is the flowchart of the improved Canny algorithm.

[0083] Figure 4 These are the results of different algorithms processing hand X-rays.

[0084] Figure 5 The processing results of the data set used by the method of the present invention are shown in Figure 1. (a)-(d) are sample images using a private database, and (e)-(h) are sample images using the RSNA database. DETAILED DESCRIPTION

[0085] The present invention is further described in detail below. A new automatic bone age assessment model TENet is proposed, which abstracts the semantic description of bone age assessment in CHN-05 and uses an algorithm to achieve the purpose of horizontal fusion of multi-feature expression. The present invention first designs a hand topology module to provide the length and position structure information of the metacarpal bones to meet the semantics of CHN-05 four-dimensional bone age assessment; secondly, an enhanced edge feature module is developed to identify the hand structure and extract effective edge information to meet the CHN-05 requirement for clear bone edges, and also enhance the accuracy of bone age assessment.

[0086] See also Figure 1-5 , a method for predicting bone age in children, comprising the following steps:

[0087] S100: Select a public dataset, which includes hand X-rays with bone age labels; randomly select some hand X-rays to form dataset A, and resize all images in A to 560×560;

[0088] S200: Establishing a TENet model, the TENet model includes a hand topology module, an edge feature enhancement module and a deep learning network improved model D, the deep learning network is a prior art, wherein the edge feature enhancement module is an improved Canny edge detection algorithm, the Canny edge detection algorithm is a prior art, the improved Canny edge detection algorithm uses a bilateral filtering denoising algorithm and an Otsu algorithm, the bilateral filtering denoising algorithm and the Otsu algorithm are prior art;

[0089] The structure of the deep learning network improved model D in S200 is to use the deep learning network as the backbone network, remove the topmost neural network layer from top to bottom, and then add a 3x3 downsampling layer, a 3x3 maximum pooling layer and a fully connected layer with 32 neurons in sequence after the bottom layer of the original backbone network, and integrate the gender information corresponding to the hand X-ray in A into the image features through the fully connected layer.

[0090] S300: Build the training set of TENet model:

[0091] S310: using an automatic color balancing algorithm to perform image preprocessing on the images in A, and obtaining a preprocessed image corresponding to each image in A; the automatic color balancing algorithm is an existing technology, and the automatic color balancing algorithm is used to perform regional adaptive filtering on the original image, complete chromatic aberration correction, obtain a spatial domain reconstructed image, highlight the features and valuable information in the image, and make the image more consistent with the perception of the human eye;

[0092] The specific steps of obtaining the pre-processed image corresponding to the image in A in S310 are as follows:

[0093] S311: Select the cth picture from A:

[0094] S312: Calculate the adaptive filtering intermediate variable R of any pixel point p in the cth image c (p);

[0095]

[0096] Among them, p and q represent two pixels on the cth image, I c (p) and I c (q) represents the grayscale difference between p and q, which represents the lateral inhibition in biological simulation. d(p,q) represents the distance metric function between p and q, which maps the regional adaptability of the filter. α (·) represents the brightness expression function;

[0097] S313: Calculate the L(p) value of the cth image. The calculation expression is as follows:

[0098]

[0099] Among them, L(p) represents the preprocessing result of pixel p, and the mapping interval of pixel p is [0,255], [minR c (p), maxR c (p)] represents the full definition domain of L(p), and the purpose of the full definition domain is to achieve global white balance of the image;

[0100] S314: repeat S312-S313, traverse all the pixels on the c-th image, pre-process all the pixels in the c-th image, and obtain a pre-processed image corresponding to the c-th image;

[0101] S320: Using a hand topology module to perform feature extraction processing on the images in A to obtain a hand topology map corresponding to each image in A;

[0102] The specific steps of obtaining the hand topology map corresponding to the image in A in S320 are as follows:

[0103] S321: Select the cth picture from A:

[0104] S322: Use the horizontal matrix to perform a plane convolution with the t-th pixel on c to obtain an approximate brightness difference G of the t-th pixel in the horizontal direction xt , the specific expression is as follows:

[0105]

[0106] Among them, G xt represents the approximate brightness difference of the t-th pixel in the horizontal direction, c t represents the tth pixel;

[0107] Use the vertical matrix to perform a plane convolution with the t-th pixel on c to obtain the approximate brightness difference G of the t-th pixel in the vertical direction yt , the specific expression is as follows:

[0108]

[0109] Among them, G yt Represents the approximate brightness difference of the t-th pixel in the vertical direction;

[0110] S323: Calculate the gradient approximation G of the t-th pixel on c t , the expression is as follows:

[0111]

[0112] Calculate the direction of the t-th pixel on c. The expression is as follows:

[0113]

[0114] Among them, α(x t ,y t ) represents the direction of the t-th pixel on c;

[0115] S324: Repeat S322-S323, traverse all the pixels on c, obtain the convolution value of each pixel, and the convolution values ​​of all the pixels form a matrix B, which is the gradient map c';

[0116] S325: Perform threshold binarization processing on c', which is a conventional technique, to obtain a rough foreground edge map of the hand skeleton of c.

[0117] S326: Use the MediaPipe application framework to predict the positions of the joint points on the rough foreground edge map of the hand backbone corresponding to c, where the MediaPipe application framework is an existing technology, to obtain a hand topology map of c.

[0118] S330: Use an edge feature enhancement module to perform edge enhancement on the preprocessed image corresponding to the image in A to obtain an edge feature enhanced image corresponding to the image in A;

[0119] The specific steps of obtaining the edge feature enhancement map corresponding to the image in A in S330 are as follows:

[0120] S331: Select the cth picture from A:

[0121] S332: Calculate the spatial proximity and pixel value similarity of each pixel point on the c-th image and perform bilateral filtering denoising on the c-th image. The expression of spatial proximity is as follows:

[0122]

[0123] Among them, i, j, k, l represent the coordinate values ​​of pixel points a and b, namely a(i, j) and b(k, l), δ d is the Gaussian standard deviation;

[0124] The expression of pixel value similarity is as follows:

[0125]

[0126] Among them, f(i,j) represents the pixel value of the image at point a(i,j), f(k,l) represents the pixel value of the image at point b(k,l), δ r is the Gaussian standard deviation, W r To calculate the difference between the near point a and the center point b;

[0127] Suppose the image c is obtained after bilateral filtering and denoising.

[0128] S333: Setting a threshold T to divide c" into background and target;

[0129] The optimal pixel threshold of c" is calculated using the Otsu algorithm. The calculation formula is as follows:

[0130] p0+p1=1; (9)

[0131]

[0132]

[0133] Among them, σ 2 represents the between-class variance, represents the average gray value of c", p0 represents the proportion of the background in the picture, p1 represents the proportion of the target in the picture, and the expression is as follows:

[0134]

[0135]

[0136] Where k = {1,…,T}, n represents the category of grayscale value, p n Represents the frequency of occurrence of each type of grayscale value, p0 represents the sum of the probability of occurrence of each type of grayscale value between 0 and k, L represents the pixel level of the image, which is generally 255, and p1 represents the sum of the probability of occurrence of each type of grayscale value between k+1 and L;

[0137] m0 represents the average gray value of the background, m1 represents the average gray value of the target, and the expression is as follows:

[0138]

[0139]

[0140] S334: Substitute formula (10) into formula (11) and simplify to obtain formula (16), which is as follows:

[0141] σ 2 =p0p1(m0-m1) 2 ; (16)

[0142] S335: Traverse all pixel points on c'', and calculate the pixel points with the same gray value on c'' in turn through formula (9) - formula (16) to obtain N σ 2 , the gray level refers to the gray level to which the pixel belongs;

[0143] S336: The obtained N σ 2 Sort in descending order and select the largest value of σ2 As the optimal pixel threshold of c", the calculation process is to loop through all pixels in the grayscale range of 0-255, so as to obtain the grayscale that maximizes formula (16) as the optimal pixel threshold of the image;

[0144] S337: Use c" processed by 336 as the edge feature enhancement image corresponding to the c-th image.

[0145] S340: Repeat S310-S330, traverse all hand X-rays in A, and obtain a preprocessing map corresponding to each image in A, a hand topology map corresponding to each image in A, and an edge feature enhancement map corresponding to each image in A;

[0146] S350: All pre-processed images, hand topology images and edge feature enhancement images obtained in S340 are combined into a training set of the TENet model, wherein the pre-processed image, hand topology image and edge feature enhancement image corresponding to each image form a training sample;

[0147] S400: Preset the maximum number of iterations, preset the initial learning rate of D to μ, and train D:

[0148] S410: taking all training samples in the training set as input of D, training D using the Adam optimizer to update the parameters of D. The Adam optimizer is a prior art, and the MAE target loss function is used in the Adam optimizer.

[0149] When the maximum number of iterations is reached, the training is stopped and the trained deep learning network improved model D' is obtained;

[0150] S420: Using the TENet model of D' as the trained TENet model;

[0151] S500: Input a child's hand X-ray film whose unknown bone age needs to be predicted into the trained TENet model, and the output is the predicted value of the child's bone age.

[0152] Experimental verification

[0153] Pediatric bone age assessment (BAA) is an important human physiological examination that can reflect human growth potential and sexual maturation trends. In clinical practice, "Chinese Adolescents and Children's Wrist Bone Maturity and Evaluation Method" (CHN-05) is a widely used method for Chinese radiologists to perform BAA. CHN-05 uses metacarpal phalangeal length (ML) and metacarpal phalangeal margin (MM) as reference standards to estimate bone age. Inspired by the semantic description of CHN-05, this paper proposes a new model, called TENet, for automatic bone age assessment. In the TENet model, a hand topology module is designed to identify the key positions of the hand for extracting structured semantics; an edge feature enhancement module is designed to provide accurate bone edge information during training. The TENet model can detect the overall edge information and the local information of the topological structure, so as to achieve the purpose of multi-feature level fusion to evaluate bone age.

[0154] Experimental results show that the TENet model achieves the most advanced model performance of 5.35 mean absolute error (MAE) on the public dataset RSNA. The design of the TENet model follows the CHN-05 semantic logic, so it has very good reliability and interpretability in clinical use.

[0155] This experiment evaluates the TENet model on a private dataset and the North American Radiological Society Bone Age Assessment (RSNA-BAA) public dataset. The private dataset contains 4954 hand X-rays in the training set, 275 hand X-rays in the validation set and test set; the RSNA-BAA public dataset contains 5611 hand X-rays in the training set, 275 hand X-rays in the validation set and test set. Before being processed by the TENet model, all hand X-rays were resized to 560×560. The mean absolute error between the ground truth bone age and the corresponding predicted bone age is also reported.

[0156] Table 1 Ablation study of the model of the present invention

[0157] EXP. OI PHI ImCAT L1 MAE 1 (Private dataset) √ √ 7.15 2 (Private Dataset) √ √ 6.92 3 (Private Dataset) √ √ √ 5.76 4(Public Dataset) √ √ √ 5.80

[0158] (EXP stands for Experiment, OI stands for Original Image, PHI stands for Pre-processed hand image, ImCAT stands for Improved adaptive thresholding of canny edge detection, L1 stands for Loss 1, MAE loss function, MAE stands for Mean absolute error, and PHI stands for Pre-processed hand image.)

[0159] Implementation details

[0160] TENet was implemented using Tensorflow 1.9 and trained on a system with NVIDIA TITANRTX GPU and 32GRAM, which took about 8 hours. We used 100 epochs to train the entire process of TENet. The batch size was 32. The initial learning rate was 3×10 -3 , after 50 epochs, it is reduced to 10 -3 After 80 epochs, the optimization is reduced to 10 times. The optimizer used is Adam.

[0161] TENet model ablation experiment

[0162] First, the role of each module in the TENet model is studied, namely the original image, preprocessed hand image (PHI) and edge feature enhancement (ImCAT). The experimental results are shown in Table 1. Without these two modules, the network degenerates to Xception, and the MAE score is 7.19 months; when the preprocessed image is applied to the model instead of the original image, the MAE score is 6.92 months; when the edge feature enhancement module is applied to the model, the MAE score is 5.76 months, and the MAE score is improved by 1.16 months; the actual experimental data illustrates the necessity of obtaining edge features and the effectiveness of the ImCAT algorithm.

[0163] Combining these two components, the TENet model achieved a MAE of 5.80 months on the public dataset. In summary, the results of the private dataset are better than those of the public dataset, which can prove the correctness of the TENet model design inspired by the CHN-05 wrist bone standard. Note: The smaller the MAE value, the better the effect.

[0164] Table 2 Ablation studies detected by ImCAT

[0165]

[0166]

[0167] Hand edge feature ablation experiment

[0168] The TENet model detects the edge features of the hand and combines the topological structure for bone age assessment. By observation, ImCAT also benefits from the design of its modules, and the performance is measured by peak signal-to-noise ratio (PSNR), mean square error (MSE), structural similarity (SSIM), root mean square error (RMSE) and mean absolute error (MAE), and the results are shown in Table 2. Taking the traditional Canny (TC) algorithm as the baseline and using four edge operators as the baseline, the ImCAT algorithm used in the present invention is significantly better than other algorithms, indicating that ImCAT has the ability to better capture the overall edge details of the hand.

[0169] The results of RSNA public and private datasets show that the TENet model achieves good performance. In addition, the TENet model is designed based on CHN-05 prior knowledge, so it has good interpretability and reliability for clinical practice. In addition, the present invention also provides a reliable experimental basis for trying to merge the feature acquisition module and the scoring regression module to form an end-to-end automatic bone age assessment framework.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution of the present invention, which should be included in the scope of the claims of the present invention.

Claims

1. A method for predicting bone age in children, characterized in that: The steps include: S100: Select a public dataset, which includes hand X-rays with bone age labels; randomly select some hand X-rays to form dataset A, and resize all images in A to 560×560; S200: Establishing a TENet model, the TENet model includes a hand topology module, an edge feature enhancement module and a deep learning network improved model D, wherein the edge feature enhancement module is an improved Canny edge detection algorithm, and the improved Canny edge detection algorithm uses a bilateral filtering denoising algorithm and an Otsu algorithm; S300: Build the training set of TENet model: S310: Preprocess the images in A using an automatic color balancing algorithm to obtain a preprocessed image corresponding to each image in A; S320: Using a hand topology module to perform feature extraction processing on the images in A to obtain a hand topology map corresponding to each image in A; S330: Use an edge feature enhancement module to perform edge enhancement on the preprocessed image corresponding to the image in A to obtain an edge feature enhanced image corresponding to the image in A; S340: Repeat S310-S330, traverse all hand X-rays in A, and obtain a preprocessing map corresponding to each image in A, a hand topology map corresponding to each image in A, and an edge feature enhancement map corresponding to each image in A; S350: All pre-processed images, hand topology images and edge feature enhancement images obtained in S340 are combined into a training set of the TENet model, wherein the pre-processed image, hand topology image and edge feature enhancement image corresponding to each image form a training sample; S400: Preset the maximum number of iterations, preset the initial learning rate of D to μ, and train D: S410: taking all training samples in the training set as input of D, training D using the Adam optimizer to update the parameters of D, wherein the MAE target loss function is used in the Adam optimizer; When the maximum number of iterations is reached, the training is stopped and the trained deep learning network improved model D' is obtained; S420: Using the TENet model of D' as the trained TENet model; S500: Input a child's hand X-ray film whose unknown bone age needs to be predicted into the trained TENet model, and the output is the predicted value of the child's bone age.

2. A method for predicting bone age in children as claimed in claim 1, characterized in that: The structure of the deep learning network improved model D in S200 is to use the deep learning network as the backbone network, remove the topmost neural network layer from top to bottom, and then add a 3x3 downsampling layer, a 3x3 maximum pooling layer and a fully connected layer with 32 neurons in sequence after the bottom layer of the original backbone network, and integrate the gender information corresponding to the hand X-ray in A into the image features through the fully connected layer.

3. A method for predicting bone age in children as claimed in claim 2, characterized in that: The specific steps of obtaining the pre-processed image corresponding to the image in A in S310 are as follows: S311: Select the cth picture from A: S312: Calculate the adaptive filtering intermediate variable R of any pixel point p in the cth image c (p); Among them, p and q represent two pixels on the cth image, I c (p) and I c (q) represents the grayscale difference between p and q, d(p,q) represents the distance metric function between p and q, S α (·) represents the brightness expression function; S313: Calculate the L(p) value of the cth image. The calculation expression is as follows: Among them, L(p) represents the preprocessing result of pixel p, and the mapping interval of pixel p is [0,255], [minR c (p), maxR c (p)] represents the full domain of L(p); S314: Repeat S312-S313, traverse all pixel points on the c-th image, pre-process all pixel points in the c-th image, and obtain a pre-processed image corresponding to the c-th image.

4. A method for predicting bone age in children as claimed in claim 3, characterized in that: The specific steps of obtaining the hand topology map corresponding to the image in A in S320 are as follows: S321: Select the cth picture from A: S322: Use the horizontal matrix to perform a plane convolution with the t-th pixel on c to obtain an approximate brightness difference G of the t-th pixel in the horizontal direction xt , the specific expression is as follows: Among them, G xt represents the approximate brightness difference of the t-th pixel in the horizontal direction, c t represents the tth pixel; Use the vertical matrix to perform a plane convolution with the t-th pixel on c to obtain the approximate brightness difference G of the t-th pixel in the vertical direction yt , the specific expression is as follows: Among them, G yt Represents the approximate brightness difference of the t-th pixel in the vertical direction; S323: Calculate the gradient approximation G of the t-th pixel on c t , the expression is as follows: Calculate the direction of the t-th pixel on c. The expression is as follows: Among them, α(x t ,y t ) represents the direction of the t-th pixel on c; S324: Repeat S322-S323, traverse all the pixels on c, obtain the convolution value of each pixel, and the convolution values ​​of all the pixels form a matrix B, which is the gradient map c'; S325: performing threshold binarization processing on c' to obtain a rough foreground edge map of the hand skeleton of c; S326: Use the MediaPipe application framework to predict the positions of the joint points on the rough foreground edge map of the hand backbone corresponding to c to obtain a hand topology map of c.

5. A method for predicting bone age in children as claimed in claim 4, characterized in that: The specific steps of obtaining the edge feature enhancement map corresponding to the image in A in S330 are as follows: S331: Select the cth picture from A: S332: Calculate the spatial proximity and pixel value similarity of each pixel point on the c-th image and perform bilateral filtering denoising on the c-th image. The expression of spatial proximity is as follows: Among them, i, j, k, l represent the coordinate values ​​of pixel points a and b, namely a(i, j) and b(k, l), δ d is the Gaussian standard deviation; The expression of pixel value similarity is as follows: Among them, f(i,j) represents the pixel value of the image at point a(i,j), f(k,l) represents the pixel value of the image at point b(k,l), δ r is the Gaussian standard deviation, W r To calculate the difference between the near point a and the center point b; Suppose the image c is obtained after bilateral filtering and denoising. S333: Setting a threshold T to divide c" into background and target; The optimal pixel threshold of c" is calculated using the Otsu algorithm. The calculation formula is as follows: p0+p1=1; (9) Among them, σ 2 represents the between-class variance, represents the average gray value of c", p0 represents the proportion of the background in the picture, p1 represents the proportion of the target in the picture, and the expression is as follows: Where k = {1,…,T}, n represents the category of grayscale value, p n Represents the frequency of occurrence of each type of gray value, p0 represents the sum of the probability of occurrence of each type of gray value between 0 and k, L represents the pixel level of the image, and p1 represents the sum of the probability of occurrence of each type of gray value between k+1 and L; m0 represents the average gray value of the background, m1 represents the average gray value of the target, and the expression is as follows: S334: Substitute formula (10) into formula (11) and simplify to obtain formula (16), which is as follows: σ 2 =p0p1(m0-m1) 2 ; (16) S335: Traverse all pixel points on c'', and calculate the pixel points with the same gray value of category on c'' in turn through formula (9)-formula (16) to obtain N σ 2 ; S336: The obtained N σ 2 Sort in descending order and select the largest value of σ 2 As the pixel optimal threshold of c”; S337: Use c" processed by 336 as the edge feature enhancement image corresponding to the c-th image.