Bone age assessment method and system combining deep learning and logic correction segmentation

By employing deep learning methods such as multi-scale candidate box generation, confidence ranking, semi-supervised pseudo-label generation, and global attention fusion, the problems of high false detection rate and poor consistency in joint detection and classification in bone age assessment systems are solved, achieving high-precision and reliable bone age assessment.

CN120977590APending Publication Date: 2025-11-18TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH +1
View PDF 17 Cites 0 Cited by

Patent Information

Application Number
CN202511486716.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing automated bone age assessment systems have high false positive rates and poor localization consistency when detecting bone joint areas in the hand. The epiphyseal maturity classification model is not sensitive enough to subtle morphological differences, resulting in low accuracy and reliability of bone age assessment.

Method used

We employ a method that combines deep learning with logical correction segmentation. By using multi-scale candidate box generation, confidence ranking, semi-supervised pseudo-label generation, and fusion of linear discriminant analysis and normalized global attention, we can improve the accuracy and consistency of joint detection and enhance the feature separability of epiphyseal maturity classification.

Benefits of technology

It significantly improves the accuracy and consistency of key joint detection for bone age, enhances the robustness of epiphyseal maturity classification, and improves the accuracy and reliability of bone age assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977590A_ABST
    Figure CN120977590A_ABST
Patent Text Reader

Abstract

The invention relates to a deep learning and logic correction segmentation combined bone age assessment method, which aims at the problems of high joint extraction false drop rate, difficulty in epiphysis maturity grade classification and the like, through a logic correction segmentation mechanism of confidence ranking, a semi-supervised pseudo-label generation strategy and a global attention enhancement module based on normalization, the bone age is assessed. And the accuracy and clinical robustness of bone age evaluation are remarkably improved. The application scenarios of the method include but are not limited to the precise medical fields of children endocrine dysplasia screening, bone age lag or advance pathological diagnosis, orthopedic treatment scheme making, growth hormone intervention curative effect evaluation and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of medical image processing and computer vision, specifically to a bone age assessment method and system that combines deep learning and logical correction segmentation. Background Technology

[0002] Bone age assessment is a crucial diagnostic tool in pediatric endocrinology, orthopedics, and growth and development management. It provides key references for predicting children's height, screening for developmental abnormalities, diagnosing endocrine disorders, and developing treatment plans. By analyzing skeletal maturity in hand X-rays, doctors can determine the difference between skeletal development level and actual age, thereby assessing growth potential and disease risk.

[0003] Traditional Greulich-Pyle (GP) atlas and Tanner-Whitehouse (TW) scoring methods are widely used in clinical practice, but they rely on manual image interpretation and experience-based judgment, resulting in high subjectivity, low efficiency, and poor consistency among different operators. With the development of deep learning and medical image processing technologies, automated bone age assessment systems have gradually become a research hotspot. These systems not only help improve diagnostic efficiency and objectivity but also provide a scalable technological path for large-scale population screening and precision medicine.

[0004] Existing automated bone age assessment systems suffer from high false positive rates and poor localization consistency during the detection phase of the hand bone and joint regions. This leads to instability in the extraction of key joints required for bone age scoring, especially when the epiphyseal boundaries are blurred or the image quality is poor, resulting in missed and false positives. Existing epiphyseal maturity classification models are not sensitive enough to subtle morphological differences and have difficulty effectively distinguishing the feature distribution of adjacent maturity stages, resulting in decreased classification accuracy. The independent optimization of the detection and classification modules lacks cross-stage information interaction and collaborative feedback mechanisms, causing early detection errors to be continuously amplified during the scoring and bone age fitting process, ultimately affecting the accuracy and clinical reliability of the overall assessment. Summary of the Invention

[0005] Based on the above description, the present invention provides a bone age assessment method and system that combines deep learning and logical correction segmentation to solve the problems of missed detections and false detections and low reliability in bone age assessment.

[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:

[0007] A bone age assessment method combining deep learning and logical correction segmentation includes the following steps:

[0008] S1. Acquire medical images of the hand, detect the bones and joints in the medical images of the hand, and filter and extract the images of the scoring joints. The extracted images of the scoring joints are the sample images.

[0009] S2. To filter and classify sample images and generate pseudo-labels based on the type of sample images;

[0010] S3. Build a bone age assessment model. The bone age assessment model has a preset bone age curve, which reflects the correspondence between the scores of each scoring joint and the actual bone age. The bone age assessment model is used to score the scoring joints in the input sample image. After obtaining the scores of each scoring joint, the bone age is obtained by combining the scores with the bone age curve.

[0011] As a preferred embodiment, step S1 includes:

[0012] S101, Multi-scale candidate generation for joint regions: Based on the bone age scoring standard, hand joints are divided into seven categories, and 21 candidate box generation strategies are designed to cover all scoring-related regions.

[0013] S102, Confidence Ranking and Logical Filtering: Rank each candidate box by confidence, select the detection boxes with higher confidence rankings, and combine them with spatial position relationships to filter out the final 13 scoring joints.

[0014] As a preferred approach, step S2 introduces a pseudo-label generation strategy based on a semi-supervised model, which includes the following specific steps:

[0015] S201. Construct a semantic description library for epiphyses: Transform 105 types of epiphyseal maturity stages into language descriptions, which will serve as text input for the language-image contrast pre-training model.

[0016] S202, Pseudo-label generation and screening: Comparative learning is performed on labeled and unlabeled data to generate pseudo-labels, and high-quality samples are screened through confidence thresholds and consistency regularization.

[0017] As a preferred approach, step S3 involves building a discriminative classifier that fuses linear discriminant analysis with a normalization-based global attention module, including the following specific steps:

[0018] S301, Normalized Channel Attention Module: Channel weight graphs are constructed using the γ parameter of the BN layer, enabling channel recalibration without the need for additional parameters. Equations (1) and (2) illustrate the process of generating weights from the BN layer. (1) (2)

[0019] In the formula, It is the input to the BN layer, that is, a mini-batch of data. This is the output of the Batch Normalization (BN) layer, which is the result after normalization, scaling, and translation. It is the mean of the mini-batch. It is a constant. It is a learnable translation factor. It is the variance of BN. It is the scaling factor for each channel. It is a normalized value. Indicates a channel;

[0020] S302. Fusion of spatial attention and channel attention to construct a global attention module: A global attention module is inserted into the bottleneck structure of ResNet50 to enhance the response in the epiphyseal region;

[0021] S303, Discriminative Dimensionality Reduction and Classifier Construction: The high-dimensional features after ResNet50 are input into the discriminative classifier, and nine epiphyseal specialist classifiers are trained respectively, and finally the overall bone age score is fitted.

[0022] A bone age assessment system combining deep learning and logical correction segmentation, for performing the method, comprising:

[0023] The joint detection module is used to detect bones and joints in medical images of the hand and to filter and extract images of scoring joints. The extracted images of scoring joints are the sample images.

[0024] The sample screening module is used to screen and classify sample images and generate pseudo-labels based on the type of sample image;

[0025] The bone age assessment model has a preset bone age curve, which reflects the correspondence between the scores of each scoring joint and the actual bone age. The bone age assessment model is used to score the scoring joints in the input sample image, and the scores of each scoring joint are combined with the bone age curve to obtain the assessed bone age.

[0026] Compared with the prior art, the technical solution of this application has the following beneficial technical effects:

[0027] This method significantly improves the accuracy and consistency of key joint detection in bone age assessment by employing confidence-ranked logical calibration joint extraction, a cross-domain semi-supervised pseudo-label generation strategy, and a fusion mechanism of linear discriminant analysis and normalized global attention. It also enhances the feature separability and robustness of epiphyseal maturity classification. This method can also be extended to other medical imaging scenarios requiring fine-grained structure recognition and stage-based grading, providing an efficient and scalable technical solution for anatomical structure localization, developmental stage assessment, and quantitative analysis of pathological processes in multimodal imaging such as X-ray, MRI, and CT. Attached Figure Description

[0028] Figure 1This is a schematic diagram of a preferred embodiment of a bone age assessment method combining deep learning and logical correction segmentation provided by the present invention.

[0029] Figure 2 A schematic diagram of the confidence ranking and logical correction rules for the logical correction segmentation method;

[0030] Figure 3 Here is a diagram of the structure of the normalized global attention mechanism;

[0031] Figure 4 This is a schematic diagram of a preferred structure for a discriminant classifier that combines linear discriminant analysis and global attention. Detailed Implementation

[0032] Example 1:

[0033] Reference Figures 1 to 4 A bone age assessment method combining deep learning and logical correction segmentation includes the following steps:

[0034] S1. Acquire medical images of the hand, detect the bones and joints in the medical images of the hand, and filter and extract images of the scoring joints. The extracted images of the scoring joints are the sample images.

[0035] In this step, a logically calibrated multi-scale joint segmentation framework is first constructed. This framework is used to identify hand joints and segment and extract images of hand joint regions. Local-global anatomical clues are extracted using a multi-scale backbone, and a logical correction strategy based on confidence ranking is proposed to replace non-maximum suppression and one-to-one head post-processing. Twenty-one candidate boxes are first generated according to seven morphological categories, and then 13 joint regions required for bone age scoring are selected by combining spatial topological relationship constraints, thereby suppressing false detections and improving localization consistency from the source.

[0036] The specific steps for building the joint detection module are as follows:

[0037] S101. Multi-scale candidate generation of joint regions: Based on the bone age scoring standard, the hand joints are divided into seven categories: 1) distal phalanges; 2) middle phalanges; 3) proximal phalanges; 4) metacarpals; 5) first metacarpal; 6) ulna; 7) radius.

[0038] For the seven types of joint regions mentioned above, a candidate box generation strategy was designed: for the middle phalanges and metacarpals, four candidate boxes were generated at the positions of the four fingers excluding the thumb, for a total of eight candidate boxes; for the distal phalanges and proximal phalanges, five candidate boxes were generated at the positions of the five fingers, for a total of ten candidate boxes; one candidate box was generated for the first metacarpal bone, for a total of one candidate box; and one candidate box was generated for the ulna and radius, for a total of two candidate boxes. This resulted in a total of 21 candidate boxes generated.

[0039] The candidate box generation strategy is based on the following: by using multi-scale and multi-location candidate boxes to cover all bone age scoring-related areas, the detection network can capture the target joint area under different hand sizes and imaging conditions, thereby improving the robustness and accuracy of the detection.

[0040] S102. Confidence Ranking and Logical Selection: The confidence of each candidate box is determined by the classification probability (softmax or sigmoid output value) output by the detection network, representing the credibility of the candidate box belonging to the target joint category. Confidence ranking is performed for each category of candidate boxes. In the logical selection stage, spatial positional constraints are introduced. Based on prior knowledge of hand anatomy, the arrangement order and spatial position of joints are ensured to be reasonable. For multiple candidate boxes of the same category, if their positions overlap or deviate from the reasonable anatomical range, only candidate boxes that match the overall skeletal position and have a higher confidence ranking are retained. Through positional constraint, 13 joint regions that meet the bone age scoring criteria are finally selected from 21 candidate boxes.

[0041] The joint regions specifically include: the distal phalanges of the first, third, and fifth fingers, totaling 3; the middle phalanges of the third and fifth fingers, totaling 2; the proximal phalanges of the first, third, and fifth fingers, totaling 3; the metacarpals of the first, third, and fifth fingers, totaling 3; and one each of the ulna and radius, totaling 2; for a total of 13.

[0042] S2. Build a sample screening module to screen and classify sample images and generate pseudo-labels based on the type of sample image.

[0043] Specifically, a pseudo-label generation strategy based on a semi-supervised model is introduced: using a language-image contrast pre-trained model as the teacher, the semantics are bridged between labeled and unlabeled clinical data through multiple rounds of self-training, and dynamic threshold screening and consistency regularization are performed on high-confidence samples to generate robust pseudo-labels to supervise subsequent grading tasks, reduce label dependence and enhance cross-domain generalization.

[0044] The introduction of a pseudo-label generation strategy based on a semi-supervised model includes the following specific steps:

[0045] S201. Construct a semantic description library for epiphyses: Transform 105 types of epiphyseal maturity stages into language descriptions, which will serve as text input for the language-image contrast pre-training model.

[0046] S202, Pseudo-label generation and filtering: Comparative learning is performed on labeled and unlabeled data to generate pseudo-labels for the unlabeled data. "Labeled data" refers to hand X-ray images and their corresponding joint region labels and bone age scores that have been professionally or manually annotated; "unlabeled data" refers to samples containing only hand X-ray images without accompanying manual annotation information.

[0047] To ensure the reliability of pseudo-labels, high-quality samples are screened using confidence thresholds and consistency regularization. Specifically: 1) Confidence screening: If the detection network's predicted probability (e.g., softmax output value) for a candidate box is higher than a preset threshold (e.g., 0.9), the pseudo-label is considered reliable; if it is lower, it is discarded. For example, if the model predicts a probability of 0.95 for a joint region, the pseudo-label is retained; if it is only 0.6, it is removed. 2) Consistency screening: The prediction results of the same unlabeled sample under different data augmentations (e.g., small-amplitude rotation, brightness perturbation) are compared. If the prediction results remain consistent in spatial location and category, the pseudo-label is considered stable and reliable; if the differences are large, it is discarded. For example, if the same X-ray film still detects the same 13 joint regions after horizontal flipping, the pseudo-label passes the consistency screening.

[0048] S3. Build a bone age assessment model. The bone age assessment model has a preset bone age curve, which reflects the correspondence between the scores of each scoring joint and the actual bone age. The bone age assessment model is used to score the scoring joints in the input sample image. After obtaining the scores of each scoring joint, the bone age is obtained by combining the scores with the bone age curve.

[0049] In this step, a discriminant classifier that integrates linear discriminant analysis and a normalization-based global attention module is built: in a convolutional network enhanced by the normalization-based global attention module, channel and spatial attention are jointly used to recalibrate feature responses, and supervised dimensionality reduction using linear discriminant analysis is performed to improve inter-class separability. Specialized classifiers are trained according to epiphyseal locations to determine maturity levels. Finally, the scores of each joint are fitted to the overall bone age through the maturity score-bone age curve, realizing precise modeling of the entire chain from joint extraction to maturity assessment.

[0050] Building a discriminative classifier that fuses linear discriminant analysis with a normalization-based global attention module involves the following specific steps:

[0051] S301, Normalized Channel Attention Module: Channel weight graphs are constructed using the γ parameter of the BN layer, enabling channel recalibration without the need for additional parameters. Equations (1) and (2) illustrate the process of generating weights from the BN layer. (1) (2)

[0052] In the formula, The input to the BN layer (a mini-batch of data). The output of the BN layer (the result after normalization, scaling, and translation). The mean of this mini-batch. A very small constant used to avoid the denominator being zero. Learnable shift parameter, each channel has an independent one. , It is the variance of BN. It is the scaling factor for each channel. It is a normalized value. Indicates a channel.

[0053] The generated weights are used to weight each channel of the input feature map, that is, to multiply the channel weights by the feature values ​​of the corresponding channels, thereby achieving channel-level recalibration.

[0054] S302. Fusion of spatial attention and channel attention to construct a global attention module: A global attention module is inserted into the bottleneck structure of ResNet50 to enhance the response in the epiphyseal region.

[0055] S303, Discriminative Dimensionality Reduction and Classifier Construction: The high-dimensional features after ResNet50 are input into the discriminative classifier, and nine epiphyseal specialist classifiers are trained respectively, and finally the overall bone age score is fitted.

[0056] According to the above implementation process, it can be combined with Figure 1 The working principle of this invention can be summarized as follows:

[0057] 1. Acquire a dataset of hand X-ray images, including 1,000 RSNA public images and 9,577 clinical images. After normalization and size adjustment, input the data into the logic calibration multi-scale detection module to extract 13 key joint regions required for bone age scoring.

[0058] 2. In Figure 2 The improved model shown is used for joint detection and epiphyseal maturity classification training and prediction. The improvements include introducing a confidence ranking-driven logical calibration mechanism to replace the traditional NMS and one-to-one head strategy, combining a semi-supervised model (i.e., the semi-supervised language-image contrast pre-trained model CLIP) to generate high-quality pseudo-labels, embedding a normalized global attention module (NGAM) into the ResNet50 backbone to enhance feature discrimination, and performing supervised dimensionality reduction through linear discriminant analysis (LDA) to improve inter-class separability.

[0059] Thus, an automatic bone age assessment model with high accuracy and low false positive rate under complex imaging conditions was obtained, which can be widely used in clinical diagnosis and developmental monitoring in pediatric endocrinology and orthopedics.

[0060] To further illustrate the experimental effects of the present invention, the experimental results are now presented:

[0061] The experiments were conducted on an Ubuntu 24.04.1 LTS system using an NVIDIA RTX 3090 (24G) GPU, implemented using PyTorch 2.4.1, Python 3.10, and CUDA 12.4. The detection module was trained for 200 epochs on pre-trained COCO weights using a momentum optimizer (momentum=0.9, weight decay=0.05) with an initial learning rate of 0.001, decaying at epochs 50 and 150. The classification module was trained for 120 epochs using an Adam optimizer (initial learning rate 1e-5, batch size=64), with pseudo-label samples introduced in three rounds of semi-supervised training.

[0062] Following the experimental procedures, tests were conducted on a mixed RSNA and clinical dataset, and the evaluation metrics are shown in the table below:

[0063]

[0064] As can be seen, the detection rate, classification accuracy and bone age prediction accuracy of the present invention are significantly better than existing methods on public datasets and large-scale clinical images, which verifies the robustness and clinical feasibility of the present invention in bone age assessment tasks.

[0065] Example 2:

[0066] A bone age assessment system combining deep learning and logical correction segmentation includes:

[0067] The joint detection module is used to detect bones and joints in medical images of the hand and to filter and extract images of scoring joints. The extracted images of scoring joints are the sample images.

[0068] The sample screening module is used to screen and classify sample images and generate pseudo-labels based on the type of sample image.

[0069] The bone age assessment model has a preset bone age curve, which reflects the correspondence between the scores of each scoring joint and the actual bone age. The bone age assessment model is used to score the scoring joints in the input sample image, and the scores of each scoring joint are combined with the bone age curve to obtain the assessed bone age.

[0070] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A bone age assessment method combining deep learning and logical correction segmentation, characterized in that, Includes the following steps: S1. Acquire medical images of the hand, detect the bones and joints in the medical images of the hand, and filter and extract the images of the scoring joints. The extracted images of the scoring joints are the sample images. S2. To filter and classify sample images and generate pseudo-labels based on the type of sample images; S3. Build a bone age assessment model. The bone age assessment model has a preset bone age curve, which reflects the correspondence between the scores of each scoring joint and the actual bone age. The bone age assessment model is used to score the scoring joints in the input sample image. After obtaining the scores of each scoring joint, the bone age is obtained by combining the scores with the bone age curve.

2. The bone age assessment method combining deep learning and logical correction segmentation according to claim 1, characterized in that, Step S1 includes: S101, Multi-scale candidate generation for joint regions: Based on the bone age scoring standard, hand joints are divided into seven categories, and 21 candidate box generation strategies are designed to cover all scoring-related regions. S102, Confidence Ranking and Logical Filtering: Rank each candidate box by confidence, select the detection boxes with higher confidence rankings, and combine them with spatial position relationships to filter out the final 13 scoring joints.

3. The bone age assessment method combining deep learning and logical correction segmentation according to claim 1, characterized in that, Step S2 introduces a pseudo-label generation strategy based on a semi-supervised model, which includes the following specific steps: S201. Construct a semantic description library for epiphyses: Transform 105 types of epiphyseal maturity stages into language descriptions, which will serve as text input for the language-image contrast pre-training model. S202, Pseudo-label generation and screening: Comparative learning is performed on labeled and unlabeled data to generate pseudo-labels, and high-quality samples are screened through confidence thresholds and consistency regularization.

4. The bone age assessment method combining deep learning and logical correction segmentation according to claim 1, characterized in that, Step S3 involves building a discriminative classifier that fuses linear discriminant analysis with a normalization-based global attention module. This includes the following specific steps: S301, Normalized Channel Attention Module: Channel weight graphs are constructed using the γ parameter of the BN layer, enabling channel recalibration without the need for additional parameters. Equations (1) and (2) illustrate the process of generating weights from the BN layer. (1) (2) In the formula, It is the input to the BN layer, that is, a mini-batch of data. This is the output of the Batch Normalization (BN) layer, which is the result after normalization, scaling, and translation. It is the mean of the mini-batch. It is a constant. It is a learnable translation factor. It is the variance of BN. It is the scaling factor for each channel. It is a normalized value, where i represents the channel; S302. Fusion of spatial attention and channel attention to construct a global attention module: A global attention module is inserted into the bottleneck structure of ResNet50 to enhance the response in the epiphyseal region; S303, Discriminative Dimensionality Reduction and Classifier Construction: The high-dimensional features after ResNet50 are input into the discriminative classifier, and nine epiphyseal specialist classifiers are trained respectively, and finally the overall bone age score is fitted.

5. A bone age assessment system combining deep learning and logical correction segmentation, used to perform the method described in any one of claims 1-4, characterized in that, include: The joint detection module is used to detect bones and joints in medical images of the hand and to filter and extract images of scoring joints. The extracted images of scoring joints are the sample images. The sample screening module is used to screen and classify sample images and generate pseudo-labels based on the type of sample image; The bone age assessment model has a preset bone age curve, which reflects the correspondence between the scores of each scoring joint and the actual bone age. The bone age assessment model is used to score the scoring joints in the input sample image, and the scores of each scoring joint are combined with the bone age curve to obtain the assessed bone age.

Citation Information

Patent Citations

  • Image processing method and device, equipment storage medium and growth and development evaluation system

    CN109949280A

  • Bone age evaluation method and device based on annotation noise correction

    CN111681203A

  • Rolling bearing fault diagnosis method fusing attention mechanism and twin network structure

    CN113191215A

  • Improved YOLOv4 network model and small target detection method

    CN114663654A

  • Hand bone age automatic evaluation and correction method based on TW3 skeletal development association rule

    CN114822837A