Age estimation method and system based on elasticity law and non-uniformly distributed X-ray human skeleton image
By introducing elasticity law-inspired biased optimization and many-to-many supervised learning into X-ray human skeleton image datasets, the accuracy and stability issues of age estimation on non-uniformly distributed datasets are solved, achieving effective representation of a small number of samples and alignment with clinical needs, thus improving the performance and interpretability of the age estimation model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIDIAN UNIV
- Filing Date
- 2026-02-03
- Publication Date
- 2026-05-15
AI Technical Summary
Existing automated age estimation methods struggle to improve the accuracy and stability of age estimation for a small number of samples when faced with non-uniformly distributed X-ray human skeleton image datasets. Furthermore, they lack prior knowledge integration for specific clinical scenarios and treatment guidelines, which limits the application of these models in real-world medical tasks.
By introducing biased optimization inspired by the elasticity law, the supervision information of samples and labels is expanded to many-to-many mutual supervision between samples. The model uses supervisory force, distance elasticity and midpoint elasticity for sample comparison learning, establishes feature queues and monotonic constraint residual blocks, and improves the model's performance on a small number of samples.
It significantly improves the age estimation performance for a small number of samples, enhances the smoothness, continuity and predictive consistency of the model, and can more accurately represent the essential change patterns of human biological age.
Smart Images

Figure CN122048885A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing, and specifically relates to an age estimation method for non-uniformly distributed X-ray human skeleton images, which can be used for disease risk assessment, personalized health management and medical research. Background Technology
[0002] Age estimation is a crucial task in predicting the physiological or pathological aging process of an individual through biomarkers. Its core value lies in characterizing age-related changes in the human body, providing quantitative evidence for disease risk assessment and personalized health management. Compared to traditional aging markers such as blood and genes, medical imaging, including MRI, CT, and X-rays, can capture multidimensional aging characteristics of the entire body's organ systems in a non-invasive and high spatial resolution manner, and has been proven to have a strong correlation with various age-related diseases such as Alzheimer's disease, osteoporosis, and cardiovascular disease. Medical data, such as X-rays, exhibit significant temporal sequence characteristics, determined by the natural characteristic of continuous changes in the human body's physiological state over time. Furthermore, over longer timescales spanning months, years, or even the entire lifespan, long-term longitudinal imaging, such as X-ray sequences, can reveal the long-term, gradual evolution of specific diseases or the stage-specific characteristics of normal human development. This temporal sequence characteristic has become an important modeling factor for age estimation using deep representation learning of medical imaging data.
[0003] In medical settings, the imbalanced distribution of data significantly impacts model performance, posing challenges to practical applications. This is because, from a medical perspective, models maintain good performance even with limited sample sizes, such as in rare disease diagnosis, in-hospital mortality prediction, and tumor detection and classification. In age estimation tasks, the non-uniform distribution of samples constitutes a major challenge for deep learning methods.
[0004] Previous research has extensively discussed the class imbalance problem in classification tasks. In these studies, methods to mitigate data imbalance are generally categorized into data-centric and algorithm-centric approaches. Data-centric approaches typically involve sampling and preprocessing of the dataset, including downsampling of the majority samples and upsampling or generation of the minority samples. Algorithm-centric approaches focus on enhancing the learning of minority samples during training to balance the model's performance on these samples. Specifically, these techniques include improving the loss function, setting cost biases, and increasing model depth.
[0005] Unlike the discrete task labels in the classification tasks mentioned above, age estimation is a continuous mapping from the image data space to the age label space. In recent years, strategies for deep imbalanced regression (DIR) in age estimation have been proposed. For example, Yang et al.'s 2021 paper at the International Conference on Machine Learning, "Delving into deep imbalanced regression," emphasized that the main factor affecting performance in DIR is the inconsistency between the representation space and the label space. In subsequent work, researchers have used this theoretical foundation to construct sample pairs based on sample label distances, and then use contrastive learning strategies to optimize the distance between samples in the feature space.
[0006] Accurate and interpretable age estimation models are crucial for characterizing human aging. Intuitive modeling from the aging representation space to the age label space enables researchers and clinicians to understand the underlying factors influencing human physiological age. Physical Information Neural Networks (PINNs), as a technique to enhance the interpretability of deep learning methods, can incorporate prior knowledge and physical laws derived from data observations, physical laws, domain expert knowledge, and semantic information into deep learning models, thereby increasing model interpretability. Research has incorporated this prior knowledge into the architectural design and training process optimization of deep learning models. This prior knowledge, explicitly or implicitly introduced into deep learning models as conditional constraints, can effectively improve model interpretability.
[0007] Patent document CN112950631A discloses an age estimation method based on saliency map constraints and X-ray cephalometric images. It uses a saliency map generated by a trained network to constrain the saliency region of the X-ray cephalometric images, providing global information for age estimation. This reduces the variance of the absolute error value for age estimation of X-ray cephalometric images of different ages, improving the accuracy and stability of age estimation using X-ray images. However, because this method does not perform additional optimization for a small number of samples, it is difficult to accurately characterize the aging features of a few samples with non-uniform distribution, thus limiting the accuracy and stability of its age estimation for a small number of samples.
[0008] In summary, existing automated age estimation methods struggle to overcome the model calibration problems in real-world medical scenarios caused by imbalanced data sample distribution. Furthermore, current deep learning models for age estimation lack effective integration of prior knowledge from specific clinical scenarios and treatment guidelines, resulting in insufficient interpretability of the human aging process, thus hindering the application of age estimation models in practical medical tasks. Summary of the Invention
[0009] The purpose of this invention is to address the shortcomings of the prior art by proposing an age estimation method and system based on the law of elasticity and non-uniformly distributed X-ray human skeletal images, so as to improve the accuracy and stability of age estimation on non-uniformly distributed human X-ray image datasets and achieve a more accurate characterization of the essential change law of human biological age in X-ray images.
[0010] The technical approach of this invention is to improve the overall performance of model age estimation, especially the performance of minority samples, by extending the "one-to-one" supervision information between samples and labels to a "many-to-many" mutual supervision. This is achieved by introducing biased optimization inspired by the laws of motion, allowing for sufficient optimization of minority samples.
[0011] Based on the above approach, the implementation scheme of the present invention includes:
[0012] 1. An age estimation network based on the law of elasticity and non-uniformly distributed X-ray images of human skeletons, characterized in that it comprises:
[0013] A feature extractor is built to extract batch features from the input image to construct the distance feature matrix and the midpoint feature matrix;
[0014] A feature queue is introduced, which uses supervision, distance elasticity and midpoint elasticity to work on the samples and perform comparative learning between samples to obtain quality-weighted biased optimization.
[0015] Using supervised force to define the basic age estimation loss Obtain the total loss of network iterative training. ;
[0016] Establish monotonic constrained residual blocks MCRes to map the feature distance between sample pairs to the partially monotonically increasing distance feature PMF between samples;
[0017] Establish mutually independent age estimation regression heads to output age estimation results, sample pair midpoints, sample pair distances, and age estimation results with monotonicity constraints;
[0018] The feature extractor, feature queue, monotonic constraint residual block, and age estimation regression head are cascaded in sequence to form an age estimation network for non-uniformly distributed X-ray human skeleton images.
[0019] Furthermore, the introduction of a feature queue is used to expand the number of sample pairs, and the samples are processed using supervised force, distance elasticity, and midpoint elasticity to perform temporal comparative learning between samples, in order to obtain quality-weighted biased optimization, which includes:
[0020] For a given batch of size B, by one-to-one pairing of samples, we obtain... 1 sample pair, with an introduced length of 1 Feature queue It follows a first-in, first-out (FIFO) update strategy to cache batch features of samples during training iterations, thus reducing the number of sample pairs from... Expand to ;
[0021] Construct a working mode that combines supervisory force, distance elasticity, and midpoint elasticity;
[0022] The above three forces work together to achieve temporal comparative learning between two samples in a sample pair. The "elastic force" comes from the difference between the label distance and the estimated distance. When the estimated distance is less than the actual distance, the elastic potential energy is released through the elastic force, generating a repulsive force between the samples. Conversely, it generates a pulling force.
[0023] Define sample The quality is The definition of sample quality Replace with medium sample Number of samples in the age group;
[0024] A biased optimization expression for mass weighting is obtained by combining sample quality and sample stress conditions.
[0025] 2. A method for age estimation of X-ray human skeletal images using the above-mentioned age estimation network, characterized in that it includes:
[0026] (1) Two datasets, lateral cephalometric radiographs covering the craniofacial, dental, and cervical spine regions and panoramic tomographic radiographs covering the dental and mandibular regions, obtained from medical institutions and aged 4 to 50 years, were used to divide the training sample set D in a 6:2:2 ratio. train Validation sample set D val Test sample set D test :
[0027] (2) Set the hyperparameters of the age estimation network, including: the number of neurons in the fully connected layers of the three age regression heads, the number of neurons in the three fully connected layers of the age regression heads with monotonicity constraints, the custom parameter γ for the total loss, the batch size B, and the length of the feature queue. Number of training iterations, learning rate ;
[0028] (3) The training sample set D train Input the age estimation network, perform iterative training and validation, and obtain a trained age estimation network model;
[0029] (4) Test sample set D test The input is fed into a trained age estimation network model, which outputs the predicted age.
[0030] Compared with the prior art, the present invention has the following advantages:
[0031] First, it significantly improves the age estimation performance for a small number of sample data.
[0032] This invention introduces the concept of sample quality and combines it with the idea of elastic potential energy in physics. It achieves biased optimization of a small number of samples by using a constraint method in which supervisory force, distance elastic force, and midpoint elastic force work together. This improves the effectiveness of age estimation methods on a small number of samples and is more in line with actual clinical needs.
[0033] Second, it has the ability to characterize aging-related changes with smoothness, continuity, and predictive consistency.
[0034] This invention improves the continuity, stationarity, and consistency of age estimation results by using distance representation between samples to establish the association between dataset samples and using label distance between samples to construct a physical information neural network for supervised temporal comparison learning, in uniform manifold approximation and projected UMAP space. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the age estimation network structure of the present invention;
[0036] Figure 2 This is a schematic diagram of the force analysis of the sample when elastic work is performed in the age estimation network of this invention;
[0037] Figure 3 This is a schematic diagram of the monotonic constraint residual block MCRes module structure in the age estimation network of this invention;
[0038] Figure 4 This is the implementation flow of the age estimation method based on an age estimation network according to the present invention;
[0039] Figure 5 These are X-ray lateral cephalometric radiographs and curved tomographic images used in the method of this invention;
[0040] Figure 6 The age and sex distribution of X-ray lateral cephalometric radiographs and curved tomographic images used in the method of this invention;
[0041] Figure 7 This is a performance improvement diagram of using the present invention to estimate age in different age ranges of lateral cephalometric radiographs;
[0042] Figure 8This is a performance improvement diagram of using the present invention to estimate the age of different age ranges of curved tomographic slice data;
[0043] Figure 9 This is a diagram illustrating the characterization effect of the present invention on aging changes in image data. Detailed Implementation
[0044] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention and not all of them. Based on the embodiments of the present invention, other embodiments obtained by those skilled in the art without creative effort should all fall within the protection scope of the present invention.
[0045] Example 1: Age estimation network based on elasticity law and non-uniformly distributed X-ray human skeleton images.
[0046] Reference Figure 1 This example includes: a feature extraction module, a feature queue, monotonic constraint residual blocks (MCRes), and an age estimation regression head. The structure and function of each module are as follows:
[0047] The feature extractor includes an initial convolutional unit, a residual unit, and a global pooling unit connected in sequence. The structure and connection relationship of each unit are as follows:
[0048] The initial convolutional unit includes a 7×7 convolutional layer with a stride of 2, a batch normalization layer, an activation layer, and a 3×3 max pooling layer, which are used to convert image information into high-dimensional information and transmit it to the residual unit.
[0049] The residual unit comprises four cascaded residual modules, used to extract features from the input high-dimensional information, generate a feature map, and transmit it to the global pooling unit, wherein:
[0050] The first residual module includes three cascaded residual blocks, and the feature map size of each residual block remains unchanged.
[0051] The second residual module includes four cascaded residual blocks, where the first residual block is downsampled and the subsequent residual blocks maintain the same feature map size.
[0052] The third residual module includes six cascaded residual blocks, where the first residual block is downsampled and subsequent residual blocks maintain the same feature map size.
[0053] The fourth residual module includes three cascaded residual blocks, where the first residual block is downsampled and the subsequent residual blocks maintain the same feature map size.
[0054] Each residual block consists of three convolutional layers: the first convolutional layer includes a 1×1 convolution, a batch normalization layer, and an activation layer; the second convolutional layer includes a 3×3 convolution, a batch normalization layer, and an activation layer; and the third convolutional layer includes a 1×1 convolution and a batch normalization layer. The input of the first convolutional layer and the output of the third convolutional layer are added element-wise, and the result after activation is the output of the residual block.
[0055] This global pooling unit includes a global average pooling layer, which performs pooling processing on the input feature map in the spatial dimension, compressing the feature map of each channel into a single feature value.
[0056] The feature queue is used to expand the number of sample pairs and to perform supervised learning, distance elasticity, and midpoint elasticity on the samples, conducting temporal comparative learning between samples to obtain quality-weighted biased optimization, referencing... Figure 2 Its specific implementation is as follows:
[0057] For a given batch of size B, by one-to-one pairing of samples, we obtain... 1 sample pair, with an introduced length of 1 Feature queue It follows a first-in, first-out (FIFO) update strategy to cache batch features of samples during training iterations, thus reducing the number of sample pairs from... Expand to ;
[0058] Building Supervisory Power Distance elasticity Midpoint elasticity The three-way collaborative work model, in which, Indicates sample Age tags, Indicates corresponding to The age regression results This represents the label distance between two samples in a sample pair. For the corresponding estimated distance, The midpoint of the labels of the two samples. The three forces mentioned above are used as the loss function to optimize the model in subsequent learning, with the midpoint of the corresponding estimation results as the midpoint.
[0059] The above three forces work together to achieve temporal comparative learning between two samples in a sample pair. The "elastic force" comes from the difference between the label distance and the estimated distance. When the estimated distance is less than the actual distance, the elastic potential energy is released through the elastic force, generating a repulsive force between the samples. Conversely, it generates a pulling force.
[0060] Define sample The mass is: ,in The number of age groups. For the sample Number of samples in the age group and This represents the mean of all age labels within the corresponding age group;
[0061] In the definition of sample quality Replace with medium sample Number of samples in the age group;
[0062] The biased optimization expression for mass weighting, which combines sample mass and sample stress conditions, is as follows:
[0063]
[0064] in, To estimate the distance loss for perceiving the temporal nature of development and aging processes, Used to select whether to use midpoint elastic force; when the distance elastic force on the sample is in the same direction as the tag supervision force. =1, otherwise =0.
[0065] Using supervised force to define the basic age estimation loss Obtain the total loss of network iterative training. The formula is as follows:
[0066]
[0067]
[0068] Where B is the batch size. The loss function of age regression results under monotonic constraints, representing the supervised force on the samples. ,in For the sample Age tags, The results of age regression with monotonicity constraints are as follows. For custom parameters, The loss function to achieve the monotonically increasing constraint.
[0069] The monotonic constrained residual blocks MCRes are used to obtain partially monotonically increasing distance features between samples. , refer to Figure 3 Its implementation includes:
[0070] First, pair the samples Feature distance ,in and Samples and The age-related changes are mapped as Partially monotonically increasing distance features: ; ,
[0071] in, For The mapping of weights in a multilayer sensing mechanism uses a loss function. Implement the monotonically increasing constraint;
[0072] Then the and Adding them together yields a partially monotonically increasing distance feature between samples. :
[0073] .
[0074] The age estimation regression head includes ψ d The distance regression head of the sample pairs with parameters , with ψ m The midpoint regression head of the sample pair with parameters Regression head of sample age with parameter ϕ Age regression head with monotonicity constraint ,in:
[0075] This sample is related to the distance regression head. Used to obtain the estimated distance between sample pairs: ;
[0076] This sample is a midpoint regression head Used to obtain the midpoint of the sample pair estimation results: ;
[0077] This sample's age regression head Used to obtain tag age estimates: ;
[0078] The age regression head with monotonicity constraints This is used to obtain age regression results with monotonicity constraints: ,
[0079] in, Batch features, i.e., samples The characteristics of age-related changes Features, i.e., samples, are cached for the feature queue. The characteristics of age-related changes ;
[0080] The above regression head , and Each layer is built using a fully connected layer, and the regression head is formed by cascading three fully connected layers. .
[0081] Example 2: A method for age estimation of X-ray human skeleton images using the age estimation network.
[0082] Reference Figure 4 The implementation steps of this example include the following:
[0083] Step 1. Obtain the dataset of X-ray images.
[0084] 1.1) Tens of thousands of lateral cephalometric radiographs covering the craniofacial, dental, and cervical spine regions were obtained from the dental hospital, with the age range of 4 to 50 years. The number of images obtained in this case was, but not limited to, 20,440. Curved tomographic radiographs covering the dental and mandibular regions were obtained, with the age range of 4 to 80 years. The number of images was, but not limited to, 102,982.
[0085] 1.2) Divide the training sample set D of both datasets into a 6:2:2 ratio. train Validation sample set D val Test sample set D test .
[0086] Step 2. Set the hyperparameters of the age estimation network.
[0087] The number of neurons in the fully connected layers of the three age regression heads is set to 512, and the number of neurons in the three fully connected layers of the age regression heads with monotonicity constraints are 512, 256 and 512 respectively.
[0088] Set the custom parameter γ for the total model loss to 0.8, the batch size B to 64, and the length of the feature queue. The training iteration count is 4096, and the number of training iterations is 90.
[0089] Set the initial learning rate It is 0.001, and decays to 0.0001 and 0.00001 in the 60th and 80th iterations, respectively.
[0090] Step 3. Train and validate the age estimation network to obtain a trained age estimation model.
[0091] 3.1) Using the Adam optimizer, the age estimation training set D obtained in 1.2) is processed. train Input into the age estimation network;
[0092] 3.2) In each training iteration, samples are randomly selected from the training set, and the total loss of the age estimation network is calculated. ;
[0093] 3.3) Based on the total loss of the age estimation network, backpropagation is performed using the Adam optimizer to update the model parameters of the age estimation network in this round and save them.
[0094] 3.4) Verify the sample set D val Input the data into the model saved after this round of updates to obtain the model's age estimation results on the validation set;
[0095] 3.5) Repeat steps 3.2)-3.4) until the predetermined number of training rounds is reached, and select the model that performs best on the validation set as the final trained age estimation model.
[0096] Step 4. Obtain the age estimate.
[0097] The test sample set D obtained in 1.2) test The input is fed into the trained age estimation model, and the output is the age estimation result.
[0098] It should be noted that the step numbers in the specification and claims of this invention are only for the purpose of clearly describing the embodiments of this invention and facilitating understanding, and their order is not limited.
[0099] The effects of this invention can be further illustrated by the following simulation experiments:
[0100] I. Simulation conditions:
[0101] The simulation test platform of this invention is a desktop computer with an Intel Core i7-9700K CPU 3.6GHz, 128GB of memory, and an Nvidia RTX4090 graphics card. The simulation platform operating system is Ubuntu 18.04, and it uses the PyTorch deep learning framework and is implemented in Python.
[0102] Data source: Reference Figure 5 The data for this experiment came from a dental hospital, including 20,440 lateral cephalometric radiographs covering the craniofacial, dental, and cervical spine regions, and 102,982 panoramic radiographs covering the teeth and mandibular regions, primarily for individuals aged 4 to 50 years. This data has been processed by professional radiologists to remove samples with incorrect imaging placement, inconsistent age, or abnormal imaging posture. (Reference) Figure 6The sample size and age distribution show that both datasets exhibit significant uneven age distribution. Specifically, the lateral cephalometric radiograph dataset is mainly distributed in the 10-25 age range, with an average sample size of 65.26 in each integer age group. The curved tomographic radiograph dataset is mainly concentrated in the 15-65 age range, with an average sample size of 200.31 in each integer age group.
[0103] To comprehensively evaluate the relationship between the performance of the age estimation model and the sample size, the two datasets were divided into three sample groups: majority samples, intermediate samples, and minority samples. The majority samples are defined as follows: the number of samples in the integer age group to which the sample belongs is more than twice the average sample size; the minority samples are defined as follows: the number of samples in the integer age group to which the sample belongs is less than half the average sample size; and the intermediate samples are defined as follows: the number of samples in the integer age group to which the sample belongs is between the two.
[0104] The mean absolute error (MAE) is used as an evaluation metric to measure the error between the estimated age and the actual age. Its calculation formula is as follows: Where N is the sample size. For sample age labels, This is the result of age estimation.
[0105] II. Simulation Content and Result Analysis:
[0106] Simulation 1: Under the simulation conditions described above, the ages of the two image datasets were estimated using both the methods presented in this invention and existing methods that perform well in age estimation tasks for the deep imbalance regression problem. The results of the age estimation evaluation metrics for different age estimation methods on the two image datasets are shown in Table 1.
[0107] Table 1. Results of age estimation evaluation indicators for different age estimation methods on two image datasets.
[0108]
[0109] The sources of the existing methods in Table 1 are as follows:
[0110] LDS: YANG Y, ZHA K, CHEN Y, et al. Delving into deep imbalancedregression[C] / / International Conference on Machine Learning. PMLR, 2021:11842-11851.
[0111] LDS+FDS: YANG Y, ZHA K, CHEN Y, et al. Delving into deep imbalancedregression[C] / / International Conference on Machine Learning. PMLR, 2021:11842-11851.
[0112] BMSE-BMC: REN J, ZHANG M, YU C, et al. Balanced mse for imbalancedvisual regression[C] / / Proceedings of the IEEE / CVF Conference on ComputerVision and Pattern Recognition. 2022: 7926-7935.
[0113] BMSE-BNI: REN J, ZHANG M, YU C, et al. Balanced mse for imbalancedvisual regression[C] / / Proceedings of the IEEE / CVF Conference on ComputerVision and Pattern Recognition. 2022: 7926-7935.
[0114] BMSE-GAI:REN J, ZHANG M, YU C, et al. Balanced mse for imbalancedvisual regression[C] / / Proceedings of the IEEE / CVF Conference on ComputerVision and Pattern Recognition. 2022: 7926-7935.
[0115] RankSim: GONG Y, MORI G, TUNG F. RankSim: Ranking similarityregularization for deep imbalanced regression[C] / / Proceedings of MachineLearning Research: Vol. 162 Proceedings of the 39th International Conference on Machine Learning. PMLR, 2022: 7634-7649.
[0116] RnC: ZHA K, CAO P, SON J, et al. Rank-n-contrast: learning continuous representations for regression [J]. Advances in Neural Information ProcessingSystems, 2024, 36.
[0117] ResNet-50: HE K, ZHANG X, REN S, et al. Deep residual learning for image recognition[C] / / Proceedings of the IEEE / CVF Conference on ComputerVision and Pattern Recognition. 2016: 770-778.
[0118] In Table 1, the best and second-best performance in each data group is indicated by bold and underline, respectively.
[0119] As shown in Table 1, the overall performance of this invention is superior to the comparative method on both datasets, and this advantage is more significant in the minority sample groups. Specifically, on the lateral cephalometric radiograph dataset, the MAE of this invention in the minority sample group is approximately 0.45 years lower than the best-performing comparative method. On the curved tomographic radiograph dataset, the MAE of this invention in the minority sample group is approximately 0.48 years lower than the best-performing comparative method. The difference in performance improvement on the two datasets also indicates that this invention is more effective for datasets with a wider age range. The performance of this invention is superior to the comparative method for the majority of samples and for the overall dataset. Specifically, the overall MAE of this invention for age estimation on lateral cephalometric radiographs and curved tomographic radiographs is reduced by 0.10 years and 0.06 years, respectively, compared to the comparative method, and the MAE for the majority sample group is reduced by 0.04 years and 0.08 years, respectively.
[0120] Due to the numerical advantage of the majority sample group, each comparison method can well characterize the aging features in these samples. The table shows that the performance improvement of the present invention for the majority sample is not significant compared to the minority sample, which also indicates that the optimization of the present invention for the minority sample is not conditional on reducing the performance of the majority sample and the dataset as a whole.
[0121] For the medium-sized sample group, the performance improvement of this invention is not significant, and is even worse than the best comparative method. This is because the medium-sized sample group does not have a numerical advantage, and there is no biased optimization for the medium-sized sample group during network training, resulting in insufficient representation of this part of the sample and ultimately lower age prediction performance.
[0122] Because of the small sample size, the performance improvement of this invention on a few sample groups does not contribute significantly to the overall performance, but it is of great significance to the individual subjects.
[0123] Simulation 2: Under the simulation conditions described above, age estimation was performed on the two image datasets using the present invention and the ResNet-50 residual regression network method, respectively. The performance improvement of the present invention relative to ResNet-50 was plotted for each age group with different numbers of test samples, i.e., the reduction in the MAE of the age estimation results. The age estimation results for the lateral cephalometric radiographs are shown below. Figure 7 The results of age estimation for curved section sections are as follows: Figure 8 ,
[0124] from Figure 7 and Figure 8As can be seen, the improvement in age estimation performance of this invention is mainly reflected in a small number of samples, consistent with the results in Table 1. Besides sample size, the age of the samples is also an important factor related to performance improvement. Specifically, the sample distribution of both datasets shows that the sample size at the two ends of the age range is much smaller than that in the middle age range. The improvement in age estimation performance of this invention is mainly reflected in older samples. For infant skeletal images, which also represent a small sample size, the performance improvement of this invention is relatively insignificant because infant development is a rapid and significant process, and their X-ray images show obvious aging characteristics, resulting in good age estimation performance for both this invention and the comparison method. However, the aging process in the human body is a relatively slow and gradual process, making age-related changes difficult to capture in the X-ray images of elderly individuals. This provides room for performance improvement for this invention. This invention effectively addresses the challenge of representing age-related changes during the slow aging process, thus showing a more significant performance improvement on older samples.
[0125] Simulation 3: Under the above simulation conditions, the age of the two image datasets was estimated using the method of this invention, the ResNet-50 residual regression network method, and the LDS+FDS method, respectively. The visualization results of the aging change characteristics extracted by each method in the uniform manifold approximation and projected UMAP spaces were then plotted. The results are as follows: Figure 9 ,
[0126] Figure 9 The color of each feature point in the scatter plot corresponds to its true age.
[0127] from Figure 9 As can be seen, this invention exhibits stronger age continuity in the UMAP space. Its feature points form a continuous gradient along the age growth direction, showing a banded distribution from childhood to old age without obvious breaks. Compared with the other two methods, this invention has significantly fewer outliers.
[0128] On the lateral cephalometric radiograph dataset, the feature points of this invention form a continuous and smooth manifold with age. Although the residual regression network method and the LDS+FDS method also show significant age correlation in the UMAP space, their continuity and smoothness are poor, and they show more outliers.
[0129] On the surface tomographic dataset, the feature point locations in the UMAP space of this invention show strong consistency with the true age, and the density distribution is also relatively consistent with the true age distribution. The residual regression network method exhibits many outliers in the UMAP feature space, primarily young age samples. The LDS+FDS method demonstrates age discontinuity, with feature points forming several meaningless and discrete clusters in the UMAP space.
[0130] Through UMAP dimensionality reduction visualization and quantitative analysis of the above-mentioned age estimation performance, it is shown that the present invention can demonstrate a more efficient age characterization capability and can more accurately characterize the essential change law of human biological age in X-ray images.
Claims
1. An age estimation network based on the law of elasticity and non-uniformly distributed X-ray human skeletal images, characterized in that, include: A feature extractor is built to extract batch features from the input image to construct the distance feature matrix and the midpoint feature matrix; A feature queue is introduced, which uses supervision, distance elasticity and midpoint elasticity to work on the samples and perform comparative learning between samples to obtain quality-weighted biased optimization. Using supervised force to define the basic age estimation loss Obtain the total loss of network iterative training. ; Establish monotonic constrained residual blocks MCRes to map the feature distance between sample pairs to the partially monotonically increasing distance feature PMF between samples; Establish mutually independent age estimation regression heads to output age estimation results, sample pair midpoints, sample pair distances, and age estimation results with monotonicity constraints; The feature extractor, feature queue, monotonic constraint residual block, and age estimation regression head are cascaded in sequence to form an age estimation network for non-uniformly distributed X-ray human skeleton images.
2. The network according to claim 1, characterized in that, The feature extractor includes an initial convolutional unit, a residual unit, and a global pooling unit. The structure and connection relationships of each unit are as follows: The initial convolutional unit includes a 7×7 convolutional layer with a stride of 2, a batch normalization layer, an activation layer, and a 3×3 max pooling layer, which are used to convert image information into high-dimensional information and transmit it to the residual unit. The residual unit includes four residual modules with different structures, which are used to extract features from the input high-dimensional information, generate feature maps, and transmit them to the global pooling unit. The global pooling unit includes a global average pooling layer, which is used to perform pooling processing on the input feature map in the spatial dimension, compressing the feature map of each channel into a single feature value. The initial convolutional unit, residual unit, and global pooling unit are cascaded in sequence to form a feature extractor.
3. The network according to claim 2, characterized in that, The four residual modules in the residual unit, each with a different structure, have the following structure and connection relationship: The first residual module includes three cascaded residual blocks, and the feature map size of each residual block remains unchanged. The second residual module includes four cascaded residual blocks, where the first residual block is downsampled and the subsequent residual blocks maintain the same feature map size. The third residual module includes six cascaded residual blocks, where the first residual block is downsampled and subsequent residual blocks maintain the same feature map size. The fourth residual module includes three cascaded residual blocks, where the first residual block is downsampled and the subsequent residual blocks maintain the same feature map size. The four residual modules are cascaded in sequence to form a residual unit.
4. The network according to claim 3, characterized in that, Each residual block in each residual module includes three convolutional layers, wherein The first convolutional layer includes a 1×1 convolution, a batch normalization layer, and an activation layer; The second convolutional layer includes a 3×3 convolution, a batch normalization layer, and an activation layer; The third convolutional layer includes a 1×1 convolution and a batch normalization layer; The three residual blocks are cascaded sequentially, and the input of the first convolutional layer and the output of the third convolutional layer are added element by element, which becomes the output of the residual block after activation.
5. The network according to claim 1, characterized in that, The feature queue is used to expand the number of sample pairs and to perform supervised force, distance elasticity, and midpoint elasticity on the samples to conduct temporal comparative learning between samples in order to obtain quality-weighted biased optimization, which includes: For a given batch of size B, by one-to-one pairing of samples, we obtain... 1 sample pair, with an introduced length of 1 Feature queue It follows a first-in, first-out (FIFO) update strategy to cache batch features of samples during training iterations, thus reducing the number of sample pairs from... Expand to ; Building Supervisory Power Distance elasticity Midpoint elasticity The three-way collaborative work model, in which, Indicates sample Age tags, Indicates corresponding to The age regression results This represents the label distance between two samples in a sample pair. For the corresponding estimated distance, The midpoint of the labels of the two samples. The three forces mentioned above are used as the loss function to optimize the model in subsequent learning, with the midpoint of the corresponding estimation results as the midpoint. The above three forces work together to achieve temporal comparative learning between two samples in a sample pair. The "elastic force" comes from the difference between the label distance and the estimated distance. When the estimated distance is less than the actual distance, the elastic potential energy is released through the elastic force, generating a repulsive force between the samples. Conversely, it generates a pulling force. Define sample The mass is: ,in The number of age groups. For the sample Number of samples in the age group and This represents the mean of all age labels within the corresponding age group; In the definition of sample quality Replace with medium sample Number of samples in the age group; The biased optimization expression for mass weighting, which combines sample mass and sample stress conditions, is as follows: ; in, To estimate the distance loss for perceiving the temporal nature of development and aging processes, Used to select whether to use midpoint elastic force; when the distance elastic force on the sample is in the same direction as the tag supervision force. =1, otherwise =0.
6. The network according to claim 1, characterized in that, The basic age estimation loss is defined using supervisory force. Obtain the total loss of network iterative training. The formula is as follows: ; ; Where B is the batch size. The loss function of age regression results under monotonic constraints, representing the supervised force on the samples. ,in For the sample Age tags, The results of age regression with monotonicity constraints are as follows. For custom parameters, The loss function to achieve the monotonically increasing constraint.
7. The network according to claim 1, characterized in that, The monotonic constrained residual block MCRes is used to obtain partially monotonically increasing distance features between samples. Its implementation includes: First, pair the samples Feature distance ,in and Samples and The age-related changes are mapped as Partially monotonically increasing distance features: ; ; in, For The mapping of weights in a multilayer sensing mechanism uses a loss function. Implement the monotonically increasing constraint; Then the and Adding them together yields a partially monotonically increasing distance feature between samples. : 。 8. The network according to claim 1, characterized in that, The establishment of mutually independent age estimation regression heads includes ψ d The distance regression head of the sample pairs with parameters , with ψ m The midpoint regression head of the sample pair with parameters ,by Regression head of sample age as parameter Age regression head with monotonicity constraint ,in: The sample pairs are distance regression heads Used to obtain the estimated distance between sample pairs: , The sample midpoint regression head Used to obtain the midpoint of the sample pair estimation results: , The sample age regression head Used to obtain tag age estimates: , The age regression head with monotonicity constraint This is used to obtain age regression results with monotonicity constraints: , in, Batch features, i.e., samples The characteristics of age-related changes Features, i.e., samples, are cached for the feature queue. The characteristics of age-related changes ; The above regression head , and Each layer is built using a fully connected layer, and the regression head is formed by cascading three fully connected layers. .
9. A method for age estimation of X-ray human skeleton images using the age estimation network described in claim 1, characterized in that, include: (1) Two datasets, lateral cephalometric radiographs covering the craniofacial, dental, and cervical spine regions and panoramic tomographic radiographs covering the dental and mandibular regions, obtained from medical institutions and aged 4 to 50 years, were used to divide the training sample set D in a 6:2:2 ratio. train Validation sample set D val Test sample set D test : (2) Set the hyperparameters of the age estimation network, including: the number of neurons in the fully connected layers of the three age regression heads, the number of neurons in the three fully connected layers of the age regression heads with monotonicity constraints, the custom parameter γ for the total loss, the batch size B, and the length of the feature queue. Number of training iterations, learning rate ; (3) The training sample set D train Input the age estimation network, perform iterative training and validation, and obtain a trained age estimation network model; (4) Test sample set D test The input is fed into a trained age estimation network model, which outputs the predicted age.
10. The method according to claim 9, characterized in that, The iterative training and validation of the age estimation network in (3) is implemented as follows: (3a) Use the Adam optimizer and set the optimizer parameters; (3b) In each training iteration, samples are randomly selected from the training set, and the total loss of the age estimation network is calculated; (3c) Based on the total loss of the age estimation network, perform backpropagation using the Adam optimizer, update the model parameters of the age estimation network in this round, and save the model parameters of the age estimation network. (3d) Verify the sample set D val Input the data into the model saved after this round of updates to obtain the model's age estimation results on the validation set; (3e) Repeat steps (3b)-(3d) until the predetermined number of training rounds is reached, and select the model that performs best on the validation set as the final trained age estimation model.