Method of generating highly consistent predicted values from planer images

TWI937509BActive Publication Date: 2026-09-01ALPHA INTELLIGENCE MANIFOLDS INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
TW113121074
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2026-09-01
Estimated Expiration
2044-06-05

AI Technical Summary

Technical Problem

Existing AI models for predicting bone mineral density (BMD) exhibit significant variability in estimates from different X-ray images taken within a short period, making them unsuitable for continuous monitoring due to unpredictable training methods, which fail to meet the accuracy standards set by the International Society for Clinical Densitometry (ISCD).

Method used

A method is introduced that includes a secondary loss term during model training, using a primary and secondary dataset to adjust parameters, where the primary dataset includes labeled training images with ground truth values and the secondary dataset comprises unlabeled image pairs with similarities, along with data augmentation techniques to enhance model consistency.

Benefits of technology

The method improves the stability and consistency of AI predictions for BMD, achieving a coefficient of variation (CV) within the recommended 1.8-2.5% for repeated measurements, suitable for continuous monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TB001908487_001
    Figure TWG2TB001908487_001
  • Figure TWG2TB001908487_002
    Figure TWG2TB001908487_002
  • Figure TWG2TB001908487_003
    Figure TWG2TB001908487_003
Patent Text Reader

Abstract

This invention relates to a method for training a prediction model to generate a primary predicted value for a primary feature of an input image. The method includes training the prediction model using a primary dataset containing labeled training images with ground truth values ​​and an auxiliary dataset containing two unlabeled training images without ground truth values. The training objective is to reduce a first loss and a second loss, wherein the first loss calculates the difference between the predicted value of the labeled training image and the training ground truth value, and the second loss calculates the difference between the predicted values ​​of the two unlabeled training images.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for generating highly consistent predictive values ​​from a planar image, particularly for generating skeletal feature values ​​related to osteoporosis and fracture risk in a subject from a radiographic image. [Previous Technology]

[0002] According to the World Health Organization (WHO), osteoporosis is a systemic bone disease characterized by decreased bone mass and deterioration of bone microstructure, leading to fragile bones and an increased risk of fractures. Patients may also experience various complications after a fracture. Because osteoporosis is usually painless and has no obvious symptoms, it is necessary to undergo examinations to determine if one has osteoporosis.

[0003] Osteoporosis can be diagnosed by measuring bone mineral density (BMD). In addition to standard methods such as dual-energy X-ray absorptiometry (DXA), other artificial intelligence-related methods can also be used to generate BMD values.

[0004] Because BMD values ​​can be used as indicators of treatment effectiveness or as a basis for continuous monitoring of patients' bone quality, they are often measured continuously, which requires a certain degree of consistency. According to the guidelines issued by the International Society for Clinical Densitometry (ISCD), when the same operator performs consecutive repeated measurements (usually 30 people, each measured twice consecutively), the accuracy of DXA measured by the coefficient of variation (CV) should be within 1.8% - 2.5%.

[0005] Currently, the use of artificial intelligence (AI) to predict bone mineral density (BMD) values ​​is mainly for screening to determine whether a subject has osteoporosis, rather than for continuous monitoring of BMD values. The main reason is that AI models may make significant differences in BMD estimates based on different X-ray images taken of the same person over a short period (typically 3-6 months) (e.g., much greater than the 1.8-2.5% limit specified by the ISCD). However, due to the unpredictable nature of AI models, these differences stemming from different AI model training methods are difficult to control. This makes it difficult to apply AI models to estimate BMD using continuous measurements.

[0006] Therefore, a new method is needed to improve the consistency of AI output, especially skeletal feature values ​​related to a subject’s osteoporosis and fracture risk. [Summary of the Invention]

[0007] To address the above problems, this invention provides a method for generating highly consistent key feature values ​​from planar images. This method introduces a second loss term in addition to the traditional first loss term during model training. Therefore, this method provides better stability for the model's prediction results.

[0008] Specifically, the present invention provides a method for training a prediction model to generate one or more principal predicted values ​​of principal features of an input image, comprising training the prediction model using a primary dataset and a secondary dataset, and reducing a total loss of the prediction model by adjusting multiple parameters in the prediction model. In this method, the primary dataset includes multiple primary learning data, each primary learning data comprising a labeled training image labeled with one or more principal ground truth values ​​of the principal features; the secondary dataset includes multiple secondary learning data, each comprising a pair of unlabeled training images, the unlabeled training image pair comprising a first unlabeled training image and a second unlabeled training image having similarity to the principal features. The total loss used to adjust the multiple parameters includes a primary loss and a secondary loss. During training, the prediction model can generate one or more primary predicted values ​​of the principal features of the labeled training image, and calculate the primary loss based on the difference between the one or more principal ground truth values ​​of the labeled training image and the one or more primary predicted values. Similarly, the prediction model can generate one or more first predicted values ​​for the principal features of a first unlabeled training image and one or more second predicted values ​​for the principal features of a second unlabeled training image. The secondary loss is calculated based on the difference between one or more first predicted values ​​of the first unlabeled training image and one or more second predicted values ​​of the second unlabeled training image. In one embodiment, the primary loss is calculated using a primary loss function, and the secondary loss is calculated using a secondary loss function. Furthermore, in a specific embodiment, the primary loss function and the secondary loss function calculate a squared loss.

[0009] In one embodiment, to ensure the similarity of the primary feature, in each of the plurality of secondary learning materials, the first unlabeled training image and the second unlabeled training image are images of the same subject taken within a predetermined time interval, wherein the primary feature is known to be constant or to vary only within the measurement limits of the measurement method during the predetermined time interval. Specifically, the predetermined time interval may be 3 months or 6 months.

[0010] In one embodiment, each of the labeled training image, the first unlabeled training image, and the second unlabeled training image is a ROI-captured image extracted from an original training image via ROI (Region of Interest), rather than using the original training image. The ROI is a region closely related to the principal feature. The ROI capture can be performed using a trained ROI capture model. Alternatively, each of the labeled training image, the first unlabeled training image, and the second unlabeled training image can also be a training image set containing the original training image and the ROI-captured image, and the prediction model can use both images in the image set to predict the principal feature value.

[0011] In addition to the primary feature, in one embodiment, one or more auxiliary features are also used, and the method further includes training the prediction model to generate one or more auxiliary predicted values ​​of the one or more auxiliary features of the input image by adjusting the multiple parameters in the prediction model to reduce the total loss of the prediction model. When training the model using auxiliary features, the one or more auxiliary features need to be associated with the primary feature. In this embodiment, the labeled training image of each of the multiple primary learning data is further labeled with one or more auxiliary ground values ​​of the one or more auxiliary features. The total loss used to adjust the multiple parameters also includes a three-dimensional loss. During training, the prediction model can also generate one or more three-dimensional predicted values ​​of the one or more auxiliary features of the labeled training image, and calculate the three-dimensional loss based on the difference between the one or more auxiliary ground values ​​and the one or more three-dimensional predicted values.

[0012] During training, data augmentation can also be applied to the prediction model. In one embodiment, the labeled training image is modified by image augmentation before the prediction model generates the one or more primary predictions. In one embodiment, the one or more primary ground truth values ​​are modified by ground truth augmentation before calculating the primary loss.

[0013] In a preferred embodiment, the key feature is a subject's bone mineral density, and the one or more key predictive values ​​are one or more bone mineral density (BMD) values. In a specific embodiment, the one or more key predictive values ​​include bone mineral density (BMD) values ​​for the entire hip, femoral neck, greater trochanter of the femur, and femoral shaft.

[0014] In one specific embodiment, the prediction model is trained to generate BMD values ​​from X-ray images. In this embodiment, the labeled training images in each first-order learning data and the first and second unlabeled training images in each second-order learning data are both X-ray images.

[0015] In one embodiment, the first unlabeled training image and the second unlabeled training image are two X-ray images taken consecutively within 3 months of the same subject.

[0016] A model for BMD prediction can use one or more auxiliary features to improve its ability to generate BMD values. In this case, the method may further include training the prediction model to generate one or more auxiliary predicted values ​​of one or more auxiliary features of the input image by adjusting the multiple parameters in the prediction model to reduce the total loss of the prediction model. For model training using auxiliary features, the one or more auxiliary features need to be associated with the subject's bone mineral density. The one or more auxiliary features may include the subject's cortical bone thickness, and the auxiliary predicted value corresponding to the cortical bone thickness is the subject's cortical bone thickness index (CTI) value. The one or more auxiliary features may also include the subject's femoral neck width, and the auxiliary predicted value corresponding to the femoral neck width is the subject's femoral neck width (FNW) value.

[0017] Typically, X-ray images taken at medical facilities are images of the entire pelvic region. To enable the prediction model to focus on areas closely related to key features, images captured by ROIs (Regions of Interest) can be used instead of the entire X-ray image as training images. In this embodiment, the labeled training images are ROI images, which are the identified ROI regions of the hip joint extracted from the original training images. Alternatively, to retain more information from the original training images, the labeled training images can be a set of training images including the original training images and the ROI images.

[0018] In one specific embodiment, before generating one or more first-order predictions through the prediction model, the original training image and the ROI capture image are modified by image enhancement. The image enhancement may be performed by cropping 0-25% of the original training image without cropping the identified ROI region, and the image enhancement may be performed by moving the identified ROI region by 0-7% in a specific direction.

[0019] In one specific embodiment, one or more principal true values ​​are modified by truth augmentation before calculating the first-order loss. The aforementioned truth augmentation can be performed by introducing a small variable into one or more principal true values. The value of this small variable may be unaffected by the numerical value of the principal true value. In one embodiment, each small variable is a value randomly selected between -0.01 g / cm² and 0.01 g / cm². In another specific embodiment, each small variable is randomly selected from a normal distribution with a population mean of 0, truncated at ±0.01 g / cm², and this normal distribution may have a standard deviation of 0.01 g / cm². Alternatively, the value of the small variable may also depend on the numerical value of the principal true value. In one embodiment, each of the one or more principal true values ​​has a value yn, and each small variable is a value randomly selected between -0.01yn and 0.01yn.

[0020] Other objects, advantages, and novel features of the present invention will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. [Brief Explanation of the Drawings]

[0021] Figure 1 shows the architecture of a prediction model trained using a labeled dataset and an unlabeled dataset.

[0022] Figure 2 shows the architecture of a prediction model trained on a labeled dataset with primary and secondary features.

[0023] Figure 3 shows the architecture of the prediction model trained using augmented data.

[0024] Figure 4 shows the prediction model connected in series with the ROI extraction model. This prediction model uses the extracted ROI as input to generate predicted values.

[0025] Figure 5A shows training images of the entire hip, labeled by a radiologist, including the left and right hip joints. Figure 5B shows training images of the hip joint region, labeled by a radiologist, including the femoral neck, greater trochanter, and femoral shaft. Figure 5C shows the ROI identification results used to generate the BMD, which identified the hip joint region as the region of interest.

[0026] Figure 6 shows an example of feeding data into the model, in which a set of first-order learning data randomly selected from the first-order dataset and a set of second-order learning data randomly selected from the second-order dataset are combined and used to calculate a total loss.

[0027] Figure 7 shows an example of calculating total loss, which includes an accuracy term and a precision term.

[0028] Figure 8 shows the feature points corresponding to the outer and inner edges of the cortical bone used to calculate the CTI value in the CTI calculation model.

[0029] Figure 9 shows an example of random displacement of the ROI. In this example, the ROI region is enlarged from the center outwards to all four sides before image cropping.

[0030] Figure 10 shows an example of data enhancement. The enhanced data from the first dataset includes enhanced image (P1') and enhanced ground truth (Bh1', Bfn1', Bgt1' and Bfs1').

Implementation Method

[0031] The terms used in this specification are intended to be interpreted in their broadest and most reasonable manner, even when used in conjunction with detailed implementation of certain specific embodiments of the present technology. Some terms may even be particularly emphasized below; however, any term intended to be interpreted in any limited manner will be specifically defined in this detailed implementation section.

[0032] The embodiments described below can be implemented by programmable circuitry programmed or configured in software and / or firmware, or entirely by special purpose circuitry, or a combination of these forms. Such special purpose circuitry (if any) can be, for example, one or more application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), etc.

[0033] The object of this invention is to provide a method for training an AI model to generate highly accurate principal predictions for the main features of an input image. In other words, the trained model should be able to produce consistent results for input images with similar principal feature values, regardless of the presence of noise in the image. This improvement can be achieved by introducing an additional "accuracy term" into the loss function during model training, which will be described in detail below.

[0034] An artificial intelligence model can learn to generate output values ​​from input images by training on a labeled training dataset. This model can be built using convolutional neural network (CNN) based algorithms (e.g., LeNet, AlexNet, VGG, GoogLeNet, ResNet, and DenseNet). Alternatively, it can use transformer-based visual algorithms (e.g., ViT). Each set of training data in the labeled training dataset contains training images labeled with their ground truth values, serving as the learning target for the model. However, artificial intelligence models trained in this way may produce highly variable outputs for similar inputs (e.g., two or more inputs considered to have nearly identical ground truth values). This variability is difficult to control due to the unpredictable nature of artificial intelligence models. One possibility for addressing this problem is to increase the size of the labeled training dataset by incorporating more training data. However, large amounts of labeled training data are not always available, often preventing the use of this method to train models.

[0035] While labeled training data is difficult to obtain in large quantities, image sets containing unlabeled images that share certain attributes or similarities in their main features are easier to obtain because they do not require labeling. Therefore, additional unlabeled training data sets can be used to train the model. Each set of training data in the unlabeled training data set is an image set consisting of two or more training images that are considered to have substantially the same main feature values. The learning objective of the model is to learn from the unlabeled training data set to produce outputs with low variability among the training images in each image set. Using this additional learning objective, the model can learn to produce consistent outputs for input images with similar main features, as shown in Figure 1.

[0036] If the intrinsic difference inferred between the principal feature values ​​of two or more images is much smaller than the accuracy error in obtaining these values, then these values ​​can be considered "identical" or "indistinguishable". For example, obtaining image pairs / sets with similar principal features is a common method for acquiring two or more images taken within a short period of time. In a preferred embodiment, the principal features are measurable human biometrics. This may include, but is not limited to, height, weight, serum albumin, and bone mineral density. The values ​​of these features may change over time, but for each feature value, a "short time period" can be found such that the natural variation of the feature value within this short time period is always less than the detection limit or measurement error of the corresponding measurement method. Since the true difference between two feature values ​​measured within a short time period exceeds the measurement error, their true difference cannot be analyzed by this measurement method and should be considered "indistinguishable" or "similar". The length of this short time period is predetermined based on (1) the maximum rate of change of the measured feature and (2) the measurement error or detection limit of the measurement method for obtaining the feature value, such that the product of the maximum rate of change and the predetermined short time period is less than the measurement error or the detection limit of the measurement method. In other words, any two feature values ​​measured from the same person within the predetermined short time period are similar to each other.

[0037] Specifically, in bone mineral density (BMD) tests measured by dual-energy X-ray absorptiometry (DXA), the error for repeated measurements by the same operator is approximately 2%. However, studies show that a person's BMD typically changes by less than 1% per year. Therefore, two X-ray images taken of the same person within 3 or 6 months can reasonably be considered "similar" because the intrinsic BMD change in the same person within 6 months is far less than 2%, a difference that DXA cannot measure considering the accuracy error of repeated measurements. Other examples include measuring the weight of the same person within one or two days, or measuring the serum albumin concentration of the same person within a few hours; in both cases, the accuracy error may be greater than the intrinsic difference itself.

[0038] In addition to training the model to produce only one output from the input, it is also possible to train the model to produce multiple outputs from the input image. If the training image label has multiple ground truth values, the model can learn to produce these values ​​simultaneously as multiple outputs. For example, the model can be trained to produce only the values ​​of the primary features as outputs, or it can be trained to produce multiple values ​​containing one or more primary feature values ​​and multiple auxiliary feature values. We can select one or more auxiliary features associated with the primary features and train the model to produce the values ​​of those auxiliary features, as shown in Figure 2.

[0039] Training the model to produce outputs other than the primary feature values ​​can serve as additional guidance for adjusting the parameters used to output the primary feature values, since the primary feature and the selected auxiliary features are interrelated. Furthermore, producing multiple values ​​for the primary feature (rather than just one) can also improve model performance, as the model can learn more feature details by producing multiple values. For example, research has shown that the values ​​of the cortical thickness index (CTI) and femoral neck width (FNW) of the hip joint are correlated with the bone mineral density (BMD) value of the hip joint region. Therefore, training a model to predict CTI and FNW values ​​other than BMD may improve the model's performance in predicting BMD values, because the three values ​​of the hip joint share some common factors that can be learned by the model during training.

[0040] In addition, various data augmentation techniques can be applied to increase the variability of training data to obtain better model generalization and performance. Data augmentation can include input image augmentation and / or ground truth augmentation, as shown in Figure 3. Since data augmentation creates artificial data in addition to the real input / measurement, industry knowledge is required to create reasonable augmented data. A "reasonable" augmented data refers to data that could actually be obtained during the data collection process. For example, reasonable augmented input images can correspond to changes during image capture, and reasonable augmented ground truth can correspond to changes in measurement values ​​caused by instrument or operational noise.

[0041] In image enhancement, the original input image can be modified through geometric transformations, color space transformations, and noise injection. Geometric transformations include rotation (rotating the image by a specified angle), flipping (reflecting the image horizontally or vertically), cropping (removing a portion of the original image), and shifting (moving the image in different directions). Color space transformations can modify the color attributes of the image, such as illumination, color saturation, and contrast. Furthermore, noise can be introduced into the image to simulate defects in the real world. The adjusted image represents a slightly modified copy of the original input image, and its true value should be the same as the original input image. For radiographic images, these adjustments simulate variations during the image capture process, such as the patient's posture, position, or offset, the X-ray energy of the instrument, and operator-induced differences. The data enhancement parameters applied in training are based on principles that follow the variations that may occur when capturing conventional X-ray images. For example, scaling / panning within 5% of the image size, rotation within 15 degrees, gamma correction within 25% to simulate overexposure / underexposure, introducing random Gaussian blur to simulate focus deviation, and introducing random Gaussian noise to simulate sensor noise.

[0042] In ground truth augmentation, the ground truth value of an input image is slightly modified within a reasonable range to simulate variability during repeated measurements on the same subject. Generally, any variation smaller than the measurement precision may be used for ground truth augmentation to simulate variability caused by repeated measurements. For example, when measuring BMD via DXA, the reasonable coefficient of variation (CV) for repeated measurements on the same patient is approximately 2%. Therefore, adjusting a specific BMD value within a 1% range can be considered an acceptable modification to create artificial data points corresponding to the original data used for model training.

[0043] In summary, using an original image and its corresponding ground truth, many modified images and many modified output values ​​can be generated, and the combination of these many modified images and many modified output values ​​can effectively expand the training dataset used for model training.

[0044] The framework of an AI model can be designed based on the attributes of the input image. It can be a single model predicting one or more output values ​​from an input image. Alternatively, it can be an integrated model connecting two or more sub-models, each handling a specific task. For example, a prediction model can be cascaded with a Region of Interest (ROI) localization / extraction model, allowing the ROI localization model to extract key regions from the input image into the prediction model, as shown in Figure 4. This allows the prediction model to focus on regions closely related to key features. The prediction model can receive the extracted ROI as its input, or it can receive both the original image and the extracted ROI as its input. In the latter case, the prediction model uses an image set containing both the original image and the extracted ROI as its input.

[0045] Specifically, an integrated AI model for predicting a subject's bone mineral density (BMD) value from planar X-ray images can be implemented by concatenating a Region of Interest (ROI) localization model with a BMD generation model. To train the integrated AI model using this framework, the ROI localization model is first built independently. This model processes the original training X-ray images and outputs ROIs (e.g., the entire hip, femoral neck, greater trochanter, and femoral shaft) associated with the BMD of the original training images. The identified ROIs are then used as training inputs to the BMD generation model to adjust its parameters. The parameters of the ROI localization model remain unchanged during BMD generation. Performing ROI localization before BMD generation may be advantageous compared to a single model that directly predicts BMD values ​​from input X-ray images, as the identified regions of interest (ROIs) may force the BMD generation model to focus on key features associated with BMD values.

[0046] The following shows an example of establishing an integrated model for generating BMD based on the above concepts, and also provides the power of the established model. In some examples, the integrated model includes a ROI localization model and a BMD generation model. In the ROI localization model, an X-ray image containing at least one hip joint is used as the original input image of the model to extract the hip joint region. The extracted ROI can be used as input to the BMD generation model to produce predicted BMD values.

[0047] 1. ROI Location Model

[0048] The object detection AI model described in US Patent Publication No. US2023 / 0029674A1 is trained to identify regions of interest (ROIs) in input X-ray images, the contents of which are incorporated herein by reference. A deep neural network (DNN) model called the You Only Look Once (YOLO) algorithm is used to train the AI ​​to recognize features and select appropriate ROIs from input images. The training dataset and training process for the ROI localization model are as follows:

[0049] (1) To prepare training data, experts used ROI markers in X-ray images as training ground values ​​for the object detection model. Clinical data of 459 pelvic X-ray images from medical institutions were converted from DICOM files to high-resolution 16-bit PNG files. Images with skeletal abnormalities such as artificial joints or fractures were excluded. Then, the brightness and contrast of the images were adjusted to a standard range. Finally, the areas of the hip joint, femoral neck, greater trochanter of the femur, and femoral shaft were marked as training ground values ​​by radiologists, as shown in Figures 5A and 5B.

[0050] (2) The object detection model that performs the YOLO algorithm is trained using labeled ROIs (boundary boxes in Figures 5A and 5B) and unlabeled corresponding X-ray images.

[0051] (3) After training, the model is able to locate the ROI in the incoming pelvic X-ray images. The ROI of the hip joint region is located from the original X-ray images, while the ROIs of the femoral neck, greater trochanter of the femur and femoral shaft are located from the ROI of the hip joint region.

[0052] Figure 5C shows the results of ROI identification used for BMD generation, with the hip joint region selected as the region of interest. The trained ROI localization model can also identify the femoral neck, greater trochanter, and femoral shaft regions. However, in this invention, the hip joint region is the only ROI required for BMD generation. The hip joint region is identified from the original input image, and the hip joint ROI is used as input in the BMD generation model described below.

[0053] 2. BMD Generation Model

[0054] US Patent Publication No. US2023 / 0029674A1 describes an AI model for generating BMD from input X-ray images, the contents of which are incorporated herein by reference. In this invention, the AI ​​model is trained to generate not only Bone Mineral Density (BMD) values, but also Cortical Bone Thickness Index (CTI) and Femoral Neck Width (FNW) values. BMD is the amount of bone mineral in bone tissue; CTI is defined as the ratio of cortical bone width minus the width of the endosteum to the width of the cortical bone 100 mm below the tip of the lesser trochanter; FNW is defined as the midpoint distance between the superior and inferior cortices of the femoral neck perpendicular to the femoral neck axis. Studies have shown that BMD, CTI, and FNW are interrelated, therefore, the BMD generation model is trained to predict the values ​​of all the above features simultaneously to improve accuracy. Here, BMD is the primary feature learned by the AI ​​model, while CTI and FNW are auxiliary features. The BMD generation model can be trained to generate only one BMD value, or it can be trained to generate multiple BMD values ​​corresponding to different regions. For example, the model can be trained to generate BMD values ​​only for the entire hip joint, or it can be trained to generate four BMD values ​​for the hip joint region being analyzed (total hip, femoral neck, greater trochanter, and femoral shaft). In this example, the model is trained using RegNetY160, a ResNet algorithm.

[0055] A primary dataset and a secondary dataset were established before training the BMD generation model, wherein the primary dataset included labeled learning data and the secondary dataset included unlabeled learning data. To establish the primary dataset, clinical data of 3,169 pelvic X-ray images and their corresponding DXA measurements from medical institutions were converted from DICOM files to high-resolution 16-bit PNG files. Images of skeletal abnormalities such as artificial joints or fractures were excluded. Image brightness and contrast were normalized using histogram normalization. For each X-ray image, a corresponding DXA report with BMD subvalues ​​for four regions (total hip, femoral neck, greater trochanter, and femoral shaft) was paired with the X-ray image. Only pairs with an interval of less than 6 months between the X-ray image and the DXA measurement were included as training data. The DXA measurement included BMD subvalues ​​for the total hip, femoral neck, greater trochanter, and femoral shaft for each X-ray training image. In addition to the BMD subvalue, the ground truth values ​​for CTI and FNW in each X-ray image are labeled by an expert or a suitable AI model. After construction, each set of first-in-first-secondary training data in the first-in-first-secondary dataset includes one original X-ray image and six ground truth values ​​(total hip BMD, femoral neck BMD, greater trochanter BMD, femoral shaft BMD, CTI, and FNW). In one embodiment, the ROI localization model described above is applied to extract the entire hip region from the original X-ray image, and the extracted ROI (rather than the original X-ray image) is used as the input image in the first-in-first-secondary dataset. In another embodiment, both the original X-ray image and the extracted ROI are used as input images.

[0056] To establish the secondary dataset, clinical data containing 3,215 pairs of X-ray images without true BMD values ​​were collected from multiple medical institutions, and 16-bit image pixel arrays were directly extracted from the DICOM archive. Each pair of X-ray images was a set of secondary learning data, which included two corresponding images of the same subject inferred to have similar BMD values. According to Berger et al. (CMAJ. 2008 Jun 17;178(13):1660-8), the annual variation in BMD of an individual is typically less than 0.01 g / cm². Considering that the variation of repeated DXA measurements by the same operator is approximately 2% (i.e., approximately 0.02 g / cm²), two X-ray images of the same subject taken within 6 months can reasonably be considered "similar" because the measurable difference is much lower than the variation between repeated measurements. Therefore, each set of secondary learning data included in the secondary dataset consists of two X-ray images of the same person taken within 6 months. Similar to the primary dataset, in one embodiment, a ROI localization model is applied to extract the entire hip region from the original X-ray images, and the extracted ROI is used as the input image for each pair of X-ray images. In yet another embodiment, both the original X-ray images and the extracted ROI are used as input images.

[0057] For model training, various methods can be used to input data into the model. In one embodiment shown in Figure 6, a set of first-order learning data randomly selected from the first-order dataset and a set of second-order learning data randomly selected from the second-order dataset are combined to calculate the total loss. After selection, the learning data from the first-order dataset is further subjected to image augmentation and ground truth augmentation to enhance the model's generalization ability. The data augmentation methods will be described in detail below. Random selection and augmentation can be performed on the fly during model training, or they can be pre-built and cached in memory to improve the loading speed during model training. In a modified embodiment, the auxiliary learning data from the auxiliary dataset is also image augmented before training.

[0058] To train the prediction model to generate one or more BMD values, a CTI value, and an FNW value, a total loss function comprising BMD accuracy, CTI accuracy, FNW accuracy, and BMD precision is used to calculate the total loss. The training objective is to minimize the calculated total loss by adjusting the parameters in the BMD generation model. In one embodiment, the BMD includes four BMD values ​​(BMD for the entire hip, femoral neck, greater trochanter, and femoral shaft), therefore the total loss function calculates the BMD accuracy and precision for all four BMD values.

[0059] Figure 7 shows an example of calculating the total loss. A selected first-order learning data set includes an X-ray image P1 with its total hip BMD value Bh1, femoral neck BMD Bfn1, greater trochanter BMD Bgt1, femoral shaft BMD Bfs1, CTI value C1, and FNW value F1. A selected second-order learning data set includes a pair of X-ray images P2 and P3 without labeled ground truth values. This selected learning data set is input into the model to generate the corresponding output. For P1, four predicted BMD values ​​(1, 2, 3), one predicted CTI, and one predicted FNW are generated. For P2, four predicted BMD values ​​(1, 2, 3) are generated, and for P3, four predicted BMD values ​​(1, 2, 3) are generated. The total loss includes an accuracy term calculated from the first-order learning data set and a precision term calculated from the second-order learning data set. In this example, the accuracy term calculates the difference between the predicted and true values, comprising a first-order loss that calculates the difference in BMD (primary feature) and a third-order loss that calculates the difference in CTI and FNW (secondary feature). The precision term calculates the difference in BMD values ​​between the two sets of predicted values ​​(,,,) and (,,,). In Figure 7, squared error (e.g., mean squared error) is used to calculate the loss for the prediction results. However, other loss functions, such as mean absolute error or mean absolute percentage error, can also be used. Furthermore, although the weights of each component in each loss, as well as the weights of the first, second, and third-order losses, are all set to one here, these weights can be adjusted as needed, or through hyperparameter optimization.

[0060] The model can then be trained using the first and second datasets. In one example, for each training input, the first learning data from the first dataset and the second learning data from the second dataset are randomly selected, and the first learning data is further subjected to image enhancement and ground truth enhancement. As described above, in one embodiment, the ROI localization model extracts the ROIs of the X-ray images in the training data, and for each training data set, both the original image and the extracted ROI are used as inputs to train the model. During training, the batch size for each iteration is set between 4 and 32. The total number of training iterations is between 10 and 500, and training stops when the loss no longer decreases further.

[0061] 3. CTI Calculation Model

[0062] Studies have found a positive correlation between femoral cortical bone thickness and the bone density (BMD) values ​​of the femoral neck, femoral shaft, greater trochanter, and the entire hip. To calculate the true CTI value on X-ray images, a CTI calculation model is described in US Patent Publication No. US2023 / 0029674A1, the contents of which are incorporated herein by reference. In short, this model is used to find feature points corresponding to the outer and inner edges of the cortical bone, as indicated by the point markings A, B, C, and D in Figure 8. The training method for the CTI calculation model is as follows:

[0063] (1) Prepare 153 hip joint X-ray images, whose feature points (i.e. points A, B, C and D in Figure 8) are marked by radiologists as training true values.

[0064] (2) Train the High-Resolution Net (HRNet) algorithm to identify feature points.

[0065] (3) The trained AI model is a feature point detection model, which can be used to find feature points in the input X-ray image.

[0066] After the above training, the CTI value can be easily calculated using the following formula: CTI = (AB - CD) / AB percentage.

[0067] The CTI calculation model is used only to provide CTI ground truth values ​​for loss calculation during BMD generative model training (rather than generating CTI values ​​as input to the BMD generative model). This is because the BMD generative model is trained to generate CTI values ​​as output, rather than using these values ​​as input. This design allows the BMD generative model to focus on the correlation between CTI values ​​and BMD values ​​during training, rather than passively receiving CTI values.

[0068] 4. Data Enhancement

[0069] As described above, the learning data from the original dataset is further subjected to image enhancement and ground truth enhancement to improve the generalization ability of the BMD generative model. In image enhancement, the input image is slightly modified through operations such as cropping, shifting, enlarging, reducing, and adjusting brightness and / or contrast. By randomly selecting the modifications applied, multiple modified images can be derived from the input image.

[0070] In one example, image enhancement includes data augmentation of the original image and the identified hip region of interest (ROI). For data augmentation of the original X-ray image, up to 25% of the input X-ray image is randomly cropped while maintaining the ROI from being cropped. This prevents the model from overfitting to the image capture and cropping characteristics of a specific medical institution. For data augmentation of the ROI, since different medical institutions or photographic styles may slightly alter the identified ROI bounding box, the identified ROI is randomly shifted by 0-7%. This mitigates adaptation and generalization problems caused by ROI identification errors. Figure 9 shows an example of random ROI shifting. In this example, the ROI region is enlarged by 7% from the center outwards along all four sides. The enlarged ROI is then randomly cropped. The cropped ROI can have the same size as the originally identified ROI to simulate pure displacement. Alternatively, the cropped ROI may be smaller or larger than the originally identified ROI to simulate additional variations in the ROI. In addition to image cropping, random scaling and rotation can also be used to simulate changes in shooting conditions.

[0071] For ground truth augmentation, the true BMD values ​​(total hip, femoral neck, greater trochanter, and femoral shaft BMD) are slightly modified within a small error range. Similar to image augmentation, multiple modified BMD values ​​can be derived from the BMD values ​​trained on the ground truth.

[0072] As described above, a small variable can be added to the original BMD value. In one example, the introduced small variable is a value randomly selected between ±0.01 g / cm². Alternatively, if the original BMD has a value yn, the introduced small variable can be a value randomly selected between ±0.01yn. More complex probability density functions can also be used. In one example, truth enhancement is performed by introducing an appropriate normal distribution variable into the original BMD value. Specifically, the introduced variable is selected based on a truncated normal distribution with a population mean of 0 and a cutoff of ±0.01 g / cm². The distribution of the added variable can be changed by altering the variance or standard deviation of the normal distribution. In one embodiment, the standard deviation of the normal distribution used is set to 1 g / cm². In another embodiment, the standard deviation of the normal distribution used is set to 0.01 g / cm², making the variable more concentrated in the central region. These changes are added to account for measurement errors that may occur in DXA photography due to location, operator skill level, and / or instrument calibration. This step aims to simulate measurement changes in the real world, avoid overfitting the model to subtle systematic errors that may exist in the training data, and thus improve the model's generalization ability.

[0073] Finally, the image enhancement and ground truth enhancement methods are combined to generate training data with both enhancements. Each randomly selected first-order training data includes an input image and a set of ground truth values. The input image is processed by ROI identification to generate an ROI image. Then, the input image and the hip joint ROI are processed by image enhancement as described above to generate randomly modified inputs and randomly modified ROIs. The ground truth BMD values ​​are also processed by ground truth enhancement as described above to generate randomly modified BMD values ​​(while CTI and FNW values ​​are the original values). The modified input image, modified ROI, and modified ground truth BMD values ​​are then used as training material in the selected first-order training data, as shown in Figure 10. Although the enhanced data is artificial, it is well representative of real-world data obtained by DXA measurements.

[0074] 5. Model Performance

[0075] To evaluate the performance of the trained model, models under different training conditions were compared. The complete model is a prediction model trained using an accuracy term loss, image augmentation, and ground truth augmentation, which produces 4 BMDs, 1 CTI, and 1 FNW output. The contributions of adding an accuracy term (second loss term), applying data augmentation to the original input X-ray image (global image augmentation), applying data augmentation to the ROI image (local image augmentation), and applying ground truth augmentation to the true BMD value (ground truth augmentation) were analyzed individually by removing one operation at a time from the complete model.

[0076] The test data included X-ray images taken from 588 people. Of these 588 people, 420 had 2 X-ray images taken in a short period of time, 77 had 3 X-ray images taken, 46 had 4 X-ray images taken, 31 had 5 X-ray images taken, and 14 had 6 or more X-ray images taken. The method described by Glüer et al. (Glüer CC, Blake G, Lu Y, Blunt BA, Jergas M, Genant HK. Osteoporos Int. 1995;5(4):262-70) in the chapter on Short-Term Precision of a Technique was used to calculate the precision error. The results are shown in Table 1. Table 1: Comparison of different models by calculating the coefficient of variation (CV) of the predicted values ​​of the test data Types of models trained CV (%) Complete model 2.78 Remove precision item 3.69 Remove global (original) image enhancement 3.42 Removal of Local Area of ​​Interest (ROI) Image Enhancement 3.20 Remove truth augmentation 2.98

[0077] The CV of the model containing all operations is 2.78%. Removing the accuracy term during training increases the CV from 2.78% to 3.69%. Omitting the data augmentation of the original X-ray images in the first dataset increases the CV to 3.42%. Omitting the data augmentation of the ROI images in the first dataset increases the CV to 3.20%. Ignoring the output augmentation of the BMD ground truth increases the CV to 2.98%. It can be seen that each operation improves the model's performance. However, adding the accuracy loss term during training contributes the most. This is consistent with our hypothesis that using unlabeled images forces the model to produce more consistent results for similar inputs.

[0078] The embodiments provided above are intended to enable any person skilled in the art to make and use the claims herein. Various modifications to these embodiments will be apparent to those skilled in the art, and the novel principles disclosed herein and the claims can be applied to other embodiments without the use of inventiveness. The invention claimed in the claims is not intended to be limited to the embodiments shown herein, but is accorded the widest scope consistent with the principles and novel features disclosed herein. Additional embodiments are contemplated to fall within the spirit and scope of the invention disclosed herein. Therefore, the invention is intended to cover modifications and variations falling within the appended claims and their equivalents.

Claims

1. A method for training a prediction model executed on a computer to generate one or more principal predicted values ​​of a principal feature of an input image, comprising training the prediction model using a first-order dataset and a second-order dataset, and reducing a total loss of the prediction model by adjusting multiple parameters in the prediction model; wherein: The primary dataset includes multiple primary learning datasets, each including a labeled training image labeled with one or more primary ground truth values ​​of the primary feature; the secondary dataset includes multiple secondary learning datasets, each including an unlabeled training image pair, the unlabeled training image pair comprising a first unlabeled training image and a second unlabeled training image having similarity to the primary feature; the total loss includes a primary loss and a secondary loss; the primary loss is calculated based on the difference between the one or more primary ground truth values ​​of the labeled training image and one or more primary predicted values, the one or more primary predicted values ​​being one or more values ​​of the primary feature generated by the prediction model; and the secondary loss is calculated based on the difference between one or more first predicted values ​​of the first unlabeled training image and one or more second predicted values ​​of the second unlabeled training image, the one or more first predicted values ​​and the one or more second predicted values ​​being one or more values ​​of the primary feature generated by the prediction model.

2. The method of claim 1, further comprising reducing the total loss of the prediction model by adjusting the plurality of parameters in the prediction model, and training the prediction model to generate one or more auxiliary predicted values ​​of one or more auxiliary features of the input image, wherein: The one or more auxiliary features are associated with the primary feature; the labeled training image of each of the multiple primary learning data is further labeled with one or more auxiliary ground values ​​of the one or more auxiliary features; the total loss further includes a three-dimensional loss; and the three-dimensional loss is calculated based on the difference between the one or more auxiliary ground values ​​of the labeled training image and one or more three-dimensional predicted values, which are one or more values ​​of the one or more auxiliary features generated by the prediction model.

3. The method of claim 1, wherein the labeled training image is modified by image enhancement before the prediction model generates the one or more first-order predictions.

4. The method of request item 1, wherein the one or more principal truth values ​​are modified by truth-enhancing before the first-order loss is calculated.

5. As in request item 1, wherein the first-order loss and the second-order loss are calculated using a squared loss function.

6. The method of claim 1, wherein in each of the plurality of secondary learning materials, the first unlabeled training image and the second unlabeled training image are images of the same subject taken at a predetermined time interval so that their main features are similar.

7. The method of request item 6, wherein the predetermined time interval is 3 months.

8. The method of claim 1, wherein each of the labeled training image, the first unlabeled training image, and the second unlabeled training image is a ROI-captured image captured by a computer from an original training image via a ROI (Region of Interest).

9. The method of claim 1, wherein each of the labeled training image, the first unlabeled training image, and the second unlabeled training image is a training image set comprising: an original training image; and an ROI capture image captured by a computer from the original training image via an ROI (Region of Interest).

10. The method of claim 1, wherein the primary feature is the bone mineral density of a subject, and the one or more primary predicted values ​​are one or more bone mineral density (BMD) values.

11. The method of claim 10, wherein the one or more principal predicted values ​​include bone mineral density (BMD) values ​​of the entire hip, femoral neck, greater trochanter of the femur and femoral shaft.

12. The method of claim 10, wherein each labeled training image in the first learning data and each first unlabeled training image and second unlabeled training image in the second learning data are X-ray images.

13. The method of claim 12, wherein the first unlabeled training image and the second unlabeled training image are two X-ray images taken consecutively within 3 months of the same subject.

14. The method of claim 10, further comprising training the prediction model to generate one or more auxiliary predicted values ​​of one or more auxiliary features of the input image by adjusting the plurality of parameters in the prediction model to reduce the total loss of the prediction model, wherein: The one or more auxiliary features are associated with the bone density of a subject; the labeled training image of each of the plurality of first-order learning data is further labeled with one or more auxiliary ground truth values ​​of the one or more auxiliary features; the total loss further includes a third-order loss; and the third-order loss is calculated based on the difference between the one or more auxiliary ground truth values ​​of the labeled training image and one or more third-order predicted values, which are the values ​​of the one or more auxiliary features generated by the prediction model.

15. The method of claim 14, wherein the one or more auxiliary features include the cortical bone thickness of the subject, and the one or more auxiliary predicted values ​​include the cortical thickness index (CTI) value of the subject.

16. The method of claim 14, wherein the one or more auxiliary features include the femoral neck width of the subject, and the one or more auxiliary predicted values ​​include the femoral neck width (FNW) value of the subject.

17. The method of claim 12, wherein the labeled training images are a training image set comprising: An original training image; and a region of interest (ROI) extraction image, which is an identified ROI region of a hip joint extracted from the original training image.

18. The method of claim 17, wherein the original training image and the ROI-captured image are modified by image enhancement before the prediction model generates the one or more first-order predictions.

19. The method of claim 18, wherein the image enhancement is performed by cropping 0-25% of the original training image without cropping to the identified ROI region, and wherein the image enhancement is performed by moving the identified ROI region by 0-7% in a specific direction.

20. The method of request item 12, wherein the one or more principal truth values ​​are modified by introducing a small variable randomly selected between -0.01 g / cm2 and 0.01 g / cm2.

Citation Information

Patent Citations

  • Osteoporosis intelligent evaluation method based on lumbar vertebra X-ray image

    CN112396591A

  • Semi-supervised learning method and device for bone mineral density estimation in hip X-ray image and storage medium

    CN116830121A

  • Bone mineral density detection image processing method and system

    CN117952962A

  • Methods of grading and monitoring osteoarthritis

    TW202341174A

  • System and method for bone fracture risk assessment

    WO2024110991A1