A method for generating highly consistent predictions from planar images

By training AI models with labeled and unlabeled image datasets and applying data augmentation, the method enhances the consistency of BMD predictions, addressing the variability issue and enabling effective continuous monitoring of osteoporosis and fracture risk.

JP7730517B1Active Publication Date: 2025-08-28ALPHA INTELLIGENCE MANIFOLDS INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024119204
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-06-06
Filing Date
2024-07-25
Publication Date
2025-08-28
Estimated Expiration
2044-07-25

AI Technical Summary

Technical Problem

Existing AI models for predicting bone mineral density (BMD) from radiological images lack consistency, leading to significant variations in output values for similar inputs, making them unsuitable for continuous monitoring of osteoporosis and fracture risk.

Method used

A method involving a predictive model trained with a first dataset of labeled images and a second dataset of unlabeled image pairs with similar features, using a combined loss function to enhance consistency, along with data augmentation and region of interest extraction to focus on key features.

Benefits of technology

The method improves the consistency of BMD predictions, reducing variability to within acceptable clinical precision limits, enabling reliable continuous monitoring of osteoporosis and fracture risk.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007730517000001_ABST
    Figure 0007730517000001_ABST
Patent Text Reader

Abstract

The present invention relates to a method for training a predictive model to generate key predictions of key features of an input image. The method includes training a predictive model using a first dataset containing labeled training images labeled with ground truth values ​​and a second dataset containing two unlabeled training images without ground truth values. The training goal is to mitigate both a first loss and a second loss, where the first loss calculates the difference between the predicted values ​​of the labeled training images and the ground truth values, and the second loss calculates the difference between the predicted values ​​of two unlabeled training images.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for generating highly consistent predictors from planar images, and in particular to a method for generating skeletal characteristics related to a subject's osteoporosis and fracture risk from radiological images. [Background technology]

[0002] According to the World Health Organization (WHO), osteoporosis is a systemic bone disease characterized by a decrease in bone mass and deterioration of bone tissue microarchitecture, leading to bone fragility and an increased risk of fracture. Patients may experience multiple complications after a fracture. Osteoporosis is usually painless and has no obvious symptoms, so to know if you have osteoporosis, you need to undergo testing.

[0003] Osteoporosis can be diagnosed by measuring bone mineral density (BMD). Besides standard methods such as dual-energy X-ray absorptiometry (DXA), other AI-related methods that generate BMD values ​​are also available.

[0004] Because BMD values ​​can be used as a treatment efficacy or follow-up indicator or as a basis for continuous monitoring of a patient's bone quality, they are often performed continuously, requiring a certain level of consistency. According to guidelines issued by the International Society for Clinical Bone Densitometry (ISCD), when the same examiner performs repeated measurements consecutively (usually 30 people, each measured twice consecutively), the precision of DXA, as measured by the coefficient of variation (CV), should be within the range of 1.8% to 2.5%.

[0005] Currently, AI is used to predict bone mineral density (BMD) values, primarily for screening to identify osteoporosis in subjects, rather than for continuous monitoring. This is primarily because the BMD estimates produced by AI models for different X-ray images of the same person taken within a short period (usually 3-6 months) can vary widely (e.g., much more than the 1.8-2.5% range specified by the ISCD). However, due to the difficult-to-interpret nature of AI models, it is difficult to control for these differences resulting from different AI model training methods. This makes it difficult to use AI models to estimate BMD from continuous measurements.

[0006] Therefore, new methods to improve the consistency of AI outputs, particularly the consistency of skeletal characteristic values ​​related to subjects' osteoporosis and fracture risk, remain desirable. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] U.S. Patent Publication No. 2023 / 0029674A1 Summary of the Invention [Problem to be solved by the invention]

[0008] To solve this problem, the present invention provides a method for generating key feature values ​​from planar images with high consistency. In this method, a second loss term is introduced in addition to the conventional first loss term during model training, thereby improving the stability of the model's prediction results. [Means for solving the problem]

[0009] Specifically, the present invention provides a method for training a predictive model to generate one or more key predicted values ​​of key features of input images, the method comprising: training the predictive model using a first dataset and a second dataset by adjusting a plurality of parameters of the predictive model to reduce a total loss of the predictive model. In this method, the first dataset includes a plurality of first training data, each of which includes labeled training images labeled with one or more key ground truth values ​​of the key features; and the second dataset includes a plurality of second training data, each of which includes a pair of unlabeled training images, each of which includes a first unlabeled training image and a second unlabeled training image that have similarity in the key features. The total loss used to adjust the plurality of parameters includes a first loss and a second loss. During training, the predictive model may generate one or more first predicted values ​​of the key features of the labeled training images, and the first loss is calculated based on a difference between the one or more key ground truth values ​​and the one or more first predicted values. Similarly, the predictive model may generate one or more first predicted values ​​of key features of a first unlabeled training image and one or more second predicted values ​​of key features of a second unlabeled training image. A second loss is calculated based on the difference between the one or more first predicted values ​​and the one or more second predicted values. In one embodiment, the first loss is calculated by a first loss function, and the second loss is calculated by a second loss function. In particular embodiments, the first loss function and the second loss function calculate a squared loss.

[0010] In one embodiment, to ensure similarity of key features, in each of the plurality of second training data, the first unlabeled training image and the second unlabeled training image are images of the same subject taken within a predetermined time interval, during which the key features are known to be constant or to vary only within the measurement limits of the measurement method. Specifically, the predetermined time interval may be three months or six months.

[0011] Instead of using the original training images, in one embodiment, each of the labeled training images, the first unlabeled training images, and the second unlabeled training images is a region of interest (ROI) extracted image extracted from the original training image through ROI extraction. An ROI is a region closely related to a key feature. The ROI extraction may be performed by a trained ROI extraction model. Alternatively, each of the labeled training images, the first unlabeled training images, and the second unlabeled training images may be a training image set including the original training image and an ROI-extracted image, and the prediction model may use both images of the image set to predict key feature values.

[0012] In one embodiment, one or more auxiliary features are used in addition to the primary features, and the method further includes training the predictive model to generate one or more auxiliary predicted values ​​for the one or more auxiliary features of the input image by adjusting multiple parameters of the predictive model to reduce a total loss of the predictive model. To use the auxiliary features in training the model, the one or more auxiliary features must be correlated with the primary features. In this embodiment, each labeled training image of the multiple first learning data is further labeled with one or more auxiliary ground truth values ​​for the one or more auxiliary features. The total loss used to adjust the multiple parameters further includes a third loss. During training, the predictive model may further generate one or more third predicted values ​​for the one or more auxiliary features of the labeled training image, and the third loss is calculated based on a difference between the one or more auxiliary ground truth values ​​and the one or more third predicted values.

[0013] During training, data augmentation may also be applied to the predictive model. In one embodiment, labeled training images are modified by image augmentation before generating one or more first predictions in the predictive model. In one embodiment, one or more primary ground truth values ​​are modified by ground truth augmentation before calculating the first loss.

[0014] In a preferred embodiment, the primary characteristic is the subject's bone mineral density, and the one or more primary predictors are one or more bone mineral density (BMD) values. In one particular embodiment, the one or more primary predictors include bone mineral density (BMD) values ​​of the total hip, femoral neck, greater trochanter, and femoral shaft.

[0015] In certain embodiments, the predictive model is trained to generate BMD values ​​from X-ray images, in which each labeled training image of the first training data and each first and second unlabeled training image of the second training data are X-ray images.

[0016] In one embodiment, the first unlabeled training image and the second unlabeled training image are two consecutive x-ray images of the same subject taken within three months.

[0017] The model for predicting BMD may use one or more auxiliary features to improve the model's ability to generate BMD values. In this case, the method may further include training the predictive model to generate one or more auxiliary predicted values ​​of one or more auxiliary features of the input image by adjusting multiple parameters of the predictive model to reduce the total loss of the predictive model. To use the auxiliary features in training the model, it is necessary that the one or more auxiliary features correlate with the subject's bone mineral density. The one or more auxiliary features may include the subject's cortical thickness, and the auxiliary predicted value corresponding to the cortical thickness is the subject's Cortical Thickness Index (CTI) value. The one or more auxiliary features may also include the subject's femoral neck width, and the auxiliary predicted value corresponding to the femoral neck width is the subject's Femoral Neck Width (FNW) value.

[0018] Typically, X-ray images taken in medical facilities are images of the entire pelvic region. To enable the predictive model to focus on areas closely related to key features, ROI (region of interest) extracted images may be used as training images instead of the entire X-ray images. In this embodiment, the labeled training images are ROI (region of interest) extracted images, which are identified ROI regions of the hip joint extracted from the original training images. Alternatively, to retain more information of the original training images, the labeled training images may be a training image set including the original training images and ROI extracted images.

[0019] In one particular embodiment, the original training images and the ROI-extracted images are modified by image augmentation before generating one or more first predictions with the predictive model. The image augmentation may be performed by cropping 0-25% of the original training images without cropping the identified ROI region, or by shifting the identified ROI region in a particular direction by 0-7%.

[0020] In one particular embodiment, one or more primary ground truth values ​​are modified by ground truth extension before calculating the first loss. The ground truth extension may be performed by introducing small variables to one or more primary ground truth values. The values ​​of the small variables may be independent of the primary ground truth values. In one embodiment, each of the small variables has a value of -0.01 g / cm 2 ~0.01g / cm 2 In one particular embodiment, each of the minor variables is a value randomly selected between ±0.01 g / cm 2 A normal distribution is randomly selected from a population mean of 0, truncated at 0.01 g / cm 2 Alternatively, the value of the small variable may vary with each of the primary ground truth values. In one embodiment, each of the one or more primary ground truths has a value y n and each of the smaller variables is -0.01y n ~0.01y nis a value chosen randomly between

[0021] Other objects, advantages and novel features of the present invention will become more apparent from the following detailed description when considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0022] [Figure 1] FIG. 1 illustrates an architecture for training a predictive model using labeled and unlabeled datasets. [Figure 2] FIG. 1 illustrates an architecture for training a predictive model using a labeled dataset labeled with both primary and secondary features. [Figure 3] FIG. 1 illustrates an architecture for training a predictive model with augmented data. [Figure 4] 1 illustrates a prediction model coupled with an ROI extraction model, which takes the extracted ROI as input to generate a prediction. [Figure 5A] A training image of a full hip joint, including both left and right hips, labeled by a radiologist. [Figure 5B] 1 is a training image of the hip joint region including the femoral neck, greater trochanter, and femoral shaft labeled by a radiologist. [Figure 5C] FIG. 10 shows the results of identifying ROIs for BMD generation, with the hip joint region identified as the region of interest. [Figure 6] FIG. 10 shows an example of incorporating data into a model, where a first set of training data randomly selected from a first dataset and a second set of training data randomly selected from a second dataset are used in combination to calculate total loss. [Figure 7] FIG. 1 illustrates an example of calculating total loss including accuracy and precision terms. [Figure 8] FIG. 10 is a diagram showing feature points corresponding to the outer and inner edges of the cortical bone used when calculating the CTI value of the CTI calculation model. [Figure 9]1 shows an example of performing random displacement on a ROI, where the ROI region is expanded from the center to all four sides before cropping the image. [Figure 10] 1 illustrates an example of data augmentation. Data augmented from a first dataset includes an augmented image (P1') and augmented ground truth values ​​(Bh1', Bfn1', Bgt1', Bfs1'). DETAILED DESCRIPTION OF THE INVENTION

[0023] The terms used in the description below are intended to be interpreted in their broadest reasonable manner, even when used in conjunction with detailed descriptions of certain specific embodiments of the technology. Although certain terms may be emphasized below, any terms intended to be interpreted in a limiting manner will be specifically defined to that effect in this detailed description section.

[0024] The embodiments described below may be implemented entirely by programmable circuitry that is programmed or configured with software and / or firmware, or entirely by special purpose circuitry, which (if any) may take the form of, for example, one or more application specific integrated circuits (ASICs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), etc.

[0025] The objective of this invention is to provide a method for training an AI model to generate key predictions for key features of input images with high accuracy. In other words, the trained model should be able to generate consistent results for input images with approximately the same values ​​for the key features, regardless of noise in the image. This improvement can be achieved by introducing an additional "accuracy term" into the loss function during model training, as described in more detail below.

[0026] An AI model may learn to generate output values ​​from input images by training on a labeled training dataset. The model may be built using convolutional neural network (CNN)-based algorithms such as LeNet, AlexNet, VGG, GoogLeNet, ResNet, and DenseNet. Transformer-based vision algorithms (e.g., ViT) may also be used. Each training dataset in the labeled training dataset contains training images labeled with their ground truth values ​​as learning targets for the model to learn from. However, AI models trained in this manner may produce highly variable outputs for similar inputs (e.g., two or more inputs may be considered to have nearly identical ground truth values). This type of variation is difficult to control due to the difficult-to-interpret nature of AI models. One way to address this issue is to increase the size of the labeled training dataset by including more training data. However, a large amount of labeled training data is not always available, and models cannot often be trained using this method.

[0027] While obtaining large amounts of labeled training data is difficult, image sets containing unlabeled images with some common characteristics or similarity in key features are much easier to obtain because they do not require labeling. Therefore, the model may be trained with additional unlabeled training datasets. Each training dataset in the unlabeled training datasets is an image set containing two or more training images in which the values ​​of key features are considered to be essentially the same. The learning goal for the model from the unlabeled training datasets is to generate output with low variance between training images within any one image set. By applying this additional goal, the model may be trained to generate consistent output for input images with similar key features, as shown in Figure 1.

[0028] Key feature values ​​of two or more images may be considered "same" or "indistinguishable" if the estimated intrinsic difference between those values ​​is much smaller than the precision error required to obtain those values. For example, two or more images taken within a short period of time is a commonly used method for obtaining image pairs / sets with similar key features. In a preferred embodiment, the key features are measurable human biological features. These may include, but are not limited to, height, weight, serum albumin, and bone mineral density. Although the values ​​of these features may change over time, a "short period" can be found for each feature value, within which the natural change in the feature value is always smaller than the detection limit or measurement error of the corresponding measurement method. Because the measurement error exceeds the true difference between two feature values ​​measured within a short period of time, the measurement method cannot analyze the true difference and the feature values ​​must be considered "indistinguishable" or "similar." The length of the short period is predetermined based on knowledge of (1) the maximum rate of change of the measured feature and (2) the measurement error or detection limit of the measurement method used to obtain the feature value, such that the product of the maximum rate of change and the predetermined short period is less than the measurement error or detection limit of the measurement method. In other words, any two feature values ​​measured by the same person over a predetermined short period will be similar to each other.

[0029] Specifically, bone mineral density (BMD) tests measured by dual-energy X-ray absorptiometry (DXA) have a variability of approximately 2% between repeated measurements by the same operator. However, research has shown that BMD change for an individual is typically less than 1% per year. Therefore, two X-ray images of the same individual taken within three or six months can reasonably be considered "similar," because the inherent BMD change for the same individual within six months is much less than 2%, which is not a measurable difference with DXA, given the precision error of repeated measurements. Other examples include measuring the weight of the same individual within a day or two, or measuring the serum albumin concentration of the same individual within a few hours; in both cases, the precision error can be greater than the inherent difference itself.

[0030] In addition to training a model to generate only one output from an input, the model may be trained to generate multiple outputs from an input image. If the training images are labeled with multiple ground truth values, the model may learn to simultaneously generate those values ​​as multiple outputs. For example, the model may be trained to generate only one value for one primary feature as an output, or it may be trained to generate multiple values ​​that include one or more values ​​for the primary feature and several values ​​for the secondary features. As shown in Figure 2, one or more secondary features that correlate with the primary feature may be selected, and the model may be trained to generate values ​​for the secondary features.

[0031] Training a model to generate outputs other than primary feature values ​​can provide additional guidance for tuning parameters to output primary feature values ​​because primary features and selected auxiliary features are interrelated. Generating several values ​​for the primary features (rather than only one primary feature value) can also improve model performance because the model can learn more feature details from generating multiple values. For example, research has shown that the cortical thickness index (CTI) and femoral neck width (FNW) values ​​of the hip joint correlate with bone mineral density (BMD) values ​​in the hip region. Therefore, training a model to predict CTI and FNW values ​​in addition to BMD can improve the model's performance in predicting BMD values. This is because the three hip joint values ​​may share some common factors that the model may learn during training.

[0032] Various types of data augmentation techniques may be applied to increase the variability in the training data to improve model generalization and performance. Data augmentation may include input image augmentation and / or ground truth augmentation, as shown in FIG. 3. Because data augmentation creates artificial data other than the actual input / measurements, domain knowledge is required to create reasonable augmented data. "Reasonable" augmented data refers to data that could have actually been acquired during data acquisition. For example, a reasonable augmented input image may correspond to variations during image acquisition, and a reasonable augmented ground truth may correspond to variations in measurements caused by instrument noise or operational noise.

[0033] In image augmentation, the original input image may be modified through geometric transformations, color space conversions, and noise injection. Geometric transformations include rotation (rotating the image by a specified angle), flipping (flipping the image horizontally or vertically), cropping (removing portions of the original image), and shifting (shifting the image in various directions). Color space conversions may modify image color characteristics, such as lighting, color saturation, and contrast. Additionally, noise may be introduced into the image to simulate real-world imperfections. The augmented image represents a copy of the original input image with slight modifications to have the same ground truth values ​​as the original input image. For radiological images, these adjustments mimic variations in image acquisition, such as patient posture, position, or misalignment, instrument X-ray energy, and operator differences. The augmentation parameters applied during training are based on the principle of conforming to variations that may occur when capturing conventional X-ray images. Examples include scaling / translating within 5% of the image size, rotating less than 15 degrees, gamma correction within 25% to mimic over / underexposure, introducing random Gaussian blur to mimic focus deviations, and introducing random Gaussian blur to mimic sensor noise.

[0034] In ground truth augmentation, the ground truth values ​​of the input image are slightly modified within a reasonable range to mimic the variability of repeated measurements of the same subject. In general, any variation smaller than the measurement precision can potentially be used in ground truth augmentation to mimic the variability caused by repeated measurements. For example, for BMD measurements by DXA, a reasonable coefficient of variation (CV) for repeated measurements of the same patient is approximately 2%. Therefore, adjusting a specific BMD value within a 1% range can be considered an acceptable modification to create artificial data points corresponding to the original data for model training.

[0035] In summary, one original image and its corresponding ground truth value may be used to generate multiple modified images and multiple modified output values, and the combination of multiple modified images and multiple modified output values ​​may effectively expand the training dataset for model training.

[0036] The framework of an AI model may be designed based on the characteristics of the input image. It may be a single model that predicts one or more output values ​​from the input image. Alternatively, it may be an integrated model that combines two or more submodels, each of which handles a specific task. For example, a predictive model may be combined with a region of interest (ROI) localization / extraction model, as shown in Figure 4, so that the ROI localization model can extract key regions of the input image for the predictive model. This allows the predictive model to focus on regions closely related to key features. The predictive model may receive the extracted ROI as input, or it may receive both the original image and the extracted ROI as input. In the latter case, the predictive model uses an image set containing both the original image and the extracted ROI as input.

[0037] Specifically, an integrated AI model for predicting a subject's bone mineral density (BMD) value from planar X-ray images may be implemented by coupling an ROI localization model with a BMD generation model. To train the integrated AI model using such a framework, an ROI localization model may first be independently established. The established ROI localization model may then process the original training X-ray images and output ROIs related to BMD in the original training images (e.g., total hip joint, femoral neck, greater trochanter, and femoral shaft). The identified ROIs are then used as training inputs for the BMD generation model to adjust its parameters. The parameters of the ROI localization model are not changed when training the BMD generation model. Performing ROI localization first before BMD generation may have advantageous effects compared to a single model that predicts BMD values ​​directly from input X-ray images, because the identified regions of interest (ROIs) force the BMD generation model to focus on key features related to BMD values.

[0038] The following provides an example of establishing an integrated model for BMD generation based on the above concepts. The effectiveness of the established model is also described. In some examples, the integrated model includes an ROI localization model and a BMD generation model. The ROI localization model uses an X-ray image including at least one hip joint as the original input image of the model to extract the hip joint region. The extracted ROI may be used as an input for the BMD generation model to generate a predicted BMD value.

[0039] 1.ROI Localization Model An object detection AI model described in U.S. Patent Publication No. 2023 / 0029674A1 (the contents of which are incorporated herein by reference) is trained to identify regions of interest (ROIs) in input X-ray images. A deep neural network (DNN) model, the You Only Look Once (YOLO) algorithm, is implemented to train the AI ​​to identify features and select appropriate ROIs from the input images. The training dataset and training workflow for the ROI localization model are as follows: (1) To prepare the training data, ROIs in X-ray images are labeled by experts as the ground truth for training the object detection model. Clinical data of 459 pelvic X-ray images from a medical facility are converted from DICOM files to high-resolution 16-bit PNG files. Images with abnormal bones, such as artificial joints and fractures, are excluded. Next, the brightness and contrast of the images are adjusted to a standard range. Finally, the regions of the hip joint, femoral neck, greater trochanter, and femoral shaft are labeled by radiologists as the training ground truth, as shown in Figures 5A and 5B. (2) Using the labeled ROIs (bounding boxes in Figures 5A and 5B) and the corresponding unlabeled X-ray images, we train an object detection model implementing the YOLO algorithm. (3) After training, the model can localize ROIs in the incoming pelvic X-ray image. The ROI for the hip joint region is located from the original X-ray image, and the ROIs for the femoral neck, greater trochanter, and femoral shaft are located from the ROI for the hip joint region.

[0040] Figure 5C shows the result of ROI identification for BMD generation, where the hip joint region is selected as the region of interest. The trained ROI localization model can also recognize the femoral neck, greater trochanter, and femoral shaft regions. However, in the present invention, the hip joint region is the only ROI required for BMD generation. The hip joint region is identified from the original input image, and the hip joint ROI is used as the input for the BMD generation model described below.

[0041] 2.BMD generation model An AI model that generates BMD from input X-ray images is described in U.S. Patent Publication No. 2023 / 0029674A1 (Patent Document 1), which is incorporated herein by reference. In this invention, the AI ​​model is trained to generate not only bone mineral density (BMD) but also cortical thickness index (CTI) and femoral neck width (FNW) values. BMD is the amount of bone mineral in bone tissue, CTI is defined as the ratio of the cortical width minus the endosteal width at a level 100 mm below the tip of the lesser trochanter to the cortical width, and FNW is defined as the midpoint distance between the upper and lower cortices of the femoral neck perpendicular to the femoral neck axis. Research has shown that BMD, CTI, and FNW are correlated, so the BMD generation model is trained to simultaneously predict the values ​​of all the above features to improve accuracy. Here, BMD is the primary feature learned by the AI ​​model, while CTI and FNW are secondary features. The BMD generation model may be trained to generate only one BMD value or multiple BMD values ​​corresponding to different regions. For example, the model may be trained to generate only the BMD value for the whole hip joint, or it may be trained to generate four BMD values ​​for the analyzed hip joint regions (the whole hip joint, the femoral neck, the greater trochanter, and the femoral shaft). In this example, a ResNet algorithm, RegNetY160, was used to train the model.

[0042] Before training the BMD generative model, two datasets were constructed: the first dataset containing labeled training data, and the second dataset containing unlabeled training data. To construct the first dataset, clinical data of 3,169 pelvic X-ray images with corresponding DXA measurements from a medical facility were converted from DICOM files to high-resolution 16-bit PNG files. Images with abnormal bones, such as artificial joints or fractures, were excluded. The image brightness and contrast were normalized using a histogram normalization method. For each X-ray image, the corresponding DXA report containing BMD subvalues ​​for four regions (total hip, femoral neck, greater trochanter, and femoral shaft) was matched to the X-ray image. Only matches with a time interval of less than six months between the X-ray and DXA measurement were included as training data. The DXA measurements included BMD subvalues ​​for the total hip, femoral neck, greater trochanter, and femoral shaft for each X-ray training image. In addition to the BMD sub-values, for each X-ray image, the ground truth values ​​of CTI and FNW are labeled by experts or a suitable AI model. After construction, each first training dataset of the first dataset includes one original X-ray image and six ground truth values ​​(total hip BMD, femoral neck BMD, greater trochanter BMD, femoral shaft BMD, CTI, and FNW). In one embodiment, the ROI localization model described above is applied to extract the total hip region of the original X-ray image, and the extracted ROI (rather than the original X-ray image) is used as the input image of the first dataset. In yet another embodiment, both the original X-ray image and the extracted ROI are used as input images.

[0043] To construct the second dataset, clinical data containing 3,215 pairs of X-ray images without BMD ground truth were collected from multiple medical institutions, and 16-bit image pixel arrays were extracted directly from DICOM files. Each pair of X-ray images constituted the second training data set, containing two corresponding images of the same subject presumed to have similar BMD values. According to Berger et al. (CMAJ.2008 Jun 17, 178(13):1660-8), the typical change in a person's BMD is 0.01 g / cm per year. 2The variability of repeated DXA measurements by the same operator is approximately 2% (approximately 0.02 g / cm 2 ), two X-ray images of the same subject taken within six months can reasonably be considered "similar," since any measurable differences are far less than the variability between repeated measurements. Thus, in the second dataset, each set included in the second training data contains two X-ray images of the same person taken within six months. As with the first dataset, in one embodiment, an ROI localization model is applied to extract the entire hip joint region of the original X-ray images, and the extracted ROI is used as the input image for each pair of X-ray images. In yet another embodiment, both the original X-ray image and the extracted ROI are used as input images.

[0044] Model training may use various techniques to incorporate data into the model. In one embodiment, as shown in FIG. 6, a first set of training data randomly selected from a first dataset and a second set of training data randomly selected from a second dataset are used in combination to calculate the total loss. After selection, the training data from the first dataset undergoes further image augmentation and ground truth augmentation to improve the model's generalization ability. Data augmentation methods are described in detail below. The random selection and augmentation may be performed on the fly during model training, or may be pre-created and cached in memory to improve loading speed during model training. In one modified embodiment, the second training data from the second dataset also undergoes image augmentation before training.

[0045] To train a predictive model to generate one or more BMD values, one CTI value, and one FNW value, the total loss function includes a BMD accuracy term, a CTI accuracy term, and an FNW accuracy term, and applies the BMD precision term to calculate the total loss. The goal of training is to attempt to reduce the calculated total loss as much as possible by adjusting the parameters of the BMD generation model. In one embodiment, the BMD includes four BMD values ​​(total hip, femoral neck, greater trochanter, and femoral shaft BMD), so the total loss function calculates the BMD accuracy and precision for all four BMD values.

[0046] FIG. 7 shows an example of calculating the total loss. The first training data selected includes an X-ray image P1, and the BMD value of the whole hip joint is B h1 , femoral neck BMD is B fn1 , the BMD of the greater trochanter is B gt1 , BMD of the femoral shaft is B fs1 , the CTI value is C1 and the FNW value is F1. The second training data selected includes a pair of X-ray images P2 and P3, and there is no labeled ground truth. The selected training data is fed into the model to generate the corresponding output. For P1, four predicted BMD TIFF0007730517000002.tif8150 and predicted CTI TIFF0007730517000003.tif8150 and predicted FNW TIFF0007730517000004.tif8150 is generated. For P2, four predicted BMD TIFF0007730517000005.tif8150 was generated, and four predicted BMDs were obtained for P3. TIFF0007730517000006.tif8150 is generated. The total loss includes an accuracy term calculated from the first training data and a precision term calculated from the second training data. In this example, the accuracy term calculates the difference between the predicted value and the ground truth, which has a first loss to calculate the difference in BMD (the main feature) and a third loss to calculate the difference in CTI and FNW (auxiliary features). The precision term is the difference between the two sets of predicted values. TIFF0007730517000007.tif8150 and Calculate the BMD value between TIFF0007730517000008.tif8150. In Figure 7, we calculate the loss of the prediction results using squared errors such as mean squared error. However, other loss functions such as mean absolute error or mean absolute percentage error can also be used. In addition, the weights of each component of each loss, as well as the weights of the first, second, and third losses, are all set to 1 here, but these weights can be adjusted based on actual needs or by optimizing hyperparameters.

[0047] The model may then be trained using the first and second datasets. In one example, for each training input, first training data from the first dataset and second training data from the second dataset are randomly selected, and the first training data is further subjected to image augmentation and ground truth augmentation. As described above, in one embodiment, the ROI localization model extracts ROIs from X-ray images in the training data, and for each training data, the model is trained using both the original image and the extracted ROI as input. During training, the batch size is set to 4 to 32 for each iteration. The total number of training iterations is 10 to 500, and the training is stopped when the loss no longer decreases.

[0048] 3.CTI calculation model The femoral cortical thickness has been found to be positively correlated with the BMD values ​​of the femoral neck, femoral shaft, greater trochanter, and total hip joint. To calculate the ground truth CTI value of an X-ray image, a CTI computation model is described in U.S. Patent Publication No. 2023 / 0029674A1, the contents of which are incorporated herein by reference. Briefly, this model is used to find feature points corresponding to the outer and inner edges of the cortical bone, as indicated by the landmarks A, B, C, and D in Figure 8. The training method for the CTI computation model is as follows. (1) Prepare 153 hip joint X-ray images containing feature points (i.e., points A, B, C, and D in Figure 8) labeled by a radiologist as ground truth for training. (2) Train a High Resolution Net (HRNet) algorithm to recognize feature points. (3) A trained AI model, which is a feature point detection model, can be used to find feature points in the input X-ray image.

[0049] Therefore, after the above training, the CTI value can be easily calculated using the following formula: CTI = (AB - CD) / AB (percent).

[0050] The CTI computation model is only used to provide the CTI ground truth values ​​to the model to calculate the loss during training of the BMD generative model (rather than generating CTI values ​​as input for the BMD generative model). This is because the BMD generative model is trained to generate CTI values ​​as output and does not consider the values ​​as input. This design allows the BMD generative model to focus on the correlation between CTI values ​​and BMD values ​​during training, rather than passively receiving CTI values.

[0051] 4. Data Augmentation As mentioned above, the training data from the first dataset is further subjected to image augmentation and ground truth augmentation to improve the generalization ability of the BMD generative model. In image augmentation, the input image is slightly modified by operations such as cropping, shifting, zooming in, zooming out, adjusting brightness and / or contrast, etc. Since the modifications to be applied are selected randomly, multiple modified images may be generated from the input image.

[0052] In one example, image augmentation involves data augmentation of the original image and the identified hip joint ROI. For the original x-ray image, data augmentation may randomly crop up to 25% of the input x-ray image, assuming the ROI remains uncropped, thereby preventing the model from overfitting to the image capture and cropping characteristics of a specific medical institution. For ROI data augmentation, because the identified ROI box may vary slightly depending on the medical institution or imaging modality, a random displacement of 0-7% may be performed on the identified ROI, thereby mitigating adaptation and generalization issues caused by ROI identification errors. Figure 9 shows an example of performing random displacement on the ROI. In this example, the ROI area is expanded by 7% from the center on all four sides. The expanded ROI is then randomly cropped. The cropped ROI may be the same size as the originally identified ROI to simulate pure displacement. Alternatively, the cropped ROI may be smaller or larger than the originally identified ROI to simulate new variations in the ROI. In addition to cropping the image, random scaling and rotation may be applied to simulate changing shooting conditions.

[0053] In ground truth augmentation, the ground truth BMD values ​​(BMD of the total hip, femoral neck, greater trochanter, and femoral shaft) are slightly modified within a small deviation. Similar to image augmentation, multiple modified BMD values ​​may result from the training ground truth BMD values.

[0054] As previously mentioned, a small variation may be introduced into the original BMD value. In one example, the small variation introduced is ±0.01 g / cm 2 Alternatively, if the original BMD value is y n If so, the small variable introduced is ±0.01y n , and a randomly selected value between . More complex probability density functions may also be applied. In one example, ground truth augmentation is performed by introducing an appropriate normally distributed variable into the original BMD values. Specifically, the variable introduced has a population mean of 0 and is within the range of ±0.01 g / cm. 2 The standard deviation of the normal distribution may be varied to change the distribution of the added variable. In one embodiment, the standard deviation of the applied normal distribution is 1 g / cm. 2 In another embodiment, the standard deviation of the applied normal distribution is set to 0.01 g / cm 2 , which further concentrates the variables in the central region. This variation is applied to account for measurement errors that may occur in DXA photography due to positioning, operator proficiency, and / or equipment calibration. The purpose of this step is to simulate real-world measurement variation and prevent the model from overfitting to unobvious systematic errors that may be present in the training data, thereby improving the model's generalization ability.

[0055] Finally, the image augmentation and ground truth augmentation methods are combined to generate training data that includes both augmentations. Each randomly selected first training data set contains one input image and one set of ground truth values. The input image undergoes ROI recognition to generate an ROI image. The input image and the hip joint ROI then undergo image augmentation as described above to generate a randomly modified input and a randomly modified ROI. The ground truth BMD values ​​also undergo ground truth augmentation as described above to generate randomly modified BMD values ​​(except for the original CTI and FNW values). Next, as shown in Figure 10, the modified input image, modified ROI, and ground truth values ​​including the modified BMD are used as training material for the selected first training data. Although the augmented data is artificial, it shows good representativeness to real-world data obtained by DXA measurements.

[0056] 5. Model Performance To evaluate the performance of the trained models, we compare them under various training conditions. The full model is a predictive model trained using the loss of accuracy terms, image augmentation, and ground truth augmentation, and produces four BMD, one CTI, and one FNW output. The contributions of adding an accuracy term (second loss term) to the loss function, applying data augmentation to the original input X-ray image (global image augmentation), applying data augmentation to the ROI image (local image augmentation), and applying ground truth augmentation to the BMD ground truth (ground truth augmentation) are analyzed separately by removing each of the operations one by one from the full model.

[0057] The test data included X-ray images from 588 individuals. Of the 588 individuals, 420 individuals had two X-ray images taken within a short period of time, 77 individuals had three X-ray images taken, 46 individuals had four X-ray images taken, 31 individuals had five X-ray images taken, and 14 individuals had six or more X-ray images taken. The accuracy error is calculated using the method described in section TIFF0007730517000009.tif18167. The results are shown in Table 1.

[0058] [Table 1]

[0059] The model including all manipulations has a CV of 2.78%. Removing the accuracy term during training increases the CV from 2.78% to 3.69%. Omitting data augmentation for the original X-ray images in the first dataset increases the CV to 3.42%. Omitting data augmentation for the ROI images in the first dataset increases the CV to 3.20%. Omitting output augmentation for the BMD ground truth values ​​increases the CV to 2.98%. We can see that each of the manipulations improves the model's performance; however, adding an accuracy loss term during training is most effective. This is consistent with our assumption that using unlabeled image pairs will encourage the model to produce more consistent results for similar inputs.

[0060] The foregoing description of the embodiments is provided to enable one skilled in the art to make and use the subject matter of the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the novel principles and subject matter disclosed herein may be applied to other embodiments without the exercise of innovative faculty. The claimed subject matter as set forth in the following claims is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. Additional embodiments are contemplated within the spirit and true scope of the disclosed subject matter. Thus, it is intended that the present invention cover modifications and variations thereof provided they come within the scope of the appended claims and their equivalents.

Claims

1. 1. A method of training a predictive model to generate one or more key predictions of key features of an input image, the method comprising: training the predictive model using a first dataset and a second dataset by adjusting a plurality of parameters of the predictive model to reduce a total loss of the predictive model; the first dataset includes a plurality of first training data, each of the first training data including a labeled training image labeled with one or more principal ground truth values ​​of the principal features; the second dataset includes a plurality of second training data, each of which includes a pair of unlabeled training images including a first unlabeled training image and a second unlabeled training image having similarity in the main features; the total loss includes a first loss and a second loss; the first loss is calculated based on a difference between one or more key ground truth values ​​of the labeled training images and one or more first predicted values, the one or more first predicted values ​​being one or more values ​​of the key features generated by the predictive model; The second loss is calculated based on a difference between one or more first predicted values ​​of the first unlabeled training images and one or more second predicted values ​​of the second unlabeled training images, wherein the one or more first predicted values ​​and the one or more second predicted values ​​are one or more values ​​of the key features generated by the predictive model. A method characterized by:

2. training the predictive model to generate one or more auxiliary predictions of one or more auxiliary features of the input image by adjusting the plurality of parameters of the predictive model to reduce the total loss of the predictive model; the one or more auxiliary features correlate with the primary feature; the labeled training images of each of the plurality of first learning data are further labeled with one or more auxiliary ground truth values ​​of the one or more auxiliary features; the total loss further includes a third loss; The third loss is calculated based on a difference between the one or more auxiliary ground truth values ​​of the labeled training images and one or more third predicted values, the one or more third predicted values ​​being one or more values ​​of the one or more auxiliary features generated by a predictive model.

2. The method of claim 1.

3. The method of claim 1 , wherein the labeled training images are modified by image augmentation before generating the one or more first predictions with the predictive model.

4. The method of claim 1 , wherein the one or more primary ground truth values ​​are modified by ground truth extension before calculating the first loss.

5. The method of claim 1 , wherein the first loss and the second loss are calculated by a squared loss function.

6. 2. The method of claim 1, wherein, in each of the plurality of second learning data, the first unlabeled training image and the second unlabeled training image are images of the same subject taken within a predetermined time interval so as to have similarity in main features.

7. 7. The method of claim 6, wherein the predetermined time interval is three months.

8. 2. The method of claim 1, wherein each of the labeled training images, the first unlabeled training images, and the second unlabeled training images is a region of interest (ROI) extracted image extracted from an original training image via ROI extraction.

9. Each of the labeled training images, the first unlabeled training images, and the second unlabeled training images is The original training images, and ROI (region of interest) extracted images extracted from the original training images via ROI extraction 2. The method of claim 1, wherein the training image set comprises:

10. 10. The method of claim 1, wherein the primary feature is the subject's bone mineral density and the one or more primary predictors are one or more bone mineral density (BMD) values.

11. 11. The method of claim 10, wherein the one or more primary predictors include bone mineral density (BMD) values ​​of the total hip, femoral neck, greater trochanter, and femoral shaft.

12. 11. The method of claim 10, wherein the labeled training images of each of the first training data and the first and second unlabeled training images of each of the second training data are X-ray images.

13. 13. The method of claim 12, wherein the first unlabeled training image and the second unlabeled training image are two consecutive x-ray images of the same subject taken within three months of each other.

14. training the predictive model to generate one or more auxiliary predictions of one or more auxiliary features of the input image by adjusting the plurality of parameters of the predictive model to reduce the total loss of the predictive model; the one or more auxiliary features correlate with bone mineral density of the subject; the labeled training images of each of the plurality of first learning data are further labeled with one or more auxiliary ground truth values ​​of the one or more auxiliary features; the total loss further includes a third loss; 11. The method of claim 10, wherein the third loss is calculated based on differences between the one or more auxiliary ground truth values ​​of the labeled training images and one or more third predicted values, the one or more third predicted values ​​being one or more values ​​of the one or more auxiliary features generated by a predictive model.

15. 15. The method of claim 14, wherein the one or more auxiliary features include a cortical thickness of the subject and the one or more auxiliary predictors include a cortical thickness index (CTI) value of the subject.

16. 15. The method of claim 14, wherein the one or more auxiliary features include a femoral neck width for the subject, and the one or more auxiliary predictors include a femoral neck width (FNW) value for the subject.

17. The labeled training images are The original training images, and An ROI-extracted image, which is an identified ROI region of the hip joint extracted from the original training image.

13. The method of claim 12, wherein the training image set comprises:

18. 18. The method of claim 17, wherein the original training images and the ROI-extracted images are modified by image augmentation before generating the one or more first predictions with the predictive model.

19. 19. The method of claim 18, wherein the image augmentation is performed by cropping 0-25% of the original training image without cropping the identified ROI region, and wherein the image augmentation is performed by shifting the identified ROI region by 0-7% in a specific direction.

20. The one or more primary ground truth values ​​are −0.01 g / cm 2 ~0.01 g / cm 2 13. The method of claim 12, wherein the method is modified by introducing a small variable chosen randomly between

Citation Information

Patent Citations

  • Image interpretation model training method and device, equipment and storage medium

    CN114332853A

  • Lung CT image segmentation model construction method and device and electronic equipment

    CN115829994A

  • Network training method and device, electronic equipment and storage medium

    CN117809131A

  • Model training method, identification method, device, storage medium and program product

    US20210406579A1

  • Supervised learning and occlusion masking for optical flow estimation

    WO2022104310A1