System and method for artificial intelligence-based fracture risk prediction using medical image
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SEOUL NAT UNIV HOSPITAL
- Filing Date
- 2026-01-23
- Publication Date
- 2026-07-30
Smart Images

Figure KR2026001389_30072026_PF_FP_ABST
Abstract
Description
System and method for predicting fracture risk using artificial intelligence-based medical imaging
[0001] The present application relates to a system and method for predicting fracture risk using artificial intelligence-based medical imaging. More specifically, the application relates to a system and method that extracts feature vectors through an artificial intelligence model from medical imaging including X-ray images, dual-energy X-ray absorptiometry (DXA) images, computed tomography (CT) images, magnetic resonance imaging (MRI) images, and ultrasound images, predicts the probability of fracture occurrence for each of multiple observation points, generates a survival curve, and quantitatively provides the risk of osteoporotic fracture at each point in time.
[0002] Cross-reference regarding related applications
[0003] This application claims priority to Korean Patent Application No. 10-2025-0010577 filed on January 23, 2025 and Korean Patent Application No. 10-2026-0012797 filed on January 22, 2026, the entire contents of which are incorporated by reference into this application.
[0004] Osteoporosis is a systemic skeletal disease characterized by weakened bones and an increased risk of fracture due to decreased bone density and the degeneration of bone microstructure. With the progression of an aging society, the prevalence of osteoporosis continues to rise, and consequently, osteoporotic fractures are emerging as a major cause of death and a factor contributing to reduced quality of life in the elderly population. In particular, since osteoporotic fractures such as vertebral, hip, humeral, and radial fractures cause high mortality rates and long-term functional impairment after occurrence, it is clinically crucial to identify high-risk groups early and provide preventive treatment before fractures occur.
[0005] Currently, the most widely used tool for assessing fracture risk in clinical practice is FRAX (Fracture Risk Assessment Tool). FRAX is a fracture risk prediction algorithm developed by the World Health Organization (WHO) that calculates the probability of major osteoporotic fractures and hip fractures occurring within the next 10 years by combining clinical risk factors such as age, gender, body mass index, history of past fractures, parental history of hip fractures, smoking status, steroid use, rheumatoid arthritis, secondary osteoporosis, alcohol consumption, and femoral neck bone density.
[0006] However, FRAX has several significant limitations. First, since FRAX provides fracture risk at a single time point of 10 years, it is difficult to assess fracture risk at various time points, such as 1, 2, 3, or 5 years. Clinically, for elderly patients or those with limited life expectancy, fracture risk within a short period (1–2 years) may be more useful for treatment decisions than the 10-year risk; however, existing tools cannot provide this time-specific risk information. Second, while bone density measurement via Dual-Energy X-ray Absorbance (DXA) is necessary for the optimal performance of FRAX, DXA equipment is expensive and is owned by only a select few medical institutions, limiting accessibility. In particular, it is difficult to perform DXA tests in primary care clinics or regions with insufficient medical resources, which limits the early screening of high-risk fracture groups. Third, even though FRAX utilizes DXA imaging information, it does not directly utilize image information but only uses clinical risk factors such as bone density values, so it cannot reflect bone microstructure or body composition-related information inherent in the image.
[0007] Meanwhile, imaging tests in the Department of Radiology, including chest X-rays, are one of the commonly performed medical imaging tests and are widely taken for various purposes, such as health checkups, pre-hospitalization examinations, and evaluation of respiratory diseases. In particular, chest X-ray images include bony structures such as ribs, thoracic vertebrae, and clavicles in addition to the lungs and heart, so there is a possibility to extract bone-related information from these images. If fracture risk could be predicted solely from routinely taken chest X-ray images without the need for separate bone density tests, accessibility to the examination would be significantly improved, making it possible to identify high-risk fracture groups in a larger number of patients.
[0008] With the recent advancement of deep learning technology, predictive studies related to osteoporosis and fractures utilizing medical imaging are being actively conducted. However, existing AI-based fracture prediction studies have limitations in that they do not adequately reflect the temporal characteristics of the fracture event, as they are limited to predicting bone mineral density (BMD) or treating the occurrence of a fracture within a specific period as a binary classification problem. Consequently, information on when a fracture occurs is not provided, and problems exist in which temporal information between subjects with different follow-up periods or censored data cannot be effectively utilized. As a result, there are limitations in distinguishing between patients with a high risk of fracture in the short term and those whose risk gradually increases over a long period, and it has been difficult to use these studies for clinically important short- and mid-term fracture risk assessments at 1, 2, 3, and 5 years.
[0009] Furthermore, existing studies have often relied on methods that indirectly calculate fracture risk by combining DXA images or clinical covariates, or on performing diagnoses after reconstructing CT images from X-ray images. Consequently, there have been limitations in directly predicting fracture risk using only routine medical images. In particular, these approaches tend to be skewed toward indicators strongly associated with bone density, failing to adequately utilize potential risk factors that are not directly correlated with bone density, such as bone microstructure, body shape, soft tissue distribution, or systemic aging and frailty.
[0010] In contrast, the present application proposes an approach that directly models the temporal characteristics of a fracture event by utilizing the occurrence of a fracture along with the follow-up period of each subject as a label. Specifically, the present application directly predicts the risk of fracture (hazard or event probability) through an artificial intelligence model having an output structure (survival head) based on survival analysis, using feature vectors extracted from medical images as input. According to this approach, the AI model can learn fracture risk by comprehensively utilizing image-based features that are not directly related to bone density, rather than merely limiting itself to bone density prediction or binary classification. Furthermore, by reflecting temporal information before and after the occurrence of the fracture, it can effectively reflect the characteristics of data in actual clinical environments, including data that has been truncated midway. As a result, the present application can not only achieve improved fracture prediction performance compared to existing methods, but also quantitatively present the probability of fracture occurrence for each of multiple observation points, such as 1 year, 2 years, 3 years, 5 years, and 10 years, thereby enabling precise risk assessment and decision-making tailored to the patient's life expectancy and clinical situation.
[0011] This application was devised to solve the problems of the prior art under the background described above, and aims to provide fracture risk information at various time points rather than a single time point by using an artificial intelligence model to predict the probability of fracture occurrence for each of multiple observation time points from medical imaging including X-ray images, dual-energy X-ray absorptiometry (DXA) images, computed tomography (CT) images, magnetic resonance imaging (MRI) images, and ultrasound images.
[0012] Another objective of the present application is to improve the accessibility of screening high-risk fracture groups by enabling the prediction of fracture risk using only medical images, such as routinely taken chest X-rays, without the need for separate clinical information surveys or bone density measurements.
[0013] Another objective of the present application is to directly predict fracture risk using an AI model that utilizes follow-up period information along with fracture occurrence as a label, unlike existing AI-based fracture prediction methods utilizing medical images, and has an output structure based on survival analysis. Through this, it is possible to reflect the characteristics of actual clinical data including time information before and after the occurrence of the fracture, and to provide a more precise and optimized fracture risk assessment compared to simple binary classification methods.
[0014] Another objective of the present application is to quantitatively provide fracture risk levels according to different time points, such as 1 year, 2 years, 3 years, 5 years, and 10 years, by predicting the probability of fracture occurrence for each of multiple observation time points through the artificial intelligence model described above. Accordingly, the present application aims to provide improved prediction performance and usability compared to existing image-based artificial intelligence methods by not being limited to prediction results at a single time point but by reflecting fracture risk levels that change over time.
[0015] Another objective of the present application is to predict the probability of occurrence at different time points for various types of osteoporotic fractures, such as vertebral fractures, hip fractures, humeral fractures, and radial fractures, and to provide customized risk information for each fracture type.
[0016] A method for predicting fracture risk using an artificial intelligence-based medical image according to one embodiment of the present application for achieving the above objective comprises: a step of receiving a medical image input, which is performed by a processor; a step of extracting a feature vector from the medical image through an artificial intelligence model; a step of predicting the probability of fracture occurrence for each of a plurality of observation points based on the feature vector; and a step of generating a survival curve using the predicted probability of fracture occurrence for each of the plurality of observation points.
[0017] In one embodiment, the medical image may include at least one of an X-ray image, a dual-energy X-ray absorptiometry (DXA) image, a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and an ultrasound image.
[0018] In one embodiment, the medical image may be a posteroanterior (Chest PA) X-ray image.
[0019] In one embodiment, the artificial intelligence model may be a feature extractor comprising at least one of a Convolutional Neural Network (CNN), a Vision Transformer (ViT), and a Residual Network (ResNet).
[0020] In one embodiment, prior to the step of extracting the feature vector, the method further includes the step of segmenting a first region of interest in the medical image to generate a first normalized image; and the step of segmenting a second region of interest in the medical image to generate a second normalized image; wherein the step of extracting the feature vector may be performed individually for each of the first normalized image and the second normalized image.
[0021] In one embodiment, the first region of interest is a lung region, and the second region of interest may be at least one of a heart region and a bone region.
[0022] In one embodiment, the steps of generating the first normalized image and generating the second normalized image may perform energy-based normalization based on the pixel value distribution characteristics of each region of interest.
[0023] In one embodiment, the step of predicting the probability of fracture occurrence may include combining the first prediction result for the first normalized image and the second prediction result for the second normalized image in an ensemble method of at least one of a soft voting method and a hard voting method.
[0024] In one embodiment, the step of predicting the probability of fracture occurrence may further combine a third prediction result for the original medical image in the ensemble manner.
[0025] In one embodiment, the artificial intelligence model may be trained using training data in which, for each of a plurality of observation points, a first label value is assigned at a point in time before the occurrence of a fracture, and a second label value is assigned at a point in time after the occurrence of a fracture.
[0026] In one embodiment, the artificial intelligence model may be trained using at least one of Survival Loss, Binary Cross Entropy Loss, and Log-likelihood Loss.
[0027] In one embodiment, the plurality of observation points may include at least two of 1 year, 2 years, 3 years, 5 years, and 10 years.
[0028] In one embodiment, the fracture is an osteoporotic fracture and may include at least one of a vertebral fracture, a hip fracture, a humerus fracture, and a radius fracture.
[0029] In one embodiment, the fracture may be defined as satisfying at least one of the following conditions: a history of surgery within a predetermined period from the date of diagnosis of the fracture; and, in the case of a vertebral fracture, a history of imaging examination within a predetermined period from the date of diagnosis of the fracture.
[0030] In one embodiment, the step of generating the survival curve may include: a step of setting a maximum observation period; a step of dividing the maximum observation period into a plurality of time intervals; and a step of calculating the probability of no fracture occurring in each of the plurality of time intervals.
[0031] In one embodiment, the step of generating the survival curve may further include the step of correcting the probability of fracture occurrence after the point in time for data that was censored at the end of observation.
[0032] In one embodiment, the step of predicting the probability of fracture occurrence comprises: a step of obtaining clinical information of a subject; and a step of predicting the probability of fracture occurrence by combining the feature vector and the clinical information; wherein the clinical information may include at least one of age, gender, weight, height, bone density value, history of past fractures, family history, smoking status, steroid use status, and rheumatoid arthritis status.
[0033] In one embodiment, the medical image may be at least one of an X-ray image, a dual-energy X-ray absorptiometry (DXA) image, a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and an ultrasound image.
[0034] In one embodiment, the chest X-ray image may include at least one of a posteroanterior (PA) image and a lateral image.
[0035] In one embodiment, when using the chest X-ray image, the fracture may include at least one of a vertebral fracture and a hip fracture.
[0036] In one embodiment, when using the chest X-ray image, the first normalized image may be generated based on the pixel value distribution of the lung region, and the second normalized image may be generated based on the pixel value distribution of at least one of the heart region, rib region, and thoracic spine region.
[0037] In one embodiment, the medical image may be a dual-energy X-ray absorptiometry (DXA) image.
[0038] In one embodiment, the DXA image may include at least one of a lumbar spine image, a femur image, and a whole body image.
[0039] In one embodiment, when using the DXA image, the step of predicting the probability of fracture occurrence can predict the probability of fracture occurrence by combining the bone mineral density (BMD) value calculated from the DXA image with the feature vector.
[0040] In one embodiment, the first region of interest in the DXA image may be a bone region, and the second region of interest may be a soft tissue region.
[0041] A fracture risk prediction system using an artificial intelligence-based medical image according to one embodiment of the present application for achieving the above objective comprises: one or more processors; and a memory storing one or more instructions executed by the one or more processors; wherein the one or more processors receive a medical image, extract a feature vector from the medical image through an artificial intelligence model, predict a fracture occurrence probability for each of a plurality of observation points based on the feature vectors, and generate a survival curve using the predicted fracture occurrence probability for each of the plurality of observation points.
[0042] In one embodiment, the medical image may include at least one of a chest X-ray image, a spine X-ray image, an abdominal X-ray image, a pelvic X-ray image, and a dual-energy X-ray absorptiometry (DXA) image.
[0043] In one embodiment, the one or more processors may divide a first region of interest in the medical image to generate a first normalized image, divide a second region of interest in the medical image to generate a second normalized image, and individually extract feature vectors for each of the first normalized image and the second normalized image.
[0044] In one embodiment, the one or more processors can predict the probability of fracture occurrence by combining the first prediction result for the first normalized image and the second prediction result for the second normalized image in at least one ensemble method among a soft voting method and a hard voting method.
[0045] In one embodiment, the artificial intelligence model may be trained using training data in which, for each of a plurality of observation times, a first label value is assigned at a time before the occurrence of a fracture and a second label value is assigned at a time after the occurrence of a fracture, and may be trained using at least one of Survival Loss, Binary Cross Entropy Loss, and Log-likelihood Loss.
[0046] In one embodiment, the medical image is a chest X-ray image, and the one or more processors may generate a first normalized image by segmenting the lung region in the chest X-ray image and generate a second normalized image by segmenting at least one of the heart region, rib region, and thoracic spine region in the chest X-ray image.
[0047] In one embodiment, the medical image is a dual-energy X-ray absorptiometry (DXA) image, and the one or more processors can predict the probability of fracture occurrence by combining the bone mineral density (BMD) value calculated from the DXA image with the feature vector.
[0048] A computer-readable non-transient recording medium according to one embodiment of the present application for achieving the above objective stores a command that enables a method for predicting fracture risk using an artificial intelligence-based medical image to be executed by one or more processors, wherein the method comprises: receiving a medical image; extracting a feature vector from the medical image through an artificial intelligence model; predicting a fracture occurrence probability for each of a plurality of observation points based on the feature vector; and generating a survival curve using the predicted fracture occurrence probability for each of the plurality of observation points.
[0049] The present application supports personalized treatment decisions tailored to the patient's life expectancy and clinical situation by predicting the probability of fracture occurrence for each of multiple observation points using an artificial intelligence model from medical images, thereby providing fracture risk information at various points in time such as 1 year, 2 years, 3 years, 5 years, and 10 years.
[0050] This application enables effective communication between medical staff and patients by intuitively visualizing changes in fracture risk over time through the generation of a survival curve using the probability of fracture occurrence for each of multiple predicted observation points.
[0051] This application significantly improves accessibility to testing by enabling the prediction of fracture risk using only routine chest X-ray images without bone density measurement, thereby allowing early screening of high-risk groups for fractures even in areas with limited medical resources.
[0052] The present application enables optimized feature extraction that reflects different tissue characteristics by segmenting different regions of interest, such as lung regions, heart regions, rib regions, and thoracic spine regions, in a chest X-ray image and performing energy-based normalization based on the pixel value distribution characteristics of each region.
[0053] This application enables the risk assessment of major osteoporotic fractures through routine examinations alone by predicting the probability of occurrence of spinal and hip fractures at different time points using chest X-ray images.
[0054] The present application can extend the clinical utility of DXA examination beyond bone density measurement to the prediction of fracture risk at different time points by using dual-energy X-ray absorptiometry (DXA) images and extracting feature vectors from at least one of the lumbar spine images, femoral images, and whole-body images to predict the probability of fracture occurrence.
[0055] The present application predicts the probability of fracture occurrence by combining feature vectors extracted from DXA images with bone density values, thereby enabling more accurate prediction by simultaneously utilizing image information and quantitative bone density information.
[0056] The present application enables a comprehensive fracture risk assessment that reflects both bone structure information and body composition information by segmenting bone regions and soft tissue regions in DXA images and extracting feature vectors individually for each region.
[0057] The present application provides improved prediction accuracy and stability compared to a single model by combining prediction results for each normalized image using an ensemble technique of soft voting or hard voting.
[0058] The present application achieves superior performance compared to existing single-point-time prediction models by providing a model optimized for time-time fracture prediction through training an artificial intelligence model using training data labeled according to whether a fracture occurred at each of multiple observation points and a survival loss function.
[0059] This application contributes to the establishment of preventive treatment strategies by providing customized risk information for each fracture type, through the individual prediction of the probability of occurrence at different time points for various types of osteoporotic fractures, such as vertebral fractures, hip fractures, humeral fractures, and radial fractures.
[0060] The present application predicts the probability of fracture occurrence by combining clinical information such as the subject's age, gender, weight, height, bone density, history of past fractures, family history, smoking status, steroid use status, and rheumatoid arthritis status with feature vectors, thereby enabling a comprehensive fracture risk assessment that integrates the clinical risk factors of the existing FRAX tool with image-based features.
[0061] To more clearly explain the exemplary embodiments of the present application, drawings necessary for the description of the embodiments are briefly introduced below. It should be understood that the drawings below are for the purpose of explaining the embodiments of the present application only and are not for the purpose of limitation. Additionally, for clarity of explanation, some elements in the drawings below may be depicted with various modifications, such as exaggeration or omission.
[0062] FIG. 1 is a block diagram of a fracture risk prediction system using artificial intelligence-based medical images according to one embodiment of the present application.
[0063] FIG. 2 is a flowchart of a method for predicting fracture risk using artificial intelligence-based medical images according to one embodiment of the present application.
[0064] FIG. 3 is a diagram showing the overall structure of an artificial intelligence model and an example of a survival curve output according to one embodiment of the present application.
[0065] FIG. 4 is a diagram illustrating the region segmentation and energy-based normalization preprocessing of a chest X-ray image according to one embodiment of the present application.
[0066] FIG. 5 is a diagram showing a survival loss design method and a label conversion process by time point according to one embodiment of the present application.
[0067] FIGS. 6a and 6b are tables showing the results of a comparison of the prediction accuracy at different time points between a deep learning model according to one embodiment of the present application and a conventional FRAX tool.
[0068] FIG. 7 is a table showing the results of a performance comparison between a single-point-time learning model and a multi-point-time learning model according to one embodiment of the present application.
[0069] FIG. 8 is a table showing the prediction performance of a chest X-ray deep learning model according to one embodiment of the present application by fracture type.
[0070] FIGS. 9a to 9d are drawings illustrating an Optical Character Recognition (OCR) process for extracting bone density values from a DXA image report according to one embodiment of the present application and the results thereof.
[0071] FIG. 10 is a block diagram showing the overall configuration of a fracture risk prediction method using DXA images according to one embodiment of the present application.
[0072] Hereinafter, some embodiments of the present application will be described in detail with reference to the exemplary drawings. In assigning reference numerals to the components of each drawing, the same components may have the same reference numeral as much as possible, even if they are shown in different drawings. Furthermore, in describing these embodiments, if it is determined that a detailed description of related known configurations or functions could obscure the essence of the technical concept, such detailed description may be omitted.
[0073] Where terms such as "includes," "has," "consists of," etc. are used in this specification, other parts may be added unless "only" is used. Where a component is expressed in the singular, it may include a plural unless there is a special explicit description.
[0074] Additionally, terms such as first, second, A, B, (a), (b), etc., may be used to describe the components of the present application. Unless otherwise specified, these terms are used merely to distinguish the components from other components, and the nature, order, sequence, or number of said components is not limited by such terms.
[0075] In this application, terms such as "unit," "module," "device," "component," or "system" are intended to refer to a combination of hardware as well as software driven by said hardware. For example, the hardware may be a data processing device including a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), or other processors.
[0076] Hereinafter, embodiments of the present application will be described in detail with reference to the attached drawings. However, the technical concept of the present application is not limited to some of the described embodiments but can be implemented in various different forms, and within the scope of the technical concept of the present application, one or more of the components among the embodiments may be selectively combined or substituted.
[0077] FIG. 1 is a block diagram of a fracture risk prediction system (100) using artificial intelligence-based medical images according to one embodiment of the present application.
[0078] Referring to FIG. 1, a fracture risk prediction system (100) according to one embodiment of the present application includes a processor (110) and a memory (120). The memory (120) stores one or more instructions executed by the processor (110). The processor (110) may include or perform the functions of a medical image input unit (130), a region segmentation unit (140), a normalization unit (150), a feature extraction unit (160), an ensemble unit (170), a clinical information combining unit (175), a prediction unit (180), a survival curve generation unit (190), and an output unit (195).
[0079] The medical image input unit (130) receives a medical image from an external source.
[0080] In one embodiment, the medical image may include at least one of an X-ray image, a dual-energy X-ray absorptiometry (DXA) image, a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and an ultrasound image. In the case of an X-ray image, it may include a posteroanterior (PA) image of the chest, a lateral image, a spine image, an abdominal image, a pelvic image, etc. In the case of a DXA image, it may include at least one of a lumbar spine image, a femur image, and a whole body image.
[0081] The region segmentation unit (140) segments a plurality of regions of interest in the input medical image. In one embodiment, the region segmentation unit (140) may segment a first region of interest in the medical image and segment a second region of interest in the medical image. When using a chest X-ray image, the first region of interest may be a lung region, and the second region of interest may be at least one of a heart region, a rib region, and a thoracic spine region. When using a DXA image, the first region of interest may be a bone region, and the second region of interest may be a soft tissue region.
[0082] The normalization unit (150) performs normalization processing for each divided region of interest. In one embodiment, the normalization unit (150) may perform energy-based normalization based on the pixel value distribution characteristics of each region of interest. In the case of a chest X-ray image, the normalization unit (150) may generate a first normalized image based on the pixel value distribution of the lung region and generate a second normalized image based on the pixel value distribution of at least one of the heart region, rib region, and thoracic spine region. Through this region-specific normalization, optimized feature extraction reflecting different tissue characteristics (air-filled lung region and high-density bone region) is possible.
[0083] The feature extraction unit (160) extracts feature vectors from a medical image using an artificial intelligence model. In one embodiment, the feature extraction unit (160) can extract feature vectors individually for each of the first normalized image and the second normalized image. Additionally, the feature extraction unit (160) can extract feature vectors for the original medical image that has not undergone normalization processing. The artificial intelligence model may be a feature extractor comprising at least one of a Convolutional Neural Network (CNN), a Vision Transformer (ViT), and a Residual Network (ResNet).
[0084] In one embodiment, the artificial intelligence model may be trained using training data in which, for each of a plurality of observation times, a first label value (e.g., 1) is assigned at a time before the occurrence of a fracture, and a second label value (e.g., 0) is assigned at a time after the occurrence of a fracture. Additionally, the artificial intelligence model may be trained using at least one of Survival Loss, Binary Cross Entropy Loss, and Log-likelihood Loss.
[0085] The ensemble unit (170) combines multiple prediction results to generate a final prediction result. In one embodiment, the ensemble unit (170) may combine a first prediction result for a first normalized image and a second prediction result for a second normalized image using at least one ensemble method among a soft voting method and a hard voting method. Additionally, the ensemble unit (170) may further combine a third prediction result for the original medical image using the ensemble method.
[0086] The clinical information combining unit (175) combines the subject's clinical information with a feature vector. In one embodiment, the clinical information may include at least one of age, gender, weight, height, bone mineral density value, history of past fractures, family history, smoking status, steroid use status, and rheumatoid arthritis status. When using DXA images, the clinical information combining unit (175) may combine the bone mineral density (BMD) value calculated from the DXA images with the feature vector. This enables a more accurate prediction by simultaneously utilizing image information and quantitative bone mineral density information.
[0087] The prediction unit (180) predicts the probability of fracture occurrence for each of a plurality of observation points based on a feature vector. In one embodiment, the plurality of observation points may include at least two of 1 year, 2 years, 3 years, 5 years, and 10 years. The fracture to be predicted is an osteoporotic fracture and may include at least one of a vertebral fracture, a hip fracture, a humerus fracture, and a radius fracture.
[0088] In one embodiment, a fracture may be defined as satisfying at least one of the following conditions: having a history of surgery within a predetermined period from the date of diagnosis of the fracture, and, in the case of a vertebral fracture, having a history of imaging examination within a predetermined period from the date of diagnosis of the fracture. The predetermined period may be, for example, 6 months, but is not limited thereto.
[0089] The survival curve generation unit (190) generates a survival curve using the probability of fracture occurrence for each of the predicted multiple observation points. In one embodiment, the survival curve generation unit (190) may set a maximum observation period, divide the maximum observation period into multiple time intervals, and calculate the probability of no fracture occurrence in each of the multiple time intervals. Additionally, the survival curve generation unit (190) may correct the probability of fracture occurrence after the corresponding point in time for data that was censored at the end of the observation.
[0090] The output unit (195) provides the user with the predicted fracture occurrence probability and the generated survival curve. In one embodiment, the output unit (195) may output the results in a manner such as a visual display, data file storage, or network transmission.
[0091] Below, a specific example of fracture risk prediction using chest X-ray images is described.
[0092] When using a chest X-ray image, the region segmentation unit (140) may segment the lung region in the chest X-ray image to generate a first normalized image, and segment at least one of the heart region, rib region, and thoracic spine region to generate a second normalized image. The lung region has relatively low pixel values because it is filled with air, while the heart region and bone region (ribs, thoracic spine) have relatively high pixel values. Therefore, by performing normalization individually according to the pixel value distribution characteristics of each region, it is possible to extract features optimized for each tissue characteristic.
[0093] Fractures predicted using chest X-ray images may include at least one of vertebral fractures and hip fractures. In particular, according to an embodiment of the present application, high prediction performance for hip fractures can be achieved even using only chest X-ray images.
[0094] Below, a specific example of fracture risk prediction using DXA imaging is described.
[0095] When using DXA images, the DXA images may include at least one of a lumbar spine image, a femur image, and a whole body image. The region segmentation unit (140) can segment the bone region in the DXA image into a first region of interest and the soft tissue region into a second region of interest. This enables a comprehensive fracture risk assessment that reflects both bone structure information and body composition information.
[0096] The clinical information combining unit (175) can predict the probability of fracture occurrence by combining bone mineral density (BMD) values calculated from DXA images with feature vectors. In DXA images, BMD, T-score, Z-score of lumbar vertebrae L1-L4, BMD, T-score, Z-score of the left and right femoral necks and the total, and TBS (Trabecular Bone Score) values of the lumbar vertebrae can be extracted. By combining this quantitative bone mineral density information with feature vectors extracted from images, the clinical utility of the existing DXA examination can be extended beyond bone mineral density measurement to the prediction of fracture risk at different time points.
[0097] Below, specific examples of multi-viewpoint learning methods are described.
[0098] An artificial intelligence model according to one embodiment of the present application adopts a learning method optimized for predicting fractures at different time points. Specifically, for each of a plurality of observation time points (e.g., 1 year, 2 years, 3 years, 5 years, 10 years), a model is trained using training data in which a first label value (e.g., 1) is assigned at a time point before the occurrence of a fracture, and a second label value (e.g., 0) is assigned at a time point after the occurrence of a fracture.
[0099] For example, if the maximum observation period is set to 10 years and the time interval is divided into 1-year increments, and a patient develops a fracture at the 2-year and 6-month mark, the patient's label can be assigned as [1, 1, 0, 0, 0, 0, 0, 0, 0, 0]. That is, since no fracture occurred during the 1st and 2nd years, the label is assigned as 1, and from the 3rd year onwards, it is assumed that a fracture occurred, so the label is assigned as 0.
[0100] In cases where no fracture occurs during the observation period (intermediate amputation, censoring), the final observation point is treated as if the event occurred, but additional processing is performed to reflect a low probability of the event occurring even after that point.
[0101] Learning using this multi-point learning method and the Survival Loss function has the effect of significantly improving prediction accuracy at each point in time compared to existing single-point prediction (e.g., predicting only whether a fracture occurs at the 10th year) or binary classification (fracture occurrence / non-occurrence).
[0102] Specific embodiments of the ensemble method are described below.
[0103] The ensemble unit (170) can combine multiple prediction results using at least one of a soft voting method and a hard voting method. A soft voting method is a method of deriving a final prediction value by calculating the average or weighted average of the probability values output by each model. A hard voting method is a method of determining the final prediction by a majority vote by voting on the prediction results of each model (e.g., high risk / low risk).
[0104] In one embodiment, the ensemble unit (170) can predict the final fracture occurrence probability by ensembling the prediction result for the lung region-based normalized image, the prediction result for the heart / bone region-based normalized image, and the prediction result for the original image. Through this ensemble method, improved prediction accuracy and stability compared to a single model can be achieved.
[0105] FIG. 2 is a flowchart of a method for predicting fracture risk using artificial intelligence-based medical images according to one embodiment of the present application.
[0106] Referring to FIG. 2, a fracture risk prediction method according to one embodiment of the present application is performed by a processor and includes the steps of receiving a medical image (S210), segmenting a region of interest (S220), generating a normalized image (S230), extracting a feature vector (S240), ensembling prediction results (S250), combining clinical information (S260), predicting the probability of fracture occurrence (S270), generating a survival curve (S280), and outputting the prediction results (S290).
[0107] In step S210, the processor receives a medical image. In one embodiment, the medical image may include at least one of an X-ray image, a dual-energy X-ray absorptiometry (DXA) image, a computed tomography (CT) image, a magnetic resonance imaging (MRI) image, and an ultrasound image. In the case of an X-ray image, it may include at least one of a posteroanterior (PA) image of the chest, a lateral image, a spine image, an abdominal image, and a pelvic image. In the case of a DXA image, it may include at least one of a lumbar spine image, a femur image, and a whole body image.
[0108] In step S220, the processor segments a region of interest in the input medical image. In one embodiment, the processor may segment a first region of interest in the medical image and a second region of interest in the medical image. When using a chest X-ray image, the first region of interest may be a lung region, and the second region of interest may be at least one of a heart region, a rib region, and a thoracic spine region. When using a DXA image, the first region of interest may be a bone region, and the second region of interest may be a soft tissue region.
[0109] In step S230, the processor generates a normalized image for the segmented regions of interest. In one embodiment, the processor may segment a first region of interest to generate a first normalized image and segment a second region of interest to generate a second normalized image. The step of generating the first normalized image and the second normalized image may perform energy-based normalization based on the pixel value distribution characteristics of each region of interest. In the case of a chest X-ray image, the first normalized image may be generated based on the pixel value distribution of the lung region and the second normalized image may be generated based on the pixel value distribution of at least one of the heart region, the rib region, and the thoracic spine region.
[0110] In step S240, the processor extracts feature vectors from the medical image through an artificial intelligence model. In one embodiment, the artificial intelligence model may be a feature extractor comprising at least one of a Convolutional Neural Network (CNN), a Vision Transformer (ViT), and a Residual Network (ResNet). The step of extracting feature vectors may be performed individually for each of the first normalized image and the second normalized image.
[0111] In one embodiment, the artificial intelligence model may be trained using training data in which, for each of a plurality of observation times, a first label value is assigned at a time before the occurrence of a fracture, and a second label value is assigned at a time after the occurrence of a fracture. Additionally, the artificial intelligence model may be trained using at least one of Survival Loss, Binary Cross Entropy Loss, and Log-likelihood Loss.
[0112] In step S250, the processor ensembles a plurality of prediction results. In one embodiment, the processor may combine a first prediction result for a first normalized image and a second prediction result for a second normalized image using at least one ensemble method among a soft voting method and a hard voting method. Additionally, the processor may further combine a third prediction result for the original medical image using the ensemble method.
[0113] In step S260, the processor acquires clinical information of the subject and combines the clinical information with a feature vector. In one embodiment, the clinical information may include at least one of age, sex, weight, height, bone mineral density values, history of past fractures, family history, smoking status, steroid use status, and rheumatoid arthritis status. When using DXA images, the processor can predict the probability of fracture occurrence by combining the bone mineral density (BMD) values calculated from the DXA images with the feature vector.
[0114] In step S270, the processor predicts the probability of fracture occurrence for each of a plurality of observation points based on a feature vector. In one embodiment, the plurality of observation points may include at least two of 1 year, 2 years, 3 years, 5 years, and 10 years. The fracture to be predicted is an osteoporotic fracture and may include at least one of a vertebral fracture, a hip fracture, a humerus fracture, and a radius fracture. When using a chest X-ray image, the fracture may include at least one of a vertebral fracture and a hip fracture.
[0115] In one embodiment, a fracture may be defined as satisfying at least one of the following conditions: having a history of surgery within a predetermined period from the date of diagnosis of the fracture, and, in the case of a spinal fracture, having a history of imaging examination within a predetermined period from the date of diagnosis of the fracture.
[0116] In step S280, the processor generates a survival curve using the probability of fracture occurrence for each of the predicted multiple observation points. In one embodiment, the step of generating the survival curve may include the step of setting a maximum observation period, the step of dividing the maximum observation period into multiple time intervals, and the step of calculating the probability of no fracture occurrence in each of the multiple time intervals. Additionally, the step of generating the survival curve may further include the step of correcting the probability of fracture occurrence after the time point for data that was censored at the end of the observation.
[0117] In step S290, the processor outputs the predicted fracture occurrence probability and the generated survival curve.
[0118] The fracture risk prediction method of FIG. 2 described above can be performed by the processor (110) of the fracture risk prediction system (100) described with reference to FIG. 1. Specifically, step S210 can be performed by the medical image input unit (130), step S220 by the region segmentation unit (140), step S230 by the normalization unit (150), step S240 by the feature extraction unit (160), step S250 by the ensemble unit (170), step S260 by the clinical information combining unit (175), step S270 by the prediction unit (180), step S280 by the survival curve generation unit (190), and step S290 by the output unit (195).
[0119] FIG. 3 is a diagram showing the overall structure of an artificial intelligence model and an example of a survival curve output according to one embodiment of the present application.
[0120] Referring to FIG. 3, a fracture risk prediction method according to one embodiment of the present application receives a chest X-ray image as input, extracts a feature vector through a vision encoder, and finally outputs a survival curve through an artificial intelligence model trained using survival loss.
[0121] Specifically, when a chest X-ray image is input into a vision encoder, the vision encoder extracts feature vectors from the image that are significant for predicting fracture risk. The vision encoder can be implemented using artificial intelligence models such as Convolutional Neural Networks (CNN), Vision Transformers (ViT), or Residual Neural Networks (ResNet). The extracted feature vectors are input into a classifier and used to predict the probability of fracture occurrence for each of multiple observation points.
[0122] The heatmap image shown in the center of Fig. 3 visualizes the areas that the artificial intelligence model focuses on when predicting fracture risk. The areas marked in red on the heatmap are the areas to which the model assigns high weights, and they mainly correspond to the thoracic spine and rib regions. This indicates that the artificial intelligence model of the present application effectively learns the characteristics of the bone region in predicting osteoporotic fractures.
[0123] The key feature of the present application is that, rather than binary classifying whether a fracture occurs at a single point in time, it generates a survival curve by predicting the probability of fracture occurrence for each of multiple observation points. In the survival curve graph shown on the right side of Fig. 3, the horizontal axis represents time (Days), and the vertical axis represents the probability of no fracture occurring (survival probability). The survival curve continuously shows how the probability of no fracture occurring changes as time progresses.
[0124] In the survival curve graph of Figure 3, the case curve represents the survival curve of the patient being predicted, and the reference curve represents the average survival curve of the reference group. If the survival curve of the patient being predicted drops sharply below the reference curve, it can be interpreted that the patient has a higher risk of fracture compared to the reference group.
[0125] In one embodiment, as illustrated in FIG. 3, the present application may provide a prediction result for a specific patient stating that "the patient is likely to experience a subsequent fracture after 852 days." Additionally, it may quantitatively provide the probability of fracture occurrence at multiple observation points, such as a fracture risk of 41.2% at 2 years, a fracture risk of 52.0% at 3 years, and a fracture risk of 52.3% at 5 years.
[0126] This information on fracture probability at specific time points is a key feature of the present application that differentiates it from existing FRAX tools, which provide only fracture risk at a single time point of 10 years. In particular, for elderly patients or patients with limited life expectancy, information on fracture risk over short periods, such as 1, 2, or 3 years, may be clinically more useful than the fracture risk after 10 years. By providing a survival curve that encompasses not only such short-term predictions but also long-term predictions, the present application supports personalized treatment decisions tailored to the patient's situation.
[0127] The artificial intelligence model of the present application is trained using Survival Loss. Survival Loss is designed based on training data in which, for each of multiple observation points, a first label value (e.g., 1) is assigned to the point in time before the occurrence of a fracture, and a second label value (e.g., 0) is assigned to the point in time after the occurrence of a fracture. For example, if the maximum observation period is set to 10 years and the time interval is divided into 1-year units, the label for a patient who has a fracture at the 2-year and 6-month mark is assigned as [1, 1, 0, 0, 0, 0, 0, 0, 0, 0]. The artificial intelligence model is trained to minimize the difference between the value predicted by the model and the correct answer, using these time-specific labels as the correct answers.
[0128] For censored data—that is, cases where no fracture occurred during the observation period—the final observation point is treated as the point in time when the event occurred; however, additional corrections are applied to reflect a lower probability of the event occurring even after that point. This processing of censored data is an important technique in survival analysis, enabling the use of information for training on patients who did not experience a fracture by the end of the observation period.
[0129] As such, learning using a multi-point learning method and a survival loss function has the effect of significantly improving prediction accuracy per point in time compared to existing single-point prediction or binary classification methods. This can be confirmed in the performance comparison results of the single-point model and the multi-point model in Figure 7, which will be described later.
[0130] FIG. 4 is a diagram illustrating the region segmentation and energy-based normalization preprocessing of a chest X-ray image according to one embodiment of the present application.
[0131] Referring to FIG. 4, a preprocessing process according to one embodiment of the present application includes segmenting a region of interest (second column) in an original chest X-ray image (first column), and performing energy-based normalization based on the pixel value distribution characteristics of each region of interest to generate a first normalized image (third column) and a second normalized image (fourth column). FIG. 4 illustrates the preprocessing process for three different patients in each row.
[0132] In the image of the second column, the lung region represents the first region of interest, and the heart region represents the second region of interest. Region segmentation can be performed automatically using a pre-trained segmentation model.
[0133] In chest X-ray images, lung regions are filled with air and have relatively low pixel values (dark areas), while heart and bone regions (ribs, thoracic vertebrae, clavicles, etc.) consist of soft or bone tissue and have relatively high pixel values (bright areas). Due to these different tissue characteristics, features of specific regions may be lost or distorted when applying a single normalization method.
[0134] To solve these problems, the present application individually performs energy-based normalization based on the pixel value distribution characteristics of each region of interest. The first normalized image in the third column is an image normalized based on the pixel value distribution of the lung region, in which the fine structure and pattern of the lung parenchyma are emphasized. The second normalized image in the fourth column is an image normalized based on the pixel value distribution of the heart region or bone region, in which bone structures such as ribs, thoracic vertebrae, and clavicles are clearly emphasized.
[0135] In one embodiment, energy-based normalization can be performed by analyzing the pixel value histogram of each region of interest to calculate the mean and standard deviation of the region, and readjusting the pixel values of the entire image based on these values. Through this, the image is transformed into a form suitable for training an artificial intelligence model while preserving the unique characteristics of each region.
[0136] As can be seen in Figure 4, in the first normalized image (third column), the shading and vascular patterns of the lung region are clearly observed, while the bone region appears relatively blurry. Conversely, in the second normalized image (fourth column), the bone density and structural features of the ribs and thoracic vertebrae are clearly observed, while the lung region appears relatively blurry. Through this region-specific normalization, it becomes possible to extract features optimized for each tissue type.
[0137] In one embodiment, a bone region (rib region, thoracic spine region) may be directly used as the second region of interest instead of the heart region. Since the heart region has pixel value distribution characteristics similar to those of the bone region, the bone structure is effectively emphasized even in an image normalized based on the heart region. Therefore, the heart region and the bone region are interchangeable as the second region of interest.
[0138] According to one embodiment of the present application, feature vectors are individually extracted for each of the first normalized image and the second normalized image, and the prediction results for each image are ensembled to predict the final fracture occurrence probability. Specifically, the first prediction result for the first normalized image and the second prediction result for the second normalized image can be combined using a soft voting method or a hard voting method. Additionally, the third prediction result for the original medical image can be further combined using an ensemble method.
[0139] These region-specific preprocessing and ensemble techniques are one of the technical differentiators of the present application, and they have the effect of improving prediction accuracy and stability compared to existing methods using a single normalization method. In the lung region-based normalized image, features reflecting changes in lung parenchyma (e.g., thoracic deformation due to osteoporosis) are extracted, and in the bone region-based normalized image, features reflecting changes in bone density and bone structure are extracted. By ensembling these complementary features, it becomes possible to predict fracture risk more accurately and robustly.
[0140] FIG. 5 is a diagram showing a survival loss design method and a label conversion process by time point according to one embodiment of the present application.
[0141] Referring to FIG. 5, for training an artificial intelligence model according to one embodiment of the present application, the "event occurrence status" and "final observation time (event occurrence time if the event occurred)" of each training data are designated as ground truths, and the artificial intelligence model is trained using a supervised learning method.
[0142] The input data shown on the left side of Fig. 5 is represented as {occurrence, 1015th day}, which indicates that a fracture event occurred in the patient and that the occurrence occurred on the 1015th day. This input data is converted into ground truth data for training an artificial intelligence model.
[0143] The example of conversion into correct answer data shown on the right side of FIG. 5 is represented as {1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0}. This is achieved by setting a maximum observation period, dividing the period into regular time intervals, and converting whether an event occurred in each time interval into a label. Specifically, a first label value (1) is assigned at a time point before the event occurs, and a second label value (0) is assigned at a time point after the event occurs.
[0144] In one embodiment, it is assumed that the maximum observation period is set to 10 years (3650 days) and the time interval is divided into units of 1 year (approx. 365 days). In the example of FIG. 5, since the fracture occurred on the 1015th day (approx. 2 years and 9 months), the label value 1 is assigned because no fracture occurred during the 1st year (365 days) and the 2nd year (730 days), and the label value 0 is assigned because the fracture is considered to have occurred from the 3rd year (1095 days) onwards. Therefore, the correct label can be configured as [1, 1, 0, 0, 0, 0, 0, 0, 0, 0]. The example of {1, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0} shown in FIG. 5 may correspond to a case where a more finely subdivided time interval (e.g., in months) is used.
[0145] The key feature of this application is that, through such time-series labeling, the AI model can learn not merely a binary classification of "whether a fracture occurs or not," but rather "until when a fracture does not occur and after what time point a fracture occurs." Through this, the probability of fracture occurrence for each of multiple observation time points can be predicted, and a survival curve can ultimately be generated.
[0146] In one embodiment, separate processing is performed on data that has been censored, i.e., data where no fracture occurred during the observation period. For censored data, the final observation point is treated as the point in time when the event occurred, but additional processing is performed to reflect a low probability of the event occurring even after that point in time. For example, in the case of a patient who was observed up to the 5th year but did not experience a fracture, a label value of 1 is assigned up to the 5th year, and for points in time after the 5th year, the weight may be lowered when calculating the loss function or separate masking processing may be performed.
[0147] During the training process of the AI model, feature vectors are extracted from medical images using an artificial neural network feature extractor (e.g., CNN, ViT, ResNet). These feature vectors are then input into a classifier to design a model capable of outputting whether an event occurred at each time point within a specified maximum observation period. At this stage, a loss function is designed to minimize the difference between the value predicted by the model and the designated ground truth, thereby updating the AI model's parameters.
[0148] In one embodiment, the loss function may include at least one of Survival Loss, Binary Cross Entropy Loss, and Log-likelihood Loss. The Survival Loss calculates the difference between the predicted value and the correct label at each time point, and is designed to enable learning specialized for survival analysis by adjusting weights based on whether censoring occurs and the timing of the event.
[0149] This multi-point labeling and survival loss-based learning method has the effect of significantly improving prediction accuracy per point in time compared to the existing single-point binary classification method. The existing method predicts only whether a fracture occurs at a specific point in time (e.g., the 10th year), so it cannot utilize information on fracture occurrence prior to that point in time, and thus has the problem of low prediction accuracy at other points in time. On the other hand, the multi-point learning method of the present application learns information from all observation points in an integrated manner, so it can achieve high prediction accuracy at any point in time, such as 1 year, 2 years, 3 years, 5 years, and 10 years.
[0150] FIGS. 6a and 6b are tables showing the results of a comparison of the prediction accuracy at different time points between a deep learning model according to one embodiment of the present application and a conventional FRAX tool. FIG. 6a shows the results from the Seoul National University Hospital (SNUH) test set, and FIG. 6b shows the results from the Seoul National University Bundang Hospital (SNUBH) test set.
[0151] Referring to Fig. 6a, the chest X-ray deep learning model (CXR DL model) of the present application achieved a C-index of 0.867 on the Seoul National University Hospital test set (n=5,000), which is significantly higher than the C-index of 0.800 of the existing FRAX tool. The C-index (Concordance index) is an indicator that evaluates the discriminative power of a model in survival analysis, and a value closer to 1 indicates superior predictive performance.
[0152] Looking at the Area Under the Curve (AUC) results by time point, the deep learning model of the present application achieved an AUC of 0.879 at 1 year, 0.878 at 2 years, 0.887 at 3 years, 0.886 at 5 years, and 0.899 at 10 years. In comparison, the FRAX tool recorded an AUC of 0.793 at 1 year, 0.805 at 2 years, 0.804 at 3 years, 0.805 at 5 years, and 0.832 at 10 years. The deep learning model of the present application showed higher prediction accuracy compared to the FRAX tool at all observation points, and the difference is particularly pronounced in short-term predictions (1 year, 2 years).
[0153] Referring to Fig. 6b, the deep learning model of the present application demonstrated superior performance compared to the FRAX tool in the Bundang Seoul National University Hospital test set (n=4,938). The deep learning model of the present application achieved a C-index of 0.747, recording a higher value compared to the FRAX tool's C-index of 0.715. In terms of AUC by time point, the deep learning model of the present application achieved an AUC of 0.771 at 1 year, 0.773 at 2 years, 0.787 at 3 years, 0.799 at 5 years, and 0.796 at 10 years, demonstrating superior performance at all time points compared to the FRAX tool's AUC of 0.704 at 1 year, 0.722 at 2 years, 0.738 at 3 years, 0.747 at 5 years, and 0.780 at 10 years.
[0154] The Seoul National University Hospital test set in Fig. 6a is an internal validation dataset collected from the same institution as the data used for model development, and the Bundang Seoul National University Hospital test set in Fig. 6b is an external validation dataset collected from an independent institution not used for model development. Although performance on external validation datasets generally tends to be somewhat lower than on internal validation datasets, the deep learning model of the present application maintained superior performance compared to the FRAX tool in external validation as well. This means that the model of the present application does not overfit to the data of a specific institution and demonstrates generalized predictive performance in various environments.
[0155] Of particular note is that the deep learning model of the present application shows greater performance improvement compared to the FRAX tool in short-term prediction (1 year, 2 years). In the Seoul National University Hospital test set, the difference in 1-year AUC is 0.086 (0.879 - 0.793), while the difference in 10-year AUC is 0.067 (0.899 - 0.832), indicating a greater improvement in performance in short-term prediction. This suggests that it can provide clinically more useful information for elderly patients or patients with limited life expectancy.
[0156] Existing FRAX tools are designed to predict only the fracture risk at the 10th year, so the accuracy of short-term prediction is relatively low. On the other hand, the deep learning model of the present application adopts the multi-point learning method described in Fig. 5 to integrally learn information from all observation points, so it can achieve high prediction accuracy at any point from 1 to 10 years.
[0157] In addition, the deep learning model of the present application achieved the above-mentioned prediction performance by using only chest X-ray images as input. Unlike the FRAX tool, which requires bone density measurements and information on various clinical risk factors, the present application enables excellent fracture risk prediction using only routinely taken chest X-ray images. This has the effect of significantly improving accessibility to examinations by enabling early screening of high-risk fracture groups even in primary care institutions without bone density measurement equipment or in regions with insufficient medical resources.
[0158] FIG. 7 is a table showing the results of a performance comparison between a single-point-time learning model and a multi-point-time learning model according to one embodiment of the present application.
[0159] Referring to Figure 7, the results of comparing the prediction performance of a single-time point model (Cross Entropy Loss) and a multi-time point model (Survival Loss) on the Seoul National University Hospital test set (n=5,000) are shown. The single-time point model is a model trained using the existing method of binary classification of whether a fracture occurs at a specific time point using Binary Cross Entropy Loss, and the multi-time point model is a model trained using the Survival Loss of the present application to learn the probability of fracture occurrence for each of multiple observation times points.
[0160] The single-time model achieved a C-index of 0.804, and recorded AUCs of 0.816 for 1 year, 0.818 for 2 years, 0.830 for 3 years, 0.818 for 5 years, and 0.830 for 10 years. On the other hand, the multi-time model achieved a C-index of 0.867, and recorded AUCs of 0.879 for 1 year, 0.878 for 2 years, 0.887 for 3 years, 0.886 for 5 years, and 0.899 for 10 years.
[0161] The multi-times model showed a performance improvement of 0.063 (0.867 - 0.804) in the C-index compared to the single-times model. In terms of AUC by time point, the multi-times model also showed superior performance compared to the single-times model at all observation points, and in particular, achieved a performance improvement of 0.063 (0.879 - 0.816) in the 1-year AUC, 0.068 (0.886 - 0.818) in the 5-year AUC, and 0.069 (0.899 - 0.830) in the 10-year AUC.
[0162] These results demonstrate that the multi-point learning method of the present application brings about a significant performance improvement compared to the existing single-point binary classification method. Since a single-point model learns only whether a fracture occurred at a specific point in time (e.g., the 10th year), it fails to fully utilize information regarding fracture occurrences prior to that point in time. For example, a patient who suffered a fracture in the 2nd year and a patient who suffered a fracture in the 9th year are both labeled as having "fracture occurred" based on the 10th year, resulting in the loss of information regarding the timing of the fracture.
[0163] On the other hand, the multi-point model of the present application learns by assigning a first label value (1) before the occurrence of a fracture and a second label value (0) after the occurrence of a fracture to each of the multiple observation points, as described in FIG. 5. Through this, the model can learn temporal information regarding "until when a fracture does not occur and after which point a fracture occurs." In addition, the survival loss function performs appropriate processing on censored data, so that information on patients who did not experience a fracture during the observation period is also effectively utilized for learning.
[0164] A particularly noteworthy point in the results of Figure 7 is that the multi-time model demonstrates excellent performance not only in long-term prediction (10-year AUC 0.899) but also in short-term prediction (1-year AUC 0.879). While the performance difference between the 5-year AUC (0.818) and 10-year AUC (0.830) of the single-time model is not significant, the multi-time model maintains consistently high prediction performance across all time points. This suggests that the multi-time learning method effectively captures fracture occurrence patterns over time.
[0165] The excellence of this multi-viewpoint learning method is one of the key technological contributions of this application. Unlike most existing fracture risk prediction studies that adopt a single-viewpoint binary classification method, this application applies a multi-viewpoint learning method based on survival analysis to medical image-based fracture prediction, thereby significantly improving prediction accuracy per viewpoint and increasing clinical utility.
[0166] FIG. 8 is a table showing the prediction performance of a chest X-ray deep learning model according to one embodiment of the present application by fracture type.
[0167] Referring to Fig. 8, the predictive performance achieved by the chest X-ray deep learning model (CXR DL model) of the present application for various fracture types in the Bundang Seoul National University Hospital test set (n=25,114) is illustrated. Fracture types are classified into total fractures (ALL), vertebral fractures, non-vertebral fractures, and hip fractures.
[0168] For all fractures (ALL, n=1,165), the deep learning model of the present application achieved a C-index of 0.810, and recorded an AUC of 0.846 at 1 year, 0.838 at 2 years, 0.846 at 3 years, 0.848 at 5 years, and 0.849 at 10 years. This indicates that the model of the present application demonstrates high prediction accuracy across all osteoporotic fractures.
[0169] For vertebral fractures (n=925), a C-index of 0.818, a 1-year AUC of 0.857, a 2-year AUC of 0.852, a 3-year AUC of 0.858, a 5-year AUC of 0.857, and a 10-year AUC of 0.858 were achieved. Vertebral fractures are the most common type of osteoporotic fracture, and it is analyzed that they showed high predictive performance because the thoracic spine region is directly observed in chest X-ray images.
[0170] For non-vertebral fractures (n=240), a C-index of 0.792, a 1-year AUC of 0.819, a 2-year AUC of 0.807, a 3-year AUC of 0.819, a 5-year AUC of 0.831, and a 10-year AUC of 0.828 were achieved. Although non-vertebral fractures include fractures in areas not directly observed in chest X-ray images, such as the humerus and radius, the model of this application demonstrated good predictive performance by learning features that reflect the systemic osteoporosis state.
[0171] Of particular note is the predictive performance for hip fractures (Hip, n=160). The deep learning model of this application achieved a C-index of 0.859, a 1-year AUC of 0.870, a 2-year AUC of 0.865, a 3-year AUC of 0.873, a 5-year AUC of 0.890, and a 10-year AUC of 0.898 for hip fractures, recording the highest predictive performance among all fracture types. Hip fractures are the type of osteoporotic fracture that causes the most severe complications, with a mortality rate of 20–30% within one year and a significant number of patients suffering from permanent functional disability. Therefore, accurate prediction of hip fractures holds great clinical significance.
[0172] It is a notable result that the chest X-ray deep learning model of the present application showed the highest predictive performance for hip fractures. Although the hip joint is not included in the imaging range of chest X-rays, the model of the present application suggests that the systemic osteoporosis status and fracture risk can be effectively inferred from the bone structures (ribs, thoracic vertebrae, clavicle, etc.) in the chest region. This is analyzed to be because the features of the bone region were effectively extracted through region-specific preprocessing and energy-based normalization described in Fig. 4.
[0173] The results of Fig. 8 demonstrate that the chest X-ray deep learning model of the present application is not specialized for a single fracture type, but can achieve high predictive accuracy for various osteoporotic fractures, including vertebral fractures and hip fractures. In particular, the fact that hip fractures can be predicted with a high accuracy of C-index 0.859 using only chest X-ray images suggests the possibility of replacing or supplementing the assessment of hip fracture risk—which previously required bone density measurement or separate pelvic imaging examinations—with only routine chest X-ray examinations.
[0174] FIGS. 9a to 9d are drawings illustrating an Optical Character Recognition (OCR) process for extracting bone density values from a DXA image report according to one embodiment of the present application and the results thereof.
[0175] Referring to Fig. 9a, the measurement results of the lumbar spine region in the DXA test report are shown. Bone density (BMD, g / cm²), percentage relative to Young-Adult (YA%), T-score, percentage relative to Age-Matched (AM%), and Z-score are recorded for each of the lumbar vertebrae L1, L2, L3, and L4. In the example of Fig. 9a, the BMD of L1 is 0.861 g / cm² and the T-score is -1.7; the BMD of L2 is 0.919 g / cm² and the T-score is -1.7; the BMD of L3 is 0.884 g / cm² and the T-score is -2.0; and the BMD of L4 is 0.956 g / cm² and the T-score is -1.4. In addition, measurements combining multiple lumbar vertebrae, such as L1-L2, L1-L3, L1-L4, L2-L3, L2-L4, and L3-L4, are also recorded.
[0176] Referring to Fig. 9b, the measurement results of the femur region in the DXA examination report are shown. BMD, YA%, T-score, AM%, and Z-score are recorded for the femoral neck, upper neck, lower neck, wards, greater troch, shaft, and total femur, respectively. In the example of Fig. 9b, the BMD of the femoral neck is 0.694 g / cm² and the T-score is -1.7, while the BMD of the total femur is 0.752 g / cm² and the T-score is -1.5. Femoral measurements can be performed on the left femur and the right femur, respectively.
[0177] Referring to Fig. 9c, the lumbar TBS (Trabecular Bone Score) measurement results from the DXA examination report are shown. TBS is a bone microstructure indicator calculated from lumbar DXA images and is used to predict fracture risk independently of bone density. In the example in Fig. 9c, the TBS of L1 is 1.279, the TBS of L2 is 1.389, the TBS of L3 is 1.481, and the TBS of L4 is 1.494. Additionally, the TBS values for the L1-L4 combination, along with their corresponding T-Score and Z-Score, are also recorded.
[0178] Referring to FIG. 9d, an OCR processing result data table according to one embodiment of the present application is illustrated. The present application performs OCR on DXA examination report images to automatically extract multiple bone density-related values for each patient and converts them into a structured data table.
[0179] In one embodiment, 28 columns are generated for each patient, which include: (1) BMD, T-score, and Z-score for each of the lumbar vertebrae L1, L2, L3, and L4 (12 columns); (2) BMD, T-score, and Z-score for the right femur neck and total (6 columns); (3) BMD, T-score, and Z-score for the left femur neck and total (6 columns); and (4) TBS values for the lumbar vertebrae L1, L2, L3, and L4 (4 columns). Additionally, clinical information included in the DXA report, such as the study date, patient identification number (PID), date of birth (DoB), weight, and height, is also extracted.
[0180] In one embodiment of the present application, OCR was performed on DXA examination results from 2008 to 2019 to construct approximately 200,000 data records for about 80,000 patients. Specifically, 198,169 lumbar spine measurement data records, 195,780 left femur measurement data records, and 2,001 right femur measurement data records were extracted.
[0181] In the DXA-based fracture risk prediction method of the present application, the bone density values extracted through the OCR process are combined with feature vectors extracted from DXA images and utilized to predict the probability of fracture occurrence. Specifically, by applying an artificial intelligence model to DXA images to extract image-based feature vectors and combining them with quantitative bone density values (BMD, T-score, Z-score, TBS, etc.) extracted through OCR, it is possible to make a comprehensive fracture risk prediction that simultaneously utilizes image information and quantitative measurement information.
[0182] This method of combining DXA images and bone density values has the effect of expanding the clinical applications of existing DXA tests—which were previously used only for measuring bone density and diagnosing osteoporosis—to include predicting fracture risk at different time points. In particular, since a large volume of accumulated DXA test results can be automatically standardized through OCR and utilized for training artificial intelligence models, it has the advantage of facilitating the construction of large-scale training data.
[0183] FIG. 10 is a block diagram showing the overall configuration of a fracture risk prediction method using DXA images according to one embodiment of the present application.
[0184] Referring to FIG. 10, a DXA-based fracture risk prediction method according to one embodiment of the present application includes two parallel paths, an image-based feature extraction path and a BMD numerical extraction path, and combines the information extracted from the two paths to finally predict the probability of fracture occurrence for each of a plurality of observation points.
[0185] First, the image-based feature extraction path is described. The DXA image input unit (1010) receives a Dual-energy X-ray Absorptiometry (DXA) image. In one embodiment, the DXA image may include at least one of a lumbar spine image, a femur image, and a whole body image. The lumbar spine image is an image used for measuring bone density of the L1-L4 vertebrae, and the femur image is an image used for measuring bone density of the femoral neck and the total femur.
[0186] The region segmentation unit (1020) segments a plurality of regions of interest in the input DXA image. In one embodiment, the region segmentation unit (1020) may segment the bone region in the DXA image into a first region of interest and the soft tissue region into a second region of interest. The bone region includes areas where bone tissue is located, such as the lumbar vertebra, femoral head, femoral neck, and greater trochanter, and the soft tissue region includes areas where tissue other than bone tissue is located, such as muscle and fat. Since the DXA image is designed to distinguish between bone tissue and soft tissue using dual-energy X-rays, the bone region and the soft tissue region can be effectively segmented using the difference in attenuation coefficients for each energy.
[0187] The normalization unit (1030) performs energy-based normalization for each segmented region of interest. In one embodiment, the normalization unit (1030) may generate a first normalized image based on the pixel value distribution of the bone region and generate a second normalized image based on the pixel value distribution of the soft tissue region. Since the bone region has a high X-ray attenuation coefficient, it exhibits relatively high pixel values, and since the soft tissue region has a low X-ray attenuation coefficient, it exhibits relatively low pixel values. By performing normalization individually according to the pixel value distribution characteristics of each region, bone structure information and body composition information are preserved in an optimized form, respectively.
[0188] The feature extraction unit (1040) extracts feature vectors from a normalized DXA image using an artificial intelligence model. In one embodiment, the artificial intelligence model may be a feature extractor comprising at least one of a convolutional neural network (CNN), a vision transformer (ViT), and a residual neural network (ResNet). The feature extraction unit (1040) may extract feature vectors individually for each of the first normalized image and the second normalized image, and the plurality of feature vectors may be combined in an ensemble method of at least one of a soft voting method and a hard voting method.
[0189] Next, the path for extracting BMD values is described. The DXA report (1015) is a report in which DXA test results are recorded, and quantitative values such as bone density (BMD), T-score, and Z-score for each measurement site are listed in a table format. The OCR processing unit (1025) performs Optical Character Recognition (OCR) on the DXA report to automatically extract values related to bone density.
[0190] The BMD value (1035) includes quantitative bone density information extracted by the OCR processing unit (1025). In one embodiment, the BMD value may include the BMD, T-score, and Z-score for each of the lumbar vertebrae L1, L2, L3, and L4; the BMD, T-score, and Z-score for the left and right femoral necks and the total femur; and the TBS (Trabecular Bone Score) values for the lumbar vertebrae L1-L4. As described with reference to FIGS. 9a through 9d, these values are automatically extracted from the DXA examination report via OCR and converted into a structured data table.
[0191] The clinical information combining unit (1050) combines the image-based feature vector extracted from the feature extraction unit (1040) with the BMD value (1035). In one embodiment, the clinical information combining unit (1050) may generate an expanded feature vector by concatenating the feature vector and the BMD value, or fuse the two pieces of information using an attention mechanism. Through this, visual features extracted from the DXA image and quantitative bone density measurements are utilized complementarily.
[0192] The prediction unit (1060) predicts the probability of fracture occurrence for each of a plurality of observation points based on combined feature information. In one embodiment, the plurality of observation points may include at least two of 1 year, 2 years, 3 years, 5 years, and 10 years. As described with reference to FIG. 5, the artificial intelligence model of the prediction unit (1060) may be trained using training data in which a first label value is assigned to a point in time before the occurrence of a fracture for each of the plurality of observation points, and a second label value is assigned to a point in time after the occurrence of a fracture. Additionally, it may be trained using at least one of Survival Loss, Binary Cross Entropy Loss, and Log-likelihood Loss.
[0193] The survival curve generation unit (1070) generates a survival curve using the probability of fracture occurrence for each of the predicted multiple observation points. The survival curve is a curve that continuously shows how the probability of no fracture occurring changes over time.
[0194] The output unit (1080) provides the user with the predicted probability of fracture occurrence at each time point and the generated survival curve.
[0195] The DXA-based fracture risk prediction method of the present application has the effect of expanding the scope of application of existing DXA examinations. While existing DXA examinations have been primarily used for bone density measurement and T-score-based osteoporosis diagnosis, according to the present application, it is possible to predict fracture risk at different time points and generate survival curves by combining DXA images and bone density values. In particular, by utilizing visual features (bone shape, bone microstructure patterns, etc.) extracted from DXA images together with quantitative bone density values, improved prediction accuracy can be achieved compared to existing methods that use only BMD values.
[0196] Additionally, the DXA-based prediction method of the present application can be performed by the fracture risk prediction system (100) described with reference to FIG. 1. Specifically, the DXA image input unit (1010) corresponds to the medical image input unit (130), the region segmentation unit (1020) corresponds to the region segmentation unit (140), the normalization unit (1030) corresponds to the normalization unit (150), the feature extraction unit (1040) corresponds to the feature extraction unit (160), the clinical information combining unit (1050) corresponds to the clinical information combining unit (175), the prediction unit (1060) corresponds to the prediction unit (180), the survival curve generation unit (1070) corresponds to the survival curve generation unit (190), and the output unit (1080) corresponds to the output unit (195).
[0197] The operation of the system and method according to the embodiments described above may be implemented at least partially by a computer or by a computer program and may be recorded on a computer-readable recording medium. For example, it may be implemented with a program product comprising a computer-readable medium containing program code, which may be executed by a processor to perform any or all of the described steps, operations, or processes.
[0198] The computer may be a computing device such as a desktop computer, laptop computer, notebook, smartphone, or similar device, or any device that may be integrated. The computer is a device having one or more alternative and special-purpose processors, memory, storage space, and networking components (either wireless or wired). The computer may run an operating system such as, for example, an operating system compatible with Microsoft Windows, Apple OS X or iOS, a Linux distribution, or Google's Android OS.
[0199] The above computer-readable recording medium includes all types of recording devices in which data that can be read by a computer is stored. Examples of computer-readable recording media include ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc. Additionally, the computer-readable recording medium may be distributed across networked computer systems, so that computer-readable code can be stored and executed in a distributed manner. Furthermore, functional programs, codes, and code segments for implementing the present embodiment will be easily understood by a person skilled in the art to which the present embodiment belongs.
[0200] The present application described above has been explained with reference to the embodiments illustrated in the drawings, but this is merely illustrative and those skilled in the art will understand that various modifications and variations of the embodiments are possible therefrom. However, such modifications should be considered to be within the technical scope of protection of the present application. Accordingly, the true technical scope of protection of the present application should be determined by the technical concept of the appended claims.
[0201] The present application relates to a fracture risk prediction system and method using artificial intelligence-based medical images. By using an artificial intelligence model from medical images to predict the probability of fracture occurrence for each of multiple observation points, it provides fracture risk information at various points in time, such as 1 year, 2 years, 3 years, 5 years, and 10 years, thereby supporting customized treatment decisions tailored to the patient's life expectancy and clinical situation.
Claims
1. A method for predicting fracture risk using artificial intelligence-based medical images performed by a processor, Step of receiving medical images; A step of extracting feature vectors from the above medical image through an artificial intelligence model; A step of predicting the probability of fracture occurrence for each of a plurality of observation points based on the above feature vector; and A step of generating a survival curve using the probability of fracture occurrence for each of the above-mentioned multiple observation points; A method for predicting fracture risk using artificial intelligence-based medical images including 2. In Paragraph 1, The medical image above is, A method for predicting fracture risk using artificial intelligence-based medical images, characterized by including at least one of X-ray images, dual-energy X-ray absorptiometry (DXA) images, computed tomography (CT) images, magnetic resonance imaging (MRI) images, and ultrasound images.
3. In Paragraph 2, A method for predicting fracture risk using an artificial intelligence-based medical image, characterized in that the medical image is a posteroanterior (Chest PA) X-ray image.
4. In Paragraph 1, The above artificial intelligence model is, A method for predicting fracture risk using artificial intelligence-based medical images, characterized by having a feature extractor comprising at least one of a Convolutional Neural Network (CNN), a Vision Transformer (ViT), and a Residual Network (ResNet).
5. In Paragraph 1, Prior to the step of extracting the above feature vector, A step of generating a first normalized image by segmenting a first region of interest in the medical image above; and A step of generating a second normalized image by segmenting a second region of interest in the medical image above; Includes more, A method for predicting fracture risk using artificial intelligence-based medical images, characterized in that the step of extracting the above feature vector is performed individually for each of the first normalized image and the second normalized image.
6. In Paragraph 5, A method for predicting fracture risk using artificial intelligence-based medical images, characterized in that the first region of interest is a lung region, and the second region of interest is at least one of a heart region and a bone region.
7. In Paragraph 5, The steps of generating the first normalized image and generating the second normalized image are: A method for predicting fracture risk using artificial intelligence-based medical images, characterized by performing energy-based normalization based on the pixel value distribution characteristics of each region of interest.
8. In Paragraph 5, The step of predicting the probability of the above-mentioned fracture occurrence is, A method for predicting fracture risk using artificial intelligence-based medical images, characterized by including the step of combining a first prediction result for the first normalized image and a second prediction result for the second normalized image using at least one ensemble method among a soft voting method and a hard voting method.
9. In Paragraph 8, The step of predicting the probability of the above-mentioned fracture occurrence is, A method for predicting fracture risk using an artificial intelligence-based medical image, characterized by further combining a third prediction result for the original medical image in the ensemble manner.
10. In Paragraph 1, The above artificial intelligence model is, A method for predicting fracture risk using artificial intelligence-based medical images, characterized by being trained using training data in which, for each of multiple observation points, a first label value is assigned at a point in time before the occurrence of a fracture and a second label value is assigned at a point in time after the occurrence of a fracture.
11. In Paragraph 10, The above artificial intelligence model is, A method for predicting fracture risk using artificial intelligence-based medical images, characterized by being trained using at least one of Survival Loss, Binary Cross Entropy Loss, and Log-likelihood Loss.
12. In Paragraph 1, The above plurality of observation points are, A method for predicting fracture risk using artificial intelligence-based medical images, characterized by including at least two of 1 year, 2 years, 3 years, 5 years, and 10 years.
13. In Paragraph 1, The above fracture is an osteoporotic fracture, A method for predicting fracture risk using artificial intelligence-based medical images, characterized by including at least one of a vertebral fracture, a hip fracture, a humerus fracture, and a radius fracture.
14. In Paragraph 13, The above fracture is, If there is a history of surgery within a specified period from the date of fracture diagnosis; and In the case of a spinal fracture, if there is a history of imaging examinations within a prescribed period from the date of diagnosis of the fracture; A method for predicting fracture risk using artificial intelligence-based medical images, characterized by satisfying at least one of the following conditions.
15. In Paragraph 1, The step of generating the above survival curve is, Step of setting the maximum observation period; A step of dividing the above maximum observation period into a plurality of time intervals; and A step of calculating the probability of no fracture occurring in each of the above plurality of time intervals; A method for predicting fracture risk using artificial intelligence-based medical images, characterized by including 16. In Paragraph 15, The step of generating the above survival curve is, A method for predicting fracture risk using artificial intelligence-based medical images, characterized by further including a step of correcting the probability of fracture occurrence after a certain point in time for data that is censored at the end of observation.
17. In Paragraph 1, The step of predicting the probability of the above-mentioned fracture occurrence is, Step of acquiring clinical information of the subject; and A step of predicting the probability of fracture occurrence by combining the above feature vector and the above clinical information; Includes, A method for predicting fracture risk using artificial intelligence-based medical images, characterized in that the above clinical information includes at least one of age, gender, weight, height, bone density value, history of past fractures, family history, smoking status, steroid use status, and rheumatoid arthritis status.
18. In a fracture risk prediction system using artificial intelligence-based medical imaging, One or more processors; and Memory for storing one or more instructions executed by the above one or more processors; Includes, The above one or more processors, Receive medical images, and Feature vectors are extracted from the above medical image using an artificial intelligence model, and Predicting the probability of fracture occurrence for each of multiple observation points based on the above feature vector, and An AI-based medical imaging fracture risk prediction system characterized by generating a survival curve using the probability of fracture occurrence for each of the predicted multiple observation points.
19. In Paragraph 18, The medical image above is, X-ray images, Dual-energy X-ray Absorptiometry (DXA) images, Computed Tomography (CT) images, A fracture risk prediction system using artificial intelligence-based medical imaging, characterized by including at least one of magnetic resonance imaging (MRI) and ultrasound imaging.
20. In Paragraph 18, The above one or more processors, A first normalized image is generated by segmenting a first region of interest in the above medical image, and A second normalized image is generated by segmenting a second region of interest in the above medical image, and A fracture risk prediction system using artificial intelligence-based medical images, characterized by individually extracting feature vectors for each of the first normalized image and the second normalized image.
21. In Paragraph 20, The above one or more processors, An AI-based medical image fracture risk prediction system characterized by predicting the probability of fracture occurrence by combining the first prediction result for the first normalized image and the second prediction result for the second normalized image using at least one ensemble method among soft voting and hard voting methods.
22. In Paragraph 18, The above artificial intelligence model is, For each of the multiple observation points, training is performed using training data in which a first label value is assigned at the point in time before the occurrence of the fracture and a second label value is assigned at the point in time after the occurrence of the fracture, and An AI-based fracture risk prediction system using medical images, characterized by being trained using at least one of Survival Loss, Binary Cross Entropy Loss, and Log-likelihood Loss.
23. A computer-readable non-transient recording medium storing instructions for performing a fracture risk prediction method using artificial intelligence-based medical images, when executed by one or more processors, The above method is, Step of receiving medical images; A step of extracting feature vectors from the above medical image through an artificial intelligence model; A step of predicting the probability of fracture occurrence for each of a plurality of observation points based on the above feature vector; and A step of generating a survival curve using the probability of fracture occurrence for each of the above-mentioned multiple observation points; A computer-readable non-transient recording medium characterized by including