Osteoporotic compression fracture screening method based on multi-modal large language model

By employing a multimodal large language model screening method, the safety and accuracy issues of early screening for osteoporotic vertebral compression fractures without imaging equipment were resolved. This enabled safe and standardized screening at home and in primary healthcare institutions, reducing costs and improving the reliability and accessibility of screening results.

CN121983318AActive Publication Date: 2026-05-05THE AFFILIATED HOSPITAL OF SOUTHWEST MEDICAL UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE AFFILIATED HOSPITAL OF SOUTHWEST MEDICAL UNIV
Filing Date
2026-04-07
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve safe, standardized, and clinically significant early screening for osteoporotic vertebral compression fractures without imaging equipment and full physician involvement. This is particularly problematic in home and primary healthcare settings, where there is a high risk of concealed misdiagnosis, inaccurate assessment, and insufficient safety.

Method used

An osteoporotic compression fracture screening method based on a multimodal large language model is adopted. Through data collection, posture sequence construction, multimodal semantic feature extraction, feature aggregation and patient-level feature vector construction, and machine learning risk modeling, a structured quantitative scoring result is generated, and the probability prediction value of osteoporotic vertebral compression fracture is output.

Benefits of technology

It enables OVCF screening without the need for specialized imaging equipment, reducing medical costs, improving the safety and reliability of screening, enhancing the standardization and repeatability of home and community screening, supporting rapid screening and tiered medical decision-making, and adapting to deployment in primary care and multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121983318A_ABST
    Figure CN121983318A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of early screening of osteoporotic compression fracture, and provides an osteoporotic compression fracture screening method based on a multi-modal large language model. A multi-modal large language model is introduced to carry out quantitative evaluation on key functions such as posture alignment, motion coordination and pain-related reactions of a patient from images and videos, and the key functions are output in a standardized scoring form, so that structural features with clear clinical significance are automatically extracted. And constructing a machine learning model based on the above structured features, evaluating the OVCF occurrence risk, and determining key discriminant factors and the action direction thereof in combination with an SHAP feature contribution analysis and decision tree visualization method, so as to realize interpretable expression of the prediction process. The technical scheme which is safe, low in cost, explainable and easy to popularize is provided for early screening of the OVCF, and the method is suitable for various application scenes such as community and home screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of early screening technology for osteoporotic vertebral compression fractures, specifically a screening method for osteoporotic compression fractures based on a multimodal large language model. Background Technology

[0002] Osteoporotic compression fracture (OVCF), also known as osteoporotic vertebral compression fracture, is one of the most common and serious complications of osteoporosis. It is mainly caused by decreased bone density and destruction of bone microstructure, leading to compression deformation of the vertebral body under minor external forces. It often results in chronic pain, spinal deformity, and impaired mobility, severely impacting patients' quality of life and increasing the risk of secondary fractures. Currently, OVCF presents three main clinical challenges: ① high risk of insidious misdiagnosis; ② a high rate of disabling disease progression and a high-risk chain of fatal complications; and ③ diagnostic resource barriers.

[0003] However, early diagnosis of OVCF mainly relies on imaging examinations such as magnetic resonance imaging (MRI), X-ray or CT. MRI is sensitive to early lesions such as bone marrow edema, but it has problems such as expensive equipment, high examination costs, complicated operation and low coverage in primary medical institutions, making it difficult to meet the needs of large-scale population screening. Although X-ray is inexpensive, it is not sensitive enough to early microfractures and is easy to miss.

[0004] In recent years, for osteoporotic vertebral compression fractures, a physical examination method based on functional movement-induced pain response has been proposed in clinical practice for preliminary screening of fracture risk without imaging. This method typically involves guiding the patient to perform specific movements such as lying supine, turning over, and sitting up, and then combining this with pain score results to determine the presence of vertebral compression fracture risk. It has the advantages of being simple to operate and relatively inexpensive. However, this type of physical examination method has not yet been widely adopted in community or home settings, and its application remains primarily limited to medical institutions. The main reasons for this are as follows: First, existing physical examination methods rely heavily on patients' subjective ratings of pain levels during the assessment process. Different patients have significant differences in their perception and expression of pain, which can easily lead to scoring biases in the absence of professional guidance, thus affecting the reliability of screening results. Second, these examinations usually have high requirements for the sequence of movements and the standardization of postures. Patients in a home environment may find it difficult to accurately master the standard movements, resulting in the examination results being greatly affected by the quality of the movements. Third, patients may experience significant pain reactions during functional movements such as turning over and sitting up, and there is even a risk of aggravating potential fractures or causing secondary injuries. In a non-medical environment, the lack of real-time assessment and intervention by professionals makes it difficult to guarantee safety.

[0005] Based on the aforementioned issues, relying solely on manual physical examinations is insufficient to meet the practical needs of home self-testing and large-scale screening at the grassroots level. With the development of artificial intelligence (AI) technology, utilizing intelligent sensing and data analysis to assist and enhance the physical examination process has become an important direction for improving the safety, objectivity, and accessibility of screening. By analyzing patients' postures, movement processes, and pain-related behavioral responses in real time, AI models can, to some extent, compensate for the subjectivity of manual assessments, enabling timely identification and alerts for high-risk movements and abnormal pain responses. Simultaneously, through comprehensive analysis of multi-dimensional behavioral characteristics, it helps transform the subjective experience in traditional physical examinations into quantifiable and reproducible objective indicators, providing a basis for subsequent risk assessment and medical decision-making.

[0006] Therefore, how to use artificial intelligence technology to comprehensively analyze the postural characteristics, movement quality, and pain-related responses of patients during physical examinations without the need for professional imaging equipment and full-time physician involvement, and to construct a safe, standardized, and clinically significant early screening method and system for osteoporotic vertebral compression fractures, has become an urgent technical problem to be solved. Summary of the Invention

[0007] The purpose of this invention is to provide a screening method for osteoporotic compression fractures based on a multimodal large language model, addressing the aforementioned problems.

[0008] The technical solution adopted in this invention is as follows: a screening method for osteoporotic compression fractures based on a multimodal large language model, comprising the following steps:

[0009] S1. Data Acquisition and Posture Sequence Construction: Acquire patient information for screening osteoporotic vertebral compression fractures; preprocess the image and video data in the patient information to obtain image and video data; perform human posture analysis on the preprocessed image and video data based on the posture estimation algorithm, extract the spatial coordinate information of key points of the human skeleton and their temporal features over time, and construct posture sequence data containing key point annotations of the human body to characterize the patient's static posture features and dynamic functional movement features;

[0010] S2. Multimodal semantic feature extraction: Design structured cue words with clear clinical semantic constraints for the posture sequence data, input the posture sequence data into a multimodal large language model, and the multimodal large language model comprehensively analyzes the postural structure features and motor behavior features of the patient in the process of maintaining posture and performing functional actions to generate a structured quantitative score result that can reflect the patient's functional status.

[0011] S3. Feature aggregation and patient-level feature vector construction: Multiple sets of structured scoring results obtained for the same patient under different perspective images and different functional action tasks are summarized. Robust statistical methods are used to aggregate the score values ​​corresponding to the same semantic indicators to form a patient-level feature vector that represents the patient's overall posture and functional status.

[0012] S4. Machine Learning Risk Modeling and Probability Prediction: Using the patient-level feature vector and basic clinical information as model input, a supervised learning classification model is constructed and trained to model and analyze whether the patient has osteoporotic compression vertebral fractures, and output the probability prediction value of the patient having osteoporotic vertebral compression fractures.

[0013] S5. Fracture Screening and Risk Classification Steps: Based on the predicted probability values, calculate the screening performance indicators under different classification thresholds, combine the Youden index to determine the optimal classification threshold for screening osteoporotic vertebral compression fractures, and output the patient's osteoporotic vertebral compression fracture risk assessment results or risk classification results accordingly.

[0014] Furthermore, in the data acquisition and posture sequence construction, the patient information includes static photos of the patient's standard standing posture, functional videos, and basic clinical information;

[0015] The patient's standard standing static body photos include front, back, and side standing static body photos of the patient;

[0016] The functional videos include dynamic functional videos of the patient lying supine, turning to the left, turning to the right, and sitting up;

[0017] The basic clinical information includes the patient's age, gender, height, weight, etc.

[0018] Furthermore, OpenPose pose estimation was used to process static photos and functional videos of patients in standard standing positions to extract human pose information and generate pose sequence data containing human key point annotations.

[0019] Furthermore, in the multimodal semantic feature extraction, based on the structured cue words, the pose sequence data is semantically scored from at least four dimensions:

[0020] Postural alignment, postural symmetry, motor coordination, and pain response;

[0021] The posture alignment is used to characterize the alignment of the patient's overall spinal force line and trunk posture.

[0022] The postural symmetry is used to characterize the consistency of the patient's left and right limbs and trunk in static structure and dynamic movement.

[0023] The motor coordination is used to characterize the patient’s ability to coordinate and control movements involving multiple joints.

[0024] The pain response is used to characterize the protective posture, hesitation, or abnormal movement patterns that a patient exhibits due to pain during the execution of an action.

[0025] Furthermore, the multimodal large language model is constrained to output structured scoring results in JSON format. The scoring results are mapped to normalized continuous values ​​0-1, where 0 represents normal and 1 represents abnormal, and are used as input features for subsequent machine learning model analysis.

[0026] Furthermore, in the feature aggregation and patient-level feature vector construction, for each patient, structured scoring results from images from different perspectives and videos of different functional tasks are summarized and robust aggregation processing is performed. For the same indicator in multiple generation or multiple perspectives, median or mean aggregation is used to reduce noise caused by model randomness.

[0027] Furthermore, in the machine learning risk modeling and probability prediction, multiple supervised learning classification models are constructed, including logistic regression, support vector machine, random forest, gradient boosting and decision tree models. The model input is the multimodal scoring result, and the output is the probability prediction value of whether osteoporotic vertebral compression fracture exists.

[0028] Furthermore, in machine learning risk modeling, hierarchical K-fold cross-validation is used for model evaluation to ensure that the proportion of cases and controls is consistent across different compromises. Each compromise model only performs feature standardization and missing value imputation on the training set and is tested independently on the validation set. The performance of the machine learning model is evaluated using the area under the curve, accuracy, F1 score, sensitivity, and specificity metrics. The diagnostic capabilities of different models are compared, and the machine learning model that balances performance and interpretability is selected as the final solution.

[0029] Furthermore, the fracture screening and risk stratification process includes the following steps:

[0030] The model output probability calculation: The machine learning model outputs the predicted probability of osteoporotic vertebral compression fracture for each patient, labeled as p, and calculates the sensitivity and specificity under multiple thresholds τ.

[0031] The Youden index is calculated, the optimal classification threshold is selected, and the threshold is adjusted according to clinical screening needs.

[0032] Furthermore, in the calculation of the Youden exponent, the Youden exponent is calculated for each threshold using the following formula:

[0033] J(τ)=Sensitivity(τ)+Specificity(τ)−1;

[0034] Choose the threshold that maximizes J(τ) as the optimal classification threshold;

[0035] Where J(τ) is the Youden index corresponding to the threshold τ, and Sensitivity(τ) is the sensitivity / true positive rate (TPR) at the threshold τ, i.e., the true positive rate; the calculation formula is:

[0036] ;

[0037] Wherein, TP represents the number of true positives and FN represents the number of false negatives;

[0038] Specificity(τ) is the specificity / true negative rate (TNR) at the threshold τ, i.e., the true negative rate;

[0039] ;

[0040] TN represents the number of true negatives, and FP represents the number of false positives.

[0041] The beneficial effects of the present invention include at least one of the following;

[0042] 1. Reducing Screening Thresholds and Medical Costs: This invention provides a screening method for osteoporotic vertebral compression fractures (OVCFs) based on multimodal large language models and machine learning. The multimodal large language model is used to automatically extract clinically significant structured features from unstructured visual data such as ordinary images and videos, enabling early screening of OVCFs without relying on professional imaging equipment such as X-rays, CT, or MRI. This significantly reduces medical costs and the investment threshold for hardware equipment in primary healthcare institutions, and improves the accessibility of screening services.

[0043] 2. Achieve automatic quantification of high-level clinical semantic features: By designing structured cue words with clinical semantic constraints, a multimodal large language model is guided to analyze and score patient posture sequence data, outputting multidimensional structured scoring indicators including posture alignment, posture symmetry, motor coordination, and pain response. This transforms functional performance and pain-related behaviors that are originally difficult to quantify into numerical features that can be used for machine learning modeling, making up for the shortcomings of traditional geometric keypoint methods that can only describe low-level posture information.

[0044] 3. Enhance model robustness and result stability: By robustly aggregating the scoring results of the same patient under different perspectives, different functional actions, or multiple model generation conditions, the randomness and noise interference of single inference are reduced by using the median or mean method, which effectively improves the stability and consistency of patient-level feature vectors, thereby improving the reliability of the overall screening results.

[0045] 4. Improve the safety of home and community screening: This invention can be operated in the context of patient home or community screening. Through continuous analysis of the patient's movement process, posture changes and pain-related responses, it can identify potential high-risk movements or abnormal behavior patterns, provide patients with risk warnings or termination suggestions, reduce the risk of secondary injury caused by improper movements or severe pain, and make up for the lack of safety guarantee in traditional physical examinations in non-clinical environments.

[0046] 5. Improve the standardization and repeatability of the screening process: Compared with physical examination methods that rely on patients' subjective pain scores or manual observation, this invention uses an algorithmic model to objectively analyze posture and movement processes, thereby achieving automated evaluation of the examination process. This reduces the differences in results caused by different patients, different operators, and different environmental conditions, and improves the standardization and repeatability of the screening process.

[0047] 6. Support for rapid screening and tiered medical treatment decisions: This invention outputs the probability prediction value of osteoporotic vertebral compression fractures through a machine learning model and determines the optimal screening threshold by combining the Youden index, thereby realizing the quantitative assessment and tiered management of patient risk. It can provide a basis for decision-making on whether further imaging examinations or medical treatment are needed, which helps to optimize the allocation of medical resources and reduce unnecessary examination and treatment burdens.

[0048] 7. Enhance model interpretability and clinical credibility: By introducing SHAP feature contribution analysis and decision tree visualization methods, the contribution of each input feature in OVCF risk prediction is quantified, and the model decision path is displayed in a rule-based form, enabling doctors to intuitively understand the basis of the model's judgment, and to verify and adjust it in combination with clinical experience, thereby improving the credibility and feasibility of artificial intelligence screening models in clinical applications.

[0049] 8. Adaptable to grassroots and multi-scenario application deployment: The data collection method and model architecture adopted in this invention have low dependence on computing resources and professional equipment. It can be deployed in grassroots medical institutions and can also be extended to community screening or home remote testing scenarios. It has good scenario adaptability and promotion value, which is conducive to building a graded screening and hierarchical management osteoporosis prevention and control system. Attached Figure Description

[0050] Figure 1 The flowchart shows a screening method for osteoporotic compression fractures based on a multimodal large language model.

[0051] Figure 2 This is a schematic diagram of a screening method for osteoporotic compression fractures based on a multimodal large language model.

[0052] Figure 3 This is a schematic diagram of the structure of an electronic device. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described in the accompanying drawings can generally be arranged and designed in various different configurations.

[0054] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0055] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.

[0056] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0057] like Figure 1 and Figure 2 As shown, the osteoporotic compression fracture screening method based on a multimodal large language model includes the following steps:

[0058] S1. Data Acquisition and Posture Sequence Construction: Acquire patient information for screening osteoporotic vertebral compression fractures; preprocess the image and video data in the patient information to obtain image and video data; perform human posture analysis on the preprocessed image and video data based on the posture estimation algorithm, extract the spatial coordinate information of key points of the human skeleton and their temporal features over time, and construct posture sequence data containing key point annotations of the human body to characterize the patient's static posture features and dynamic functional movement features;

[0059] S2. Multimodal semantic feature extraction: Design structured cue words with clear clinical semantic constraints for the posture sequence data, input the posture sequence data into a multimodal large language model, and the multimodal large language model comprehensively analyzes the postural structure features and motor behavior features of the patient in the process of maintaining posture and performing functional actions to generate a structured quantitative score result that can reflect the patient's functional status.

[0060] S3. Feature aggregation and patient-level feature vector construction: Multiple sets of structured scoring results obtained for the same patient under different perspective images and different functional action tasks are summarized. Robust statistical methods are used to aggregate the score values ​​corresponding to the same semantic indicators to form a patient-level feature vector that represents the patient's overall posture and functional status.

[0061] S4. Machine Learning Risk Modeling and Probability Prediction: Using the patient-level feature vector and basic clinical information as model input, a supervised learning classification model is constructed and trained to model and analyze whether the patient has osteoporotic compression vertebral fractures, and output the probability prediction value of the patient having osteoporotic vertebral compression fractures.

[0062] S5. Fracture Screening and Risk Classification Steps: Based on the predicted probability values, calculate the screening performance indicators under different classification thresholds, combine the Youden index to determine the optimal classification threshold for screening osteoporotic vertebral compression fractures, and output the patient's osteoporotic vertebral compression fracture risk assessment results or risk classification results accordingly.

[0063] The purpose of this design is to provide a screening method for osteoporotic compression fractures based on a multimodal large language model. The multimodal large language model is used to extract structured clinical features from unstructured visual data, enabling OVCF screening without the need for professional imaging. This directly reduces medical costs and the equipment investment threshold for primary healthcare institutions. For example, in home or community settings, users can complete data collection through mobile phones or home cameras, extending screening from hospitals to the home environment and significantly improving accessibility.

[0064] In this embodiment, the patient information in the data acquisition and posture sequence construction includes the patient's standard standing static posture photos, functional videos, and basic clinical information;

[0065] The patient's standard standing static body photos include front, back, and side standing static body photos of the patient;

[0066] The functional videos include dynamic functional videos of the patient lying supine, turning to the left, turning to the right, and sitting up;

[0067] The basic clinical information includes the patient's age, gender, height, weight, etc.

[0068] Meanwhile, OpenPose pose estimation was used to extract human pose information from the patient's standard standing static photos and functional videos, and pose sequence data containing human key point annotations was generated.

[0069] In this embodiment, to enable the multimodal large language model to stably and repeatedly output clinically meaningful structured features from unstructured visual data such as patient images and videos, a structured cue word design method based on clinical semantic constraints is proposed. The cue words are not free-response descriptions, but rather constrain the reasoning process of the multimodal large language model through role limitations, task limitations, scoring dimension limitations, and output format limitations, ensuring that its output meets the requirements of consistency, interpretability, and quantifiability in medical screening scenarios. Example (using a lateral static body photograph as an example):

[0070] You are a spinal surgeon who is skilled at assessing changes in physiological curvature using lateral radiographs.

[0071] [Input] A lateral photograph (taken from the side) of a human body in a standard standing spinal posture. The image has been annotated with skeletal key points using OpenPose.

[0072] Please evaluate the following indicators, rating each one on a scale of 0–1 (0 for normal, 1 for severe abnormality), and provide an explanation.

[0073] Output metrics and definitions:

[0074] 1. Thoracic Kyphosis: Whether the thoracic kyphosis is within the normal range; 0 = moderate, 1 = excessive kyphosis or straightening.

[0075] 2. Lumbar lordosis: Whether the lumbar lordosis is normal; 0 = good curvature, 1 = flat or reverse curvature.

[0076] 3. Trunk Inclination: The angle at which the trunk leans forward relative to the vertical; 0 = upright, 1 = significantly forward lean.

[0077] 4. Head Position: Whether the earlobe is in front of the acromion; 0 = aligned, 1 = head protrusion.

[0078] 5. Sagittal Balance: Is the horizontal distance between C7 and S1 too large? 0 = good balance, 1 = unbalanced.

[0079] Output format requirements:

[0080] Please output in standard JSON format, with each metric including a score and explanation.

[0081] Example output:

[0082] {

[0083] "ThoracicKyphosis":{"score":0.25,"explanation":"Slightly increased thoracic kyphosis but not abnormal"}

[0084] "LumbarLordosis":{"score":0.30,"explanation":"decreased lumbar lordosis"},

[0085] "TrunkInclination":{"score":0.32,"explanation":"Significant forward lean of the torso"},

[0086] "HeadPosition":{"score":0.28,"explanation":"Head protruding forward"},

[0087] "SagittalBalance":{"score":0.30,"explanation":"Mild sagittal imbalance"}

[0088] }

[0089] Note: Only return JSON, do not add any extra text.

[0090] When the multimodal large language model scores using prompts, it uses four dimensions—posture alignment, posture symmetry, motor coordination, and pain response—to score the posture sequence data.

[0091] The posture alignment is used to assess the overall alignment of the patient's spine.

[0092] The postural symmetry is used to assess the consistency between the patient's left and right limb and trunk structures and movement patterns.

[0093] The motor coordination is used to assess a patient's ability to coordinate multiple joint movements.

[0094] The pain response is used to assess protective postures and abnormal movement patterns during patient movements.

[0095] Furthermore, the multimodal large language model is constrained to output structured scoring results in JSON format. The scoring results are mapped to normalized continuous values ​​0-1, where 0 represents normal and 1 represents abnormal, and are used as input features for subsequent machine learning model analysis.

[0096] The purpose of this design is to guide multimodal large language models to perform quantitative scoring through structured cue word design. Specifically, this includes setting the model to analyze from the perspective of orthopedic or sports medicine experts, and clearly specifying evaluation dimensions such as posture alignment, symmetry, motor coordination, and pain response. At the same time, the model is constrained to output structured scoring results in JSON format to improve the parsing and reproducibility of the results.

[0097] It should be noted that OVCF leads to changes in spinal stability, abnormal posture, and motor dysfunction. Its biomechanical changes are mainly manifested in postural structural deviation, left-right asymmetry, motor control impairment, and pain-related compensatory behaviors. Therefore, a postural function scoring system is constructed from four dimensions: postural alignment, postural symmetry, motor coordination, and pain response, to comprehensively characterize the features of spinal functional impairment. Postural alignment is used to assess the overall alignment of the spinal force lines and is an important indicator reflecting spinal stability; postural symmetry is used to assess the consistency of left and right limb and trunk structures and movement patterns, reflecting functional compensation and neuromuscular dysfunction; motor coordination is used to assess multi-joint coordinated movement ability, reflecting the state of neuromuscular control function; and pain response is used to assess the potential pain compensation behaviors reflected in protective postures and abnormal movement patterns during movement.

[0098] Based on the previous two steps, static posture photographs (front, back, and lateral views) of the patient in a standard standing position were collected to represent static posture structure information, and videos of functional movements (such as turning over and sitting up) were collected to represent dynamic motor function information. Human keypoint sequences were extracted using a posture estimation algorithm. A multimodal large language model was used to perform structured scoring on both images and videos, extracting feature indicators such as posture alignment, posture symmetry, motor coordination, and pain response. Subsequently, the features from the images and videos were concatenated and aggregated to construct a patient-level multimodal feature vector, which was used for subsequent machine learning model training and OVCF risk prediction.

[0099] In this embodiment, to verify the clinical effectiveness and interpretability of the scores generated by the multimodal large language model, a quantitative correlation analysis was further performed on the model output scores and corresponding text explanations. For key indicators such as spinal verticality, scapular symmetry, trunk tilt, sagittal balance, movement completion time, and thoracic kyphosis angle, physical quantities explicitly reported in the model explanation text (such as lateral offset in centimeters, tilt angle, and movement time in seconds) were extracted and correlated with the normalized scores output by the model, as shown in the table below:

[0100] Table 1. Interpretability association results based on key features.

[0101]

[0102] Data shows that the Pearson correlation coefficient between spinal verticality score and C7-S1 lateral offset reached 0.982 (95% confidence interval: 0.975–0.987), the correlation coefficient between scapular symmetry score and scapular height difference was 0.971 (95% confidence interval: 0.961–0.979), and the correlation coefficient between trunk tilt score and trunk tilt angle was 0.978 (95% confidence interval: 0.970–0.984). These results demonstrate that the multimodal large language model does not produce random or uninterpretable scoring, but rather stably maps physical measurements observed in images and videos into quantitative scores with clear clinical semantics, and its output shows a high degree of consistency with objective clinical measurement standards.

[0103] In this embodiment, between feature aggregation and patient-level feature vector construction (S3) and machine learning risk modeling (S4), in order to realize the quantitative mapping relationship between the structured scoring results output by the multimodal large language model and the risk prediction of osteoporotic vertebral compression fracture, the intermediate feature calculation process and the model training process are described by the following mathematical modeling.

[0104] First, for each patient, an initial score set is constructed based on structured scoring results generated by a multimodal large language model under different perspective images (anterior, posterior, lateral) and videos of different functional movements (supine, rolling over, sitting up, etc.). Let the first... The patient in The evaluation criteria (including postural alignment, postural symmetry, motor coordination, and pain response, etc.) are as follows: The score result of the observation is in This indicates the number of observations corresponding to this indicator.

[0105] To reduce the impact of randomness in multiple inference processes of multimodal large language models and improve feature stability, robust aggregation of scoring results for the same indicator under different observation conditions is performed to obtain patient-level feature representations:

[0106] ;

[0107] in, Indicates the first The patient in The final characteristic value on each evaluation indicator Indicates the first The patient in The first indicator The score for each observation. This represents the total number of observations under the same indicator.

[0108] Furthermore, the aggregated results of all evaluation indicators are concatenated to construct a patient-level multimodal feature vector:

[0109] ;

[0110] in, Indicates the first The feature vector of each patient To determine the number of feature dimensions, the feature vector not only includes posture function scores but can also further integrate basic clinical information of the patient, including age, gender, height, and weight, thereby forming an extended feature representation:

[0111] ;

[0112] in This represents the expanded feature vector after fusion. This represents a vector of the patient's clinical information.

[0113] Based on this, the patient-level feature vector Risk modeling is performed using a supervised learning classification model. Taking a logistic regression model as an example, its function for predicting the probability of osteoporotic vertebral compression fractures is defined as:

[0114] ;

[0115] in, Indicates the first The predicted probability of an osteoporotic vertebral compression fracture in a patient. This indicates the true label (1 indicates a fracture, 0 indicates no fracture). These are the model weight parameters. Indicates the transpose term. This is a bias term.

[0116] During model training, the model parameters are optimized by minimizing the difference between the predicted probabilities and the true labels. The number of training samples is defined as... The real label is The binary classification cross-entropy loss function is used as the optimization objective, and its expression is:

[0117] ;

[0118] Where N represents the total number of training samples, Indicates the true label, This represents the probability predicted by the model.

[0119] During parameter optimization, an iterative update method based on gradient descent is used to update the model parameters, and the update rules are as follows:

[0120] ;

[0121] in, The learning rate is used to control the step size for updating model parameters. Loss functions (such as cross-entropy loss) are measures of the difference between model predictions and true labels. This represents the loss function. For weight vector The partial derivative, or gradient, represents the value when... Loss when small changes occur The rate of change indicates the direction in which the loss is reduced most quickly; It is a loss function For bias terms The partial derivatives of the loss function with respect to the parameters The gradient.

[0122] Furthermore, when using a gradient boosting model for risk modeling, an ensemble model is constructed by progressively learning residuals, and its iterative update process is represented as follows:

[0123] ;

[0124] in, Indicates the first The overall model after the next iteration; It is the first The previous ensemble model after the next iteration represents the predictive power that the current model has accumulated; It is the first The new base learner added in the next iteration represents the fit to the current model. The residual between the predicted value and the actual value (i.e., the direction of gradient descent); It is the first Individual learners The weighting coefficients are used to control the contribution of newly added base learners to the final model and prevent overfitting.

[0125] Through the above feature calculation and model training process, the structured scoring results output by the multimodal large language model are transformed into patient-level features with clear mathematical expressions, and further mapped to the risk probability of osteoporotic vertebral compression fractures, realizing an end-to-end quantitative modeling process from unstructured visual data to clinical risk assessment results.

[0126] In this embodiment, in the feature aggregation and patient-level feature vector construction, for each patient, the structured scoring results from images from different perspectives and videos of different functional tasks are summarized and robust aggregation processing is performed. For the same indicator in multiple generation or multiple perspectives, the median or mean aggregation method is used to reduce the noise caused by the randomness of the model.

[0127] Meanwhile, in the machine learning risk modeling and probability prediction, a variety of supervised learning classification models are constructed, including logistic regression, support vector machine, random forest, gradient boosting and decision tree models. The input of the model is the multimodal scoring result, and the output is the probability prediction value of whether osteoporotic vertebral compression fracture exists.

[0128] Furthermore, in machine learning risk modeling, hierarchical K-fold cross-validation is used for model evaluation to ensure that the proportion of cases and controls is consistent across different compromises. Each compromise model only performs feature standardization and missing value imputation on the training set and is tested independently on the validation set. The performance of the machine learning model is evaluated using the area under the curve, accuracy, F1 score, sensitivity, and specificity metrics. The diagnostic capabilities of different models are compared, and the machine learning model that balances performance and interpretability is selected as the final solution.

[0129] In this embodiment, the fracture screening and risk grading step includes the following steps:

[0130] The model output probability calculation: The machine learning model outputs the predicted probability of osteoporotic vertebral compression fracture for each patient, labeled as p, and calculates the sensitivity and specificity under multiple thresholds τ.

[0131] The Youden exponent is calculated for each threshold using the following formula:

[0132] J(τ)=Sensitivity(τ)+Specificity(τ)−1;

[0133] Choose the threshold that maximizes J(τ) as the optimal classification threshold;

[0134] Where J(τ) is the Youden index corresponding to the threshold τ, and Sensitivity(τ) is the sensitivity / true positive rate (TPR) at the threshold τ, i.e., the true positive rate; the calculation formula is:

[0135] ;

[0136] Wherein, TP represents the number of true positives and FN represents the number of false negatives;

[0137] Specificity(τ) is the specificity / true negative rate (TNR) at the threshold τ, i.e., the true negative rate;

[0138] ;

[0139] Wherein, TN represents the number of true negatives and FP represents the number of false positives;

[0140] Clinically-oriented threshold optimization: The threshold is adjusted according to clinical screening needs.

[0141] The purpose of this design is to quantify the contribution of each input feature to the predicted probability through SHAP analysis and decision tree visualization, thereby revealing the positive or negative impact of features such as posture alignment, symmetry, motor coordination, and pain response on the model's prediction results. For each patient, the predicted probability can be decomposed into a combination of baseline risk and the contributions of each feature; the SHAP value is used to represent the marginal contribution of a single feature to the predicted probability. This interpretation mechanism supports both global feature importance analysis and individual-level prediction interpretation, improving the transparency and clinical credibility of the model's prediction results.

[0142] In this embodiment, five models are used: Decision Tree, Gradient Boosting, Logistic Regression, Random Forest, and Radial Basis Function Support Vector Machine (SVM_RBF). The cross-validation performance is assessed based on Area Under the Curve (AUC), Standard Deviation (std), Accuracy, F1 Score, Sensitivity, and Specificity, as shown in the table below:

[0143] Table 2 compares the cross-validation performance of OVCF classification candidate models based on multimodal scoring.

[0144]

[0145] Among them, Random Forest performed best in several core metrics: AUC value of 0.95: possesses the strongest overall discrimination ability; Accuracy (0.88) and F1 value (0.89) are the highest: performs best in classification accuracy and balance between precision and recall; Sensitivity (0.82) and specificity (0.84) are balanced and excellent: has strong ability to identify both positive and negative samples; Standard deviation is generally less than 0.02: model stability is excellent, and cross-validation results fluctuate very little.

[0146] Gradient Boosting's overall performance is second only to Random Forest: AUC value of 0.90: strong overall discrimination ability; sensitivity (0.81): good performance; outstanding ability to identify positive samples; accuracy and F1 value are slightly lower than Random Forest, but stability is also very good.

[0147] The decision tree performed at a moderate level: the AUC value was 0.88, indicating that the overall discrimination ability was acceptable, and the accuracy and F1 score were good, but the sensitivity (0.72) was relatively low, and the ability to identify positive samples was average.

[0148] Logistic Regression and Radial Basis Kernel Support Vector Machine (SVM_RBF) performed relatively poorly: their AUC values ​​were 0.78 and 0.79, respectively. Their overall discrimination ability was weak, and their accuracy, F1 score and other indicators were also at a low level. The sensitivity (0.73) and specificity (0.78) of logistic regression were not ideal.

[0149] The standard deviation (std) of all models is less than 0.04, indicating that the performance fluctuations are relatively small under cross-validation. Random forest and gradient boosting tree have the best stability, with standard deviations generally around 0.01.

[0150] In this embodiment, a hyperparameter scan of the decision tree was performed under the settings of maximum depth and node size. The maximum depth, minimum number of leaf samples, minimum number of split samples, area under the curve (AUC), accuracy, F1 score, sensitivity, and specificity were statistically analyzed, as shown in the table below:

[0151] Table 3 shows the optimization of decision tree hyperparameters based on maximum depth and node size settings.

[0152]

[0153] From the table above, the maximum depth is the key influencing factor overall: as the maximum depth increases from 1 to 5, the model's core metrics such as AUC, accuracy, and F1 score continuously improve. Performance peaks at depth 5: with a maximum depth of 5, a minimum number of samples per leaf node of 5, and a minimum number of samples per split of 2, the AUC reaches its highest value of 0.95, and the accuracy (0.88) and F1 score (0.89) also reach optimal levels. Performance declines when depth increases to 6: when the depth increases from 5 to 6, all metrics show a decline, indicating that the model is overfitting.

[0154] Impact of minimum splitting sample count: At a maximum depth of 5, increasing the minimum splitting sample count from 2 to 10 causes a significant drop in AUC from 0.95 to 0.82, indicating that a smaller splitting threshold can capture more effective features. Impact of minimum leaf node sample count: Increasing the leaf node sample count from 1 to 5 significantly improves model performance, indicating that appropriately increasing the leaf node sample count can enhance model stability.

[0155] Optimal configuration: Maximum depth = 5, minimum number of leaf nodes = 5, minimum number of split samples = 2. Advantages: With this configuration, the model achieves the best balance in terms of discriminative ability (AUC), classification accuracy (accuracy, F1 score), and the ability to distinguish between positive and negative samples (sensitivity, specificity).

[0156] When the depth is 1 or 2, all metrics are around 0.6, indicating that the model is too simple and underfitting. When the depth increases to 6, the performance drops significantly, indicating that the model begins to overfit and its generalization ability deteriorates. When the maximum depth is 5, the model achieves the optimal balance between fitting ability and generalization ability.

[0157] like Figure 3 The diagram illustrates the structure of an electronic device according to an embodiment of this application. The electronic device includes a memory and a processor. The memory stores computer-readable instructions, and the processor executes the computer-readable instructions stored in the memory to implement a data processing method as described in any of the above embodiments.

[0158] In this embodiment of the application, the electronic device further includes a bus and a computer program stored in the memory and executable on the processor, such as a program for a data processing method.

[0159] Combination Figure 3 The memory in the electronic device stores a plurality of computer-readable instructions to implement a data processing method, and the processor can execute the plurality of instructions to implement the method.

[0160] Specifically, the processor's implementation method for the above instructions can be found in the description of the relevant steps in the corresponding embodiment of the figure, and will not be repeated here.

[0161] Those skilled in the art will understand that the schematic diagram is merely an example of an electronic device and does not constitute a limitation on the electronic device. The electronic device may be a bus-type structure or a star-type structure. The electronic device may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, the electronic device may also include input / output devices, network access devices, etc.

[0162] It should be noted that electronic devices are merely examples. Other existing or future electronic products that are suitable for this application should also be included within the scope of protection of this application and are incorporated herein by reference.

[0163] The memory includes at least one type of readable storage medium, which can be non-volatile or volatile. The readable storage medium includes flash memory, portable hard drives, multimedia cards, card-type memory (e.g., SD or DX memory), magnetic storage, magnetic disks, optical disks, etc. In some embodiments, the memory can be an internal storage unit of an electronic device, such as a portable hard drive. In other embodiments, the memory can be an external storage device of the electronic device, such as a plug-in portable hard drive, smart memory card (SMC), secure digital card (SD), flash card, etc. The memory can be used not only to store application software and various types of data installed on the electronic device, such as the code for an osteoporotic compression fracture screening method based on a multimodal large language model, but also to temporarily store data that has been output or will be output.

[0164] In some embodiments, a processor may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions. This includes combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processing units, and various control chips. The processor is the control unit of the electronic device, connecting various components of the device through various interfaces and lines. It executes programs or modules stored in the memory (e.g., programs for screening osteoporotic compression fractures based on multimodal large language models) and calls data stored in the memory to perform various functions and process data within the electronic device.

[0165] The processor executes the operating system of the electronic device and various installed applications. The processor executes the applications to implement the steps in the above embodiments of the osteoporotic compression fracture screening method based on multimodal large language models, such as the steps shown in the figure.

[0166] For example, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program in an electronic device. For example, the computer program may be divided into a receiving module, a preprocessing module, a projection module, and a determining module.

[0167] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the osteoporotic compression fracture screening method based on a multimodal large language model described in the various embodiments of this application.

[0168] When modules / units integrated into an electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.

[0169] This application provides a computer-readable storage medium storing computer-readable instructions. These instructions are executed by a processor in an electronic device to implement the osteoporotic compression fracture screening method based on a multimodal large language model as described in any of the above embodiments. It can be applied to one or more electronic devices. An electronic device is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0170] Electronic devices can be any electronic product that allows human-computer interaction with a customer, such as personal computers, tablets, smartphones, personal digital assistants (PDAs), game consoles, interactive network television (IPTV), smart wearable devices, etc.

[0171] Electronic devices may also include network devices and / or client devices. The network devices include, but are not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0172] The networks in which electronic devices are located include, but are not limited to, the Internet, wide area networks, metropolitan area networks, local area networks, and virtual private networks (VPNs).

[0173] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, and other memory.

[0174] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.

[0175] The bus can be a Peripheral Component Interconnect Standard (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus. The bus is configured to implement the connection and communication between the memory and at least one processor, etc.

[0176] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0177] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0178] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0179] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in the specification may also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.

[0180] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A screening method for osteoporotic compression fractures based on a multimodal large language model, characterized in that, Includes the following steps: S1. Data Acquisition and Posture Sequence Construction: Obtain patient information for screening osteoporotic vertebral compression fractures, and perform preprocessing operations on the image data and video data in the patient information to obtain image and video data; Human pose analysis is performed on preprocessed image and video data based on pose estimation algorithms. Spatial coordinate information of human skeletal key points and their temporal features over time are extracted. Pose sequence data containing human key point annotations are constructed to characterize the static pose features and dynamic functional movement features of patients. S2. Multimodal semantic feature extraction: Design structured cue words with clear clinical semantic constraints for the posture sequence data, input the posture sequence data into a multimodal large language model, and the multimodal large language model comprehensively analyzes the postural structure features and motor behavior features of the patient in the process of maintaining posture and performing functional actions to generate a structured quantitative score result that can reflect the patient's functional status. S3. Feature aggregation and patient-level feature vector construction: Multiple sets of structured scoring results obtained for the same patient under different perspective images and different functional action tasks are summarized. Robust statistical methods are used to aggregate the score values ​​corresponding to the same semantic indicators to form a patient-level feature vector that represents the patient's overall posture and functional status. S4. Machine Learning Risk Modeling and Probability Prediction: Using the patient-level feature vector and basic clinical information as model input, a supervised learning classification model is constructed and trained to model and analyze whether the patient has osteoporotic compression vertebral fractures, and output the probability prediction value of the patient having osteoporotic vertebral compression fractures. S5. Fracture Screening and Risk Classification Steps: Based on the predicted probability values, calculate the screening performance indicators under different classification thresholds, combine the Youden index to determine the optimal classification threshold for screening osteoporotic vertebral compression fractures, and output the patient's osteoporotic vertebral compression fracture risk assessment results or risk classification results accordingly.

2. The method for screening osteoporotic compression fractures based on a multimodal large language model according to claim 1, characterized in that, In the data acquisition and posture sequence construction, patient information includes static photos of the patient's standard standing posture, functional videos, and basic clinical information; The patient's standard standing static body photos include front, back, and side standing static body photos of the patient; The functional videos include dynamic functional videos of the patient lying supine, turning to the left, turning to the right, and sitting up; The basic clinical information includes the patient's age, gender, height, and weight.

3. The method for screening osteoporotic compression fractures based on a multimodal large language model according to claim 2, characterized in that, The OpenPose pose estimation process was used to extract human pose information from static photos and functional videos of patients in standard standing positions and generate pose sequence data containing human key point annotations.

4. The method for screening osteoporotic compression fractures based on a multimodal large language model according to claim 1, characterized in that, In the multimodal semantic feature extraction, based on the structured cue words, the pose sequence data is semantically scored from at least four dimensions: Postural alignment, postural symmetry, motor coordination, and pain response; in, The posture alignment is used to characterize the alignment of the patient's overall spinal alignment and trunk posture. The postural symmetry is used to characterize the consistency of the patient's left and right limbs and trunk in static structure and dynamic movement. The motor coordination is used to characterize the patient’s ability to coordinate and control movements involving multiple joints. The pain response is used to characterize the protective posture, hesitation, or abnormal movement patterns that a patient exhibits due to pain during the execution of an action.

5. The method for screening osteoporotic compression fractures based on a multimodal large language model according to claim 4, characterized in that, The multimodal large language model is constrained to output structured scoring results in JSON format. The scoring results are mapped to normalized continuous values ​​0-1, where 0 represents normal and 1 represents abnormal, and are used as input features for subsequent machine learning model analysis.

6. The method for screening osteoporotic compression fractures based on a multimodal large language model according to claim 1, characterized in that, In the feature aggregation and patient-level feature vector construction, for each patient, structured scoring results from images from different perspectives and videos of different functional tasks are summarized and robust aggregation processing is performed. For the same indicator in multiple generation or multiple perspectives, median or mean aggregation is used to reduce noise caused by model randomness.

7. The method for screening osteoporotic compression fractures based on a multimodal large language model according to claim 1, characterized in that, In the machine learning risk modeling and probability prediction, various supervised learning classification models are constructed, including logistic regression, support vector machine, random forest, gradient boosting and decision tree models. The input of the model is the multimodal scoring results, and the output is the probability prediction value of whether osteoporotic vertebral compression fracture exists.

8. The method for screening osteoporotic compression fractures based on a multimodal large language model according to claim 7, characterized in that, In machine learning risk modeling, hierarchical K-fold cross-validation is used for model evaluation to ensure that the proportion of cases and controls is consistent across different compromises. Each compromise model is only standardized for features and imputed for missing values ​​on the training set and tested independently on the validation set. The performance of the machine learning model is evaluated using the area under the curve, accuracy, F1 score, sensitivity, and specificity metrics. The diagnostic capabilities of different models are compared, and the machine learning model that balances performance and interpretability is selected as the final solution.

9. The method for screening osteoporotic compression fractures based on a multimodal large language model according to claim 1, characterized in that, The fracture screening and risk stratification process includes the following steps: The model output probability calculation: The machine learning model outputs the predicted probability of osteoporotic vertebral compression fracture for each patient, labeled as p, and calculates the sensitivity and specificity under multiple thresholds τ. The Youden index is calculated, the optimal classification threshold is selected, and the threshold is adjusted according to clinical screening needs.

10. The method for screening osteoporotic compression fractures based on a multimodal large language model according to claim 9, characterized in that, In the calculation of the Youden exponent, the Youden exponent is calculated for each threshold using the following formula: J(τ)=Sensitivity(τ)+Specificity(τ)−1; Choose the threshold that maximizes J(τ) as the optimal classification threshold; Where J(τ) is the Youden index corresponding to the threshold τ, and Sensitivity(τ) is the sensitivity / true positive rate (TPR) at the threshold τ, i.e., the true positive rate; the calculation formula is: ; Wherein, TP represents the number of true positives and FN represents the number of false negatives; Specificity(τ) is the specificity / true negative rate (TNR) at the threshold τ, i.e., the true negative rate; ; TN represents the number of true negatives, and FP represents the number of false positives.

Citation Information

Patent Citations

  • Information processing method and electronic device

    CN105844134A

  • Vertebral compression fracture multi-modal intelligent diagnosis system based on deep learning

    CN113384261A

  • Lumbar pain prediction system

    CN119049720A

  • Recognition model, training method and prediction method for osteoporosis and osteoporosis centrum compression fracture

    CN119446561A

  • Osteoporosis diagnosis method based on image recognition

    CN120495295A