A spine curvature measurement method, device and equipment based on a visual large model
Patent Information
- Application Number
- CN202610096226.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-23
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-01-23
AI Technical Summary
该过程繁琐耗时,其结果高度依赖操作者的主观经验和判断,准确性低
[0009] This disclosure also provides a computer-readable storage medium storing a computer program for performing the spinal curvature measurement method based on a large visual model as provided in this disclosure.
Smart Images

Figure CN122156060B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of medical image processing technology, and in particular to a method, apparatus and device for measuring spinal curvature based on a large visual model. Background Technology
[0002] Scoliosis is a common three-dimensional spinal deformity, and its diagnosis and severity assessment are crucial in clinical practice. Accurate measurement of spinal curvature angles based on imaging is the core basis for assessing the degree of scoliosis and developing a treatment plan. Currently, the most common clinical method for angle measurement is full-spine X-ray.
[0003] In traditional measurement methods, physicians need to manually identify key vertebrae and specific anatomical points representing scoliosis characteristics on long-span images containing multiple vertebrae, based on experience. They then draw measurement baselines and finally use protractors to calculate the angle. This process is tedious and time-consuming, and the results are highly dependent on the operator's subjective experience and judgment, resulting in low accuracy.
[0004] In automated measurement methods, relying on traditional image processing algorithms or dedicated deep learning models designed for specific tasks often results in insufficient generalization ability and robustness when faced with challenges such as complex and variable image quality in clinical settings, significant anatomical differences between individuals, and key points being obscured by foreign objects or implants. Summary of the Invention
[0005] To solve the above-mentioned technical problems, or at least partially solve them, this disclosure provides a method, apparatus, and device for measuring spinal curvature based on a large visual model.
[0006] This disclosure provides a method for measuring spinal curvature based on a large visual model. The method includes: constructing an integrated progressive spinal measurement framework; wherein the integrated progressive spinal measurement framework includes a feature extraction module and a progressive task decoder module; the feature extraction module is constructed based on a large visual model, and the progressive task decoder module is constructed using multiple lightweight dedicated task heads; inputting a target image into the integrated progressive spinal measurement framework, using the feature extraction module to extract multi-scale high-resolution feature maps of the target image, using the progressive task decoder module to perform a spinal measurement task on the multi-scale high-resolution feature maps to obtain vertebral key point information; and using the vertebral key point information to calculate the angle information of spinal curvature.
[0007] This disclosure also provides a spinal curvature measurement device based on a large visual model. The device includes: a construction module for constructing a feature extraction module based on the large visual model; a progressive task decoder module using a lightweight dedicated task head; and an integrated progressive spinal measurement framework combining the feature extraction module and the progressive task decoder module; a detection module for inputting a target image into the integrated progressive spinal measurement framework, extracting multi-scale high-resolution feature maps of the target image using the feature extraction module, and performing a spinal measurement task on the multi-scale high-resolution feature maps using the progressive task decoder module to obtain vertebral key point information; and a calculation module for calculating the angle information of spinal curvature using the vertebral key point information.
[0008] This disclosure also provides a computing device, which includes: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the spinal curvature measurement method based on a large visual model as provided in this disclosure.
[0009] This disclosure also provides a computer-readable storage medium storing a computer program for performing the spinal curvature measurement method based on a large visual model as provided in this disclosure.
[0010] In this implementation, a feature extraction module is built on top of a large visual model, providing pixel-level detail support for vertebral angle localization, thus improving measurement accuracy and generalization ability. Simultaneously, a progressive task decoder module is constructed based on a lightweight dedicated task head. Combined with the feature extraction module and the progressive task decoder module, a highly integrated large model system is generated for fully automated spinal curvature detection. This improves model stability and reliability. Furthermore, multiple task heads share the multi-scale, high-resolution, dense features output by the feature extraction module, avoiding computational redundancy from repeatedly extracting features for each task, simplifying clinical operations, and improving diagnostic efficiency. Compared to existing neural network models used in spinal curvature measurement, this application adopts a scheme based on a large visual model combined with a lightweight task head. This enhances generalization ability, improves robustness, and optimizes performance in handling complex tasks while reducing computational redundancy, thereby achieving more efficient measurement and analysis. Attached Figure Description
[0011] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0012] Figure 1A simplified flowchart illustrating a method for measuring spinal curvature provided in this embodiment of the present disclosure; Figure 2 This is a schematic flowchart illustrating a method for measuring spinal curvature provided in an embodiment of the present disclosure. Figure 3 A schematic diagram of a feature extraction module provided in an embodiment of this disclosure; Figure 4 A schematic diagram of a progressive task decoder module provided in an embodiment of this disclosure; Figure 5 A flowchart illustrating an integrated progressive spine measurement framework training method provided in this embodiment of the disclosure; Figure 6 A schematic diagram of a digital X-ray image of a vertebra provided in an embodiment of this disclosure; Figure 7 A schematic diagram illustrating the Cobb angle annotation of a vertebra as provided in an embodiment of this disclosure; Figure 8 A schematic diagram of a spinal curvature measurement device based on a visual large model provided in this disclosure embodiment; Figure 9 This is a schematic diagram of the structure of a computing device provided in an embodiment of the present disclosure. Detailed Implementation
[0013] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0014] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0015] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0019] To achieve automated measurement of spinal curvature angles, a geometric derivation method based on vertebral key point detection is employed. Because its measurement process is logically similar to manual measurement by clinicians, the results are easily interpretable and readily understood and accepted by clinicians. However, its measurement accuracy heavily relies on the accuracy of the front-end key point detection; any minute deviation in coordinate prediction can be amplified in subsequent geometric calculations, leading to significant errors in the angle results and low accuracy. Methods based on overall spinal morphology analysis perform well on images with clear vertebral boundaries and complete contours. However, their accuracy is influenced by both segmentation and fitting algorithms. In complex situations with poor image quality, occlusion, or severe deformities, segmentation errors and fitting biases can compound, resulting in low stability. End-to-end direct regression avoids the propagation of multi-step errors but lacks interpretability. Indirect assessment methods based on body surface features suffer from low measurement accuracy and reliability. Therefore, current methods generally suffer from low measurement accuracy and low efficiency.
[0020] To address the aforementioned issues, this disclosure provides a method for measuring spinal curvature based on a large visual model. A feature extraction module is constructed based on the large visual model; an integrated progressive spinal measurement framework is built, comprising a feature extraction module and a progressive task decoder module. The feature extraction module is built based on the large visual model, and the progressive task decoder module utilizes multiple lightweight dedicated task heads. The target image is input into the integrated progressive spinal measurement framework. The feature extraction module extracts multi-scale high-resolution feature maps from the target image, and the progressive task decoder module performs spinal measurement tasks on the multi-scale high-resolution feature maps to obtain vertebral key point information. The vertebral key point information is then used to calculate the angle information of spinal curvature. By constructing a feature extraction module based on the large visual model, pixel-level detail support is provided for vertebral corner point localization, improving measurement accuracy and generalization ability. Meanwhile, a progressive task decoder module is built based on a lightweight dedicated task head. Combined with the feature extraction module and the progressive task decoder module, a highly integrated large model system is generated for fully automated spinal curvature detection. This improves the stability and reliability of the model. Furthermore, multiple task heads share the multi-scale, high-resolution, dense features output by the feature extraction module, avoiding computational redundancy of repeatedly extracting features for each task, simplifying clinical operations, and improving diagnostic efficiency.
[0021] The method will be described below with reference to specific embodiments.
[0022] Figure 1 This is a simplified flowchart illustrating a method for measuring spinal curvature according to an embodiment of this disclosure. The method can be executed by a spinal curvature measurement device based on a large visual model, wherein the device can be implemented in software and / or hardware, and is generally integrated into a computing device. Figure 1 As shown, the method includes: S101. Construct an integrated progressive spine measurement framework based on a large visual model.
[0023] The integrated progressive spine measurement framework based on a large visual model includes a connected feature extraction module and a progressive task decoder module.
[0024] The feature extraction module is constructed by adjusting the basic large-scale visual model. The large-scale visual model is an artificial intelligence model based on a deep learning architecture, using visual data such as images or videos as its core input, and possessing the ability to extract large-scale visual features and process complex tasks.
[0025] Large-scale models are models trained on massive amounts of data, typically with hundreds of millions or more parameters. As an advanced form of neural network model, their core advantage lies in leveraging the large number of parameters and massive pre-training data to achieve generalization capabilities, superior generalization performance, and broad task adaptability that are difficult for traditional neural networks to attain. Compared to traditional neural network models, large-scale models have the following advantages: Stronger cross-task generalization ability: Traditional neural network models are typically designed for specific tasks and trained for a single task, while large models are pre-trained on massive amounts of unlabeled data, enabling them to learn general underlying rules for language, vision, or multimodal learning. They can quickly adapt to multiple downstream tasks with only small-sample fine-tuning or even zero-sample hints, without having to repeatedly build the model for each task. Lower downstream task adaptation cost: Traditional neural network model training often requires a large amount of labeled data, resulting in high labeling costs and long cycles; while large models only require a small number of labeled samples to complete the adaptation, significantly reducing the dependence on labeled data and the overall development cost.
[0026] With superior ability to handle complex tasks: Large models, with their massive parameter scale and rich pre-trained knowledge, can capture deep semantic relationships in the data, thereby achieving more advanced and complex functions.
[0027] It has stronger robustness and fault tolerance: Due to the diversity and scale of training data, large models have stronger noise tolerance, can effectively handle non-standard inputs, and output stable and accurate results under interference. Its robustness is significantly better than that of traditional neural networks.
[0028] It has more flexible deployment and expansion potential: large models support a variety of optimization techniques such as model compression, quantization, and distillation, and can flexibly adjust the model size and computing resources according to different hardware environments, ensuring performance while taking into account efficiency and energy consumption.
[0029] In one possible implementation, when the large visual model includes a classification head, a feature pyramid transformer is used to replace the classification head of the large visual model to construct a feature extraction module. For example, the large visual model is a DINOv2 model.
[0030] The feature extraction module includes an image embedding layer (Patch Embedding), multiple encoders, and a feature pyramid transformer. The image embedding layer is used for image size segmentation, the multiple encoders are used for multi-level feature extraction, and the feature pyramid transformer is used for feature fusion output. The input of the feature extraction module is the target image, and the output is a multi-scale high-resolution feature map. The original classification head output of the large-scale visual model cannot meet the high-resolution detail requirements of dense prediction tasks. The feature pyramid transformer replaced in this application implements a cross-scale bidirectional feature interaction mechanism, which enhances the intelligence level of feature fusion.
[0031] The progressive task decoder module is built upon a lightweight, dedicated task header.
[0032] In one possible implementation, the spine measurement task includes posture recognition, spine region detection, and vertebral key point localization. The corresponding lightweight dedicated task head consists of a serially connected posture recognition head, spine region detection head, and vertebral key point localization head.
[0033] The body position recognition head is a global classifier used for body position classification, including frontal and lateral positions, which is the standing posture of the target object when the image is captured. The body position recognition head includes a Global Average Pooling Layer (GAP) and a Multi-Layer Perceptron (MLP). The GAP is used to extract global statistical features, and the MLP is used for body position classification. The input to the body position recognition head is the output of the feature extraction module, and the output is the body position classification result. The body position recognition head is lightweight and efficient, with few parameters and fast computation speed, quickly providing overall body position judgment and offering important spatial context guidance information for subsequent tasks.
[0034] The spine region detection head is used to detect regions of interest (ROIs) in the spine. It comprises a Region Proposal Network (RPN), which has two branches: bounding box regression and foreground / background classification. The bounding box regression branch delineates and crops the bounding boxes of the output spine region, while the foreground / background classification branch outputs confidence scores for the foreground or background. The inputs to the spine region detection head are the outputs of the feature extraction module and the posture recognition head; the output is the spine region. This head accurately detects the spine region from the global image, effectively eliminating interference from a large amount of irrelevant background (such as ribs and pelvis), providing a precise processing area for the next task.
[0035] The vertebral keypoint localization head is used to perform pixel-level precise calibration of keypoints within the detected spinal region. It is a heatmap-based dense prediction head. The vertebral keypoint localization head includes deconvolutional and convolutional layers. The deconvolutional layer is used for feature refinement and upsampling, while the convolutional layer is used for feature transformation and dimensionality reduction to obtain a heatmap. The peak positions in the heatmap are the precise coordinates of the keypoints. The input to the vertebral keypoint localization head is the output of the spinal region detection head, and the output is the vertebral keypoint information. The vertebral keypoint localization head can achieve sub-pixel-level detection accuracy and can simultaneously output the corner coordinates, confidence scores, and anatomical landmarks of all vertebrae.
[0036] The feature extraction module and the progressive task decoder module are combined to construct an integrated progressive spine measurement framework based on a large visual model.
[0037] S102. Input the target image into the integrated progressive spine measurement framework, use the feature extraction module to extract multi-scale high-resolution feature maps of the target image, and use the progressive task decoder module to perform a spine measurement task on the multi-scale high-resolution feature maps to obtain vertebral key point information.
[0038] This application utilizes a trained, integrated progressive spine measurement framework to detect key vertebral points in target images. The target images are digital radiography (DR) images of the vertebral bodies.
[0039] Please see Figure 2 , Figure 2 This is a schematic flowchart illustrating a method for measuring spinal curvature provided in an embodiment of the present disclosure.
[0040] like Figure 2 As shown, the integrated progressive spine measurement framework also includes an image preprocessing module. When the target image is input into the integrated progressive spine measurement framework, the image preprocessing module first preprocesses and standardizes the target image. Specifically, the image preprocessing module adjusts the contrast, size, and grayscale of the target image to ensure uniformity in size and grayscale, thereby eliminating image differences caused by different devices and shooting parameters. This provides standardized, high-quality input data for subsequent model processing, improving the model's stability and generalization ability.
[0041] Furthermore, the preprocessed target image is input into a feature extraction module based on a large visual model, which extracts multi-scale high-resolution feature maps of the target image.
[0042] In one possible implementation, the large visual model is the open-source visual foundation model DINOv2 (Distillation with No Labels version 2)ViT-B / 14 released by Meta AI.
[0043] The standard visual large-scale model DINOv2 ViT has an encoder-classifier head structure, where the encoder generates feature maps at different stages, and the classifier head outputs a single feature vector or a low-resolution feature map. This application removes the original classifier head from the DINOv2 ViT visual large-scale model and introduces a learnable Feature Pyramid Transformer (FPT). The core of the FPT is a lightweight Transformer decoder with a built-in cross-scale bidirectional attention mechanism to receive and fuse feature maps generated at different stages. A feature extraction module is constructed using the updated visual large-scale model. For details, please refer to [link to relevant documentation]. Figure 3 , Figure 3 This is a schematic diagram of a feature extraction module provided in an embodiment of the present disclosure.
[0044] like Figure 3 As shown, the target image is input into the image embedding layer, which performs size segmentation processing to divide the target image into multiple fixed-size target image patches. Then, linear projection is used to convert each target image patch into a token vector. Multiple target image patches are converted into target sequence data, realizing the transformation of two-dimensional image data into one-dimensional sequence data, which prepares for subsequent Transformer encoding processing.
[0045] Multiple ViT encoders extract features from the target sequence data, resulting in multi-level features. Specifically, multiple ViT encoders, such as ViT encoders Block1, 2, ..., N, each include a multi-head self-attention mechanism and a feed-forward neural network, enabling them to progressively capture global contextual information and high-level semantic features of the image, providing powerful initial representation capabilities. For example... Figure 3 As shown, the ViT encoder Block1 is used to extract features from the target image sequence data to obtain shallow features containing high-resolution details and mid-level features containing semantics and details. ViT encoder Block2 to ViT encoder BlockN are used to extract features from the target image sequence data multiple times to obtain deep features containing high semantics or low resolution.
[0046] The Feature Pyramid Transformer performs lightweight cross-scale feature fusion on the aforementioned multi-level features, outputting a multi-scale high-resolution feature map. Specifically, the Feature Pyramid Transformer includes a sampling layer, a feature concatenation layer, a lightweight Transformer decoder layer, and an output layer. The sampling layer unifies the resolution of the multi-level features. For example, an upsampling layer can be used to upsample the multi-level features, or a downsampling layer can be used to downsample the multi-level features, resulting in multiple features of the same scale. The feature concatenation layer concatenates multiple features of the same scale. The lightweight Transformer decoder of the Feature Pyramid Transformer includes a cross-scale bidirectional attention mechanism to decode the concatenated features. The output layer decomposes and outputs the feature map, producing a dense multi-scale high-resolution feature map.
[0047] In this implementation, the FPT of this application, through its built-in cross-scale bidirectional attention mechanism, enables global interaction and information complementarity between feature maps of different depths and scales. Deep semantic features guide shallow features to perform semantic enhancement, while shallow detail features supplement deep features with spatial localization information, achieving true and dynamic deep fusion. This solves the problem of relatively independent feature maps at different levels in the original ViT encoder, lacking explicit information exchange. Finally, the feature extraction module of this application outputs multi-scale high-resolution dense feature maps, solving the problem of single feature vectors or low-resolution feature maps output by the original network, while also providing an information foundation for downstream tasks. The output of FPT maintains high resolution in the spatial dimension, providing crucial detail support for subsequent pixel-level localization tasks; at the same time, it integrates all semantic information from low to high levels in the channel dimension, providing rich and unified contextual representations for various tasks such as classification, detection, and segmentation.
[0048] Furthermore, the multi-scale high-resolution feature maps are input into a progressive task decoder module, which performs a spine measurement task on the multi-scale high-resolution feature maps to obtain vertebral body key point information. For details, please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a schematic diagram of a progressive task decoder module provided in an embodiment of this disclosure.
[0049] like Figure 4 As shown, the multi-scale high-resolution feature map is first input into the body position recognition head. The body position recognition head performs body position classification on the multi-scale high-resolution feature map, determines whether the spine is captured in an anteroposterior or lateral position, and outputs the body position classification result.
[0050] Specifically, the global average pooling layer compresses the multi-scale high-resolution feature maps in the spatial dimension, extracting global statistical features; the multilayer perceptron maps the global statistical features to a probability distribution of body position categories, and outputs the body position classification result based on the probability distribution of body position categories. The body position classification result includes anteroposterior and lateral positions. In one possible implementation, the multilayer perceptron is a two-layer perceptron.
[0051] Furthermore, the multi-scale high-resolution feature maps and body position classification results are input into the spinal region detection head, which performs spinal detection and outputs the spinal region.
[0052] Specifically, the 3×3 convolutional layer performs convolution processing on the multi-scale high-resolution feature maps and body position classification results, the bounding box regression branch in the region proposal network detects the spine region box, and the foreground and background classification branch in the region proposal network detects the foreground and background confidence.
[0053] Furthermore, the spinal region is input into the vertebral body key point positioning head, which performs key point detection on the spinal region and outputs vertebral body key point information.
[0054] Specifically, two deconvolutional layers refine and upsample the features of the spinal region; a 1×1 convolutional layer performs feature transformation and dimensionality reduction to obtain a heatmap. The peak positions in the heatmap are the precise coordinates of the key points, and the output includes information such as vertebral corner coordinates, confidence scores, and anatomical identifiers. The vertebral corner coordinates include the precise coordinates of the four corner points of each vertebra. Vertebral identifiers are such as T1, T2, ..., L5. For example, MT represents the main thoracic segment, PT represents the upper thoracic segment, L represents the lumbar segment, and TL represents the thoracolumbar segment.
[0055] In this implementation, the three task heads operate in a strictly sequential manner, forming an efficient "coarse-to-fine" analysis chain. The body position recognition results provide prior knowledge for region detection, while the region detection results provide precise processing areas for keypoint localization. The network design in this application not only conforms to the logical sequence of clinical diagnosis but also reduces the complexity of the problem through staged processing.
[0056] S103. Calculate the angle information of spinal curvature using the key point information of the vertebral body.
[0057] Specifically, the angle information of spinal curvature is calculated by combining the key point information of the vertebral body and the position classification results of the position recognition head. Among them, the angle information of spinal curvature is the Cobb angle.
[0058] In one possible implementation, when the body position classification result of the body position recognition head indicates that the target image is an orthogonal image, the Cobb angle of the spinal curvature is calculated using a matrix analysis method based on vector computation.
[0059] First, the vertebral midline is generated. Specifically, the coordinates of the four corner points of each vertebra are extracted, and the average coordinates of the two points on its left and two points on its right are calculated respectively. The left average point is connected to the right average point to form a midline segment representing the direction of the vertebra, and the vector of the vertebral midline is calculated.
[0060] Next, the Cobb angle matrix is calculated. Specifically, the angle between the vectors of the midline segments of every two vertebrae is calculated using the cosine formula, resulting in a 17×17 dimensional Cobb angle matrix. This Cobb angle matrix can comprehensively record the potential angles between any two vertebrae.
[0061] Next, the type of spinal curvature and spinal segmentation are determined. Specifically, the maximum value in the matrix is identified as the target Cobb angle, and the overall curvature type of the spine, such as straightening or S-shape, is determined based on the location of the target Cobb angle.
[0062] Finally, based on the curvature type, the target Cobb angle, and the corresponding upper and lower vertebrae, the spine is divided into multiple physiological segments, and the Cobb angle of each segment is output for subsequent assessment of spinal curvature. These physiological segments include the main thoracic segment (MT), the upper thoracic segment (PT), and the thoracolumbar / lumbar segment (TL / L).
[0063] In one possible implementation, when the body position classification result of the body position recognition head indicates that the target image is a lateral image, the Cobb angle of the spinal curvature is calculated using a fixed physiological segmentation method.
[0064] For example, based on common anatomical knowledge, the spine is automatically divided into five consecutive vertebral segments according to preset rules, where T2-T5 is the upper thoracic segment, T2-T10 is the main thoracic segment, T5-T10 is the lower thoracic segment, T10-L2 is the thoracolumbar segment, and T12-S1 is the lumbar segment.
[0065] Then, the Cobb angle for each vertebral segment is calculated. Specifically, for each vertebral segment, the angle between the corresponding endplate lines of the upper and lower vertebrae of that segment is calculated.
[0066] Finally, the Cobb angle of each vertebral segment is output for subsequent assessment of the curvature of the spine.
[0067] In this implementation, a feature pyramid transformer is introduced on top of the DINOv2 ViT large-scale visual model. A lightweight Transformer decoder enables bidirectional interaction and fusion of multi-scale features, overcoming the bottleneck of traditional ViT in dense prediction. This generates multi-scale, high-resolution, dense feature maps, providing pixel-level detail support for vertebral corner localization and improving measurement accuracy and generalization ability. Simultaneously, a serial task head is designed from global to local, with upstream task results providing spatial priors for downstream tasks. This achieves seamless integration of the entire process from global posture recognition and spinal region detection to vertebral key point localization, replacing the cumbersome process of traditional multi-tool, manual operation, simplifying clinical procedures, and improving diagnostic efficiency. Combining the feature extraction module and the progressive task decoder module to generate a highly integrated large-scale model system for fully automated spinal curvature detection improves model stability and reliability. Furthermore, multiple task layers share the multi-scale, high-resolution, dense features output by the feature extraction module, avoiding computational redundancy from repeatedly extracting features for each task, simplifying clinical procedures, and improving diagnostic accuracy and efficiency.
[0068] Figure 5 This is a flowchart illustrating an integrated progressive spine measurement framework training method provided in an embodiment of this disclosure. This method can be executed by a spinal curvature measurement device based on a large visual model, wherein the device can be implemented in software and / or hardware, and is generally integrated into a computing device. Figure 5 As shown, the method includes: S501, a feature extraction module for training large visual models based on open-source pre-trained models.
[0069] The training module for feature extraction is transformed from recognizing natural images to recognizing medical images.
[0070] First, a training set for the feature extraction module is constructed. Specifically, digital radiography (DR) images from various manufacturers' full-spine digital X-ray imaging equipment are acquired, along with images from various imaging protocols (e.g., different kilovolt and milliampere-second imaging protocols) and images from various clinical cases (e.g., different ages, body types, and pathological degrees). Further, basic preprocessing, such as grayscale adjustment based on window width and window level and size normalization, is performed on these images to construct the training set for the feature extraction module. This method ensures the diversity and robustness of the data itself, and the images do not require manual annotation such as keypoints or segmentation masks, thus reducing the cost and barrier to data preparation.
[0071] Secondly, the training weights of an open-source pre-trained large-scale visual model are reused to determine the weights of the initial feature extraction module. The initial feature extraction module is then trained using a digital X-ray image dataset through self-supervised learning to obtain a fine-tuned feature extraction module.
[0072] Specifically, the feature extraction module is trained using a self-supervised learning framework. By constructing a student model and a teacher model, the student model learns to predict the feature representations generated by the teacher model for the same DR image from different enhancement perspectives, such as random cropping, brightness and contrast adjustments, and simulated noise. During the learning process, the model focuses on learning the inherent anatomical semantics and tissue texture features in the image that are unaffected by these random transformations. Anatomical semantics include spinal sequence, vertebral morphology, and pelvic contour, while tissue texture features include cortical bone, trabeculae, and soft tissue contrast. Finally, the trained feature extraction module is generated after deep learning of the DR image.
[0073] The above training process guides the transformation of the large visual model obtained from training on general natural images into a feature extraction module capable of understanding medical images, achieving a leap across professional knowledge domains. The trained feature extraction module possesses powerful and general DR image feature extraction capabilities, enabling a deep understanding of the imaging characteristics and anatomical structures of DR images, and providing a high-quality feature foundation for all downstream specific tasks.
[0074] S502, Load the fine-tuned parameters of the feature extraction module, initialize the parameters of the progressive task decoder module, and train the progressive task decoder module in the integrated progressive spine measurement framework.
[0075] The integrated progressive spine measurement framework focuses on specific clinical applications of measuring the Cobb angle of the spine.
[0076] First, a training set for the progressive task decoder module is constructed. Specifically, spinal X-ray images are acquired, and image-level positional classification labels, spinal region labels, vertebral key point labels, and anatomical markers are annotated in the spinal X-ray images to construct a digital spinal X-ray dataset. Furthermore, data augmentation techniques such as rotation, translation, and scaling are applied to the images in the dataset to improve the generalization ability of the trained model.
[0077] Next, the parameters of the feature extraction module after fine-tuning in step S501 are loaded, the parameters of all task heads in the progressive task decoder module are initialized, and the integrated progressive spine measurement framework is determined. The digitized spinal X-ray dataset is input into the integrated progressive spine measurement framework, and the progressive task decoder module is trained end-to-end using a loss function.
[0078] Specifically, a loss function for the progressive task decoder module is constructed. Sub-loss functions are designed for the body position recognition head, the spine region detection head, and the vertebral key point localization head. The three sub-loss functions are then weighted and summed to obtain the loss function for the progressive task decoder module. The loss function for the progressive task decoder module is as follows: for:
[0079] in, Let be the loss function for the body position recognition head, corresponding to the optimization objective of the body position recognition head. Body position recognition is a global classification task, and this application uses the standard cross-entropy loss function. Its calculation formula is:
[0080] In the formula, N is the number of samples in the training batch, and C is the total number of body position categories. For example, if the body position categories include frontal and lateral positions, then C=2. The true class label (one-hot encoded) for sample i. This is the probability predicted by the model that the sample belongs to class c. This loss function enables the progressive task decoder module to correctly predict the subject position in the image.
[0081] in, This is the sub-loss function for spine region detection in the spine region detection head, and this loss term corresponds to the optimization objective of the spine region detection head. The spine region detection head is based on a region proposal network structure, which includes two branches: bounding box regression and foreground / background classification. This application uses two parts of loss to constitute the sub-loss function for spine region detection. Its calculation formula is as follows:
[0082] In the formula, The foreground / background classification loss is used to determine whether the anchor box contains the target spine. The bounding box regression loss is used to fine-tune the position and size of the anchor boxes to more closely fit the actual spine region. For example, the bounding box regression loss is a Smooth L1 Loss. The spine region detection sub-loss function ensures that the progressive task decoder module can accurately locate the region of interest in the spine.
[0083] in, This is the vertebral keypoint localization sub-loss function for the vertebral keypoint localization head, and this loss term corresponds to the optimization objective of the vertebral keypoint localization head. The vertebral keypoint localization output is a heatmap representing the keypoint location. This application uses mean squared error loss to measure the pixel-level difference between the predicted heatmap and the actual heatmap. Its calculation formula is as follows:
[0084] In the formula, N is the number of samples in the training batch, K is the total number of key points, and is the sum of all vertebral corner points. and These are the heatmaps predicted by the model and the Gaussian distribution heatmaps generated centered on the actual keypoints, respectively. This loss is fundamental to achieving pixel-level precise localization and ultimately accurate calculation of the Cobb angle.
[0085] Furthermore, loss weight hyperparameters are set for the position classification sub-loss function, the spinal region detection sub-loss function, and the vertebral key point localization sub-loss function, respectively. , , This loss weight hyperparameter balances the losses from different tasks. By adjusting the contribution ratio of each sub-loss to the total loss, it addresses potential differences in the magnitude of losses from different tasks. By setting these weights appropriately, it ensures that the three sub-tasks converge in a balanced and stable manner.
[0086] This implementation employs a two-stage training paradigm: self-supervised fine-tuning of the feature extraction module and supervised training of the progressive task decoder module. During fine-tuning of the feature extraction module, the powerful generalization ability gained from large-scale pre-training can be reused, enabling stable adaptation to different devices, imaging qualities, and clinical images with artifacts or anatomical variations. This improves the model's stability and reliability in the face of unknown data, enhancing its clinical value. By utilizing unlabeled data to mine general features, the reliance on expensive and time-consuming manual annotation is greatly reduced, improving efficiency. Self-supervised learning allows the model to undergo self-supervised domain-adaptive pre-training, circumventing the bottleneck of scarce medical image annotation data and avoiding performance degradation caused by direct migration from the natural image domain to the medical image domain. It also provides stronger tolerance to image noise and artifacts, improving robustness. Furthermore, the "generalist first, specialist later" training process aligns with cognitive principles, making the model's learning path clear and highly interpretable. During the training of the progressive task decoder module, three types of annotation information—image-level body position labels, region-level spine labels, and pixel-level keypoint labels—were used to construct datasets. The weights of the fine-tuned feature extraction module were frozen, and the three task heads in the progressive task decoder module were jointly trained in an end-to-end manner. The different task heads formed a back-to-back feedback, realizing the collaborative optimization of the three tasks of body position classification, region detection, and keypoint localization. This ensured the accuracy and efficiency of the entire process from image-level classification to pixel-level localization and improved the generalization ability of the overall model.
[0087] Therefore, the spinal curvature measurement method based on a large visual model in this application utilizes the powerful feature extraction and contextual understanding capabilities of the large model to develop an intelligent algorithm that surpasses traditional dedicated models and is closer to the cognitive level of human doctors. Ultimately, it achieves a repeatable and highly interpretable angle measurement method, thereby reducing human error, improving diagnostic efficiency and reliability, and providing a more feasible technical solution for large-scale spinal health screening.
[0088] For example, please refer to Figure 6 , Figure 6 This is a schematic diagram of a digital X-ray image of a vertebra provided in an embodiment of this disclosure. Figure 6 As shown, Figure 6 (a) is a digital X-ray image of the vertebral body in an anteroposterior view. Figure 6 (b) is a lateral vertebral radiograph. Key point information was detected and the Cobb angle of each vertebral segment was calculated using the integrated progressive spinal measurement frame of this application on both images. Please refer to [link to relevant documentation]. Figure 7 , Figure 7 This is a schematic diagram illustrating the Cobb angle annotation of a vertebra as provided in an embodiment of this disclosure. Figure 7 As shown, Figure 7 (a) is Figure 6 (a) The Cobb angles of the vertebral bodies in the anteroposterior position are marked as follows: the Cobb angle of the T1-T5 vertebral segment is 26.74°, the Cobb angle of the T5-T11 vertebral segment is 43.59°, and the Cobb angle of the T11-L4 vertebral segment is 34.74°. Figure 7 (b) is Figure 6 (b) The Cobb angles of the lateral vertebral bodies are marked as follows: the Cobb angle of the T2-T5 vertebral segment is 18.6°, the Cobb angle of the T5-T12 vertebral segment is 29.95°, the Cobb angle of the T10-L2 vertebral segment is 1.11°, the Cobb angle of the T2-T12 vertebral segment is 45.04°, and the Cobb angle of the T12-S1 vertebral segment is 59.87°.
[0089] To achieve the above embodiments, this disclosure also proposes a spinal curvature measurement device based on a large visual model.
[0090] Figure 8 This is a schematic diagram of a spinal curvature measurement device based on a large visual model, provided as an embodiment of this disclosure. This device can be implemented by software and / or hardware, and is generally integrated into a computing device. Figure 8 As shown, the spinal curvature measurement device based on a large visual model includes: Module 801 is used to build an integrated progressive spine measurement framework. The integrated progressive spine measurement framework includes a feature extraction module and a progressive task decoder module. The feature extraction module is built based on a large visual model, and the progressive task decoder module is built using multiple lightweight dedicated task heads.
[0091] The detection module 802 is used to input the target image into the integrated progressive spine measurement framework, extract multi-scale high-resolution feature maps of the target image using the feature extraction module, and perform a spine measurement task on the multi-scale high-resolution feature maps using the progressive task decoder module to obtain vertebral key point information.
[0092] The calculation module 803 is used to calculate the angle information of spinal curvature using the key point information of the vertebral body.
[0093] In one possible implementation, the visual large model-based spinal curvature measurement device further includes: The first training module is used to construct a digital X-ray image dataset based on various full-spine digital X-ray imaging devices, various imaging protocols, and various clinical pathologies; the weights of the initial feature extraction module are determined using the weights of an open-source visual large model with pre-trained weights; and the initial feature extraction module is trained using the digital X-ray image dataset through self-supervised learning to obtain a fine-tuned feature extraction module.
[0094] In one possible implementation, the visual large model-based spinal curvature measurement device further includes: The second training module is used to design sub-loss functions for the posture recognition head, spine region detection head, and vertebral key point localization head, respectively. The three sub-loss functions are weighted and summed to obtain the loss function for the progressive task decoder module. A digital X-ray dataset of the spine labeled with posture classification, spine region, and vertebral key point labels is constructed. The parameters of the fine-tuned feature extraction module are loaded, the parameters of the progressive task decoder module are initialized, and the integrated progressive spine measurement framework is determined. The digital X-ray dataset of the spine is input into the integrated progressive spine measurement framework, and the progressive task decoder module is trained using the loss function.
[0095] In one possible implementation, the building module 801 includes: The building unit is used to replace the classification head of the large visual model with a feature pyramid transformer to build the feature extraction module.
[0096] In one possible implementation, the detection module 802 includes: The segmentation unit is used to segment the target image by size using the image embedding layer in the feature extraction module to obtain multiple target image blocks, and then uses linear projection to convert the multiple target image blocks into target sequence data.
[0097] The extraction unit is used to extract features from the target sequence data using multiple encoders in the feature extraction module to obtain multi-level features; each encoder includes a multi-head self-attention mechanism and a feedforward neural network.
[0098] The fusion unit is used to perform cross-scale feature fusion on multi-level features using the feature pyramid transformer in the feature extraction module, and output multi-scale high-resolution feature maps.
[0099] In one possible implementation, the fusion unit includes: The sampling subunit is used to upsample multi-level features using the upsampling layer of the feature pyramid transformer to obtain multiple features of the same scale; or to downsample multi-level features using the downsampling layer of the feature pyramid transformer to obtain multiple features of the same scale.
[0100] The splicing subunit is used to splice multiple features of the same scale using the feature splicing layer of the feature pyramid transformer.
[0101] The decoding subunit is used to decode the concatenated features using a lightweight Transformer decoder with a feature pyramid transform, and output a multi-scale high-resolution feature map; the lightweight Transformer decoder includes a cross-scale bidirectional attention mechanism.
[0102] In one possible implementation, the detection module 802 includes: The classification unit is used to classify body positions using a body position recognition head on multi-scale high-resolution feature maps and output the body position classification results.
[0103] The region detection unit is used to perform spine detection on multi-scale high-resolution feature maps and positional results using the spine region detection head, and outputs the spine region.
[0104] The key point detection unit is used to detect key points in the spinal region using the vertebral key point positioning head and output vertebral key point information.
[0105] In one possible implementation, the classification unit includes: The extraction sub-unit is used to compress multi-scale high-resolution feature maps in the spatial dimension using the global average pooling layer of the body position recognition head, and to extract global statistical features.
[0106] The classification subunit is used to map global statistical features to a probability distribution of body position categories using the multilayer perceptron of the body position recognition head, and outputs the body position classification result based on the probability distribution of body position categories; the body position classification result includes frontal and lateral positions.
[0107] In one possible implementation, the region detection unit includes: The region detection subunit is used to perform spine region detection on multi-scale high-resolution feature maps and positional results using the region proposal network of the spine region detection head, to obtain the spine region bounding box and the corresponding foreground and background confidence, and to output the spine region based on the spine region bounding box and the corresponding foreground and background confidence.
[0108] In one possible implementation, the key point detection unit includes: The key point detection subunit is used to refine and upsample the features of the spinal region using the deconvolution layer of the vertebral key point localization head, perform feature transformation and dimensionality reduction using the convolution layer of the vertebral key point localization head to obtain a heatmap, and output vertebral key point information based on the heatmap; wherein, the vertebral key point information includes the coordinates of the vertebral corner points and the confidence level.
[0109] The spinal curvature measurement device based on a large visual model provided in this disclosure can execute the spinal curvature measurement method based on a large visual model provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method.
[0110] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program / instructions that, when executed by a processor, implement the spinal curvature measurement method based on a large visual model as described above.
[0111] Figure 9 This is a schematic diagram of the structure of a computing device provided in an embodiment of the present disclosure.
[0112] The following is a detailed reference. Figure 9 The diagram illustrates a structural schematic suitable for implementing the computing device 900 in the embodiments of this disclosure. The computing device 900 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The computing device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0113] like Figure 9 As shown, the computing device 900 may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from memory 908 into random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the computing device 900. The processor 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0114] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows computing device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 A computing device 900 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0115] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a memory 908, or installed from a ROM 902. When the computer program is executed by the processor 901, it performs the functions defined in the visual large model-based spinal curvature measurement method of embodiments of this disclosure.
[0116] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0117] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0118] The aforementioned computer-readable medium may be included in the aforementioned computing device; or it may exist independently and not assembled into the computing device.
[0119] The aforementioned computer-readable medium carries one or more programs that, when executed by the computing device, cause the computing device to perform the aforementioned method for measuring spinal curvature based on a large visual model.
[0120] The computing device can be programmed with computer program code in one or more programming languages or a combination thereof to perform the operations of this disclosure. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0122] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0123] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0124] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0125] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0126] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0127] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for measuring spinal curvature based on a large visual model, characterized in that, The method includes: An integrated progressive spine measurement framework is constructed, comprising a feature extraction module and a progressive task decoder module. The feature extraction module uses a feature pyramid transformer to replace the classification head of the large visual model, and the progressive task decoder module is constructed using a serially connected position recognition head, a spine region detection head, and a vertebral key point localization head. The target image is input into the integrated progressive spine measurement framework. The feature extraction module extracts multi-scale high-resolution feature maps of the target image. The progressive task decoder module performs a spine measurement task on the multi-scale high-resolution feature maps to obtain vertebral key point information. The angle of spinal curvature is calculated using the key point information of the vertebral body. The step of extracting multi-scale high-resolution feature maps of the target image using the feature extraction module includes: The target image is segmented by size using the image embedding layer in the feature extraction module to obtain multiple target image blocks, and the multiple target image blocks are converted into target sequence data using linear projection. The target sequence data is subjected to feature extraction using multiple encoders in the feature extraction module to obtain multi-level features; wherein each encoder includes a multi-head self-attention mechanism and a feedforward neural network; The multi-level features are upsampled using the upsampling layer of the feature pyramid transformer to obtain multiple features of the same scale; or the multi-level features are downsampled using the downsampling layer of the feature pyramid transformer to obtain multiple features of the same scale. The feature pyramid transformer is used to stitch together the multiple features of the same scale using its feature stitching layer. The concatenated features are decoded using the lightweight Transformer decoder of the feature pyramid transform to output the multi-scale high-resolution feature map; wherein, the lightweight Transformer decoder includes a cross-scale bidirectional attention mechanism.
2. The method for measuring spinal curvature based on a large visual model according to claim 1, characterized in that, The construction of the integrated progressive spine measurement framework subsequently includes: A digital X-ray image dataset was constructed based on a variety of full-spine digital X-ray imaging devices, multiple imaging protocols, and various clinical pathologies. The weights of the initial feature extraction module are determined using an open-source visual large model with pre-trained weights; The initial feature extraction module is trained using the digital X-ray image dataset through self-supervised learning to obtain the fine-tuned feature extraction module.
3. The method for measuring spinal curvature based on a large visual model according to claim 2, characterized in that, The method further includes: Sub-loss functions are designed for the body position recognition head, the spine region detection head, and the vertebral key point localization head, respectively. The three sub-loss functions are weighted and summed to obtain the loss function of the progressive task decoder module. Construct a digital X-ray dataset of the spine labeled with body position classification labels, spinal region labels, and vertebral key point labels; Load the fine-tuned parameters of the feature extraction module, initialize the parameters of the progressive task decoder module, and determine the integrated progressive spine measurement framework; The digitized X-ray dataset of the spine is input into the integrated progressive spine measurement framework, and the progressive task decoder module is trained using the loss function.
4. The method for measuring spinal curvature based on a large visual model according to claim 1, characterized in that, The progressive task decoder module includes a position recognition head, a spine region detection head, and a vertebral key point localization head connected in series. The progressive task decoder module is used to perform a spine measurement task on the multi-scale high-resolution feature map to obtain vertebral key point information, including: The body position recognition head is used to classify the multi-scale high-resolution feature map into body positions, and the body position classification result is output. The spinal region detection head is used to detect the spine based on the multi-scale high-resolution feature map and the body position results, and the spinal region is output. The key point positioning head of the vertebral body is used to detect key points in the spinal region and output the key point information of the vertebral body.
5. The method for measuring spinal curvature based on a large visual model according to claim 4, characterized in that, The step of using the posture recognition head to classify the multi-scale high-resolution feature map for posture and outputting the posture classification result includes: The global average pooling layer of the body position recognition head is used to compress the multi-scale high-resolution feature map in the spatial dimension and extract global statistical features; The multilayer perceptron of the body position recognition head is used to map the global statistical features into a probability distribution of body position categories, and the body position classification result is output based on the probability distribution of the body position categories; the body position classification result includes frontal and lateral positions. The process of using the spinal region detection head to perform spinal detection on the multi-scale high-resolution feature map and the body position classification result, and outputting the spinal region, includes: The region proposal network of the spinal region detection head is used to detect the spinal region on the multi-scale high-resolution feature map and the body position classification result, so as to obtain the spinal region box and the corresponding foreground and background confidence. The spinal region is output based on the spinal region box and the corresponding foreground and background confidence. The step of using the vertebral body key point positioning head to detect key points in the spinal region and outputting the vertebral body key point information includes: The spinal region is refined and upsampled using the deconvolution layer of the vertebral key point localization head, and its features are transformed and dimensionality reduced using the convolution layer of the vertebral key point localization head to obtain a heatmap. The vertebral key point information is then output based on the heatmap. The vertebral key point information includes the coordinates of the vertebral corner points and the confidence level.
6. A spinal curvature measurement device based on a large visual model, characterized in that, The device includes: A construction module is used to build an integrated progressive spine measurement framework; wherein, the integrated progressive spine measurement framework includes a feature extraction module and a progressive task decoder module; the feature extraction module uses a feature pyramid transformer to replace the classification head of the large visual model, and the progressive task decoder module is constructed using a serially connected position recognition head, a spine region detection head, and a vertebral key point localization head; The detection module is used to input the target image into the integrated progressive spine measurement framework, extract multi-scale high-resolution feature maps of the target image using the feature extraction module, and perform a spine measurement task on the multi-scale high-resolution feature maps using the progressive task decoder module to obtain vertebral key point information. The extraction of multi-scale high-resolution feature maps of the target image using the feature extraction module includes: segmenting the target image by size using the image embedding layer in the feature extraction module to obtain multiple target image patches, and converting the multiple target image patches into target sequence data using linear projection; and extracting features from the target sequence data using multiple encoders in the feature extraction module to obtain... The system employs a multi-level feature map, wherein each encoder includes a multi-head self-attention mechanism and a feedforward neural network. The multi-level features are upsampled using the upsampling layer of the feature pyramid transformer to obtain multiple features at the same scale; or the multi-level features are downsampled using the downsampling layer of the feature pyramid transformer to obtain multiple features at the same scale; the multiple features at the same scale are concatenated using the feature concatenation layer of the feature pyramid transformer; and the concatenated features are decoded using the lightweight Transformer decoder of the feature pyramid transformer to output the multi-scale high-resolution feature map. The lightweight Transformer decoder includes a cross-scale bidirectional attention mechanism. The calculation module is used to calculate the angle information of spinal curvature using the key point information of the vertebral body.
7. A computing device, characterized in that, The computing device includes: a processor; a memory for storing executable instructions of the processor; the processor for reading the executable instructions from the memory and executing the instructions to implement the spinal curvature measurement method based on a large visual model as described in any one of claims 1 to 5.