Three-dimensional spine assessment method and system based on RGB images and anatomical constraints

Through the three-dimensional spine evaluation method based on RGB images and anatomical constraints, RGB images are captured using smart devices, combined with depth estimation and generation of adversarial networks, a high-precision three-dimensional spine model is generated, which solves the subjective error and radiation risk problems of traditional methods, and achieves non-invasive and low-cost scoliosis screening and monitoring.

CN120298415BActive Publication Date: 2025-08-19HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510793254.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-08-19
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

The existing scoliosis screening and monitoring methods have large subjective errors, high costs and radiation risks, making it difficult to achieve non-invasive and convenient large-scale screening and dynamic monitoring.

Method used

Using a three-dimensional spine evaluation method based on RGB images and anatomical constraints, RGB images are captured through intelligent devices, combined with lightweight segmentation networks, depth estimation models and generative adversarial networks, a high-precision three-dimensional spine model is generated, and biomechanical characteristics are evaluated.

Benefits of technology

It has achieved non-invasive, low-cost, high-precision three-dimensional reconstruction and health assessment of the spine, supporting extensive screening and dynamic monitoring in primary medical scenarios, and reducing screening thresholds and health risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298415B_ABST
    Figure CN120298415B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional spine assessment method and system based on RGB images and anatomical constraints, belonging to the field of medical image analysis. The method includes: pre-processing and anatomical symmetry correction of an upright back RGB image taken by a smart mobile device in a single shot; extracting regions of interest of scapula and iliac crest landmarks using a lightweight segmentation network; generating an initial depth map through a depth estimation model that integrates a vertebral attention mechanism and a multi-constraint loss function; calibrating the initial depth map to physiological scales based on height and weight parameters; using a cascaded GAN to eliminate clothing interference and reconstruct the three-dimensional surface of the back; adaptively adjusting NURBS parameters to fit the three-dimensional curve of the spine based on biomechanical characteristics; and outputting the coronal Cobb angle, trunk rotation angle, and sagittal curvature parameters. The present invention solves the radiation risk, subjective error, and clothing interference problems of traditional methods through monocular depth estimation and dynamic assessment based on anatomical constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image analysis and computer-aided diagnosis, and specifically relates to a three-dimensional spine assessment method and system based on RGB images and anatomical constraints, which is particularly suitable for the screening and quantitative assessment of spinal morphological abnormalities such as scoliosis and kyphosis. Background Art

[0002] Scoliosis is a common spinal deformity characterized by abnormal curvature of the spine in the coronal, sagittal, and transverse planes. Studies have shown that the prevalence of adolescent idiopathic scoliosis (AIS) is approximately 2-4%, and its incidence is increasing with increased poor posture, such as prolonged sitting and slouching. Failure to promptly screen and intervene can lead to worsening spinal deformity, impaired lung function, and even chronic pain, severely impacting adolescents' physical and mental health and quality of life.

[0003] Scoliosis is characterized by progressive development and may rapidly worsen during adolescence. Therefore, long-term monitoring and multiple measurements are required to accurately assess the trend of spinal changes. Current screening methods mainly rely on visual examination, spinal angle measuring instruments (such as spinal goniometers), three-dimensional reconstruction, X-ray examinations and other methods. However, manual measurement is subject to subjective errors, three-dimensional reconstruction equipment is expensive, and although X-ray examinations can provide high-precision assessments, they pose radiation risks and are not suitable for frequent monitoring. In addition, long-term follow-up costs are high. In the screening and follow-up management of scoliosis, there is an urgent need for a non-invasive, convenient, low-cost screening technology that can be easily promoted on a large scale to achieve early detection, regular monitoring and individualized health management, thereby effectively reducing the risk of disease progression and improving the health level of adolescents. Summary of the Invention

[0004] To address the above technical issues, this paper proposes a three-dimensional spinal assessment method and system based on RGB images and anatomical constraints. This method achieves high-precision three-dimensional reconstruction of the spine using only a single RGB image, and further assesses spinal morphology. This technology not only avoids the radiation risks associated with X-rays but also provides a simple and accurate assessment of spinal health without the need for expensive equipment, demonstrating its broad potential for application and widespread adoption.

[0005] The technical solutions that can achieve the purpose of the invention include:

[0006] A three-dimensional spine assessment method based on RGB images and anatomical constraints, the method comprising:

[0007] S1: The smart mobile device takes a single upright back RGB image as required and inputs the subject's height and weight parameters;

[0008] S2: Preprocessing and anatomical symmetry correction of the upright back RGB image taken by a smart mobile device in a single shot;

[0009] S3: Extracting the scapula and iliac crest landmark regions of interest from the preprocessed and rectified RGB images using a lightweight segmentation network;

[0010] S4: Based on the extracted region of interest image, an initial depth map is generated by integrating the vertebral attention mechanism and the depth estimation model with a multi-constraint loss function. The multi-constraint loss function includes the statistical shape model (SSM) projection constraint.

[0011] S5: Physiological scale calibration of the initial depth map based on height and weight parameters;

[0012] S6: Input the initial depth map after physiological scale calibration into the cascade generative adversarial network to eliminate clothing interference and obtain the back 3D surface reconstruction data;

[0013] S7: Based on the 3D surface reconstruction data of the back and combined with the biomechanical characteristics, the NURBS parameters are adaptively adjusted to fit the 3D curve of the spine;

[0014] S8: Output the coronal Cobb angle, trunk rotation angle and sagittal curvature parameters based on the three-dimensional curve of the spine.

[0015] A three-dimensional spine assessment system based on RGB images and anatomical constraints, comprising:

[0016] Image acquisition module: used to obtain standardized upright back RGB images of subjects through smart mobile devices and upload the image data to the cloud computing server for processing;

[0017] The image preprocessing and posture correction module is used to enhance and denoise the acquired upright back RGB images and adjust the shooting posture error based on anatomical landmarks to ensure standardized input of the spinal region;

[0018] The anatomical landmark detection module is used to automatically detect the anatomical landmarks of the scapula and iliac crest based on a deep learning model and extract the spinal region of interest (ROI) to provide a basis for subsequent modeling;

[0019] A monocular depth estimation module, which generates an initial depth map by integrating a vertebral attention mechanism with a depth estimation model that uses a multi-constraint loss function, including a statistical shape model (SSM) projection constraint.

[0020] The scale calibration module is used to perform scale correction on the initial depth map based on the subject's height, weight, and BMI physiological parameters to ensure that the spine measurement conforms to the individual's anatomical proportions;

[0021] Clothing interference removal and skin surface reconstruction module, which is used to remove clothing occlusion interference based on a generative adversarial network and restore the true three-dimensional shape of the back skin;

[0022] The spine 3D modeling module is used to generate an accurate 3D spine curve model based on the 3D morphological data of the back skin through curve fitting method, and can adapt to different degrees of spinal morphological changes;

[0023] An evaluation parameter calculation module is used to calculate and output the coronal Cobb angle, trunk rotation angle, and sagittal curvature parameters based on the generated three-dimensional spinal curve model;

[0024] The system output module is used to return the analysis results to the smart mobile device in a visual manner for display, and provide an automated diagnostic report containing spinal deformity assessment data for reference by medical professionals or individual users.

[0025] Beneficial effects:

[0026] The proposed method and system for 3D spine assessment based on RGB images and anatomical constraints offer the following advantages over existing technologies: Based on RGB depth estimation technology and statistical anatomical model constraints, a deep estimation network integrating statistical shape model (SSM) constraints with a vertebral attention mechanism enhances regional spinal anatomical perception. A dynamic calibration strategy driven by physiological parameters such as height and BMI addresses individual anatomical differences and errors in obese patients. A cascaded GAN technique, combined with clothing interference elimination and physiological surface reconstruction, improves the accuracy and fidelity of 3D back morphology reconstruction. Based on NURBS curve fitting and biomechanical property optimization, a high-precision 3D spine model is adaptively generated, enabling the joint assessment of multiple parameters such as the coronal Cobb angle and trunk rotation angle. Compared to traditional methods, this method requires no specialized equipment or radiological examinations. Using only a single back image captured by a smart device, it can perform non-invasive and low-cost 3D spine reconstruction and health assessment. This supports widespread screening and dynamic monitoring in primary care settings, significantly reducing screening barriers and health risks, and providing technical support for the early detection and intervention of spinal deformities. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 2. A schematic flow chart of a method for three-dimensional spine assessment based on RGB images and anatomical constraints according to an embodiment of the present invention;

[0028] Figure 2 2 is a module diagram of a three-dimensional spine assessment system based on RGB images and anatomical constraints according to an embodiment of the present invention.

[0029] Figure numerals: image acquisition module 210, image preprocessing and posture correction module 220, anatomical landmark detection module 230, monocular depth estimation module 240, scale calibration module 250, clothing interference elimination and skin surface re-evaluation module 260, spine three-dimensional evaluation module 270, evaluation parameter calculation module 280, system output module 290. DETAILED DESCRIPTION

[0030] In order to make those skilled in the art more clearly understand the technical solution of the present invention, the specific implementation of the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the following description is only a specific example of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention.

[0031] The present invention discloses a three-dimensional spine assessment method and system based on RGB images and anatomical constraints. A single back image taken by a smart device can complete non-invasive and low-cost three-dimensional spine reconstruction and health assessment.

[0032] Figure 1 This is a flow chart of a spine three-dimensional assessment method based on RGB images and anatomical constraints according to an embodiment of the present invention. Figure 1 The specific steps of the method are described. Among them, the specific application of the present invention is demonstrated by using a mobile phone and a WeChat applet as an example. Figure 1 As shown, the method includes the following steps:

[0033] S1: The smart mobile device captures a single upright back RGB image as required and inputs the subject's height and weight parameters. In an embodiment, the smart mobile device can be, for example, a smartphone or tablet. For example, the mobile device opens a mini-program on the phone, captures a single upright back RGB image as required, and inputs the subject's height and weight parameters.

[0034] Step S1 may specifically include:

[0035] 1) Enter the subject's height, weight, and personal identity information;

[0036] 2) The smart mobile device (such as a mobile phone camera) is placed on a horizontal plane based on the height of the subject's iliac crest, and the camera is taken in a direction perpendicular to the plane of the human back at the midpoint of the iliac crest line. The shooting distance is satisfy:

[0037] ,

[0038] in: is the height of the subject, in meters.

[0039] The subject should place their feet together with their toes touching the reference surface, with the heels no more than 5 cm apart. Their upper limbs should hang naturally, with their palms touching the outside of their thighs and their thumbs facing forward. The subject should wear tight clothing (such as sports tights) to ensure that the spine area is not significantly obstructed, facilitating the subsequent 3D modeling process. Tight clothing should fit the body closely, avoiding being too loose or too thick, to ensure a clear visual outline of the spine.

[0040] 3) Before shooting, use the built-in gyroscope and accelerometer data of the smart mobile device (mobile phone) to detect the pitch and roll angles of the smart mobile device (mobile phone). By obtaining the angle information of the smart mobile device in real time, ensure that the smart mobile device is kept vertical for shooting. If the pitch or roll angle of the smart mobile device exceeds the preset range , the system will issue a warning prompt to remind users to adjust the phone posture to ensure that the angle of the smart mobile device is correct when shooting.

[0041] S2: Preprocess and perform anatomical symmetry correction on the upright back RGB image captured by a smart mobile device, including:

[0042] Preprocessing includes: image enhancement and denoising preprocessing of the upright back RGB images using bilateral filtering, adaptive histogram equalization (CLAHE) and frequency domain Wiener filtering to improve the clarity and contrast of the back anatomical features.

[0043] Anatomical symmetry correction specifically includes:

[0044] 1) Using the improved high-resolution network (HRNet-W32) to detect key anatomical feature points on the back of the preprocessed RGB image, including the lower corner of the scapula ( ), and the upper edge of the left and right iliac crests , obtain the two-dimensional image coordinates of the above feature points;

[0045] 2) Based on the two-dimensional coordinates of the upper edge of the iliac crest, define and calculate the coronal symmetry axis of the human back , the specific mathematical expression is:

[0046] ,

[0047] in: , Represents the two-dimensional pixel coordinate components in the image plane, corresponding to the horizontal and vertical directions of the image, the upper edge point of the left iliac crest , upper edge of the right iliac crest , , , .

[0048] 3) Calculate the angle between the coronal symmetry axis of the human back and the vertical direction of the image , defined as:

[0049] ,

[0050] When the angle When the preset threshold is not exceeded , then no rotation correction is performed and the subsequent depth estimation and three-dimensional evaluation phase is directly entered. When the preset threshold is exceeded , by constructing and performing the rotational affine transformation matrix Perform posture correction on RGB images:

[0051] ,

[0052] in: is the rotational affine transformation matrix, , The image center is set to ensure that the iliac crest reference line is parallel to the horizontal axis of the image after correction, thereby eliminating the global tilt error caused by perspective distortion caused by shooting angle, distance and posture.

[0053] S3: Extract the scapula and iliac crest landmark regions of interest from the preprocessed and rectified RGB images using a lightweight segmentation network, including:

[0054] 1) Based on the standardized back image corrected in step S2, an improved lightweight U-Net++ network is used, combining an encoder, a decoder, and a multi-level output layer to achieve accurate segmentation of regions of interest (ROIs) of landmarks such as the scapula, iliac crest, and spinous processes;

[0055] 2) The encoder of the lightweight U-Net++ network uses the MobileOne-S3 module to enhance local and global feature extraction capabilities and optimize computational efficiency. Its block structure is as follows:

[0056] Branch 1: 3×3 depthwise separable convolution → batch normalization → SiLU activation;

[0057] Branch 2: 1×1 point-wise convolution → batch normalization → SiLU activation → 3×3 depth-wise separable convolution.

[0058] The above two branch feature maps are fused element by element and then output as a coded feature map.

[0059] The convolution kernel parameters and batch normalization layer parameters of the module are automatically learned and determined by the back propagation algorithm during the network training process; the number of convolution kernels (number of channels) in the convolution module is a hyperparameter of the network structure, usually in the range of 32 to 128. The optimal combination is determined by cross-validation experiments. In this embodiment, 64 is taken as the optimal solution.

[0060] 3) Decoder of lightweight U-Net++ network fused with dynamic upsampling module ( ), restore the image resolution through learnable parameter interpolation to enhance the model's adaptability to various image inputs and ensure the preservation of detail information during the segmentation process. The specific calculation method is:

[0061] ,

[0062] in: Represents the feature map after upsampling; is the input feature map; For the learnable weights of the interpolation branches; For the The upsampling result of the nearest neighbor interpolation calculation; is the number of interpolation branches, determined by experimental cross-validation (recommended value =2).

[0063] 4) The output layer of the lightweight U-Net++ network predicts multiple ROI masks in parallel, including: Represents the scapular spine region; represents the posterior superior iliac spine region; Represents the continuous line of the spinous process of the spine.

[0064] 5) The lightweight U-Net++ network uses the Dice loss function to perform pixel-level optimization on the scapula and iliac crest ROIs to improve segmentation accuracy of anatomical landmarks. The application of the Dice loss function ensures high accuracy and consistency of the segmentation results, thereby improving the accuracy of subsequent processing stages. The Dice loss function is defined as follows:

[0065] ,

[0066] in: and are the height and width of the segmented image respectively; Indicates the corresponding position The true labeled value of Indicates that the network predicts at pixel location ε is the smoothing coefficient, which is 1×10 -6 .

[0067] 6) The lightweight U-Net++ network shares the backbone feature extraction layer with the key anatomical feature point detection network HRNet-W32 used for iliac crest baseline correction described in step S2 to achieve a unified anatomical structure understanding and ensure that the segmentation results are aligned with the globally corrected human posture, thereby improving the accuracy of subsequent depth estimation and 3D assessment.

[0068] S4: Based on the extracted ROI image, an initial depth map is generated by integrating the vertebral attention mechanism and the depth estimation model with a multi-constraint loss function. The multi-constraint loss function includes the statistical shape model (SSM) projection constraint, including:

[0069] 1) The ROI mask output from step S3 (including the scapula, iliac crest, and spinous process) and the RGB image corrected from step S2 are input into a dual-branch depth estimation network structure.

[0070] 2) The dual-branch depth estimation network estimates depth information from two perspectives: appearance features and anatomical constraints, achieving an organic fusion of anatomical constraints and appearance features;

[0071] Branch 1: (Appearance feature extraction): retain the MiDaS encoder-decoder structure and output relative depth map ; and embed the frequency domain de-artifacting module (FDM) in the 4th layer of the decoder:

[0072] ,

[0073] in: A dynamically generated high-frequency noise mask is used to selectively filter frequency domain artifacts in RGB images. It is element-by-element multiplication, shielding high frequencies, It is a frequency domain artifact removal module used to suppress high-frequency noise artifacts in RGB images. is the convolutional layer, For the inverse discrete cosine transform, restore from the frequency domain to the spatial domain, is the discrete cosine transform, which converts the image from the spatial domain to the frequency domain. This is the feature map output by the 4th layer of the decoder of branch 1. The convolution kernel parameters are automatically determined during network training.

[0074] This branch extracts deep features through the global information of the image.

[0075] Branch 2: (Anatomical Constraint Fusion): A vertebral attention mechanism is introduced after the fourth decoder layer in branch 1 to enhance the depth prediction accuracy of the vertebral region. By increasing the depth prediction weights of vertebral regions such as the scapula, iliac crest, and spinous process, the network is more sensitive to key spinal regions and improves the accuracy of depth estimation.

[0076] Branch 1 and branch 2 are fused through channel splicing:

[0077] ,

[0078] in: is the feature map output by branch 1, is the feature map output by branch 2, For channel splicing operation, is the channel number 1×1 convolution, To fuse features, the initial depth map is obtained by optimizing the multi-constraint loss function .

[0079] 3) Vertebrae attention mechanism: Enhances the depth prediction weight of the vertebrae region to ensure that the network is more sensitive to learning about the vertebrae region and improves the depth estimation accuracy.

[0080] enter:

[0081] ,

[0082] in: are the height, width and number of channels of the feature map respectively. It is the feature map output by the 4th layer of the decoder of branch 1.

[0083] Anatomical position encoding: The positions of vertebral regions such as the scapula, iliac crest, and spinous process are used to generate anatomical position encoding, which is input into the vertebral attention mechanism as an anatomical prior to further enhance the network's attention to the spinal region.

[0084] Vertebral attention mechanism generates vertebral probability distribution map :

[0085] , ,

[0086] in: represents the pixel position in the feature map, Corresponding to the width direction, Corresponding to height direction; Indicates the vertebral index; For the Standard projection position of each vertebra; is the vertebral size correlation parameter, satisfying , is the estimated value of vertebral height; To segment the confidence function, control the weight attenuation of the low confidence cone, It means the Confidence of each vertebra;

[0087] Attention calculation of vertebral attention mechanism:

[0088] , , ,

[0089] ,

[0090] in: , , is the linear transformation matrix; is the channel number normalization factor to avoid the gradient vanishing or exploding problem. is the feature map output by the 4th layer of the decoder of branch 1, is the vertebral probability distribution map generated by the vertebral attention mechanism, It is a normalized exponential function that converts attention weights into probability distribution.

[0091] Output: Generate enhanced feature map:

[0092] ,

[0093] in, Normalize the layer.

[0094] 3) The multi-constraint loss function

[0095] ,

[0096] in:

[0097] : Depth estimation error loss, used to minimize the difference between predicted depth and true depth;

[0098] : Statistical Shape Model (SSM) projection loss ensures that depth estimation conforms to the cone statistical model;

[0099] : Loss of left and right hemisymmetry, constraining the symmetry of the spinal morphology and improving deep stability;

[0100] : Key point loss, used to optimize the positions of key points such as scapula, iliac crest, and spinous process;

[0101] : is the weight coefficient, usually set to λ∈(0,1). The specific value within the range can be tuned according to experiments to optimize the depth estimation accuracy.

[0102] 4) Depth estimation error loss : Using scale-invariant logarithmic error:

[0103] ,

[0104] in: , are the height and width of the image, the total number of pixels : The real depth map in pixels The depth value at To predict the depth map at pixel The depth value at Represents two different sets of pixel coordinates in an image, used to traverse all possible pixel pairs. is a balancing factor to optimize the local depth relationship constraint.

[0105] 5) SSM projection loss : Force the depth map to satisfy the projection constraints of the pyramidal statistical shape model:

[0106] ,

[0107] in: is the total number of vertebral types (usually 17 major vertebrae); The first in the statistical shape model (SSM) The three-dimensional coordinates of the vertebra; It is a perspective projection operation; The network prediction Vertebral segmentation results. is the absolute depth map.

[0108] 6) Loss of left-right hemisymmetry : Penalize left-right half depth asymmetry:

[0109] ,

[0110] in: is the depth value of the pixels in the left half of the image; is the depth value of its mirror-symmetrical position; , are the height and width of the image respectively; is the threshold for symmetry loss; is a weighting factor based on different regions of the spine, used to dynamically adjust the symmetry loss.

[0111] 7) Keypoint Loss :

[0112] ,

[0113] in: is the number of key points (such as scapula, iliac crest, spinous process, etc.); The coordinates of the actual key points such as scapula, iliac crest, and spinous process; : Coordinates of keypoints predicted by the network.

[0114] S5: Perform physiological scale calibration on the initial depth map based on height and weight parameters, including:

[0115] 1) Height-vertebral size mapping

[0116] Based on the statistical model, using the height of the subjects Predict vertebral height , establish the regression relationship between vertebral height and body height:

[0117] ,

[0118] in: is the predicted vertebral height; is the height of the subject; , is the regression coefficient, which is set according to statistical data.

[0119] 2) Adaptive height scale calibration calculation

[0120] To calibrate the depth estimation scale, calculate the height scale calibration factor :

[0121] , vertebral weight ,

[0122] Calculate the height-calibrated depth map based on the height-calibration factor :

[0123] ,

[0124] in: For the The vertical coordinate of the vertebra in the upright back RGB image, the midline of the spine is , is the vertebral size correlation parameter, Vertebral height predicted based on regression model; Standard vertebral height in the statistical shape model SSM; The default value is 5, and 17 is used if the image is complete; is the initial depth map processed in step S4;

[0125] 3) BMI compensation mechanism

[0126] To correct the depth estimation error in obese patients (BMI ≥ 25), fat compensation is added to the depth after height scale calibration:

[0127] ,

[0128] in: It is a depth map calibrated with height and weight scales; The depth map is calibrated for height scale; is the fat compensation coefficient; is the vertical coordinate of the pixel point in the upright back RGB image; is the mean vertical coordinate of the iliac crest ROI (iliac crest is used as the waist reference); is the Gaussian distribution parameter that controls the fat compensation range.

[0129] S6: Input the physiological scale calibrated depth map into the cascaded generative adversarial network to eliminate clothing interference and obtain the back 3D surface reconstruction data:

[0130] The depth map calibrated with height and weight scale in step S5 is input into the two-stage generator in the cascaded generative adversarial network architecture. , , achieving the elimination of clothing interference and three-dimensional surface reconstruction of the back, including:

[0131] Phase 1 (texture generation): Building a generator based on partial convolution , fill in the area blocked by clothing;

[0132] Phase II (Physiological Protuberance Synthesis): Based on anatomical priors and biomechanical modeling, the skin surface morphology is modified.

[0133] 1) Phase 1: Eliminate clothing interference

[0134] Generator :

[0135] ,

[0136] in, (Encoder layer): Input image feature extraction module, used to identify clothing boundaries and occluded areas; (Decoder layer): Recover the occluded skin texture information based on the encoder output.

[0137] Generator The calculations used in include:

[0138] Deformable Convolution:

[0139] ,

[0140] in represents the pixel position in the feature map, is the convolution kernel size, suitable for local feature extraction; is the weight of the convolution kernel; is the pixel position; is the offset of the convolution kernel; The convolution kernel offset variation allows the convolution kernel to be adaptively adjusted to improve feature learning capabilities; It is a modulation parameter that controls the weight influence of different areas.

[0141] Attention-gated fusion:

[0142] ,

[0143] ,

[0144] in: Calculates the weight function for the attention mechanism; balance factor Used to balance feature information at different layers, It is to upsample the previous decoder output.

[0145] 2) Second stage: 3D back surface reconstruction

[0146] Costal arch modeling: simulates the human costal arch curve to improve the physiological rationality of skin reconstruction.

[0147] ,

[0148] in: , is the regression coefficient, parameter is the subject's height (cm), range [150, 200]; The body mass index range is [15, 35]. is the output of the costal arch model.

[0149] Erector Spinae Stress Simulation: This step estimates changes in skin surface morphology based on a muscle stress model. This step simulates the effects of muscle activity and stress distribution on skin surface morphology. This step combines muscle elasticity and stress distribution models to modify the skin surface to conform to physiological structure.

[0150] ,

[0151] in: stress that indicates changes in the skin's surface; represents the base force when there is no muscle contraction; is the muscle elasticity parameter, which indicates the elasticity of the muscle; is the muscle stress elasticity parameter, which represents the elastic behavior of the muscle under stress; is the muscle fiber activation function.

[0152] 3) Multi-scale identification mechanism

[0153] The multi-scale discriminator improves the realism of the generated skin and ensures that it matches the anatomical curvature of the real skin. The structure includes:

[0154] Global Discriminator : Input 512×512 image and evaluate the overall authenticity;

[0155] Local Discriminator : Identify details such as clothing seams in a 128×128 local area;

[0156] Physiological plausibility discriminator : Calculate the KL divergence between the generated skin curvature and the real anatomical database to evaluate the physiological rationality of the generated image:

[0157] ,

[0158] in: , are the curvature distributions of the generated image and the reference database, respectively. is the physiological plausibility loss function, which represents the difference between the generated image and the reference data.

[0159] S7: Based on the 3D surface reconstruction data of the back and combined with biomechanical characteristics, the NURBS parameters are adaptively adjusted to fit the 3D curve of the spine, including:

[0160] 1) NURBS-based spine curve fitting

[0161] Based on the back three-dimensional surface reconstruction data in step S6, a non-uniform rational B-spline curve (NURBS) is constructed to fit the center line of the human spine, thereby improving the accuracy and smoothness of the three-dimensional reconstruction.

[0162] ,

[0163] in:

[0164] : The spinal curve in parameters The location, Represents the normalization parameter, ranging from .

[0165] : No. The pixel positions of the control points, is the total number of control points, representing key locations of the spine (scapula, iliac crest, spinous process, etc.).

[0166] : B-spline basis function, is the control point index, is the curve order, commonly =3, determines the local smoothness and fitting ability.

[0167] : Control point weights, used to adjust local curvature to ensure a more reasonable fitting of the spine shape.

[0168] 2) Parameter adjustment strategy

[0169] Optimizing NURBS node vectors , to adapt to the individual spinal anatomy and ensure that the curve is more accurate in key positions (such as cervical vertebra C7, thoracic vertebrae T1-T12, and lumbar vertebrae L1-L5).

[0170] ,

[0171] in:

[0172] : After optimization The location of the node.

[0173] : No. Known vertebral position points (obtained by 3D modeling).

[0174] : Euclidean distance between vertebrae.

[0175] : The total distance between all vertebrae is normalized.

[0176] This optimization method adaptively adjusts the node density to provide a more accurate curve description in high-curvature areas of the spine (such as the thoracic curvature area) while reducing redundant calculations in low-curvature areas (such as the lumbar spine).

[0177] 3) Weight distribution strategy

[0178] The influence of control points is dynamically adjusted according to the spinal morphology to improve fitting accuracy and ensure that the curve accurately matches the actual anatomical structure in key anatomical areas (such as high curvature points of scoliosis).

[0179] ,

[0180] in:

[0181] : No. The weights of the control points are adjusted to ensure a more accurate fit in areas with high spinal curvature.

[0182] : No. The local curvature value of each control point.

[0183] : The mean curvature of all control points (used for normalization).

[0184] : Standard deviation of the curvature of all control points (controls overall smoothness).

[0185] When the spine has high curvature (such as in patients with scoliosis), the strategy automatically increases the weight to ensure that the fitting curve is more realistic. When the spine is straight, the weight fluctuation is reduced to ensure a smooth transition of the curve.

[0186] 4) Dynamic Constraint Optimization

[0187] The curve smoothness, anatomical consistency and data fitting are optimized simultaneously, so that the fitting curve conforms to the individual physiological structure while minimizing the error.

[0188] ,

[0189] in:

[0190] (Curve Smoothness): Avoid excessive bends and ensure that they conform to the natural curve of the spine.

[0191] (Anatomical Prior Constraint): Ensures that the curve conforms to the true spinal anatomy.

[0192] (Data fitting error): Ensures that the curve fits the spine point cloud data obtained from the depth estimation.

[0193] , , is the weight coefficient, usually set at ∈(0, 1).

[0194] is the loss function, To minimize the loss function, are the shape control points of the spine curve, is the weight of the control point.

[0195] This optimization strategy can adapt to the spinal morphology of different individuals and provide reasonable three-dimensional reconstruction in both normal and spinal abnormality populations.

[0196] S8: Outputs coronal Cobb angle, trunk rotation angle, and sagittal curvature parameters based on the 3D spine curve, including:

[0197] 1) Cobb angle assessment (based on the spinal curve fitting results, the end vertebrae are automatically detected and the Cobb angle is calculated to assess the severity of scoliosis).

[0198] End vertebra identification: extracting the extreme curvature points along the spinal curve , as the landmark points of the upper and lower vertebrae.

[0199] , ,

[0200] in: : The point of maximum curvature on the spinal curve (end vertebra position). : Coronal spinal curvature function, indicating the curve in The degree of local curvature.

[0201] Plane Fitting: Based on the 3D point cloud data near the end vertebra, the maximum tilt angle is calculated to determine the actual angle of the end vertebra. Principal Component Analysis (PCA) is performed on the point cloud data to extract the main directions.

[0202] ,

[0203] ,

[0204] in: is the point cloud data near the upper vertebra, is the point cloud data near the lower vertebra, is the point cloud range offset near the end vertebra, : The direction vector of the end plate of the upper and lower vertebrae.

[0205] Cobb angle calculation: Calculate the angle between the upper and lower vertebral endplate lines.

[0206] ,

[0207] in:

[0208] : Normalized modulus length in the endplate direction.

[0209] 2) Angle of trunk rotation (ATR) assessment (based on the point cloud of the back surface, the maximum height difference between the left and right sides of the spine is calculated to assess the torsional deformity caused by scoliosis).

[0210] Back surface point cloud reconstruction: Based on depth estimation and 3D reconstruction, the spine midline point and its corresponding left and right highest points on the back are extracted.

[0211] Select the iliac crest point (PSIS) as the benchmark and fit the horizontal reference plane:

[0212] ,

[0213] in: 、 、 are the components of the plane normal vector, is the distance between the plane and the origin, 、 、 are the coordinates of each point in the point cloud.

[0214] exist In the horizontal reference plane, scan the highest points on both sides of the spine:

[0215] ,

[0216] in: The height of the highest point on the left and right back surfaces.

[0217] ATR calculation: Calculate the maximum left-right height difference of the back protrusion and convert it into an angle.

[0218] ,

[0219] in: : Maximum left-right height difference; : Horizontal projection spacing.

[0220] 3) Sagittal curvature assessment (calculation of the curvature of the thoracic and lumbar spine, and assessment of lordosis and kyphosis).

[0221] The overall curvature of different segments of the spine is calculated based on the arc length parameterized curvature.

[0222] ,

[0223] in: : thoracic kyphosis angle; : Lumbar lordosis angle; Sagittal plane The arc length parameterizes the curvature. - It is from the first to the twelfth thoracic vertebrae. It is the first lumbar vertebra. It is the first sacral vertebra.

[0224] like Figure 2 As shown, an embodiment of the present invention further provides a three-dimensional spine assessment system based on RGB images and anatomical constraints, specifically comprising:

[0225] Image acquisition module 210: used to obtain a standardized upright back RGB image of the subject through a smart mobile device (such as a smartphone or tablet), and upload the image data to a cloud computing server for processing;

[0226] Image preprocessing and posture correction module 220, for enhancing and denoising the acquired upright back RGB image, and adjusting the shooting posture error based on anatomical landmarks to ensure standardized input of the spinal region;

[0227] The anatomical landmark detection module 230 is used to automatically detect anatomical landmarks such as the scapula and iliac crest based on a deep learning model and extract the spinal region of interest (ROI) to provide a basis for subsequent evaluation;

[0228] Monocular depth estimation module 240, used to generate an initial depth map by integrating a vertebral attention mechanism and a multi-constraint loss function (including a statistical shape model SSM projection constraint) into a depth estimation model;

[0229] a scale calibration module 250 for performing scale correction on the initial depth map based on physiological parameters such as the subject's height, weight, and BMI to ensure that the spinal measurements conform to the individual's anatomical proportions;

[0230] Clothing interference removal and skin surface re-evaluation module 260, for removing clothing occlusion interference based on a generative adversarial network (GAN) and restoring the true three-dimensional morphology of the back skin;

[0231] The spine 3D assessment module 270 is used to generate an accurate spine 3D curve model based on the back skin 3D morphological data through a curve fitting method, and adapt to different degrees of spine morphological changes;

[0232] An evaluation parameter calculation module 280 is used to calculate and output spinal physiological curve parameters such as the coronal Cobb angle, the trunk rotation angle (ATR), and the sagittal curvature based on the generated three-dimensional spinal curve model;

[0233] The system output module 290 is used to return the analysis results to the smart mobile device in a visual manner for display, and provide an automated diagnostic report containing spinal deformity assessment data for reference by medical professionals or individual users.

[0234] The above describes the basic principles, main features, and advantages of the present invention. It should be understood that the embodiments of the present invention are only intended to illustrate the technical solutions of the present invention and are not intended to limit the present invention in any way. Any technical solutions obtained by equivalent substitution or equivalent transformation fall within the scope of protection of the present invention.

Claims

1. A three-dimensional spine assessment method based on RGB images and anatomical constraints, characterized by: The method comprises: S1: The smart mobile device takes a single upright back RGB image as required and inputs the subject's height and weight parameters; S2: Preprocessing and anatomical symmetry correction of the upright back RGB image taken by a smart mobile device in a single shot; S3: Extracting the scapula and iliac crest landmark regions of interest from the preprocessed and rectified RGB images using a lightweight segmentation network; S4: Based on the extracted region of interest image, an initial depth map is generated by integrating the vertebral attention mechanism and the depth estimation model with a multi-constraint loss function. The multi-constraint loss function includes the statistical shape model (SSM) projection constraint. S5: Physiological scale calibration of the initial depth map based on height and weight parameters; S6: Input the physiological scale calibrated depth map into the cascade generative adversarial network to eliminate clothing interference and obtain the back 3D surface reconstruction data; S7: Based on the 3D surface reconstruction data of the back and combined with the biomechanical characteristics, the NURBS parameters are adaptively adjusted to fit the 3D curve of the spine; S8: Output the coronal Cobb angle, trunk rotation angle and sagittal curvature parameters based on the three-dimensional curve of the spine.

2. The method for three-dimensional spine assessment based on RGB images and anatomical constraints according to claim 1, characterized in that: Step S1 includes: 1) Enter the subject's height, weight, and personal identity information; 2) The camera of the smart mobile device is located on a horizontal plane based on the height of the subject's iliac crest. The camera is taken along the direction perpendicular to the plane of the human back at the midpoint of the iliac crest line. The shooting distance is satisfy: , in: is the subject's height, in meters; 3) Before shooting, the device uses the built-in gyroscope and accelerometer data to detect the pitch and roll angles of the smart mobile device. By obtaining the angle information of the smart mobile device in real time, the device is ensured to maintain a vertical shooting position. If the pitch or roll angle of the smart mobile device exceeds the preset range, an early warning prompt is issued.

3. The method for three-dimensional spine assessment based on RGB images and anatomical constraints according to claim 1, characterized in that: In step S2, The preprocessing includes: performing image enhancement and denoising preprocessing on the upright RGB back image taken based on step S1 by using bilateral filtering, adaptive histogram equalization and frequency domain Wiener filtering; Anatomical symmetry correction includes: The pre-processed upright back RGB image was processed using an improved high-resolution network to detect key anatomical feature points on the back, including the inferior angle of the scapula, the upper edge of the left iliac crest, and the upper edge of the right iliac crest, and obtain the two-dimensional image coordinates of these feature points. Based on the two-dimensional coordinates of the upper edge of the iliac crest, the coronal symmetry axis of the human back is defined and calculated; Calculate the angle between the symmetry axis and the vertical direction of the upright back RGB image If the angle is less than or equal to the preset threshold, no rotation correction is performed and the subsequent depth estimation and three-dimensional evaluation stages are directly entered. If the angle exceeds the preset threshold, posture correction is performed by rotating the affine transformation matrix to adjust the posture in the upright back RGB image so that the iliac crest baseline is parallel to the horizontal axis of the image.

4. The method for three-dimensional spine assessment based on RGB images and anatomical constraints according to claim 1, characterized in that: In step S3, the lightweight segmentation network is used to extract the ROI of the scapula and iliac crest landmarks, including: 1) Based on the RGB image corrected in step S2, an improved lightweight U-Net++ network is used, combining an encoder, a decoder, and a multi-level output layer to achieve accurate segmentation of the ROI of the scapula, iliac crest, and spinous process landmarks; 2) The encoder of the lightweight U-Net++ network uses the MobileOne-S3 module, whose block structure is: Branch 1 includes, in order: 3×3 depthwise separable convolution, batch normalization, and SiLU activation; Branch 2 includes, in order: 1×1 point-wise convolution, batch normalization, SiLU activation, and 3×3 depth-wise separable convolution; 3) The decoder of the lightweight U-Net++ network is integrated with a dynamic upsampling module to restore image resolution through learnable parameter interpolation; 4) The output layer of the lightweight U-Net++ network predicts multiple region of interest masks in parallel, including: Represents the scapular spine region; represents the posterior superior iliac spine region; Represents the continuous line of the spinous process of the vertebra; 5) The lightweight U-Net++ network uses the Dice loss function to perform pixel-level optimization on the ROIs of the scapula and iliac crest regions; 6) The lightweight U-Net++ network shares the backbone feature extraction layer with the key anatomical feature point detection network HRNet-W32 used for iliac crest baseline correction described in step S2.

5. The method for three-dimensional spine assessment based on RGB images and anatomical constraints according to claim 1, characterized in that: Step S4 includes: 1) The ROI mask output from step S3 and the RGB image corrected by S2 are input into the dual-branch depth estimation network structure; 2) The dual-branch depth estimation network: Branch 1: Based on the MiDaS encoder-decoder structure, extract the relative depth map from the RGB image corrected in step S2 ; Branch 2: A vertebral attention mechanism is introduced after the fourth layer of the decoder in branch 1. By increasing the depth prediction weights of the scapula, iliac crest, and spinous process vertebrae, this ensures that the network is more sensitive to learning key areas of the spine and improves the accuracy of depth estimation. Branch 1 and branch 2 are fused through channel splicing: , in: is the feature map output by branch 1, is the feature map output by branch 2, For channel splicing operation, is the channel number 1×1 convolution, To fuse features, the initial depth map is obtained by optimizing the multi-constraint loss function ; 3) The vertebral attention mechanism: enter: , in: are the height, width and number of channels of the feature map, respectively. It is the feature map output by the 4th layer of the decoder of branch 1; Anatomical position encoding: The positions of the scapula, iliac crest, and spinous process of the vertebra are used to generate anatomical position encoding, which is used as anatomical prior input into the vertebral attention mechanism; Vertebral attention mechanism generates vertebral probability distribution map : , , in: represents the pixel position in the feature map, Corresponding to the height direction, Corresponding to the width direction; Indicates the vertebral index; For the Standard projection position of each vertebra; is the vertebral size correlation parameter; To segment the confidence function, control the weight attenuation of the low confidence cone, It means the Confidence of each vertebra; Attention calculation of vertebral attention mechanism: , , , , in: , , is the linear transformation matrix; is the channel number normalization factor to avoid the gradient disappearance or explosion problem, It is the feature map output by the 4th layer of the decoder of branch 1; is the vertebral probability distribution map generated by the vertebral attention mechanism, It is a normalization function that converts the attention weight into a probability distribution; Output: Generate enhanced feature map: , in: is the layer normalization operation.

6. The method for three-dimensional spine assessment based on RGB images and anatomical constraints according to claim 5, characterized in that: In step S4, the multi-constraint loss function is: , in: : Depth estimation error loss, used to minimize the difference between predicted depth and true depth; : Statistical shape model SSM projection loss ensures that the depth estimation conforms to the vertebral statistical model; : Loss of left-right hemisymmetry, constraining the symmetry of the spine; : Keypoint loss, used to optimize the positions of keypoints of the scapula, iliac crest, and spinous process; : is the weight coefficient; Among them: depth estimation error loss : Using scale-invariant logarithmic error: , in: , are the height and width of the corrected RGB image, and the total number of pixels : The real depth map in pixels The depth value at To predict the depth map at pixel The depth value at Represents two different sets of pixel coordinates in an image, used to traverse all possible pixel pairs; is the balancing factor to optimize the local depth relationship constraint; SSM projection loss : Force the depth map to satisfy the projection constraints of the pyramidal statistical shape model: , in: is the total number of vertebral categories; is the first The three-dimensional coordinates of the vertebra; It is a perspective projection operation; The network prediction The vertebral segmentation results are: is the absolute depth map; Loss of left-right hemisymmetry : Penalize left-right half depth asymmetry: , in: is the depth value of the pixels in the left half of the corrected RGB image; is the depth value of its mirror-symmetrical position; , are the height and width of the image respectively; is the threshold for symmetry loss; is a weighting factor based on different regions of the spine, used to dynamically adjust the symmetry loss; Keypoint loss : , in: is the number of key points; are the coordinates of the actual key points of the scapula, iliac crest, and spinous process; : Coordinates of keypoints predicted by the network.

7. The method for three-dimensional spine assessment based on RGB images and anatomical constraints according to claim 6, characterized in that: Step S5 includes: 1) Height-vertebral size mapping Based on the statistical model, using the height of the subjects Predict vertebral height , establish the regression relationship between vertebral height and body height: , in: is the predicted vertebral height; is the height of the subject; , is the regression coefficient, which is set according to statistical data; 2) Adaptive height scale calibration calculation To calibrate the depth estimation scale, calculate the height scale calibration factor : , vertebral weight , Calculate the height-calibrated depth map based on the height-calibration factor : , in: For the The vertical coordinate of the vertebra in the upright back RGB image, the midline of the spine is , is the vertebral size correlation parameter, Vertebral height predicted based on regression model; Standard vertebral height in the statistical shape model SSM; is the number of visible vertebrae; is the initial depth map processed in step S4; 3) BMI compensation mechanism Add fat compensation to depth after height scale calibration: , in: Depth map calibrated for height and weight scale; The depth map is calibrated for height scale; is the fat compensation coefficient; is the vertical coordinate of the pixel point in the upright back RGB image; is the mean vertical coordinate of the iliac crest ROI; is the Gaussian distribution parameter that controls the fat compensation range.

8. The method for three-dimensional spine assessment based on RGB images and anatomical constraints according to claim 7, characterized in that: Step S6 includes: The depth map calibrated with height and weight scale in step S5 is input into the two-stage generator in the cascaded generative adversarial network architecture. , , eliminating clothing interference and achieving 3D surface reconstruction of the back, including: 1) Phase 1: Eliminate clothing interference In the first stage, the generator is constructed by partial convolution , used to restore the skin texture in the clothing-occluded area of the image, generator The encoder-decoder architecture is used to extract features from the input image and reconstruct occluded areas. The encoder is used to extract features from the input image and identify areas occluded by clothing and their boundaries. The decoder is used to restore the occluded skin texture information based on the encoder output. 2) Second stage: 3D back surface reconstruction In the second phase, the generator The morphology of the skin surface is modified by combining anatomical priors and biomechanical modeling. Costal arch modeling is used to calculate the costal arch curve based on the subject's height and BMI to ensure that the generated skin surface meets the requirements of the human anatomical structure. Erector spinae stress simulation: simulates the effects of muscle activity and stress distribution on skin surface morphology, and combines muscle elasticity and stress distribution models to modify the skin surface to make it conform to the rationality of physiological structure; 3) Multi-scale identification mechanism The realism of the generated skin is improved through a multi-scale discriminator, ensuring that the anatomical curvature matches that of real skin, including: Global Discriminator : Input 512×512 image and evaluate the overall authenticity; Local Discriminator : Identify clothing seam details for a 128×128 local area; Physiological plausibility discriminator : The KL divergence between the generated skin curvature and the real anatomical database is calculated to evaluate the physiological rationality of the generated image.

9. The method for three-dimensional spine assessment based on RGB images and anatomical constraints according to claim 1, characterized in that: Step S7 includes: 1) NURBS spine curve fitting: Based on the back 3D surface reconstruction data in step S6, a NURBS curve is constructed to fit the centerline of the human spine; 2) Parameter adjustment strategy: Optimizing NURBS node vectors , to suit individual spinal anatomy: 3) Weight distribution strategy: Dynamically adjust the influence of control points based on spinal morphology to improve fitting accuracy and ensure that the curve matches the actual anatomy in key anatomical areas; 4) Dynamic constrained optimization: Optimize curve smoothness, anatomical consistency, and data fit to make the fitting curve conform to the individual physiological structure while minimizing errors.

10. The method for three-dimensional spine assessment based on RGB images and anatomical constraints according to claim 1, characterized in that: Step S8 includes: 1) Cobb angle assessment: Based on the spinal curve fitting results, the end vertebrae are automatically detected and the Cobb angle is calculated to assess the severity of scoliosis; 2) Trunk rotation angle assessment: Based on the back surface point cloud, the maximum height difference between the two symmetrical positions of the spine is calculated to assess the torsional deformity caused by scoliosis; 3) Sagittal curvature assessment: Calculate the curvature of the thoracic and lumbar spine and assess the lordosis and kyphosis of the spine.

11. A three-dimensional spine assessment system based on RGB images and anatomical constraints, characterized by: include: Image acquisition module: used to obtain standardized upright back RGB images of subjects through smart mobile devices and upload the image data to the cloud computing server for processing; The image preprocessing and posture correction module is used to enhance and denoise the acquired upright back RGB images and adjust the shooting posture error based on anatomical landmarks to ensure standardized input of the spinal region; The anatomical landmark detection module is used to automatically detect the anatomical landmarks of the scapula and iliac crest based on a deep learning model and extract the spinal region of interest (ROI) to provide a basis for subsequent modeling; A monocular depth estimation module generates an initial depth map by integrating a vertebral attention mechanism with a depth estimation model that uses a multi-constraint loss function, including a statistical shape model (SSM) projection constraint. The scale calibration module is used to perform scale correction on the initial depth map based on the subject's height, weight, and BMI physiological parameters to ensure that the spine measurement conforms to the individual's anatomical proportions; Clothing interference removal and skin surface assessment module, which is used to remove clothing occlusion interference based on a generative adversarial network and restore the true three-dimensional morphology of the back skin; The 3D spine assessment module is used to generate an accurate 3D spine curve model based on the 3D morphological data of the back skin through curve fitting methods, and can adapt to different degrees of spinal morphological changes; An evaluation parameter calculation module is used to calculate and output the coronal Cobb angle, trunk rotation angle, and sagittal curvature parameters based on the generated three-dimensional spinal curve model; The system output module is used to return the analysis results to the smart mobile device in a visual manner for display, and provide an automated diagnostic report containing spinal deformity assessment data for reference by medical professionals or individual users.

Citation Information

Patent Citations

  • Scoliosis screening method based on 2D RGB image

    CN115526845A

  • Spine three-dimensional structure reconstruction method based on deep learning

    CN116402954A