Spine three-dimensional evaluation method and system based on RGB image and anatomical constraint
Through a method based on RGB image and anatomical constraints, using deep learning and generative adversarial networks to perform three-dimensional reconstruction of the spine, the high cost and radiation risk problems of scoliosis screening and follow-up in the prior art are solved, and a non-invasive, low-cost high-precision assessment is achieved.
Patent Information
- Application Number
- CN202510793254.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing scoliosis screening and follow-up methods have problems such as large subjective errors, high cost, radiation risk and inappropriate monitoring, and lack of non-invasive, convenient and low-cost high-precision evaluation technology.
Through an RGB image and anatomical constraint method, intelligent mobile devices are used to capture back images, combined with deep learning and generative adversarial networks, spine 3D reconstruction and evaluation, including image preprocessing, anatomical marker detection, depth estimation, scale calibration and clothing interference elimination, to generate an accurate spine 3D curve model.
It realizes high-precision three-dimensional spine reconstruction and health assessment without expensive equipment, reduces the cost of screening and monitoring, reduces radiation risks, and supports widespread applications in primary medical scenarios.
Smart Images

Figure CN120298415A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image analysis and computer-aided diagnosis, and specifically relates to a three-dimensional spine assessment method and system based on RGB images and anatomical constraints, which is particularly suitable for the screening and quantitative assessment of spinal morphological abnormalities such as scoliosis and kyphosis. Background Art
[0002] Scoliosis is a common spinal deformity, which is mainly manifested as abnormal curvature of the spine in the coronal, sagittal and horizontal planes. Studies have shown that the prevalence of adolescent idiopathic scoliosis (AIS) is about 2-4%, and the incidence rate is on the rise with the increase of bad postures such as sitting for a long time and bowing the head. If not screened and intervened in time, it may lead to worsening of spinal deformity, affect lung function, and even cause chronic pain, seriously affecting the physical and mental health and quality of life of adolescents.
[0003] Scoliosis is characterized by progressive development, and the condition may rapidly worsen during adolescence, so long-term monitoring and multiple measurements are required to accurately assess the trend of spinal changes. Current screening methods mainly rely on visual examination, spinal angle measuring instruments (such as spinal goniometers), three-dimensional reconstruction, X-ray examination and other methods. However, manual measurement has subjective errors, and three-dimensional reconstruction equipment is expensive. Although X-ray examination can provide high-precision evaluation, it has radiation risks and is not suitable for frequent monitoring. The cost of long-term follow-up is high. In the screening and follow-up management of scoliosis, there is an urgent need for a non-invasive, convenient, low-cost and easy-to-large-scale screening technology to achieve early detection, regular monitoring and individualized health management, thereby effectively reducing the risk of disease progression and improving the health level of adolescents. Summary of the invention
[0004] In response to the above technical problems, the present invention proposes a three-dimensional spine assessment method and system based on RGB images and anatomical constraints. This method can achieve high-precision three-dimensional reconstruction of the spine with only one ordinary RGB image, and further perform spinal morphology assessment. This technology not only avoids the radiation risk in X-rays, but also provides a simple and accurate spinal health assessment without the need for expensive equipment, and has broad application potential and promotion value.
[0005] The technical solutions that can achieve the purpose of the invention include:
[0006] A three-dimensional spine assessment method based on RGB images and anatomical constraints, the method comprising:
[0007] S1: The smart mobile device takes a single upright back RGB image as required and inputs the height and weight parameters of the subject;
[0008] S2: Preprocess the frontal back RGB image taken by the intelligent mobile device once and correct the anatomical symmetry;
[0009] S3: Extract the regions of interest of the scapula and iliac crest landmark points in the preprocessed and corrected RGB image through a lightweight segmentation network;
[0010] S4: Based on the extracted region-of-interest image, generate an initial depth map through a depth estimation model that fuses the vertebral attention mechanism and a multi-constraint loss function, and the multi-constraint loss function includes the statistical shape model SSM projection constraint;
[0011] S5: Perform physiological scale calibration on the initial depth map based on the height and weight parameters;
[0012] S6: Input the initial depth map after physiological scale calibration into a cascaded generative adversarial network to eliminate clothing interference and obtain the three-dimensional surface reconstruction data of the back;
[0013] S7: Based on the three-dimensional surface reconstruction data of the back, adaptively adjust the NURBS parameters to fit the three-dimensional curve of the spine in combination with biomechanical characteristics;
[0014] S8: Output the coronal Cobb angle, trunk rotation angle, and sagittal curvature parameters based on the three-dimensional curve of the spine.
[0015] A three-dimensional spine evaluation system based on RGB images and anatomical constraints, including:
[0016] Image acquisition module: used to obtain the standardized frontal back RGB image of the subject through an intelligent mobile device and upload the image data to the cloud computing server for processing;
[0017] Image preprocessing and pose correction module, used to enhance and denoise the obtained frontal back RGB image, and adjust the shooting pose error based on anatomical landmark points to ensure the standardized input of the spine area;
[0018] Anatomical landmark detection module, used to automatically detect the scapula and iliac crest anatomical landmark points based on a deep learning model and extract the region of interest ROI of the spine to provide a basis for subsequent modeling;
[0019] Monocular depth estimation module, used to generate an initial depth map through a depth estimation model that fuses the vertebral attention mechanism and a multi-constraint loss function, and the multi-constraint loss function includes the statistical shape model SSM projection constraint;
[0020] Scale calibration module, used to perform scale correction on the initial depth map based on the physiological parameters of the subject's height, weight, and BMI to ensure that the spine measurement conforms to the individual anatomical proportion;
[0021] The clothing interference elimination and skin surface reconstruction module is used to remove the interference of clothing occlusion based on the generative adversarial network and restore the true three-dimensional shape of the back skin;
[0022] The three-dimensional spinal modeling module is used to generate an accurate three-dimensional spinal curve model based on the three-dimensional back skin morphology data through the curve fitting method and adapt to different degrees of spinal morphological changes;
[0023] The evaluation parameter calculation module is used to calculate and output the coronal Cobb angle, trunk rotation angle, and sagittal curvature parameters based on the generated three-dimensional spinal curve model;
[0024] The system output module is used to return the analysis results to the intelligent mobile device for display in a visual manner and provide an automated diagnostic report containing spinal deformity evaluation data for medical professionals or individual users to reference.
[0025] Beneficial effects:
[0026] The three-dimensional spinal evaluation method and system based on RGB images and anatomical constraints provided by the present invention have the following advantages compared with the prior art: Based on RGB depth estimation technology and statistical anatomical model constraints, by fusing a depth estimation network with statistical shape model (SSM) constraints and vertebral attention mechanism, the anatomical perception ability of the spinal region is enhanced; A dynamic calibration strategy driven by physiological parameters such as height and BMI is adopted to solve the problems of individual anatomical differences and errors in obese patients; Through the cascaded GAN technology combined with clothing interference elimination and physiological surface reconstruction, the accuracy and authenticity of the back three-dimensional shape restoration are improved; Based on NURBS curve fitting and biomechanical property optimization, a high-precision three-dimensional spinal model is adaptively generated to realize the joint evaluation of multiple parameters such as the coronal Cobb angle and trunk rotation angle. Compared with traditional methods, the present invention does not require professional equipment or radiological examinations, and can complete non-invasive and low-cost three-dimensional spinal reconstruction and health assessment only through a single back image taken by an intelligent device, support extensive screening and dynamic monitoring in primary medical scenarios, significantly reduce the screening threshold and health risks, and provide technical support for the early detection and intervention of spinal deformities. Description of the Drawings
[0027] Figure 1 It is a schematic flowchart of a three-dimensional spinal evaluation method based on RGB images and anatomical constraints according to an embodiment of the present invention;
[0028] Figure 2 It is a module diagram of a three-dimensional spinal evaluation system based on RGB images and anatomical constraints according to an embodiment of the present invention.
[0029] Figure numerals: image acquisition module 210, image preprocessing and posture correction module 220, anatomical landmark detection module 230, monocular depth estimation module 240, scale calibration module 250, clothing interference elimination and skin surface re-evaluation module 260, spine three-dimensional evaluation module 270, evaluation parameter calculation module 280, system output module 290. DETAILED DESCRIPTION
[0030] In order to make those skilled in the art understand the technical solution of the present invention more clearly, the specific implementation of the present invention is further described in detail below in conjunction with the drawings and embodiments. It should be understood that the following description is only a specific example of the technical solution of the present invention and is not intended to limit the scope of protection of the present invention.
[0031] The present invention discloses a three-dimensional spine assessment method and system based on RGB images and anatomical constraints. Non-invasive and low-cost three-dimensional spine reconstruction and health assessment can be completed through a single back image taken by a smart device.
[0032] Figure 1 A schematic diagram of a three-dimensional spine assessment method based on RGB images and anatomical constraints according to an embodiment of the present invention. Figure 1 The specific steps of the method are described. Among them, a mobile phone and a WeChat applet are used as examples to demonstrate the specific application of the present invention. Figure 1 As shown, the method comprises the following steps:
[0033] S1: The smart mobile device takes a single upright back RGB image as required and inputs the height and weight parameters of the subject. In an embodiment, the smart mobile device may be, for example, a smart phone, a tablet, etc. For example, a small program is opened on a mobile phone, and a single upright back RGB image that meets the requirements is taken as required and the height and weight parameters of the subject are input.
[0034] Step S1 may specifically include:
[0035] 1) Enter the subject's height, weight and personal identity information;
[0036] 2) The smart mobile device (such as a mobile phone camera) is located on a horizontal plane based on the height of the subject's iliac crest, and the camera is taken in a direction perpendicular to the plane of the human back at the midpoint of the iliac crest line. The shooting distance is satisfy:
[0037] ,
[0038] in: is the height of the subject, in meters.
[0039] Among them: The subject stands with feet together and toes touching the reference plane, the heel spacing is no more than 5 cm, the upper limbs hang naturally, the palms are attached to the outside of the thighs and the thumbs face forward; The subject needs to wear tight-fitting clothes (such as sports tight-fitting clothes) to ensure that there is no obvious occlusion in the spinal area, which is convenient for the subsequent 3D modeling process. The tight-fitting clothes should fit the body, avoiding loose or thick clothes to ensure that the spinal contour is clearly visible;
[0040] 3) Before shooting, use the gyroscope and accelerometer data built into the smart mobile device (mobile phone) to detect the pitch angle and roll angle of the smart mobile device (mobile phone). By obtaining the angle information of the smart mobile device in real time, ensure that the smart mobile device maintains a vertical shooting. If the pitch angle or roll angle of the smart mobile device exceeds the preset range , the system will issue a warning prompt to remind the user to adjust the posture of the mobile phone to ensure that the angle of the smart mobile device is correct during shooting.
[0041] S2: Preprocess and anatomically symmetrically correct the frontal back RGB image taken by the smart mobile device once, including:
[0042] The preprocessing includes: performing image enhancement and denoising preprocessing on the captured frontal back RGB image using bilateral filtering, adaptive histogram equalization (CLAHE), and frequency domain Wiener filtering to improve the clarity and contrast of the back anatomical features.
[0043] The anatomical symmetry correction specifically includes:
[0044] 1) Use the improved high-resolution network (HRNet-W32) to detect the key anatomical feature points of the back in the preprocessed RGB image, including the inferior angle points of the scapula ( ), and the upper edge points of the left and right iliac crests , and obtain the two-dimensional image coordinates of the above feature points;
[0045] 2) Based on the two-dimensional coordinates of the upper edge points of the iliac crests, define and calculate the coronal plane symmetry axis of the human back , and the specific mathematical expression is:
[0046] ,
[0047] Among them: , represents the two-dimensional pixel coordinate components in the image plane, corresponding to the horizontal and vertical directions of the image respectively, the upper edge point of the left iliac crest , the upper edge point of the right iliac crest , , , .
[0048] 3) Calculate the angle between the symmetry axis of the human back's coronal plane and the vertical direction of the image , defined as:
[0049] ,
[0050] When the angle does not exceed the preset threshold , then no rotation correction is performed, and directly enter the subsequent depth estimation and three-dimensional evaluation stage. When the said angle exceeds the preset threshold , perform pose correction on the RGB image by constructing and executing a rotation affine transformation matrix :
[0051] ,
[0052] Where: is the rotation affine transformation matrix, , is the center of the picture, to ensure that the iliac crest reference line is parallel to the horizontal axis of the image after correction, so as to eliminate the global tilt error caused by the perspective distortion caused by the shooting angle, distance and pose.
[0053] S3: Extract the regions of interest of the scapula and iliac crest landmark points in the preprocessed and corrected RGB image through a lightweight segmentation network, including:
[0054] 1) Based on the standardized back image corrected in step S2, adopt an improved lightweight U-Net++ network, combined with an encoder, a decoder and a multi-level output layer, to achieve accurate segmentation of the regions of interest (ROIs) of landmark points such as the scapula, iliac crest and spinal spinous process;
[0055] 2) The encoder of the lightweight U-Net++ network adopts the MobileOne-S3 module, which is used to enhance the local and global feature extraction capabilities and optimize the calculation efficiency. Its block structure is:
[0056] Branch 1: 3×3 depthwise separable convolution → batch normalization → SiLU activation;
[0057] Branch 2: 1×1 pointwise convolution → batch normalization → SiLU activation → 3×3 depthwise separable convolution.
[0058] The feature maps of the above two branches are fused by element-wise addition and then the encoded feature map is output.
[0059] The convolution kernel parameters and batch normalization layer parameters of the module are automatically learned and determined by the backpropagation algorithm during network training; the number of convolution kernels (number of channels) in the convolution module belongs to the hyperparameters of the network structure, usually in the range of 32 to 128, and the best combination is determined through cross-validation experiments. In this embodiment, 64 is taken as the optimal solution.
[0060] 3) The decoder of the lightweight U-Net++ network integrates a dynamic upsampling module ( ), and restores the image resolution through learnable parameter interpolation to enhance the adaptability of the model to various image inputs and ensure the retention of detailed information during the segmentation process. The specific calculation method is as follows:
[0061] ,
[0062] where: represents the feature map after upsampling; is the input feature map; is the learnable weight of the th interpolation branch; is the upsampling result of the th nearest neighbor interpolation calculation; is the number of interpolation branches, determined through experimental cross-validation (recommended value = 2).
[0063] 4) The output layer of the lightweight U-Net++ network predicts multiple ROI masks in parallel, including: represents the acromion region; represents the posterior superior iliac spine region; represents the continuous line of vertebral spinous processes.
[0064] 5) The lightweight U-Net++ network uses the Dice loss function to optimize the ROI of the scapula and iliac crest regions at the pixel level to improve the segmentation accuracy of the anatomical landmark regions. The application of the Dice loss function ensures the high precision and consistency of the segmentation results, thereby improving the accuracy of the subsequent processing stage. The Dice loss function is defined as follows:
[0065] ,
[0066] where: and are the height and width of the segmented image respectively; represents the true annotation value at the corresponding position ; represents the value of the network prediction at the pixel position ; ε is a smoothing coefficient, and the value in this embodiment is 1×10 -6 .
[0067] 6) The lightweight U-Net++ network shares the backbone feature extraction layer with the key anatomical feature point detection network HRNet-W32 used for the iliac crest baseline correction in step S2 to achieve unified understanding of anatomical structures, ensure that the segmentation results are aligned with the globally corrected human body postures, and thus improve the accuracy of subsequent depth estimation and three-dimensional evaluation.
[0068] S4: Based on the extracted region-of-interest images, generate an initial depth map through a depth estimation model that fuses the vertebral attention mechanism and a multi-constraint loss function. The multi-constraint loss function includes the statistical shape model SSM projection constraint, including:
[0069] 1) Input the region-of-interest mask map (including the scapula, iliac crest, and spinal spinous process regions) output in step S3 and the RGB image corrected in S2 into the dual-branch depth estimation network structure;
[0070] 2) The dual-branch depth estimation network estimates depth information from two perspectives of appearance features and anatomical constraints respectively, realizing the organic integration of anatomical constraints and appearance features;
[0071] Branch 1: (Appearance feature extraction): Retain the MiDaS encoder-decoder structure and output a relative depth map ; and embed a frequency-domain artifact removal module (FDM) in the 4th layer of the decoder:
[0072] ,
[0073] where: is a dynamically generated high-frequency noise mask for selectively filtering frequency-domain artifacts in the RGB image. is an element-wise multiplication to mask high frequencies, is the frequency-domain artifact removal module for suppressing high-frequency noise artifacts in the RGB image, is a convolutional layer, is the inverse discrete cosine transform to recover from the frequency domain to the spatial domain, is the discrete cosine transform to convert the image from the spatial domain to the frequency domain, is the feature map output by the 4th layer of the decoder of Branch 1. The convolutional kernel parameters are automatically determined during network training.
[0074] This branch extracts depth features through the global information of the image.
[0075] Branch 2: (Anatomical constraint fusion): Introduce the vertebral attention mechanism after the 4th layer of the decoder in Branch 1 to enhance the depth prediction accuracy of the vertebral region. By enhancing the depth prediction weights of vertebral regions such as the scapula, iliac crest, and spinal spinous process, ensure that the network is more sensitive to the learning of key regions of the spine and improve the accuracy of depth estimation.
[0076] Branch 1 and Branch 2 are spliced and fused through a channel:
[0077] ,
[0078] Where: is the feature map output by Branch 1, is the feature map output by Branch 2, is the channel splicing operation, is the 1×1 convolution of the number of channels, is the fused feature, optimized by a multi-constraint loss function to obtain the initial depth map .
[0079] 3) Vertebral attention mechanism: Enhance the depth prediction weight of the vertebral region, ensure that the network is more sensitive to the learning of the vertebral region, and improve the depth estimation accuracy.
[0080] Input:
[0081] ,
[0082] Where: are the height, width, and number of channels of the feature map respectively. is the feature map output by the 4th layer of the decoder of Branch 1.
[0083] Anatomical position encoding: Generate anatomical position encoding using the positions of vertebral regions such as the scapula, iliac crest, and spinal spinous process, and input it as an anatomical prior into the vertebral attention mechanism to further enhance the network's attention to the spinal region.
[0084] The vertebral attention mechanism generates a vertebral probability distribution map :
[0085] , ,
[0086] Where: represents the pixel position in the feature map, corresponds to the width direction, corresponds to the height direction; represents the th vertebral index; is the standard projection position of the th vertebra; is the vertebral size correlation parameter, satisfying , is the estimated vertebral height value; is the segmentation confidence function, controlling the weight decay of low-confidence vertebrae, is the confidence representing the th vertebra;
[0087] Attention calculation of the vertebral attention mechanism:
[0088] , , ,
[0089] ,
[0090] Where: , , is a linear transformation matrix; is the channel number normalization factor to avoid the problem of gradient vanishing or explosion. is the feature map output by the 4th layer of the decoder in branch 1, is the vertebral body probability distribution map generated by the vertebral attention mechanism, is the normalized exponential function, which calculates the attention weight and converts it into a probability distribution.
[0091] Output: Generate an enhanced feature map:
[0092] ,
[0093] Where, is layer normalization.
[0094] 3) The multi-constraint loss function
[0095] ,
[0096] Where:
[0097] : Depth estimation error loss, which is used to minimize the difference between the predicted depth and the true depth;
[0098] : Statistical shape model (SSM) projection loss, which ensures that the depth estimation conforms to the vertebral body statistical model;
[0099] : Left and right half-side symmetry loss, which constrains the symmetry of the spinal column morphology and improves the depth stability;
[0100] : Key point loss, which is used to optimize the positions of key points such as the scapula, iliac crest, and spinal spinous process;
[0101] : is the weight coefficient, usually set in the range of λ∈(0,1), and the specific value can be adjusted according to experiments to optimize the depth estimation accuracy.
[0102] 4) Depth estimation error loss : Adopt scale-invariant logarithmic error:
[0103] ,
[0104] where: , are the height and width of the image, respectively, and the total number of pixels : is the depth value of the real depth map at pixel ; is the depth value of the predicted depth map at pixel ; represents two different sets of pixel coordinates in the image, used to traverse all possible pixel pairs. is the balance factor to optimize the local depth relationship constraint.
[0105] 5) SSM projection loss : Forces the depth map to satisfy the projection constraint of the vertebral statistical shape model:
[0106] ,
[0107] where: is the total number of vertebral classes (usually 17 main vertebrae); is the three-dimensional coordinate of the th vertebra in the statistical shape model (SSM); is the perspective projection operation; is the segmentation result of the th vertebra predicted by the network. is the absolute depth map.
[0108] 6) Left-right hemisymmetry loss : Penalizes left-right depth asymmetry:
[0109] ,
[0110] where: is the depth value of the pixels in the left half of the image; is the depth value at its mirror-symmetric position; , are the height and width of the image, respectively; is the threshold of the symmetry loss; is the weighting factor based on different regions of the spine, used to dynamically adjust the symmetry loss.
[0111] 7) Key point loss :
[0112] ,
[0113] where: The number of key points (such as the scapula, iliac crest, spinal spinous process, etc.); The coordinates of the actual key points such as the scapula, iliac crest, spinal spinous process, etc.; : The key point coordinates predicted by the network.
[0114] S5: Physiological scale calibration of the initial depth map based on height and weight parameters, including:
[0115] 1) Height-vertebral size mapping
[0116] Based on a statistical model, using the height of the subject Predict the vertebral height , and establish a regression relationship between the vertebral height and the height:
[0117] ,
[0118] Among them: Is the predicted vertebral height; Is the height of the subject; , Is the regression coefficient, set according to statistical data.
[0119] 2) Adaptive height scale calibration calculation
[0120] To calibrate the depth estimation scale, calculate the height scale calibration factor :
[0121] , vertebral weight ,
[0122] Calculate the depth map after height scale calibration according to the height scale calibration factor :
[0123] ,
[0124] Among them: Is the Vertical coordinate of the th vertebral in the upright back RGB image, with the spinal midline as Is the vertebral size correlation parameter, Based on the vertebral height predicted by the regression model; The standard vertebral height in the statistical shape model SSM; Is the number of visible vertebrae, default is 5, if the image is complete, 17 is used; Is the initial depth map processed by step S4;
[0125] 3) BMI compensation mechanism
[0126] To correct the depth estimation error of obese patients (BMI ≥ 25), add fat compensation to the depth after height scale calibration:
[0127] ,
[0128] where: is the depth map after calibration of height and weight scale; is the depth map after height scale calibration; is the fat compensation coefficient; is the longitudinal coordinate of the pixel point in the frontal back RGB image; is the average ordinate of the iliac crest ROI (the iliac crest is used as the waist benchmark); is the Gaussian distribution parameter for controlling the fat compensation range.
[0129] S6: Input the depth map after physiological scale calibration into the cascaded generative adversarial network to eliminate clothing interference and obtain the back three-dimensional surface reconstruction data:
[0130] Input the depth map after height and weight scale calibration in step S5 into the two-stage generator in the cascaded generative adversarial network architecture , , to achieve eliminating clothing interference and three-dimensional surface reconstruction of the back, including:
[0131] The first stage (texture generation): Construct a generator based on partial convolution , and complete the clothing occlusion area;
[0132] The second stage (physiological bulge synthesis): Based on anatomical prior and biomechanical modeling, correct the skin surface morphology.
[0133] 1) The first stage: Eliminate clothing interference
[0134] Generator :
[0135] ,
[0136] where, (encoder layer): Input image feature extraction module, used to identify clothing boundaries and occlusion areas; (decoder layer): According to the encoder output, restore the occluded skin texture information.
[0137] The calculations adopted in generator include:
[0138] Deformable convolution:
[0139] ,
[0140] where Indicates the pixel position in the feature map, is the convolutional kernel size, applicable to local feature extraction; are the weights of the convolutional kernel; is the pixel position; is the offset of the convolutional kernel; is the convolutional kernel offset variation, allowing the convolutional kernel to adaptively adjust and improve the feature learning ability; is the modulation parameter, controlling the weight influence of different regions.
[0141] Attention gating fusion:
[0142] ,
[0143] ,
[0144] Among them: is the attention mechanism weight calculation function; balance factor is used to balance the feature information of different layers, is the upsampling of the output of the previous decoder.
[0145] 2) The second stage: 3D surface reconstruction of the back
[0146] Costal arch modeling: Simulate the curve of the human costal arch to improve the physiological rationality of skin reconstruction.
[0147] ,
[0148] Among them: , are the regression coefficients, parameters is the height (cm) of the subject, in the range [150, 200]; is the body mass index range [15, 35]. is the output of the costal arch model.
[0149] Erector spinae muscle stress simulation: Estimate the morphological changes of the skin surface according to the muscle stress model. Simulate the influence of muscle activity and stress distribution on the skin surface morphology. This step combines the muscle elasticity and stress distribution models to correct the skin surface to make it conform to the physiological structure rationality.
[0150] ,
[0151] Among them: represents the stress of the skin surface change; represents the basic force without muscle contraction; is the muscle elasticity parameter, representing the elasticity of the muscle; It is a muscle stress elasticity parameter, representing the elastic behavior of muscles under stress; It is a muscle fiber activation function.
[0152] 3) Multi-scale discrimination mechanism
[0153] Improve the authenticity of the generated skin through a multi-scale discriminator to ensure matching with the anatomical curvature of real skin. The structure includes:
[0154] Global discriminator : Input a 512×512 image to evaluate the overall authenticity;
[0155] Local discriminator : For a 128×128 local area, identify details such as clothing seams;
[0156] Physiological rationality discriminator : Calculate the KL divergence between the generated skin curvature and the real anatomical database to evaluate the physiological rationality of the generated image:
[0157] ,
[0158] Among them: , are the curvature distributions of the generated image and the reference database respectively, is the physiological rationality loss function, representing the difference between the generated image and the reference data.
[0159] S7: Based on the three-dimensional surface reconstruction data of the back, combined with biomechanical characteristics, adaptively adjust the NURBS parameters to fit the three-dimensional curve of the spine, including:
[0160] 1) NURBS-based spine curve fitting
[0161] Construct a non-uniform rational B-spline curve (NURBS) based on the three-dimensional surface reconstruction data of the back in step S6 to fit the centerline of the human spine, improving the accuracy and smoothness of the three-dimensional reconstruction.
[0162] ,
[0163] Among them:
[0164] : The position of the spine curve at parameter is, represents the normalized parameter, with a range of .
[0165] : The pixel position of the th control point, is the total number of control points, representing the key positions of the spine (scapula, iliac crest, spinal spinous process, etc.).
[0166] : B-spline basis function, is the control point index, is the curve order, commonly = 3, which determines the local smoothness and fitting ability.
[0167] : Control point weights, used to adjust the local curvature to ensure a more reasonable fit to the spinal morphology.
[0168] 2) Parameter adjustment strategy
[0169] Optimize the knot vector of NURBS , to adapt to the individual spinal anatomy and ensure more accurate curves at key positions (such as cervical vertebra C7, thoracic vertebrae T1 - T12, lumbar vertebrae L1 - L5).
[0170] ,
[0171] where:
[0172] : The position of the th optimized knot.
[0173] : The th known vertebral position point (obtained from 3D modeling).
[0174] : The Euclidean distance between vertebrae.
[0175] : Normalization of the total distance between all vertebrae.
[0176] This optimization method adaptively adjusts the knot density, providing a higher-precision curve description in the high-curvature regions of the spine (such as the thoracic curvature area), while reducing redundant calculations in the low-curvature regions (such as the lumbar spine).
[0177] 3) Weight assignment strategy
[0178] Dynamically adjust the influence of control points according to the spinal morphology to improve the fitting accuracy and ensure that the curve can accurately match the actual anatomy in key anatomical regions (such as the high-curvature points of scoliosis).
[0179] ,
[0180] where:
[0181] : The weight of the th control point, ensuring a higher-precision fit in the high-curvature interval of the spine.
[0182] : The local curvature value of the th control point.
[0183] : The mean value of the curvatures of all control points (used for normalization).
[0184] : The standard deviation of the curvatures of all control points (controlling the overall smoothness).
[0185] When the curvature of the spine is high (such as in scoliosis patients), this strategy automatically increases the weight to ensure that the fitted curve is more in line with the actual situation. When the spine is relatively straight, the weight fluctuation is reduced to ensure a smooth transition of the curve.
[0186] 4) Dynamic constraint optimization
[0187] Optimize the curve smoothness, anatomical consistency and data fitting degree simultaneously, so that the fitted curve conforms to the individual physiological structure while minimizing the error.
[0188] ,
[0189] Among them:
[0190] (Curve smoothness): Avoid excessive tortuosity and ensure compliance with the natural curve of the spine.
[0191] (Anatomical prior constraint): Ensure that the curve conforms to the true spinal anatomical morphology.
[0192] (Data fitting error): Ensure that the curve fits the spinal point cloud data obtained from depth estimation.
[0193] , , is the weight coefficient, usually set in ∈ (0, 1).
[0194] is the loss function, is to minimize the loss function, is the shape control point of the spinal curve, is the weight of the control point.
[0195] This optimization strategy can adapt to the spinal morphology of different individuals and can provide reasonable three-dimensional reconstructions in both normal populations and populations with spinal abnormalities.
[0196] S8: Output the coronal Cobb angle, trunk rotation angle and sagittal curvature parameters based on the three-dimensional spinal curve, including:
[0197] 1) Cobb angle assessment (Based on the spinal curve fitting results, automatically detect the end vertebrae and calculate the Cobb angle to evaluate the severity of scoliosis).
[0198] End vertebrae identification: Extract the curvature extreme points along the spinal curve as the landmark points of the upper and lower end vertebrae.
[0199] , ,
[0200] Among them: : The point of maximum curvature on the spinal curve (end vertebra position). : The spinal curvature function in the coronal plane, representing the local bending degree of the curve at .
[0201] Plane fitting: Based on the 3D point cloud data near the end vertebrae, calculate the maximum tilt angle to determine the actual angle of the end vertebrae. Perform principal component analysis (PCA) on the point cloud data to extract the main direction.
[0202] ,
[0203] ,
[0204] Among them: is the point cloud data near the upper end vertebra, is the point cloud data near the lower end vertebra, is the point cloud range offset near the end vertebra, : The endplate direction vectors of the upper and lower end vertebrae.
[0205] Cobb angle calculation: Calculate the angle between the endplate lines of the upper and lower end vertebrae.
[0206] ,
[0207] Among them:
[0208] : The normalized modulus length of the endplate direction.
[0209] 2) Assessment of the angle of trunk rotation (ATR) (Based on the back surface point cloud, calculate the maximum height difference between the left and right sides of the spine to evaluate the torsional deformity caused by scoliosis).
[0210] Reconstruction of the back surface point cloud: Based on depth estimation and 3D reconstruction, extract the spinal midline points and their corresponding highest points on the left and right backs.
[0211] Select the iliac crest point (PSIS) as the reference and fit the horizontal reference plane:
[0212] ,
[0213] Wherein: , , are the components of the plane normal vector, is the distance between the plane and the origin, , , are the coordinates of each point in the point cloud.
[0214] In the horizontal reference plane, scan the highest points on both the left and right sides of the spine:
[0215] ,
[0216] Wherein: The height values of the highest points on the left and right surfaces of the back.
[0217] ATR calculation: Calculate the maximum left - right height difference of the back protrusion and convert it into an angle.
[0218] ,
[0219] Wherein: : The maximum left - right height difference; : The horizontal projection distance.
[0220] 3) Sagittal plane curvature assessment (calculate the curvature of the thoracic - lumbar segment of the spine and evaluate the lordosis and kyphosis conditions of the spine).
[0221] Based on arc - length parameterized curvature, calculate the overall bending degree of different segments of the spine.
[0222] ,
[0223] Wherein: : Thoracic kyphosis angle; : Lumbar - sacral lordosis angle; is the arc - length parameterized curvature at the sagittal plane position. - is from the first thoracic vertebra to the twelfth thoracic vertebra, is the first lumbar vertebra, is the first sacral vertebra.
[0224] As Figure 2 shown, an embodiment of the present invention further provides a three - dimensional spine assessment system based on RGB images and anatomical constraints, specifically including:
[0225] Image acquisition module 210: It is used to obtain the standardized upright back RGB image of the subject through an intelligent mobile device (such as a smart phone, tablet), and upload the image data to the cloud computing server for processing;
[0226] The image preprocessing and pose correction module 220 is used to enhance and denoise the acquired frontal back RGB image, and adjust the shooting pose error based on anatomical landmark points to ensure standardized input of the spinal region;
[0227] The anatomical landmark detection module 230 is used to automatically detect anatomical landmark points such as the scapula and iliac crest based on a deep learning model, and extract the region of interest (ROI) of the spine to provide a basis for subsequent evaluation;
[0228] The monocular depth estimation module 240 is used to generate an initial depth map through a depth estimation model that fuses the vertebral attention mechanism and a multi-constraint loss function (including the statistical shape model SSM projection constraint);
[0229] The scale calibration module 250 is used to perform scale correction on the initial depth map based on physiological parameters such as the subject's height, weight, and BMI to ensure that spinal measurements conform to individual anatomical proportions;
[0230] The clothing interference elimination and skin surface re-evaluation module 260 is used to remove the interference of clothing occlusion based on a generative adversarial network (GAN) and restore the true three-dimensional shape of the back skin;
[0231] The spinal three-dimensional evaluation module 270 is used to generate an accurate spinal three-dimensional curve model based on the three-dimensional shape data of the back skin through a curve fitting method and adapt to different degrees of spinal shape changes;
[0232] The evaluation parameter calculation module 280 is used to calculate and output spinal physiological curve parameters such as the coronal Cobb angle, axial trunk rotation (ATR), and sagittal curvature based on the generated spinal three-dimensional curve model;
[0233] The system output module 290 is used to return the analysis results to the intelligent mobile device for display in a visual manner and provide an automated diagnostic report containing spinal deformity evaluation data for reference by medical professionals or individual users.
[0234] The above describes the basic principles, main features, and advantages of the present invention. It should be understood that the embodiments of the present invention are only used to illustrate the technical solutions of the present invention and are not subject to any form of limitation. Any technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A three-dimensional spinal evaluation method based on RGB images and anatomical constraints, characterized in that, The method includes: S1: The intelligent mobile device takes a single frontal back RGB image as required and inputs the height and weight parameters of the subject to be measured; S2: Preprocess and correct the anatomical symmetry of the frontal back RGB image taken by the intelligent mobile device once; S3: Extract the regions of interest of the scapula and iliac crest landmark points in the preprocessed and corrected RGB image through a lightweight segmentation network; S4: Based on the extracted region-of-interest image, generate an initial depth map through a depth estimation model that fuses the vertebral attention mechanism and a multi-constraint loss function, and the multi-constraint loss function includes the projection constraint of the statistical shape model SSM; S5: Physiologically scale-calibrate the initial depth map based on the height and weight parameters; S6: Input the physiologically scale-calibrated depth map into a cascaded generative adversarial network to eliminate clothing interference and obtain the three-dimensional surface reconstruction data of the back; S7: Based on the three-dimensional surface reconstruction data of the back, adaptively adjust the NURBS parameters to fit the three-dimensional curve of the spine in combination with biomechanical characteristics; S8: Output the coronal Cobb angle, trunk rotation angle, and sagittal curvature parameters based on the three-dimensional curve of the spine.
2. The three-dimensional spinal evaluation method based on RGB images and anatomical constraints according to claim 1, characterized in that Step S1 includes: 1) Input the height, weight, and personal identity information of the subject; 2) The camera of the intelligent mobile device is located on the horizontal plane based on the height of the subject's iliac crest, and takes pictures in the direction perpendicular to the plane where the human back is located at the midpoint of the iliac crest connection line. The shooting distance satisfies: , Wherein: is the height of the subject, in meters; 3) Before shooting, use the gyroscope and accelerometer data built into the intelligent mobile device to detect the pitch angle and roll angle of the intelligent mobile device. By obtaining the angle information of the intelligent mobile device in real time, ensure that the intelligent mobile device maintains vertical shooting. If the pitch angle or roll angle of the intelligent mobile device exceeds the preset range, issue a warning prompt.
3. The three-dimensional spine evaluation method based on RGB images and anatomical constraints according to claim 1, characterized in that In step S2, The preprocessing includes: performing image enhancement and denoising preprocessing on the frontal back RGB back image taken based on step S1 using bilateral filtering, adaptive histogram equalization, and frequency-domain Wiener filtering; The anatomical symmetry correction includes: Using an improved high-resolution network to process the preprocessed frontal back RGB image, detecting the key anatomical feature points of the back, including the inferior angle points of the scapula, the upper edge points of the left iliac crest, and the upper edge points of the right iliac crest, and obtaining the two-dimensional image coordinates of these feature points; Based on the two-dimensional coordinates of the upper edge points of the iliac crest, define and calculate the coronal symmetry axis of the human back; Calculate the angle between the axis of symmetry and the vertical direction of the RGB image of the upright back , if the angle is less than or equal to a preset threshold, no rotation correction is performed and the process directly enters the subsequent depth estimation and 3D evaluation stages. If the angle exceeds the preset threshold, pose correction is performed through a rotation affine transformation matrix to adjust the pose in the RGB image of the upright back so that the iliac crest reference line is parallel to the horizontal axis of the image.
4. The three-dimensional spinal evaluation method based on RGB images and anatomical constraints according to claim 1, wherein In the said step S3, using a lightweight segmentation network to extract the regions of interest of the scapula and iliac crest landmark points includes: 1) Based on the RGB image corrected in step S2, adopt an improved lightweight U-Net++ network, combine the encoder, decoder, and multi-level output layer to achieve accurate segmentation of the regions of interest of the scapula, iliac crest, and vertebral spinous process landmark points; 2) The encoder of the lightweight U-Net++ network adopts the MobileOne-S3 module, and its block structure is: Branch 1 sequentially includes: 3×3 depthwise separable convolution, batch normalization, SiLU activation; Branch 2 sequentially includes: 1×1 pointwise convolution, batch normalization, SiLU activation, 3×3 depthwise separable convolution; 3) The decoder of the lightweight U-Net++ network fuses the dynamic upsampling module to restore the image resolution by means of learnable parameter interpolation; 4) The output layer of the lightweight U-Net++ network predicts multiple region-of-interest masks in parallel, including: Representing the acromion region; Representing the posterior superior iliac spine region; Representing the continuous line of spinal spinous processes; 5) The lightweight U-Net++ network uses the Dice loss function to perform pixel-level optimization on the regions of interest in the scapula and iliac crest areas; 6) The lightweight U-Net++ network shares the backbone feature extraction layer with the key anatomical feature point detection network HRNet-W32 used for iliac crest baseline correction described in step S2.
5. The three-dimensional spinal evaluation method based on RGB images and anatomical constraints according to claim 1, characterized in that Step S4 includes: 1) Input the region-of-interest mask map output by step S3 and the RGB image corrected in S2 into the dual-branch depth estimation network structure; 2) The dual-branch depth estimation network: Branch 1: Extract a relative depth map from the RGB image corrected in step S2 based on the MiDaS encoder-decoder structure ; Branch 2: Introduce a vertebral attention mechanism after the 4th layer of the decoder in Branch 1. By enhancing the depth prediction weights of the scapula, iliac crest, and vertebral spinous process areas, ensure that the network is more sensitive to the learning of key areas of the spine and improve the accuracy of depth estimation; Branch 1 and Branch 2 are fused through channel concatenation: , Wherein: is the feature map output by Branch 1, is the feature map output by Branch 2, is the channel concatenation operation, is the 1×1 convolution of the number of channels, is the fused feature, which is optimized by the multi-constraint loss function to obtain the initial depth map ; 3) The vertebral attention mechanism: Input: , Wherein: are respectively the height, width and number of channels of the feature map, is the feature map output by the fourth layer of the decoder of Branch 1; Anatomical position encoding: Generate anatomical position encoding using the positions of the scapula, iliac crest, and vertebral spinous process areas and input it as an anatomical prior into the vertebral attention mechanism; Vertebral attention mechanism generates vertebral probability distribution map : , , Wherein: represents the pixel position in the feature map, corresponding to the height direction, corresponding to the width direction; represents the th vertebral index; is the standard projection position of the th vertebra; is the vertebral size correlation parameter; is the segmentation confidence function, which controls the weight attenuation of low-confidence vertebrae, is the confidence representing the th vertebra; Attention calculation of the vertebral attention mechanism: , , , , Wherein: , , is a linear transformation matrix; is a channel number normalization factor to avoid the problem of gradient vanishing or explosion, is the feature map output by the 4th layer of the decoder of Branch 1; is the vertebral body probability distribution map generated by the vertebral attention mechanism, is a normalization function that converts the attention weights into a probability distribution; Output: Generate an enhanced feature map: , Wherein: is a layer normalization operation.
6. The three-dimensional spinal evaluation method based on RGB images and anatomical constraints according to claim 5, wherein In step S4, the multi-constraint loss function is: , Where: : Depth estimation error loss, which is used to minimize the difference between the predicted depth and the true depth; : Statistical shape model SSM projection loss to ensure that depth estimation conforms to the vertebral statistical model; : Loss of left-right symmetry, restricting the symmetry of the spinal column's shape; : Key point loss, used to optimize the positions of key points such as the scapula, iliac crest, and spinal spinous process; : is the weight coefficient; Among them: depth estimation error loss : using scale-invariant logarithmic error: , Wherein: , are the height and width of the corrected RGB image, respectively, and the total number of pixels : is the depth value of the ground truth depth map at pixel ; is the depth value of the predicted depth map at pixel ; represents two different sets of pixel coordinates in the image, which are used to traverse all possible pixel pairs; is the balance factor to optimize the local depth relationship constraint; SSM Projection Loss : Enforcing the projection constraint of the depth map to satisfy the statistical shape model of the frustum , Wherein: is the total number of vertebra categories; is the three-dimensional coordinates of the th vertebra in the statistical shape model SSM; is the perspective projection operation; is the segmentation result of the th vertebra predicted by the network, is the absolute depth map; Loss of left-right hemispheric symmetry : Penalize left-right hemispheric depth asymmetry: , Wherein: is the depth value of the pixels in the left half of the corrected RGB image; is the depth value at its mirror-symmetric position; , are the height and width of the image respectively; is the threshold of symmetry loss; is the weighting factor based on different regions of the spine for dynamically adjusting symmetry loss; Key point loss : , Wherein: is the number of key points; are the coordinates of the actual key points of the scapula, iliac crest, and spinal spinous process; are the key point coordinates predicted by the network.
7. The three-dimensional spinal evaluation method based on RGB images and anatomical constraints according to claim 6, wherein Step S5 includes: 1) Height-vertebral size mapping Based on a statistical model, using the height of the subject to predict the vertebral body height , and establish the regression relationship between the vertebral body height and the height: , Wherein: is the predicted vertebral body height; is the subject's height; , is the regression coefficient, set according to statistical data; 2) Adaptive height scale calibration calculation To calibrate the depth estimation scale, a height scale calibration factor is calculated : , vertebral body weight , Calculate the depth map calibrated by the height scale according to the height scale calibration factor : , Wherein: is the ordinate of the th vertebral body in the upright back RGB image, and the spinal midline is , is the vertebral body size correlation parameter, the vertebral body height predicted based on the regression model; the standard vertebral body height in the statistical shape model SSM; is the number of visible vertebral bodies; is the initial depth map processed through step S4; 3) BMI compensation mechanism Add fat compensation to the depth after height scale calibration: , Wherein: is the depth map after calibration of height and weight scale; is the depth map after calibration of height scale; is the fat compensation coefficient; is the longitudinal coordinate of the pixel point in the RGB image of the standing back; is the average value of the vertical coordinates of the iliac crest ROI; is the Gaussian distribution parameter for controlling the fat compensation range.
8. The three-dimensional spine assessment method based on RGB images and anatomical constraints according to claim 7, characterized in that, Step S6 includes: Input the depth map calibrated by the height and weight scale in step S5 into the two-stage generator in the cascaded generative adversarial network architecture , , eliminate clothing interference and achieve 3D surface reconstruction of the back, including: 1) First stage: Eliminate clothing interference In the first stage, a generator is constructed through partial convolution , which is used to restore the skin texture of the clothing-occluded area in the image. The generator extracts the features of the input image and reconstructs the occluded area through an encoder-decoder architecture; among them, the encoder is used to extract the features in the input image, identify the occluded area and its boundary, and the decoder is used to restore the occluded skin texture information according to the output of the encoder; 2) Second stage: Back three-dimensional surface reconstruction In the second stage, the generator combines anatomical priors and biomechanical modeling to correct the morphology of the skin surface. By modeling the costal arch, the costal arch curve is calculated based on the subject's height and BMI to ensure that the generated skin surface meets the requirements of human anatomical structures; Erector spinae muscle stress simulation: Simulate the effects of muscle activity and stress distribution on the skin surface morphology, and combine the muscle elasticity and stress distribution model to correct the skin surface to make it conform to the rationality of the physiological structure; 3) Multi-scale discrimination mechanism Improve the authenticity of the generated skin through a multi-scale discriminator to ensure that it matches the anatomical curvature of the real skin, including: Global discriminator : Input a 512×512 image and evaluate the overall authenticity; Local discriminator : Identify the clothing seam details for a 128×128 local area; Physiological Rationality Discriminator : Calculate the KL divergence between the generated skin curvature and the real anatomical database to evaluate the physiological rationality of the generated image.
9. The three-dimensional spinal evaluation method based on RGB images and anatomical constraints according to claim 1, characterized in that Step S7 includes: 1) NURBS spine curve fitting: Based on the back three-dimensional surface reconstruction data in step S6, construct a NURBS curve to fit the centerline of the human spine; 2) Parameter adjustment strategy: Optimize the knot vector of NURBS , to adapt to the individual spinal anatomy: 3) Weight assignment strategy: Dynamically adjust the influence of control points according to the spine morphology to improve the fitting accuracy and ensure that the curve can match the actual anatomical structure in key anatomical areas; 4) Dynamic constraint optimization: Optimize the curve smoothness, anatomical consistency, and data fitting degree to make the fitting curve conform to the individual physiological structure while minimizing the error.
10. The three-dimensional spinal evaluation method based on RGB images and anatomical constraints according to claim 1, characterized in that, Step S8 includes: 1) Cobb angle evaluation: Based on the spine curve fitting result, automatically detect the end vertebrae and calculate the Cobb angle to evaluate the severity of scoliosis; 2) Trunk rotation angle evaluation: Based on the back surface point cloud, calculate the maximum height difference between symmetric positions on both sides of the spine to evaluate the torsional deformity caused by scoliosis; 3) Sagittal plane curvature evaluation: Calculate the curvature of the thoracolumbar segment of the spine to evaluate the lordosis and kyphosis of the spine.
11. A three-dimensional spinal evaluation system based on RGB images and anatomical constraints, characterized in that Includes: Image acquisition module: Used to obtain the standardized upright back RGB image of the subject through an intelligent mobile device and upload the image data to the cloud computing server for processing; Image preprocessing and pose correction module, which is used to enhance and denoise the acquired upright back RGB image, and adjust the shooting pose error based on anatomical landmark points to ensure the standardized input of the spinal region; Anatomical landmark detection module, which is used to automatically detect the anatomical landmark points of the scapula and iliac crest based on a deep learning model, and extract the region of interest (ROI) of the spine to provide a basis for subsequent modeling; Monocular depth estimation module, which is used to generate an initial depth map through a depth estimation model that fuses the vertebral attention mechanism and a multi-constraint loss function. The multi-constraint loss function includes the projection constraint of the statistical shape model (SSM); Scale calibration module, which is used to perform scale correction on the initial depth map based on the physiological parameters of the subject's height, weight, and BMI to ensure that spinal measurements conform to the individual anatomical proportion; Clothing interference elimination and skin surface evaluation module, which is used to remove the interference of clothing occlusion based on a generative adversarial network and restore the true three-dimensional shape of the back skin; Spinal three-dimensional evaluation module, which is used to generate an accurate spinal three-dimensional curve model through a curve fitting method based on the three-dimensional shape data of the back skin and adapt to different degrees of spinal shape changes; Evaluation parameter calculation module, which is used to calculate and output the coronal Cobb angle, trunk rotation angle, and sagittal curvature parameters based on the generated spinal three-dimensional curve model; System output module, which is used to return the analysis results to the intelligent mobile device for display in a visual way and provide an automated diagnostic report containing spinal deformity evaluation data for medical professionals or individual users to refer to.
Citation Information
Patent Citations
Scoliosis cobb angle measurement method based on three-dimensional image and multi-layer perception
CN114081471A
Scoliosis screening method based on 2D RGB image
CN115526845A
Spine 3D modeling method and system
CN115908717A
Spine three-dimensional structure reconstruction method based on deep learning
CN116402954A
Spine three-dimensional dynamic reconstruction method and system
CN116524124A
Cited By
Intelligent teenager scoliosis monitoring system based on multi-modal sensing
CN121059146A
Teenager idiopathic scoliosis rehabilitation evaluation method and system based on three-dimensional spine dynamic test method
CN121549773A
Bone evaluation method based on neural network enhancement and finite element analysis
CN121616708A
Cervical vertebra MRI image automatic segmentation method and system based on attention mechanism
CN122199581A
Femoral cortical bone parameter quantitative analysis method and system based on three-dimensional medical image
CN122265273A