Rapid Gaussian splashing human body surface grid reconstruction method for scoliosis diagnosis and treatment

By introducing a reference object method or height method into 3DGS technology for size calibration, and combining it with video acquisition from consumer-grade devices, the problem of insufficient size accuracy in existing technologies is solved, generating a high-precision and convenient three-dimensional human body model that is suitable for medical applications such as scoliosis diagnosis and treatment.

CN121937664APending Publication Date: 2026-04-28SHANGHAI ZHIHUI MEDICAL TECHNOLOGY CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI ZHIHUI MEDICAL TECHNOLOGY CO LTD
Filing Date
2025-12-29
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing methods for human body surface reconstruction based on 3D Gaussian sputtering (3DGS) are convenient to operate and inexpensive, but the reconstruction results lack real dimensional references and cannot meet the strict requirements for precise geometric parameters in medical applications such as scoliosis diagnosis and treatment. Although professional 3D scanning equipment has dimensional accuracy, it is difficult to popularize due to its high cost and complex operation.

Method used

By capturing video using consumer devices such as mobile phones and combining it with a specific size calibration mechanism, a high-precision 3D human body model with real physical dimensions is generated using the reference object method or height method for size calibration. This includes camera parameter calculation, multi-constraint composite loss function optimization, TSDF fusion algorithm, and size transformation steps.

Benefits of technology

While maintaining convenience, the generated 3D human body model is highly consistent with the real human body in terms of geometry and size, and can output reliable key parameters such as trunk circumference and limb length, meeting the needs of clinical diagnosis and orthopedic brace design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937664A_ABST
    Figure CN121937664A_ABST
Patent Text Reader

Abstract

The invention discloses a rapid Gaussian splashing human body surface grid reconstruction method for scoliosis diagnosis and treatment. The method comprises the steps that a static human body surrounding video is acquired, a plurality of key frames are extracted, and scene sparse point clouds corresponding to the processed key frames and camera parameters corresponding to the frames are calculated based on a re-projection error of a minimized three-dimensional point; analyzing a scene sparse point cloud and camera parameters based on a pre-trained large model, generating an initial Gaussian distribution point set, obtaining an initialized surface primitive, generating a rendered image, a depth map and a normal graph of each view angle based on a differentiable rendering pipeline, and optimizing surface primitive parameters by using a multi-constraint composite loss function; integrating the depth map corresponding to the optimized primitive set with camera parameters to generate a human body surface triangular mesh; and calculating a scaling scale of the virtual size and the real size by adopting a reference object method or a height method, and applying the scaling scale to the triangular mesh on the surface of the human body to obtain a scoliosis patient body three-dimensional model with the real physical size.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human body three-dimensional modeling technology, specifically to a rapid Gaussian splashing method for reconstructing human body surface meshes for the diagnosis and treatment of scoliosis. Background Technology

[0002] Accurate 3D reconstruction of scenes and high-quality synthesis of new views are core tasks of computer vision and graphics, with broad application prospects in fields such as virtual reality, autonomous driving, and healthcare. In healthcare scenarios, accurate 3D surface models of the human body can provide precise quantitative data references for disease tracking and management during the treatment of scoliosis and for the production of "spinal orthotics (braces)".

[0003] In recent years, the emergence of Neural Radiation Field (NeRF) technology has greatly promoted the development of 3D reconstruction. By representing a scene as a continuous radiation field, it has achieved photorealistic rendering effects for the first time. However, NeRF relies on large multilayer perceptrons (MLPs), which incur huge computational overhead and have a lengthy training process, making it difficult to meet the real-time requirements of medical applications. Subsequent improvements, such as InstantNGP, have significantly shortened the training time through techniques such as hash grids, but their core focus remains on view synthesis rather than accurate geometric surface reconstruction.

[0004] To obtain accurate geometric surfaces, neural surface reconstruction methods have emerged. These methods (such as Neuus and Neuralangelo) recover the 3D surface of a scene while achieving high-quality rendering through techniques like parametric symbolic distance functions (SDF). Although these methods have made significant progress in reconstruction accuracy, the problem of excessive time consumption for single-scene reconstruction still exists, limiting their application in large-scale human modeling and rapid clinical assessment.

[0005] Recently, 3D Gaussian Splatting (3DGS) technology has brought about a new breakthrough. This technology represents a scene as tens of thousands of anisotropic 3D Gaussian primitives and initializes them using Structure-from-Motion (SfM), thereby greatly shortening training and rendering time and enabling real-time interaction. However, the Gaussian point cloud distribution generated by the basic 3DGS method is relatively messy, making it difficult to form a geometrically consistent and smooth surface. This limits its value in applications requiring accurate surface information, such as scoliosis screening and assessment, posture modeling, and orthopedic brace design. Although subsequent improvements (such as SuGaR and GOF) have improved surface geometry quality by introducing regularization terms, they generally suffer from a fatal flaw: insufficient dimensional accuracy of the reconstructed model. Key parameters such as trunk circumference and limb length deviate significantly from the real human body. In the context of scoliosis diagnosis and treatment, distorted dimensional data may lead to improper brace customization or incorrect assessment of treatment effects, and may even worsen the condition.

[0006] In contrast, while professional 3D scanning equipment (such as laser scanners and structured light scanners) can provide high-precision human body size data and provide reliable support for clinical diagnosis, such equipment is usually expensive, bulky and complex to operate, making it difficult to popularize outside of professional medical institutions.

[0007] Therefore, existing technologies present a clear contradiction: on the one hand, 3DGS-based methods offer the convenience, low cost, and high versatility of reconstruction via multi-view photography using mobile phones, but lack the necessary dimensional accuracy; on the other hand, specialized scanning equipment provides the dimensional accuracy required for clinical applications, but lacks convenience and versatility. The core challenge in promoting the widespread application of this technology in the medical and health field (especially in the diagnosis and treatment of scoliosis) lies in how to integrate the advantages of both—that is, maintaining the convenience of 3DGS while ensuring that the reconstructed human body surface has realistic and accurate dimensions.

[0008] While existing methods for human surface reconstruction based on 3D Gaussian sputtering (3DGS) are convenient and inexpensive, the reconstruction results lack realistic dimensional references and cannot guarantee dimensional accuracy. Consequently, they fail to meet the stringent requirements for precise geometric parameters in medical applications such as scoliosis diagnosis and treatment. Meanwhile, specialized 3D scanning equipment capable of providing precise dimensions is difficult to popularize due to its high cost and complex operation. Summary of the Invention

[0009] The purpose of this invention is to propose a rapid Gaussian splashing method for reconstructing human body surface meshes for the diagnosis and treatment of scoliosis, comprising the following steps: S1, Based on the video recording device, video is recorded around the human target to be reconstructed to obtain a static human body surround video; S2, perform frame sampling on the static human body surround video, extract several key frames, process the key frames, and calculate the scene sparse point cloud and camera parameters corresponding to each key frame after processing based on minimizing the reprojection error of the three-dimensional points. S3 analyzes the sparse point cloud and camera parameters of the scene based on the pre-trained large model, generates an initial Gaussian distribution point set, obtains the initial surface primitives, generates the rendered images, depth maps and normal maps of each viewpoint based on the differentiable rendering pipeline, and optimizes the surface primitive parameters using the multi-constraint composite loss function to obtain the optimized primitive set that fits the human body surface. S4, the depth map corresponding to the optimized primitive set and the camera parameters are integrated by the TSDF fusion algorithm with truncated signed distance function to extract the zero level set and generate a closed topological triangular mesh of the human body surface. S5. Calculate the scaling ratio between the virtual and real dimensions using the reference method or height method, and apply the scaling ratio to the triangular mesh on the human body surface to obtain a three-dimensional body model of a scoliosis patient with real physical dimensions.

[0010] In a preferred embodiment of the present invention, the calculation of camera parameters and sparse point cloud of the scene in step S2 specifically includes: Based on the pinhole camera model, feature point matching is performed on keyframes to construct the projection relationship between a 3D point X∈ℝ³ and the pixel x on the image plane, as shown in the following formula: x̂ = K(RX + t); Where K is the camera intrinsic parameter, and R and t are the rotation matrix and translation vector of the camera extrinsic parameter, respectively; The projection error for the actual detected pixel x is calculated based on the Bundle Adjustment optimization algorithm. The objective function formula is as follows: ; The optimal camera parameter set and sparse point cloud are calculated based on minimizing the projection error.

[0011] In a preferred embodiment of the present invention, the multi-constraint composite loss function includes photometric loss, depth concentration loss, and normal smoothing loss, as detailed below; Photometric loss is used to constrain the difference between the rendered RGB image and the actual input video frame. The photometric loss function is as follows: ; The depth concentration loss causes the depth distribution of Gaussian pixels to concentrate towards the real surface. For the ray of pixel p, the depth concentration loss function is: in It is a mixed weight, and z is the Gaussian depth; The normal smoothing loss is used to align the surface normal of primitives with the surface normal calculated from the gradient of the rendering depth map. Let the normal of primitive i be ni, and the neighborhood set be N(i). The normal smoothing loss function is: .

[0012] In a preferred embodiment of the present invention, the differentiable rendering calculation process is as follows: Differentiable rendering employs a volume rendering model based on ray integrals. For pixel p, the corresponding camera ray r(t) = o + td. All Gaussian elements intersecting the ray are sorted by depth. Color accumulation uses differentiable alpha blending. ; o is the camera origin, d is the ray direction, t is the ray parameter variable, i.e., the distance along the ray direction; C(p) is the color accumulation value. Gauss color value, , Gauss Contribution at pixel p For Gauss Opacity Gauss The probability density of the Gaussian distribution at pixel p.

[0013] In a preferred embodiment of the present invention, the truncated TSDF fusion algorithm in step S4 is integrated as follows: TSDF fusion assigns a sign function with truncation distance to multi-view depth maps in units of voxel x: Where d(x) is the signed distance from the voxel center to the depth surface, and μ is the truncation threshold. Weighted averaging is used when fusing multiple viewpoints. Extracting the zero level set As the reconstructed surface, a triangular mesh of the human body surface with a closed topology is obtained.

[0014] In a preferred embodiment of the present invention, the reference object method in step S5 specifically includes: The pre-trained image segmentation model processes all sampled video frames to identify and segment the pixel region where the reference object is located, which is a rectangular card. Record the two-dimensional pixel coordinates of the two vertices of the diagonal of the rectangular region of the reference object; Based on the camera parameters and corresponding depth map of the video frame, two two-dimensional pixels are converted into three-dimensional coordinate points in the reconstructed three-dimensional space through back projection calculation. Calculate the Euclidean distance between two 3D coordinate points. The distance is the length of the diagonal of the reference object in the virtual size space. Compare the virtual length with the actual physical length of the diagonal of the reference object to calculate the scaling scale from the virtual size to the actual size. By applying the scale bar to all vertices of the human body surface mesh and performing size transformation, a human body model with the correct physical dimensions is obtained.

[0015] In a preferred embodiment of the present invention, the back projection calculation method is as follows: Back projection uses depth D(u,v) to map pixel (u,v) back to 3D space, as shown in the following formula: The absolute position of the point in the reconstructed coordinate system is obtained based on the camera's extrinsic parameters; The two pixel vertices of the reference object are back-projected to obtain X1 and X2 respectively, and the Euclidean distance is ||X1−X2||2, which is the diagonal of the virtual space reference object.

[0016] In a preferred embodiment of the present invention, the height method in step S5 specifically includes: Record the actual height of the person being photographed as a benchmark for their true dimensions; Principal component analysis is applied to the coordinates of all vertices of the virtual human body surface mesh to find the principal direction with the largest variance. Project all vertices onto the principal direction with the largest variance. The difference between the maximum and minimum values ​​after projection is the virtual height of the current model. By comparing the calculated virtual height with the actual height of the person being photographed, a scale for size transformation can be obtained; Apply the scale bar to all vertices of the mesh to obtain a human body surface mesh with the correct dimensions.

[0017] In a preferred embodiment of the present invention, the virtual height calculation method is as follows: PCA is used to find the principal orientation of a virtual human body mesh. Let the set of all vertices be . The covariance matrix is: in Given the mean of the vertex coordinates, solve for the eigenvalue decomposition Cv = λv, where v is the eigenvector to be solved for, and C is the covariance matrix. The eigenvalues ​​to be solved are the eigenvectors. The largest eigenvalue corresponds to the eigenvector v1, which is the height direction. Project all points onto the v1 direction, and take the maximum difference between each pair of projected values ​​to obtain the virtual height.

[0018] Due to the application of the above technical solution, the present invention has the following advantages compared with the prior art: Building upon convenient video capture via consumer-grade devices like smartphones, this invention introduces a specific size calibration mechanism to ensure that the final generated 3D human body surface mesh model not only maintains high geometric fidelity but also adheres to the dimensions of a real human body. This allows for the conversion of videos captured by ordinary mobile devices into high-precision 3D human body models with true physical dimensions. This invention eliminates the need for expensive and bulky professional 3D scanning equipment; data acquisition can be completed using only everyday consumer-grade recording devices such as smartphones. This significantly lowers the barrier to entry for the technology, breaking the limitation that professional equipment is confined to medical institutions.

[0019] Easy to operate and user-friendly: The data collection process is simplified to recording a video around the human body. The operation is simple and intuitive, requiring no professional technical skills from the user. This convenience allows users to perform self-modeling anytime, anywhere, facilitating long-term, high-frequency tracking of changes in body shape.

[0020] High dimensional accuracy, meeting clinical needs: The core advantage of this invention lies in its dimensional calibration using a reference object method or height method, which solves the problem of insufficient dimensional accuracy commonly found in existing rapid reconstruction techniques based on 3DGS. This ensures that the final generated mesh model has physical dimensions consistent with the real human body, and can output reliable key geometric parameters such as trunk circumference and limb length. This provides crucial data support for the quantitative diagnosis of scoliosis, the design of personalized orthotic braces, and the assessment of rehabilitation progress. Attached Figure Description

[0021] Figure 1 This diagram illustrates the overall process of size calibration using the reference object method or height method provided in an embodiment of the present invention. Figure 2 This diagram illustrates the principle of using principal component analysis (PCA) to calculate the height of a model, as provided in an embodiment of the present invention. Detailed Implementation

[0022] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0023] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0024] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0025] like Figure 1 As shown, Figure 1 This invention demonstrates the complete steps from video acquisition to the final generation of a realistically sized mesh, including camera parameter calculation, 3DGS 3D reconstruction, reference object segmentation and size calculation, principal component analysis (PCA) to calculate virtual height, scale generation, and size transformation. This invention also provides a rapid Gaussian splashing method for human surface mesh reconstruction for scoliosis diagnosis and treatment, comprising the following steps: S1, Based on the video recording device, video is recorded around the human target to be reconstructed to obtain a static human body surround video; Specifically, the collected videos must adhere to the following specifications: (1) Human posture and environmental requirements: Stillness of posture: Throughout the entire shooting process, the target to be reconstructed must remain still to avoid the impact of motion blur on the accuracy of the model.

[0026] Lighting conditions: The shooting environment should be bright and evenly lit, avoiding overexposure, strong shadows or direct light on the lens, to ensure that clear surface texture details are captured.

[0027] (2) Camera movement and framing requirements: Complete surround: The camera must make at least three smooth, slow circular movements around the human body. Multiple shots provide richer perspective information, helping to improve the robustness of camera parameter calculation and the integrity of reconstruction.

[0028] Centering the subject: During the shooting process, the human subject should always be fully within the camera's field of view and kept as centered as possible in the frame.

[0029] Include a top-down view: In a surround shot, the top of the human head must be captured completely at least once around the perimeter. This view is crucial for subsequent size calibration using the height method.

[0030] Video resolution: To ensure clarity of details, it is recommended that the video resolution be no less than 1920x1080 pixels.

[0031] (3) Requirements for placing reference objects: If the "reference object method" is subsequently used for size transformation, a reference object with known dimensions needs to be introduced during the shooting. This reference object should be a rectangular hard card that is not easily deformed (such as an ID card, bank card, etc.).

[0032] When filming, the reference object should be flatly pasted on the front of the human torso (such as the chest and abdomen), and it should be ensured that it remains still and clearly visible throughout the entire video recording process, especially appearing completely from the first frame of the video.

[0033] S2, perform frame sampling on the static human body surround video, extract several key frames, process the key frames, and calculate the scene sparse point cloud and camera parameters corresponding to each key frame after processing based on minimizing the reprojection error of the three-dimensional points.

[0034] Specifically, frame sampling is performed from the input surround video to extract several key frames. It is recommended to sample about 120 frames to balance computational efficiency and information density.

[0035] Subsequently, these sampled frames were processed using the classic Structure-from-Motion (SfM) technique. The SfM algorithm simplifies the camera model to a pinhole camera model. By matching feature points between consecutive frames, it can simultaneously estimate the sparse 3D point cloud of the scene and the camera parameters corresponding to each frame, specifically including the camera's intrinsic parameters (focal length, optical center) and extrinsic parameters (position coordinates and rotation attitude in 3D space).

[0036] S3 analyzes the sparse point cloud and camera parameters of the scene based on the pre-trained large model, generates an initial Gaussian distribution point set, obtains the initial surface primitives, generates the rendered images, depth maps and normal maps of each viewpoint based on the differentiable rendering pipeline, and optimizes the surface primitive parameters using the multi-constraint composite loss function to obtain the optimized primitive set that fits the human body surface. Specifically, the 3D reconstruction step based on surface primitives is the core of the technical solution. By optimizing a set of 3D Gaussian surface primitives to accurately fit the human body surface, a high-fidelity geometric model is finally generated.

[0037] (1) Initialization: The sparse point cloud output by the SfM algorithm is used as the initial 3D position of the Gaussian surface primitives. At the same time, based on the projection position of the point cloud in the original video frame, the color information of the corresponding pixel is extracted, and each primitive is given initial appearance attributes (such as color, opacity, etc.).

[0038] To improve initialization quality, a pre-trained large model can be introduced to generate an initial Gaussian distributed point set G0, which is then aligned with the sparse point cloud P output by the SfM model, resulting in a more thorough initialization. This initialization reduces the number of reconstruction iterations, thereby improving reconstruction efficiency.

[0039] (2) Differentiable Rendering: A differentiable rendering pipeline is constructed from 3D primitives to a 2D image. For any viewpoint (i.e., any frame of video), a ray is emitted from the camera center of that viewpoint, passing through every pixel on the image plane. The rendering pipeline accurately calculates the intersection of this ray with all surface primitives in the scene and sorts these intersections according to their depth values. To ensure that the human outline remains clear under various viewpoints, a low-pass filtering mechanism is introduced during the rendering process to avoid artifacts caused by shape degradation of primitives under tilted viewpoints. Finally, through the alpha blending algorithm, the color and opacity of all primitives intersecting with the ray are depth-weighted and blended to generate the rendered RGB image, depth map, and normal map of that viewpoint on the image plane.

[0040] In the computational process, differentiable rendering employs a volume rendering model based on ray integrals. For pixel p, the corresponding camera ray r(t) = o + td, and all Gaussian elements intersecting the ray are sorted by depth. Color accumulation utilizes differentiable alphablending. ;

[0041] C(p) is the cumulative color value. Gauss color value, , Gauss Contribution at pixel p For Gauss Opacity Gauss The Gaussian probability density at pixel p is given by this formula, which ensures that the rendering result can be gradient-calculated with respect to the Gaussian parameters, thus supporting subsequent optimization.

[0042] (3) Multi-constraint optimization: To ensure that the reconstruction results achieve high accuracy in both appearance and geometry, a multi-constraint composite loss function is defined for optimization: Photometric loss: Used to constrain the difference between the rendered RGB image and the real input video frame. This loss combines L1 loss and structural similarity (SSIM) loss, ensuring pixel-level color consistency while maintaining the overall structure and texture similarity of the image, thus ensuring the realistic appearance of the reconstructed model.

[0043] Photometric loss is calculated using a combination loss formula, as follows: ;

[0044] The L1 term constrains pixel-by-pixel color differences, while SSIM maintains the similarity of overall structure, brightness, and texture, ensuring that the reconstructed surface appearance remains realistic and consistent. photo For the imaging loss function that needs to be optimized , These are hyperparameters used to balance two terms in the function, where I is the rendering frame. These are real frames.

[0045] Depth Concentration Loss: This loss function constrains the distribution of primitives along the direction of each rendering ray, causing primitives to concentrate and shrink towards the real human body surface location, avoiding the primitives from being diffusely distributed in space, thus forming a more compact and clear geometric surface.

[0046] The depth concentration loss aims to concentrate the depth distribution of Gaussian elements towards the real surface. For a ray of pixel p, the definition is: in, It is a depth-focused loss, for two Gaussian elements G on a ray from the camera to the pixel. i and G j , These are their respective mixed weights, which are numerically equal to their contribution. , These correspond to their depth values ​​in the camera coordinate system. This loss drives the primitives to shrink to a uniform surface, rather than being diffusely distributed in space.

[0047] Normal smoothing loss: Used to align the surface normals of primitives with the surface normals calculated from the gradient of the rendered depth map. This constraint ensures the local continuity and smoothness of the reconstructed surface normals, avoiding irregular bumps or noise, and making the transition of human surfaces (especially skin) more natural.

[0048] Normal smoothing is achieved by constraining the normal consistency of locally adjacent primitives. Let the normal of primitive i be n. i Let the neighborhood set be N(i), defined as follows: For normal smoothing loss function, , This represents the normal vector of the Gaussian element. This constraint effectively smooths the local geometry, ensuring that the final mesh is free from jitter and noise.

[0049] (4) Iterative Updates and Dynamic Adjustments: Using the backpropagation algorithm and optimizers such as Adam, the parameters (position, rotation, scale, color, and opacity) of each surface primitive are iteratively updated based on the gradient of the composite loss function. Simultaneously, the number of primitives is dynamically adjusted during optimization: primitives with low contribution to rendering (such as those with opacity below a certain threshold) are periodically removed to reduce computational redundancy; and primitives are adaptively added in areas where the model deviates significantly from the real image to enhance the ability to capture local details. This iterative process continues until the loss function converges, ultimately yielding an optimized set of surface primitives that accurately fits the real human body surface.

[0050] S4 integrates the depth map corresponding to the optimized primitive set with the camera parameters through the TSDF fusion algorithm with truncated signed distance function, extracts the zero level set, and generates a closed topological triangular mesh of the human body surface.

[0051] Specifically, after obtaining the optimized set of surface primitives, this step converts it into a standard three-dimensional mesh format to facilitate subsequent measurements and applications.

[0052] Using the established differentiable rendering framework, a high-precision depth map is generated for each sampled video frame. Then, a truncated signed distance function (TSDF) fusion algorithm is employed to integrate the depth maps from all viewpoints along with their corresponding camera parameters. The TSDF algorithm constructs a voxel mesh in 3D space and projects the information from each depth map into this mesh. By weighted averaging of information from multiple viewpoints, a zero-level set representing the human body surface is extracted, thus generating a triangular mesh of the human body surface with a complete and closed topology. The resulting mesh is geometrically accurate, but its size is relative and unitless virtual.

[0053] TSDF fusion assigns a sign function with truncation distance to multi-view depth maps in units of voxel x: This is the truncation function, where the last two terms represent the upper and lower bounds of the function. When the value of the first term exceeds the upper or lower bound, the value of the upper or lower bound is taken, respectively. Here, d(x) is the signed distance from the voxel center to the depth surface, and μ is the truncation threshold. Weighted averaging is used when fusing multiple viewpoints. The weighted signed distance function represents the multi-view fusion. This represents the weight of each viewpoint. The signed distance function for each viewpoint, where i represents the viewpoint, extracts the zero-level set. As the reconstructed surface, a triangular mesh of the human body surface with a closed topology is obtained.

[0054] S5 uses the reference method or height method to calculate the scaling ratio between the virtual size and the real size, and applies the scale to the triangular mesh on the human body surface to obtain a three-dimensional body model of a scoliosis patient with real physical size.

[0055] It should be noted that size transformation transforms the generated virtual size mesh into a correct size mesh with real physical units (such as millimeters) by introducing real-world scale information.

[0056] According to an embodiment of the present invention, the calculation of camera parameters and sparse point cloud of the scene in step S2 specifically includes: Based on the pinhole camera model, feature point matching is performed on keyframes to construct the projection relationship between a 3D point X∈ℝ³ and the pixel x on the image plane, as shown in the following formula: x̂ = K(RX + t); Where K is the camera intrinsic parameter, and R and t are the rotation matrix and translation vector of the camera extrinsic parameter, respectively; The projection error for the actual detected pixel x is calculated based on the Bundle Adjustment optimization algorithm. The objective function formula is as follows: ; This represents the projection error function, for each viewpoint i, This represents the coordinates of a 3D point X in a 2D image coordinate system. and Let represent the camera extrinsic rotation matrix and translation vector for this viewpoint, respectively. The optimal camera parameter set and sparse point cloud are calculated based on minimizing the projection error.

[0057] According to an embodiment of the present invention, the reference method in step S5 specifically includes: The pre-trained image segmentation model processes all sampled video frames to identify and segment the pixel region where the reference object is located. The reference object is a rectangular card. Record the two-dimensional pixel coordinates of the two vertices of the diagonal of the rectangular region of the reference object; Based on the camera parameters and corresponding depth map of the video frame, two two-dimensional pixels are converted into three-dimensional coordinate points in the reconstructed three-dimensional space through back projection calculation. Calculate the Euclidean distance between two 3D coordinate points. The distance is the length of the diagonal of the reference object in the virtual size space. Compare the virtual length with the actual physical length of the diagonal of the reference object to calculate the scaling scale from the virtual size to the actual size.

[0058] By applying the scale bar to all vertices of the human body surface mesh and performing size transformation, a human body model with the correct physical dimensions is obtained.

[0059] According to an embodiment of the present invention, the back projection calculation method is as follows: Back projection uses depth D(u,v) to map pixel (u,v) back to 3D space, as shown in the following formula: The absolute position of the point in the reconstructed coordinate system is obtained based on the camera's extrinsic parameters; The two pixel vertices of the reference object are back-projected to obtain X1 and X2 respectively, and the Euclidean distance is ||X1−X2||2, which is the diagonal of the virtual space reference object.

[0060] According to an embodiment of the present invention, the height method in step S5 specifically includes: Record the actual height of the person being photographed as a benchmark for their true dimensions; Principal component analysis is applied to the coordinates of all vertices of the virtual human body surface mesh to find the principal direction with the largest variance. Project all vertices onto the principal direction with the largest variance. The difference between the maximum and minimum values ​​after projection is the virtual height of the current model. By comparing the calculated virtual height with the actual height of the person being photographed, a scale for size transformation can be obtained; Apply the scale bar to all vertices of the mesh to obtain a human body surface mesh with the correct dimensions.

[0061] like Figure 2 As shown, Figure 2 The image shows a human body mesh model, where the red vertical line represents the first principal component direction (i.e., the height direction) calculated by PCA. The blue dot at the top of the model's head and the green dot at the bottom of the feet represent the two endpoints of the projection of all vertices in this direction, and the vertical distance between these two points is the virtual height of the model.

[0062] According to an embodiment of the present invention, the virtual height calculation method is as follows: PCA is used to find the principal orientation of a virtual human body mesh. Let the set of all vertices be . The covariance matrix is: in Given the mean of the vertex coordinates, solve for the eigenvalue decomposition Cv = λv, where v is the eigenvector to be solved for, and C is the covariance matrix. These are the eigenvalues ​​to be solved. The largest eigenvalue corresponds to the eigenvector v1, which is the height direction. Project all points onto the v1 direction, and take the maximum difference between each pair of projected values ​​to obtain the virtual height.

[0063] To verify the effectiveness and accuracy of the method of this invention, we conducted a series of experiments and compared the reconstruction results with the real model obtained by high-precision scanning. The experimental data strongly demonstrate the superiority of this method.

[0064]

[0065] Table 1 compares the reconstruction accuracy of the two size calibration methods proposed in this invention, including the reference method and the height method, in multiple cases. Accuracy evaluation metrics include chamfer distance and F1 score. A smaller chamfer distance indicates a higher surface fit between the reconstructed model and the real model; a higher F1 score indicates better model integrity and accuracy.

[0066] As shown in Table 1, both methods can achieve high-precision 3D reconstruction. In most cases, the height method exhibits superior performance, with its chamfer distance generally lower than that of the reference method. For example, in Case 8, the chamfer distance of the height method is 260.14, significantly better than the 363.40 of the reference method. This indicates that aligning height using principal component analysis (PCA) can achieve very robust and accurate global size scaling.

[0067]

[0068] As shown in Table 2, we further evaluated the precision (P), recall (R), and F1 score (F1) of the method under different voxel sizes (10 mm, 15 mm, 20 mm). The voxel size represents the allowable error range during evaluation.

[0069] Experimental results show that as the allowable error range is widened (voxel size increases from 10mm to 20mm), the P / R / F1 scores of all cases are significantly improved. For example, in Case 8, the F1 score using the height method increased from 0.4643 under the 10mm tolerance to 0.7062 under the 20mm tolerance.

[0070] This trend strongly demonstrates that the surface reconstructed by this method is geometrically highly consistent with the real surface, and the error of most surface points is within the clinically acceptable range.

[0071] In one specific embodiment of the present invention, the present invention describes a complete process of size calibration using the height method, which is mainly divided into two parts: user-side operation and server-side calculation.

[0072] User-side data collection and uploading specifically includes: The user uses a regular smartphone as the recording device.

[0073] Preliminary preparation: Users first measure and record their exact height, for example, 175.0 cm.

[0074] Video shooting: The user stands still in an environment with even lighting and a relatively simple background. An assistant holds a mobile phone and smoothly and slowly circles around the user to shoot. The shooting process must meet the following conditions: The camera rotates around the user at least three times.

[0075] The user's entire body (from head to toe) is always fully within the phone's screen and is roughly centered on the screen.

[0076] At least one of the shots must clearly capture the user's overhead view.

[0077] Set the video resolution to 1920x1080 pixels or higher.

[0078] Data Upload: After the video is taken, the user uploads the recorded video file and the previously recorded height data (175.0 cm) to the server for processing through an application or webpage.

[0079] The server-side automated 3D reconstruction and dimensional calibration are as follows: After receiving user data, the server automatically executes the following calculation process.

[0080] The camera parameter calculation is as follows: The server first samples 120 frames of still images evenly from the uploaded video.

[0081] Next, the SfM (Structure from Motion) algorithm is run to process these 120 still images. The algorithm will output the internal and external camera parameters (focal length, position, rotation, etc.) of each frame as well as a sparse 3D point cloud of the scene.

[0082] The high-precision 3D reconstruction process is as follows: Initialization: Using the sparse point cloud generated by SfM as the initial position, create a set of three-dimensional Gaussian surface primitives.

[0083] Iterative Optimization: Entering the core optimization loop. In each iteration, a rendered image and depth map are generated based on the current primitive parameters through the differentiable rendering pipeline. Then, the photometric loss, depth concentration loss, and normal smoothing loss between the rendered result and the real video frame are calculated. Finally, backpropagation and the Adam optimizer are used to update the parameters (position, shape, color, etc.) of all primitives based on the loss gradients, and primitives are dynamically added or removed to optimize details.

[0084] The loop continues until the loss function converges, at which point an optimized set of Gaussian elements that can accurately describe the human body surface is obtained.

[0085] Exporting the human body surface mesh specifically includes: Using the optimized Gaussian set, a high-precision depth map is rendered for each camera viewpoint.

[0086] Using TSDF (Signed Distance Function with Truncated Section) fusion technology, depth maps from all perspectives and their corresponding camera parameters are integrated to ultimately extract and generate a geometrically accurate and topologically complete triangular mesh of the human body surface (e.g., in STL format). At this point, the mesh size is a unitless "virtual size".

[0087] The specific dimensional transformations include: Calculating Virtual Height: The server applies PCA (Principal Component Analysis) to the coordinates of all vertices of the virtual size mesh generated in the previous step. The algorithm automatically finds the principal direction with the largest variance (i.e., the direction of the human body's height). By calculating the difference between the maximum and minimum values ​​of the projections of all vertices in this direction, the virtual height of the model is obtained, for example, "1250.0" units.

[0088] Calculate the scale: Compare the user's actual height (175.0 cm, i.e. 1750.0 mm) with the calculated virtual height (1250.0) to obtain the scale: Scale = Actual height / Virtual height = 1750.0 / 1250.0 = 1.4.

[0089] Application transformation: Multiply the 3D coordinates of all vertices of the mesh model by the scale (1.4) to complete the size transformation.

[0090] The output includes: the server saves the final human body surface mesh file with true physical dimensions. This file can now be used directly for 3D printing and design of personalized orthopedic braces. Users can download or view the model online on the client side.

[0091] Through the above embodiments, the present invention successfully transforms an ordinary mobile phone into a high-precision, remotely operable 3D human body scanning device, providing a convenient, low-cost, and dimensionally accurate solution for application scenarios such as scoliosis.

[0092] In summary, this invention successfully combines the convenience and efficiency of 3D Gaussian sputtering technology with the dimensional accuracy required for clinical applications, and its effectiveness and reliability have been rigorously verified through experiments. It not only provides a novel, low-cost, and high-precision technical tool for personalized treatment of scoliosis, but also brings broad application prospects to other medical and health fields that require precise human body modeling.

[0093] Building upon convenient video capture via consumer-grade devices like smartphones, this invention introduces a specific size calibration mechanism to ensure that the final generated 3D human body surface mesh model not only maintains high geometric fidelity but also adheres to the dimensions of a real human body. This allows for the conversion of videos captured by ordinary mobile devices into high-precision 3D human body models with true physical dimensions. This invention eliminates the need for expensive and bulky professional 3D scanning equipment; data acquisition can be completed using only everyday consumer-grade recording devices such as smartphones. This significantly lowers the barrier to entry for the technology, breaking the limitation that professional equipment is confined to medical institutions.

[0094] Easy to operate and user-friendly: The data collection process is simplified to recording a video around the human body. The operation is simple and intuitive, requiring no professional technical skills from the user. This convenience allows users to perform self-modeling anytime, anywhere, facilitating long-term, high-frequency tracking of changes in body shape.

[0095] High dimensional accuracy, meeting clinical needs: The core advantage of this invention lies in its dimensional calibration using a reference object method or height method, which solves the problem of insufficient dimensional accuracy commonly found in existing rapid reconstruction techniques based on 3DGS. This ensures that the final generated mesh model has physical dimensions consistent with the real human body, and can output reliable key geometric parameters such as trunk circumference and limb length. This provides crucial data support for the quantitative diagnosis of scoliosis, the design of personalized orthotic braces, and the assessment of rehabilitation progress.

[0096] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0097] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A rapid Gaussian splashing method for reconstructing human body surface meshes for scoliosis diagnosis and treatment, characterized in that, Includes the following steps: S1, Based on the video recording device, video is recorded around the human target to be reconstructed to obtain a static human body surround video; S2, perform frame sampling on the static human body surround video, extract several key frames, process the key frames, and calculate the scene sparse point cloud and camera parameters corresponding to each key frame after processing based on minimizing the reprojection error of the three-dimensional points. S3 analyzes the sparse point cloud and camera parameters of the scene based on the pre-trained large model, generates an initial Gaussian distribution point set, obtains the initial surface primitives, generates the rendered images, depth maps and normal maps of each viewpoint based on the differentiable rendering pipeline, and optimizes the surface primitive parameters using the multi-constraint composite loss function to obtain the optimized primitive set that fits the human body surface. S4, the depth map corresponding to the optimized primitive set and the camera parameters are integrated by the TSDF fusion algorithm with truncated signed distance function to extract the zero level set and generate a human surface triangular mesh with a closed topology. S5. Calculate the scaling ratio between the virtual and real dimensions using the reference method or height method, and apply the scaling ratio to the triangular mesh on the human body surface to obtain a three-dimensional body model of a scoliosis patient with real physical dimensions.

2. The rapid Gaussian splashing human body surface mesh reconstruction method for scoliosis diagnosis and treatment as described in claim 1, characterized in that, Step S2, which involves calculating camera parameters and sparse point clouds of the scene, specifically includes: Based on the pinhole camera model, feature point matching is performed on keyframes to construct the projection relationship between a 3D point X∈ℝ³ and the pixel x on the image plane, as shown in the following formula: x̂ = K(RX + t); Where K is the camera intrinsic parameter, and R and t are the rotation matrix and translation vector of the camera extrinsic parameter, respectively; The projection error for the actual detected pixel x is calculated based on the Bundle Adjustment optimization algorithm. The objective function formula is as follows: ; The optimal camera parameter set and sparse point cloud are calculated based on minimizing the projection error.

3. The rapid Gaussian splashing human body surface mesh reconstruction method for scoliosis diagnosis and treatment as described in claim 1, characterized in that, The multi-constraint composite loss function includes photometric loss, depth concentration loss, and normal smoothing loss, as detailed below; Photometric loss is used to constrain the difference between the rendered RGB image and the actual input video frame. The photometric loss function is as follows: ; The depth concentration loss causes the depth distribution of Gaussian pixels to concentrate towards the real surface. For the ray of pixel p, the depth concentration loss function is: in It is a mixed weight, and z is the Gaussian depth; The normal smoothing loss is used to align the surface normal of primitives with the surface normal calculated from the gradient of the rendering depth map. Let the normal of primitive i be ni, and the neighborhood set be N(i). The normal smoothing loss function is: 。 4. The rapid Gaussian splashing human body surface mesh reconstruction method for scoliosis diagnosis and treatment as described in claim 3, characterized in that, The calculation process for differentiable rendering is as follows: Differentiable rendering employs a volume rendering model based on ray integrals. For pixel p, the corresponding camera ray r(t) = o + td. All Gaussian elements intersecting the ray are sorted by depth. Color accumulation uses differentiable alpha blending. ; o is the camera origin, d is the ray direction, t is the ray parameter variable, i.e., the distance along the ray direction; C(p) is the color accumulation value. Gauss color value, , Gauss Contribution at pixel p For Gauss Opacity Gauss The probability density of the Gaussian distribution at pixel p.

5. The rapid Gaussian splashing human body surface mesh reconstruction method for scoliosis diagnosis and treatment as described in claim 4, characterized in that, In step S4, the TSDF fusion algorithm with truncated symbolic distance function is integrated, and the process is as follows: TSDF fusion assigns a sign function with truncation distance to multi-view depth maps in units of voxel x: Where d(x) is the signed distance from the voxel center to the depth surface, and μ is the truncation threshold. Weighted averaging is used when fusing multiple viewpoints. Extracting the zero level set As the reconstructed surface, a triangular mesh of the human body surface with a closed topology is obtained.

6. The rapid Gaussian splashing method for human body surface mesh reconstruction for scoliosis diagnosis and treatment as described in claim 1, characterized in that, The reference method described in step S5 specifically includes: The pre-trained image segmentation model processes all sampled video frames to identify and segment the pixel region where the reference object is located, which is a rectangular card. Record the two-dimensional pixel coordinates of the two vertices of the diagonal of the rectangular region of the reference object; Based on the camera parameters and corresponding depth map of the video frame, two two-dimensional pixels are converted into three-dimensional coordinate points in the reconstructed three-dimensional space through back projection calculation. Calculate the Euclidean distance between two 3D coordinate points. The distance is the length of the diagonal of the reference object in the virtual size space. Compare the virtual length with the actual physical length of the diagonal of the reference object to calculate the scaling scale from the virtual size to the actual size. By applying the scale bar to all vertices of the human body surface mesh and performing size transformation, a human body model with the correct physical dimensions is obtained.

7. The rapid Gaussian splashing human body surface mesh reconstruction method for scoliosis diagnosis and treatment as described in claim 6, characterized in that, The back projection calculation method is as follows: Back projection uses depth D(u,v) to map pixel (u,v) back to 3D space, as shown in the following formula: The absolute position of the point in the reconstructed coordinate system is obtained based on the camera's extrinsic parameters; X represents the three-dimensional spatial coordinates of the pixel (u,v). The two pixel vertices of the reference object are back-projected to obtain X1 and X2 respectively. The Euclidean distance is ||X1−X2||2, which is the diagonal of the virtual spatial reference object.

8. The rapid Gaussian splashing method for human surface mesh reconstruction for scoliosis diagnosis and treatment as described in claim 1, characterized in that, Step S5, the height method, specifically includes: Record the actual height of the person being photographed as a benchmark for their true dimensions; Principal component analysis is applied to the coordinates of all vertices of the virtual human body surface mesh to find the principal direction with the largest variance. Project all vertices onto the principal direction with the largest variance. The difference between the maximum and minimum values ​​after projection is the virtual height of the current model. By comparing the calculated virtual height with the actual height of the person being photographed, a scale for size transformation can be obtained; Apply the scale bar to all vertices of the mesh to obtain a human body surface mesh with the correct dimensions.

9. The rapid Gaussian splashing method for human body surface mesh reconstruction for scoliosis diagnosis and treatment as described in claim 8, characterized in that, The virtual height calculation method is as follows: PCA is used to find the principal orientation of a virtual human body mesh. Let the set of all vertices be . The covariance matrix is: in Given the mean of the vertex coordinates, solve for the eigenvalue decomposition Cv = λv, where v is the eigenvector to be solved for, and C is the covariance matrix. The eigenvalues ​​to be solved are the eigenvectors. The largest eigenvalue corresponds to the eigenvector v1, which is the height direction. Project all points onto the v1 direction, and take the maximum difference between each pair of projected values ​​to obtain the virtual height.