A multi-modal adolescent idiopathic scoliosis screening method based on back RGB-D images

By combining RGB-D images and an improved MSCA-YOLO network, the trunk rotation angle and back contour asymmetry index are calculated, which solves the problems of high professionalism, strong subjectivity and low efficiency of existing AIS screening methods, and realizes radiation-free, fast and accurate screening for adolescent idiopathic scoliosis.

CN120543912BActive Publication Date: 2026-07-24UNIV OF ELECTRONICS SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2025-05-12
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing AIS screening methods require highly specialized screening physicians, are highly subjective, involve cumbersome procedures, are inefficient, and have poor robustness, making it difficult to achieve efficient and accurate screening for adolescent idiopathic scoliosis under radiation-free conditions.

Method used

A multimodal screening method based on back RGB-D images was adopted, which utilizes the complementarity of RGB-D data to calculate the maximum/minimum trunk rotation angle and back contour asymmetry index, and combines the improved MSCA-YOLO network to detect key points on the back, so as to achieve radiation-free screening for adolescent idiopathic scoliosis.

Benefits of technology

It achieves rapid, efficient, and robust AIS screening, accurately identifying scoliosis in large-scale screenings, reducing reliance on professionals and subjective influence, and improving the objectivity and accuracy of screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543912B_ABST
    Figure CN120543912B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal juvenile idiopathic scoliosis screening methods based on back RGB-D image, belong to image processing and scoliosis screening technical field, specifically include: acquisition subject is under the RGB-D image pair of Adams forward bending posture, RGB color image is artificially labeled, construct and train back key point detection model, establish the calculation rule of maximum trunk rotation angle, minimum trunk rotation angle and back profile asymmetry index, establish AIS classification rule based on RGB-D dataset, and then complete the AIS classification of the subject to be detected.The application is a kind of non-contact radiation-free AIS screening method, utilizes the complementary mode of two kinds of modal data of RGB-D, realizes the comprehensive utilization of back two-dimensional, three-dimensional information, makes up the problem of feature missing under single mode, with the characteristics of fast and efficient, robust objective.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of image processing and scoliosis screening technology, specifically relating to a multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images. Background Technology

[0002] Scoliosis refers to the lateral curvature of one or more vertebral segments in the spine, deviating from the body's midline in the coronal plane. Idiopathic scoliosis is a type of scoliosis that does not resolve spontaneously and has an unknown cause; it is the most common type of scoliosis among adolescents. Idiopathic scoliosis has become one of the major health problems facing adolescents.

[0003] Adolescent idiopathic scoliosis (AIS) is a progressive, long-term disease. In its early stages, the spinal deformity is not obvious. With the rapid growth of bones during adolescence, the deformity develops rapidly, often manifesting as uneven shoulders, asymmetrical chest, and spinal deviation from the midline. In severe cases, back deformities, trunk imbalance, and spinal cord and nerve damage may occur, causing significant harm to the patient's physical and mental health. Furthermore, the cost of treatment for AIS differs greatly between the early and later stages. In the early stages, the scoliosis angle is generally less than 20°, which can be improved through functional exercises; if the angle is greater than 30°, costly surgical treatment is required. Therefore, early detection and timely intervention are crucial for patients.

[0004] In clinical settings, radiological examinations expose adolescents to harmful radiation, increasing their risk of cancer. Therefore, full-dose X-rays are often not chosen for AIS screening. Currently, AIS screening methods mainly include physical screening and image screening.

[0005] Physical screening includes standing examination, forward flexion test, trunk rotation measurement, etc. It requires professionals to directly observe and touch the back of the subject for classification. This method requires high professionalism from the screening medical staff, is cumbersome and inefficient, and is easily affected by subjectivity.

[0006] Image screening uses various non-contact sensors, such as laser scanners and RGB cameras, to collect data on the subject's back, which is then processed and analyzed using traditional or deep learning methods. Among traditional methods, the moiré pattern method utilizes the interference effect of overlapping grating patterns to generate moiré patterns for asymmetry detection. A specific light source perpendicular to the back projects the moiré pattern onto the back area, and changes in the shape of the moiré pattern reflect the severity of scoliosis. This method can detect very small back asymmetries, but it is highly dependent on the screening environment, including lighting, and is prone to high false positives, making it unsuitable for large-scale screening. Hideki et al., in their study "Automated noninvasive detection of idiopathic scoliosis in children and adolescents: A principle validation study," developed an automated asymmetry assessment system for idiopathic scoliosis (AIS) using a three-dimensional depth sensor. Using a three-dimensional asymmetry analysis algorithm, they analyzed the point cloud of forward flexion posture, completing the scanning and analysis of a single subject within 1.5 seconds. The sensitivity for moderate to severe cases reached 97%. However, the subjects in this experiment were all suspected AIS patients; in actual screening where healthy adolescents make up a larger proportion, its sensitivity and accuracy may be even worse.

[0007] Deep learning methods for AIS image screening can be categorized into three types: convolutional neural network-based classification methods, generative classification methods, and object detection-based classification methods. Convolutional neural network-based scoliosis classification optimizes network parameters by learning from a large number of AIS positive and negative samples, ultimately classifying new samples. Generative classification methods utilize generative models to reconstruct the shape of the spine in two or three dimensions, calculating the Cobb angle based on the reconstructed spinal curve to classify the severity of scoliosis. Object detection-based classification methods focus on detecting key back features, using predictions of these features for rapid scoliosis screening. However, convolutional neural network-based methods require a large number of high-quality labeled samples, and the spatial invariance assumptions and relatively fixed structure of convolution make it difficult for the model to handle rotated or deformed data, lacking flexibility. Generative and object detection-based methods rely too heavily on bony points in the back, making it difficult to capture effective features for screening when the spine's bony features are not prominent.

[0008] Therefore, there is an urgent need to propose a radiation-free, low-cost, efficient, and robust AIS screening method. Summary of the Invention

[0009] To address the problems of high professional requirements for screening doctors, strong subjectivity, cumbersome procedures, low efficiency, and poor robustness in current AIS screening methods, this invention provides a multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images. As a non-contact, radiation-free AIS screening method, it utilizes the complementary nature of RGB-D modal data to achieve comprehensive utilization of two-dimensional and three-dimensional information from the back, compensating for the feature loss problem under single modality, and is characterized by speed, efficiency, robustness, and objectivity.

[0010] The technical solution adopted in this invention is as follows:

[0011] A multimodal screening method for adolescent idiopathic scoliosis based on RGB-D images of the back includes the following steps:

[0012] Step 1: Calibrate the depth camera, obtain the camera intrinsic parameter matrix, and use the depth camera to acquire multiple RGB-D image pairs of subjects in the Adams flexion posture, including RGB color images and depth images, to form an RGB-D dataset; the subjects are adolescents of different ages and genders.

[0013] Step 2: Experts manually annotate the RGB color images of each subject in the RGB-D dataset, including three parts:

[0014] 1) Label the human body region with a rectangle and name the category "back";

[0015] 2) Mark the locations of 8 key points on the back: left acromion (category name: acromion_left); right acromion (category name: acromion_right); cervical vertebra (category name: cervical7); sixth thoracic vertebra (category name: thoracic6); fourth lumbar vertebra (category name: lumbar4); left posterior superior iliac spine (category name: iliacus_left); right posterior superior iliac spine (category name: iliacus_right); sacral vertebra (category name: sacral).

[0016] 3) Label each subject with their AIS positive or negative category, including AIS negative and AIS positive;

[0017] All RGB color images and their corresponding annotation files containing bounding boxes and keypoint locations are divided into training and validation sets;

[0018] Step 3: Based on the YOLO11 deep neural network, construct a back keypoint detection model, and train it based on the training set and validation set to obtain the trained back keypoint detection model.

[0019] Step 4: Establish the subject's maximum trunk rotation angle (ART) max Minimum trunk rotation angle ART min The calculation rules for the Profile Asymmetry Index (PAI) are as follows:

[0020] Step 4.1: Input the subject's RGB color image into the trained back key point detection model and output the coordinates of 8 back key points in the pixel coordinate system;

[0021] Step 4.2: Calculate the width of the spine and its surrounding region using the x-axis coordinates of the left and right posterior superior iliac spines. ART ;

[0022] Step 4.3: Use an interpolation-based method to fill holes in the subject's depth image to obtain a filled depth image;

[0023] Step 4.4: Based on the camera intrinsic parameter matrix, RGB color image, and filled depth image, acquire point cloud image; transform the coordinates of the thoracic vertebra, sixth thoracic vertebra, fourth lumbar vertebra, and sacral vertebra from pixel coordinate system to point cloud coordinate system; based on the camera intrinsic parameter matrix, Width... ART The point cloud image of the spine and its associated regions is obtained by using the coordinates of the thoracic protuberance and sacrum in the point cloud coordinate system. Then, based on the y-axis coordinates of the thoracic protuberance, the sixth thoracic vertebra, the fourth lumbar vertebra, and the sacrum in the point cloud coordinate system, the point cloud image of the spine and its associated regions is divided into three sub-regions.

[0024] Step 4.5: Fit the point cloud in the three sub-regions to a three-dimensional plane and obtain the equation coefficients of the three-dimensional plane corresponding to each sub-region;

[0025] Step 4.6: For each sub-region, calculate the angle between its corresponding 3D plane and the x-axis of the point cloud coordinate system; take the maximum angle between the three sub-regions as ART. max The minimum included angle is ART min ;

[0026] Step 4.7: Calculate the PAI based on the coordinates of the 8 back keypoints in the pixel coordinate system;

[0027] Step 5: Based on Step 4, calculate the ART of all subjects in the RGB-D dataset. max ART min Combined with PAI, and the AIS positive / negative categories and key point locations marked in step 2, according to ART max ART minBased on the distribution of PAI in AIS positive and negative categories, a threshold 'a' for trunk rotation angle and a threshold 'b' for back PAI classification were set. Considering that the coordinates of key back points are obtained through model inference in practical applications, which may have some error compared to the annotations of professional spine surgeons, the thresholds 'a' and 'b' were appropriately adjusted to increase the number of suspected AIS positive categories to avoid missed detections. The specific AIS classification rules established are as follows:

[0028] 1) The criteria for classification as positive for AIS are met:

[0029]

[0030] 2) A candidate is classified as a suspected AIS positive if any of the following criteria are met:

[0031] PAI≥(b+2)and ART max ≥(a-4)

[0032] (b+1)≤PAI<(b+2)and ART max ≥ART min ≥(a-3)

[0033] b≤PAI<(b+1)and ART max ≥(a-1)

[0034] ART max ≥(a+5)and ART min ≥a

[0035] a≤ART max <(a+5)and ART min >(a-1) and PAI≥(b-1)

[0036] 3) When the criteria for AIS positive or suspected AIS positive are not met, the classification is AIS negative;

[0037] Step 6: For the subjects to be tested, which are independent of the RGB-D dataset, first use a depth camera to acquire RGB-D image pairs in the Adams flexion posture, including RGB color images and depth images; then, according to Step 4, calculate the ART of the subjects to be tested. max ART min And PAI; finally, based on the AIS classification rules, the subjects to be tested were divided into three categories: AIS negative, suspected AIS positive, and confirmed AIS positive.

[0038] Furthermore, in step 4.2, Width ART The calculation formula is:

[0039] WidthART =1 / 2×(|ileft) x -iright x |)

[0040] In the formula, ileft x The x-axis coordinate of the left posterior superior iliac spine; iright x The x-axis coordinate is the right posterior superior iliac spine.

[0041] Furthermore, the specific process of filling the void in step 4.3 is as follows:

[0042] Obtain the pixel matrix corresponding to the subject's depth image. Using each pixel in the pixel matrix as the center pixel, perform filling processing sequentially. Specifically, when the pixel value of the center pixel is 0, calculate the average value of the non-zero pixel values ​​within its 3×3 region kernel as the new pixel value of the center pixel. If the center pixel value is not 0, no processing is performed. This process is repeated for each pixel to complete one iteration of filling, resulting in a new pixel value matrix after filling. A second iteration of filling is performed based on the new pixel value matrix until the preset maximum number of iterations is reached. Once the iteration is complete, the filled depth image is obtained.

[0043] Furthermore, step 4.5 employs the Random Sample Consensus (RANSAC) algorithm for point cloud fitting.

[0044] Furthermore, the specific process for calculating PAI in step 4.7 is as follows:

[0045] The least squares method was used to fit a straight line to four points: the thoracic vertebra, the sixth thoracic vertebra, the fourth lumbar vertebra, and the sacrum. This straight line was used as the spinal direction vector. The acromion direction vector formed by the left and right acromions was obtained. The angle β1 between the acromion direction vector and the spinal direction vector was calculated. The absolute value of the difference between β1 and 90° was used as the shoulder asymmetry index. The iliac crest direction vector formed by the left and right posterior superior iliac crests was obtained. The angle β2 between the iliac crest direction vector and the spinal direction vector was calculated. The absolute value of the difference between β2 and 90° was used as the lumbar asymmetry index. The maximum value between the shoulder asymmetry index and the lumbar asymmetry index was taken as the PAI (Post-operative Asymmetry Index).

[0046] Furthermore, the back key point detection model described in step 3 includes the skeletal structure, neck structure, and detection head structure;

[0047] The backbone structure comprises seven layers for downsampling and feature extraction of the input RGB color image. The first layer is a Conv (convolutional) block; the second and third layers each consist of a Conv block and a C2F (cross-stage local network) layer; the fourth and fifth layers each consist of a Conv block and a C3K2 (cross-level channel grouping and rearrangement) layer; and the sixth and seventh layers are respectively an SPPF (fast spatial pyramid pooling) layer and a C2PSA (cross-level pyramid slice attention) layer.

[0048] The neck structure is used to aggregate feature maps of different scales and pass these enhanced feature maps to the detection head;

[0049] The detection head structure is used to perform keypoint regression on the multi-scale features output by the neck structure, including three keypoint prediction branches FPC_KeyPoint with the same structure. Each keypoint prediction branch FPC_KeyPoint includes a human region detection branch, a category prediction branch, and a keypoint coordinate prediction branch.

[0050] The loss function of the back keypoint detection model consists of five parts: Box Loss, Keypoint Loss, Classification Loss, Objectness Loss, and Distribution Focal Loss.

[0051] Furthermore, between the first and second layers of the backbone structure, there is also an MSCA (Multi-Scale Convolutional Attention) block with residual connections. Specifically, the 11×11 large convolution is decomposed into two smaller convolutions of 3×3 and 5×5 using the explicit decomposition mechanism of convolution. These smaller convolutions are used to process the output features of the first layer, resulting in two feature maps with different receptive fields. After feature concatenation, max pooling, and average pooling, spatial attention is calculated. The spatial attention is then converted into a spatial attention map using a 7×7 convolution. The spatial attention map is then used to perform weighted calculations and summation on the feature maps of different scales. The output features of the first layer are then connected to the weighted summation features as the input features of the second layer.

[0052] Furthermore, the C2F layer in the second and third layers is preferably an MSCAC2F layer. Specifically, based on the C2F layer, an MSCA block with residual connections is used to replace the two convolutional layers in the original bottleneck layer. The MSCA block with residual connections is used to perform multi-scale convolution operations on the split features to obtain two feature maps with different receptive fields, and to perform spatial attention calculations on them.

[0053] Furthermore, the keypoint coordinate prediction branch adopts a grouped parallel structure, dividing the eight back keypoints into four groups. Specifically, the left and right acromions are grouped as the first group, the metatarsals and sacral vertebrae as the second group, the sixth thoracic vertebra and the fourth lumbar vertebra as the third group, and the left and right posterior superior iliac crests as the fourth group. The four groups of back keypoints are detected using parallel sub-paths with the same structure. Each parallel sub-path consists of a Conv block encapsulated with a standard convolution, batch normalization layer, and Swish activation function, an FPConv block with feature multiplication operation, and a standard 1×1 convolution. The output features of the four parallel sub-paths are concatenated and reshaped in the channel dimension to obtain the prediction result of each keypoint prediction branch FPC_KeyPoint. Finally, the prediction results of the three keypoint prediction branches FPC_KeyPoint are concatenated in the channel dimension. After dimensional transformation and decoding operations, the prediction results are used as the keypoint coordinate prediction results, i.e., the coordinates of the eight back keypoints in the pixel coordinate system.

[0054] Furthermore, the FPConv block with feature multiplication operation uses the implicit dimensionality increase idea of ​​features. First, the input features are fed into the Conv block, and the output features of the Conv block are processed using two parallel fully connected branches. One of the fully connected branches has an activation function set, while the other does not. Then, the output features of the two branches are multiplied element-wise to map the low-dimensional Conv block output features to a higher-dimensional nonlinear implicit space. Finally, the output is connected to the input features through a batch normalization layer.

[0055] Furthermore, the keypoint localization loss includes mean squared error (MSE) and adaptive wing loss;

[0056] The MSE is calculated based on the squared Euclidean distance between the predicted point and the actual point, where the predicted point is the predicted result of the key point coordinates, and the actual point is the position of 8 manually marked back key points.

[0057] The formula for calculating Adaptive Wing Loss is as follows:

[0058]

[0059] In the formula, y and These represent the actual heatmap and the model-predicted heatmap, respectively; ω, θ, and ∈ are all hyperparameters with values ​​greater than 0, where ω controls the weights and determines the impact of small errors, θ is the threshold for distinguishing between large and small errors, and ∈ controls the smoothness of the loss function; α is a hyperparameter with a value greater than 2, determining the shape of the loss function when the error is small; A and C are intermediate variables determined by ω, θ, and , as shown in the formula:

[0060] A=ω(1 / (1+(θ / ω)(α-y) ))(α-y)((θ / ω) (α-y-1) (1 / ω)

[0061] C=(θA-ωln(1+(θ / ω) (a-y) )).

[0062] Furthermore, the formula for calculating the loss function Loss of the back keypoint detection model is as follows:

[0063] Loss = G box ·L box +G kpmse ·L mse +G kpaw ·L awing +G kobj ·L kobj +G cl ·L cl +G dfl ·L dfl

[0064] In the formula, L box L represents the bounding box regression loss; mse L represents the mean square error; awing Indicates Adaptive WingLoss; L cl L represents classification loss; kobj Indicates confidence loss; L dfl G represents the distribution focal point loss; box G kpmse G kpaw G cl G kobj G dfl L respectively box L mse L awing L cl L kobj and L dfl Gain parameters.

[0065] Furthermore, G box G kpmse G kpaw G cl G kobj G dfl The values ​​are 7.5, 12.0, 1.0, 0.5, 1.0, and 1.5.

[0066] The beneficial effects of this invention are as follows:

[0067] 1. This invention proposes a multimodal screening method for adolescent idiopathic scoliosis based on back RGB-D images. As a non-contact, radiation-free AIS screening method, it adopts a complementary approach of RGB-D and two-modal data, calculates the maximum / minimum trunk rotation angle using a simulated trunk rotation measurement instrument, and then calculates the back contour asymmetry index to achieve comprehensive utilization of information on the asymmetry of the three-dimensional point cloud and the two-dimensional contour, thereby comprehensively measuring the degree of asymmetry in the subject's back region. This compensates for the feature loss problem under a single modality, and thus enables rapid, efficient, robust, and objective low-cost, radiation-free large-scale AIS screening.

[0068] 2. Preferably, this invention proposes an improved MSCA-YOLO network based on the YOLO11 deep neural network as a back keypoint detection model. Specifically, by using MSCA blocks with residual connections, the receptive field is expanded and dynamically adjusted while reducing the number of parameters in large kernel convolution; the MSCA2F layer is used instead of the C2F layer to expand the receptive field of the convolution kernel, reduce the inherent limitations of convolutional structures in processing rotated and deformed data, and improve the keypoint detection accuracy when the back region has a large degree of tilt and asymmetry; a grouped parallel structure is adopted in the keypoint coordinate prediction branch, and FPConv blocks with feature multiplication operations are used to effectively increase the feature dimension and improve the detection head accuracy. Attached Figure Description

[0069] Figure 1 This is an overall flowchart of the multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images proposed in Embodiment 1 of the present invention;

[0070] Figure 2 This is a schematic diagram of the overall architecture of the back key point detection model proposed in Embodiment 1 of the present invention;

[0071] Figure 3 This is a structural diagram of the MSCA block and the MSCAC2F block proposed in Embodiment 1 of the present invention;

[0072] Figure 4 This is a structural diagram of the FPConv block proposed in Embodiment 1 of the present invention;

[0073] Figure 5 This is a schematic diagram of the back key point coordinate inference results and back point cloud region output by the back key point detection model in Embodiment 1 of the present invention; wherein, (a) is an RGB image; (b) is the back key point inference result; (c) is a point cloud image; and (d) is a point cloud image of the spine and its associated regions.

[0074] Figure 6 This is a schematic diagram illustrating the calculation of the back contour asymmetry index (PAI) in Embodiment 1 of the present invention.

[0075] Figure 7 This is a schematic diagram illustrating the division of a point cloud image of the spine and its associated regions into three sub-regions in Embodiment 1 of the present invention.

[0076] Figure 8 This is a sample image of subjects classified as AIS negative, suspected AIS positive, and confirmed AIS positive under the established AIS classification rules in Embodiment 1 of the present invention. Detailed Implementation

[0077] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and represented herein can generally be arranged and designed in various different configurations.

[0078] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0079] Example 1

[0080] This embodiment provides a multimodal AIS screening method based on back RGB-D images, the overall process of which is as follows: Figure 1 As shown, in the quantification of two-dimensional asymmetry in the back region, the back contour asymmetry index is calculated using the coordinates of key points on the back. In the quantification of three-dimensional asymmetry in the back region, the trunk rotation angle is calculated using a three-dimensional plane fitted with point clouds of the spine and its associated regions. Then, the degree of scoliosis is comprehensively measured based on these two indicators. The following is an explanation of this embodiment with reference to the accompanying drawings.

[0081] The multimodal AIS screening method based on back RGB-D images specifically includes the following steps:

[0082] Step 1: Calibrate the depth camera and obtain the camera intrinsic parameter matrix. Use the depth camera to acquire 1500 RGB-D image pairs of subjects in the Adams flexion posture, including RGB color images and depth images, both 480×640 in size and pixel aligned, to form an RGB-D dataset; the subjects are adolescents of different ages and genders.

[0083] Step 2: Utilize experts (such as professional spine surgeons) to manually annotate the RGB color images of each subject in the RGB-D dataset. This includes three parts:

[0084] 1) Label the human body region with a rectangle and name the category "back";

[0085] 2) Mark the locations of 8 key points on the back: left acromion (category name: acromion_left); right acromion (category name: acromion_right); cervical vertebra (category name: cervical7); sixth thoracic vertebra (category name: thoracic6); fourth lumbar vertebra (category name: lumbar4); left posterior superior iliac spine (category name: iliacus_left); right posterior superior iliac spine (category name: iliacus_right); sacral vertebra (category name: sacral).

[0086] 3) Label each subject with their AIS positive or negative category, including AIS negative and AIS positive;

[0087] All RGB color images and their corresponding annotation files containing bounding boxes and keypoint locations are randomly divided into training and validation sets in an 8:2 ratio.

[0088] Step 3: Based on the YOLO11 deep neural network, construct the overall architecture as follows: Figure 2 The improved MSCA-YOLO network shown is used as a back keypoint detection model and is trained based on the training set and validation set to obtain the trained back keypoint detection model.

[0089] The back key point detection model includes the skeletal structure (i.e., Figure 2 The MSCA-YOLO backbone and neck structure (i.e.) Figure 2 The MSCA-YOLO neck and detection head structure (i.e. Figure 2 (MSCA-YOLO detection head in the middle);

[0090] The backbone structure includes an 8-layer structure for downsampling and feature extraction of the input RGB color image;

[0091] The first layer is a Conv block, which uses standard convolution, batch normalization (BN), and the Swish activation function (Sigmoid Linear Unit, SiLU) to capture image features and achieve downsampling.

[0092] The second layer is an MSCA block with residual connections, structured as follows: Figure 3As shown in the inner diagram, multi-scale convolutional attention is used to calculate the attention of the feature blocks extracted in the first layer. Specifically, the 11×11 large convolution is decomposed into two smaller convolutions of 3×3 and 5×5 using the explicit decomposition mechanism of convolution. These smaller convolutions are used to process the output features of the first layer, resulting in two feature maps with different receptive fields. After feature concatenation, max pooling, and average pooling operations, spatial attention is calculated. The spatial attention is converted into a spatial attention map using a 7×7 convolution. The spatial attention map is then used to perform weighted calculations and summation on the feature maps of different scales. To avoid gradient vanishing and gradient exploding, the output features of the first layer are connected to the weighted summation features as the input features of the second layer.

[0093] Both the third and fourth layers consist of a Conv block and an MSCAC2F layer; the structure of the MSCAC2F layer is as follows: Figure 3 As shown, based on the C2F layer of the YOLO11 deep neural network, an MSCA block with residual connections is used to replace the two convolutional layers in the original bottleneck layer. Compared with the original bottleneck layer structure, the MSCA block with residual connections performs multi-scale convolution operations on the split features to obtain two feature maps with different receptive fields, and performs spatial attention calculation on them. This MSCA2F layer expands the receptive field of the convolution kernel, realizes dynamic adjustment of the receptive field, effectively reduces the inherent limitations of convolution in processing rotated and deformed data, and improves the key point detection accuracy when the back region has a large degree of tilt and asymmetry.

[0094] Both the fifth and sixth layers consist of a Conv block and a C3K2 layer;

[0095] The seventh layer is the SPPF layer;

[0096] The eighth layer is the C2PSA layer;

[0097] The neck structure is used to aggregate feature maps of different scales and pass these enhanced feature maps to the detection head;

[0098] The detection head structure is used to perform keypoint regression on the multi-scale features output by the neck structure. It includes three keypoint prediction branches FPC_KeyPoint with the same structure, which process feature maps of different sizes respectively. Each keypoint prediction branch FPC_KeyPoint includes a human region detection branch, a category prediction branch, and a keypoint coordinate prediction branch.

[0099] Among them, the human region detection pathway and the category prediction pathway have the same structure as those in the YOLO11 deep neural network;

[0100] To reduce the coupling between keypoints on the back, the keypoint coordinate prediction path adopts a grouped parallel structure, dividing the eight back keypoints into four groups: the left and right acromion as the first group, the metatarsals and sacrum as the second group, the sixth thoracic vertebra and the fourth lumbar vertebra as the third group, and the left and right posterior superior iliac crests as the fourth group. The four groups of back keypoints are detected using parallel sub-paths with identical structures. Each parallel sub-path consists of a Conv block encapsulated with a standard convolution, batch normalization layer, and Swish activation function, an FPConv block with feature multiplication, and a standard 1×1 convolution. The output features of the four parallel sub-paths are concatenated and reshaped in the channel dimension to obtain the prediction result of each keypoint prediction branch, FPC_KeyPoint. Finally, the prediction results of the three keypoint prediction branches, FPC_KeyPoint, are concatenated in the channel dimension. After dimensional transformation and decoding, this becomes the predicted result of the keypoint coordinates, i.e., the coordinates of the eight back keypoints in pixel coordinates.

[0101] The structure of the FPConv block with feature product operation is as follows: Figure 4 As shown, the input features are first fed into a Conv block containing a 3×3 standard convolution. The features processed by the Conv block are then input into two fully connected branches, where the multilayer perceptron ratio is set to 3. The left fully connected branch uses ReLU6 as the activation function, while the right branch does not have an activation function. Subsequently, the output features of the two fully connected branches are multiplied element-wise, passed through a batch normalization layer, and then concatenated with the input features before output. The FPConv block with feature multiplication operation utilizes the implicit dimensionality upscaling idea of ​​features to map the low-dimensional Conv block output features to a higher-dimensional nonlinear implicit space, effectively amplifying the feature dimension and improving the accuracy of the detection head.

[0102] The loss function of the back keypoint detection model consists of five parts: bounding box regression loss, keypoint localization loss, classification loss, confidence loss, and distribution focus loss. Except for the keypoint localization loss, the other four parts are consistent with the loss function in the YOLO11 deep neural network.

[0103] The keypoint localization loss includes MSE and Adaptive Wing Loss;

[0104] The MSE is calculated based on the squared Euclidean distance between the predicted point and the actual point, using the following formula:

[0105] MSE=(x pred -x gt ) 2 +(y pred -y gt ) 2

[0106] In the formula, (x pred ,y pred (x) represents the predicted coordinates of the keypoint; the predicted point is the predicted coordinate of the keypoint. gt ,y gt () represents the coordinates of the actual points, which are the locations of 8 manually labeled key points on the back.

[0107] The Adaptive Wing Loss is a loss function proposed based on MSE loss. It overcomes the insensitivity of MSE loss to small errors and the imbalance between foreground and background pixel categories. Its calculation formula is as follows:

[0108]

[0109] In the formula, y and These represent the actual heatmap and the model-predicted heatmap, respectively; ω, θ, and ∈ are all hyperparameters with values ​​greater than 0, where ω controls the weights and determines the impact of small errors, θ is the threshold for distinguishing between large and small errors, and ∈ controls the smoothness of the loss function; α is a hyperparameter with a value greater than 2, determining the shape of the loss function when the error is small; A and C are intermediate variables determined by ω, θ, and , as shown in the formula:

[0110] A=ω(1 / (1+(θ / ω) (α-y) ))(α-y)((θ / ω) (α-y-1) (1 / ω)

[0111] C=(θA-ωln(1+(θ / ω) (α-y) )).

[0112] In this embodiment, ω = 14, θ = 0.5, ∈ = 1, α = 2.1;

[0113] Therefore, the formula for calculating the loss function Loss of the back keypoint detection model is as follows:

[0114] Loss = G box ·L box +G kpmse ·L mse +G kpaw ·L awing +G kobj ·L kobj +G cl ·L cl +G dfl ·L dfl

[0115] In the formula, L box L represents the bounding box regression loss; mse Indicates MSE; L awingIndicates Adaptive Wing Loss; L cl L represents classification loss; kobj Indicates confidence loss; L dfl G represents the distribution focal point loss; box G kpmse G kpaw G cl G kobj G dfl L respectively box L mse L awing L cl L kobj and L dfl Gain parameters.

[0116] In this embodiment, G box G kpmse G kpaw G cl G kobj G dfl The values ​​are 7.5, 12.0, 1.0, 0.5, 1.0, and 1.5.

[0117] Step 4, based on Figures 5-7 Establish the maximum trunk rotation angle of the subject in ART max Minimum trunk rotation angle ART min The calculation rules and specific process for the Back Contour Asymmetry Index (PAI) are as follows:

[0118] Step 4.1, as follows Figure 5 The RGB color image of the subject shown in (a) is input into the trained back keypoint detection model, which outputs the coordinates of eight back keypoints in pixel coordinates, as shown in (a). Figure 5 As shown in (b);

[0119] Step 4.2: Calculate the width of the spine and its surrounding region using the x-axis coordinates of the left and right posterior superior iliac spines. ART The calculation formula is:

[0120] Width ART =1 / 2×(|ileft) x -irigt x |)

[0121] In the formula, ileft x The x-axis coordinate of the left posterior superior iliac spine; iright x The x-axis coordinate of the right posterior superior iliac spine;

[0122] Step 4.3: Use an interpolation-based method to fill holes in the subject's depth image to obtain a filled depth image. The specific process is as follows:

[0123] Obtain the pixel matrix corresponding to the subject's depth image. Using each pixel in the pixel matrix as the center pixel, perform filling processing sequentially. Specifically, when the pixel value of the center pixel is 0, calculate the average value of the non-zero pixel values ​​within its 3×3 kernel region as the new pixel value of the center pixel. If the center pixel value is not 0, no processing is performed. This process is repeated for each pixel to complete one iteration of filling, resulting in a new pixel value matrix after filling. A second iteration of filling is performed based on the new pixel value matrix until a preset maximum number of iterations is reached. Once the iteration is complete, the filled depth image is obtained.

[0124] Step 4.4: Based on the camera intrinsic parameter matrix, RGB color image, and filled depth image, obtain the point cloud image, such as... Figure 5 As shown in (c); the coordinates of the thoracic vertebra, sixth thoracic vertebra, fourth lumbar vertebra, and sacrum are transformed from the pixel coordinate system to the point cloud coordinate system. The transformation formula is:

[0125] z = d / s

[0126] x = (uc) x )·z / f x ,y=((vc y )·z / f y

[0127] In the formula, (u,v) are the coordinates in the pixel coordinate system; (x,y,z) are the coordinates in the point cloud coordinate system; and the principal optical center (c) is the coordinate of the pixel coordinate system. x ,c y ), focal length (f x ,f y Both the depth scaling factor s and the depth scaling factor s are known values ​​related to the depth camera;

[0128] In this embodiment, (c x ,c y )=(233.6612,321.2410),(f x ,f y )=(590.7570,590.2881), s=1000;

[0129] Based on camera intrinsic parameter matrix, Width ART And the coordinates of the protuberance and sacrum in the point cloud coordinate system, to obtain, as Figure 5The back point cloud region shown in (d) is denoted as ROI_ART. Then, based on the y-axis coordinates of the thoracic vertebrae, the sixth thoracic vertebra, the fourth lumbar vertebra, and the sacrum in the point cloud coordinate system, the point cloud image of the spine and its associated regions is divided into three sub-regions, with the division rules as follows: Figure 7 As shown, the area from the thoracic vertebra to the sixth thoracic vertebra is designated as sub-region 1, the area from the sixth thoracic vertebra to the fourth lumbar vertebra is designated as sub-region 2, and the area from the fourth lumbar vertebra to the sacrum is designated as sub-region 3.

[0130] Step 4.5: Using the random sample consensus algorithm, fit the point clouds in the three sub-regions into a 3D plane, and obtain the equation coefficients of the 3D plane corresponding to each sub-region, such as... Figure 7 The three-dimensional plane equations corresponding to the three sub-regions of the subject example shown are as follows:

[0131] 0.00x + 0.50y + 0.87z + 0.75 = 0

[0132] 0.04.x ​​- 0.21y + 0.98z + 0.87 = 0

[0133] 0.01x - 0.53y + 0.85z + 0.72 = 0

[0134] Step 4.6: For each sub-region, calculate the angle between its corresponding 3D plane and the x-axis of the point cloud coordinate system; take the maximum angle between the three sub-regions as ART. max The minimum included angle is ART min ;

[0135] Step 4.7: Calculate the PAI based on the coordinates of the 8 back keypoints in the pixel coordinate system. The specific process is as follows:

[0136] like Figure 6 As shown, the least squares method is used to fit four points—the thoracic vertebra, the sixth thoracic vertebra, the fourth lumbar vertebra, and the sacrum—to a straight line, which is then used as the spinal direction vector. The acromion direction vector formed by the left and right acromions is obtained, and the angle β1 between the acromion direction vector and the spinal direction vector is calculated. The absolute value of the difference between β1 and 90° is used as the shoulder asymmetry index. The iliac crest direction vector formed by the left and right posterior superior iliac crests is obtained, and the angle β2 between the iliac crest direction vector and the spinal direction vector is calculated. The absolute value of the difference between β2 and 90° is used as the lumbar asymmetry index. The maximum value between the shoulder asymmetry index and the lumbar asymmetry index is taken as the PAI (Post-operative Asymmetry Index).

[0137] Step 5: Based on Step 4, calculate the ART of all subjects in the RGB-D dataset. max ART min Combined with PAI, the AIS positive / negative category annotations and key point locations marked in step 2, and based on ART maxART min Based on the distribution of PAI in AIS positive and negative categories, a threshold 'a' for trunk rotation angle and a threshold 'b' for back PAI classification were set. Considering that the coordinates of key back points are obtained through model inference in practical applications, which may have some error compared to the annotations of professional spine surgeons, the thresholds 'a' and 'b' were appropriately adjusted to increase the number of suspected AIS positive categories to avoid missed detections. The specific AIS classification rules established are as follows:

[0138] 1) The criteria for classification as positive for AIS are met:

[0139]

[0140] 2) A candidate is classified as a suspected AIS positive if any of the following criteria are met:

[0141] PAI≥(b+2)and ART max ≥(a-4)

[0142] (b+1)≤PAI<(b+2)and ART max ≥ART min ≥(a-3)

[0143] b≤PAI<(b+1)and ART max ≥(a-1)

[0144] ART max ≥(a+5)and ART min ≥a

[0145] a≤ART max <(a+5)and ART min >(a-1) and PAI≥(b-1)

[0146] 3) When the criteria for AIS positive or suspected AIS positive are not met, the classification is AIS negative;

[0147] In this embodiment, a = 5.0, b = 3.0.

[0148] Step 6: For the subjects to be tested, which are independent of the RGB-D dataset, first use a depth camera to acquire RGB-D image pairs in the Adams flexion posture, including RGB color images and depth images; then, according to Step 4, calculate the ART of the subjects to be tested. max ART min And PAI; finally, based on the AIS classification rules, the subjects to be tested were divided into three categories: AIS negative, suspected AIS positive, and confirmed AIS positive.

[0149] To verify the practicality of the multimodal AIS screening method based on back RGB-D images provided in this embodiment, AIS positive and negative classification was performed on data from 700 subjects independent of the RGB-D dataset. Figure 8 The examples shown are subject samples classified as AIS negative, suspected AIS positive, and confirmed AIS positive according to the classification criteria of this invention. The first row contains subject samples classified as AIS negative, the second row contains subject samples classified as suspected AIS positive, and the third row contains subject samples classified as confirmed AIS positive. This demonstrates that the method of this embodiment can rapidly, efficiently, robustly, and objectively achieve low-cost, radiation-free large-scale AIS screening.

[0150] The above embodiments are only for illustrating the principles and advantages of the present invention, and are not intended to limit the present invention. They are only for helping to understand the principles of the present invention. The scope of protection of the present invention is not limited to the above configurations and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the disclosed technology without departing from the essence of the present invention, but they are still within the scope of protection of the present invention.

Claims

1. A multimodal screening method for adolescent idiopathic scoliosis based on RGB-D images of the back, characterized in that, Includes the following steps: Step 1: Calibrate the depth camera, obtain the camera intrinsic parameter matrix, and use the depth camera to acquire multiple RGB-D image pairs of subjects in the Adams flexion posture, including RGB color images and depth images, to form an RGB-D dataset; the subjects are adolescents of different ages and genders. Step 2: Experts manually annotate the RGB color images of each subject in the RGB-D dataset, including three parts: 1) Label the human body region with a rectangle and name the category "back"; 2) Mark the locations of 8 key points on the back: left acromion (category name: acromion_left); right acromion (category name: acromion_right); cervical vertebra (category name: cervical7); sixth thoracic vertebra (category name: thoracic6); fourth lumbar vertebra (category name: lumbar4); left posterior superior iliac spine (category name: iliacus_left); right posterior superior iliac spine (category name: iliacus_right); sacral vertebra (category name: sacral). 3) Label each subject with their AIS positive or negative category, including AIS negative and AIS positive; All RGB color images and their corresponding annotation files containing bounding boxes and keypoint locations are divided into training and validation sets; Step 3: Based on the YOLO11 deep neural network, construct a back keypoint detection model, and train it based on the training set and validation set to obtain the trained back keypoint detection model. Step 4: Establish the subject's maximum trunk rotation angle (ART) max Minimum trunk rotation angle ART min The calculation rules and specific process for the Back Contour Asymmetry Index (PAI) are as follows: Step 4.1: Input the subject's RGB color image into the trained back key point detection model and output the coordinates of 8 back key points in the pixel coordinate system; Step 4.2: Calculate the width of the spine and its surrounding region using the x-axis coordinates of the left and right posterior superior iliac spines. ART ; Step 4.3: Use an interpolation-based method to fill holes in the subject's depth image to obtain a filled depth image; Step 4.4: Based on the camera intrinsic parameter matrix, RGB color image, and filled depth image, acquire point cloud image; transform the coordinates of the thoracic vertebra, sixth thoracic vertebra, fourth lumbar vertebra, and sacral vertebra from pixel coordinate system to point cloud coordinate system; based on the camera intrinsic parameter matrix, Width... ART The point cloud image of the spine and its associated regions is obtained by using the coordinates of the thoracic protuberance and sacrum in the point cloud coordinate system. Then, based on the y-axis coordinates of the thoracic protuberance, the sixth thoracic vertebra, the fourth lumbar vertebra, and the sacrum in the point cloud coordinate system, the point cloud image of the spine and its associated regions is divided into three sub-regions. Step 4.5: Fit the point cloud in the three sub-regions to a three-dimensional plane and obtain the equation coefficients of the three-dimensional plane corresponding to each sub-region; Step 4.6: For each sub-region, calculate the angle between its corresponding 3D plane and the x-axis of the point cloud coordinate system; take the maximum angle between the three sub-regions as ART. max The minimum included angle is ART min ; Step 4.7: Calculate the PAI based on the coordinates of the 8 back keypoints in the pixel coordinate system; Step 5: Based on Step 4, calculate the ART of all subjects in the RGB-D dataset. max ART min Combined with PAI, and the AIS positive / negative categories and key point locations marked in step 2, according to ART max ART min Based on the distribution of PAI in AIS positive and negative categories, a trunk rotation angle threshold 'a' and a back PAI classification threshold 'b' are set, and AIS classification rules are established as follows: 1) The criteria for classification as positive for AIS are met: 2) A candidate is classified as a suspected AIS positive if any of the following criteria are met: PAI≥(b+2)and ART max ≥(a-4) (b+1)≤PAI<(b+2)and ART max ≥ART min ≥(a-3) b≤PAI<(b+1)and ART max ≥(a-1) ART max ≥(a+5)and ART min ≥a a≤ART max <(a+5)and ART min >(a-1)and PAI≥(b-1) 3) When the criteria for AIS positive or suspected AIS positive are not met, the classification is AIS negative; Step 6: For the subjects to be tested, which are independent of the RGB-D dataset, first use a depth camera to acquire RGB-D image pairs in the Adams flexion posture, including RGB color images and depth images; then, according to Step 4, calculate the ART of the subjects to be tested. max ART min And PAI; finally, based on the AIS classification rules, the subjects to be tested were divided into three categories: AIS negative, suspected AIS positive, and confirmed AIS positive.

2. The multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images according to claim 1, characterized in that, The back key point detection model described in step 3 includes the skeletal structure, neck structure, and detection head structure; The backbone structure comprises a 7-layer structure for downsampling and feature extraction of the input RGB color image; wherein, the first layer is a Conv block; the second and third layers are both composed of Conv blocks and C2F layers; the fourth and fifth layers are both composed of Conv blocks and C3K2 layers; the sixth and seventh layers are SPPF layers and C2PSA layers, respectively. The neck structure is used to aggregate feature maps of different scales and transmit them to the detection head; The detection head structure is used to perform keypoint regression on the multi-scale features output by the neck structure, including three keypoint prediction branches FPC_KeyPoint with the same structure. Each keypoint prediction branch FPC_KeyPoint includes a human region detection branch, a category prediction branch, and a keypoint coordinate prediction branch. The loss function of the back keypoint detection model consists of five parts: bounding box regression loss, keypoint localization loss, classification loss, confidence loss, and distribution focus loss.

3. The multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images according to claim 2, characterized in that, Between the first and second layers of the backbone structure, there is also an MSCA block with residual connections. Specifically, the 11×11 large convolution is decomposed into two smaller convolutions of 3×3 and 5×5 using the explicit decomposition mechanism of convolution. These smaller convolutions are used to process the output features of the first layer, resulting in two feature maps with different receptive fields. After feature concatenation, max pooling, and average pooling, spatial attention is calculated. The spatial attention is converted into a spatial attention map using a 7×7 convolution. The spatial attention map is then used to perform weighted calculations and summation on the feature maps of different scales. The output features of the first layer are then connected to the weighted summation features as the input features of the second layer.

4. The multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images according to claim 3, characterized in that, The C2F layer in the second and third layers is the MSCAC2F layer. Specifically, based on the C2F layer, the two convolutional layers in the original bottleneck layer are replaced by an MSCA block with residual connections. The MSCA block with residual connections is used to perform multi-scale convolution operations on the split features to obtain two feature maps with different receptive fields, and to perform spatial attention calculations on them.

5. The multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images according to claim 4, characterized in that, The keypoint coordinate prediction branch adopts a grouped parallel structure, dividing the eight back keypoints into four groups: the left and right acromion as the first group, the protuberance and sacrum as the second group, the sixth thoracic vertebra and the fourth lumbar vertebra as the third group, and the left and right posterior superior iliac spines as the fourth group. The four groups of back keypoints are detected using parallel sub-paths with identical structures. Each parallel sub-path consists of a Conv block encapsulated with a standard convolution, batch normalization layer, and Swish activation function, an FPConv block with feature multiplication, and a standard 1×1 convolution. The output features of the four parallel sub-paths are concatenated and reshaped in the channel dimension to obtain the prediction result of each keypoint prediction branch, FPC_KeyPoint. Finally, the prediction results of the three keypoint prediction branches, FPC_KeyPoint, are concatenated in the channel dimension. After dimensional transformation and decoding, this is used as the prediction result of the keypoint coordinates, i.e., the coordinates of the eight back keypoints in the pixel coordinate system.

6. The multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images according to claim 5, characterized in that, The specific process of the FPConv block with feature product operation is as follows: First, the input features are fed into the Conv block. The output features of the Conv block are processed by two parallel fully connected branches. One of the fully connected branches has an activation function set, while the other does not. Then, the output features of the two branches are multiplied element by element. After passing through a batch normalization layer and connecting with the input features, the output is output.

7. The multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images according to claim 2, characterized in that, The keypoint localization loss includes mean square error and Adaptive Wing Loss; MSE is calculated based on the squared Euclidean distance between the predicted point and the actual point, where the predicted point is the predicted result of the key point coordinates, and the actual point is the position of 8 manually marked back key points. The formula for calculating Adaptive Wing Loss is as follows: In the formula, y and These represent the actual heatmap and the model-predicted heatmap, respectively; ω, θ, and ∈ are hyperparameters with values ​​greater than 0; α is a hyperparameter with a value greater than 2; A and C are intermediate variables determined by ω, θ, and α, as shown in the formula: A=ω(1 / (1+(θ / ω) (α-y) ))(α-y)((θ / ω) (α-y-1) (1 / h) C=(θA-ωln(1+(θ / ω) (α-y) ))。 8. The multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images according to any one of claims 1 to 7, characterized in that, In step 4.2, Width ART The calculation formula is: Width ART =1 / 2×(|ileft x -iright x |) In the formula, ileft x The x-axis coordinate of the left posterior superior iliac spine; iright x The x-axis coordinate is the right posterior superior iliac spine.

9. The multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images according to any one of claims 1 to 7, characterized in that, The specific process for calculating PAI in step 4.7 is as follows: The least squares method was used to fit a straight line to four points: the thoracic vertebra, the sixth thoracic vertebra, the fourth lumbar vertebra, and the sacrum. This straight line was used as the spinal direction vector. The acromion direction vector formed by the left and right acromions was obtained. The angle β1 between the acromion direction vector and the spinal direction vector was calculated. The absolute value of the difference between β1 and 90° was used as the shoulder asymmetry index. The iliac crest direction vector formed by the left and right posterior superior iliac crests was obtained. The angle β2 between the iliac crest direction vector and the spinal direction vector was calculated. The absolute value of the difference between β2 and 90° was used as the lumbar asymmetry index. The maximum value between the shoulder asymmetry index and the lumbar asymmetry index was taken as the PAI (Post-operative Asymmetry Index).

10. The multimodal adolescent idiopathic scoliosis screening method based on back RGB-D images according to any one of claims 1 to 7, characterized in that, The specific process of filling the void in step 4.3 is as follows: Obtain the pixel matrix corresponding to the subject's depth image. Take each pixel in the pixel matrix as the center pixel and perform filling processing in sequence. Specifically, when the pixel value of the center pixel is 0, calculate the average value of the pixels with non-zero pixel values ​​in its 3×3 region kernel as the new pixel value of the center pixel. If the center pixel value is not 0, no processing is performed. Repeat this process for each pixel to complete one iteration of filling and obtain the new pixel value matrix after filling. The process of filling in the image is repeated twice based on the new pixel value matrix until the preset maximum number of iterations is reached. Once the iteration is complete, the filled depth image is obtained.