Scoliosis-Assisted Screening Method Based on Symmetric Contrastive Learning and Contour Perception
By using symmetric contrast learning and contour perception methods in scoliosis screening, ordinary RGB cameras collect forward-flexion action videos and extract image features and contour edge masks, solving the problems of high hardware cost and low recognition accuracy in the prior art, and achieving efficient and accurate scoliosis screening.
Patent Information
- Application Number
- CN202411859739.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The existing scoliosis screening methods have problems such as high hardware cost, complex implementation process and low recognition accuracy, and have failed to make full use of dynamic forward bending video and contour edge features.
Using a method based on symmetric contrast learning and contour perception, the forward bending action video was collected through ordinary RGB cameras, and the image features and contour edge mask were extracted using the Swin Transformer codec. The model was trained to improve the recognition accuracy of scoliosis.
It reduces the cost of hardware equipment, simplifies the screening process, improves the accuracy of scoliosis identification, reduces the possibility of missed examinations, and is suitable for large-scale adolescent screening.
Smart Images

Figure CN119810046B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning and image processing, and particularly relates to a method for assisting in scoliosis screening, which can be used for remote large-scale scoliosis screening in hospitals. Background Art
[0002] Scoliosis is a three-dimensional deformity involving abnormal curvature and rotation of the spine in multiple dimensions. Adolescent idiopathic scoliosis (AIS) is the most common type, usually affecting adolescents aged 10 to 18. Statistical data shows that the global incidence of AIS is about 1% to 3% and shows an increasing trend year by year. For scoliosis patients, the severity of the impact varies from person to person. The mild ones may only show low back pain or mild neurological symptoms, while the severe ones may have limited cardiopulmonary development due to thoracic deformity, leading to cardiopulmonary insufficiency and even life-threatening. In addition, the appearance deformity caused by scoliosis often has an adverse impact on the mental health of patients. Many patients become self-abased due to body changes and then have social barriers.
[0003] Due to the high incidence and potential serious hazards of scoliosis, early detection, early diagnosis and early treatment are crucial. Timely intervention can not only reduce the impact of the disease on physical health, but also help patients regain confidence and avoid deepening of psychological trauma. In view of this, promoting the normalization and systematization of scoliosis screening is of great significance for reducing the incidence of scoliosis and its adverse effects.
[0004] The patent document with the publication number CN114287915A discloses a non-invasive scoliosis screening method and system based on back color images. The implementation scheme is as follows: First, the RGB and depth images of the back of the human body are collected by an RGB-D camera and saved as a color image and a depth map respectively; then, a human body segmentation model is obtained by training the Mask-RCNN network, and the background-free human color image is segmented by using this model; then, a back recognition model is obtained by training the YOLOv5 network, and the back region is intercepted to obtain a back color image; further, the back depth map is obtained, and the maximum ATR angle value is calculated; finally, a classification standard is formulated and the category is labeled, and the back color image and the category label are input into the EfficientNet network for training to obtain a spine classification model. However, this method has the problems of high price and high cost because it needs to use an RGB-D camera. At the same time, since the classification standard and category labeling need to be based on the calculated maximum ATR angle value during the implementation process, the implementation process is complex, the efficiency is low, and the recognition accuracy is not high.
[0005] The patent document with the publication number CN116258897A discloses a spinal scoliosis screening method based on back images. The implementation steps of this method include: First, construct a network model and prepare two spinal scoliosis datasets; then, input the first dataset into the network model for training, and optimize the model by selecting the minimization of the loss function and the optimal evaluation index; then, use the second dataset to fine-tune the model to obtain stable model parameters; finally, solidify the model parameters. During screening, only by inputting the image into the model can the screening result be obtained. Although this method can achieve large-scale spinal scoliosis screening with high accuracy and high efficiency, due to its dependence on the scale and quality of the dataset, if the sample size in the dataset is insufficient or the sample annotation is inaccurate, it will affect the accuracy and generalization ability of the model.
[0006] In addition, since the above existing method only relies on the static RGB images of the human back when standing, it does not analyze the dynamic forward flexion video of the examinee and does not obtain the edge features related to the judgment of spinal scoliosis. Therefore, the information for judging human spinal scoliosis is incomplete, resulting in possible missed detections. Summary of the Invention
[0007] The purpose of the present invention is to address the deficiencies of the above existing technologies and propose a spinal scoliosis assisted screening method based on symmetric contrast learning and contour perception, so as to reduce the screening cost by using a low-cost ordinary RGB camera to collect the forward flexion action video of the examinee, making it easier to popularize and promote; by analyzing the spinal dynamic information of the examinee in the ordinary RGB video of the forward flexion action and capturing the contour edge features of the examinee's shoulders, waistlines, pelvises, and backs, improve the accuracy of spinal scoliosis recognition and reduce the possibility of missed detections.
[0008] To achieve the above purpose, the implementation solution of the present invention includes the following steps:
[0009] 1. A spinal scoliosis assisted screening method based on symmetric contrast learning and contour perception, characterized by including the following:
[0010] (1) Collect videos and divide the training set D t and the test set D v :
[0011] (1a) Use an ordinary RGB camera to collect forward flexion test videos of different ages, genders, and spinal scoliosis degrees to construct a spinal scoliosis screening task dataset D;
[0012] (1b) Perform edge mask annotation and class label annotation on the video data in the dataset D to obtain its edge mask image and class label;
[0013] (1c) Preprocess the video data and the edge mask image by flipping to obtain the left image and its right-flipped image, as well as the corresponding left-edge mask annotation image and the right-flipped edge mask image;
[0014] (1d) Divide all the left and right image data into a training set D t and a test set D v ;
[0015] (2) Use the Swin Unet segmentation network including the Swin Transformer encoder-decoder as the contour edge prediction network;
[0016] (3) Train the Swin Unet contour edge prediction network using contrastive learning:
[0017] (3a) Input the training set D t into the contour edge prediction network in multiple batches to obtain its semantic features and the edge prediction mask image;
[0018] (3b) Calculate the symmetry contrastive learning loss value L ita between the semantic features of the left image and the right-flipped image of each sample in the same batch and the Dice loss value L D of the edge prediction mask image and the edge mask label;
[0019] (3c) Calculate the loss function L1 = αL ita + βL D , and perform backpropagation on the Swin Unet contour edge prediction network by minimizing L1 to adjust the model weight parameters of the network, where α + β = 1, and α and β are weight coefficients;
[0020] (3d) Repeat steps (3a) to (3d) until the number of training times reaches the set threshold or the value of the loss function converges to obtain the contour edge prediction model weights;
[0021] (4) Construct a scoliosis screening network based on symmetric contrastive learning and contour perception, which uses two-path Swin Transformer encoders, a feature decoder, and one-path contour feature encoder;
[0022] (5) Train the scoliosis screening network based on symmetric contrastive learning and contour perception:
[0023] (5a) Load the contour edge prediction model weights obtained in (3d) into the Swin Transformer encoders and the feature decoder, and freeze their weights;
[0024] (5b) Input the training set D tThe left image and the right flipped image in are input into the scoliosis screening network based on symmetric contrast learning and contour perception in multiple batches, semantic features of the left and right images and features of the superimposed image of the left and right edge masks are extracted, and they are concatenated to obtain fused features;
[0025] (5d) The fused features are input into the existing multi-layer perceptron module, and then the predicted class results of all samples in this batch are obtained through the softmax operation;
[0026] (5e) The predicted class results of all samples in this batch and the actual class labels are input into the Focal loss function and the cross-entropy loss function, and the Focal loss L foc and the cross-entropy loss L cls are calculated;
[0027] (5f) Calculate the loss function L2 = γL foc + δL cls , and perform backpropagation on the scoliosis screening network based on symmetric contrast learning and contour perception by minimizing L2, so as to adjust the model weight parameters of the network, where γ + δ = 1, and γ and δ are weight coefficients;
[0028] (5g) Repeat steps (5b) to (5f) until the number of training times reaches the set threshold or the value of the loss function converges, and obtain the trained scoliosis screening model based on symmetric contrast learning and contour perception;
[0029] (6) Input the test set D v into the trained scoliosis screening model based on symmetric contrast learning and contour perception, and use the visual physical examination strategy in the model to obtain the spinal state:
[0030] If all four test indicators of a certain sample in the test set D v are normal, it is confirmed that the spine of this sample is normal;
[0031] If there is one abnormal test indicator in a certain sample in the test set D v , it is confirmed that this sample has scoliosis.
[0032] The present invention has the following advantages compared with the prior art:
[0033] 1. Through the Swin Transformer encoder-decoder structure, the present invention extracts image features and contour edge masks, introduces the contour edge mask loss, which can prompt the model to enhance the contour edge features of the key areas of scoliosis; and judges the symmetry by extracting the superimposed contour edge features through the Swin encoder, which can constrain the model to recognize the morphological differences between the left and right sides of the spine.
[0034] 2. The present invention uses the Focal loss function to perform model constraint processing on the class imbalance problem, enabling the model to focus on classifications with fewer classes such as equal shoulder heights and difficult-to-classify samples.
[0035] 3. Since the present invention constructs an intra-sample symmetry contrast loss, the contour edge prediction model compares the symmetry differences between the left and right sides of the same sample, thereby effectively learning the symmetry features within the sample. Through the symmetry features, the screening accuracy of the scoliosis screening model can be significantly improved.
[0036] 4. Since the present invention uses an ordinary RGB camera to collect scoliosis data videos, compared with the existing method of using an RGB-D camera for scoliosis screening, the hardware device cost required is greatly reduced, making it suitable for large-scale screening of teenagers. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is the implementation flowchart of the present invention;
[0038] Figure 2 is the structural block diagram of the scoliosis screening model in the present invention;
[0039] Figure 3 is the edge mask image of the shoulder region prediction in the present invention;
[0040] Figure 4 is the edge mask image of the waistline region prediction in the present invention;
[0041] Figure 5 is the edge mask image of the pelvis region prediction in the present invention;
[0042] Figure 6 is the edge mask image of the back region prediction in the present invention;
[0043] Figure 7 is the superimposed image of the left and right masks of the shoulder region in the present invention;
[0044] Figure 8 is the superimposed image of the left and right masks of the waistline region in the present invention;
[0045] Figure 9 is the superimposed image of the left and right masks of the pelvis region in the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0046] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] It should be noted that the step numbers in the specification, claims and claims of the present invention are only for clearly describing the implementation solutions of the present invention for easy understanding, and their sequence numbers are not limited.
[0048] Refer to Figure 1 , the implementation steps of this example are as follows:
[0049] Step 1: Use an ordinary RGB camera to collect forward flexion test videos of different ages, genders, and degrees of scoliosis, and construct a scoliosis screening task dataset D.
[0050] (1.1) Set the acquisition requirements for the forward flexion test:
[0051] The upper body of male test subjects is bare, and the upper body of female test subjects wears underwear;
[0052] The test subjects take off their shoes, face a clean wall, take a natural standing posture, with their feet shoulder-width apart, look straight ahead, let their arms hang naturally, and the palms face inwards;
[0053] (1.2) Before the acquisition starts, inform the test subjects of the acquisition requirements for the forward flexion test to ensure that the forward flexion test standards are available during the official recording;
[0054] (1.3) After the acquisition starts, the test subjects start the forward flexion movement, including straightening the knees, closing the feet, standing at attention, straightening the arms and clasping the hands, lowering the head and then slowly bending forward to about 90°, gradually placing the clasped hands between the knees, and the tester shoots the overall forward flexion movement video of the test subjects and saves it to the dataset D.
[0055] Step 2: Preprocess the data and divide it into a training set D t and a test set D v .
[0056] (2.1) Perform edge mask annotation on the images of the video samples in the dataset D:
[0057] 2.1.1) Perform pixel-level annotation on the first frame image of the video samples in the dataset D, that is, annotate the edges of the waistline, shoulders, and pelvic contours, and generate binary mask labels for the waistline area and shoulder area in this frame image;
[0058] 2.1.2) For the images of the bending-down actions of the video samples in dataset D, perform pixel-level annotation every ten frames, mark the edges of the back contour, and generate binary mask labels for the back regions in these images;
[0059] (2.2) Set up a visual physical examination strategy to perform label annotation on the video samples in dataset D:
[0060] 2.2.1) Observe whether the shoulders of the person being measured in the standing image are at the same height. If the shoulders are at the same height, label it as normal shoulders; if the shoulders are not at the same height, label it as abnormal shoulders;
[0061] 2.2.2) Observe whether the two side waistlines of the person being measured in the standing image are symmetric. If the two side waistlines are symmetric, label it as normal waistline; if the two side waistlines are not symmetric, label it as abnormal waistline;
[0062] 2.2.3) Observe whether the left and right pelvises of the person being measured in the standing image are symmetric. If the left and right pelvises are symmetric, label it as normal pelvis; if the left and right pelvises are not symmetric, label it as abnormal pelvis;
[0063] 2.2.4) Observe the condition of the back of the person being measured by observing the overall forward flexion action video, and label the back image frames with left-right asymmetry during the bending-down process as razor-back labels, and label the left-right symmetric back image frames as normal-back labels at every other frame.
[0064] (2.3) Perform flipping preprocessing on the video data and edge mask images:
[0065] 2.3.1) Divide the first-frame standing RGB image of the video in dataset D and its annotated mask image into left and right parts centered on the human body, and also divide the back RGB image of the bending-down action of the video in dataset D and its annotated mask image into left and right parts centered on the human body;
[0066] 2.3.2) Horizontally flip the right-side RGB image and the right-side mask image of the standing and bending-down actions to obtain the right-side flipped RGB image and the right-side flipped mask image;
[0067] 2.3.3) Adjust the sizes of the left-side RGB image, the right-side flipped RGB image, the left-side mask image, and the right-side flipped mask image to 224×224 to conform to the input size of the model;
[0068] (2.4) Randomly divide all the samples in dataset D into a training set D t and a test set D v .
[0069] Step 3, construct a Swin Unet contour edge prediction network.
[0070] (3.1) Adopt the existing Swin Transformer encoder, which is equipped with four layers of Swin Transformer block encoding modules, used to efficiently capture multi-scale features of the input image by using hierarchical feature extraction and sliding window self-attention mechanism;
[0071] (3.2) Select the existing Swin Transformer decoder, which is equipped with four layers of Swin Transformer block decoding modules, and gradually restore the spatial details of the image through layer-by-layer upsampling and multi-scale fusion;
[0072] (3.3) Perform skip connections between the outputs of each layer of the Swin Transformer encoder and the inputs of the corresponding layers of the decoder to construct a Swin Unet contour edge prediction network. This network fully combines the powerful feature extraction ability of the encoder and the layer-by-layer detail restoration ability of the decoder to achieve multi-scale encoding and decoding of the input image, ensuring the efficient combination of global information and local details, so as to achieve high-precision prediction of the contour edge.
[0073] Step 4, train the Swin Unet contour edge prediction network using contrastive learning.
[0074] (4.1) Input the training set D t into the contour edge prediction network in multiple batches to obtain the semantic features of the image and the predicted mask image containing the waistline, shoulders and back edges:
[0075] 4.1.1) Input the left image and the right flipped image of the training set D t into the Swin Transformer encoder of the Swin Unet contour edge prediction network, and gradually reduce the spatial size of the feature map through the hierarchical structure of the encoder to extract four layers of multi-scale feature images of 56×56, 28×28, 14×14, and 7×7; perform pooling operation on the features of the 7×7 scale to obtain the semantic features of the image;
[0076] 4.1.2) Input the four layers of multi-scale feature maps extracted in step (4.1.1) into the Swin Transformer decoder for upsampling, and then generate a segmentation image with the same size as the input image by fusing these multi-scale features, and finally obtain the edge prediction mask images of the waistline area, shoulder area and back area;
[0077] (4.2) Calculate the symmetry contrastive learning loss value L ita :
[0078] 4.2.1) Take the semantic features of the left image of each sample in the same batch as the anchor point. When the class annotation label of the sample is normal, take the semantic features of the right flipped image as the first positive sample z of the anchor point p1 ; when the class annotation label of the sample is abnormal, take the semantic features of the right flipped image as the first negative sample z of the anchor point n1 ;
[0079] 4.2.2) Copy the anchor point image features as the second positive sample z of the anchor point p2 ;
[0080] 4.2.3) Rotate the anchor point image features by a small angle of 5° to 10°, and take the rotated image features as the second negative sample z of the anchor point n2 ;
[0081] 4.2.4) Input the semantic features of the left image, the semantic features of the right flipped image, and the image features of the small angle rotation of the anchor point image into the intra-sample contrastive learning loss function, and calculate the contrastive learning loss value L between the semantic features of the left image and the right flipped image of each sample in the same batch ita :
[0082]
[0083] where K is the set of the semantic features of the left image, the semantic features of the right flipped image, and the image features of the small angle rotation of the anchor point image in the same batch; k represents the k-th element of the set K; z k is the feature of the anchor point image of the k-th sample; P(k) is the set of the positive sample picture features of the k-th sample; A(k) is the set of all image features of the k-th sample except the anchor point; p represents the p-th element of the set P(k); a represents the a-th element of the set A(k); z p =z p1 +z p2 , z a =z p1 +z p2 +z n1 +z n2 ; τ is the temperature coefficient, which is used to control the discrimination of the model for negative samples;
[0084] (4.3) Calculate the Dice loss value L of the marginal prediction mask image and the marginal mask label of the same batch D :
[0085] 4.3.1) Calculate the number of intersection pixels |M∩G| of the marginal prediction mask image M and the marginal mask label G;
[0086] 4.3.2) Calculate the Dice loss value L of each regionD :
[0087]
[0088] Among them, |M∩G| represents the number of intersection pixels of the edge prediction mask image M and the edge mask label G, |M| represents the total number of pixels of the edge prediction mask image M, and |G| represents the total number of pixels of the edge mask label G;
[0089] (4.4) Calculate the loss function L1 = αL ita +βL D , and perform backpropagation on the Swin Unet contour edge prediction network by minimizing L1, so as to adjust the model weight parameters of the network, where α + β = 1, and α and β are weight coefficients used to balance the influence of the two parts of the loss on model training;
[0090] (4.5) Repeat steps (4.1) to (4.4) until the number of training times reaches the set threshold or the value of the loss function converges, and obtain the contour edge prediction model weights.
[0091] Step 5, construct a scoliosis screening network based on symmetric contrast learning and contour perception.
[0092] (5.1) Perform skip connections between each Swin Transformer block encoding module in the Swin Transformer image feature encoder and each Swin Transformer block encoding module in the Swin Transformer image feature decoder to form a path of encoder-decoder structure;
[0093] (5.2) Copy the path of encoder-decoder structure in (5.1) to obtain a second path of encoder-decoder structure, and place the two paths of encoder-decoder structures side by side in a way of contrast learning to form a siamese network for contrast learning, and the two paths of encoder-decoder share weights;
[0094] (5.3) Connect a path of contour feature encoder to the backend of the siamese network to form a scoliosis screening network based on symmetric contrast learning and contour perception, as Figure 2 shown.
[0095] Step 6, train the scoliosis screening network based on symmetric contrast learning and contour perception:
[0096] (6.1) Load the weights of the contour edge prediction model obtained in (4.5) into the Swin Transformer image feature encoder and feature decoder, and freeze their weights, which will not be updated during the subsequent training process.
[0097] (6.2) Input the left image and the right flipped image into the Swin Transformer image encoder to obtain four - layer multi - scale features of 56×56, 28×28, 14×14, and 7×7. Perform pooling operations on the left - and right - image features at the 7×7 scale to obtain the left - image semantic feature F image-L and the right - flipped - image semantic feature F image-R ;
[0098] (6.3) Input the left - and right - image multi - scale features into the feature decoder for up - sampling and decoding to obtain the edge mask image F mask-L of the left image with the same scale as the original image (224×224) and the edge mask image F mask-R of the right - flipped image. Among them, the edge mask images of the shoulders, waistline, pelvis, and back regions are respectively as shown in Figure 3 , Figure 4 , Figure 5 , Figure 6 ;
[0099] (6.4) After adding the left - and right - edge mask images pixel - by - pixel and then performing normalization processing, that is, when the pixel value is greater than 1, set it to 1, so as to obtain the superimposed image of the left - and right - edge masks. Among them, the superimposed images of the left - and right - edge masks in the shoulders, waistline, and pelvis regions are respectively as shown in Figure 7 , Figure 8 , Figure 9 ;
[0100] (6.5) Extract the features F edge of the superimposed image through the contour encoder from the left - and right - edge mask superimposed image;
[0101] (6.6) Perform a concatenation operation on the semantic features F image-K , F image-R of the left - and right - images and the features F edge of the left - and right - edge mask superimposed image to obtain the fused feature F final ; Input the fused feature into the existing multi - layer perceptron module, and then through the softmax operation, obtain the predicted class results of all samples in this batch;
[0102] (6.7) Input the predicted class results of all samples in this batch and the actual class labels into the Focal loss function to calculate the value of the Focal loss L foc :
[0103]
[0104] Among them, B is the number of samples in a single batch in the training set, and C is the number of classes;
[0105] θ cis the weight of category c, used to balance the influence between different categories. In this example, it is set to but not limited to 0.2;
[0106] μ is a parameter that adjusts the influence of the easy - hard sample pair on the loss. In this example, its value is set to but not limited to 3;
[0107] y i,c is the actual category label of the i - th sample. If sample i belongs to category c, then y i,c = 1, otherwise y i,c = 0;
[0108] pred i,c is the predicted probability that the i - th sample belongs to category c obtained after MLP and softmax operations;
[0109] (6.8) Input the predicted category results and actual category labels of all samples in this batch into the cross - entropy loss function to calculate the cross - entropy loss L cls :
[0110]
[0111] where B is the number of samples in the batch;
[0112] C is the number of categories;
[0113] y i ∈{0,1} represents the actual label of the i - th sample;
[0114] represents the predicted probability that the i - th sample is classified correctly;
[0115] (6.9) According to the results of steps (6.7) and (6.8), calculate the loss function L2 of the scoliosis screening network based on symmetric contrast learning and contour perception:
[0116] L2 = γL foc +δL cls ;
[0117] (6.10) Perform backpropagation on the scoliosis screening network based on symmetric contrast learning and contour perception by minimizing L2 to adjust the model weight parameters of the network, where γ + δ = 1, and γ and δ are weight coefficients;
[0118] (6.11) Repeat steps (6.2) - (6.10) until the number of training times reaches the set threshold or the value of the loss function converges, and obtain the trained scoliosis screening model based on symmetric contrast learning and contour perception.
[0119] Step 7, use the test set D vInput it into the trained scoliosis screening model based on symmetric contrast learning and contour perception, and use the visual physical examination strategy to obtain the spinal status of the test set samples.
[0120] (7.1) Input the test set D v into the trained scoliosis screening model based on symmetric contrast learning and contour perception, and output the classification results of whether the shoulders of the test set samples are at the same height, whether the waistlines are symmetric, whether the pelvises are symmetric, and whether there is a rib hump, that is, output 1 for normal and 0 for abnormal;
[0121] (7.2) According to the classification results, use the visual physical examination strategy to obtain the spinal status of the test set samples:
[0122] If for a certain sample in the test set D v the classification results of the shoulders, waistlines, pelvises, and rib humps are all normal, then confirm that the spine of this sample is normal;
[0123] If for a certain sample in the test set D v there is one abnormal classification result among the shoulders, waistlines, pelvises, and rib humps, then confirm that this sample has scoliosis.
[0124] The effects of the present invention can be further illustrated by the following test experiments:
[0125] I. Test conditions:
[0126] The test equipment is a computer and a camera, which are connected and set up 2 to 3 meters directly behind the person being tested for scoliosis.
[0127] The test data is the forward bending test videos of 204 tested persons collected in a certain hospital. All the tested persons are labeled according to the visual physical examination, and the labels are all marked by hospital doctors.
[0128] The simulation test platform uses an Intel(R) Core(TM) CPU E5-2683 v4@2.10GHz, equipped with an NVIDIA TITAN Xp graphics card, and the memory capacity is 12GB. This platform is based on the Linux operating system and is implemented using the Python language.
[0129] Evaluation indicators: Conduct the "visual physical examination" test to obtain a series of results, including whether the shoulders are at the same height, whether the two waistlines are symmetric, whether the left and right pelvises are at the same height, and whether there is a rib hump deformity, etc., and judge whether there is scoliosis.
[0130] II. Test content and results:
[0131] Under the above test conditions, the forward flexion test video data of 204 tested persons was processed by using the method of the present invention. 164 samples were used as the training set to train the scoliosis screening model constructed by the present invention, and the remaining 40 samples were tested by using the trained scoliosis screening model, that is, whether the shoulders are at the same height, whether the two waistlines are symmetrical, whether the left and right pelvises are symmetrical, whether there is a rib hump deformity, and whether there is scoliosis were tested, and the respective accuracy rates, precision rates, recall rates and F1 scores were calculated. The results are shown in Table 1.
[0132] Table 1 Scoliosis Screening Experiment Index Table
[0133]
[0134] As can be seen from Table 1, the accuracy rates and F1 scores of all test results reached 85% or above. In terms of judging whether a person has scoliosis, the accuracy rate also reached 85%, and the F1 score reached 91.5%. This shows that the present invention has excellent accuracy and reliability in screening scoliosis, providing strong support for the early detection and screening of scoliosis.
Claims
1. A scoliosis auxiliary screening method based on symmetric contrast learning and contour perception, characterized in that: These include: (1) Collect videos and divide them into training sets and test set : (1a) Use a common RGB camera to collect flexion test videos of different ages, genders, and scoliosis degrees to construct a scoliosis screening task dataset ; (1b) For the dataset Perform edge mask annotation and category label annotation on the video data in the image to obtain its edge mask image and category label; (1c) performing flip preprocessing on the video data and the edge mask image to obtain a left image and a right flipped image thereof, as well as a corresponding left edge mask annotated image and a right flipped edge mask image; (1d) Divide all left and right image data into training sets in a ratio of 8:2 and test set ; (2) Using the Swin Unet segmentation network including the Swin Transformer encoder-decoder as the contour edge prediction network; (3) Use contrastive learning to train the Swin Unet contour edge prediction network: (3a) The training set Divide into multiple batches and input them into the contour edge prediction network to obtain its semantic features and edge prediction mask images; (3b) Calculate the symmetry comparison learning loss value between the semantic features of the left image and the right flipped image of each sample in the same batch And the Dice loss value of the edge prediction mask image and the edge mask label ; (3c) Calculate the loss function , by minimizing Back propagation is performed on the Swin Unet contour edge prediction network to adjust the model weight parameters of the network, where , and is the weight coefficient; (3d) Repeat steps (3a) to (3d) until the number of training times reaches a set threshold or the value of the loss function converges, and obtain the weight of the contour edge prediction model; (4) Construct a scoliosis screening network based on symmetric contrast learning and contour perception, which uses two SwinTransformer encoders, feature decoders and one contour feature encoder; The implementation is as follows: (4a) Each layer of the Swin Transformer encoder is connected to each layer of the corresponding feature decoder through a skip connection to form a codec structure; (4b) duplicating one codec structure in (4a) to obtain a second codec structure, and placing the two codec structures in parallel through contrastive learning to form a twin network for contrastive learning, and the two codecs share weights; (4c) Connecting a contour feature encoder to the back end of the twin network to form a scoliosis screening network based on symmetric contrast learning and contour perception; (5) Training of scoliosis screening network based on symmetric contrast learning and contour perception: (5a) Load the contour edge prediction model weights obtained in (3d) into the Swin Transformer encoder and feature decoder, and freeze their weights; (5b) The training set The left image and the right flipped image in the image are divided into multiple batches and input into the scoliosis screening network based on symmetric contrast learning and contour perception, and the semantic features of the left and right images and the features of the left and right edge mask superimposed images are extracted, and they are spliced to obtain the fusion features; (5c) Input the fused features into the existing multi-layer perceptron module, and then perform a softmax operation to obtain the predicted category results of all samples in the batch; (5d) Input the predicted category results and actual category labels of all samples in the batch into the Focal loss function and the cross entropy loss function to calculate the Focal loss and cross entropy loss ; (5e) Calculate the loss function , by minimizing Back propagation is performed on the scoliosis screening network based on symmetric contrast learning and contour perception to adjust the model weight parameters of the network, where , and is the weight coefficient; (5f) Repeat steps (5b) to (5f) until the number of training times reaches a set threshold or the value of the loss function converges, thereby obtaining a trained scoliosis screening model based on symmetric contrast learning and contour perception; (6) Test set The data is input into the trained scoliosis screening model based on symmetric contrast learning and contour perception, and the visual physical examination strategy in the model is used to obtain the spinal status: If the test focus If all four test indicators of a sample are normal, it is confirmed that the spine of the sample is normal; If the test focus If one test indicator in a sample is abnormal, it is confirmed that the sample suffers from scoliosis.
2. The method according to claim 1, characterized in that In step (1a), the flexion test videos of people of different ages, genders, and scoliosis degrees are collected. Before the test begins, the subjects are first informed of the key points of the flexion test, which include: standing with straight knees, feet together, and standing posture, with arms straight and palms together, and when bending down, lowering the head and then slowly bending forward until the upper body is 90 degrees with the ground, so as to ensure that the flexion test standard is available during the formal recording; then a common RGB camera is used to record the overall flexion action video of the subject for storage.
3. The method according to claim 1, characterized in that In step (1b), the dataset The edge mask and category label annotation of the video data in are implemented as follows: (1b1) For the data set The first frame of the video sample is annotated at the pixel level, that is, the edges of the waistline, shoulders and pelvic contour are annotated, and the binary mask labels of the waistline area and shoulder area in the frame image are generated. The bending action images of the video samples are annotated pixel-wise every ten frames to mark the edges of the back contour and generate binary mask labels for the back area in these images; (1b2) Set up a visual physical examination strategy for the dataset Label the video samples in: By observing the standing image of the overall forward bending action video of the tested person, three conditions are marked: whether the shoulders of the tested person are at the same height, whether the waist lines on both sides of the sample are symmetrical, and whether the left and right pelvises of the sample are at the same height; By observing the overall forward bending movement video of the person being tested, the condition of his / her back is found out, and whether the person being tested has a razor back in the forward bending test is marked. The back image frames with left-right asymmetry during the bending process are marked as razor back labels, and the left-right symmetrical back image frames are marked as normal back labels every other frame.
4. The method according to claim 1, characterized in that: Step (1c) performs flip preprocessing on the video data and edge mask image, which is implemented as follows: (1c1) The data set The first frame standing RGB image and its annotated mask image in the video are centered on the human body and divided into left and right parts. The back RGB image and its annotated mask image of the bending action in the video are also centered on the human body and divided into left and right parts; (1c2) Horizontally flip the right RGB image and the right mask image of the standing and bending positions; (1c3) Resize the left RGB image, the right flipped RGB image, the left mask image, and the right flipped mask image to .
5. The method according to claim 1, characterized in that In step (3a), the training set Divide into multiple batches and input them into the contour edge prediction network to obtain its semantic features and edge prediction mask image, which is implemented as follows: (3a1) The left image and the right flipped image in the training set are input into the Swin Transformer encoder of the Swin Unet contour edge prediction network. The spatial size of the feature map is gradually reduced through a hierarchical structure to extract The multi-scale features of the image are The scale features are pooled to obtain their semantic features; (3a2) The multi-scale features of the image obtained in (3a1) are input into the Swin Transformer decoder, and these multi-scale features are gradually upsampled into segmented images of the same size as the input image to obtain edge prediction mask images of the waistline area, shoulder area and back area.
6. The method according to claim 1, characterized in that In step (3b), the symmetry comparison learning loss value between the semantic features of the left image and the right flipped image of each sample in the same batch is calculated. , the implementation steps include the following: (3b1) The semantic features of the left image of each sample in the same batch are used as anchor points. When the category label of the sample is normal, the semantic features of the right flipped image are used as the first positive sample of the anchor point. ; When the category label of the sample is abnormal, the semantic features of the right-flipped image are used as the first negative sample of the anchor point ; (3b2) Copy the anchor image features as the second positive sample of the anchor ; (3b3) Rotate the anchor image feature by a small angle of 5°~10° and use the rotated image feature as the second negative sample of the anchor ; (3b4) Input the semantic features of the left image, the semantic features of the right-flipped image, and the image features of the anchor image rotated at a small angle into the symmetry contrast learning loss function, and calculate the contrast learning loss value between the semantic features of the left image and the right-flipped image of each sample in the same batch : , in, It is the set of semantic features of the left image, the right-flipped image, and the image features of the anchor image rotated at a small angle in the same batch; Representing a collection No. elements; It is Features of the anchor images of samples; It is A collection of positive sample image features of samples; It is The set of all image features of a sample except the anchor point; p represents a set No. elements; a represents a set No. elements; = , = ; is the temperature coefficient, which is used to control the model's discrimination against negative samples.
7. The method according to claim 1, characterized in that In step (3b), the Dice loss value of the edge prediction mask image and edge mask label of the same batch is calculated. , implemented as follows: (3b5) Calculate edge prediction mask image and edge mask labels The number of intersection pixels ; (3b6) According to the number of intersection pixels Calculate the Dice loss value for each region : ; in, Represents edge prediction mask image The total number of pixels, Represents the edge mask label The total number of pixels.
8. The method according to claim 1, characterized in that Step (5b) extracts the semantic features of the left and right images and the features of the left and right edge mask superimposed images, which is implemented as follows: Multiple batches of left-side images and right-side flipped images are input into the scoliosis screening network, and the multi-scale features and semantic features of the left and right images are first extracted through the SwinTransformer encoder; The multi-scale features and semantic features of the left and right images are extracted and then decoded by the input feature decoder to obtain the edge mask images of the left and right images, and the superimposed image of the left and right edge masks is obtained by pixel-by-pixel addition operation and normalization processing; The superimposed image of the left and right edge masks is finally passed through a contour encoder to output the features of the superimposed image.
9. The method according to claim 1, characterized in that: Focal loss is calculated in step (5e) and cross entropy loss , the formulas are as follows: ; ; in, is the number of samples in a single batch in the training set, is the number of categories; Yes Category The weights are used to balance the impact between different categories; is a parameter that adjusts the impact of easy and difficult samples on the loss; It is The actual category label of the sample. If the sample Belongs to category ,but ,otherwise ; is obtained after MLP and softmax operation. samples belong to the category The predicted probability of Indicates The actual labels of the samples; Indicates The predicted probability of a sample being classified correctly.
Citation Information
Patent Citations
Non-invasive scoliosis screening method and system based on back color image
CN114287915A
Scoliosis screening method based on back image
CN116258897A
Scoliosis identification method based on double contrast learning
CN118781630A