Scoliosis Recognition Method Based on Dual Contrast Learning

By adopting a dual-contrast learning-based identification method in scoliosis screening, combined with video analysis and deep learning networks, the problems of low accuracy and high cost in the existing technology are solved, and more efficient and accurate scoliosis screening is achieved.

CN118781630BActive Publication Date: 2025-05-27XIDIAN UNIV +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410872666.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2025-05-27
Estimated Expiration
2044-07-01

AI Technical Summary

Technical Problem

The prior art has low accuracy in scoliosis screening, and relying on static 2D RGB images failed to fully capture dynamic spinal curve information, and the equipment cost is high and the applicability is limited.

Method used

The scoliosis recognition method based on dual contrast learning is adopted. By analyzing the subject's ordinary RGB video during forward flexion movements, the spine dynamic information is extracted, and combined with the Swin Transformer twin network and MMPose pose estimation algorithm, the feature extraction capability of key point areas is enhanced and the cost of hardware equipment is reduced.

Benefits of technology

It improves the accuracy of scoliosis identification, reduces the possibility of missed examinations, reduces screening costs, makes it easier to popularize and promote, and provides fine-grained identification results to help doctors accurately evaluate scoliosis abnormalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118781630B_ABST
    Figure CN118781630B_ABST
Patent Text Reader

Abstract

The present invention discloses a scoliosis recognition method based on dual contrast learning, which mainly solves the problems of high recognition cost and low accuracy in the prior art. The implementation solution is as follows: collect forward flexion test videos of different ages, genders, and scoliosis degrees; set a five-step strategy for visual physical examination to label all the collected forward flexion test video samples; construct a Swin Transformer siamese network based on dual contrast learning; obtain the left and right parts of the key area based on the skeletal joints, and input them into the siamese network to extract the left and right features and overall features of the samples; enhance the left, right, and overall features of the image through the skeletal feature heat map; use the enhanced features to train the siamese network; input the test set into the trained siamese network to detect the spinal recognition result. The present invention reduces the cost of scoliosis recognition and improves the recognition accuracy, and can be used for hospital physical examinations or initial examinations by patients on their own spinal conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of deep learning and image processing, and particularly relates to a method for scoliosis recognition, which can be used for hospital physical examinations or initial examinations by patients on their own spinal conditions. Background Art

[0002] Adolescent idiopathic scoliosis, abbreviated as "adolescent scoliosis", is the main spinal disease encountered by today's adolescent population. Under normal circumstances, the human spine is straight when viewed from the front or back. If there is a "C" or "S" shaped bend to the left or right, and a standing full-spine X-ray shows a lateral curvature of more than 10 degrees in the spine, it is scoliosis. Scoliosis can lead to spinal deformities, uneven shoulders and backs, unequal scapulae, pelvic tilt, asymmetrical waistlines, and abnormal forms such as rib humps. At the same time, it affects functions such as mobility. It can even cause some mental illnesses, such as extreme introversion, depression, self-isolation, etc., due to the abnormal appearance and the inability to integrate into normal work and life. Examination results of patients with early-onset scoliosis show that the number of their alveoli is lower than that of normal people, the alveoli are over-inflated or atrophied, affecting lobes or the whole lung, and the pulmonary artery diameter is also much lower than that of peers. The thoracic volume of scoliosis patients decreases, and the thoracic volume during both inhalation and exhalation is lower than that of the normal control group. Scoliosis affects gas exchange, including local ventilation, blood flow, ventilation-perfusion ratio, diffusion, etc. It is prone to respiratory disorders such as shortness of breath and gasping, and affects blood circulation, thus affecting cardiopulmonary function.

[0003] According to statistics, the global prevalence rate ranges from 0.5% to 5.2%. Scoliosis has now become a very common phenomenon and has become the third major disease endangering children and adolescents after obesity and myopia. Traditional scoliosis screening methods, such as visual inspection and X-ray examination, provide convenient examination methods for the medical field. Visual inspection is a simple basic clinical examination, which is visually inspected by professional therapists and is a commonly used detection method for screening early scoliosis. In the forward flexion test, the tested person relaxes their hands and bends the torso forward. The doctor visually inspects the back of the tested person, and any asymmetry in the back indicates a positive test result. In the forward flexion test, rotational asymmetry of the back is easily observed, so there are many false positive and false negative results in the forward flexion test. Through the forward flexion test, doctors can observe the typical changes in the scoliosis posture and movement of the tested person. This test can only tell doctors whether the tested person may have scoliosis, and cannot accurately evaluate the degree of scoliosis of the patient. Therefore, it is also necessary to combine X-ray films to measure the degree of scoliosis - that is, the Cobb's angle, so as to completely determine whether scoliosis is present. However, due to the relatively low positive predictive value of scoliosis, many unnecessary referrals and X-ray exposures have been caused, which not only increases medical costs but also causes unnecessary psychological and physiological pressure on patients.

[0004] The patent document with the publication number CN115526845A discloses a spinal scoliosis screening method based on 2D RGB images. This method first obtains the 2D RGB image of the human back, and then uses the RVM human segmentation model to segment the human body area. Next, the Yolo V5 model is used to detect the back image, and SEResNet is used for key point recognition. Four regions are divided through these key points, so as to fit the spinal line and calculate the maximum offset distance between the fitted spinal line and the normal spinal line. Then, according to the connection line between the midpoint of the two inner shoulder points and the coccyx point, each region is divided into two left and right sub-regions, and the area difference between the two sub-regions is calculated and normalized as the contrast between the left and right regions. Finally, the contrast and the maximum offset distance are input into the discriminator network, and the discriminator outputs the back type. However, since this patent method judges spinal scoliosis by judging the contrast and the maximum offset distance between the left and right regions, not only is the judgment complicated, but the recognition accuracy is also relatively low.

[0005] The patent document with the publication number CN114287915A discloses a non-invasive spinal scoliosis screening method and system based on back color images. This method first collects the RGB-D image of the back of the human body and saves it as a depth image and a color image. Then, the Mask-RCNN network is used for training to obtain a human body segmentation model. Next, the YOLOv5 network is trained to obtain a back recognition model. According to the corresponding relationship between the depth image and the color image, the back region in the depth image is extracted to obtain the depth map of the human back, and the maximum ATR angle value is calculated. Finally, a classification standard is formulated according to the ATR angle value and category labels are marked, and the back color image and the category labels are input into the EfficientNet network for training to obtain a spinal classification model. However, this method uses an RBG-D camera, which is expensive and has a high cost. At the same time, the implementation process is complex and the efficiency is low, and the recognition accuracy is not high.

[0006] In addition, the above existing methods only rely on the static 2D RGB images of the human back when standing, without analyzing the dynamic forward flexion video of the subject, ignoring the subtle manifestations of spinal scoliosis in the bending action state. Therefore, the information for judging human spinal scoliosis is incomplete, resulting in possible missed detections. Summary of the Invention

[0007] The purpose of the present invention is to address the above deficiencies of the existing technologies, and propose a spinal scoliosis recognition method based on dual contrast learning, so as to analyze the spinal dynamic information of the subject in the ordinary RGB video of the forward flexion action, capture the posture changes of the subject in the forward flexion action, improve the accuracy of spinal scoliosis recognition, and reduce the possibility of missed detections; by using a low-cost ordinary RGB camera, reduce the screening cost and make it easier to popularize and promote.

[0008] To achieve the above object, the implementation scheme of the present invention includes the following:

[0009] (1) Collect forward flexion test videos of different ages, genders, and degrees of scoliosis;

[0010] (2) Set a five-step strategy for visual physical examination, and label all the collected forward flexion test video samples, that is, label whether the shoulders of the samples are at the same height, whether the left and right scapulas of the samples are at the same height, whether the bilateral lumbar depressions of the samples are symmetrical, whether the left and right pelvises of the samples are at the same height, and whether there is a rib hump in the forward flexion test for these 5 conditions;

[0011] (3) Divide all the labeled forward flexion test videos into multiple batches for input, identify the skeletal joint points according to the existing pose estimation algorithm MMpose, and extract the key regions of scoliosis for each batch of samples to be measured, that is, extract the left and right parts of the regions including the shoulders, scapulas, back, waistlines, and pelvis:

[0012] (3a) Extract the standing frames of the tested persons from the forward flexion test videos;

[0013] (3b) Use the Faster RCNN method to perform object detection on the human body, and use the MMPose pose estimation algorithm to extract the coordinates of the skeletal key points such as the shoulders and hips for the detected bounding boxes, and intercept the human body regions including the shoulders, scapulas, bilateral lumbar depressions, pelvis, and back;

[0014] (3c) Calculate the center point of the human body according to the extracted skeletal key points of the shoulders and hips, and divide the human body region to be measured into left and right parts by the vertical line passing through the center point;

[0015] (3d) Divide these batches of samples into a training set and a test set according to a ratio of 8:2;

[0016] (4) Based on the existing Swin Transformer network, construct a SwinTransformer siamese network based on dual contrast learning, perform dual contrast on the overall features between samples and the left and right part features within samples, that is, between-class and within-class features, and the two branches of this siamese network share the same weight;

[0017] (5) Train the Swin Transformer siamese network based on dual contrast learning:

[0018] (5a) Input the left and right parts of the human body regions of the samples in the same batch in the training set into the two branches of the SwinTransformer siamese network respectively, extract the features of the left and right parts respectively, and perform a Concatenate operation on them to obtain the overall human body features of this batch of samples;

[0019] (5b) Use the MMPose pose estimation algorithm to output the feature heatmaps of the shoulders and pelvis key points in the human body region of this batch of samples, and enhance the features of the left and right parts and the overall features obtained in step (5a) through this heatmap;

[0020] (5c) Calculate the intra-class contrastive learning loss value L between the enhanced left and right part features of each sample in the same batch 1 , and calculate the inter-class contrastive learning loss value L between the overall feature of a single enhanced sample in the same batch and the overall features of other enhanced samples 2 ;

[0021] (5d) Input the overall image features enhanced in step (5c) into the MLP, and then through the softmax operation, obtain the predicted class results of all samples in this batch. Then input them together with the actual class labels of the samples in this batch into the cross-entropy loss function to calculate the value of the cross-entropy loss L 3 ;

[0022] (5e) Perform backpropagation on the Swin Transformer Siamese network, calculate the loss function L = αL 1 + βL 2 + γL 3 , where α + β + γ = 1. Adjust the weight parameters of the model by minimizing L until the number of training times reaches the set threshold or the value of the loss function converges to obtain the trained scoliosis screening model;

[0023] (6) Input the test set into the trained scoliosis screening model, and use the five-step visual physical examination strategy in the model to obtain the spinal state:

[0024] If all five test indicators of a sample in the test set are normal, confirm that the spine of this sample is normal;

[0025] If there is one abnormal test indicator in a sample in the test set, confirm that this sample has scoliosis.

[0026] The present invention has the following advantages compared with the prior art:

[0027] 1. It can provide fine-grained recognition results:

[0028] Based on the "five-step visual physical examination method" for clinical diagnosis of scoliosis, the present invention can obtain specific abnormal manifestation areas through in-depth analysis of video data, so as to describe in detail the external manifestations of scoliosis. Such fine-grained recognition results help doctors accurately evaluate the abnormal manifestation sites of patients' scoliosis and provide important references for doctors to subsequently judge the specific categories of scoliosis.

[0029] 2. By analyzing the dynamic forward flexion video of the subject, the present invention determines whether the back is symmetric during the downward bending process, which can provide an important basis for judging whether the tested person has scoliosis.

[0030] 3. The present invention fuses the feature heatmaps of the two shoulders and pelvis key points output by the pose estimation and the image features extracted by the Swin Transformer through element-wise multiplication, enhancing the features of the model in the key point areas, seamlessly integrating the pose information into the image representation, and significantly improving the feature extraction ability of the model for the key point areas.

[0031] 4. Since the present invention constructs the intra-class contrast loss and the inter-class contrast loss to constrain the model to learn the relevant information for scoliosis judgment, the intra-class contrast loss can be used to compare the symmetry differences between the left and right parts of the same sample to learn the symmetry information between the left and right under the sample, and the inter-class contrast loss can be used to compare the features of different category samples to learn the similarity between samples within the same category and the differences between samples of different categories.

[0032] 5. The present invention conducts intelligent identification of scoliosis based on ordinary RGB videos, so the cost of the required hardware equipment is very low, which is suitable for large-scale screening of teenagers. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is the implementation flowchart of the present invention;

[0034] Figure 2 is the structural block diagram of the scoliosis screening model in the present invention;

[0035] Figure 3 is the heatmap of the two-shoulder bone points in the present invention;

[0036] Figure 4 is the heatmap of the pelvis bone points in the present invention;

[0037] Figure 5 is the standing schematic diagram of the tested person in the forward flexion test in the present invention;

[0038] Figure 6 is the bending schematic diagram of the tested person in the forward flexion test in the present invention.

[0039] Figure 7 is the confusion matrix diagram of the test set results in the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0040] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0041] It should be noted that the step numbers in the specification, claims and claims of the present invention are only for clearly describing the implementation solutions of the present invention for easy understanding, and the order of their serial numbers is not limited.

[0042] Refer to Figure 1 , the implementation steps of this example are as follows:

[0043] Step 1, collect forward flexion test videos of different ages, genders and degrees of scoliosis:

[0044] (1.1) Set the collection requirements for the forward flexion test:

[0045] The upper body of male test subjects is bare, and the upper body of female test subjects wears underwear;

[0046] The test subjects take off their shoes, face a clean wall, take a natural standing posture, with their feet shoulder-width apart, look straight ahead, let their arms hang naturally, and the palms face inwards;

[0047] (1.2) Before the collection starts, inform the test subjects of the collection requirements for the forward flexion test to ensure that the forward flexion test standards are available during the formal recording:

[0048] (1.3) After the collection starts, the test subjects start the forward flexion movement, including straightening the knees, closing the feet, standing at attention, straightening the arms and clasping the hands, lowering the head and then slowly bending forward to about 90°, and gradually placing the clasped hands between the knees. The tester shoots the overall forward flexion movement video of the test subjects for preservation.

[0049] Step 2, set the five-step strategy for visual physical examination, and label all the forward flexion test video samples collected.

[0050] The first step, observe whether the shoulders of the test subjects in the standing image are at the same height. If the shoulders are at the same height, label them as normal shoulders. If the shoulders are not at the same height, label them as abnormal shoulders;

[0051] The second step, observe whether the left and right scapulas of the test subjects in the standing image are at the same height. If the scapulas are at the same height, label them as normal scapulas. If the scapulas are not at the same height, label them as abnormal scapulas;

[0052] In the third step, observe whether the two lumbar concavities of the person being measured in the standing image are symmetrical. If the two lumbar concavities are symmetrical, it is labeled as normal lumbar concavity; if the two lumbar concavities are asymmetrical, it is labeled as abnormal lumbar concavity;

[0053] In the fourth step, observe whether the left and right pelvises of the person being measured in the standing image are symmetrical. If the left and right pelvises are symmetrical, it is labeled as normal pelvis; if the left and right pelvises are asymmetrical, it is labeled as abnormal pelvis;

[0054] In the fifth step, find the condition of the back of the person being measured by observing the overall forward flexion movement video of the person being measured, and label the back image frames with left - right asymmetry during the bending process as the "razor - back" label, and label the back image frames with left - right symmetry at every other frame as the "normal - back" label, where the state of bending is as Figure 6 shown.

[0055] Step 3: Divide all the forward flexion test videos after annotation into multiple batches for input. According to the existing pose estimation algorithm MMpose, identify the skeletal joint points, and extract the key regions of scoliosis for each batch of measured samples, that is, extract the left and right parts of the regions including the shoulders, scapulas, back, waistlines, and pelvis regions respectively.

[0056] (3.1) Extract the first frame from the forward flexion test video as the standing image of the person being measured, as Figure 5 shown;

[0057] (3.2) Use Faster RCNN to perform object detection on the forward flexion test video and the standing image to obtain the human detection box results:

[0058] (3.2.1) Generate a series of candidate regions that may contain a human body through the Region Proposal Network (RPN). These regions are marked as potential target positions in the image;

[0059] (3.2.2) Use the convolutional neural network VGG16 trained with human image data to extract high - level features of these candidate regions, and input these features into the classification layer and regression layer of this network;

[0060] (3.2.3) Determine whether each candidate region contains a human body through the classification layer, determine the coordinates of the bounding box containing the human body through the regression layer, and use non - maximum suppression on the classification and regression results to remove redundant detection boxes to obtain the most accurate human detection box results.

[0061] (3.3) Use the MMPose pose estimation algorithm to generate the coordinates of the two shoulders (X 1 , Y 1 ), (X 2 , Y 2 ) and the coordinates of the hip key points (X 3 , Y3 ),(X 4 ,Y 4 );

[0062] (3.4)Extract the areas including the shoulders, scapulae, bilateral lumbar concavities, pelvis, and the back area:

[0063] (3.4.1)Extract the coordinate points A at the four corners in the human body rectangular area including the shoulders, scapulae, bilateral lumbar concavities, and pelvis according to the shoulder and hip coordinates and the proportional relationship 1 , A 2 , A 3 , A 4 , which are respectively expressed as:

[0064] A 1 = (X 1 - 0.25×(X 2 - X 1 ), Y 1 - 0.25×(Y 3 - y 1 ))

[0065] A 2 = (X 2 + 0.25×(X 2 - X 1 ), Y 1 - 0.25×(Y 3 - Y 1 ))

[0066] A 3 = (X 1 - 0.25×(X 2 - X 1 ), Y 3 + 0.25×(Y 3 - Y 1 ))

[0067] A 4 = (X 2 + 0.25×(X 2 - X 1 ), Y 3 + 0.25×(Y 3 - Y 1 ));

[0068] (3.4.2)Extract the coordinate points B at the four corners of the back rectangular area according to the shoulder and hip coordinates and the proportional relationship 1 , B 2 , B 3 , B 4 , which are respectively expressed as:

[0069] B1 =(X 3 -0.25×(X 4 -X 3 ),Y 3 -0.5×(X 4 -X 3 ))

[0070] B 2 =(X 3 +0.25×(X 2 -X 1 ),Y 3 -0.5×(X 4 -X 3 ))

[0071] B 3 =(X 3 -0.25×(X 4 -X 3 ),Y 3 )

[0072] B 4 =(X 3 +0.25×(X 2 -X 1 ),Y 4 );

[0073] (3.5) Calculate the center point coordinates of the human body:

[0074]

[0075] (3.6) Divide the human body area to be measured into left and right parts by the vertical lines of the center point coordinates X and Y.

[0076] Step 4, construct a Swin Transformer Siamese network based on dual contrast learning.

[0077] Based on the existing Swin Transformer network, construct a Swin Transformer Siamese network based on dual contrast learning. As Figure 2 shown, the left and right branches respectively use the Swin Transformer feature encoder to extract the features of the left and right sides of the sample image, and perform a Concatenate operation on them to obtain the overall feature of the sample image. Perform dual contrast on the overall feature between samples and the left and right part features within the sample, that is, between-class and within-class features. The two branches of this Siamese network share the same weight;

[0078] Step 5, train the Swin Transformer Siamese network based on dual contrast learning:

[0079] (5.1) Input the left and right parts of the human body regions of the same batch of samples in the training set into the two branches of the SwinTransformer Siamese network respectively. After passing through the four-layer cascaded feature extractors in each of the two branches of the Swin Transformer network, extract the features of the left and right parts respectively, and perform a Concatenate operation on them to obtain the overall human body features of this batch of samples;

[0080] (5.2) Use the trained HRnet network in the MMPose library to extract the bone features of the human body region; then perform a convolution operation on this feature to generate heatmaps of the bone key points of the shoulders and hips, and the brighter regions in this image represent the high confidence of the positions of the bone key points, and its output scale is 224×224, as Figure 3 、 Figure 4 shown.

[0081] (5.3) Enhance the features of the left and right parts and the overall features through this heatmap:

[0082] (5.3.1) Perform a 4×4 convolution operation on the bone feature heatmap to reduce its scale from 224×224 to 56×56, and then match its dimension with the feature extracted by the first layer of the Swin Transformer network obtained in step (5.1);

[0083] (5.3.2) Perform a 2×2 average pooling operation on the reduced bone feature heatmap to reduce its scale from 56×56 to 28×28, 14×14, and 7×7 in sequence, and then match its dimension with the features extracted by the second, third, and fourth layers of the Swin Transformer network obtained in step (5.1);

[0084] (5.3.3) Perform an element-wise multiplication operation on each layer of the features output from the Swin Transformer network in step (5.1) and the bone features with matching scales extracted in (5.3.1) and (5.3.2) to seamlessly integrate the bone information into each layer of the features of the Swin Transformer, and obtain the enhanced features of the left and right parts.

[0085] (5.4) Calculate the intra-class contrastive learning loss value L between the enhanced left and right part features of each sample in the same batch 1 :

[0086] (5.4.1) Use the enhanced left part feature of each sample image in the same batch as the anchor point, and construct positive and negative samples for intra-class contrastive learning of the right part image feature according to the class label of the sample:

[0087] If the class annotation label of the sample is normal, the right - hand - side image features are used as the first positive sample for intra - class contrastive learning;

[0088] If the class annotation label of the sample is abnormal, the right - hand - side image features are used as the first negative sample for intra - class contrastive learning;

[0089] (5.4.2) Horizontally flip the anchor image features, and use the horizontally - flipped image features as the second positive sample for intra - class contrastive learning;

[0090] (5.4.3) Horizontally flip and then perform a small - angle rotation on the anchor image features in sequence, and use the flipped and rotated image features as the second negative sample for intra - class contrastive learning;

[0091] (5.4.4) Input the features of the left and right parts of the enhanced image, the horizontally - flipped image of the anchor, and the small - angle - rotated image into the intra - class contrastive learning loss function, and calculate the intra - class contrastive learning loss value \(L\) between the features of the left and right parts of each sample after enhancement in the same batch 1 :

[0092]

[0093] where \(K\) is the set of the features of the left and right parts of the enhanced image, the horizontally - flipped image features of the anchor, and the small - angle - rotated image features in the same batch, \(z\) k is the feature of the anchor image of the \(k\) - th sample, \(P(k)\) is the set composed of the first and second positive - sample picture features of the \(k\) - th sample, \(A(k)\) is the set of all image features except the anchor in the \(k\) - th sample, \(z\) p represents the feature of the positive sample, \(z\) a represents the feature representations of all positive and negative samples; \(\tau\) is the temperature coefficient, which is used to control the discrimination degree of the model for negative samples;

[0094] (5.5) Calculate the inter - class contrastive learning loss value \(L\) between the overall features of a single sample after enhancement and the overall features of other samples after enhancement in the same batch 2 :

[0095] (5.5.1) Use the overall image features of a single sample after enhancement in this batch as the anchor;

[0096] (5.5.2) Use the overall image features of other samples with the same class label as the anchor sample in this batch after enhancement as the positive samples for inter - class contrastive learning, and use the overall image features of other samples with different class labels from the anchor sample in this batch after enhancement as the negative samples for inter - class contrastive learning;

[0097] (5.5.3) Input the anchors and positive and negative samples of this batch into the inter-class contrastive learning, and calculate the inter-class contrastive learning loss value L between the overall feature of a single sample after enhancement and the overall features of other samples after enhancement in this batch. 2 :

[0098]

[0099] Where K is the set of overall image features after enhancement in the same batch, z k is the feature when the enhanced image feature of the k-th sample is used as an anchor, P(k) is the set of positive samples when the enhanced image feature of the k-th sample is used as an anchor, A(k) is the set of positive and negative samples when the enhanced image feature of the k-th sample is used as an anchor, z p represents the feature of the positive sample, z a represents the feature representations of all positive and negative samples; τ is the temperature coefficient, which is used to control the discrimination of the model for negative samples.

[0100] (5.6) Input the overall image features enhanced in step (5.3) into the multi-layer perceptron MLP, and then through the softmax operation, obtain the predicted class results of all samples in this batch;

[0101] (5.7) Input the sample predicted class results of this batch and the actual class labels together into the cross-entropy loss function, and calculate the value of the cross-entropy loss L 3 :

[0102]

[0103] Where B is the number of samples in the batch, and C is the number of classes. y i,c is the actual class label of the i-th sample. If sample i belongs to class c, then y i,c = 1, otherwise y i,c = 0; p i,c is the predicted probability that the i-th sample belongs to class c obtained through the MLP and softmax operations.

[0104] (5.8) Perform backpropagation on the Swin Transformer Siamese network, and calculate the loss function L = αL 1 + βL 2 + γL 3 , where α + β + γ = 1. Adjust the weight parameters of the model by minimizing L until the number of training times reaches the set threshold or the value of the loss function converges, and obtain the trained scoliosis screening model.

[0105] Step 6, input the test set into the trained scoliosis screening model, and use the five-step visual physical examination strategy in the model to obtain the spinal state:

[0106] If the five test indicators of a sample in the test set are all normal, it is confirmed that the spine of the sample is normal;

[0107] If there is one abnormal test indicator in a sample in the test set, it is confirmed that the sample has scoliosis.

[0108] The effect of the present invention can be further illustrated by the following test experiments:

[0109] I. Test conditions:

[0110] The test equipment is a computer and a camera. The computer and the camera are connected and set up 2 to 3 meters directly behind the person being tested for scoliosis.

[0111] The test data is the forward flexion test videos of 90 tested persons collected in a certain hospital. All the tested persons are labeled according to the five-step visual physical examination method, and the labels are all marked by hospital doctors.

[0112] The simulation test platform uses an Intel(R) Core(TM) CPU E5-2683v4@2.10GHz, equipped with an NVIDIA TITAN Xp graphics card, and the memory capacity is 12GB. This platform is based on the Linux operating system and is implemented using the Python language.

[0113] Evaluation indicators: Conduct the "five-step visual physical examination method" test to obtain a series of results, including whether the shoulders are at the same height, whether the left and right scapulas are at the same height, whether the bilateral lumbar depressions are symmetrical, whether the left and right pelvises are at the same height, and whether there is a rib hump deformity, etc., and judge whether there is scoliosis.

[0114] II. Test content and results:

[0115] Under the above test conditions, the method of the present invention is used to process the forward flexion test video data of 90 tested persons. 72 samples are used as the training set to train the scoliosis recognition model constructed by the present invention. The trained scoliosis recognition model is used to test the remaining 18 samples, that is, to test whether the shoulders are at the same height, whether the scapulas are at the same height, whether the bilateral lumbar depressions are symmetrical, whether the left and right pelvises are symmetrical, whether there is a rib hump deformity, and whether there is scoliosis. The confusion matrix of the test results is as Figure 7 shown, and the respective accuracies are calculated. The results are as follows:

[0116]

[0117] The data in Table 1 clearly show that the method of the present invention demonstrated highly reliable performance in the tests. The accuracy rate of all test results reached over 83%, and especially in the aspect of judging whether there is scoliosis, the accuracy rate reached 94.44%. This indicates that the present invention has excellent accuracy and reliability in identifying scoliosis, providing strong support for the early detection and screening of scoliosis.

Claims

1. A scoliosis recognition method based on double contrast learning, characterized in that: These include: (1) Collect flexion test videos of different ages, genders, and degrees of scoliosis; (2) A five-step visual physical examination strategy was set up to label all collected forward flexion test video samples, namely, whether the shoulders of the sample were at the same height, whether the left and right shoulder blades of the sample were at the same height, whether the two sides of the sample's lumbar concavities were symmetrical, whether the left and right pelvis of the sample were at the same height, and whether the sample had a razor back in the forward flexion test; (3) All the annotated forward flexion test videos are divided into multiple batches for input. The bone joints are identified according to the existing posture estimation algorithm MMpose, and the key areas of scoliosis are extracted for each batch of tested samples, that is, the left and right parts including the shoulder, scapula, back, waistline, and pelvic area are extracted respectively: (3a) Extracting the standing frame of the subject from the forward bending test video; (3b) The Faster RCNN method is used to detect the human body, and the coordinates of the key bone points such as the shoulders and hips are extracted from the detected bounding box using the MMPose posture estimation algorithm, and the human body area including the shoulders, shoulder blades, waist concavities on both sides, pelvis and back is cut out; (3c) calculating the center point of the human body according to the extracted key points of the shoulder and hip bones, and dividing the human body area to be measured into left and right parts by a vertical line passing through the center point; (3d) Divide the samples of these batches into training sets and test sets in a ratio of 8:2; (4) Based on the existing Swin Transformer network, a Swin Transformer twin network based on double contrast learning is constructed to perform double contrast on the overall features between samples and the left and right partial features within the samples. The two branches of the twin network share the same weight; (5) Training the Swin Transformer twin network based on dual contrast learning: (5a) The left and right parts of the human body region of the same batch of samples in the training set are input into the two branches of the Swin Transformer twin network respectively, the features of the left and right parts are extracted respectively, and the concatenation operation is performed on them to obtain the overall human body features of the batch of samples; (5b) Using the MMPose posture estimation algorithm to output the characteristic heat map of the key points of the shoulders and pelvis in the human body area of ​​the batch of samples, the characteristics of the left and right parts and the overall characteristics obtained in step (5a) are enhanced through the heat map; (5c) Calculate the intra-class contrast learning loss value L1 between the left and right partial features of each sample after enhancement in the same batch, and calculate the inter-class contrast learning loss value L2 between the overall features after enhancement of a single sample and the overall features after enhancement of other samples in the same batch; (5d) Input the overall image features enhanced in step (5c) into the MLP, and then perform a softmax operation to obtain the predicted category results of all samples in the batch, and then input them together with the actual category labels of the samples in the batch into the cross entropy loss function to calculate the value of the cross entropy loss L3; (5e) Back-propagating the Swin Transformer twin network, calculating the loss function L = αL1 + βL2 + γL3, where α + β + γ = 1, and adjusting the weight parameters of the model by minimizing L until the number of training times reaches the set threshold or the value of the loss function converges, thereby obtaining a trained scoliosis screening model; (6) Input the test set into the trained scoliosis screening model and use the five-step visual physical examination strategy in the model to derive the spinal status: If all five conditions of a sample in the test set are normal, it is confirmed that the spine of the sample is normal; If one condition is abnormal in a sample in the test set, the sample is confirmed to have scoliosis.

2. The method according to claim 1, characterized in that In step (1), flexion test videos of different ages, different genders, and different degrees of scoliosis are collected, and the implementation is as follows: Before the test begins, determine the key points of the forward flexion test for the person being tested to ensure that the forward flexion test standards are available during the formal recording; The person being tested begins the forward bending movement, including straightening the knees, putting the feet together, standing at attention, stretching the arms and putting the palms together, lowering the head and slowly bending forward to about 90 degrees, putting the palms together and gradually placing them between the knees; Record the overall forward bending movement video of the person being tested and save it.

3. The method according to claim 1, characterized in that In step (2), whether the sample has a razor back in the forward bending test is marked by observing the overall forward bending movement video of the subject to find out the condition of his back, and marking the back image frames with left-right asymmetry during the bending process as razor back labels, and marking the left-right symmetrical back image frames as normal back labels every other frame.

4. The method according to claim 1, characterized in that: In step (3b), the Faster RCNN method is used to detect human targets, which is implemented as follows: (3b1) Generate a series of candidate regions that may contain human bodies through the region proposal network RPN. These regions are marked as potential target locations in the image; (3b2) using the convolutional neural network VGG16 trained with human image data to extract high-level features of these candidate regions, and inputting these features into the classification layer and regression layer of the network; (3b3) The classification layer is responsible for determining whether each candidate region contains a human body, and the regression layer performs bounding box regression on the candidate region to further accurately locate the boundary of the human body. The classification and regression results are suppressed using non-maximum values ​​to remove redundant detection frames and retain only the most accurate human detection results.

5. The method according to claim 1, characterized in that: Step (3b) uses the MMPose posture estimation algorithm to extract the coordinates of the key points of the bones of the shoulders and hips from the detected bounding box. The key point prediction layer in MMPose is applied to the feature map to generate the coordinates of the shoulders (X1, Y1), (X2, Y2) and the coordinates of the key points of the hips (X3, Y3), (X4, Y4), and the center coordinates of the human body are obtained: X=(X1+X2+X3+X4) / 4, Y=(Y1+Y2+Y3+Y4) / 4, used to cut the left and right parts of the human body; In step (3b), the human body region including shoulders, shoulder blades, waist concave areas on both sides, and pelvis is cut out by cutting out the four corner coordinate points A1, A2, A3, and A4 of the human body rectangular region including shoulders, shoulder blades, waist concave areas on both sides, and pelvis through the position and proportional relationship of the key points, which are respectively expressed as: A1=(X1-0.25×(X2-X1), Y1-0.25×(Y3-Y1)), A2=(X2+0.25×(X2-X1),Y1-0.25×(Y3-Y1)), A3=(X1-0.25×(X2-X1),Y3+0.25×(Y3-Y1)), A4=(X2+0.25×(X2-X1), Y3+0.25×(Y3-Y1)); The human body region of the back is captured in step (3b) by capturing the coordinate points B1, B2, B3, and B4 of the four corners of the rectangular region of the back through the information of the key points, which are respectively: B1=(X3-0.25×(X4-X3),Y3-0.5×(X4-X3)) B2=(X3+0.25×(X2-X1),Y3-0.5×(X4-X3)) B3=(X3-0.25×(X4-X3),Y3) B4=(X3+0.25×(X2-X1),Y4).

6. The method according to claim 1, characterized in that In step (5b), the MMPose posture estimation algorithm is used to output the characteristic heat map of the shoulder and pelvis key points in the human body area of ​​the batch of samples. The HRnet network trained in the MMPose library is first used to extract the bone features of the human body area; then the features are convolved to generate an image of the confidence of the shoulder and hip bone key points, and the brighter areas in the image represent high confidence in the position of the bone key points, and the output scale is 224×224.

7. The method according to claim 1, characterized in that In step (5b), the features of the left and right parts and the overall features are enhanced by using the heat map, as follows: (5b1) Perform a 4×4 convolution operation on the bone feature heat map to reduce its scale from 224×224 to 56×56 to match the feature dimension extracted by the first layer of the Swin Transformer network; (5b2) Perform a 2×2 average pooling operation on the reduced skeleton feature heat map to reduce its scale from 56×56 to 28×28, 14×14, and 7×7 in sequence to match the feature dimensions extracted by the second, third, and fourth layers of the Swin Transformer network; (5b3) Perform element-by-element multiplication of each layer of features output from the Swin Transformer network in step (4) with the scale-matched bone features extracted in steps (5b1) and (5b2) to seamlessly integrate the bone information into each layer of features of the Swin Transformer to obtain enhanced left and right part features.

8. The method according to claim 1, characterized in that In step (5c), the intra-class contrast learning loss value L1 between the left and right part features of each sample in the same batch after enhancement is calculated, which is implemented as follows: (5c1) The enhanced left half of each sample image in the same batch is used as an anchor point. When the class label of the sample is normal, the right half of the image feature is used as a positive sample for intra-class contrast learning; when the class label of the sample is abnormal, the right half of the image feature is used as a negative sample for intra-class contrast learning; (5c2) The image features obtained by horizontally flipping the anchor image features are used as positive samples for intra-class contrastive learning; (5c3) The anchor image features are horizontally flipped and then rotated at a small angle as negative samples for intra-class contrastive learning; (5c4) Input the features of the left and right parts of the enhanced image, the image with the anchor point horizontally flipped, and the image rotated at a small angle into the intra-class contrastive learning loss function, and calculate the intra-class contrastive learning loss value L1 between the features of the left and right parts of each sample after enhancement in the same batch: Among them, K is the set of left and right parts of the enhanced image, the image with horizontal flip of the anchor point, and the image features of small-angle rotation in the same batch, z k is the feature of the anchor image of the kth sample, P(k) is the set of positive sample image features of the kth sample, A(k) is the set of all image features in the kth sample except the anchor point, z p Represents the characteristics of the positive sample, z a Represents the feature representation of all positive and negative samples; τ is the temperature coefficient, which is used to control the model's discrimination of negative samples.

9. The method according to claim 1, characterized in that: In step (5c), the inter-class contrast learning loss value L2 between the overall features of a single sample after enhancement and the overall features of other samples after enhancement in the same batch is calculated as follows: (5c5) taking the enhanced overall image features of a single sample in the batch of samples as an anchor point; (5c6) The overall image features of other samples in the batch with the same category label as the anchor sample are enhanced as positive samples for inter-class contrast learning, and the overall image features of other samples in the batch with different category labels as the anchor sample are enhanced as negative samples for inter-class contrast learning; (5c7) Input the anchor points and positive and negative samples of this batch into the inter-class contrastive learning, and calculate the inter-class contrastive learning loss value L2 between the overall features of a single sample after enhancement and the overall features of other samples after enhancement in this batch: Among them, K is the set of enhanced overall image features in the same batch, z k is the feature of the enhanced image feature of the kth sample as the anchor point, P(k) is the set of positive samples when the enhanced image feature of the kth sample is used as the anchor point, A(k) is the set of positive and negative samples when the enhanced image feature of the kth sample is used as the anchor point, z p represents the characteristics of the positive sample, z a Represents the feature representation of all positive and negative samples; τ is the temperature coefficient, which is used to control the model's discrimination of negative samples.

10. The method according to claim 1, characterized in that The value of the cross entropy loss L3 is calculated in step (5d) as follows: Where B is the number of samples in the batch, C is the number of categories, and y i,c is the actual category label of the i-th sample. If sample i belongs to category c, then y i,c =1, otherwise y i,c =0; p i,c It is the predicted probability that the i-th sample belongs to category c after MLP and softmax operation.

Citation Information

Patent Citations

  • Scoliosis screening method based on 2D RGB image

    CN115526845A

  • Non-invasive scoliosis screening method and system based on back color image

    CN114287915A

  • Automatic detection system for scoliosis

    CN114983396A