Back segmentation method based on improved mask rcnn
By improving the Mask RCNN model, the automatic segmentation of the naked back of the human body is solved by using the explicit vision center (EVC) module, and the problem of time-consuming and labor-intensive manual evaluation and low accuracy in manual labeling in scoliosis screening is solved, achieving an efficient and accurate screening process.
Patent Information
- Application Number
- PCT/CN2023/136325
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-12
AI Technical Summary
The prior art relies on manual evaluation in scoliosis screening, which is time-consuming and labor-intensive, and the accuracy of manual labeling of images is low, affecting the results of subsequent algorithms.
By improving the Mask RCNN model, the explicit vision center (EVC) module is used to obtain features that have both global information and local information, and automatically segment the exposed back of the human body to reduce dependence on manual annotation.
It achieves efficient and accurate scoliosis screening process, reduces the time and energy of manual labeling, and improves the accuracy of subsequent algorithms.
Smart Images

Figure CN2023136325_12062025_PF_FP_ABST
Abstract
Description
A back segmentation method based on improved Mask RCNN Technical Field
[0001] The present invention belongs to the technical field of medical image segmentation in computer vision, and in particular relates to a back segmentation method based on an improved Mask RCNN. Background Art
[0002] Scoliosis refers to a C-shaped or S-shaped spine to the left or right. Adolescent idiopathic scoliosis (AIS) is one of the most common types of scoliosis, and approximately 2% to 3% of adolescents worldwide suffer from idiopathic scoliosis. Today, many different types of assessment methods have been proposed for scoliosis, among which visual inspection, lordosis test, and ripple image inspection are more commonly used scoliosis assessment methods. The assessment results obtained by applying these direct and easy-to-implement methods provide doctors with a basis for screening for scoliosis. However, manually using these methods to assess scoliosis in subjects is time-consuming and labor-intensive for both doctors and subjects.
[0003] With the continuous advancement of deep learning algorithms in computer vision, numerous object detection and instance segmentation algorithms have been developed. Many of these algorithms have been successfully applied to various medical scenarios, such as screening for eye diseases using retinal lesion images, diagnosing rare genetic diseases using facial images, and diagnosing cardiovascular and cerebrovascular diseases using retinal images. However, there is currently limited research on the application of computer vision algorithms to scoliosis screening. Developing scoliosis screening algorithms has two main benefits. First, by using object detection and instance segmentation algorithms, computers can determine whether a patient has scoliosis based on back images, rather than visually observing the patient's back. This significantly simplifies the screening process. Second, since the computer only needs to analyze images of the patient's exposed back to complete the screening task, the spinal Cobb angle can be calculated, eliminating the need for X-rays and significantly reducing the risk of cancer in non-scoliotic individuals with negative screening results. However, scoliosis screening requires a large number of labeled images of the exposed, flexed back. Obtaining these large numbers of images typically requires numerous physicians and volunteers with specialized expertise to label the collected images, which is extremely time-consuming and labor-intensive. More importantly, the accuracy of the manual annotation process can significantly impact the accuracy of subsequent region of interest extraction, thus affecting the results of the screening algorithm in unexpected ways. To address this issue, the present invention fine-tunes the improved Mask RCNN model using annotated naked forward-flexed back images. This fine-tuned model can automatically segment the subject's back region, further improving the efficiency and accuracy of subsequent scoliosis screening.
[0004] Mask R-CNN is a framework based on Faster R-CNN. This framework adds a fully connected segmentation network after the base feature network, transforming the original classification + regression task into a classification + regression + segmentation task. Mask R-CNN uses the same two-stage approach as Faster R-CNN. The first stage scans the image and generates proposals (regions that may contain an object). The second stage classifies the proposals and generates bounding boxes (bboxes) and masks. In Mask RCNN, the image is input into a feature extraction network to extract features and obtain the corresponding feature map. The generated feature map is then input into a region proposal network (RPN) to select the target region. The RPN is a lightweight neural network that maps all pixels in the shared feature map output by the backbone network back to the corresponding receptive field in the input image. It then creates rectangular boxes (anchors) of varying sizes and aspect ratios on the input image, with the center of the receptive field at the center. These anchors overlap to cover the image as much as possible. The RPN generates two outputs for each anchor: the anchor category, used to distinguish foreground from background, and the bounding box accuracy, which better fits the object. By using the prediction of RPN, the best anchor containing the target can be selected and its position and size can be fine-tuned. If there are multiple anchors overlapping each other, the anchor with the highest foreground score is retained as the proposal through non-maximum suppression. After the bounding box fine-tuning step in RPN, the region of interest (RoI) mapped to the feature map by the proposal can be of different sizes. However, since the classifier can only process fixed-size inputs and cannot handle multi-size inputs well, RoIAlign is needed to solve this problem. This method samples at different points in the RoI feature map and applies bilinear interpolation to resize the RoI to a fixed size. Finally, the RoI adjusted to a fixed size is input into two parallel branches respectively to classify the RoI, perform bbox regression, and generate a mask.
[0005] Mask RCNN uses a feature extraction network composed of a convolutional neural network (CNN) and a feature pyramid network (FPN). The convolutional layer extracts features, and the intermediate output is a bottom-up feature map of different scales, which is laterally connected with the top-down feature map of the FPN through feature fusion. CNN-based methods are widely used, but due to their local perception and parameter sharing characteristics, the network is affected by a limited receptive field, making it difficult to capture long-distance dependencies. Therefore, the features output by the network may lose global information of the image, which affects the accuracy of the image segmentation results. The FPN network focuses on the interaction of features between (different) layers, but ignores the relationship between features within (the same) layer, which has been proven to be beneficial for visual recognition tasks.
[0006] Summary of the Invention
[0007] In order to overcome the shortcomings of the prior art, the present invention provides a back segmentation method based on an improved Mask RCNN, which includes the following main steps:
[0008] Step 1: Data collection: Collect images of teenagers with exposed backs flexed forward from different regions, ages, genders, and scoliosis levels to obtain a color image dataset.
[0009] Step 2: Labeling: Manually label the exposed back area of the human body in the color image to obtain the labeling file corresponding to the original color image.
[0010] Step 3: Build an improved Mask RCNN framework. This framework consists of four modules: a feature extraction network, an RPN, RoIAlign, and a fully connected network. The feature extraction network uses a ResNet-based feature pyramid structure. An explicit visual center (EVC) module is added between the feature extraction module and the RPN module to obtain features that contain both global and local information. This module consists of two parallel blocks: a lightweight multi-layer perceptron (MLP) and a learnable visual center (LVC) mechanism.
[0011] Among them, the lightweight MLP is used to capture the global long-distance dependencies (i.e., global information) of the output features of the feature extraction network, while LVC is used to retain the local corner area information (i.e., local information) of the output features of the feature extraction network to aggregate the local area features within the layer. The resulting feature maps of these two blocks are connected together along the channel dimension and passed to the downstream RPN model as the output of EVC.
[0012] Step 4: Model training. The original color image and the annotation file are input into the improved Mask-RCNN network for training to obtain the human naked back segmentation model.
[0013] Step 5: Apply the model to obtain the segmented image. The human body color image to be segmented is input into the back segmentation model trained in step 4 to segment the exposed back of the human body, and then the background is filled with black to obtain a background-free color image of the human body. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] FIG1 shows a flow chart of a back segmentation method based on an improved Mask RCNN proposed in the present invention;
[0015] FIG2 is a schematic diagram of the local structure of a back segmentation model based on an improved Mask RCNN in an embodiment of the present invention;
[0016] FIG3 is a schematic diagram showing the principle of an explicit vision center (EVC) module in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The preferred embodiments of the present invention are further described below with reference to the accompanying drawings and examples.
[0018] The flowchart shown in FIG1 shows the specific process of the entire implementation of the present invention:
[0019] Step 1: Data collection. Collect images of teenagers with exposed backs flexed forward from different regions, ages, genders, and scoliosis levels to obtain a color image dataset.
[0020] This step includes:
[0021] Step 1.1: Use a non-invasive, non-contact camera, such as an infrared depth camera, to capture a color image of the back of the subject. The subject should bend their upper body forward, with no clothing obstructing the view. The image should be taken from directly above the subject, fully exposing the anatomical landmarks of the back. Save the captured color image to a local computer.
[0022] Step 2: Labeling: Manually label the exposed back area of the human body in the color image to obtain the labeling file corresponding to the original color image.
[0023] This step includes:
[0024] Step 2.1: Use Labelme annotation software to annotate the exposed back of the human body in the original color image in the format of point annotation, name the annotated area "back", and save the coordinate information and name of each point in the annotation file; or use Labelme annotation software to annotate the exposed back of the human body in the original color image in the format of box annotation, name the annotated area "back", and save the coordinate information and name of each point in the annotation file, then pass the annotation file into the SAM model, so that the model outputs the back area in each annotated box, and finally add the coordinate information of each output point to the annotation file. This can reduce the workload of manual annotation to a certain extent.
[0025] Step 3: Build an improved Mask RCNN framework.
[0026] This step includes:
[0027] Step 3.1, the Mask RCNN model structure is used as the framework, which consists of four modules: feature extraction network, RPN, RoIAlign and fully connected network. The feature extraction network uses the feature pyramid network based on ResNet. The features of each layer output by the network are X i(i=0,1,2,3,4), the corresponding spatial sizes are the input image
[0028] Step 3.2, as shown in Figure 2, adds an explicit visual center (EVC) module between the feature extraction module and the RPN module, so that the feature extraction network outputs features with both global and local information;
[0029] The input of the EVC module is the top-level feature X4 of the ResNet output after being processed by the Stem block, denoted as X in The Stem block consists of a 7×7 convolution with an output channel of 256, a batch normalization layer (BN), and an activation function layer, so that X in It can be calculated by the following formula: in =σ(BN(Conv 7×7 (X4))). (1)
[0030] X in The two parallel blocks of lightweight MLP and LVC are placed in the EVC module at the same time to obtain global information and local information respectively.
[0031] The lightweight MLP consists of two residual modules: a module based on depthwise separable convolution (DeepwiseConvolution) and a module based on channel MLP. The input of the MLP module is the output of the depthwise separable convolution module. Both modules undergo channel scaling, DropPath, and residual connection operations to improve feature generalization and robustness. The depthwise separable convolution can improve feature expression while reducing computational cost. The input of this module is the feature map X that has been group normalized. in , the above process can be expressed as:
[0032] Compared with spatial MLP, channel MLP can not only effectively reduce computational complexity, but also meet the requirements of general visual tasks. The input of this module is the output of the depthwise separable convolution module after group normalization (GN). The above process can be expressed as:
[0033] LVC is an encoder with an intrinsic dictionary, which consists of an inherent codebook B = {b1, b2, ..., b K ,} and a set of learnable visual center scaling factors S={s1,s2,…,s k ,}. This module uses a set of convolutional layers to process the X output by the Stem block. inEncode and process the encoded features using a CBR block (consisting of a 3×3 convolution, a BN layer, and a ReLU activation function). After the above steps, the encoded features is input into the codebook, and then a set of scaling factors S are used to calculate b in sequence k Relative to The position information of all points on the image is obtained. Therefore, the information of the kth codeword relative to the entire feature image can be calculated by the following formula:
[0034] in It is a feature map The i-th pixel on b k is the kth learnable visual codeword, s k is the kth scaling factor, yes The position information of each pixel relative to the kth codeword, K is the total number of visual centers, N is the feature map The total number of pixels on the graph. Then, φ is used to fuse all the pixels on the graph. k , where φ contains a batch normalization layer and a mean layer with a ReLU activation function. The complete information of the entire feature image about the K codewords is calculated as follows:
[0035] After obtaining the codebook output according to the above operations, e is input into a fully connected layer and a 1×1 convolutional layer to further predict the prominent key class features. in The local corner area features of the input features and the scaling factor coefficients are subjected to channel multiplication and channel addition. The above operation can be expressed as:
[0036] Among them, δ(·) is the sigmoid function, Is channel multiplication. Then, the feature X output by the Stem block in Channel addition is performed between the local corner area feature Z and the local corner area feature Z, and the formula is:
[0037] Finally, the output of the EVC module is the feature X4′ obtained by concatenating the result features of the lightweight MLP and LVC blocks along the channel dimension. The calculation formula can be expressed as: X4′=cat(MLP(X in ),LVC(X in )). (8)
[0038] In step 3.3, the features containing the spatial explicit visual center output by the EVC module are adjusted from top to bottom through the feature pyramid to adjust all the previous shallow features X3, X2, X1, and X0, and generate features X3′, X2′, X1′, X0′ and X4′ as the input of the RPN network.
[0039] Step 4: Model training. The original color image and the annotation file are input into the improved Mask-RCNN network for training to obtain the human naked back segmentation model.
[0040] This step is the data processing process of the improved Mask RCNN network, including:
[0041] Step 4.1: The feature extraction network of the EVC module is added to the image input to extract features and obtain the corresponding feature map;
[0042] Step 4.2: The feature map is input into RPN to generate K proposals for each image and map these proposals to the feature map to generate the corresponding RoI;
[0043] Step 4.3: The feature map with RoI is input into the RoIAlign layer to adjust each RoI to a fixed size.
[0044] In step 4.4, the fully connected network is used to classify the RoIs, regress the positions of the annotation boxes, and then segment the back of the human body in each annotation box.
[0045] Step 5: Apply the model to obtain a segmented image. Input the color image to be segmented into the back segmentation model trained in step 4 to segment the exposed back of the human body, and then fill the background with black to obtain a color image of the human body without background.
[0046] This step includes:
[0047] In step 5.1, all original color images are input into the trained human body segmentation model, which will output the human back mask in all images. The human back area is white, and the corresponding R, G, and B values are all 255. The background area is black, and the corresponding R, G, and B values are all 0.
[0048] In step 5.2, the value of each pixel in the mask is compared with the value of the pixel at the corresponding position in the original color image, and the smaller value is taken to obtain a color image of the human back with the background filled with black.
Claims
1. An improved Mask RCNN-based back segmentation method, characterized in that, it includes the following steps: Step 1: Collect data; collect the forward-bending bare-back human body images of teenagers with different regions, ages, genders, and scoliosis degrees to obtain a color image dataset; Step 2: Label; manually label the bare-back area of the human body in the color image to obtain the annotation file corresponding to the original color image; Step 3: Build an improved Mask RCNN model; build a Mask RCNN framework composed of four modules: a feature extraction network, RPN, RoIAlign, and a fully connected network, where the feature extraction network uses a ResNet-based feature pyramid structure; add an Explicit Visual Center (EVC) module between the feature extraction module and the RPN module to obtain features with both global and local information. This module consists of two parallel blocks, namely a lightweight multi-layer perceptron (MLP) and a learnable visual center (LVC) mechanism; Step 4: Train the model; input the original color image and the annotation file into the improved Mask RCNN network for training to obtain a human bare-back segmentation model; Step 5: Apply the model to obtain the segmented image; input the color image to be segmented into the back segmentation model trained in Step 4 to segment the bare-back area of the human body, and then fill the background with black to obtain a background-free human color image.
2. An improved Mask RCNN-based back segmentation method according to claim 1, characterized in that, in Step 2, the bare-back area of the human body in the color image is manually labeled to obtain an annotation file; Step 2 further includes: Step 2.1, use the Labelme annotation software to label the bare-back part of the human body in the original color image in the format of point annotation, name the annotation area as back, and save the coordinate information and name of each point in the annotation file; or use the Labelme annotation software to label the bare-back part of the human body in the original color image in the format of box annotation, name the annotation area as back, and save the coordinate information and name of each point in the annotation file, then input the annotation file into the SAM model to make the model output the back area in each annotation box, and finally add the coordinate information of each output point to the annotation file, which can reduce the workload of manual annotation to a certain extent.
3. An improved Mask RCNN-based back segmentation method according to claim 1, characterized in that, in Step 3, an improved Mask RCNN model is built; Step 3 further includes: Step 3.1, with the Mask RCNN model structure as the framework, which is composed of four modules: a feature extraction network, RPN, RoIAlign, and a fully connected network. Among them, the feature extraction network adopts a feature pyramid network based on ResNet, and the features of each layer output by this network are X i (i = 0, 1, 2, 3, 4), and their corresponding spatial dimensions are those of the input image Step 3.2, add an Explicit Visual Center (EVC) module between the feature extraction module and the RPN module to obtain features with both global and local information; The input of the EVC module is the topmost feature X output by ResNet that has been processed by the Stem block 4 , denoted as X in ; The Stem block includes a 7×7 convolution with 256 output channels, a batch normalization layer (BN), and an activation function layer. Thus, X in can be calculated by the following formula: X in = σ(BN(Conv 7×7 (X 4 ))); (1) X in The lightweight MLP and LVC, two parallel blocks, are simultaneously placed into the EVC module to obtain global information and local information respectively; Among them, the lightweight MLP consists of two residual modules, namely the module based on depthwise convolution and the module based on channel MLP. The input of the MLP module is the output of the depthwise convolution module. Both of these modules have undergone channel scaling, DropPath, and residual connection operations to improve feature generalization and robustness. Depthwise convolution can improve feature expression ability while reducing computational cost. The input of this module is the feature map X that has undergone group normalization processing. in , the above process can be expressed as: Compared with the spatial MLP, the channel MLP can not only effectively reduce the computational complexity but also meet the requirements of general vision tasks. The input of this module is the output of the depthwise separable convolution module after group normalization (GN). The above process can be expressed as: LVC is an encoder with an inherent dictionary, consisting of an inherent codebook B = {b 1 , b 2 , …, b K ,} and a set of learnable visual center scaling factors S = {s 1 , s 2 , …, s k ,}; this module uses a set of convolutional layers to encode the X in output by the Stem block, and uses a CBR block (consisting of a 3×3 convolution, a BN layer, and a ReLU activation function) to process the encoded features; after the above steps, the encoded features is input into the codebook, and then b is calculated successively using a set of scaling factors S k relative to The position information of all points above, whereby the information of the k-th codeword relative to the entire feature image can be calculated by the following formula: Among them Is the feature map The i-th pixel point on, b k is the k-th learnable visual codeword, s k is the k-th scaling factor, Yes The position information of each pixel on it relative to the k-th codeword, where K is the total number of visual centers and N is the feature map The total number of pixel points on; then, use φ to fuse all e k , where φ includes a batch normalization layer and a mean layer with a ReLU activation function; the complete information of the entire feature image regarding K codewords is calculated as follows: After obtaining the output of the codebook according to the above operations, e is input into a fully connected layer and a 1×1 convolutional layer to further predict prominent key class features; then, X from the Stem block in The input features and the local corner region features of the scaling factor coefficients are subjected to channel multiplication and channel addition, and the above operations can be expressed as: where δ(·) is the sigmoid function, is channel multiplication; then, a channel addition is performed between the feature X output by the Stem block in and the local corner region feature Z, and its formula is: Finally, the output of the EVC module is the feature X 4 ′ obtained by concatenating the resulting features of the lightweight MLP and LVC blocks along the channel dimension, and its calculation formula can be expressed as: X 4 ′ = cat(MLP(X in ), LVC(X in )); (8) Step 3.3, the features containing the spatial explicit visual center output by the EVC module adjust all the previous shallow features X 3 , X 2 , X 1 , X 0 through the top-down manner of the feature pyramid, and generate the features X 3 ′, X 2 ′, X 1 ′, X 0 ′ and X 4 ′ together serve as the input to the RPN network.
4. An improved Mask RCNN-based back segmentation method according to claim 1, characterized in that, In step 4, the improved Mask RCNN network is trained to obtain a human bare back segmentation model; Step 4 further includes: Step 4.1, the feature extraction network with the EVC module is added to the image input for feature extraction to obtain the corresponding feature map; Step 4.2, the feature map is input into the RPN to generate K proposals for each image and map these proposals onto the feature map to generate the corresponding RoIs; Step 4.3, the feature map with RoIs is input into the RoIAlign layer to adjust each RoI to a fixed size; Step 4.4, the fully connected network is used to classify the RoIs, and the position of the annotation box is regressed, and then the human back is segmented in each annotation box.
Citation Information
Patent Citations
Mask-RCNN-based Gaofen-3 SAR image road detection method
CN110852176A
Construction method of image segmentation model and image segmentation method and system
CN111161290A
Non-invasive scoliosis screening method and system based on back color image
CN114287915A
Automatic visual identification method and sorting system
CN116213306A
Medical image segmentation method based on u-net
US20220309674A1
Cited By
Multi-modal medical image segmentation method based on frequency domain perception fusion
CN121685970A
A multi-modal medical image segmentation method based on frequency domain perception fusion
CN121685970B