Endoscopic lumbosacral plexus nerve root segmentation method and system based on semi-supervised learning

By adopting a network model based on semi-supervised learning under the endoscopy, combining the CNN-Transformer encoder and dual attention module, dual flow perturbation of labelless data, the problem of lack of high-precision lumbosacral plexus nerve root segmentation method in the prior art is solved, and high-precision segmentation of nerve roots in spinal endoscopic surgery is achieved.

CN119992091APending Publication Date: 2025-05-13BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510094077.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art lacks endoscopic lumbosacral plexus nerve root segmentation method based on semi-supervised learning, making it difficult to achieve high-precision segmentation through a small amount of labeled data, especially in the case of narrow visual field and blurred nerve root edges during spinal endoscopic surgery.

Method used

Using a network model based on semi-supervised learning, two identical branch models are constructed, dual-stream perturbation (weak perturbation and strong perturbation) are performed on label-free data, combined with the CNN-Transformer encoder and the dual attention module, model training is performed to achieve high-precision segmentation of nerve roots.

Benefits of technology

High-precision segmentation of nerve roots through a small amount of labeled data is achieved, which improves the accuracy of nerve root detection in spinal endoscopic surgery and enhances the safety and accuracy of the surgery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992091A_ABST
    Figure CN119992091A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical image processing, and particularly relates to an endoscopic lumbosacral plexus nerve root segmentation method and system based on semi-supervised learning. An endoscopic lumbosacral plexus nerve root segmentation method based on semi-supervised learning specifically comprises the steps that a semi-supervised learning network model is constructed, and the semi-supervised learning network model comprises two same branch models; double-flow disturbance, namely weak disturbance and strong disturbance, is carried out on the label-free data; inputting the tagged data set xl into one branch of the network model, inputting the weak disturbance untagged data xw and the strong disturbance untagged data xs into the other branch of the network model, and carrying out model training; and segmenting the endoscope data of the lower lumbosacral plexus nerve root by using the trained network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and in particular relates to an endoscopic lumbar sacral plexus nerve root segmentation method and system based on semi-supervised learning. Background Art

[0002] Minimally invasive endoscopic spinal surgery is performed in a highly restricted surgical field with a working channel as narrow as 8 mm adjacent to critical structures such as nerve roots and blood vessels. In order to perform the surgery safely and effectively, there is an urgent need for advanced computer-assisted technologies to provide real-time intraoperative support, such as nerve root detection, which has the potential to improve surgical precision and patient outcomes.

[0003] However, due to the particularity of medical images, it is very difficult to obtain sufficient data, especially spinal endoscopic images that must be obtained through clinical surgery. In addition, annotating medical images requires the expertise of experienced clinicians and requires a lot of time and cost investment. Therefore, it is extremely challenging to develop a robust model from a small amount of labeled data. There is a lack of research on lumbar sacral nerve root segmentation algorithms for spinal endoscopic surgery, so only recently published methods that are closer to this solution are listed.

[0004] 202310129552.1 "Semi-supervised medical image segmentation method based on interaction between Transformer and CNN" introduces Transformer into the semi-supervised segmentation framework of medical images, designs the C2T module and T2C module for feature interaction between the Transformer branch and the CNN branch, and adds feature consistency distribution constraints. Through this cross-teaching method, the semi-supervised framework after the introduction of the Transformer branch is more stable and produces more accurate pseudo-labels.

[0005] 202310129552.1 By combining the Transformer branch and the CNN branch, a more stable semi-supervised learning framework was constructed, and more accurate pseudo labels were generated. The network performance was tested using the skin disease segmentation dataset and the heart segmentation dataset. However, the field of view of spinal endoscopic surgery is narrow and the edges of the nerve roots are blurred. Therefore, it is unknown whether the detection accuracy of the algorithm meets the requirements in the face of the above situation.

[0006] There is no endoscopic lumbar plexus nerve root segmentation method based on semi-supervised learning in the prior art. Summary of the invention

[0007] In view of this, the present invention provides an endoscopic lumbar sacral plexus nerve root segmentation method and system based on semi-supervised learning, which can achieve high-precision segmentation of nerve roots with a small amount of labeled data.

[0008] The technical solution for implementing the present invention is as follows:

[0009] An endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning, the specific process is as follows:

[0010] Construct a semi-supervised learning network model, including two identical branch models;

[0011] Two-stream perturbations, i.e. weak perturbation and strong perturbation, are performed on unlabeled lumbar plexus nerve root endoscopy data.

[0012] The labeled lumbar plexus nerve root endoscopy dataset x l Input to one branch of the network model, the weakly perturbed unlabeled data x w and strongly perturbed unlabeled data x s Input to another branch of the network model for model training;

[0013] The trained network model is used to segment the lumbar plexus nerve root endoscopic data.

[0014] Optionally, the present invention adds feature disturbance to the weakly disturbed unlabeled data and uses it as sample data for training the neural network.

[0015] Optionally, the strong perturbation described in the present invention is: using an adaptive patch splicing method to identify difficult areas and simple areas of unlabeled data, using the identified areas as patches for cutting and mixing operations, and performing strong perturbations on the unlabeled data.

[0016] Optionally, the specific implementation process of the strong disturbance described in the present invention is:

[0017] First, for an input unlabeled image of size h×w, it is divided into n×n patches and the learning difficulty score r of each patch is calculated. m ;

[0018] Secondly, the learning difficulty score r m Sort by, calculate the first β% of the results k hard As a difficult module M hard , then β% calculation result k easy As a simple M easy ;

[0019]

[0020] Finally, regions with different learning difficulties are randomly selected as patches for cutting and mixing operations, and cut to generate enhanced images to obtain strongly perturbed unlabeled data x s1 and strongly perturbed unlabeled data x s2 .

[0021] Optionally, the present invention further adjusts the corresponding pseudo-labels to be consistent with the modification of the cutting and mixing operations, and the obtained mixed patch cutting and mixing area is represented as follows:

[0022] x s1 and x s2 Together they formed M easy (i,j) and M hard (i,j) together form M,

[0023]

[0024] in, represents the i-th unlabeled image, whose pseudo label is M represents a binary mask, Mosaic represents the image after mosaic processing,

[0025] Optionally, the overall objective function of the network model described in the present invention is a combination of unsupervised loss and supervised loss.

[0026]

[0027] The loss function of the supervised algorithm is defined as:

[0028]

[0029] Among them, N is the number of samples with labeled data, C is the number of categories in the segmentation task, and y l (i,j) represents the true label of the i-th sample in the j-th category, p l (i,j) represents the predicted probability of the i-th sample in the j-th category;

[0030] The loss function of the semi-supervised algorithm is defined as:

[0031]

[0032] Where H() represents the entropy that minimizes the difference between two probability distributions, τ represents the confidence threshold, and B u represents the batch size of unlabeled data for computing the loss;

[0033] The dice loss in the semi-supervised algorithm is defined as:

[0034]

[0035] in, Indicates DICE loss.

[0036] Optionally, each branch of the semi-supervised learning network model of the present invention includes an encoder ε and a decoder

[0037] Optionally, the encoder of the present invention is a CNN-Transformer structure, and a dual attention module is introduced;

[0038] In the encoder, the image is input into the dual attention module to extract local features through convolution and downsampling, and then reshaped into a series of image blocks (patches) and input into the Transformer module to capture global semantic information through linear projection and position encoding.

[0039] Optionally, the dual attention module described in the present invention includes a PAM module and a CAM module. The PAM module captures global dependencies and local spatial relationships through 1×1 convolution and 3×3 convolution respectively, and the CAM module dynamically enhances channel-level feature expression through inter-channel relationship calculation and adaptive weighting mechanism.

[0040] In a second aspect, the present invention provides an endoscopic lumbar sacral plexus nerve root segmentation system based on semi-supervised learning, the system comprising a semi-supervised learning network model, the network model comprising two identical branch models; the system is trained using the above method.

[0041] Beneficial effects:

[0042] First, the present invention can achieve high-precision segmentation of nerve roots through a small amount of labeled data. The proposed method can not only make up for the current insufficient research in this field, but also assist surgical robots to complete safe operations.

[0043] Second, the introduction of the CNN-Transformer encoder of the present invention connects global context clues with image-specific structural information by capturing global information and extracting local spatial and channel features, thereby enhancing feature extraction capabilities and improving segmentation accuracy.

[0044] Third, the introduction of the dual attention module of the present invention generates a more context-aware and discriminative feature representation for the segmentation task, which can effectively improve the accuracy of the depth estimation algorithm and make it more suitable for clinical use.

[0045] Fourth, this method can provide accurate lumbar plexus nerve root segmentation results, which can be subsequently expanded to various complex tissue segmentations in spinal endoscopic surgery. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 This is a framework diagram of the existing semi-supervised medical image segmentation method based on the CNN-Transformer structure;

[0048] Figure 2 This is a schematic diagram of the structure of the semi-supervised learning network model of the present invention;

[0049] Figure 3 This is a schematic diagram of image segmentation after training with different amounts of labeled data using the present invention in this example. DETAILED DESCRIPTION

[0050] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0051] It should be noted that the following embodiments and features in the embodiments may be combined with each other in the absence of conflict; and, based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making any creative work are within the scope of protection of the present disclosure.

[0052] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.

[0053] The present application embodiment provides an endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning, such as Figure 2 As shown, the input dataset of this method includes a labeled dataset of lower lumbar plexus nerve root endoscopy and lower lumbar plexus nerve root endoscopy unlabeled dataset The model training is guided by a small amount of labeled data. The specific process of this method is as follows:

[0054] Construct a semi-supervised learning network model, including two identical branch models;

[0055] Two-stream perturbations, i.e. weak perturbation and strong perturbation, are performed on unlabeled lumbar plexus nerve root endoscopy data.

[0056] The labeled lumbar plexus nerve root endoscopy dataset x l Input to one branch of the network model, the weakly perturbed unlabeled data xw and strongly perturbed unlabeled data x s Input to another branch of the network model for model training;

[0057] The trained network model is used to segment the lumbar plexus nerve root endoscopic data.

[0058] When this embodiment is specifically implemented, weak disturbances may be, for example, cropping and rotation; strong disturbances may be, for example, color jittering; in addition, this embodiment uses labeled data to perform normal supervised learning, so that the present invention only requires a small amount of labeled data to achieve high-precision segmentation of nerve roots. This method can not only make up for the current lack of research in this field, but also assist surgical robots in completing safe operations.

[0059] In another embodiment of the present invention, feature disturbance is added to the weakly disturbed unlabeled data, for example, dropout or adding uniform noise, and it is input into the neural network during training to achieve feature enhancement.

[0060] In yet another embodiment of the present invention, the strong perturbation is to use an adaptive patch splicing method to identify difficult and simple areas of unlabeled data, and use the identified areas as patches for cutting and blending operations, thereby enhancing the learning process.

[0061] In the data enhancement stage, the Cutmix operation is usually used when strong perturbations are performed on unlabeled data. However, in the early training stage and when encountering difficult samples, the extensive application of this operation to unlabeled data may introduce model bias. To solve this problem, this embodiment adopts an adaptive patch stitching method to identify difficult areas and simple areas. These identified areas will be used as patches for cutting and mixing operations, thereby enhancing the learning process.

[0062] In another embodiment of the present application, the specific execution process of the strong disturbance is:

[0063] For an input unlabeled image of size h×w, it is segmented into n×n patches; each patch contains C classes in the segmentation task; p m (c) represents the predicted probability of category c in the mth patch; logC represents the normalized term of the maximum information entropy value; the difficulty score of this region is defined as r m :

[0064]

[0065] Difficulty score r m Sort by, calculate the first β% of the results k hard As a difficult module M hard , then β% calculation result k easy As a simple M easy :

[0066] sorted_indices[i] = argsort( m [i,:]

[0067] Among them, argsort is a sort of difficulty score r m The index obtained by sorting the values ​​in ascending order.

[0068]

[0069] In order to enhance the understanding ability of the network model, patches with different learning difficulties are randomly selected and cut to generate enhanced images, and the strongly perturbed unlabeled data x is obtained. s1 and strongly perturbed unlabeled data x s2 .

[0070] The percentage β% of the above-mentioned difficult modules and simple modules to the total number of modules can be adjusted according to different tasks, and can also be added to the algorithm model as a training parameter.

[0071] In addition, the corresponding pseudo-labels are adjusted to be consistent with these modifications. The resulting mixed patch cut mixed region is represented as follows:

[0072]

[0073] in, represents the i-th unlabeled image, whose pseudo label is M represents a binary mask, Mosaic represents the mosaicked image, and the symbol ⊙ represents element-by-element multiplication.

[0074] x s1 and x s2 Together they formed M easy (i,j) and M hard (i,j) together form M, that is,

[0075]

[0076]

[0077] In another embodiment of the present application, each branch of the semi-supervised learning network model includes an encoder ε and a decoder

[0078] For labeled data x l After the output of the network model p l :

[0079]

[0080] For the weakly perturbed unlabeled data x w The output pseudo label p of the network model w :

[0081]

[0082] For the weakly perturbed unlabeled data x w After adding feature perturbations, for example, dropout or adding uniform noise, the output p fp ;

[0083] For the strongly perturbed unlabeled data x s After the output of the network model p s :

[0084]

[0085] The overall objective function of the network model is a combination of unsupervised loss and supervised loss.

[0086]

[0087] The loss function of the supervised algorithm is defined as:

[0088]

[0089] Among them, N is the number of samples with labeled data, C is the number of categories in the segmentation task, and y l (i,j) represents the true label of the i-th sample in the j-th category, p l (i,j) represents the predicted probability of the i-th sample in the j-th category.

[0090] The loss function of the semi-supervised algorithm is defined as:

[0091]

[0092] Among them, τ represents the confidence threshold, which is set to 0.95, H() represents the entropy that minimizes the difference between the two probability distributions, and B u Indicates the batch size of unlabeled data for computing the loss.

[0093] The dice loss in the semi-supervised algorithm is defined as:

[0094]

[0095] in, represents the DICE loss, τ represents the confidence threshold, which is set to 0.95, and B u Indicates the batch size of unlabeled data for computing the loss.

[0096] In another embodiment of the present application, the encoder combines the self-attention mechanism of the Transformer with the local feature extraction capability of the CNN. In the encoder, the input image first extracts local features through convolution and downsampling, and then is reshaped into a series of image patches and input into the Transformer module, and global semantic information is captured through linear projection and position encoding to enhance the representation capability of the model.

[0097] In the feature extraction process, dual attention modules (PAM and CAM) are used to improve feature quality. The PAM module captures global dependencies and local spatial relationships through 1×1 convolution and 3×3 convolution respectively, while the CAM module dynamically enhances channel-level feature expression through inter-channel relationship calculation and adaptive weighting mechanism.

[0098] F 1 =Conv(PAM(F in ))

[0099] F 2 =Conv(CAM(F in ))

[0100] Then introduce the adaptive fusion strategy:

[0101] F out =Conv[αF 1 +(1-α)F 2 ]

[0102] Among them, F in is the input feature, Conv is the convolution operation, and α is a learnable parameter used to balance the contribution of position and channel enhancement features.

[0103] In this embodiment, the dual attention module can be replaced by other attention mechanisms such as SENet, CBAM module, etc. In the field of deep learning image processing, any model that can achieve depth estimation is applicable.

[0104] In the decoder stage, the framework uses segmentation heads, convolution modules and multi-attention jump connections to further improve the robustness of features and model performance. The overall design effectively combines global and local features, enhancing the stability and accuracy of medical image segmentation tasks.

[0105] like Figure 3 As shown, this embodiment constructs 700 images as a spinal endoscopic surgery dataset, the first column is the original image, the second column is the test result after the model is trained when the number of labeled images is 30, the third column is the test result after the model is trained when the number of labeled images is 50, and the fourth column is the test result after the model is trained when the number of labeled images is 70.

[0106] The method of the present invention has high lumbar sacral plexus nerve root segmentation accuracy. The method can subsequently perform a single task (nerve root segmentation) or face multiple tasks (nerve root segmentation and ligamentum flavum segmentation, etc.) by expanding the data set.

[0107] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. An endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning, characterized in that: The specific process is: Construct a semi-supervised learning network model, including two identical branch models; Two-stream perturbations, i.e. weak perturbation and strong perturbation, are performed on unlabeled lumbar plexus nerve root endoscopy data. The labeled lumbar plexus nerve root endoscopy dataset x l Input to one branch of the network model, the weakly perturbed unlabeled data x w and strongly perturbed unlabeled data x s Input to another branch of the network model for model training; The trained network model is used to segment the lumbar plexus nerve root endoscopic data.

2. The endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning according to claim 1, characterized in that: Feature perturbations are added to the weakly perturbed unlabeled data, and the data is used as sample data for training the neural network.

3. The endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning according to claim 2, characterized in that: The strong perturbation is: using an adaptive patch splicing method to identify difficult areas and simple areas of unlabeled data, using the identified areas as patches for cutting and mixing operations, and performing strong perturbations on the unlabeled data.

4. The endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning according to claim 3, characterized in that: The specific implementation process of the strong disturbance is: First, for an input unlabeled image of size h×w, it is divided into n×n patches and the learning difficulty score r of each patch is calculated. m ; Secondly, the learning difficulty score r m Sort by, calculate the first β% of the results k hard As a difficult module M hard , then β% calculation result k easy As a simple M easy ; Finally, regions with different learning difficulties are randomly selected as patches for cutting and mixing operations, and cut to generate enhanced images to obtain strongly perturbed unlabeled data x s1 and strongly perturbed unlabeled data x s2 .

5. The endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning according to claim 4, characterized in that: The corresponding pseudo-labels are further adjusted to be consistent with the modifications of the cutting and blending operations, and the resulting mixed patch cut-mixed region is represented as follows: x s1 and x s2 Together they formed M easy (i, j) and M hard (i, j) together form M, in, represents the i-th unlabeled image, whose pseudo label is M represents a binary mask, and Mosaic represents the image after mosaic processing.

6. The endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning according to claim 4, characterized in that: The overall objective function of the network model is a combination of unsupervised loss and supervised loss. The loss function of the supervised algorithm is defined as: Where N is the number of samples with labeled data, C is the number of categories in the segmentation task, yl(i, j) represents the true label of the i-th sample in the j-th category, and pl(i, j) represents the predicted probability of the i-th sample in the j-th category; The loss function of the semi-supervised algorithm is defined as: Where H() represents the entropy that minimizes the difference between two probability distributions, τ represents the confidence threshold, and B u represents the batch size of unlabeled data for calculating the loss; p w Represents the unlabeled data x after weak perturbation w The pseudo label output by the network model, p fP Represents the unlabeled data x after weak perturbation w Add feature perturbations to the pseudo labels output by the network model. Represents the strongly perturbed unlabeled data x s1 The pseudo label output by the network model, Represents the strongly perturbed unlabeled data x s2 Pseudo labels output by the network model; The dice loss in the semi-supervised algorithm is defined as: in, Indicates DICE loss.

7. The endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning according to claim 1, characterized in that: Each branch of the semi-supervised learning network model includes an encoder ε and a decoder 8. The endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning according to claim 7, characterized in that: The encoder is a CNN-Transformer structure, and a dual attention module is introduced; In the encoder, the image is input into the dual attention module to extract local features through convolution and downsampling, and then reshaped into a series of image patches and input into the Transformer module to capture global semantic information through linear projection and position encoding.

9. The endoscopic lumbar sacral plexus nerve root segmentation method based on semi-supervised learning according to claim 8, characterized in that: The dual attention module includes a PAM module and a CAM module. The PAM module captures global dependencies and local spatial relationships through 1×1 convolution and 3×3 convolution respectively. The CAM module dynamically enhances channel-level feature expression through inter-channel relationship calculation and adaptive weighting mechanism.

10. An endoscopic lumbar plexus nerve root segmentation system based on semi-supervised learning, characterized in that: It comprises a semi-supervised learning network model, wherein the network model comprises two identical branch models; and is obtained by training using any one of the methods in claims 1-9 above.