Ear cartilage and substructure image segmentation method and system based on matching and segmentation cascade deep learning network
By using a deep learning network based on matching and segmentation cascades, combined with magnetic resonance imaging and deep learning technology, automatic and efficient segmentation of auricular cartilage and its substructures was achieved. This solved the problem of surgical damage and aesthetic effects relying on the surgeon's experience in autologous rib cartilage sculpting scaffold reconstruction, and provided a high-quality 3D bioprinting model.
Patent Information
- Application Number
- CN202110539024.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-18
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2041-05-18
AI Technical Summary
Existing techniques for autologous rib cartilage sculpting framework reconstruction of the auricle have the problems of significant surgical trauma and aesthetic results that depend on the surgeon's experience. Furthermore, manually labeling and segmenting auricular cartilage images is time-consuming and labor-intensive, making it difficult to obtain high-quality segmentation results efficiently.
A deep learning network based on matching and segmentation cascades was used to acquire the outer ear contour image through magnetic resonance imaging. The auricular cartilage and its substructures were automatically segmented using a fully convolutional neural network and an encoder-decoder U-shaped framework deep learning network. Combined with a small amount of manual label training, non-rigid registration and fine segmentation were achieved.
It enables automatic and efficient segmentation of auricular cartilage and its substructures under ultra-short echo time-series imaging, reducing manual segmentation time and improving segmentation accuracy and efficiency, thus providing high-quality models for 3D bioprinting.
Smart Images

Figure CN115375696B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of processing of medical images, and in particular to an ear cartilage and substructure image segmentation method and system based on a matching and segmentation cascade deep learning network. BACKGROUND
[0002] Congenital microtia is a common congenital maxillofacial deformity, and according to epidemiological statistics, the incidence rate in China is 5.18 / 10000. This disease is generally manifested as severe auricle hypoplasia, which has an important influence on the appearance, function and mental health of patients.
[0003] At present, using autologous rib cartilage to carve a support to reconstruct the external ear is one of the main plastic surgical treatment methods for congenital microtia. However, this plastic surgery causes additional harm to the child when obtaining the rib cartilage, and the aesthetic effect of the surgery depends largely on the experience of the surgeon. With the development of tissue engineering and 3D bioprinting technology, this method of precisely dispensing cells, matrix and biomaterials for bioprinting can not only improve the construction method of the ear cartilage support, but also precisely control the structure of the ear cartilage support. 3D bioprinting follows a three-dimensional model designed based on biomedical images, and completes the support construction task by layering histological assembly. The method of supervised learning can automatically obtain an ear cartilage model available for 3D bioprinting, but this method relies on a large number of manual labels to improve segmentation accuracy. Manual labeling relies on expert manual annotation, which requires a lot of time and effort, and how to produce high-quality segmentation results with only a small amount of labels is a difficult problem.
[0004] The above statements of background art are only for the convenience of deep understanding of the technical solutions of the present application (the technical means used, the technical problems solved and the technical effects produced, etc.), and should not be regarded as acknowledging or in any form implying that the message constitutes the prior art known to those skilled in the art. SUMMARY
[0005] In view of the actual problems existing in the prior art, the present application provides a solution for constructing an ear cartilage morphological structure model suitable for tissue engineering and 3D bioprinting based on magnetic resonance imaging suitable for ear cartilage.
[0006] According to one embodiment of the present application, a deep learning network based ear cartilage and substructure image segmentation method based on matching and segmentation cascade is provided, which comprises the following steps: acquiring a UTE sequence magnetic resonance image of an external ear contour; pre-processing the acquired UTE sequence magnetic resonance image of the external ear contour; taking the pre-processed UTE sequence image as input, training a registration network in a deep learning based ear cartilage and substructure image segmentation system based on matching and segmentation cascade, thereby obtaining a deformation field, and applying the deformation field to a manual label image of a reference image or a manual label image of a template image to obtain a coarse label image, and using the obtained ear cartilage coarse label image and ear cartilage substructure coarse label image to train a segmentation network in the deep learning based ear cartilage and substructure image segmentation system based on matching and segmentation cascade, thereby performing image segmentation on the ear cartilage and ear cartilage substructure, respectively; wherein the deep learning based ear cartilage and substructure image segmentation system based on matching and segmentation cascade comprises a first registration network, a second registration network and a segmentation network; randomly selecting one image from the pre-processed UTE ear cartilage image as a moving image, and randomly selecting another image as a fixed image, to form a set of input images as input of the first registration network, obtaining a first deformation field through the first registration network, and performing spatial transformation on the moving image using the first deformation field to obtain a deformed moving image; taking the deformed moving image and the fixed image as input image set of the second registration network, obtaining a second deformation field through the second registration network, and performing spatial transformation on the deformed moving image using the second deformation field to finally obtain a moving image matched with the fixed image, thereby completing the training of the first registration network and the second registration network; after the training of the first registration network and the second registration network is completed, selecting an image with clear and complete ear contour shape and structure from the moving image as a reference image, or using part or all of the data to form a template image through inter-group matching, and performing manual segmentation on the reference image or the template image to obtain a manual label image of the reference image or a manual label image of the template image; taking the reference image or the template image as a moving image, and taking the pre-processed UTE sequence image to be segmented into ear cartilage and substructure as a fixed image, obtaining a first deformation field and a second deformation field using the trained first registration network and second registration network, and performing spatial transformation on the manual label image of the reference image or the manual label image of the template image using the first deformation field and the second deformation field to obtain a coarse label image corresponding to the fixed image, thereby training the segmentation network using the coarse label image.
[0007] Preferably, the first registration network takes a full convolutional neural network as a basic structure, which includes a down-sampling module and an up-sampling interpolation module, and the loss function of the first registration network is a global cross-correlation coefficient loss function; the second registration network and the segmentation network take an encoding-decoding U-shaped framework as a backbone network, and each includes an encoder and a decoder; the encoder and the decoder of the second registration network have a skip connection therebetween, and the loss function of the second registration network is a local cross-correlation coefficient loss function; the encoder and the decoder of the segmentation network have a skip connection therebetween, and an attention module is added at each skip connection layer, and the loss function of the segmentation network is a Dice similarity coefficient loss function.
[0008] Preferably, when the UTE sequence magnetic resonance image of the external ear contour is acquired, the echo time is set to 0-1 ms according to the nuclear magnetic scanning device used; the UTE sequence magnetic resonance image includes axial images, sagittal images and coronal images; the preprocessing of the acquired UTE sequence magnetic resonance image of the external ear contour includes image format conversion, renaming, turning and rigid registration.
[0009] Preferably, the manually segmented ear cartilage label image is obtained according to the manual operation of the operator; when the ear cartilage label image is segmented according to the manual operation of the operator, the ear cartilage, the surrounding fat and the connective tissue are all presented as bright signals in the preprocessed UTE sequence magnetic resonance image to different degrees, and the skin and other tissues are relatively dark, so that the ear cartilage boundary is segmented.
[0010] Preferably, the manually segmented ear cartilage substructure label image is obtained according to the manual operation of the operator; when the ear cartilage substructure label image is segmented according to the manual operation of the operator, a total of twelve ear cartilage substructure images are segmented in the order of helix, triangular fossa and antihelix segmentation, helix and antihelix segmentation, cymba concha segmentation, tragus segmentation, antitragus, antitragus and intertragic notch segmentation, and external auditory canal and concha cavity segmentation.
[0011] According to one embodiment of the present application, an ear cartilage and substructure image segmentation system based on a matching and segmentation cascade deep learning network is provided, which comprises the following modules: an acquisition module that acquires a magnetic resonance image of a UTE sequence of an external ear contour; a preprocessing module that pre-processes the acquired magnetic resonance image of the UTE sequence of the external ear contour; a training module that takes the pre-processed UTE sequence image as input, trains a registration network in the ear cartilage and substructure image segmentation system based on a matching and segmentation cascade deep learning, thereby obtaining a deformation field, applies the deformation field to a manual label image of a reference image or a manual label image of a template image to obtain a coarse label image, and trains a segmentation network in the ear cartilage and substructure image segmentation system based on a matching and segmentation cascade deep learning using the obtained ear cartilage coarse label image and ear cartilage substructure coarse label image, respectively, thereby performing image segmentation on the ear cartilage and ear cartilage substructure, respectively; wherein the ear cartilage and substructure image segmentation system based on a matching and segmentation cascade deep learning comprises a first registration network, a second registration network, and a segmentation network; a random image is selected from the pre-processed UTE ear cartilage image as a moving image, and another random image is selected as a fixed image to form a set of input images as input to the first registration network, and a first deformation field is obtained through the first registration network, and the moving image is spatially transformed using the first deformation field to obtain a deformed moving image; the deformed moving image and the fixed image are taken as an input image set of the second registration network, and a second deformation field is obtained through the second registration network, and the deformed moving image is spatially transformed again using the second deformation field to finally obtain a moving image that matches the fixed image, thereby completing the training of the first registration network and the second registration network; after the training of the first registration network and the second registration network is completed, an image with clear and complete ear contour morphology and structure is selected from the moving image as a reference image, or a template image is constructed using part or all of the data through an inter-group matching method, and a manual segmentation is performed on the reference image or the template image to obtain a manual label image of the reference image or a manual label image of the template image; the reference image or the template image is taken as a moving image, and the pre-processed UTE sequence image that needs to be segmented into ear cartilage and substructure is taken as a fixed image, the first deformation field and the second deformation field are obtained using the trained first registration network and the second registration network, the manual label image of the reference image or the manual label image of the template image is spatially transformed using the first deformation field and the second deformation field to obtain a coarse label image corresponding to the fixed image, and the coarse label image is used to train the segmentation network.
[0012] Preferably, the first registration network takes a full convolutional neural network as a basic structure, which includes a down-sampling module and an up-sampling interpolation module, and the loss function of the first registration network is a global cross-correlation coefficient loss function; the second registration network and the segmentation network take an encoding-decoding U-shaped framework as a backbone network, and each includes an encoder and a decoder; the encoder and the decoder of the second registration network have a skip connection, and the loss function of the second registration network is a local cross-correlation coefficient loss function; the encoder and the decoder of the segmentation network have a skip connection, and an attention module is added at each skip connection layer, and the loss function of the segmentation network is a Dice similarity coefficient loss function.
[0013] Preferably, when the UTE sequence magnetic resonance image of the external ear contour is acquired, the echo time is set to 0-1 ms according to the nuclear magnetic scanning device used; the UTE sequence magnetic resonance image includes axial images, sagittal images and coronal images; the preprocessing of the acquired UTE sequence magnetic resonance image of the external ear contour includes image format conversion, renaming, turning and rigid registration.
[0014] Preferably, the manually segmented ear cartilage label image is obtained according to the manual operation of the operator; when the ear cartilage label image is segmented according to the manual operation of the operator, the ear cartilage, the surrounding fat and the connective tissue are all presented as bright signals in different degrees in the preprocessed UTE sequence magnetic resonance image, and the skin and other tissues are relatively dark, so that the ear cartilage boundary is segmented.
[0015] Preferably, the manually segmented ear cartilage substructure label image is obtained according to the manual operation of the operator; when the ear cartilage substructure label image is segmented according to the manual operation of the operator, a total of twelve ear cartilage substructure images are segmented in the order of helix, triangular fossa and antihelix segmentation, helix and antihelix segmentation, cymba concha segmentation, tragus segmentation, antitragus, antitragus and intertragic notch segmentation, and external auditory canal and concha cavity segmentation.
[0016] The present application adopts the above technical scheme, and has the following beneficial effects:
[0017] On the basis of ear cartilage ultra-short echo time sequence imaging and very small amount of manual segmentation results, the present application obtains a model capable of automatically segmenting high-quality ear cartilage and substructure morphology through a cascaded deep learning network, and the method and system provide an artificial intelligent solution for automatically and efficiently obtaining ear cartilage models for 3D biological printing. BRIEF DESCRIPTION OF DRAWINGS
[0018] Exemplary embodiments of the present application will be described in greater detail below with reference to the accompanying drawings. For the purpose of clarity, the same components in the drawings are denoted by identical reference numerals. It is noted that the drawings merely show schematic representations of embodiments of the present application and are not necessarily drawn to scale. In the drawings:
[0019] Figure 1 is a processing flow chart of an auricular cartilage and substructure image segmentation method and system based on a matching and segmentation cascade deep learning network according to an embodiment of the present application;
[0020] Figure 2 is a process schematic diagram of pre-processing of original data by an auricular cartilage and substructure image segmentation method and system based on a matching and segmentation cascade deep learning network according to an embodiment of the present application;
[0021] Figures 3a to 3d is a schematic diagram of a pre-processing process of a magnetic resonance image of a UTE sequence obtained;
[0022] Figures 4a to 4d shows an auricular cartilage image segmented from a magnetic resonance image of a UTE sequence according to an embodiment of the present application;
[0023] Figures 5a to 5d shows an auricular cartilage image segmented from a magnetic resonance image of a UTE sequence according to an embodiment of the present application;
[0024] Figure 6 shows a flow chart of manual segmentation of auricular cartilage substructures according to an embodiment of the present application;
[0025] Figure 7a and Figure 7b shows a schematic diagram of an auricular cartilage image and substructure images segmented from a magnetic resonance image of a UTE sequence;
[0026] Figure 8 shows an architecture schematic diagram of an auricular cartilage and substructure image segmentation system based on a matching and segmentation cascade deep learning according to an embodiment of the present application;
[0027] Figures 9a to 9c and Figures 10a to 10c shows, as an example of two data, a superimposed image of a pre-processed moving image and a fixed image, a superimposed image of a moving image and a fixed image after registration by a first registration network, and a superimposed image of a moving image and a fixed image after registration by a second registration network;
[0028] Figure 11a and Figure 11b and Figure 12a and Figure 12bThree-dimensional display figures of ear cartilage deep learning network segmentation and manual segmentation of two examples of data are shown.
[0029] Figure 13a and Figure 13b and Figure 14a and Figure 14b Three-dimensional display figures of ear cartilage substructure deep learning network segmentation and manual segmentation of two examples of data are shown. DETAILED DESCRIPTION
[0030] The embodiments of the present application are described in detail below, which are implemented on the premise of the technical solutions of the present application, and detailed implementation modes and specific operation processes are given, but the protection scope of the present application is not limited to the following embodiments.
[0031] Figure 1 is a processing flow chart of an ear cartilage and its substructure image segmentation method and system based on a matching and segmentation cascade deep learning network according to the embodiments of the present application. The working principle of the ear cartilage and its substructure image segmentation method and system based on the matching and segmentation cascade deep learning network according to the present application is as follows, which is mainly divided into three steps (or modules).
[0032] The acquisition step (module) acquires the magnetic resonance image of the UTE sequence of the external ear contour;
[0033] The preprocessing step (module) pre-processes the acquired magnetic resonance image of the UTE sequence of the external ear contour;
[0034] The training step (module) takes the pre-processed UTE sequence image as input, trains the registration network in the ear cartilage and its substructure image segmentation system based on the matching and segmentation cascade deep learning, thereby obtaining the deformation field, applies the deformation field to the manual label image or the template image of the manual label image of the reference image to obtain the coarse label image, and respectively trains the segmentation network in the ear cartilage and its substructure image segmentation system based on the matching and segmentation cascade deep learning using the obtained ear cartilage coarse label image and ear cartilage substructure coarse label image, thereby respectively performing image segmentation on the ear cartilage and the ear cartilage substructure.
[0035] The manually segmented ear cartilage and its substructure label image is obtained manually according to the pre-prepared operation specification by a professional with relevant background knowledge.
[0036] The processing of each step (module) of the ear cartilage and its substructure image segmentation method and system based on the matching and segmentation cascade deep learning network of the present application is described in detail below.
[0037] Magnetic resonance imaging (MRI) is an imaging modality with no ionizing radiation, excellent soft tissue contrast, and rich imaging sequences. The signal intensity of the transverse relaxation time of the ear cartilage tissue is between muscle or fat tissue and bone tissue, and it changes exponentially in ultra-short echo time (UTE) time (<1 ms) and short echo time (1-10 ms), especially in UTE range, while the signal intensity changes slowly in long echo time (>10 ms). In other words, by setting the echo time (TE) below the short echo time, the shorter the TE, the better the enhancement effect of ear cartilage imaging.
[0038] In an embodiment of the present application, when acquiring the UTE sequence, the TE can be set to 0-1 ms, preferably 0.1-0.5 ms, according to the nuclear magnetic scanning device used. Table 1 below is an exemplary parameter selection for the acquired UTE sequence magnetic resonance image of the external ear contour, and the present application is not limited thereto.
[0039] [Table 1]
[0040] Parameters UTE sequence Repetition time, TR (ms) 5.58 Echo time, TE (ms) 0.14 Flip angle (degrees) 15 Slice thickness, Th (mm) 1.2 Inter-slice spacing (mm) 0.6 In-plane resolution (mm) 0.25×0.25 Matrix size (voxels) 800×800×80
[0041] Figure 2 The ear cartilage and its substructure image segmentation method and system based on the matching and segmentation cascade deep learning network according to the present application are shown to preprocess the acquired UTE sequence magnetic resonance image.
[0042] As Figure 2 shown, the ear cartilage and its substructure image segmentation method and system based on the matching and segmentation cascade deep learning network according to the present application can be applied to the UTE sequence magnetic resonance image of the external ear contour, and its data types can include but are not limited to DICOM (Digital Imaging and Communications in Medicine) and NIFTI (Neuroimaging Informatics Technology Initiative).
[0043] The preprocessing of the acquired UTE sequence magnetic resonance image of the external ear contour includes format conversion, renaming, steering, and rigid registration of the image.
[0044] According to the embodiment of the present application, if the data that can be used is DICOM format data, it is converted into NIFTI format data; if the data that can be used is NIFTI format data, no format conversion processing is performed. The warping is to transfer the image to the standard space and unify the voxel size. The rigid registration is to select an image with clear auricle structure and moderate shape and size as a fixed image, and the rest of the pre-processing images are all taken as moving images, which are uniformly moved to the same position as the fixed image through rotation and translation.
[0045] Figures 3a to 3d is a schematic diagram of a pre-processing process of a magnetic resonance image of a UTE sequence. Figures 3a to 3d The original UTE sequence magnetic resonance images (from left to right, the axial image, the sagittal image, and the coronal image), the images obtained after format conversion (from left to right, the axial image, the sagittal image, and the coronal image), the images obtained after warping (from left to right, the axial image, the sagittal image, and the coronal image), and the images obtained after rigid registration (i.e., the images obtained after pre-processing, from left to right, the axial image, the sagittal image, and the coronal image) are shown respectively.
[0046] In the embodiment of the present application, it is necessary to manually segment the auricle cartilage and its substructures of the reference image or template image in advance, and then generate the coarse label of the supervised segmentation network by performing two-stage registration on the label image. The manual label segmentation is based on the scanning scheme of the UTE sequence, and fine substructure manual segmentation can be performed on the basis of auricle cartilage segmentation.
[0047] In the magnetic resonance image of the UTE sequence, the cartilage, the surrounding fat, and the connective tissue all present bright signals to different degrees, and the skin and other tissues have relatively dark signals, so that the auricle cartilage boundary can be completely outlined.
[0048] Therefore, based on the imaging effect diagram of the auricle cartilage and the earlobe of the magnetic resonance image of the UTE sequence, the auricle cartilage image can be manually segmented.
[0049] Figures 4a to 4d and Figures 5a to 5d Two examples of data are taken as examples to show the auricle cartilage images segmented based on the magnetic resonance image of the UTE sequence according to the embodiment of the present application. Among them, Figure 4a and Figure 5a The axial image of the manual segmentation of the auricle cartilage is shown, Figure 4b and Figure 5b The sagittal image of the manual segmentation of the auricle cartilage is shown; Figure 4c and Figure 5c The coronal image of the manual segmentation of the auricle cartilage is shown; Figure 4d and Figure 5dA three-dimensional schematic diagram showing manual segmentation of ear cartilage.
[0050] Therefore, according to the embodiment of the present application, the ear cartilage can be imaged in high definition and non-invasively by UTE sequence scanning.
[0051] In the following, the segmented ear cartilage image will be further segmented into twelve substructures, i.e., the antihelix, the antihelix crus, the triangular fossa, the helix, the helix crus, the cymba concha, the scapha, the tragus, the antitragus, the intertragic notch, the external auditory canal, and the concha cavity.
[0052] According to the embodiment of the present application, as shown in Figure 6 the segmentation of the ear cartilage substructures is mainly performed in six steps, i.e., the antihelix, the triangular fossa, and the antihelix crus segmentation, the helix and the helix crus segmentation, the cymba concha segmentation, the scapha segmentation, the tragus, the antitragus, and the intertragic notch segmentation, and the external auditory canal and the concha cavity segmentation. Among them, the substructures of the ear cartilage are divided into an upper half and a lower half with the helix crus as the dividing line, the boundaries of the two important mechanical support structures of the helix and the antihelix are defined in the segmentation of the upper half, so as to ensure the correct segmentation of the remaining substructures, and then the antihelix, the antihelix crus, the triangular fossa, the helix, the helix crus, the cymba concha, the scapha, the tragus, the antitragus, the intertragic notch, the external auditory canal, and the concha cavity are segmented in sequence according to the anatomical definition of each structure and the corresponding features in the label and combined with the MRI image. Among them, the first to fourth steps are the segmentation steps of the upper half of the ear cartilage substructures, and the fifth and sixth steps are the segmentation steps of the lower half of the ear cartilage substructures.
[0053] Figure 7a and Figure 7b The ear cartilage image segmented based on the UTE sequence magnetic resonance image and the schematic diagram of the ear cartilage substructures are shown. According to the embodiment of the present application, the ear cartilage can be imaged in high definition based on the magnetic resonance technology by the UTE sequence magnetic resonance image, so that the ear cartilage image can be manually segmented, and the fine substructure can be manually segmented based on the ear cartilage segmentation.
[0054] The accuracy of the fully supervised network depends on the number and accuracy of the manual labels, but the manual labels require relevant experienced experts and are very time-consuming. The embodiment of the present application realizes the automatic segmentation of the ear cartilage and its substructures based on the matching and segmentation cascaded deep learning network, and only a small amount of individual reference or template label is needed to realize the ear cartilage segmentation of all data. In the framework of the registration network, the multi-level neural network is also used in the form of cascade to realize the non-rigid registration, which ensures the accuracy while the required time is much shorter than that required by the traditional registration method, and finally the attention U-Net is cascaded to optimize the segmentation result.
[0055] Figure 8An architecture schematic diagram of an ear cartilage and its substructure image segmentation system based on matching and segmentation cascade deep learning according to an embodiment of the present application is shown. As shown in Figure 8 The ear cartilage and its substructure image segmentation system based on matching and segmentation cascade deep learning can include a registration network and a segmentation network. Due to the large difference in the anatomical structure of ear cartilage between individuals, it is difficult to complete large-scale deformation using a single-stage matching network, so the registration network is divided into two stages of large deformation preliminary registration and small deformation registration. The registration network according to the present application can include an unsupervised deformable first registration network based on a convolutional neural network and an unsupervised deformable second registration network based on a convolutional neural network.
[0056] A 3D MRI image is randomly selected from the preprocessed UTE ear cartilage image as a moving image, and a 3D MRI image is randomly selected as a fixed image. The selected moving image and fixed image form a group of input images for input to the down-sampling module or encoder of the first registration network. Multiple random extractions can be performed according to the above process to obtain multiple groups of training images, which are used to train the first registration network. Here, the number of training image combinations is at least 100 or more, and in the present embodiment, 360 groups are selected.
[0057] The first deformation field is obtained through the first registration network, the spatial transformation of the input moving image is performed using the first deformation field, and the deformed moving image (i.e., the new moving image) is obtained. The deformed moving image and the input fixed image are used as the input image group of the second registration network to obtain the second deformation field. The deformed moving image (i.e., the new moving image) is spatially transformed again using the second deformation field, and finally the moving image matched with the fixed image is obtained, and the training of the first registration network and the second registration network is completed.
[0058] After the training of the first registration network and the second registration network is completed, an image with clear and complete auricle morphology and structure is selected from the finally obtained moving image matched with the fixed image as a reference image, or a template image is formed by using part or all of the data (UTE image) through the method of groupwise registration, and manual segmentation is performed on the reference image or the template image to obtain the manual label image of the reference image or the manual label image of the template image. The manual label image of the reference image or the manual label image of the template image includes the manually segmented ear cartilage label image and the manually segmented ear cartilage substructure label image.
[0059] The reference image or template image is taken as a moving image, and the preprocessed UTE sequence image requiring ear cartilage and substructure segmentation is taken as a fixed image. The first registration network and the second registration network are trained to obtain the deformation field of two stages. The manual label image of the reference image or the manual label image of the template image is continuously spatially transformed by the first deformation field and the second deformation field to obtain a coarse label image corresponding to the fixed image. The coarse label image includes an ear cartilage coarse label image and an ear cartilage substructure coarse label image, and the coarse label image is used to train a subsequent segmentation network.
[0060] The manual label image of the reference image or the manual label image of the template image is obtained by manual segmentation by a professional with relevant background knowledge according to a pre-prepared operation specification. When the ear cartilage label image is segmented according to the manual operation of the operator, the ear cartilage boundary is segmented based on the fact that the ear cartilage, the surrounding fat, and the connective tissue all appear as bright signals to different extents in the pre-processed UTE sequence magnetic resonance image, and other tissues such as the skin are relatively dark. When the ear cartilage substructure label image is segmented according to the manual operation of the operator, a total of twelve ear cartilage substructure images are segmented in the order of helix, triangular fossa and antihelix foot segmentation, helix and antihelix foot segmentation, cymba concha segmentation, tragus segmentation, antitragus, antitragus and intertragic notch segmentation, and external auditory canal and concha segmentation.
[0061] The segmentation network takes the pre-processed UTE ear cartilage image as input, and is supervised trained by the ear cartilage coarse label image and the ear cartilage substructure coarse label image, so that the trained segmentation network can realize automatic segmentation of the ear cartilage and its substructure.
[0062] The first registration network based on unsupervised deformation of the convolutional neural network is a coarse registration / global large deformation registration network. The moving image and the fixed image are input into the first registration network for training to obtain the first deformation field (DVF1), and then spatial transformation (Spatial Transform) is performed to obtain the registered image corresponding to the individual data. The first registration network obtains the first deformation field by maximizing the similarity between the moving image and the fixed image. The first deformation field contains the displacement of each pixel point in the moving image in the x, y, and z axial directions. The first deformation field is applied to the moving image to displace and interpolate each voxel in the moving image, complete the deformation, and accurately match the position and shape of the moving image and the fixed image.
[0063] The first registration network takes a fully convolutional neural network (FCN) as a basic structure, and includes a downsampling module and an upsampling interpolation module. The loss function can use a global correlation coefficient (Global Correlation coefficient) to realize global large deformation registration.
[0064] The second registration network based on unsupervised deformation of the convolutional neural network is a fine registration / local small deformation registration network. The registered image obtained through the first registration network and the spatial transformation is taken as a new moving image, and the reference image or the template image is still taken as a fixed image. The new moving image and the fixed image are input into the second registration network according to the input mode of the first registration network for training, a second deformation field is obtained, and then a spatial transformation is performed to obtain a registered image corresponding to individual data. The second registration network obtains the second deformation field by maximizing the similarity between the new moving image and the fixed image. The second deformation field contains the displacement of each pixel point in the new moving image in the x, y and z axial directions. The deformation field is applied to the new moving image, and each voxel point in the new moving image is displaced and interpolated to complete the deformation, so that the new moving image and the fixed image are accurately matched in position.
[0065] The second registration network is a U-shaped network with an encoder and a decoder, and has a skip connection design between the encoder and the decoder. According to the embodiment of the application, the local correlation coefficient can be used as a loss function to achieve further local small deformation registration.
[0066] Figures 9a to 9c and Figures 10a to 10c Two examples of data are taken as examples to show the superimposed images of the preprocessed moving image and the fixed image, the superimposed images of the moving image and the fixed image registered through the first registration network, and the superimposed images of the moving image and the fixed image registered through the second registration network. Among them, Figure 9a and Figure 10a The preprocessed magnetic resonance images of the UTE sequence are shown (from left to right, the axial image, the sagittal image and the coronal image), Figure 9b and Figure 10b The superimposed images of the moving image and the fixed image registered through the first registration network are shown (from left to right, the axial image, the sagittal image and the coronal image), Figure 9c and Figure 10c The superimposed images of the moving image and the fixed image registered through the second registration network are shown (from left to right, the axial image, the sagittal image and the coronal image).
[0067] According to the embodiment of the present application, the reference image or template image is taken as the moving image, the preprocessed UTE image requiring ear cartilage and substructure segmentation is taken as the fixed image, the two-stage deformation field is obtained by using the trained first registration network and second registration network, the manual label image corresponding to the reference image or template image is continuously subjected to spatial transformation by using the first deformation field and the second deformation field to obtain the coarse label of the corresponding fixed image, and the subsequent segmentation network is trained by using the coarse label.
[0068] The manual label of the reference image or template image includes an ear cartilage label image and a substructure label image, and the coarse label of the ear cartilage and substructure is obtained by two-stage registration for any group of external ear contour UTE images requiring segmentation. In the present application, the coarse label image of the ear cartilage is used to train the network, so as to segment the ear cartilage image, and the coarse label image of the ear cartilage substructure is used to train the network, so as to segment the ear cartilage substructure image.
[0069] The segmentation network is a U-shaped network with an encoder and a decoder, and has a skip connection design between the encoder and the decoder, and an attention module is added at each skip connection layer, which is arranged between the encoder and the decoder in the same segmentation network, more encoding layer information is introduced, the segmentation network is optimized, and the final segmentation result is obtained. According to the embodiment of the present application, the attention mechanism module can be an attention gate module, and the loss function can be a Dice coefficient loss function.
[0070] The ear cartilage and substructure image segmentation system based on the matching and segmentation cascade deep learning according to the embodiment of the present application only needs to realize the ear cartilage segmentation of all data by using a small amount of individual reference or template label. In the framework of the registration network, the cascade mode of the multi-level neural network is used to realize the non-rigid registration, the accuracy is guaranteed, and the required time is much less than that of the traditional registration mode, and finally the cascade segmentation network is used to optimize the segmentation result.
[0071] Figure 11a and Figure 11b and Figure 12a and Figure 12b The three-dimensional display diagrams of the ear cartilage deep learning network segmentation and manual segmentation are shown by taking two data as examples. Among them, Figure 11a and Figure 12a is a three-dimensional display diagram of ear cartilage manual segmentation; Figure 11b and Figure 12b is a three-dimensional display diagram of ear cartilage deep learning network segmentation.
[0072] Figure 13a and Figure 13b andFigure 14a and Figure 14b Three-dimensional display diagrams of ear cartilage substructure deep learning network segmentation and manual segmentation are shown by taking two data as examples. Among them, Figure 13a and Figure 14a Three-dimensional display diagrams of ear cartilage substructure manual segmentation are shown; Figure 13b and Figure 14b Three-dimensional display diagrams of ear cartilage substructure deep learning network segmentation are shown.
[0073] In the present application, the result index of ear cartilage segmentation can be evaluated by using the Dice similarity coefficient (DSC) :
[0074]
[0075] It is calculated that the result index DSC mean of the ear cartilage and its substructure image segmentation system based on the matching and segmentation cascade deep learning of the present application is 83.7%.
[0076] In addition, the result index of ear cartilage substructure segmentation can be evaluated by using the Dice similarity coefficient (DSC) and 95% Hausdorff surface distance (HSD95), wherein the HSD95 is calculated as follows:
[0077]
[0078] For a structure, S is the point set of the structure in the segmentation map, R is the point set of the structure in the standard annotation map, and are the corresponding edge point sets respectively. Wherein, |V| represents the number of points in a point set V, d m (v,V) represents the minimum Euclidean distance between point v and all points in point set V, inf 5%,v∈V (d m ) represents the lower bound of the maximum d m value of the first 5% in the point set V.
[0079] In addition, the result index of different network models can be evaluated by using the volumetric similarity (VS), wherein the VS is calculated as follows:
[0080]
[0081] It is calculated that the segmentation result index of different network models is shown in Table 2 as follows.
[0082] [Table 2]
[0083]
[0084] In addition, the ear cartilage and substructure image segmentation system based on the matching and segmentation cascade deep learning of the application has the following segmentation result indicators for different parts of the ear cartilage substructure as shown in Table 3.
[0085] The ear cartilage and substructure image segmentation method and system based on the matching and segmentation cascade deep learning network of the application can achieve an average DSC of 83.7% for ear cartilage segmentation and an average DSC of 67.6% for ear cartilage substructure segmentation. The segmentation results can be directly converted into STL format models for general 3D printing and 3D bioprinting.
[0086] [Table 3]
[0087] Ear cartilage substructure DSC (%) HSD95 VS (%) 1 Helix 64.7 1.93 86.1 2 Scapha 51.3 2.70 71.6 3 Antihelices 68.3 1.83 89.7 4 Helices 77.6 1.32 92.7 5 Triangular fossa 62.7 2.02 83.2 6 Concha 67.1 1.88 88.3 7 Interconchal incisure 78.8 1.25 93.5 8 Conchae 69.2 1.67 90.3 9 Helix crus 67.5 1.82 88.6 10 Cymba conchae 72.2 1.46 91.2 11 Cymbiform cavity 65.2 1.90 86.5 12 External auditory canal 66.9 1.89 87.1
[0088] Although the example methods of the application described above are represented as a series of operations for the sake of clarity of description, it is not intended to limit the order of execution of the steps, and each step can be executed simultaneously or in a different order as required. In order to implement the method according to the application, the steps shown can further include other steps, can include the remaining steps other than certain steps, or can include other additional steps other than certain steps.
[0089] The various embodiments of the application are not an exhaustive list of all possible combinations, but are intended to describe representative aspects of the application, and what is described in the various embodiments can be applied independently or in combination of two or more.
[0090] In addition, various embodiments of the application can be implemented by hardware, firmware, software, or a combination thereof. The hardware can be implemented by one or more of a graphic processor (GPU), an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a general purpose processor, a controller, a microcontroller, a microprocessor, etc.
[0091] The scope of the application is intended to include software or machine executable instructions (e.g., operating systems, applications, firmware, programs, etc.) and non-transitory computer-readable media that store such software or instructions, which cause operations in accordance with various embodiments to be performed on an apparatus or computer.
[0092] The above description presented in the exemplary embodiments is merely intended to illustrate the technical solutions of the present application, and is not intended to be complete, nor intended to limit the present application to the precise forms described. Obviously, many changes and modifications are possible according to the above teachings for those skilled in the art. The exemplary embodiments are selected and described in order to explain the specific principles of the present application and its practical applications, so that other skilled in the art can easily understand, implement and utilize various exemplary embodiments of the present application and various selected forms and modified forms thereof. The scope of protection of the present application is intended to be defined by the appended claims and their equivalents.
Claims
1. A method for image segmentation of auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades, characterized in that, Includes the following steps: Obtain magnetic resonance images of the UTE sequence of the outer ear contour; The magnetic resonance images of the UTE sequence of the acquired outer ear contour were preprocessed; The preprocessed UTE sequence image is used as input to train the registration network in the deep learning-based image segmentation system of ear cartilage and its substructures based on matching and segmentation cascade, thereby obtaining the deformation field. The deformation field is applied to the manually labeled image of the reference image or the manually labeled image of the template image to obtain the coarse label image. The obtained coarse label images of ear cartilage and ear cartilage substructures are used to train the segmentation network in the deep learning-based image segmentation system of ear cartilage and its substructures based on matching and segmentation cascade, thereby performing image segmentation of ear cartilage and ear cartilage substructures respectively. The deep learning-based image segmentation system for auricular cartilage and its substructures, based on matching and segmentation cascades, includes: a first registration network, a second registration network, and a segmentation network. One image is randomly selected from the preprocessed UTE auricular cartilage images as the moving image, and another image is randomly selected as the fixed image to form a set of input images, which are used as the input of the first registration network. The first deformation field is obtained through the first registration network, and the moving image is spatially transformed using the first deformation field to obtain the deformed moving image. The deformed moving image and the fixed image are used as the input image group of the second registration network. The second deformation field is obtained through the second registration network. The deformed moving image is then spatially transformed again using the second deformation field to finally obtain the moving image that matches the fixed image, thus completing the training of the first registration network and the second registration network. After the first and second registration networks are trained, images with clear and complete auricle shape and structure are selected from the moving images as reference images, or a template image is constructed using part or all of the data through inter-group matching. Manual segmentation is then performed on the reference image or the template image to obtain the manually labeled image of the reference image or the manually labeled image of the template image. Using the reference image or template image as the moving image and the preprocessed UTE sequence image that needs to be segmented into ear cartilage and substructures as the fixed image, the first deformation field and the second deformation field are obtained by using the trained first registration network and the second registration network. The manually labeled image of the reference image or the manually labeled image of the template image are spatially transformed using the first deformation field and the second deformation field to obtain the coarse label image of the corresponding fixed image, and the segmentation network is trained using the coarse label image.
2. The method for image segmentation of auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades as described in claim 1, characterized in that, The first registration network uses a fully convolutional neural network as its basic structure, which includes a downsampling module and an upsampling interpolation module. The loss function of the first registration network is the global cross-correlation coefficient loss function. The second registration network and segmentation network use an encoder-decoder U-shaped framework as the backbone network, and each includes an encoder and a decoder. The encoder and decoder of the second registration network have skip connections, and the loss function of the second registration network is the local cross-correlation coefficient loss function; The segmentation network has skip connections between the encoder and decoder, and an attention module is added to each skip connection layer. The loss function of the segmentation network is the Dice similarity coefficient loss function.
3. The method for image segmentation of auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades as described in claim 1, characterized in that, When acquiring magnetic resonance images of the UTE sequence of the outer ear contour, the echo time is set to 0-1ms depending on the MRI scanning equipment used; The magnetic resonance images of the UTE sequence include axial images, sagittal images, and coronal images; Preprocessing of the acquired UTE sequence magnetic resonance images of the outer ear contour includes: format conversion, renaming, reorientation, and rigid registration of the images.
4. The method for image segmentation of auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades as described in claim 1, characterized in that, The manually segmented ear cartilage label images were obtained based on manual operations by the operator; When segmenting auricular cartilage labeled images based on manual operation by the operator, the boundaries of auricular cartilage are segmented based on the fact that in the magnetic resonance images of the preprocessed UTE sequence, auricular cartilage, surrounding fat and connective tissue all appear as bright signals to varying degrees, while skin and other tissues are relatively dark.
5. The method for image segmentation of auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades according to claim 1, characterized in that, The manually segmented auricular cartilage substructure label images were obtained based on manual operation by the operator; When segmenting the labeled images of auricular cartilage substructures according to the operator's manual operation, a total of twelve auricular cartilage substructure images are segmented in the following order: antihelix, triangular fossa and crus of antihelix, helix and crus of helix, cymba conchae, scaphoid fossa, tragus, antitragus and intertragusal notch, and external auditory canal and conchae cavity.
6. An image segmentation system for auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades, characterized in that, Includes the following modules: The acquisition module acquires magnetic resonance images of the UTE sequence of the outer ear contour; The preprocessing module preprocesses the magnetic resonance images of the UTE sequence of the acquired outer ear contour. The training module takes the preprocessed UTE sequence image as input and trains the registration network in the deep learning-based image segmentation system for ear cartilage and its substructures based on matching and segmentation cascade to obtain a deformation field. The deformation field is applied to the manually labeled image of the reference image or the manually labeled image of the template image to obtain a coarse label image. The obtained coarse label images of ear cartilage and ear cartilage substructures are used to train the segmentation network in the deep learning-based image segmentation system for ear cartilage and its substructures based on matching and segmentation cascade, thereby performing image segmentation on ear cartilage and ear cartilage substructures respectively. The deep learning-based image segmentation system for auricular cartilage and its substructures, based on matching and segmentation cascades, includes: a first registration network, a second registration network, and a segmentation network. One image is randomly selected from the preprocessed UTE auricular cartilage images as the moving image, and another image is randomly selected as the fixed image to form a set of input images, which are used as the input of the first registration network. The first deformation field is obtained through the first registration network, and the moving image is spatially transformed using the first deformation field to obtain the deformed moving image. The deformed moving image and the fixed image are used as the input image group of the second registration network. The second deformation field is obtained through the second registration network. The deformed moving image is then spatially transformed again using the second deformation field to finally obtain the moving image that matches the fixed image, thus completing the training of the first registration network and the second registration network. After the first and second registration networks are trained, images with clear and complete auricle shape and structure are selected from the moving images as reference images, or a template image is constructed using part or all of the data through inter-group matching. Manual segmentation is then performed on the reference image or the template image to obtain the manually labeled image of the reference image or the manually labeled image of the template image. Using the reference image or template image as the moving image and the preprocessed UTE sequence image that needs to be segmented into ear cartilage and substructures as the fixed image, the first deformation field and the second deformation field are obtained by using the trained first registration network and the second registration network. The manually labeled image of the reference image or the manually labeled image of the template image are spatially transformed using the first deformation field and the second deformation field to obtain the coarse label image of the corresponding fixed image, and the segmentation network is trained using the coarse label image.
7. The image segmentation system for auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades as described in claim 6, characterized in that, The first registration network uses a fully convolutional neural network as its basic structure, which includes a downsampling module and an upsampling interpolation module. The loss function of the first registration network is the global cross-correlation coefficient loss function. The second registration network and segmentation network use an encoder-decoder U-shaped framework as the backbone network, and each includes an encoder and a decoder. The encoder and decoder of the second registration network have skip connections, and the loss function of the second registration network is the local cross-correlation coefficient loss function; The segmentation network has skip connections between the encoder and decoder, and an attention module is added to each skip connection layer. The loss function of the segmentation network is the Dice similarity coefficient loss function.
8. The image segmentation system for auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades as described in claim 6, characterized in that, When acquiring magnetic resonance images of the UTE sequence of the outer ear contour, the echo time is set to 0-1ms depending on the MRI scanning equipment used; The magnetic resonance images of the UTE sequence include axial images, sagittal images, and coronal images; Preprocessing of the acquired UTE sequence magnetic resonance images of the outer ear contour includes: format conversion, renaming, reorientation, and rigid registration of the images.
9. The image segmentation system for auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades as described in claim 6, characterized in that, The manually segmented ear cartilage label images were obtained based on manual operations by the operator; When segmenting auricular cartilage labeled images based on manual operation by the operator, the boundaries of auricular cartilage are segmented based on the fact that in the magnetic resonance images of the preprocessed UTE sequence, auricular cartilage, surrounding fat and connective tissue all appear as bright signals to varying degrees, while skin and other tissues are relatively dark.
10. The image segmentation system for auricular cartilage and its substructures based on a deep learning network of matching and segmentation cascades according to claim 6, characterized in that, The manually segmented auricular cartilage substructure label images were obtained based on manual operation by the operator; When segmenting the labeled images of auricular cartilage substructures according to the operator's manual operation, a total of twelve auricular cartilage substructure images are segmented in the following order: antihelix, triangular fossa and crus of antihelix, helix and crus of helix, cymba conchae, scaphoid fossa, tragus, antitragus and intertragusal notch, and external auditory canal and conchae cavity.
Citation Information
Patent Citations
Automatic segmentation of brain tumor images based on convolution neural network
CN109035263A
Full-automatic registration and segmentation method for multi-parameter magnetic resonance image
CN111260700A