A deep learning-based oral and maxillofacial multi-modal image fusion and segmentation method and system

By using deep learning models to automatically process CT and MRI images, the problem of time-consuming, labor-intensive, and inconsistent processes in existing technologies has been solved, achieving efficient and accurate image fusion and segmentation, and supporting precise surgical planning for oral and maxillofacial tumors.

CN122199471APending Publication Date: 2026-06-12PEKING UNIV SCHOOL OF STOMATOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PEKING UNIV SCHOOL OF STOMATOLOGY
Filing Date
2026-03-13
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing CT/MRI multimodal image fusion and segmentation techniques rely on human experience, are time-consuming and labor-intensive, produce inconsistent results, and lack efficient and automated fusion and segmentation models tailored to the complex anatomical features of the oral and maxillofacial region.

Method used

A deep learning-based image fusion and segmentation method is adopted. After preprocessing CT and MRI images, they are input into a deep learning model for registration and fusion to segment a three-dimensional model of the target structure and assist in virtual surgical planning.

Benefits of technology

It achieves high-precision and automated image fusion and segmentation, significantly shortens processing time, outputs stable and reliable 3D segmentation models, assists in precise surgical design, and improves clinical work efficiency and safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122199471A_ABST
    Figure CN122199471A_ABST
Patent Text Reader

Abstract

The application discloses a kind of oral cavity jaw face multi-modal image fusion and segmentation method and system based on deep learning, belong to oral cavity jaw face image processing field.Method includes: S1: obtaining the computer tomography (CT) image and magnetic resonance imaging (MRI) image of the same patient's oral cavity jaw face region;S2: pre-processing, to obtain the standardized CT image and MRI image;S3: the standardized CT image and MRI image are input into the pre-trained deep learning image fusion model for registration and fusion, to obtain multi-modal fusion image;S4: the multi-modal fusion image is input into the pre-trained deep learning image segmentation model, and the three-dimensional model of at least one target structure is segmented and output.The present application overcomes the defects of the prior art, such as strong subjectivity, time-consuming, inconsistent results, etc., and provides stable and reliable technical support for clinical precise treatment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of oral and maxillofacial image processing technology, and more specifically to a method and system for multimodal image fusion and segmentation of the oral and maxillofacial region based on deep learning. Background Technology

[0002] Oral and maxillofacial tumors (especially midfacial tumors) often invade both bone and surrounding soft tissues, and this region has a complex anatomy, adjacent to important blood vessels and nerves. Therefore, accurate preoperative assessment of tumor location, extent of invasion, and involvement of adjacent important structures through medical imaging is crucial for determining safe resection margins and reducing surgical risks.

[0003] Currently, clinical practice primarily relies on single-modal imaging such as computed tomography (CT) and magnetic resonance imaging (MRI) for diagnosis and treatment planning. CT can clearly display bony structures but lacks sufficient resolution for soft tissues; MRI excels at depicting soft tissues and tumor boundaries but is poor at displaying bone. This limitation of single-modal imaging makes it difficult for physicians to comprehensively and intuitively assess the true extent of a tumor's involvement in both soft and hard tissues. To address this issue, multimodal image fusion and segmentation technology has emerged. By integrating complementary information from CT and MRI, it aims to achieve precise segmentation and three-dimensional visualization of the target area, thereby improving the accuracy and personalization of diagnosis and treatment throughout the entire process from early diagnosis and treatment planning to efficacy evaluation.

[0004] However, existing CT / MRI multimodal image fusion and segmentation technologies have significant shortcomings: First, the operation process is highly dependent on human experience, usually requiring doctors to manually perform point-to-point registration or boundary delineation in specialized software, a time-consuming and laborious process that often takes several hours to process a set of images; second, the consistency and reproducibility of the results are poor, with significant differences in fusion and segmentation accuracy between different operators, and even between the same operator processing the same set of images at different times; third, the operating software is complex, and training a doctor to be proficient in multimodal processing requires a long period of time, which limits the popularization and application of this technology in clinical practice.

[0005] With the advancement of artificial intelligence technology, deep learning has shown great potential in medical image analysis. Although some studies have attempted to apply deep learning to image segmentation of head and neck tumors, the following limitations still exist: 1) Existing studies mostly focus on the final segmentation results, lacking systematic optimization and evaluation of the multimodal image fusion strategy itself, and failing to explore the synergistic relationship between fusion quality and subsequent segmentation efficiency; 2) In terms of modality selection, most studies are based on PET / CT, while research on deep learning fusion and segmentation of CT and MRI dual modalities, which are more anatomically specific for oral and maxillofacial tumors, is still relatively scarce; 3) The model application is relatively singular, and a systematically compared and screened "fusion-segmentation" integrated model combination has not yet been formed for the complex anatomical characteristics of the oral and maxillofacial region (such as the simultaneous inclusion of rigid bones and non-rigid soft tissues), and there is a lack of systematic evaluation of the application effect of this combination in real clinical scenarios.

[0006] Therefore, how to provide a technical solution that can automate, accurately, and efficiently complete the fusion and segmentation of CT / MRI multimodal images of oral and maxillofacial tumors is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a method and system for multimodal image fusion and segmentation of the oral and maxillofacial region based on deep learning, so as to overcome the defects of existing technologies such as strong subjectivity, time consumption and inconsistent results, and provide stable and reliable technical support for precise clinical treatment.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: A deep learning-based method for multimodal image fusion and segmentation of the oral and maxillofacial region includes the following steps: S1: Acquire computed tomography (CT) and magnetic resonance imaging (MRI) images of the oral and maxillofacial region of the same patient; S2: Preprocess the computed tomography (CT) images and magnetic resonance imaging (MRI) images respectively to obtain standardized CT images and MRI images; S3: Input the standardized CT and MRI images into a pre-trained deep learning image fusion model for registration and fusion to obtain multimodal fused images; S4: Input the multimodal fused image into a pre-trained deep learning image segmentation model to segment at least one three-dimensional model of the target structure and output it.

[0009] A deep learning-based multimodal image fusion and segmentation system for the oral and maxillofacial region includes: Data acquisition module: Acquires computed tomography (CT) images and magnetic resonance imaging (MRI) images of the oral and maxillofacial region of the same patient; Preprocessing module: preprocesses the computed tomography (CT) images and magnetic resonance imaging (MRI) images respectively to obtain standardized CT images and MRI images; Image fusion module: Standardized CT and MRI images are input into a pre-trained deep learning image fusion model for registration and fusion to obtain multimodal fused images; Image segmentation module: Input the multimodal fused image into a pre-trained deep learning image segmentation model, segment out at least one three-dimensional model of the target structure and output it.

[0010] Furthermore, it also includes: Decision-making support module: Based on the segmented 3D model of the target structure, perform virtual surgical planning, determine the tumor resection boundary and / or plan the surgical path.

[0011] As can be seen from the above technical solution, compared with the prior art, the present invention provides a method and system for multimodal image fusion and segmentation of the oral and maxillofacial region based on deep learning, with the following specific beneficial effects: (1) By employing an optimized combination of deep learning models, this invention demonstrates outstanding performance on key evaluation metrics. In the fusion stage, the model achieves a higher level of local fusion accuracy (FI) for tumors; in the segmentation stage, the segmentation Dice coefficient for key structures such as tumors, maxillae, mandibles, and blood vessels is higher, and the 95% Hausdorff distance (HD95) and mean surface distance (MSD) are lower. This means that the segmentation boundary is more closely aligned with the actual anatomical structure, providing a more reliable imaging basis for precise surgical planning.

[0012] (2) This invention automates the traditional, cumbersome, and highly manual multimodal processing workflow. The deep learning model can achieve rapid fusion and segmentation in seconds to tens of seconds. The overall processing time (including necessary manual review and correction) is orders of magnitude shorter than that of purely manual methods (which usually take several hours), significantly improving clinical work efficiency and better adapting to the tight pace of preoperative preparation.

[0013] (3) Based on the objective processing of deep learning algorithms, this invention effectively eliminates the subjective differences between different operators and the fluctuations caused by the different states of the same operator at different times. For the same input data, the model can output highly consistent, stable and repeatable fusion and segmentation results, which improves the standardization level and reliability of the diagnosis and treatment process.

[0014] (4) This invention not only outputs fused images, but also directly provides precise three-dimensional segmentation models of key anatomical structures. Based on this, doctors can directly design virtual surgeries, intuitively plan osteotomy lines and resection ranges, and pre-mark intraoperative biopsy locations. Preliminary clinical applications show that this technology can effectively assist in intraoperative margin control, improve the negative margin rate, and thus is expected to improve the complete tumor resection rate and patient prognosis.

[0015] (5) This invention is presented in the form of packaged software or system, which simplifies the operation process. Clinicians do not need to undergo long-term training in complex software operation to quickly obtain high-quality multimodal fusion and segmentation results using this tool, which is conducive to the promotion and application of this advanced technology in medical institutions at all levels. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the method flow provided by the present invention; Figure 2 This is a schematic diagram of the operation path of a deep learning image fusion model. Figure 3 Flowchart of the VoxelMorph model; Figure 4 This is a diagram of the VoxelMorph network structure. Figure 5 Flowchart of the TransMorph model; Figure 6 Here is a diagram of the TransMorph network structure; Figure 7 Here is a flowchart of the NeMAR model; Figure 8 This is a diagram of the NeMAR network structure. Figure 9 Here is a flowchart of the BFM model; Figure 10 Here is a diagram of the BFM network structure; Figure 11 A schematic diagram of the operation path of a deep learning image segmentation model; Figure 12 A 3D U-Net network diagram; Figure 13 This is a 3D UX-Net network diagram; Figure 14This is a diagram of the nnU-Net network. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] See Figure 1 This invention discloses a deep learning-based method for multimodal image fusion and segmentation of the oral and maxillofacial region, comprising the following steps: S1: Acquire computed tomography (CT) and magnetic resonance imaging (MRI) images of the oral and maxillofacial region of the same patient; S2: Preprocess the computed tomography (CT) images and magnetic resonance imaging (MRI) images respectively to obtain standardized CT images and MRI images; S3: Input the standardized CT and MRI images into a pre-trained deep learning image fusion model for registration and fusion to obtain multimodal fused images; S4: Input the multimodal fused image into a pre-trained deep learning image segmentation model to segment at least one three-dimensional model of the target structure and output it.

[0020] The deep learning image fusion model is an unsupervised neural modality-independent registration (NeMAR) model.

[0021] Specifically, in the surgical treatment of oral and maxillofacial tumors, the foundation of virtual surgical design is accurate multimodal image fusion, followed by precise and rapid multimodal image segmentation based on image fusion, resulting in the segmentation and annotation of the tumor and important anatomical structures. The accuracy of multimodal image fusion directly affects the subsequent segmentation accuracy, which in turn affects the accuracy of virtual surgical design, thus impacting the surgical outcome. This invention aims to establish and validate a deep learning model for multimodal image fusion suitable for oral and maxillofacial tumors, thereby improving the accuracy of imaging examinations and laying a solid foundation for subsequent image segmentation.

[0022] This invention focuses on multimodal image registration and fusion between CT and MRI for oral and maxillofacial tumors, aiming to construct a reproducible and deployable technical approach in clinical workflows. To this end, three mainstream models are compared: rigid registration and fusion using traditional non-deep learning models Elastix, ANTs, and NiftyReg; rigid registration and fusion using deep learning models VoxelMorph, TransMorph, and NeMAR; and non-rigid registration and fusion using the deep learning model BFM. Under unified data and workflow, the applicability and advantages and disadvantages of each method are analyzed and clarified, and based on this, a comparative framework and implementation path for oral and maxillofacial tumor scenarios are established.

[0023] The validation and evaluation system uses Normalized Mutual Information (NMI) and the Structural Similarity Index Measure (SSIM) to measure overall consistency across modal intensity levels, Average Symmetric Surface Distance (ASD / ASSD) to characterize boundary geometric errors, and Recall / Precision and Fusion Index (FI) to evaluate the spatial overlap and detection balance of anatomical targets after registration and fusion. Simultaneously, model runtime is recorded as one of the indicators of model usability. This combination of indicators addresses clinical concerns regarding global alignment, local alignment, and target consistency, and also provides a benchmark for comparing the clinical applications of different methods.

[0024] For multimodal image fusion of CT and MRI in oral and maxillofacial tumors, this invention constructs a fusion model within a deep learning framework and provides an operable technical route. Through a cross-comparison of traditional methods and the deep learning model, it elucidates the construction and optimization strategies of the fusion process, such as... Figure 2 As shown.

[0025] This embodiment selected oral and maxillofacial tumor cases treated in the Department of Oral and Maxillofacial Surgery of a certain hospital from January 2016 to August 2025. All imaging data were anonymized and informed consent was waived after evaluation by the ethics committee. Inclusion criteria: (1) The tumor was located in the deep oral cavity (gingiva of the upper / lower mandibular posterior teeth, soft palate, tongue root or parapharynx, etc.) or the deep maxillofacial region (maxilla, mandibular ramus, zygomatic bone or skull base-infratemporal fossa, etc.), involving two or more anatomical regions; (2) The patient had undergone CT and MRI examinations before surgery and had imaging data in two modalities of digital imaging and communication in medicine (DICOM) format; (3) The scanning range of the imaging examination was above the infraorbital margin and below the chin apex; (4) The tumor was primary or recurrent, with or without pathological diagnosis. Exclusion criteria: (1) Scanning parameters (slice thickness, pixel size, etc.) of the imaging data cannot be obtained or are incomplete; (2) The time interval between two imaging examinations is greater than 20 days, in which case the tumor may have deformed between different images due to volume growth.

[0026] Since there is no explicit limit on the number of cases during the training process of deep learning models, based on the actual clinical situation, this invention intends to include 50 pairs of paired CT and MRI monomodal images, totaling 11,636 tomographic images, including 8,817 CT tomographic images and 2,819 MRI tomographic images.

[0027] CT scan parameters were as follows: scanning voltage 120kV, scanning current 140-160mA, field of view 200mm, pixel matrix 512×512, slice thickness 1.25mm. All patients underwent ceCT scans after iodine contrast agent injection. MRI T1-weighted scan parameters were as follows: scanning magnetic field strength 1.5T, repetition time (TR) 600ms, echo time (TE) 16ms, flip angle 120°, field of view 250mm, pixel matrix 320×320, slice thickness 2mm. MRI T2-weighted scan parameters were as follows: scanning magnetic field strength 1.5T, repetition time 5000ms, echo time 99ms, flip angle 125°, field of view 230mm-260mm, pixel matrix 320×320, slice thickness 4mm. DICOM format data of the patient's imaging examinations were acquired and imported into BrainLAB iPlan CMF 3.0 software (BrainLAB, Germany). Axial images from enhanced CT and T2-weighted fat-suppressed fast spin echo (T2 fs FSE) sequences from MRI are extracted. These parameter settings ensure soft tissue contrast while balancing scan time and patient comfort, making them particularly suitable for image analysis of complex structures in the oral and maxillofacial region.

[0028] By querying patient diagnostic information, extracting inherent image features, and reviewing the specific content of the images, the following data were obtained for each pair of monomodal images: (1) Tumor nature: ① Benign; ② Malignant; (2) Tumor location: ① Maxilla: including tumors located in the maxilla, upper gingiva, soft palate, zygomatic bone, parapharynx, and skull base-infratemporal fossa; ② Mandible: tumors located in the mandible, lower gingiva, tongue, and floor of mouth; (3) Presence or absence of artifacts: ① Yes: artifacts are present in single-modal images; ② No: no artifacts are seen in any single-modal images; (4) Presence or absence of positional changes: ① Yes: images show that the patient's position was not consistent when receiving different modal image scans (e.g., inconsistent head position, upper and lower teeth not in centric relation, etc.); ② No: no such changes were found.

[0029] This embodiment obtains the following data for each pair of monomodal images by labeling the tumor layer by layer and extracting the inherent features of the images from the DICOM data: Tumor volume: the average volume of the two three-dimensional tumor models generated from the tumor range labeled layer by layer on the monomodal image.

[0030] After image data acquisition, systematic data preprocessing is required to ensure the accuracy and stability of subsequent analysis and registration fusion algorithms. All acquired DICOM format images are first standardized and pre-screened within the mitk (Medical Imaging Interaction Toolkit) system to remove obvious artifacts, motion artifacts, and acquisition defects, ensuring that the data included in the analysis has good spatial continuity and structural integrity, and contributing to improved image feature analysis and model repeatability.

[0031] Regarding image spatial standardization, all CT and MRI images must be uniformly resampled to the same voxel resolution and spatial size. According to CT and MRI technical specifications, images are resampled to 1.0 mm³ isotropic voxels, with a unified spatial matrix of 256×256×256, ensuring pixel-level registration and fusion of data from different modalities and time points. By utilizing the resampling function of SimpleITK (Insight Segmentation and Registration Toolkit) and the logic of precisely configuring the "physical space—voxel grid—interpolator," this invention achieves precise alignment of multimodal images, effectively eliminating spatial resolution differences caused by different devices and acquisition conditions.

[0032] In image registration and fusion, the CT and MRI data are first coarsely registered (rigid registration) to achieve preliminary rigid alignment (translation and rotation) between the two sets of images in three-dimensional space. This process uses SimpleITK's VersorRigid3DTransform + Mattes MI for rigid coarse registration. After coarse registration, manual checks are performed to confirm that all original data have been correctly pre-registered. Based on this, affine registration or manual fine-tuning based on feature points is further employed to address any residual errors, ensuring that the spatial positions of key anatomical structures such as the skull base, maxilla, and mandible are highly consistent in the two sets of images. This process conforms to the classic hierarchical registration and fusion paradigm and theoretical framework of "from rigid to flexible, from coarse to fine".

[0033] Furthermore, to reduce interference from non-anatomically relevant areas, all images require Region of Interest (ROI) extraction. Target regions (jawbone, tumor, blood vessels) can be segmented and cropped semi-automatically or manually to remove irrelevant tissue or air pockets, thereby improving the targeting and effectiveness of subsequent analysis and algorithm training. This is a common method in MRI craniotomy / non-brain tissue removal and foreground cropping in deep learning workflows.

[0034] To address the differences in grayscale distribution across multimodal images, intensity normalization is necessary. This includes methods such as linear normalization, Z-score standardization, and histogram matching to reduce the impact of grayscale differences between different modalities and patients on the performance of registration and learning algorithms. For MRI data, N4ITK bias field correction is often used to improve intensity stability and comparability. In the context of radiomics research, the necessity of intensity normalization and its positive impact on feature robustness are supported by numerous empirical studies. Furthermore, a training:validation:test set ratio of 7:2:1 is used to validate the model.

[0035] After the above multi-step data preprocessing, the resulting image data has a uniform spatial resolution, consistent anatomical structure distribution, and standardized grayscale characteristics, providing high-quality and reusable basic data for subsequent three-dimensional registration and fusion, deep learning algorithm training, and surgical navigation system development. The standardized and streamlined preprocessing also directly improves the scientific rigor and reproducibility of the research, and lays a solid foundation for multi-center clinical research and big data analysis.

[0036] Elastix, ANTs, and NiftyReg were chosen for comparison because they represent classic paradigms of traditional multimodal registration and fusion, are mature in CT / MRI scenarios, are open-source and reproducible, offer stable accuracy, transparent parameters, and are easy to implement clinically. At the clinical level, all three have been extensively validated on multimodal data, demonstrating stable rigid / affine pose correction, interpretable transformation parameters, ease of quality control and traceability, and seamless integration with intraoperative navigation procedures. Therefore, they can serve as directly deployable rigid fusion baselines and provide reliable controls for subsequent deep learning model evaluation.

[0037] Elastix is ​​an open-source medical image registration toolkit based on ITK, supporting various registration and fusion models such as rigid, affine, and B-spline. It is widely used due to its flexibility and robustness. In multimodal image registration and fusion, such as CT and MRI, Elastix can employ evaluation criteria like NMI to address intensity differences between modalities, exhibiting good robustness. Elastix uses a multi-resolution pyramid and multi-stage optimization strategy to improve fusion accuracy while controlling computational load. The Elastix tool provides a convenient scripting interface, facilitating use in clinical work or research workflows. Its open-source community is also active, with abundant plugin extensions. Overall, Elastix demonstrates a balanced performance in fusion accuracy, robustness, and practicality, making it one of the commonly used benchmark methods for multimodal rigid registration and fusion. Elastix's rigid registration and fusion function has the advantages of maturity, stability, and efficiency in the field of medical image processing, supporting multimodal, scalable, and flexible modular design. Its innovation lies in its highly parameterized and configurable design, allowing users to easily customize the optimal registration and fusion scheme for specific application scenarios. Elastix has proven its accuracy and robustness, but it is prone to local extrema when there is initial misalignment or large multimodal differences. It is also highly sensitive to parameters such as multiresolution and multi-iteration, and its configuration is cumbersome. Over-setting parameters may also produce unreasonable deformations. The main bottleneck in clinical application lies in the reliance on parameters and human experience.

[0038] ANTs (Advanced Normalization Tools) is another well-known open-source tool in the field of medical image fusion, widely recognized for its top-tier accuracy in multiple evaluations. For rigid registration fusion, ANTs also provides multimodal similarity metrics (such as Mattes mutual information) and multi-scale optimization schemes, enabling high-precision CT / MRI fusion. Especially after fully optimizing parameters (e.g., using multi-level pyramids and gradient step size tuning) and combining mutual information as a similarity metric, ANTs typically achieves high-precision fusion results. ANTs' fusion algorithm has a solid theoretical foundation (such as the symmetric normalization model). ANTs is widely recognized as a benchmark tool in neuroimaging, and multiple studies recommend it as one of the most accurate fusion schemes. In summary, ANTs excels in fusion accuracy and alignment of anatomical details, precisely aligning the bony anatomy and tumor location in the maxillofacial region in multimodal rigid registration fusion, providing a reliable basis for subsequent tumor annotation.

[0039] NiftyReg is an open-source registration and fusion tool developed by institutions such as UCL (University College London), aiming to achieve higher operational efficiency while maintaining accuracy. It supports rigid, affine, and free-form registration and fusion, and provides CUDA (Compute Unified Device Architecture) GPU (Graphics Processing Unit) acceleration options. In rigid registration and fusion of CT / MRI oral and maxillofacial images, NiftyReg utilizes block matching and gradient descent-based optimizations to improve computational speed to some extent. Due to its GPU parallel computing implementation, NiftyReg runs quickly on medium-sized datasets, and the registration and fusion results are stable. Its open-source BSD (Berkeley Software Distribution) license and cross-platform support also facilitate researchers' use in their custom workflows. Overall, NiftyReg provides a trade-off between accuracy and efficiency in traditional algorithms, achieving a certain level of acceleration and convenience while ensuring the accuracy of multimodal rigid registration and fusion.

[0040] Rigid registration unifies images from different time phases / modalities into the same coordinate system through rotation and translation, enabling better registration of rigid structures (jawbone). We selected three models: VoxelMorph, TransMorph, and NeMAR, because they represent three main lines of deep learning rigid multimodal fusion: the efficient unsupervised CNN baseline (VoxelMorph, inference in seconds, clinically feasible), the robust Transformer architecture with large pose (large initial position differences) (TransMorph, better accuracy and stability), and the "modal translation + single-modal fine-fit" paradigm (NeMAR, directly reducing CT / MRI modal differences and improving cross-modal alignment). Together, these three models cover the key dimensions of efficiency, accuracy, and robustness, and are open-source and reproducible, serving as representative deep model comparisons and clinical alternatives. Clinically: VoxelMorph has low hardware requirements, making it suitable for rapid initial fusion in outpatient batches, which can significantly reduce pre-annotation waiting time and improve clinical efficiency; TransMorph is more robust to large pose differences / FOV (Field of View) differences, has a low failure rate, and can output displacement fields for quality control and review, which is beneficial for obtaining more stable and consistent coordinates in complex positions and mixed bone-soft tissue areas of the oral and maxillofacial region; NeMAR directly addresses CT / MRI grayscale differences with "modal translation + single-modal fine matching", improves boundary consistency and tumor neighborhood availability (such as local tumor FI enhancement), reduces annotation bias caused by cross-modal mismatch, and is more likely to maintain consistency in cross-center / cumulative scenarios.

[0041] VoxelMorph is a pioneering method in the field of deep learning image registration and fusion, offering a novel solution to the speed bottlenecks and parameter tuning problems of traditional methods. VoxelMorph directly learns the mapping from moving images to the deformation field of fixed images through a convolutional neural network, allowing traditional iterative optimization to be "pre-emptively" implemented during the training phase, thus achieving real-time registration and fusion during inference. Compared to Elastix / ANTs, which requires hundreds of iterations for each image pair, VoxelMorph, after training, only needs a single forward propagation for new CT / MRI images to output rigid transformation parameters and deformation fields. This makes VoxelMorph highly clinically promising for applications such as intraoperative navigation requiring rapid alignment. Secondly, VoxelMorph employs an unsupervised learning paradigm, not relying on manually labeled registration ground truth values, but training the network through a similarity metric between images (based on pixel intensity loss). For data such as CT / MRI images of oral and maxillofacial tumors, which lack pre-labeled registration ground truth values, VoxelMorph can fully utilize unlabeled image pairs for training, learning implicit intermodal alignment patterns. Furthermore, VoxelMorph achieves accuracy close to traditional optimization algorithms. This performance improvement is attributed to its unsupervised learning method, which is not only fast but also requires no annotation information, offering significant advantages over traditional registration methods that require optimization of metric functions. In the fusion of oral and maxillofacial tumor images, it demonstrates superior fusion results for tumor regions. Moreover, VoxelMorph is open-source (MIT license), well-maintained by the community, and has numerous variants and extensions (such as adding multi-resolution and regularization terms) to enhance robustness and accuracy. Overall, VoxelMorph successfully addresses the inefficiency of traditional methods, achieving an efficient, automated, and remarkably accurate fusion solution.

[0042] VoxelMorph's network employs a convolutional neural network structure based on U-Net. Its input is a fixed image. and moving images The stitched data is processed by an encoder to extract multi-scale features, and then by a decoder to gradually restore the spatial resolution and generate a deformation field. Specifically, the network parameterization function... Where u is the deformation displacement field, and θ is the network parameters (convolution kernel weights). The final deformation mapping is:

[0043] Where Id is the unit mapping. The encoder progressively downsamples the feature maps through 3D convolutional layers, extracting higher-level spatial semantic features at each layer. The decoder fuses high-level and low-level features through upsampling, convolution, and skip connections, progressively generating a high-resolution deformation field; see details below. Figure 3 and Figure 4This structure combines deep global features with shallow fine-grained features, enabling registration and fusion to balance global consistency and local accuracy. To ensure the physical validity of the deformation field, the network outputs a displacement field. The receptive field range covers the expected maximum anatomical displacement range, thus enabling effective registration and fusion of a wide range of structural differences.

[0044] TransMorph is a cutting-edge method proposed in recent years that introduces Transformer into medical image registration and fusion. Building upon convolutional networks such as VoxelMorph, it further addresses the issues of limited local receptive fields and insufficient capture of complex deformations. The TransMorph model employs a hybrid architecture, combining the powerful feature extraction capabilities of Transformer with the local processing advantages of convolutional neural networks to acquire global information and output a fused field. Unlike VoxelMorph, which relies solely on local convolutions, Transformer's self-attention mechanism can cover long-range dependencies in images, thus better matching corresponding anatomical structures across large distances. This is highly advantageous for rigid fusion of oral and maxillofacial CT / MRI images. Even if the two modalities have significant differences in pose or different morphologies of tissues surrounding the tumor area, TransMorph can find the correct alignment by capturing global information, reducing the risk of getting trapped in local optima. TransMorph also provides a diffeomorphic version to ensure the topological validity of the output deformation field, avoiding the micro-local folding problem that VoxelMorph may exhibit. This topological constraint is equivalent to ensuring a strictly rigid body transformation in rigid registration fusion, enhancing the physical plausibility and stability of the results. TransMorph also introduces a Bayesian variant to quantify the uncertainty of the registration fusion results. This is important in clinical applications: when there are blurred areas or artifacts between MRI and CT, uncertainty prediction can prompt physicians to pay attention to possible fusion error areas (tumors, blood vessels). In summary, TransMorph, by introducing Transformer technology, significantly improves the ability of deep fusion networks to model global deformations in medical image registration tasks, achieving a leading position in accuracy and robustness, and providing a powerful tool for handling complex multimodal registration fusion tasks.

[0045] The core structure of TransMorph consists of a Swing Transformer encoder and a convolutional decoder. The input is a moving image. and fixed image The concatenation of these elements forms a 2×H×W×D volume graph. The encoder divides it into non-overlapping 3D patches (typically 4×4×4 in size), each patch is unfolded into a token, and after linear projection, mapped to the embedding dimension C.

[0046] in , It is a linear projection matrix.

[0047] Subsequently, a multi-layered Swin Transformer is used, with each window performing a self-attention mechanism, allowing the encoder to capture the global semantics of the image. In the decoding stage, the deformation field is concatenated and the spatial resolution is gradually restored through multi-scale convolutions and skip connections, outputting the deformation displacement field u. The final generated deformation map is:

[0048] Where p represents the voxel coordinates. For the homeomorphic variant version, the deformation field is converted into a velocity field v, and a topologically ordered deformation map is obtained through an exponential mapping:

[0049] Meanwhile, the convolutional layers and skip connection structure of the decoder ensure the synchronous transmission of global and local information. (See also) Figure 5 and Figure 6 As shown.

[0050] The basic unsupervised loss function of TransMorph includes an image similarity term and a deformation field smoothing term:

[0051] Mutual information (MI) can be used as a similarity measure:

[0052] in and These are the edge entropy values ​​for a fixed image and a moving image, respectively. It is the joint entropy.

[0053] Or locally normalized cross-correlation (CC). The smoothness term measures the gradient of the displacement field:

[0054] For homeomorphic variants, the deformation model becomes the velocity field v, generated via an exponential mapping. And ensure that its Jacobian determinant is always positive. The corresponding loss is followed by a topological order-preserving constraint after the MI and smoothing terms.

[0055] In the Bayesian version, the encoder learns a distributed deformation field parameterized as and Deformation mapping is generated through sampling. Add KL divergence regularization to the loss function:

[0056] Where q is the approximate posterior and p is the prior (e.g., a standard normal distribution); this mechanism enables the model to output an uncertainty estimation graph. Enhance clinical interpretability.

[0057] TransMorph was the first to successfully apply the Transformer architecture to 3D medical image registration and fusion tasks. Its larger effective receptive field addresses the local limitations of CNNs, and it achieves global and local alignment through a hybrid structure. It proposes a homeomorphic variant for topological order preservation and a Bayesian variant for assessable uncertainty, further expanding the theoretical and practical boundaries. Experimental validation on multiple real-world datasets shows that, compared to VoxelMorph and other Transformer architectures, TransMorph exhibits significant advantages in fusion accuracy (e.g., Dice coefficients), deformation reasonableness, smoothness, and smoothness, and has become a leading method in this field. While TransMorph raises the registration ceiling through long-range modeling, the trade-offs include high parameter / memory and data requirements, complex training and parameter tuning, greater sensitivity to noise and offset, and a lack of explicit cross-modal mechanisms. In small datasets or highly heterogeneous scenarios, additional regularization and modality adaptation are often required.

[0058] NeMAR (Unsupervised Neural Modality-Agnostic Registration) is an unsupervised deep learning method proposed to address the challenge of multimodal image registration and fusion. It fundamentally solves the shortcomings of previous methods in cross-modal cases by introducing an "image-to-image translation" network. In the rigid registration and fusion of CT and MRI images of oral and maxillofacial tumors, the significant differences in grayscale distribution across different modalities pose challenges to directly calculating similarity (whether traditional mutual information or deep learning's L2 / L1). NeMAR innovatively employs a geometry-preserving generative adversarial network: first, an image translation network T is trained to map CT to MRI, ensuring that the output of T preserves the geometry of the input image. Then, a fusion network R is trained to align the translated CT with the original MRI. During this process, accurate single-modal similarity metrics (such as overlap or L2 loss) are used, avoiding the difficulty of directly designing cross-modal similarity functions. This two-step approach enables NeMAR to implicitly align modal differences, decomposing the complex modal alignment problem into a process of "first unifying the modalities, then registering and fusing." Its significant advantages are: it does not require paired fused CT / MRI training samples; during training, only unregistered data of each modality are needed, and modality transformation is learned through a GAN (Generative Adversarial Network); during inference, a "pseudo-MRI" is generated first, and then the deformation is calculated, which is only one more forward pass than a regular network, and can still be completed within a few seconds. NeMAR annotation maintains the high-speed advantage of deep learning methods while solving the core challenge of modality differences, providing a new solution for CT / MRI registration of oral and maxillofacial tumors.

[0059] NeMAR's network architecture is based on the typical U-Net architecture, enabling it to capture both local and global information within an encoder-decoder structure. Specifically, the input is a fixed image. With moving images splicing tensor Multi-scale features are extracted progressively by the encoder, then fused and restored using a decoder and skip connections, ultimately outputting the deformation displacement field. (See [link to documentation]). Figure 7 and Figure 8 The corresponding deformation mapping is:

[0060] Where p represents the coordinates of the original image. This design combines global context and local structure, enabling the model to register overall anatomical differences while maintaining the accuracy of local detail alignment.

[0061] To ensure sufficient information sharing between layers, skip connections are used between the encoder and decoder, and the feature maps output by each layer have multi-scale receptive fields. This method of mixing shallow features and deep semantics enables the output deformation to have both global consistency (such as overall brain alignment) and local finesse (such as brain region boundary alignment).

[0062] NeMAR's loss function design reflects a balance between image similarity and the physical plausibility of deformation: Image similarity loss (unsupervised)

[0063] Alternatively, normalized cross-correlation (NCC) can be used to better suit datasets with significant modal differences. This ensures that, after registration, the moving and stationary plots are highly consistent in grayscale values.

[0064] Smoothing regularization loss:

[0065] The gradient term u is used to constrain the deformation field and avoid unreasonable abrupt changes, thereby ensuring the continuity and smoothness of the registration mapping.

[0066] Total loss function form: Combining the above terms, the total objective function of NeMAR is:

[0067] The parameters λ and γ control the weights between different terms.

[0068] Through the above design, VoxelMorph can simultaneously ensure the accuracy of fusion and the physical rationality of deformation, and further improve the fusion accuracy at key locations through optional structural supervision.

[0069] NeMAR's contribution lies in transforming the traditional pairwise optimization problem of deformable fusion of medical images into a mapping function paradigm of one-time learning plus continuous fast inference, thereby achieving a dual improvement in speed and accuracy. Its U-Net-based structural design balances multi-scale features and is highly adaptable to various medical imaging tasks. The loss function integrates image similarity, deformation smoothness, and optional anatomical structure supervision, resulting in excellent performance in integrity and detail alignment. Although NeMAR can improve alignment on CT / MRI, the "translation + registration" dual-network introduces higher training and parameter tuning complexity. Geometric preservation is highly dependent on the translation subnetwork, and modality retraining is required, with limited transfer to new modalities. Therefore, manual quality control is still necessary for clinical applications to suppress artifacts and mismatch risks.

[0070] Non-rigid registration uses topology-preserving local deformation to finely align anatomical structures that change with time and position, enabling precise identification of tumor and organ displacement and deformation, thus achieving better registration of non-rigid tissues.

[0071] BFM (Segmentation-Registration Combined) was chosen as the representative non-rigid model because it deeply couples structural priors (segmentation subnets) with deformation estimation (registration subnets): in multimodal scenarios, it uses goal-oriented features to smooth the non-rigid deformation field, taking into account both local boundaries and overall topological consistency; it has the advantage of being clinically feasible and easy to integrate with clinical segmentation workflows, serving as a strong baseline and practical alternative for non-rigid multimodal fusion. Its clinical value lies in: the integrated output can be directly used as the initial annotation draft, significantly shortening preoperative waiting time; the structural prior focuses on tumor boundaries and high-risk structures (such as tumors, arteries and veins), reducing HD95 / MSD and physician modification rates while ensuring local fit; its smooth deformation field, constrained by Jacobian regularization, is more conducive to reversibility and topological rationality, facilitating label backpropagation and error tracing; in addition, segmentation guidance can mitigate the uncertainty caused by CT / MRI grayscale differences, is more stable across centers and sequences, and seamlessly integrates with existing clinical segmentation workflows.

[0072] In the field of deep learning, BFM proposes an innovative multi-task dual-fusion framework that integrates anatomical segmentation and image registration tasks into a unified network. The BFM method specifically addresses the limitations of traditional registration algorithms. For example, it no longer relies on manually designed similarity metrics but instead measures image similarity through learned high-level features, effectively mitigating the impact of intensity differences in multimodal images. In multimodal image fusion of oral and maxillofacial tumors, especially in CT and MRI images, due to the consistency of anatomical structures and grayscale differences, introducing an anatomical segmentation sub-network (SegNet) can provide modality-independent structural information, thus significantly improving the accuracy and robustness of cross-modal registration fusion. Secondly, BFM employs end-to-end training, shifting iterative optimization to the training phase. During inference, only one forward propagation is needed to obtain the registration result, significantly improving operational efficiency. Compared to the time-consuming pairwise optimization of ANTs / Elastix, BFM inference can be completed in seconds, showing potential for real-time clinical image fusion. Furthermore, through multi-scale feature bidirectional fusion (fusion of structural features and deformation fields), BFM can simultaneously output anatomical structure segmentation results and registration fusion transformation parameters, completing the segmentation task while maintaining fusion accuracy. These improvements enable BFM to demonstrate advantages in rigid registration fusion tasks for CT / MRI of oral and maxillofacial tumors: by segmenting and locating blood vessels, tumors, and bony structures, BFM can more accurately align the corresponding anatomical locations on CT and MRI, overcoming the problems caused by grayscale differences in traditional methods.

[0073] The BFM network architecture consists of three main parts: a segmentation sub-network (SegNet), a registration sub-network (RegNet), and a multi-task connection module (MTC). The design of this network structure fully considers the potential correlation between the segmentation and registration tasks and promotes the flow of information between them by sharing feature representations.

[0074] SegNet (Segmentation Subnetwork): A network for medical image segmentation, based on the classic U-Net architecture, consisting of an encoder and a decoder. SegNet effectively extracts semantic features from images and generates pixel-level segmentation masks. SegNetr, an improved version of SegNet, performs exceptionally well in medical image segmentation. Through dynamic local-global interactions and information-preserving skip connections, it achieves segmentation performance comparable to state-of-the-art methods while significantly reducing parameters and computational complexity.

[0075] RegNet (Registration Subnetwork) is a deep learning model specifically designed to learn the deformation fields between medical images of different modalities to achieve structural alignment. This model ensures image consistency by predicting the deformation fields between images and is trained using an artificial displacement vector field to improve registration accuracy. RegNet is also based on an encoder-decoder architecture and learns the deformation field by calculating the similarity between images.

[0076] MTC Module (Multi-Task Connection Module): This is a key innovative module of BFM, designed to improve the mutual assistance effect between tasks through feature sharing across tasks. The MTC module includes: (1) Spatial Attention Fusion Module (SAF): enhances local features in the segmentation task through spatial weighting, helping the registration task to better understand structural information; (2) Multi-Scale Spatial Attention Fusion Module (MSAF): fuses spatial attention information at different resolutions to further enhance the multi-scale expression of features; (3) Velocity Field Fusion Module (VFF): provides accurate local deformation information in the registration process by fusing velocity field information, ensuring the topological consistency of the registration results.

[0077] This multi-task connectivity module ensures collaborative optimization between tasks by transferring feature representations between segmentation and registration fusion tasks. Each module is designed to address the issues of insufficient information sharing and unbalanced optimization difficulty that may arise during independent task optimization; see details below. Figure 9 and Figure 10 As shown.

[0078] The training objective of BFM is to simultaneously optimize the performance of segmentation and registration fusion tasks; therefore, its loss function consists of three main parts: segmentation loss ( ), registration loss ( ), and multi-task learning loss ( ).

[0079] Segmentation loss typically employs Dice loss or cross-entropy loss to optimize segmentation accuracy. Dice loss measures the overlap between predicted segments and ground truth labels, and is defined as follows:

[0080] in, and These are the binarized masks for predicted segmentation and ground truth labels, respectively.

[0081] The registration loss consists of an image similarity metric and a deformation field smoothness constraint. The similarity metric often uses mean squared error (MSE) or mutual information (MI) to evaluate the difference between the registered image and the stationary image. The smoothness constraint prevents unnatural deformation by regularizing the gradient of the deformation field and is typically calculated using the following formula:

[0082] in, It is an image similarity measure (such as mutual information or mean square error). It is a smoothness constraint on the deformation field. and It is the weighting coefficient.

[0083] The multi-task learning loss combines the losses from segmentation and registration tasks, aiming to promote complementarity between the two tasks through joint optimization. Its form is:

[0084] This loss function combines the losses of the two tasks, so that the segmentation task can not only optimize its own loss, but also improve the registration accuracy with the guidance of the registration task, while the registration task can achieve better results with the help of the segmentation task.

[0085] The overall objective function of BFM combines segmentation loss, registration loss, and multi-task learning loss, and the specific formula is as follows:

[0086] in, It is registration loss. It is a segmentation loss. It is multi-task learning loss. It is a deformation field. and These are two separate pieces of content. Indicates function composition. and These are the weighting coefficients for the losses of each part.

[0087] BFM innovatively integrates structural and deformation information, employing a multi-scale feature fusion module (MTC) to strengthen the correlation between segmentation and registration fusion tasks at the network architecture level. Similar to the PEMMA framework, it leverages the modularity of the transformer architecture to achieve parameter-efficient model adaptation. This fusion not only improves the performance of medical image registration and segmentation but also minimizes cross-modal entanglement, enhancing the consistency of overall registration and individual registration of bone structures, tumors, and blood vessels in oral and maxillofacial tumor registration. Through joint optimization via multi-task learning, BFM demonstrates excellent performance in experiments across multiple modalities, including MRI, CT, and ultrasound images. This aligns with the achievements of multimodal technology research in medical imaging, which combines various imaging methods such as CT, MRI, and PET to provide more comprehensive information for clinical diagnosis and treatment. Its introduced spatial attention fusion and velocity field fusion modules effectively share information in segmentation and registration fusion tasks, improving their respective accuracy and ensuring the smoothness and topological consistency of the deformation field. The model's advantage lies in the fact that joint training allows segmentation and registration fusion tasks to mutually reinforce each other, significantly improving task execution efficiency and result quality. Although BFM improves performance through the coupling of "segmentation + registration", as a weakly supervised method, it is highly dependent on a large number of high-quality annotations. The large training and parameter tuning of the model is computationally intensive and complex. Inaccurate segmentation subnets can also transmit errors to registration, causing alignment deviations. Therefore, additional quality control and validation are required in data-scarce and clinical applications.

[0088] The fusion quality evaluation indicators include Normalized Mutual Information (NMI), Structural Similarity Index (SSIM), Average Surface Distance (ASD / ASSD), Recall, Precision, and Fusion Index (FI) for tumor regions.

[0089] In the context of CT / MRI multimodal image fusion for oral and maxillofacial tumors, this embodiment demonstrates a multi-objective task of accuracy, efficiency, and robustness. The overall ranking is: NiftyReg < ANTs < Elastix < BFM < VoxelMorph < TransMorph < NeMAR. NeMAR was ultimately chosen as the most suitable clinical implementation solution, achieving the optimal solution in terms of overall consistency (NMI), local critical area alignment (FI), cross-case robustness, and interpretability.

[0090] This invention compares three types of multimodal image fusion models, taking into account seven indicators: NMI, SSIM, ASD, Recall, Precision, FI, and time. The NeMAR model selected in this invention achieves optimal local fusion while ensuring geometric rationality, and also takes into account global similarity, robustness, and efficiency. Therefore, it is the best performing scheme overall and has the greatest potential for clinical application.

[0091] Specifically, the deep learning image segmentation model is the 3D UX-Net model.

[0092] This invention focuses on the automatic segmentation of multimodal images of oral and maxillofacial tumors. It aims to accurately segment tumors and key anatomical structures (jawbone and blood vessels) based on complementary information from CT and MRI, providing a reliable basis for preoperative assessment, virtual surgical design and intraoperative navigation.

[0093] Based on CT / MRI fused images and their corresponding masks (gold standard), data reading and preprocessing were first completed. Then, 3D U-Net, 3D UX-Net, and nnU-Net were used to train image segmentation models and generate segmentation results. Finally, standard evaluation metrics were used for comparative analysis to select the best-performing model or model combination, which was then packaged into a program for clinical application. The technical route is as follows: Figure 11 As shown.

[0094] The study subjects were selected from the same data mentioned above, namely oral and maxillofacial tumor cases treated in the Department of Oral and Maxillofacial Surgery of a certain hospital from January 2016 to August 2025.

[0095] The tumor, maxilla, mandible, and arteries and veins were annotated based on multimodal image fusion results. The "expert-annotated ground truth" refers to the tumor segmentation results based on multimodal CT / MRI image fusion, performed by a licensed oral and maxillofacial surgeon with 7 years of experience in oral and maxillofacial tumor resection and virtual surgical planning. Furthermore, all manually annotated results underwent final review by two senior physicians with 26 and 11 years of experience respectively in oral and maxillofacial tumor surgery before being used for model training.

[0096] To ensure the accuracy of image segmentation results, each image and its corresponding label need to be meticulously processed to ensure a one-to-one correspondence. Medical image processing technology is then used to convert the results to NIfTI (Neuroimaging Informatics Technology Initiative) format for subsequent storage and analysis. This step involves not only format conversion but also the uniform adjustment of image coordinate systems and orientations to ensure consistency across all images in subsequent processing. Furthermore, in medical image processing, to ensure data consistency and accuracy, the fused images need to be resampled to unify voxel spacing and size. Next, to ensure the stability and effectiveness of the subsequent training process, this invention scientifically and rationally divides the entire dataset in a 7:2:1 ratio to form training, validation, and test sets. This division ratio not only provides sufficient sample support for model training but also allows for continuous optimization and adjustment of the model through effective feedback from the validation and test sets, laying a solid foundation for the final training results.

[0097] 3D U-Net is a deep learning network architecture specifically designed for 3D medical image segmentation tasks; it is a 3D extension of the classic U-Net. Unlike traditional 2D segmentation methods, 3D U-Net can directly process medical images containing 3D voxel information, such as CT, MRI, and PET scans. Compared to 2D convolution, 3D convolution can simultaneously capture contextual information in the lateral, longitudinal, and depth directions, which is particularly important for the segmentation of organs, tissues, or lesion regions, as these structures typically have complex morphologies and continuity in 3D space. 3D U-Net inherits the basic idea of ​​U-Net, namely, achieving effective feature fusion from local details to global context through a symmetrical encoder-decoder structure combined with skip connections.

[0098] See Figure 12 From an overall architectural perspective, 3D U-Net can be divided into three main parts: the encoder (contracting path), the decoder (expanding path), and skip connections. During training, 3D U-Net requires a suitable loss function to guide the network's learning. Medical image segmentation often faces class imbalance problems, such as the background region having far more pixels than the target organ or lesion region. Relying solely on cross-entropy loss can easily lead the model to predict the background class, thereby reducing the segmentation accuracy of the foreground region.

[0099] 3D UX-Net, a novel deep convolutional neural network architecture, inherits the symmetric encoder-decoder architecture of the classic U-Net and optimizes feature extraction and fusion by introducing more efficient structures and convolutional modules to adapt to the characteristics of 3D volumetric data. Compared to traditional 2D segmentation networks, 3D UX-Net directly utilizes the contextual relationships between spatial voxels through 3D convolution and 3D upsampling operators, thereby preserving the geometric structure and local details of the target region while extracting global semantic information. This feature enables the network to exhibit good adaptability and accuracy in data segmentation tasks such as CT and MRI.

[0100] See Figure 13 In terms of network structure, 3D UX-Net employs a symmetrical network structure consisting of alternating multi-layer encoders and decoders, and achieves multi-scale information extraction and fusion through multi-stage feature compression and reconstruction. During the encoding stage, the input data first undergoes several layers of 3D convolution and normalization operations to extract low-level edge features and texture information. As the network depth increases, multiple 3D max-pooling operations (MaxPool3d) gradually reduce the spatial resolution of the feature maps and increase the number of channels, thereby obtaining higher-level abstract features. During the decoding stage, deconvolution (ConvTranspose3d) is used to gradually restore the spatial resolution, and the corresponding encoder features are concatenated with the upsampled features along the channel dimension. This feature fusion mechanism preserves the high-resolution detail information from the encoding stage, helping the network obtain more accurate boundary predictions in the segmentation of complex anatomical structures.

[0101] In the feature extraction module design, 3D UX-Net employs a dual-convolutional residual convolutional block (ConvBlock). Each convolutional block contains two consecutive 3D convolutions, instance normalization (InstanceNorm3d), and a LeakyReLU non-linear activation function. This type of convolutional block enhances feature representation while maintaining gradient stability. The network consists of five encoding layers and five decoding layers, with the highest-level feature compression and abstraction performed at the bottom bottleneck layer, thus forming the network's global semantic feature representation.

[0102] In its loss function design, 3D UX-Net considers the class imbalance problem in medical image segmentation. To simultaneously optimize the overlap of prediction results and the classification probability distribution, the network uses a joint form of Dice loss and cross-entropy loss as the optimization objective. Dice loss directly measures the spatial overlap between the prediction and the ground truth label, and is defined as:

[0103] in and These are the binary values ​​of the predicted voxel and the true labeled voxel, respectively. The smoothing factor is used. Cross-entropy loss measures the difference between the predicted probability distribution and the true distribution, and its formula is:

[0104] in For the number of categories, and These represent the true label and the predicted label in the category, respectively. The probability of [something] is [something]. By combining the two, the model can both enhance the segmentation sensitivity of small target structures and maintain the stability of the training process.

[0105] 3D UX-Net is an enhanced version of the classic 3D U-Net encoding and decoding framework. Through a deeper and wider network design and stronger convolutional modules (such as dual convolutional ConvBlock combined with normalization and non-linear activation), it expands channels and layers at the encoding end to capture richer multi-scale anatomical features. At the decoding end, it uses deconvolution and skip connections to progressively restore resolution, achieving more comprehensive multi-scale information fusion. Compared to the traditional 3D U-Net, it has stronger feature representation and non-linear modeling capabilities, thus better characterizing boundaries and fine structures, improving the accuracy of small target recognition and overall segmentation. However, 3D UX-Net also has some shortcomings. Training and inference are slower in resource-constrained scenarios; it relies more heavily on high-quality, large-scale data and is prone to overfitting when data is insufficient; and due to the lack of automated structures, its performance still heavily depends on manual hyperparameter tuning. Furthermore, its insufficient adaptive modeling of multimodal differences makes it prone to performance degradation when modal differences are large.

[0106] See Figure 14 In terms of network design, nnU-Net adopts a symmetrical encoder-decoder structure. The encoder extracts features progressively through multi-layer 3D convolutions (Conv3d), instance normalization (InstanceNorm3d), and LeakyReLU nonlinear activation, and expands the receptive field using multiple downsampling operations. The decoder restores spatial resolution through upsampling and convolution, and fuses shallow and deep features through skip connections. The Generic_UNet module used in the code contains a 5-layer pooling structure, with each stage consisting of two stacked convolutional units, each of which is a combination of 3D convolution-normalization-activation. This design can efficiently capture multi-scale contextual information while preserving edge and detail features, ensuring spatial accuracy of segmentation. The network output layer is a 1×1×1 convolutional kernel used to map multi-channel features to the corresponding class probability distribution.

[0107] nnU-Net is a segmentation framework that achieves "full-process automated adaptation" based on 3D U-Net. It automatically determines key hyperparameters such as network depth, convolutional kernels, input patch size, and batch size according to data characteristics, and combines automatic preprocessing, data augmentation, training, and post-processing strategies to achieve out-of-the-box usability, minimal hyperparameter tuning, and strong generalization on datasets with different modalities, resolutions, and organ sizes. At the same time, it retains the versatility of the encoder-decoder architecture, making it easy to seamlessly extend it with new modules such as attention and Transformer without rebuilding the data processing pipeline. Although nnU-Net reduces hyperparameter tuning costs through automation, its high computational cost and long cycle time due to multi-scale and multi-version training and fusion, as well as its reliance on a large amount of labeled data, limit its segmentation accuracy and practicality in resource-constrained or small-sample scenarios.

[0108] Specifically, on the mandibular bone dataset, NeMAR+3D UX-Net ranked first across all categories, with a Dice coefficient of 0.9743 (the highest among multiple models), an HD95 value of 0.9000 mm (the lowest among multiple models), an MSD value of 1.0829 mm (the lowest among multiple models), a Recall value of 0.9883 (the highest among multiple models), and a Precision value of 0.9609 (the highest among multiple models). These data demonstrate that the mandibular bone boundary closely matches the actual cortical bone, with almost no missed segments and minimal missegmentation. This can significantly reduce the time doctors spend revising the data.

[0109] In summary, this embodiment evaluates the application of 3D U-Net, 3D UX-Net, and nnU-Net in the segmentation of oral and maxillofacial tumors, maxilla and mandible, and arteriovenous images based on seven multimodal image fusion models: Elastix, ANTs, NiftyReg, VoxelMorph, TransMorph, NeMAR, and BFM. It demonstrates that the NeMAR+3D UX-Net selected in this invention performs best overall, achieving a higher Dice and a lower HD95 / MSD.

[0110] Following step S4, the following is also included: Based on the segmented 3D model of the target structure, virtual surgical planning is performed to determine the tumor resection boundary and / or plan the surgical path.

[0111] On the other hand, embodiments of the present invention also disclose a deep learning-based multimodal image fusion and segmentation system for the oral and maxillofacial region, comprising: Data acquisition module: Acquires computed tomography (CT) images and magnetic resonance imaging (MRI) images of the oral and maxillofacial region of the same patient; Preprocessing module: preprocesses the computed tomography (CT) images and magnetic resonance imaging (MRI) images respectively to obtain standardized CT images and MRI images; Image fusion module: Standardized CT and MRI images are input into a pre-trained deep learning image fusion model for registration and fusion to obtain multimodal fused images; Image segmentation module: Input the multimodal fused image into a pre-trained deep learning image segmentation model, segment out at least one three-dimensional model of the target structure and output it.

[0112] Furthermore, it also includes: Decision-making support module: Based on the segmented 3D model of the target structure, perform virtual surgical planning, determine the tumor resection boundary and / or plan the surgical path.

[0113] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0114] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A deep learning-based method for multimodal image fusion and segmentation of the oral and maxillofacial region, characterized in that, Includes the following steps: S1: Obtain computed tomography (CT) and magnetic resonance imaging (MRI) images of the oral and maxillofacial region of the same patient; S2: Preprocess the computed tomography (CT) images and magnetic resonance imaging (MRI) images respectively to obtain standardized CT images and MRI images; S3: Input the standardized CT and MRI images into a pre-trained deep learning image fusion model for registration and fusion to obtain multimodal fused images; S4: Input the multimodal fused image into a pre-trained deep learning image segmentation model to segment out at least one three-dimensional model of the target structure and output it.

2. The method for multimodal image fusion and segmentation of the oral and maxillofacial region based on deep learning according to claim 1, characterized in that, The target structures include at least one of the following: tumor, maxilla, mandible, artery, and vein.

3. The method for multimodal image fusion and segmentation of the oral and maxillofacial region based on deep learning according to claim 1, characterized in that, The preprocessing in step S2 includes: The CT and MRI images are spatially normalized to have the same voxel resolution and spatial size.

4. The method for multimodal image fusion and segmentation of the oral and maxillofacial region based on deep learning according to claim 1, characterized in that, The deep learning image fusion model is an unsupervised NeMAR model.

5. The method for multimodal image fusion and segmentation of the oral and maxillofacial region based on deep learning according to claim 1, characterized in that, Step S3 also includes: After fusing the images using the deep learning image fusion model, a fusion quality evaluation index is calculated. The fusion quality evaluation index includes at least one of Normalized Mutual Information (NMI), Structural Similarity Index (SSIM), Average Surface Distance (ASD), Recall, Precision, and Fusion Index (FI) for tumor regions.

6. The method for multimodal image fusion and segmentation of the oral and maxillofacial region based on deep learning according to claim 1, characterized in that, The deep learning image segmentation model is the 3D UX-Net model.

7. The method for multimodal image fusion and segmentation of the oral and maxillofacial region based on deep learning according to claim 1, characterized in that, Following step S4, the following is also included: Based on the segmented 3D model of the target structure, virtual surgical planning is performed to determine the tumor resection boundary and / or plan the surgical path.

8. A deep learning-based multimodal image fusion and segmentation system for the oral and maxillofacial region, characterized in that, include: Data acquisition module: Acquires computed tomography (CT) images and magnetic resonance imaging (MRI) images of the oral and maxillofacial region of the same patient; Preprocessing module: preprocesses the computed tomography (CT) images and magnetic resonance imaging (MRI) images respectively to obtain standardized CT images and MRI images; Image fusion module: Standardized CT and MRI images are input into a pre-trained deep learning image fusion model for registration and fusion to obtain multimodal fused images; Image segmentation module: Input the multimodal fused image into a pre-trained deep learning image segmentation model, segment out at least one three-dimensional model of the target structure and output it.

9. A deep learning-based multimodal image fusion and segmentation system for the oral and maxillofacial region according to claim 8, characterized in that, Also includes: Decision-making support module: Based on the segmented 3D model of the target structure, perform virtual surgical planning, determine the tumor resection boundary and / or plan the surgical path.