Temporomandibular joint image synthesis system, method, storage medium

CN122798640APending Publication Date: 2026-09-22ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611251788.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-18
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

颞下颌关节影像合成系统、方法、存储介质,以解决现有技术中存在的技术问题:

Benefits of technology

(1)本发明首次在TMJ区域引入条件扩散模型实现跨模态影像合成,有效克服了CBCT无法直接可视化关节盘等软组织的临床模态局限。相较于传统生成对抗网络,本系统结合了感兴趣区域自适应裁剪策略与扩散解析框架,有效抑制了生成过程中的解剖畸变与模式崩溃,使合成影像的峰值信噪比等客观影像质量评价指标得到显著提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122798640A_ABST
    Figure CN122798640A_ABST
Patent Text Reader

Abstract

The application discloses a temporomandibular joint image synthesis system, method and storage medium, belongs to the technical field of medical devices, and aims to solve the problems of weak detection capability of an existing cone beam CT in oral diagnosis, subjectivity in inferring whether a disc is displaced through indirect changes of a joint gap, low accuracy and the like. Technical solution points are as follows: the system comprises a cross-modal image collaborative preprocessing module, a three-dimensional rigid registration operator is used to realize anatomical space alignment, and a self-adaptive region of interest extraction algorithm is used to uniformly resample images to a target resolution voxel space; a multi-dimensional feature extraction and coding module comprises an image encoder and a semantic encoder; a diffusion generation core module is composed of a denoising autoencoder and a conditional control branch; and a space-time continuity guiding module introduces a sliding window mechanism to take continuous multiple frames of source modal image slices as input constraints, and ensures the spatial smoothness of a synthesized image through multi-frame feature fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical device technology, specifically relating to a temporomandibular joint image synthesis system, method, and storage medium. Background Technology

[0002] Temporomandibular joint disorders (TMDs), as one of the most common chronic diseases of the oral and maxillofacial region, have a profound impact on human mastication, speech, and facial aesthetics. Their pathological features typically involve changes in the anatomical relationships between the articular disc, condyle, and glenoid fossa, with articular disc displacement (ADD) being the most representative pathological type in clinical practice. Epidemiological studies show that approximately 35% of ADD patients are asymptomatic in the early stages, leading to delays in intervention and subsequently causing irreversible jaw deformities or severe joint dysfunction.

[0003] In modern medical imaging diagnostics, magnetic resonance imaging (MRI), with its superior soft tissue resolution, can clearly visualize the fine structures of soft tissues such as articular discs, synovium, and ligaments, and is widely recognized as the "gold standard" for assessing the location and internal anatomy of the temporomandibular joint (TMJ). MRI can precisely capture the millimeter-level morphology and thickness of the articular disc and its positional relationship with the condyle and glenoid fossa. It can also clearly detect early pathological manifestations such as joint effusion, synovial inflammation, and disc tears, providing reliable imaging evidence for the classification, assessment, and treatment planning of atopic dermatitis (ADD). However, the widespread application of MRI in oral clinical practice faces several practical challenges: First, MRI equipment is expensive to purchase, with a single unit typically costing several million yuan, and the daily maintenance costs are also considerable. As a result, it is usually only equipped in large tertiary hospitals and specialized dental hospitals, making it difficult for primary dental clinics and community health institutions to afford. Secondly, MRI acquisition cycles are relatively long, with a single scan typically taking 10-20 minutes. Patients are required to keep their heads still and mouths closed during the scan, which is extremely challenging for TMD patients with limited mouth opening, joint pain, or anxiety. Insufficient patient cooperation can lead to blurred images and affect diagnostic accuracy. In addition, MRI is unevenly distributed in the population, especially in remote areas and primary healthcare settings where coverage is extremely low, which further limits its feasibility as a routine early screening method for TMDs and makes it difficult to achieve early detection and early intervention of the disease.

[0004] Compared to MRI, cone-beam computed tomography (CBCT) has become widely used in oral diagnosis due to its unique technological advantages, becoming a commonly used imaging method in oral clinical practice. CBCT uses a cone-shaped X-ray beam for scanning, offering core advantages such as high spatial resolution (up to 0.1 mm), relatively low radiation dose (only 1 / 10 to 1 / 5 of traditional CT), fast imaging speed (only a few seconds per scan), and ease of operation. It excels in assessing bone tissue lesions such as condylar bone erosion, osteophyte formation, subchondral bone sclerosis, and joint space narrowing, providing accurate imaging diagnostic evidence for TMD-related bone tissue abnormalities.

[0005] However, due to the limitations of CBCT's physical imaging principles, its ability to detect soft tissues (such as articular discs, synovium, and ligaments composed of collagen fibers) is extremely weak. It cannot directly display the morphology and position of the articular disc. Clinicians often have to rely on indirect changes in the joint space (such as widening of the anterior joint space and narrowing of the posterior joint space) and their own clinical experience to infer whether the joint disc is displaced and the degree of displacement. This empirical diagnostic method is highly subjective, and its diagnostic accuracy is greatly affected by factors such as the doctor's clinical experience and differences in judgment criteria. It often fails to meet the needs of early and accurate diagnosis and treatment of TMDs, and is prone to missed diagnoses, misdiagnoses, and delays in the best intervention time.

[0006] In recent years, deep learning technology has made significant progress in medical image synthesis, diagnosis, and analysis due to its powerful feature extraction and image generation capabilities, providing a new technical approach to address the pain points in TMD (tumor lesioning) image diagnosis. Among these, Generative Adversarial Networks (GANs), as an early mainstream generative model, have been attempted to be applied to cross-modal synthesis of medical images, aiming to achieve conversion between different image modalities. However, GANs often face several technical bottlenecks when processing small and low-contrast anatomical structures like TMJ (tumor lesioning): instability easily occurs during model training, resulting in blurred structural details in the generated images; mode collapse is prone to occur, meaning the generated images lack diversity and cannot cover individual anatomical differences among patients; simultaneously, the generated images often lack physical realism and have insufficient consistency with the anatomical structures of real MRI images, making it difficult to meet the precise requirements of clinical diagnosis.

[0007] Diffusion models, as an emerging paradigm in generative artificial intelligence, have gradually replaced GANs as a research hotspot in medical image synthesis since their inception due to their unique technical advantages. This model simulates the stochastic dynamics of noise addition and denoising, progressively adding Gaussian noise to the original clear image until it is completely blurred, and then reconstructing the clear image through a reverse denoising process. The generated images have higher structural fidelity and perceptual similarity, better preserving subtle anatomical features and effectively avoiding problems such as mode collapse and training instability inherent in GANs. In the field of medical imaging, diffusion models have been successfully applied to tasks such as cross-modal synthesis, image restoration, and lesion segmentation in CT and MRI, demonstrating enormous clinical application potential.

[0008] Relevant patent documents retrieved: The patent, published in China with publication number CN119941889A and dated May 6, 2025, discloses a system and method for reconstructing pseudo-MRI images using CBCT based on artificial intelligence. The system includes a data acquisition and preprocessing module, a pseudo-MRI image generation module, and a central control module. By introducing deep learning technology, the system analyzes the preprocessed cone-beam computed tomography (CBCT) image data to generate pseudo-MRI images with high resolution and high contrast. The deep learning technology employed is the diffusion model MC. Any one of IDDPM, Boundary Condition Diffusion Model, and Conditional Generative Adversarial Network.

[0009] The prior art represented by the aforementioned documents has at least the following unresolved technical problems or defects: Problems such as unstable model training, mode collapse, and lack of physical realism in generated images are addressed by evidence that existing technology CN119941889A employs three models: the diffusion model MC. IDDPM, boundary conditional diffusion models, and conditional generative adversarial networks (GANs) are all used. However, for small anatomical structures with extremely low contrast, such as the temporomandibular joint, the training of GANs often suffers from instability and collapse. While diffusion models offer a certain degree of fidelity and perceptual similarity, existing technology CN119941889A does not provide a specific implementation plan for combining diffusion models with CBCT images.

[0010] Addressing the challenges of TMD imaging diagnosis, combining diffusion models with the specific anatomical features of TMJs to achieve accurate translation from CBCT to high-quality MRI has become a hot topic and a difficult point in current precision oral medicine research. On the one hand, the anatomical structure of TMJs is complex and delicate, with structures such as the articular disc and condyle being tiny, and significant individual differences in anatomical morphology among different patients. Diffusion models need to accurately capture these specific features to ensure that the generated MRI images can accurately reproduce the position, shape, and anatomical relationship of the articular disc with surrounding structures. On the other hand, it is necessary to solve the domain difference problem in cross-modal synthesis, namely, the imaging mechanisms and grayscale distributions of CBCT (mainly bone tissue) and MRI (mainly soft tissue) are significantly different. Model optimization is needed to achieve accurate mapping between the two modalities, so that the generated MRI images are not only visually consistent with real MRI, but also meet the clinical diagnostic needs for assessing key information such as the position of the articular disc and pathological lesions. Successfully translating CBCT into MRI will fully leverage the widespread availability of CBCT and the soft tissue diagnostic advantages of MRI, enabling primary dental clinics to obtain precise articular disc location information through routine CBCT examinations. This will provide strong support for the early screening, accurate diagnosis, and individualized treatment of TMDs, and promote the development of precision dentistry.

[0011] In solving the above problems or overcoming the above defects, the present invention encountered the following difficulties and obstacles: Traditional image generation networks are prone to detail blurring and training instability when dealing with the unique and minute anatomical structures of the temporomandibular joint (TMJ). When further introducing diffusion models for cross-modal (CBCT to MRI) generation, the significant differences in physical imaging mechanisms between the two modalities make it difficult for the base model to simultaneously achieve high-precision reconstruction of minute soft tissues (such as articular discs) and accurate alignment of three-dimensional spatial structures. Summary of the Invention

[0012] The purpose of this invention is to provide: A temporomandibular joint image synthesis system, method, and storage medium are provided to address the technical problems existing in the prior art. Cone-beam CT has extremely weak detection capabilities for soft tissues and cannot directly visualize soft tissue structures such as articular discs. Clinicians can only make empirical diagnoses based on indirect changes in the joint space, and the accuracy is insufficient to meet the needs of early and precise treatment. Magnetic resonance imaging (MRI) equipment has extremely high purchase and maintenance costs, long acquisition cycles, and uneven distribution per unit population, which limits its feasibility as a routine screening method. Existing generative adversarial networks often face problems such as unstable model training, pattern collapse, and a lack of physical realism in generated images when dealing with small and low-contrast anatomical structures such as the temporomandibular joint, making it impossible to achieve accurate translation from cone-beam CT to high-quality MRI.

[0013] Terminology Explanation: Unless otherwise defined, all technical terms herein have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains.

[0014] It should be understood that the above brief description and the following detailed description are exemplary and for illustrative purposes only, and do not limit the subject matter of the invention in any way. In this invention, the singular is used in conjunction with the plural unless otherwise specifically stated. It should also be noted that, unless otherwise stated, the use of “or” or “or” means “and / or”. Furthermore, the use of the term “comprising” and other forms such as “including,” “containing,” and “contains” are not limiting.

[0015] Unless specifically defined herein, the use of various commercially available products herein employs standard techniques. For example, it may be implemented using the manufacturer's instructions for use, or in accordance with methods known in the art or the description of this invention. The techniques and methods described herein can generally be implemented according to conventional methods well known in the art, based on the descriptions in the various general and more specific documents cited and discussed in this specification.

[0016] The terms “optional / arbitrary” or “optionally / arbitrarily” mean that the event or situation described below may or may not occur, including both the occurrence and non-occurrence of the event or situation.

[0017] The term "cone-beam computed tomography" used in this article refers to a computed tomography imaging device that uses a cone-shaped X-ray beam and an area array detector to perform a single rotational scan around the target area to reconstruct a three-dimensional tomographic image. It is commonly used for rapid three-dimensional imaging of small-area anatomical structures such as the oral and maxillofacial region and the temporomandibular joint.

[0018] The term “three-dimensional rigid registration operator” used in this article refers to a three-dimensional spatial coordinate mapping calculation operator that only includes translation and rotation transformations and does not cause deformation. It is used to align two sets of three-dimensional images as a whole, maintaining the relative size and shape of the anatomical structure throughout the process.

[0019] The term "source modal image" used in this paper refers to the original modal medical image that serves as the input benchmark and provides prior information on anatomical structures in cross-modal image conversion and registration tasks. In this scheme, it mainly refers to CBCT images.

[0020] The term “target modal contrast” used in this article refers to the tissue grayscale differentiation features of the target imaging modality (MRI) to be generated and aligned, including the grayscale difference and boundary differentiation between different anatomical tissues such as bone, soft tissue, and joint cavity.

[0021] The term "anatomical constraint" used in this article refers to the model loss term and transformation constraints set based on the inherent spatial position, shape, and size relationship of the human physiological and anatomical structure, to prevent results that violate physiological structure, such as bone misalignment and tissue deformation, from occurring during the registration / generation process.

[0022] The term "sliding window mechanism" used in this paper refers to a block-based processing strategy that uses a fixed-size step size to slide and capture image patches on a 3D / 2D image, performing feature extraction, noise prediction, or registration calculations in each patch, and finally stitching them together to obtain a complete output image. This strategy is used to reduce the computational memory overhead of large-volume images.

[0023] The term “Markov chain forward noise addition and reverse noise reduction” used in this article refers to: the core calculation process of the diffusion model; the forward process gradually adds Gaussian noise to the clean image along the Markov chain until a pure noise map is obtained; the reverse process gradually removes noise through a neural network to restore a clear target modal medical image.

[0024] The term "conditional control network branch" used in this paper refers to an independent branch network set up outside the main network of the diffusion model. It is used to extract the anatomical features of the source modality CBCT and combine them with high-level semantic priors as conditional information to guide the inverse denoising process to generate MRI images that conform to the anatomical structure.

[0025] The term "continuous three-layer CBCT" used in this article refers to three consecutive adjacent cone-beam CT slices along the scanning thickness direction, forming a three-dimensional local image block, which is used for model learning of inter-layer anatomical continuity features and improving the consistency of three-dimensional structural reconstruction.

[0026] The term "high-level semantic prior" used in this paper refers to the global anatomical semantic features extracted by the pre-trained visual encoder, which include global structural information such as skeletal contours, joint positions, and soft tissue distribution, and serve as prior conditions to constrain image generation.

[0027] The term "conditional control branch locked diffusion model" used in this paper refers to a lightweight training model architecture in which the weights of the backbone diffusion inverse denoising network are trained with fixed values, and only the branch parameters of the conditional control network are updated. It reuses mature denoising capabilities and only optimizes cross-modal conditional mapping relationships to reduce training costs.

[0028] The term "Elastix toolkit" used in this article refers to an open-source medical image registration algorithm toolkit that integrates multiple registration transformation operators such as rigid, affine, and elastic models for the standardized alignment of 3D medical images, serving as a benchmark tool for comparing the registration effects of this scheme.

[0029] The term "CBCT" used in this article refers to Cone Beam Computed Tomography, which is a type of computed tomography imaging.

[0030] The term "MRI" used in this article refers to Magnetic Resonance Imaging, which relies on magnetic fields and radio frequency pulses to acquire high-resolution images of soft tissues, serving as the source modality input image for this scheme.

[0031] The term "TransTMJ-Net" used in this paper refers to the Transformer backbone neural network designed for cross-modal image transformation of the temporomandibular joint (TMJ), which integrates multi-scale feature extraction, anatomical constraint modules, and conditional diffusion control branches.

[0032] The term "CLIP encoder" used in this paper refers to a visual feature encoder based on contrastive language-image pre-training, which can extract general global high-level semantic features of images and provide anatomical semantic priors for this model.

[0033] The term "ImageNet" as used in this article refers to a publicly available natural image classification dataset with a scale of millions of images, used to pretrain various visual backbone networks and provide general low-level visual feature extraction capabilities for medical imaging tasks.

[0034] The term "epoch" used in this article refers to a training cycle in which a neural network completely traverses the entire training dataset. Each completed epoch represents a round of forward propagation and backward parameter update for all samples.

[0035] The term "CycleGAN" used in this paper refers to Cyclic Consistency Generative Adversarial Network, which achieves cross-modal image transformation without pairing data by relying on bidirectional generation and cyclic loss, and serves as the baseline model for comparison with this scheme.

[0036] The term "GAN" used in this article refers to Generative Adversarial Network, a basic deep learning framework that uses a generator and a discriminator trained adversarially to achieve image generation and modality transformation.

[0037] The term “Pix2Pix” used in this paper refers to a conditional generative adversarial network trained under paired image supervision, which uses paired source-target modal images to complete one-to-one image conversion.

[0038] The term "U-Net generator" used in this paper refers to a U-shaped network with encoder downsampling, decoder upsampling, and skip connection structure as a GAN model generator, which is good at preserving fine-grained anatomical boundary information in medical images.

[0039] The term "PatchGAN discriminator" used in this paper refers to a discriminative network that does not output global true / false labels but only distinguishes between true and false images of local image patches, focusing on constraining the quality of local texture and contrast detail generation. In a first aspect, this invention provides: Temporomandibular joint image synthesis system, including: A cross-modal image collaborative preprocessing module is used to perform spatial mapping between source and target modal images. This module utilizes a three-dimensional rigid registration operator, using preset anatomical landmarks as anchor points to achieve anatomical spatial alignment, and employs an adaptive region of interest extraction algorithm to uniformly resample the source and target modal images to the target resolution voxel space. The cross-modal image collaborative preprocessing module is connected to a multi-dimensional feature extraction and encoding module via data connectivity. The multidimensional feature extraction and encoding module includes an image encoder and a semantic encoder. The image encoder is used to project source modality image slices onto a latent feature space. The semantic encoder is based on a contrastive language-image pre-trained model to extract global semantic feature vectors related to anatomical sites. The multidimensional feature extraction and encoding module is connected to the diffusion generation core module through data connection. The diffusion generation core module, based on a diffusion model architecture, consists of a denoising autoencoder and a conditional control branch. The conditional control branch adjusts the denoising trajectory in real time through residual connections, so that the generated image maintains the contrast of the target modality while following the anatomical constraints provided by the source modality image. The diffusion generation core module is connected to the spatiotemporal continuity guidance module through a data connection. The spatiotemporal continuity guidance module introduces a sliding window mechanism during the generation process, using continuous multi-frame source modal image slices as input constraints, and ensuring the spatial smoothness of the synthesized target modal image at the longitudinal anatomical level through multi-frame feature fusion.

[0040] Through the above technical solution, this invention, for the first time, combines a diffusion model with a conditional control branch in the field of cross-modal medical image generation. It accurately captures the anatomical constraints of bone tissue in the source modality image through zero-convolutional layers and fuses global visual semantic priors extracted from a contrastive language-image pre-training model through a cross-attention mechanism. During the iterative process of forward noise addition and reverse denoising in the Markov chain, it achieves a nonlinear mapping from Gaussian noise to the soft tissue features of the target modality. Compared to traditional generative adversarial networks, this invention effectively overcomes problems such as unstable model training, mode collapse, and lack of physical realism in generated images. Simultaneously, through a three-dimensional rigid registration algorithm with multiple anatomical landmarks, it eliminates the differences in patient positioning when acquired from different devices, providing a high-quality paired dataset for model training. Furthermore, through a continuous multi-frame sliding window spatiotemporal continuity guidance mechanism and adaptive cropping of the region of interest and voxel normalization preprocessing strategies, this invention ensures the anatomical coherence and topological consistency of the synthesized image in three-dimensional space, while eliminating non-target background interference, significantly improving generation quality and training efficiency.

[0041] The essential difference between the conditional diffusion generation model used in this invention and the conventional general diffusion model is that the conventional model is dominated by "unconstrained random probability generation", which has a high degree of spatial uncontrollability and feature divergence in the image reconstruction process, and is very easy to produce non-physiological artifacts (i.e. feature fabrication) in small anatomical areas; while the model of this invention constructs a "spatial anchoring cross-modal translation with strong physical constraints" architecture.

[0042] The specific technical implementation and effects are as follows: Rigid Topological Constraints and Spatial Anchoring: This invention innovatively introduces a Conditional Control Network branch (ControlNet network branch), injecting the local spatiotemporal features of three consecutive CBCT layers as an insurmountable rigid topological boundary into the generative network. This mechanism forces the network to strictly use the three-dimensional spatial positions of existing bony structures (condyles, glenoid fossae, etc.) as anchor points when generating MRI soft tissue (such as articular discs) from Gaussian noise through inverse denoising, achieving perfectly pixel-level anatomical alignment.

[0043] Eliminating structural drift and artifact fabrication: Through the synergistic constraint of the above dual conditions (global visual semantics and local continuous spatial features), the diffusion model of this invention completely eliminates the inherent defects of "anatomical drift" and "detail fabrication" in general generative large models, ensuring that every pixel in the cross-modal mapping process (from CBCT to synthetic MRI) has a solid physical basis, thereby meeting the clinical needs for sub-millimeter-level precision in auxiliary diagnosis.

[0044] Based on further solutions to the technical problems of the present invention, or simultaneous solutions to multiple technical problems, the preferred solution in the technical solution provided in the first aspect of the present invention includes: First preferred option: The training logic of the diffusion generation core module follows the forward noise addition and reverse noise reduction principle of Markov chains. In the reverse noise reduction stage, the system uses a cross-attention mechanism to fuse the semantic feature vector with the latent features of the image, guiding the noise predictor to perform dissecting feature correction in each time step.

[0045] Through this mechanism, the present invention introduces high-level semantic priors during the generation process, which enhances the morphological rationality and positional accuracy of the synthesized images on key anatomical structures such as articular discs and condyles, and solves the problem of inconsistency in anatomical topology logic that may be caused by simply generating based on pixel-level features.

[0046] Second preferred option: The diffusion generation core module uses Markov chain noise addition and inverse Gaussian distribution denoising algorithms to iteratively train the model, resulting in the optimized TransTMJ-Net conditional diffusion generation model.

[0047] Markov chain noise addition: to achieve progressive hierarchical noise addition, generate multi-gradient noisy image training samples, and construct a stable positive diffusion link from clear to pure noise; Inverse Gaussian distribution denoising: accurately fits the multi-layer cumulative noise distribution, calculates the denoising loss to guide the network iterative optimization, and reversely realizes noise stripping and restores complete standard medical images while preserving subtle anatomical features.

[0048] Furthermore, the Markov chain noise addition specifically includes: In the forward diffusion process, the generation process of the noisy image is defined as a Markov chain, based on a pre-defined variance table. Noise that follows a standard normal distribution is gradually added to the input data. (0, I); Noisy image at step n Depends on the image from the previous step ,satisfy:

[0049] Therefore, the noise addition formula for step n is:

[0050] in, This represents the original image without added noise.

[0051] Furthermore, the inverse Gaussian distribution denoising algorithm is as follows: In the reverse denoising process, the diffusion generation core module recursively calculates an inverse Gaussian distribution from... Reverse derivation to target soft tissue image Furthermore, a reverse diffusion process is introduced during the reverse denoising process:

[0052] in, Characterizes the noise image given at the current nth step. and cross-modal conditional input At that time, the network predicts a clearer image from the previous step (n-1 step). The conditional probability distribution, This represents the Gaussian (normal) distribution function. This indicates that the model was generated by TransTMJ-Net (with parameters). The predicted Gaussian distribution mean is essentially the network's prediction based on the current noise level. and CBCT bony boundary constraints The calculated "denoising direction and soft tissue structure characteristics" The variance (covariance matrix) of the predicted Gaussian distribution represents the confidence and uncertainty of the current step's denoising generation.

[0053] This technical solution addresses the technical issues of "the weak detection capability of existing cone-beam CT in oral diagnosis, and the subjectivity and low accuracy of inferring whether the joint disc has shifted through indirect changes in the joint space" and further solves the technical problem of "how to establish a stable cross-modal image translation mechanism with high anatomical fidelity to suppress mode collapse and non-physiological artifacts in the generation process of small, low-contrast tissues".

[0054] Third preferred option: During the generation of synthetic MRI images, the basic parameters of the conditional control branch locked diffusion model (the original StableDiffusion model) are trained only on additional copies of the neural network layers introduced to process the conditional inputs.

[0055] Specifically, the diffusion model serves as the underlying pre-training foundation of this invention, providing general visual prior knowledge for high-fidelity image denoising and reconstruction. During the generation of synthetic MRI images, the ControlNet branch of this invention strictly locks (freezes) the basic network parameters of the original Stable Diffusion model, and only performs specialized fine-tuning training on additional neural network layer copies introduced to handle specific conditional inputs.

[0056] This technical solution addresses the technical problems of "the weak detection capability of existing cone-beam CT in oral diagnosis, and the subjectivity and low accuracy of inferring whether the joint disc has shifted through indirect changes in the joint space" and further solves the technical problems of "how to ensure the training stability of the model when processing limited medical image data, and how to reduce the number of parameters that need to be updated".

[0057] Fourth preferred option: When the cross-modal image collaborative preprocessing module resamples to the target resolution voxel space, it follows the mapping rules below: Before training and inference, DICOM format images with spatial resolution matching are acquired; for MRI images, their voxel intensity is first truncated to the range of [0, 3500], and then normalized to the interval of [-1, 1]; simultaneously, for cone-beam CT images, their voxel intensity is normalized from the original range of [0, 1023] to the interval of [-1, 1].

[0058] Fifth preferred option: The three-dimensional rigid registration operator employs a multi-level optimization strategy based on the mutual information similarity criterion: First, the Elastix toolkit was used to perform voxel-level rigid registration of CBCT and MRI sequences; Secondly, precise three-dimensional orientation matching is performed based on multiple anatomical landmarks in the axial, sagittal, and coronal planes; through manual fine-tuning, the anteroposterior direction of the CBCT image is aligned with the superior-inferior direction of the MRI image until the mutual information between the two sets of images reaches its maximum value.

[0059] By employing this voxel standardization strategy, this invention addresses the problem of model training instability caused by differences in grayscale value distribution among images of different modalities, accelerates model convergence, and reduces the impact of imperfect registration.

[0060] Furthermore, several anatomical landmarks include: the center of the left condyle, the center of the right condyle, the left zygomaticotemporal suture, and the right zygomaticotemporal suture.

[0061] Sixth preferred option: The region of interest is the area centered on the core anatomical structures of the temporomandibular joint with a resolution of 256×256.

[0062] Furthermore, the region of interest is resampled to a 512×512 image and used as input data for training the conditional diffusion generation module.

[0063] This technical solution addresses the problems of "weak detection capability of existing cone-beam CT in oral diagnosis, and subjectivity and low accuracy in inferring whether the joint disc has shifted through indirect changes in the joint space" and further solves the problem of "how to eliminate interference from non-target brain tissue and improve spatial registration error tolerance".

[0064] Seventh preferred option: The spatiotemporal continuity guidance module uses continuous multi-frame source modal image slices as input constraints, and the intermediate layer slices are input to the semantic encoder to extract global semantic features. The spatial features of the multi-frame slices are fused through the conditional control branch.

[0065] Through this spatiotemporal continuity guidance mechanism, the present invention solves the problem of anatomical discontinuity between adjacent slices that may be caused by independent generation layer by layer, ensures the anatomical coherence of the synthesized target modal image sequence in three-dimensional space, and avoids morphological abrupt changes or positional jumps in soft tissues such as articular discs between adjacent slices.

[0066] Furthermore, the spatiotemporal continuity guidance module uses three consecutive source modal image slices as input constraints.

[0067] After acquiring volumetric data by performing a 3D scan of the temporomandibular joint (TMJ) using a CBCT device, the system digitally cuts the tissue at equal intervals along a specific anatomical axis (such as the sagittal plane) according to a preset slice thickness, generating a sequence of two-dimensional tomographic slices with a strict spatial topological order. For the target slice (defined as the intermediate layer, i.e., layer z), the module simultaneously extracts its spatially adjacent upper layer (layer z-1) and lower layer (layer z+1), and these three layers together constitute the joint input unit for feature extraction.

[0068] The advantages of using a "continuous three-layer" technology: Capturing 3D anatomical trends: This compensates for the lack of depth information in single-layer 2D slices, enabling diffusion models to effectively extract the morphological gradient trends of complex bony structures such as condyles in the axial direction.

[0069] Strengthening spatial boundary constraints: The upper and lower slices provide physical topological constraints for the reverse generation of soft tissues (such as articular discs) in the middle layer, effectively suppressing anatomical distortions and artifacts that are prone to occur in diffusion models, and ensuring the realism and coherence of synthetic MRI images in three-dimensional space.

[0070] Balancing computational power and generation performance: It avoids computational overload and model training crashes caused by directly processing the full amount of 3D data, while giving the model the ability to understand 3D context with extremely low additional computational overhead.

[0071] Secondly, the present invention provides: A method for synthesizing temporomandibular joint images, executed on the aforementioned temporomandibular joint image synthesis system, includes: S1. The cross-modal image collaborative preprocessing module obtains the source modal image sequence, performs rigid registration based on grayscale information, and extracts local voxel blocks with key anatomical structures as the core. S2. The multidimensional feature extraction and encoding module extracts spatial constraint features from the preprocessed source modal image input condition control branch. At the same time, it extracts the center slices of multiple consecutive frames of source modal images and inputs them into the image encoder of the vision-language pre-trained model to directly extract high-dimensional implicit semantic embeddings. S3. The diffusion generation core module starts from Gaussian noise in the diffusion model, and under the dual guidance of the spatial constraint features and the semantic embedding, iteratively predicts and removes noise to restore the potential features of the target modal image. S4. The diffusion generation core module decodes the latent features into pixel-level images using a pre-trained decoder to generate a synthetic target modal image.

[0072] This method corresponds one-to-one with the technical solution of the aforementioned system, achieving high-fidelity cross-modal generation from source modal images to target modal images. This image-based strategy for direct extraction of implicit semantic features avoids the subjectivity and inefficiency of manual annotation, while ensuring the accurate correspondence between semantic embedding vectors and anatomical sites.

[0073] Furthermore, the local voxel block extracted in step S1 has a resolution of 256×256×N. The rigid registration is performed by maximizing the mutual information value between the source modal image and the target modal image until the mutual information value reaches a preset threshold.

[0074] This specific parameter setting ensures the coverage and registration accuracy of the region of interest, providing high-quality input data for subsequent feature extraction and image generation.

[0075] Preferably, the number of denoising steps in the inference stage of step S3 is set to 10 to 30, and the anatomical features are corrected using a cross-attention mechanism within each time step.

[0076] Thirdly, the present invention provides: A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described cross-modal medical image synthesis method.

[0077] The storage medium may be a non-volatile storage medium, including but not limited to at least one of read-only memory, random access memory, disk, optical disk, USB flash drive, and solid-state drive.

[0078] The present invention has at least the following beneficial effects: (1) This invention introduces a conditional diffusion model in the TMJ region for the first time to achieve cross-modal image synthesis, effectively overcoming the limitation of CBCT in directly visualizing soft tissues such as joint discs. Compared with traditional generative adversarial networks, this system combines an adaptive cropping strategy for regions of interest with a diffusion parsing framework, effectively suppressing anatomical distortion and modality collapse during the generation process, and significantly improving objective image quality evaluation indicators such as peak signal-to-noise ratio of the synthesized images.

[0079] (2) This invention innovatively introduces a three-dimensional rigid registration algorithm with multiple anatomical landmarks. It employs a multi-level optimization strategy based on mutual information similarity criteria. First, a registration toolkit is used to perform voxel-level rigid registration between the source and target modal images. Then, manual fine-tuning is performed based on multiple anatomical landmarks in the axial, sagittal, and coronal planes until the mutual information between the source and target modal images reaches its maximum value. This registration strategy eliminates the patient's positional differences when acquired using different devices, achieving precise three-dimensional spatial alignment between the source and target modal images. This provides a high-quality paired dataset for model training and significantly improves the anatomical consistency of the synthesized images.

[0080] (3) The synthetic MRI generated by this invention not only achieves high-fidelity reconstruction of the structure but also highly preserves pathological details, resulting in a significant leap in the accuracy of radiologists' auxiliary diagnosis of anterior disc displacement (ADD). Furthermore, the system's built-in automated diagnostic classifier also demonstrates excellent diagnostic performance for ADD, approaching the gold standard performance of real MRI images. Therefore, this invention provides a low-cost, high-precision auxiliary screening solution for TMJ disorder in medical institutions with limited MRI resources without requiring additional expensive equipment investment, possessing extremely high clinical translation and application value.

[0081] (4) This invention extracts global semantic feature vectors related to anatomical sites by contrasting a language-image pre-training model, and fuses semantic features with latent image features through a cross-attention mechanism, guiding the diffusion model to correct anatomical features at each time step. This semantic prior guidance mechanism solves the problem of inconsistency in anatomical topology logic that may be caused by generating images solely based on pixel-level features, enhances the morphological rationality and positional accuracy of synthetic images on key anatomical structures such as articular discs and condyles, and improves the clinical usability of the generated images.

[0082] (5) This invention uses a continuous multi-frame sliding window spatiotemporal continuity guidance mechanism to take continuous multi-frame source modal image slices as input constraints, and intermediate layer slices as input semantic encoders to extract global semantic features. The spatial features of multi-frame slices are fused through conditional control branches. This mechanism solves the problem of anatomical discontinuity between adjacent slices that may be caused by independent generation layer by layer, ensures the anatomical coherence of the synthesized target modal image sequence in three-dimensional space, avoids morphological abrupt changes or positional jumps of soft tissues such as articular discs between adjacent slices, and improves the three-dimensional consistency of the synthesized image.

[0083] (6) This invention employs an adaptive region-of-interest (ROI) cropping and voxel normalization preprocessing strategy. It crops an ROI of a preset resolution centered on key anatomical structures and resamples it to the target resolution. Simultaneously, it truncates and normalizes the voxel intensities of the target and source modal images to a preset normalization interval, eliminating slices lacking core anatomical structures at the beginning and end of the sequence. This preprocessing strategy solves the model training instability problems caused by non-target background interference and inconsistent voxel intensity distribution. ROI cropping eliminates redundant information, voxel normalization accelerates model convergence, improving training efficiency and generation quality, while reducing the impact of imperfect registration.

[0084] (7) The synthetic magnetic resonance imaging (sMRI) generated by this invention exhibits the following significant technical effects and advantages in improving the auxiliary diagnostic efficiency of temporomandibular joint (TMJ) disc displacement: Traditional pain points: When relying solely on traditional CBCT images, clinicians find it difficult to accurately identify the location of the joint disc due to the lack of soft tissue details. The diagnostic accuracy is only around 50% (close to random guessing), and the sensitivity is extremely low.

[0085] When using the synthetic MRI generated by this invention to assist in diagnosis, the overall diagnostic accuracy rate significantly increased from 50.3% to 86.3% to 89.0%.

[0086] Efficacy in targeting key pathologies: Particularly for the most common clinical condition, anterior disc displacement (ADD), the diagnostic accuracy rate reached 91.7% with the aid of synthetic MRI, successfully identifying the vast majority of displacement cases. This demonstrates that the images generated by this invention provide crucial soft tissue visual details, greatly enhancing physicians' ability to identify diseases. Attached Figure Description

[0087] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0088] Figure 1 This is a structural block diagram of the temporomandibular joint image synthesis system of the present invention; Figure 2 A schematic diagram of the preprocessing for precise 3D registration of cross-modal images; Figure 3 Diagram of the core diffusion network architecture of TransTMJ-Net; Figure 4 This is a diagram showing the comparison of cross-modal generation effects. Detailed Implementation

[0089] The following non-limiting embodiments are intended to enable those skilled in the art to gain a more comprehensive understanding of the present invention, but do not limit the invention in any way. The following content is merely an exemplary description of the scope of protection claimed by the present invention, and those skilled in the art can make various changes and modifications to the present invention based on the disclosed content, and such changes should also fall within the scope of protection claimed by the present invention.

[0090] The present invention will be further described below by way of specific embodiments. Unless otherwise specified, all instruments, devices, equipment and other items used in the embodiments of the present invention are obtained through conventional commercial means.

[0091] This invention provides a temporomandibular joint image synthesis system, method, and storage medium, which are particularly suitable for high-fidelity generation of temporomandibular joint cone-beam CT to MRI images.

[0092] The core improvement of this invention lies in the first-ever introduction of the TransTMJ-Net conditional diffusion generation architecture into the field of cross-modal medical image generation. By integrating a pre-trained diffusion backbone network and a conditional control sub-network, it accurately captures the anatomical constraints of bone tissue in the source modality using zero-convolutional layers. Furthermore, it fuses textual semantic priors extracted from a contrastive language-image pre-trained model through a cross-attention mechanism. During the iterative processes of forward noise addition and reverse denoising in the Markov chain, it achieves a nonlinear mapping from Gaussian noise to the soft tissue features of the target modality. Compared to traditional generative adversarial networks, this invention effectively overcomes problems such as unstable model training, modality collapse, and a lack of physical realism in the generated images.

[0093] Meanwhile, this invention innovatively introduces a three-dimensional rigid registration algorithm with multiple anatomical landmarks. By precisely aligning four anchor points—the bilateral condylar centers and the zygomaticotemporal suture—in the axial, sagittal, and coronal planes, it eliminates the positional differences of patients acquired using different devices, providing a high-quality paired dataset for model training. Furthermore, through a three-frame sliding window spatiotemporal continuity guidance mechanism, adaptive region-of-interest cropping, and voxel normalization preprocessing strategies, this invention ensures the anatomical coherence and topological consistency of the synthesized images in three-dimensional space, while eliminating non-target background interference, significantly improving generation quality and training efficiency.

[0094] Example 1 This embodiment provides a temporomandibular joint image synthesis system, such as Figure 1 As shown, the system includes: A cross-modal image collaborative preprocessing module is used to perform spatial mapping between source and target modal images. This module utilizes a three-dimensional rigid registration operator, using preset anatomical landmarks as anchor points to achieve anatomical spatial alignment, and employs an adaptive region of interest extraction algorithm to uniformly resample the source and target modal images to the target resolution voxel space. The cross-modal image collaborative preprocessing module is connected to a multi-dimensional feature extraction and encoding module via data connectivity. The multidimensional feature extraction and encoding module includes an image encoder and a visual semantic encoder. The image encoder is used to project source modality image slices onto a latent feature space. The visual semantic encoder is based on a contrastive language-image pre-trained model to extract global semantic feature vectors related to anatomical sites. The multidimensional feature extraction and encoding module is connected to the diffusion generation core module through data connection. The diffusion generation core module, based on a diffusion model architecture, consists of a denoising autoencoder and a conditional control branch. The conditional control branch adjusts the denoising trajectory in real time through residual connections, so that the generated image maintains the contrast of the target modality while following the anatomical constraints provided by the source modality image. The diffusion generation core module is connected to the spatiotemporal continuity guidance module through a data connection. The spatiotemporal continuity guidance module introduces a sliding window mechanism during the generation process, using continuous multi-frame source modal image slices as input constraints, and ensuring the spatial smoothness of the synthesized target modal image at the longitudinal anatomical level through multi-frame feature fusion.

[0095] Example 2 This embodiment provides a temporomandibular joint image synthesis system, such as Figure 1 As shown, the system includes: A cross-modal image collaborative preprocessing module is used to perform spatial mapping between source and target modal images. This module utilizes a three-dimensional rigid registration operator, using preset anatomical landmarks as anchor points to achieve anatomical spatial alignment, and employs an adaptive region of interest extraction algorithm to uniformly resample the source and target modal images to the target resolution voxel space. The cross-modal image collaborative preprocessing module is connected to a multi-dimensional feature extraction and encoding module via data connectivity. The multidimensional feature extraction and encoding module includes an image encoder and a semantic encoder. The image encoder is used to project source modality image slices onto a latent feature space. The visual semantic encoder is based on a contrastive language-image pre-trained model to extract global semantic feature vectors related to anatomical sites. The multidimensional feature extraction and encoding module is connected to the diffusion generation core module through data connection. The diffusion generation core module, based on a diffusion model architecture, consists of a denoising autoencoder and a conditional control branch. The conditional control branch adjusts the denoising trajectory in real time through residual connections, so that the generated image maintains the contrast of the target modality while following the anatomical constraints provided by the source modality image. The diffusion generation core module is connected to the spatiotemporal continuity guidance module through a data connection. The spatiotemporal continuity guidance module introduces a sliding window mechanism during the generation process, using continuous multi-frame source modal image slices as input constraints, and ensuring the spatial smoothness of the synthesized target modal image at the longitudinal anatomical level through multi-frame feature fusion.

[0096] The cross-modal image collaborative preprocessing module is used to acquire cone-beam computed tomography (CBCT) images and magnetic resonance imaging (MRI) images of the temporomandibular joint. Specifically, multiple sets of dual-joint CBCT and MRI images were selected, and a total of 739 cases of image data including normal morphology and disc displacement status were selected. The data were then de-identified.

[0097] The MRI scanner was a GE Discovery 3.0T, and the scanning sequence was two-dimensional fast spin echo oblique sagittal proton density-weighted imaging. The CBCT scanner was a NewTom, with a voltage of 110kV, a voxel size of 0.3mm, and a scanning range of 17cm × 19cm. Image data for each case were exported in DICOM format with a spatial resolution of 512 × 512 × n voxels.

[0098] The cross-modal image collaborative preprocessing module also includes an image registration unit and an image cropping and normalization unit. The image registration unit employs an image registration method based on the mutual information similarity criterion. For example... Figure 2 As shown, voxel-level rigid registration of CBCT and MRI sequences was first performed using 3D-Slicer software and the Elastix toolkit. Subsequently, two oral radiology specialists manually fine-tuned the data in the axial, sagittal, and coronal planes based on four anatomical landmarks.

[0099] These four anatomical landmarks are: the center of the left condyle, the center of the right condyle, the left zygomaticotemporal suture, and the right zygomaticotemporal suture. Figure 2 Key anatomical landmarks are marked with small yellow circles. Through manual fine-tuning, the anteroposterior orientation of the CBCT is aligned with the superior-inferior orientation of the MRI until the mutual information value between the two sets of images reaches its maximum.

[0100] Before registration, the mutual information value between CBCT and MRI images was 0.65, and the spatial alignment error was 2.8 mm. After registration, the mutual information value increased to 0.93, and the spatial alignment error decreased to 0.4 mm. Figure 2 In the merged image at the bottom, there are obvious red and blue artifacts before registration, but the artifacts disappear after registration, which verifies the registration accuracy.

[0101] In addition, registration was performed based on three structures: condyle, external auditory canal, and articular tubercle. The registration results are shown in Table 1. Among them, the 95% Hausdorff distance is used to measure the maximum deviation of the segmented contour boundary and evaluate the contour edge fit. The smaller the value, the better the segmented anatomical contour fits the real anatomical contour edge. The Dice similarity coefficient is the overlap index. The closer the value is to 1, the higher the overlap between the algorithm segmented region and the real anatomical region, and the better the segmentation effect. The target registration error TRE represents the spatial distance error between the predicted anatomical landmark coordinates and the real landmarks, and is expressed as median, percentile, and extreme value.

[0102] Table 1

[0103] The median overall registration error was less than 0.71 mm, which meets the accuracy requirements for clinical image registration.

[0104] The image cropping and normalization unit performs unified preprocessing on the registered images. For MRI images, the system first truncates the voxel intensity to the range of 0 to 3500, and then normalizes it to the interval of -1 to 1. For CBCT images, the system normalizes the voxel intensity from the original range of 0 to 1023 to the interval of -1 to 1.

[0105] To eliminate interference from non-target brain tissue and improve spatial registration tolerance, this unit uses the core anatomical structure of the temporomandibular joint as the center, cropping out a region of interest with a resolution of 256×256, and resampling it to 512×512 as the model input. After removing slices lacking core anatomical structures at the beginning and end of the sequence, 1136 pairs of paired data were finally obtained.

[0106] This voxel standardization and region of interest clipping strategy can improve the accuracy of training results and reduce noise that may be generated in other regions.

[0107] The multidimensional feature extraction and encoding module includes an image encoder and a semantic encoder. This module uses preprocessed three consecutive CBCT slices as spatial input.

[0108] The image encoder projects CBCT slices into a latent feature space. Specifically, the preprocessed CBCT images are input into the ControlNet branch, and spatial constraint features are extracted using a convolutional neural network.

[0109] The semantic encoder, based on a contrastive language-image pre-trained model, extracts global semantic feature vectors related to the anatomical locations of the temporomandibular joint. Specifically, the intermediate layer of three consecutive slices is input into the image encoder of the pre-trained CLIP model. This image encoder directly extracts high-dimensional implicit semantic embedding vectors by parsing the visual attributes of the anatomical locations in the input image.

[0110] For example, for a sagittal section sequence of the temporomandibular joint, the system directly extracts its central sagittal section and inputs it into the image branch of the pre-trained CLIP encoder. The CLIP encoder implicitly encodes the slice into a 512-dimensional high-dimensional semantic embedding vector by directly parsing the low-level visual anatomical properties of the slice. This vector serves as a global visual semantic prior, guiding the subsequent diffusion generation process.

[0111] The diffusion generation core module is the core of the system's image generation. For example... Figure 3 As shown, this module adopts the diffusion model as its basic architecture and introduces a conditional control network branch.

[0112] Figure 3 The left side shows the U-Net backbone network, which contains multiple codec modules and uses a pre-trained StableDiffusion model as the diffusion backbone. The right side shows the ControlNet conditional control branch, which receives CBCT spatial feature constraints by constructing trainable replicas.

[0113] The ControlNet branch locks the fundamental parameters of the original Stable Diffusion model, training only copies of the additional neural network layers introduced to handle conditional inputs. This design ensures the model's training stability when processing limited medical image data and effectively reduces the number of parameters that need to be updated, avoiding the risk of overfitting in small sample scenarios.

[0114] The ControlNet branch uses zero-convolutional layers to receive the spatial feature constraints of the aforementioned CBCT. The weights of the zero-convolutional layers are initialized to zero, so as not to interfere with the backbone network in the early stages of training. As training progresses, the network gradually learns how to inject the bone tissue features of CBCT into the diffusion process.

[0115] The Stable Diffusion model uses the AdamW optimizer with a learning rate of 0.0001, 20 inference denoising steps, a training batch size of 16, a total of 64,000 training iterations, a regularization coefficient of 0.7, a decay coefficient every 2,000 iterations, and an epoch of 100.

[0116] The training logic of this module follows the principle of forward noise addition and reverse noise removal of Markov chains. During the forward diffusion process, the generation process of the noisy image is defined as a Markov chain, based on a pre-defined variance table. Noise that follows a standard normal distribution is gradually added to the input data. (0, I); Noisy image at step n Depends on the image from the previous step ,satisfy:

[0117] Therefore, the noise addition formula for step n is:

[0118] in, This represents the original image without added noise.

[0119] In the reverse denoising process, the diffusion generation core module recursively calculates an inverse Gaussian distribution from... Reverse derivation to target soft tissue image Furthermore, a reverse diffusion process is introduced during the reverse denoising process:

[0120] in, Characterizes the noise image given at the current nth step. and cross-modal conditional input The time-space network predicts a clearer image from the previous step (step n-1). The conditional probability distribution, This represents the conditional input information, which guides the model to generate specific MRI images from CBCT source images. This represents the Gaussian (normal) distribution function. This indicates that the model was generated by TransTMJ-Net (with parameters). The predicted Gaussian distribution mean is essentially the network's prediction based on the current noise level. and CBCT bony boundary constraints The calculated "denoising direction and soft tissue structure characteristics" The variance (covariance matrix) of the predicted Gaussian distribution represents the confidence and uncertainty of the current step's denoising generation.

[0121] Specifically, in the decoder part of the U-Net backbone network, each decoder module contains a cross-attention layer. This layer receives semantic embedding vectors from the CLIP encoder as keys and values, receives latent image features from the current layer as queries, and calculates semantic guidance weights through an attention mechanism to incorporate high-level semantic priors into the image features.

[0122] Through this mechanism, the present invention introduces high-level semantic priors during the generation process, which enhances the morphological rationality and positional accuracy of the synthesized images on key anatomical structures such as articular discs and condyles, and solves the problem of inconsistency in anatomical topology logic that may be caused by simply generating based on pixel-level features.

[0123] The denoising steps in the inference phase are set to 20. Finally, the latent features are decoded into pixel-level images using a pre-trained decoder to generate a synthetic MRI with high anatomical realism.

[0124] The spatiotemporal continuity guidance module introduces a sliding window mechanism during the generation process. Specifically, three consecutive CBCT slices are used as input constraints, intermediate slices are fed into the CLIP encoder to extract semantic features, and the spatial features of the three slices are fused using a ControlNet branch.

[0125] This mechanism ensures the spatial smoothness of synthetic MRI in the longitudinal anatomical plane, avoiding abrupt changes in articular disc morphology between adjacent slices. Through spatiotemporal feature fusion, the anatomical coherence of the synthetic MRI sequence in three-dimensional space is guaranteed.

[0126] Example 3 This embodiment provides a method for synthesizing temporomandibular joint images, executed on the aforementioned temporomandibular joint image synthesis system, including: S1. The cross-modal image collaborative preprocessing module obtains the source modal image sequence, performs rigid registration based on grayscale information, and extracts local voxel blocks with key anatomical structures as the core. S2. The multidimensional feature extraction and encoding module extracts spatial constraint features from the preprocessed source modal image input condition control branch. At the same time, it extracts the center slices of multiple consecutive frames of source modal images and inputs them into the image encoder of the vision-language pre-trained model to directly extract high-dimensional implicit semantic embeddings. S3. The diffusion generation core module starts from Gaussian noise in the diffusion model, and under the dual guidance of the spatial constraint features and the semantic embedding, iteratively predicts and removes noise to restore the potential features of the target modal image. S4. The diffusion generation core module decodes the latent features into pixel-level images using a pre-trained decoder to generate a synthetic target modal image.

[0127] Example 4 This embodiment provides comparative experimental data with existing technologies to verify the technical effects of the present invention.

[0128] The testing conditions were as follows: 739 cases of data provided in Example 2; comparison methods included CycleGAN, Pix2Pix, and the TransTMJ-Net of this invention. Evaluation metrics included peak signal-to-noise ratio, structural similarity index, Frechet Inception Distance, expert diagnostic accuracy, condylar complete reconstruction rate, and positive rate of articular disc formation.

[0129] CycleGAN is a cycle-consistent GAN that does not require paired data and is widely used for cross-modal image transformation. Pix2Pix is ​​a conditional GAN ​​based on paired data, employing a U-Net generator and a PatchGAN discriminator.

[0130] The results of the comparative experiments are shown in Table 2: Table 2

[0131] As shown in Table 2, the peak signal-to-noise ratio (PSNR) of the present invention TransTMJ-Net is improved by 5.4 dB compared to CycleGAN, representing an improvement of 42.6%. Compared to Pix2Pix, the PSNR is improved by 3.5 dB, representing an improvement of 24%.

[0132] In terms of Learning-Based Perceptual Image Patch Similarity (LPIPS), this invention reduces the similarity by 0.015 compared to CycleGAN, a reduction of 5.1%. Compared to Pix2Pix, it reduces the similarity by 0.012, a reduction of 4.3%.

[0133] Regarding the mean absolute error (MAE), this invention reduces it by 0.078 compared to CycleGAN, a reduction of 46.1%. Compared to Pix2Pix, it reduces it by 0.021, a reduction of 16.4%.

[0134] Regarding mean squared error (MSE), this invention reduces it by 0.041 compared to CycleGAN, a reduction of 70.7%. Compared to Pix2Pix, it reduces it by 0.021, a reduction of 55.3%.

[0135] Regarding spectral residual similarity (SRSIM), this invention achieves an improvement of 0.081 compared to CycleGAN, representing a 10.6% improvement. Compared to Pix2Pix, it achieves an improvement of 0.049, representing a 6.2% improvement.

[0136] In terms of Feature Similarity Index (FSIM), this invention improves upon CycleGAN by 0.093, representing a 13.8% improvement. Compared to Pix2Pix, it improves by 0.053, representing a 7.4% improvement.

[0137] like Figure 4 As shown, the cross-modal generation effect is compared between normal and anterior displacement of the articular disc. Figure 4 The image is divided into three rows: the first row is the input CBCT image, the second row is the real MRI image, and the third row is the synthetic MRI generated by this invention.

[0138] At different slice locations, the synthetic MRI generated by this invention is highly consistent with real MRI in key anatomical features such as condylar contour, articular disc morphology, and joint space. Especially in cases of anterior disc displacement, the generated images accurately reconstruct the abnormal positional relationship between the articular disc and the condyle, verifying the superior performance of this invention in preserving pathological details.

[0139] The comparative experiments above demonstrate that the present invention, by introducing a conditional diffusion generation architecture, multi-marker registration, CLIP semantic guidance, spatiotemporal continuity guidance, region of interest cropping, and voxel standardization, comprehensively surpasses existing technologies in terms of objective image quality indicators and clinical diagnostic efficacy, achieving unexpected technical effects.

[0140] In particular, the present invention achieves a 100% positive rate for the formation of condylar structures and a 94.8% positive rate for the formation of articular discs. The latter is 13 to 21 percentage points higher than the prior art, which fully verifies the technical advantages of the present invention in processing low-contrast micro-anatomical structures and solves the bottleneck problem of soft tissue formation in the prior art.

[0141] Furthermore, using the synthetic MRI generated by this invention to assist in diagnosis, two medical image readers respectively used CBCT images and the CBCT+MRI synthetic images of this invention for diagnosis. The final accuracy, sensitivity, and specificity data are shown in Table 3. It can be seen that using the synthetic MRI of this invention to assist in diagnosis significantly increased the accuracy from 50.3% to 89.0% and from 51.7% to 86.3% (P<0.05). In particular, for the most common clinical condition, anterior disc displacement (ADD), the diagnostic accuracy rate of doctors reached 91.7% with the assistance of synthetic MRI, successfully identifying the vast majority of displacement cases. This indicates that the images generated by this invention provide crucial soft tissue visual details, greatly enhancing doctors' ability to identify diseases.

[0142] Table 3

[0143] It is understood that the backbone network of the diffusion model in the above embodiments is not limited to Stable Diffusion, and can also adopt other diffusion model architectures such as DDPM, DDIM, Latent Diffusion Model, as long as it can achieve the reverse denoising generation from Gaussian noise to the target image.

[0144] Obviously, semantic encoders are not limited to CLIP; other vision-language pre-trained models such as ALIGN, BLIP, and CoCa can also be used, as long as they can extract global semantic feature vectors related to anatomical sites.

[0145] Understandably, the registration tool is not limited to Elastix. Other medical image registration software such as ANTs, SimpleITK, and NiftyReg can also be used, as long as they can achieve rigid three-dimensional registration based on the mutual information criterion.

[0146] It is understandable that the resolution of the region of interest cropping is not limited to 256×256 upsampled to 512×512, but can also be adjusted to 128×128 upsampled to 256×256, or 512×512 upsampled to 1024×1024, depending on the specific application scenario.

[0147] Obviously, the voxel intensity cutoff range is not limited to 0 to 3500 for MRI and 0 to 1023 for CBCT. It can also be adaptively adjusted according to the gray value distribution of different scanning devices, such as 0 to 4000 for MRI and 0 to 2047 for CBCT.

[0148] Understandably, the number of frames in a continuous slice is not limited to 3 frames. Five, seven, or more frames can also be used to guide spatiotemporal continuity. The more frames there are, the better the spatial smoothness, but the computational complexity also increases accordingly.

[0149] Obviously, this invention is not limited to CBCT to MRI synthesis of the temporomandibular joint, but can also be extended to cross-modal image generation of other anatomical sites or other modal combinations, as long as parameters such as registration anchor points, region of interest cropping regions, and semantic descriptors are adjusted.

[0150] Example 5 This embodiment provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described cross-modal medical image synthesis method.

[0151] The storage medium may be a non-volatile storage medium, including but not limited to at least one of read-only memory, random access memory, disk, optical disk, USB flash drive, and solid-state drive.

[0152] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0153] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention do not depart from the essence and scope of the technical solution of the present invention.

Claims

1. A temporomandibular joint image synthesis system, characterized in that, include: A cross-modal image collaborative preprocessing module is used to perform spatial mapping between source and target modal images. This module utilizes a three-dimensional rigid registration operator, using preset anatomical landmarks as anchor points to achieve anatomical spatial alignment, and employs an adaptive region of interest extraction algorithm to uniformly resample the source and target modal images to the target resolution voxel space. The cross-modal image collaborative preprocessing module is connected to a multi-dimensional feature extraction and encoding module via data connectivity. The multidimensional feature extraction and encoding module includes an image encoder and a semantic encoder. The image encoder is used to project source modality image slices onto a latent feature space. The semantic encoder is based on a contrastive language-image pre-trained model to extract global semantic feature vectors related to anatomical sites. The multidimensional feature extraction and encoding module is connected to the diffusion generation core module through data connection. The diffusion generation core module, based on the diffusion model architecture, consists of a denoising autoencoder and a conditional control branch. The conditional control branch adjusts the denoising trajectory in real time through residual connections, so that the generated image maintains the contrast of the target modality while following the anatomical constraints provided by the source modality image. The diffusion generation core module is connected to the spatiotemporal continuity guidance module via a data connection; The spatiotemporal continuity guidance module introduces a sliding window mechanism during the generation process, using continuous multi-frame source modal image slices as input constraints, and ensuring the spatial smoothness of the synthesized target modal image at the longitudinal anatomical level through multi-frame feature fusion.

2. The system according to claim 1, characterized in that, The training logic of the diffusion generation core module follows the forward noise addition and reverse noise reduction principle of Markov chains. In the reverse noise reduction stage, the system uses a cross-attention mechanism to fuse the semantic feature vector with the latent features of the image, guiding the noise predictor to perform dissecting feature correction in each time step.

3. The system according to claim 2, characterized in that, The conditional control branch locks the basic parameters of the pre-trained diffusion model, trains only a copy of the neural network layer used to process the conditional input, and utilizes the spatial feature constraints of the source modal image received by the zero convolutional layer.

4. The system according to claim 2, characterized in that, The forward noise addition to the Markov chain specifically includes: In the forward diffusion process, the generation process of the noisy image is defined as a Markov chain, based on a pre-defined variance table. Noise that follows a standard normal distribution is gradually added to the input data. (0, I); Noisy image at step n Depends on the image from the previous step ,satisfy: Therefore, the noise addition formula for step n is: in, This represents the original image without added noise.

5. The system according to claim 2, characterized in that, The inverse denoising of the Markov chain specifically includes: In the reverse denoising process, the diffusion generation core module recursively calculates an inverse Gaussian distribution from... Reverse derivation to target soft tissue image Furthermore, a reverse diffusion process is introduced during the reverse denoising process: in, Characterizes the noise image given at the current nth step. and cross-modal conditional input At that time, the network predicts a clearer image in step n-1. The conditional probability distribution, Represents the Gaussian distribution function. This represents the mean of the Gaussian distribution predicted by the TransTMJ-Net generative model, which is essentially the network's prediction based on the current noise. and CBCT bony boundary constraints The calculated denoising direction is related to the soft tissue structure characteristics. This represents the variance of the predicted Gaussian distribution, and denoises the confidence and uncertainty of the current step's denoising generation.

6. The system according to claim 1, characterized in that, The cross-modal image collaborative preprocessing module includes a voxel normalization rule. For the target modal image, its voxel intensity is truncated to the voxel intensity truncation range of the target modal image and then normalized to a preset normalization interval. For the source modal image, its voxel intensity is normalized from the voxel intensity truncation range of the source modal image to the preset normalization interval. The region of interest extraction algorithm removes slices that lack core anatomical structures at the beginning and end of the sequence.

7. The system according to claim 6, characterized in that, Before training and inference, DICOM format images with spatial resolution matching are acquired; for MRI images, their voxel intensity is first truncated to the range of [0, 3500], and then normalized to the interval of [-1, 1]; simultaneously, for cone-beam CT images, their voxel intensity is normalized from the original range of [0, 1023] to the interval of [-1, 1].

8. The system according to claim 6, characterized in that, The three-dimensional rigid registration operator adopts a multi-level optimization strategy based on the mutual information similarity criterion. First, it uses a registration toolkit to perform voxel-level rigid registration between the source modal image and the target modal image. Then, it performs manual fine-tuning in the axial, sagittal and coronal directions based on multiple anatomical landmarks. Registration is performed by maximizing the mutual information value between the source modal image and the target modal image until the mutual information value reaches a preset threshold.

9. The system according to claim 8, characterized in that, Several anatomical landmarks include: the center of the left condyle, the center of the right condyle, the left zygomaticotemporal suture, and the right zygomaticotemporal suture.

10. The system according to claim 8, characterized in that, The region of interest is the area centered on the core anatomical structures of the temporomandibular joint with a resolution of 256×256.

11. The system according to claim 1, characterized in that, The spatiotemporal continuity guidance module uses continuous multi-frame source modal image slices as input constraints, and the intermediate layer slices are input to the semantic encoder to extract global semantic features. The spatial features of the multi-frame slices are fused through the conditional control branch.

12. The system according to claim 11, characterized in that, The spatiotemporal continuity guidance module uses three consecutive source modal image slices as input constraints.

13. A method for synthesizing temporomandibular joint images, performed on the temporomandibular joint image synthesis system according to any one of claims 1-12, characterized in that, include: S1. The cross-modal image collaborative preprocessing module obtains the source modal image sequence, performs rigid registration based on grayscale information, and extracts local voxel blocks with key anatomical structures as the core. S2. The multidimensional feature extraction and encoding module extracts spatial constraint features from the preprocessed source modal image input condition control branch; at the same time, it extracts the center slices of multiple consecutive frames of source modal images and inputs them into the image encoder of the vision-language pre-trained model to directly extract high-dimensional implicit semantic embedding vectors. S3. The diffusion generation core module starts from Gaussian noise in the diffusion model, and under the dual guidance of the spatial constraint features and the semantic embedding, iteratively predicts and removes noise to restore the potential features of the target modal image. S4. The diffusion generation core module decodes the latent features into pixel-level images using a pre-trained decoder to generate a synthetic target modal image.

14. The method according to claim 13, characterized in that, The local voxel block extracted in step S1 has a resolution of 256×256×N. The rigid registration is performed by maximizing the mutual information value between the source modal image and the target modal image until the mutual information value reaches a preset threshold.

15. The method according to claim 13, characterized in that, In step S3, the number of denoising steps in the inference stage is set to 10 to 30, and the cross-attention mechanism is used to correct the anatomical features within each time step.

16. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the temporomandibular joint image synthesis method of claim 13.

Citation Information

Patent Citations

  • CBCT pseudo nuclear magnetic image reconstruction system and method based on artificial intelligence

    CN119941889A