A fusion processing method based on multimodal oral image data
By combining rigid and non-rigid registration, CycleGAN model and wavelet transformation, local registration error and complex deformation problems of multimodal oral image data are solved, and high-precision and efficient fusion of multimodal image data are achieved.
Patent Information
- Application Number
- CN202510795584.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-06-16
AI Technical Summary
In the prior art, the local registration error and complex deformation adaptability of multimodal oral image data are insufficient, resulting in low multimodal fusion accuracy and efficiency.
A strategy of combining rigid registration and non-rigid registration is adopted, and a structural mask is generated by combining 3D Canny edge detection and dynamic threshold segmentation technology. Multimodal image fusion is used to optimize image data through wavelet transformation and feature decoupling technology to establish a spatiotemporal alignment and hierarchical feature fusion model.
It significantly improves the accuracy and efficiency of multimodal image fusion, effectively corrects complex deformation, inhibits metal artifacts, and maintains the clinical availability and accuracy of the image.
Smart Images

Figure CN120318093B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to a fusion processing method based on multimodal oral image data. Background Art
[0002] In oral medicine, clinicians often rely on multiple imaging data, such as X-rays, CT scans (e.g., CBCT), and oral endoscopy images, to comprehensively assess a patient's oral health. However, each modality has its own strengths and weaknesses, and when used alone, they often fail to provide complete diagnostic information.
[0003] Chinese invention patent publication number CN118537699B provides a method for fusing and processing multimodal oral image data. Initial multimodal oral image acquisition is performed using a 3D oral scanner and a cone-beam computed tomography (CBCT) device to generate multimodal oral image data. The multimodal oral image data is then regionally aligned to generate aligned multimodal oral image data. Multimodal oral image matching analysis is performed on the aligned multimodal oral image data to generate multimodal oral image matching data. Multimodal oral image mapping and fusion processing is then performed on the multimodal oral image matching data using a preset CycleGAN model to generate fused oral image data. Finally, fused oral image feature mining is performed on the fused oral image data to generate fused oral image feature data.
[0004] Although affine transformation is used to achieve spatial alignment of multimodal images, the internal structure of the oral cavity has complex nonlinear deformations. Affine transformation cannot completely correct such deformations, which may lead to local registration errors and insufficient adaptability to complex deformations. Summary of the Invention
[0005] This application provides a fusion processing method based on multimodal oral image data, which solves the problems of local alignment errors and insufficient adaptability to complex deformations in the existing technology, and achieves the technical effect of improving the accuracy and efficiency of multimodal fusion.
[0006] The present application provides a fusion processing method based on multimodal oral image data, the method comprising:
[0007] S1: Acquire multimodal oral image data, and obtain a structure mask and a grayscale image based on the multimodal oral image data; the multimodal oral image data includes 3D surface oral image data and CBCT oral image data;
[0008] S2: Perform rigid registration on the structure mask to obtain alignment parameters, and perform non-rigid registration on the grayscale image to obtain an enhanced grayscale image; input the structure mask, alignment parameters and enhanced grayscale image into the CycleGAN model to obtain a multimodal registered fusion image;
[0009] S3: Perform wavelet transform on the multimodal registered fusion image to obtain multi-scale frequency domain data, and annotate the multi-scale frequency domain data to obtain annotated areas;
[0010] The multimodal oral imaging data collected at different times are used as time series imaging data;
[0011] Perform spatiotemporal alignment based on the annotated regions and time-series image data to obtain a multimodal spatiotemporal registration dataset;
[0012] S4: Feature decoupling is performed based on the multimodal spatiotemporal registration dataset to obtain hierarchical features; regional adaptive fusion is performed based on the hierarchical features, structural masks and grayscale images to obtain a multimodal optimized fused image.
[0013] Furthermore, the generation of the structure mask includes:
[0014] 3D Canny edge detection and topology restoration are used on 3D surface oral image data to obtain three-dimensional edge contours; dynamic threshold segmentation and connected domain analysis are used on CBCT oral image data to obtain bone structure segmentation results; the three-dimensional edge contours and bone structure segmentation results are superimposed to generate a structural mask.
[0015] Furthermore, the 3D Canny edge detection includes:
[0016] Three-dimensional directional gradient detection is used on multimodal oral imaging data to calculate the brightness change intensity in the X, Y, and Z axis directions to obtain the gradient amplitude map and direction map;
[0017] Based on the gradient magnitude map and direction map, the edge is refined to a single pixel width through non-maximum suppression to obtain a refined edge map;
[0018] The gradient threshold is set based on the edge map to retain the boundary features between the crown edge and the gingiva and obtain the three-dimensional edge contour.
[0019] Furthermore, the rigid registration is to obtain the bone contour features based on the structure mask, and then calculate the rigid transformation parameters to obtain the alignment parameters; and obtain the enhanced grayscale image through non-rigid registration with the grayscale image according to the alignment parameters.
[0020] Furthermore, the CycleGAN model includes:
[0021] The structure mask and alignment parameters are used as the structure channel data, and the enhanced grayscale image is used as the texture channel data to input into the CycleGAN model;
[0022] Among them, the structure channel uses a convolutional network to extract anatomical contour features; the texture channel uses a residual network to extract multi-scale texture features; the anatomical contour features are fused with the multi-scale texture features to obtain a multimodal registration fusion image.
[0023] Furthermore, the wavelet transform optimization includes:
[0024] Perform three-level wavelet decomposition on the multimodal registered fusion image to obtain low-frequency components and high-frequency sub-bands;
[0025] Identify the artifact area in the high-frequency sub-band, determine the artifact interference level based on the artifact area, and then set the value range of the high-frequency weight coefficient and the value range of the normal area weight coefficient;
[0026] Sharpen the low-frequency components, enhance the clarity of bone contours, and suppress interference from artifact areas;
[0027] The weight-adjusted high-frequency subband is fused with the sharpened low-frequency component by inverse wavelet transform to obtain multi-scale frequency domain data.
[0028] Furthermore, the marked area includes:
[0029] The multi-scale frequency domain data was divided into 8×8 pixel blocks. The deformation field residual within each block was calculated based on the alignment parameters. The area with a deformation error greater than 0.3 mm was identified as an uncorrected nonlinear deformation area. The sliding window standard deviation of the HU value of each block was calculated. The area with a standard deviation exceeding 200 was identified as a metal artifact interference area. The nonlinear deformation area and the metal artifact interference area together constituted the annotation area.
[0030] Furthermore, the feature decoupling step includes:
[0031] The global structural features of the structure mask are extracted through a convolutional neural network, and the symmetry error with the gold standard is constrained to be less than or equal to 0.3mm.
[0032] The local detail features of the grayscale image are extracted through the residual network, keeping the edge gradient amplitude no less than 90% of the original image; the spatial attention mechanism is used to dynamically fuse global structural features and local detail features to obtain hierarchical features; a unified spatial benchmark is provided for time series images to reduce the negative impact of artifacts on fusion quality.
[0033] Furthermore, the spatiotemporal alignment step includes:
[0034] The non-rigid deformation field of the soft tissue deformation is obtained by non-rigid registration calculation, and the composite alignment parameter is obtained according to the alignment parameter and the non-rigid deformation field;
[0035] Based on composite alignment parameters, the time-series image data are registered to the static reference space to form a rigid registration grouping. A unified spatial coordinate system is established through the elastic deformation field, and the rigid registration error tolerance is less than or equal to 0.1mm. Based on the feature data extracted from the annotated area, a non-rigid compensation grouping is constructed, and a recurrent neural network is used to predict the local progressive deformation field, with a maximum deformation of less than or equal to 2.5mm. A hierarchical registration strategy is used to achieve precise alignment of bones and soft tissues across time points. The rigid registration grouping maintains the consistency of the spatial position of the implant, and the non-rigid compensation grouping corrects gingival deformation, outputting a multimodal spatiotemporal registration dataset.
[0036] Furthermore, the region adaptive fusion step includes:
[0037] Based on hierarchical features, the multimodal spatiotemporal registration dataset is segmented into [10, 15] functional regions;
[0038] Compute spatial proximity edges based on structure masks;
[0039] Calculate feature similarity edges based on multi-scale frequency domain data;
[0040] According to the spatial proximity edge and feature similarity edge, a regional correlation matrix is established;
[0041] Based on the regional correlation matrix, the attention network is used to dynamically allocate modal weights to obtain a multimodal optimized fusion image.
[0042] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0043] By adopting a strategy that combines rigid registration with non-rigid registration, combined with 3D Canny edge detection and dynamic threshold segmentation technology, the problem of complex deformation correction between multimodal images is effectively solved; by constructing a dual-channel CycleGAN model, the biological characteristics of the gingival texture are restored while maintaining the spatial position accuracy of the implant, significantly improving the clinical usability of multimodal fusion images; using local wavelet optimization and feature decoupling fusion mechanism, metal artifacts are effectively suppressed, while computational efficiency is increased through regional adaptive fusion strategy; by establishing a correlation model between spatiotemporal alignment and hierarchical features, the technical effect of improving multimodal fusion accuracy and efficiency is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a flow chart of a fusion processing method based on multimodal oral image data in an embodiment of the present invention;
[0045] Figure 2 This is a feature decoupling flow chart in an embodiment of the present invention. DETAILED DESCRIPTION
[0046] To facilitate understanding of the present invention, the present application will be described more comprehensively below with reference to the relevant drawings; the drawings show preferred embodiments of the present invention, but the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, the purpose of providing these embodiments is to enable a more thorough and comprehensive understanding of the disclosed content of the present invention.
[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains; the terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more of the associated listed items.
[0048] Example 1: Figure 1 As shown, a fusion processing method based on multimodal oral image data includes:
[0049] S1: Acquire multimodal oral image data, and obtain a structure mask and a grayscale image based on the multimodal oral image data; the multimodal oral image data includes 3D surface oral image data and CBCT oral image data;
[0050] Specifically, a 3D oral scanner uses optical imaging technology to capture the three-dimensional geometry and texture information of the inner surface of the oral cavity to obtain 3D surface oral image data; a cone-beam CT device uses the collaborative work of a cone X-ray beam and a two-dimensional flat-panel detector to collect multi-angle projection data and ultimately obtain CBCT (Cone Beam Computed Tomography) oral image data.
[0051] The generation of the structure mask includes: applying 3D Canny edge detection and topological repair processing to 3D surface oral image data to obtain a three-dimensional edge contour; applying dynamic threshold segmentation and connected domain analysis to CBCT oral image data to obtain a bone structure segmentation result; and superimposing the three-dimensional edge contour and the bone structure segmentation result to generate a structure mask.
[0052] The 3D Canny edge detection includes: using three-dimensional directional gradient detection on multimodal oral imaging data, calculating the brightness change intensity in the X, Y, and Z axis directions to obtain a gradient amplitude map and a direction map; based on the gradient amplitude map and the direction map, refining the edge to a single pixel width through non-maximum suppression to obtain a refined edge map; setting a gradient threshold based on the edge map to retain the boundary characteristics between the crown edge and the gums to obtain a three-dimensional edge contour.
[0053] Specifically, based on the 3D surface oral imaging data, 3D Canny edge detection (used to detect structural boundaries in volume data) and topology restoration are used to further process the data;
[0054] Using 3D Canny edge detection, we first remove tiny noise points (denoising) from the scanned data, such as fluctuations caused by equipment noise or patient movement, and use a "blur filter" to smooth the three-dimensional data as a whole. This blurring is not uniform, but rather gives higher weights to close areas to ensure that the edges are not overly blurred. Then, we calculate the gradient and detect the intensity of the brightness change in the X, Y, and Z directions respectively. The slopes in the three directions are combined to obtain the gradient amplitude map and direction map, marking all possible edge positions. Through the brightness changes in the three-dimensional direction, we find the steep areas in the data, such as the edge of the tooth and the gingival junction. Secondly, through non-maximum suppression refinement, we check whether each point along the gradient direction is the steepest position in that direction. If so, it is retained as an edge point; otherwise, it is eliminated, thereby refining the edges of the thick lines to a single pixel width to avoid overlap or blurring, and obtaining a refined edge map. Finally, we set the gradient threshold, which is formulated as follows:
[0055] ,
[0056] in, is the high / low threshold, is the global maximum value of the gradient amplitude of all data, is the proportional coefficient, the high threshold is [0.6, 0.8], and the low threshold is [0.2, 0.4];
[0057] Based on the gradient threshold, valid edges are screened. Areas with very obvious brightness changes, such as the border between the crown and the gum, are retained directly. Areas with less obvious brightness changes, such as the enamel surface texture, are retained if they are connected to strong edges, otherwise they are treated as noise and removed. The final output is a continuous, high-precision 3D edge contour.
[0058] To address problems that may arise after 3D Canny edge detection, topology repair is used for optimization: first, the edge is expanded outward by 1-2 pixels to connect the broken areas, and then the expanded edge is contracted inward to restore the original width, thereby obtaining a complete and smooth anatomical structure outline. This can repair the problem of breakpoints that may occur after edge detection; second, adjacent edge points are classified as the same connected domain, and the volume of each connected domain is calculated. If the volume is too small, such as smaller than the size of a tooth tip, it is determined to be noise and deleted, which can fix the problem of misidentifying small noise points as edges.
[0059] Based on CBCT oral image data, threshold segmentation and connected domain analysis are used to further process the data:
[0060] First, a HU threshold (Hounsfield Unit, a standardized unit of measurement used to quantify tissue density in CT images) is set, such as HU>400 for bone tissue, HU=300-800 for cancellous bone, and HU>1500 for cortical bone. To distinguish implants from fillings, a dynamically adjusted threshold is used, such as HU>2000. Voxels (volume elements, the smallest unit in a three-dimensional digital image) are classified, and each voxel is traversed. If its HU value is above the threshold, it is marked as 1 (hard tissue), otherwise it is marked as 0 (soft tissue). For artifacts around implants or fillings, a local threshold is raised, such as HU>500, to avoid misclassification. For low-contrast areas such as the gums, a gradient threshold is used, such as an edge gradient>10HU / mm, to assist in segmentation.
[0061] To remove isolated small areas caused by noise and retain anatomically significant continuous bone structures, a connected domain analysis was performed: first, connected regions were marked, and voxels were considered connected if they were adjacent in three-dimensional space, in terms of faces, edges, and vertices. Starting from a seed point, all connected 1 (hard tissue) voxels were grouped into the same region. Second, the volume was filtered to exclude regions with a volume <10 mm³. Finally, the regional volume was obtained, which is formulated as follows: Regional volume = number of voxels The volume of a single voxel.
[0062] The three-dimensional edge contour obtained by optimizing the 3D surface oral image data is superimposed with the bone structure obtained by optimizing the CBCT oral image data to generate a structural mask that combines surface details and internal anatomy (a binary template commonly used in image processing, medical image analysis, and computer vision to accurately mark or isolate specific target areas in an image).
[0063] S2: Perform rigid registration on the structure mask to obtain alignment parameters, and perform non-rigid registration on the grayscale image to obtain an enhanced grayscale image; input the structure mask, alignment parameters and enhanced grayscale image into the CycleGAN model to obtain a multimodal registered fusion image;
[0064] Specifically, since 3D surface oral imaging data often suffers from unbalanced grayscale distribution due to uneven lighting and enamel reflection, histogram equalization is used to enhance image contrast and specular reflection suppression to eliminate highlight interference:
[0065] The labeling area includes: dividing the multi-scale frequency domain data into 8×8 pixel blocks, calculating the deformation field residual within each block based on the alignment parameters, and determining the area with a deformation error greater than 0.3mm as an uncorrected nonlinear deformation area; at the same time, calculating the sliding window standard deviation of the HU value of each block (a statistical indicator of the degree of HU value fluctuation in the detection area), and identifying areas with a standard deviation that suddenly increases by more than 200 as metal artifact interference areas; the nonlinear deformation area and the metal artifact interference area together constitute the labeling area.
[0066] Specifically, an 8×8 pixel block corresponds to a physical size of 1.6×1.6 mm, covering a single anterior tooth margin or posterior tooth cusp anatomical unit, ensuring that a single block contains all anatomical landmarks. The sliding window standard deviation is a statistical indicator of the degree of fluctuation in the HU value of the detection area.
[0067] Calculate the grayscale histogram (counting the frequency of each grayscale value 0-255 in the image) and CDF (cumulative distribution function, which accumulates the histogram and indicates the proportion of pixels with grayscale values less than or equal to a certain value to the total pixels) for each local area separately. To prevent local over-enhancement, limit the number of pixels of a certain grayscale value in each local area histogram to no more than a threshold, such as Clip Limit = 2.0 (clip limit, a parameter used to limit the strength of histogram equalization in image processing). Pixels exceeding the threshold are clipped and evenly distributed to other grayscale levels. Recalculate the CDF using the restricted histogram, normalize the CDF to 0-255, and generate the grayscale mapping value. The formula is:
[0068] ,
[0069] in, is the grayscale mapping value of grayscale value k, is the cumulative distribution function value of the gray value k, ranging from [0,1], The minimum valid CDF value in the current block. 255 is to convert the [0,1] interval into the actual image grayscale value [0,255]. Round ensures that the grayscale value is an integer.
[0070] The mapping results of adjacent blocks are bilinearly interpolated to eliminate block artifacts and generate the final grayscale image.
[0071] There are deviations in the HU values of different cone-beam CT devices. In order to map the HU values to a unified standard space and ensure data consistency across devices, an affine transformation is performed on the HU value of each voxel: ,in, is the standardized HU value, which represents the calibrated tissue density value that can be compared across devices; a is the slope of the linear transformation, which is used to adjust the proportional relationship of the HU value; is the HU value in the original scan data, which is directly generated by the device; b is the intercept of the linear transformation, which is used to translate the overall offset of the HU value; when the device is scanning, the water phantom is placed next to the patient, and its true HU value is 0. The air area outside the scan has a true HU value of -1000. The HU value of the measured water phantom ( ) and the measured HU value of air ( ) into the equation: , you can get the values of parameters a and b.
[0072] The rigid registration is to obtain the bone contour features based on the structure mask, and then calculate the rigid transformation parameters to obtain the alignment parameters; and obtain the enhanced grayscale image through non-rigid registration with the grayscale image according to the alignment parameters.
[0073] Specifically, because rigid tissues such as bones and teeth are morphologically stable and will not deform in a short period of time, rigid registration is used to eliminate global position differences through translation, rotation, and scaling operations to ensure that the bone contours of 3D surface oral image data and CBCT oral image data are accurately aligned, and rigid transformation parameters (rotation matrix, translation vector, scaling factor) are obtained; after rigid registration eliminates global differences, soft tissues (gingiva, mucosa, tumors) may undergo local displacement due to body position, respiration, treatment deformation or equipment resolution differences. Non-rigid registration aligns these dynamically changing structures through an elastic deformation field (used to describe the local displacement of each voxel in the image, thereby aligning images of different times, modalities or states to the same anatomical space) to obtain a non-rigid deformation field; alignment parameters are composed of rigid transformation parameters and non-rigid deformation fields; the grayscale images of the optimized image data are fused, retaining the bone structure at low frequencies and the soft tissue texture at high frequencies, thereby obtaining enhanced grayscale images.
[0074] The CycleGAN model includes: inputting the structure mask and alignment parameters as structure channel data and the enhanced grayscale image as texture channel data into the CycleGAN model;
[0075] Among them, the structure channel uses a convolutional network to extract anatomical contour features; the texture channel uses a residual network to extract multi-scale texture features; the anatomical contour features are fused with the multi-scale texture features to obtain a multimodal registration fusion image.
[0076] Specifically, the CycleGAN model uses a structural mask and alignment parameters as data for the structure channel, and an enhanced grayscale image as data for the texture channel. Structural consistency and texture fidelity constraints are added to achieve fine-grained control of the generated image quality. The structure channel uses convolutional layers to extract anatomical contour features, focusing on the generation of rigid structures such as bone boundaries and implant paths. Skip connections preserve the geometric information of the original mask to prevent deformation. The texture channel uses a residual network to extract multi-scale texture features, capturing details from macroscopic gingival morphology to microscopic enamel cracks. Within the deep network of the generator, a channel attention module dynamically assigns weights to structural and texture features. For example, at the bone-soft tissue interface, the model enforces structural constraints; in pure soft tissue regions, it prioritizes texture generation. Structural consistency constraints enforce strict matching of the bone or tooth contours in the generated image with the input structural mask to avoid issues such as implant offset or root misalignment. Comparing the alignment parameters between the input and generated data ensures that the model does not introduce unphysical spatial distortions. Texture fidelity constraints include extracting high-level features from the generated image using a pre-trained neural network and comparing it with the target image to ensure the biological plausibility of the texture. Texture coherence is also checked at different scales. For example, when observing microcracks in tooth enamel under a magnifying glass, the generated result must be consistent with the microstructure of the real image. The generated image is then remapped back to the original input domain, ensuring the reversibility and logical consistency of the generation process. This ensures the spatial accuracy of anatomical boundaries and preserves the biological authenticity of soft tissues, ultimately resulting in a multimodal registered fusion image.
[0077] S3: Perform wavelet transform on the multimodal registered fusion image to obtain multi-scale frequency domain data, and annotate the multi-scale frequency domain data to obtain annotated areas;
[0078] The wavelet transform includes: performing three-level wavelet decomposition on the multimodal registered fusion image to obtain low-frequency components and high-frequency sub-bands; identifying artifact areas in the high-frequency sub-bands, determining the artifact interference level based on the artifact areas, and then setting the high-frequency weight coefficient value range and the normal area weight coefficient value range; sharpening the low-frequency components to enhance the clarity of bone contours and suppress artifact area interference; performing inverse wavelet transform fusion on the weight-adjusted high-frequency sub-bands and the sharpened low-frequency components to obtain multi-scale frequency domain data.
[0079] Specifically, wavelet transform is performed on the multimodal registered fusion image, which can effectively suppress artifacts without sacrificing the overall quality. First, the generated image is decomposed into sub-bands of different scales, including low-frequency components (overall structure, such as bone contours) and high-frequency components (details and noise, such as gum texture, artifacts); secondly, the abnormally high response area in the high-frequency sub-band (such as the bright spot corresponding to the metal artifact) is regarded as the artifact area; the artifact interference level is determined based on the artifact area, including level one (2-5 mm, local artifacts caused by small metal objects), level two (5-10 mm, star-shaped artifacts caused by medium metal restorations or high-density materials), level three (10-15 mm, extensive artifacts caused by large metal implants or dense metal fillings), the high-frequency weight is reduced according to the artifact interference level (the weight is used to specifically suppress artifacts or enhance effective information, and can be adjusted through model adaptive adjustment or manual adjustment) to weaken noise and artifacts, with the first-level value range being [0.4, 0.5], the second-level value range being [0.3, 0.4], and the third-level value range being [0.3, 0.4]. In normal areas, the high-frequency weight is maintained or enhanced within the range of [0.8, 1.0] to preserve true details. Then, the low-frequency subband is slightly sharpened to enhance the clarity of the anatomical structure. Finally, the adjusted high-frequency and low-frequency components are fused through inverse wavelet transform to obtain multi-scale frequency domain data.
[0080] Based on multi-scale frequency domain data, the algorithm manually labels key regions to obtain annotated areas. These include artifact interference, nonlinear deformation, diagnostic sensitivity, and feature conflict, ensuring accurate optimization targets. The image is segmented into multiple blocks, and only the annotated regions are subjected to wavelet transformation and weight adjustment, while other regions remain unchanged to reduce memory usage. A progressive optimization approach is employed, first performing low-frequency corrections on the annotated regions to eliminate structural distortions, followed by fine-tuning for high-frequency artifacts to avoid over-processing and loss of detail. This effectively suppresses artifacts while reducing global computational overhead.
[0081] For example, a patient who needs to undergo multiple posterior tooth implants receives both a 3D intraoral scan and a CBCT scan before the operation. Due to the presence of old metal restorations in the patient's mouth, metal artifacts are generated in the CBCT image; at the same time, the patient's slight swallowing during the scan causes the tongue to shift, resulting in nonlinear deformation differences between the gingival contour of the 3D scan and the bone structure of the CBCT. Using the method of Example 1, the complex deformation problem caused by patient movement is solved through the synergistic effect of rigid registration and non-rigid registration; the dual-channel CycleGAN uses the deep bone information of CBCT to reconstruct the key anatomical structures obscured by artifacts while retaining the surface accuracy of the 3D scan. However, due to the lack of hierarchical registration and nonlinear deformation correction, the traditional method has a risk of implant perforation of the nerve canal as high as 15% in similar cases. Example 1 reduces this risk to less than 2%.
[0082] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:
[0083] This application uses 3D Canny edge detection to extract contours, combined with CBCT dynamic thresholding to generate structural masks to address the misalignment between 3D scanned surface data and CBCT anatomical landmarks. Rigid and non-rigid registration are used to optimize registration efficiency. The CycleGAN model improves training efficiency and result accuracy through structural and texture channels. Wavelet transforms are used to achieve effective artifact suppression. Existing technologies rely on a single affine transformation for global alignment and are unable to handle the elastic deformation of soft tissue surrounding implants.
[0084] Example 2: In Example 1, CycleGAN's structural consistency constraints may result in loss of texture detail at complex anatomical interfaces (such as the crown and gums). Furthermore, the design does not explicitly separate global and local features, which can lead to structural distortion. To address this issue, this example further improves Example 1.
[0085] S4: Feature decoupling based on multimodal spatiotemporal registration dataset to obtain hierarchical features;
[0086] The feature decoupling step includes: extracting the global structural features of the structure mask through a convolutional neural network, constraining its symmetry error with the gold standard to be less than or equal to 0.3mm; extracting the local detail features of the grayscale image through a residual network, maintaining the edge gradient amplitude at least 90% of the original image; using a spatial attention mechanism to dynamically fuse the global structural features and local detail features to obtain hierarchical features; and providing a unified spatial benchmark for time series images to reduce the negative impact of artifacts on fusion quality.
[0087] like Figure 2As shown, the features are decoupled into local detail features and global structural features, collectively referred to as hierarchical features. Through independent optimization and dynamic fusion, the collaborative expression ability of multimodal data is improved. Based on the data of structural mask and rigid registration, a convolutional neural network (a deep learning model specifically for processing spatially structured data) is used to extract global structural features. The attention mechanism focuses on key anatomical landmarks to obtain vectors of global structural features, such as bone density distribution and symmetry. Based on enhanced grayscale images and annotated areas, a residual network is used to extract high-frequency details, and adaptive filtering is used to suppress noise to obtain tensors of local detail features, such as gingival texture gradient and microcrack morphology. Global structural features and local detail features are optimized independently. The Dice coefficient of global features and gold standards (such as CBCT bone segmentation) is forced to be ≥0.95. The symmetry error between the generated image and the gold standard is ≤0.3mm to maintain the consistency of the anatomical structure, thereby completing global structural optimization. By comparing the texture distribution of the generated image and the real image, the microstructure of the soft and hard tissues is ensured to conform to the anatomical laws. The three indicators of brightness, contrast, and structural similarity are comprehensively used to quantify the difference between the generated image and the real image. By maximizing the edge gradient amplitude of the generated image, the sharpness of key boundaries such as enamel-gingiva and bone-soft tissue is ensured. In the artifact area, gradient loss can reduce edge diffusion and maintain the sharp contours of the anatomical structure, thereby completing local detail optimization.
[0088] Generate a spatial attention map based on global features to identify key anatomical areas; generate a channel attention vector based on local features to enhance detail areas; then dynamically fuse global features with local features to obtain the final feature representation, which is formulated as follows:
[0089] ,
[0090] in, is the fused feature representation, is a global feature, is a local feature, is a weight coefficient used to control the contribution ratio of global features to local features. Its value range is [0.6, 0.8]. The closer the value is to 0.8, the more emphasis is placed on global features, and vice versa. For example, lowering α in the artifact area can reduce the amplification of noise by the global structure and adapt to the needs of different clinical scenarios.
[0091] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:
[0092] This application adopts feature decoupling, independently optimizes global structure and local details, and dynamically fuses weights to ensure that global interference is reduced in key areas such as artifacts and texture authenticity is improved; combined with the edge gradient maximization constraint in local detail optimization, effective high-frequency information is targetedly enhanced, and at the same time, the global feature weight is actively reduced in the artifact area through the spatial attention mechanism to reduce noise amplification; convolutional networks are used to extract global features and residual networks to extract local features, and dynamic fusion of spatial attention is used to solve the deformation accumulation problem at the implant-soft tissue junction.
[0093] Example 3: In Example 2, feature decoupling only addresses static data from a single point in time, failing to consider the temporal correlation of multiple examination images. This results in an inability to capture dynamic changes such as implant displacement or bone resorption. Non-rigid registration also addresses soft tissue deformation within a single scan, lacking a correction mechanism for progressive anatomical changes during treatment. To address this issue, this example further improves Example 2.
[0094] Step S3 further includes: using the multimodal oral image data collected at different times as time-series image data; performing spatiotemporal alignment based on the marked areas and the time-series image data to obtain a multimodal spatiotemporal registration data set;
[0095] The spatiotemporal alignment step includes: obtaining a non-rigid deformation field of soft tissue deformation through non-rigid registration calculation, and obtaining composite alignment parameters based on the alignment parameters and the non-rigid deformation field; based on the composite alignment parameters, aligning the time series image data to the static reference space to form a rigid registration grouping, establishing a unified spatial coordinate system through the elastic deformation field, and the rigid registration error tolerance is less than or equal to 0.1mm; constructing a non-rigid compensation grouping based on feature data extracted from the annotated area, and using a recurrent neural network to predict the local progressive deformation field, with a maximum deformation amount less than or equal to 2.5mm; achieving precise alignment of bones and soft tissues across time point images through a hierarchical registration strategy, wherein the rigid registration grouping maintains the consistency of the spatial position of the implant, and the non-rigid compensation grouping corrects gingival deformation, and outputs a multimodal spatiotemporal registration dataset.
[0096] Specifically, the multimodal oral imaging data collected at different times are used as time-series imaging data. Based on the annotated regions and time-series imaging data, spatiotemporal alignment (synchronizing data or events in different time series to ensure that they are aligned on the time axis) is used to divide the data into rigid registration groups and non-rigid compensation groups.
[0097] The non-rigid deformation field of soft tissue deformation is obtained by non-rigid registration. The composite alignment parameter is obtained based on the alignment parameter and the non-rigid deformation field, and its formula is:
[0098] ,
[0099] in, is the coordinate of a point, is the composite alignment parameter, is the alignment parameter, is a non-rigid deformation field.
[0100] Based on the composite alignment parameters, all time series data are registered to the static reference space. Data that eliminate body position differences are grouped as rigid registration data to establish a unified spatial coordinate system for all time series data. The error tolerance is ≤0.1mm to ensure that the anatomical structures are comparable in spatial dimensions. Data that have been deformed by the deformation field are grouped as non-rigid compensation data. Under the unified spatial reference, local deformation in dynamically changing areas is allowed, with a maximum deformation of ≤2.5mm. Rigid registration eliminates global offsets, and non-rigid compensation retains local dynamics, forming a hierarchical grouping optimization.
[0101] Based on the registered data, the patient's images from multiple examinations are used as dynamic time series data to capture the dynamic changes in their structure. A single high-precision CT image is used as static structural data to provide a spatial benchmark to ensure the anatomical rationality of cross-temporal data alignment. A recurrent neural network (a neural network specifically designed to process sequence data, whose core feature is the ability to "memorize" historical information and capture the time dependency in the data) is used to capture dynamic time series features and obtain time encoding vectors (such as bone resorption rate and implant displacement). A convolutional neural network is used to capture static structural features and obtain spatial encoding vectors (such as bone density distribution and symmetry parameters). The dynamic time series features are fused with the static structural features and weighted to obtain the formula:
[0102] ,
[0103] in, represents the fusion feature, is the static structural feature, is the dynamic time series feature, is spatiotemporal attention, with a value range of [0,1]. The closer the value is to 1, the more static structural features are used to stabilize the anatomical region and strengthen the constraints of the reference structure; the closer the value is to 0, the more dynamic temporal features are used to dominate, and it is used for regions with significant changes, focusing on temporal evolution features.
[0104] A hierarchical registration strategy is used to achieve precise alignment of bones and soft tissues across time points. Rigid registration grouping maintains the consistency of implant spatial position, while non-rigid compensation grouping corrects gingival deformation, outputting a multimodal spatiotemporal registration dataset.
[0105] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:
[0106] This application adopts spatiotemporal alignment, divides time series images into rigid groups and non-rigid groups, extracts dynamic time series features through recurrent neural networks, and realizes precise alignment of cross-time data; sets non-rigid compensation groups to allow dynamic adjustment of local deformation fields, and realizes precise compensation of progressive deformation through cumulative calculation of deformation fields; establishes a layered registration strategy, rigid grouping forces errors to ensure baseline consistency, and non-rigid grouping establishes regional deformation associations through graph convolutional networks to suppress error propagation.
[0107] Example 4: In Example 3, spatiotemporal attention is a global parameter, unable to distinguish the differences in sensitivity between bone and soft tissue regions to CT and 3D scan data, resulting in excessive weakening of CT features in the implant area. The local deformation field of the non-rigid compensation grouping fails to consider the motion consistency of adjacent regions, potentially leading to conflicting deformation directions between the gums and crowns. The artifact suppression predicted by the recurrent neural network applies to the entire image, resulting in residual artifacts around metal implants. To address these issues, this example further improves Example 3.
[0108] Step S4 further includes: performing regional adaptive fusion based on the hierarchical features, the structural mask and the grayscale image to obtain a multimodal optimized fused image;
[0109] The regional adaptive fusion steps include: dividing the multimodal spatiotemporal registration dataset into [10, 15] functional regions based on hierarchical features; calculating spatial proximity edges based on structural masks; calculating feature similarity edges based on multi-scale frequency domain data; establishing a regional association matrix based on spatial proximity edges and feature similarity edges; and dynamically allocating modal weights using an attention network based on the regional association matrix to obtain a multimodal optimized fusion image.
[0110] Specifically, based on hierarchical features, structural masks and grayscale images, geometric features such as curvature and density distribution, and texture features such as gray-level co-occurrence matrix and local binary pattern are extracted from the hierarchical features;
[0111] The K-means clustering algorithm is used to construct a graph structure based on feature similarity and divide it into K regions, such as k=10, with a value range of [10,15]. The regional label map is output to mark the region to which each voxel belongs. This achieves the effect of automatically dividing functional regions according to image features and realizes unsupervised region segmentation.
[0112] Specifically, based on the divided regions, inter-regional feature dependencies are modeled. Each region is treated as a node, and feature vectors, such as texture, volume, and spatial position, are used as node features. Edge connections are based on two independent conditions; an edge is established if either condition is met. Spatial proximity edges are calculated using Euclidean distance with a threshold of 5mm, and an edge is established if the distance between the geometric centers of two regions is ≤5mm. Feature similarity edges are calculated using cosine similarity with a threshold of 0.7, and an edge is established if the cosine similarity of the feature vectors of two regions is ≥0.7. Spatial proximity edges and feature similarity edges are input into a graph convolutional network, and the output is a region association matrix.
[0113] Specifically, based on the regional correlation matrix, regional adaptive fusion is performed to implement a differentiated fusion strategy and optimize local accuracy. First, the regional importance weight is calculated through the attention mechanism, and its formula is:
[0114] ,
[0115] in, For the region The weight of For the region A set of connected neighbor regions; is the element of the regional correlation matrix, which represents the correlation strength between regions i and j; is a fully connected layer, which transforms the feature vector of region j into Map to scalar; is a normalization function that maps weights to [0,1] and sums to 1.
[0116] The modality weight is dynamically adjusted according to regional characteristics. For example, the CT weight of the bone area is high, and the value range is [0.7, 0.9]; the 3D scan weight of the soft tissue area is high, and the value range is [0.6, 0.8]. The formula of regional importance weight is integrated to obtain the formula:
[0117] ,
[0118] in, is the pixel value of the fused image at position x; M is the number of modalities; is the contribution weight of mode m in region i; is the pixel value of modality m at position x; by independently calculating the weight of each region, differential fusion is achieved to obtain a multi-modal optimized fusion image.
[0119] The technical solutions in the above embodiments of the present application have at least the following technical effects or advantages:
[0120] This application uses K-means clustering to divide the region, enforces CT weights in the bone tissue region, and accurately preserves the internal trabecular structure; constructs a regional association matrix, and avoids anatomical structure dislocation through dual constraints of spatial proximity and feature similarity; automatically identifies implant areas with HU>2000 in the regional segmentation stage, and specifically reduces the weight of high-frequency components to improve the artifact elimination rate.
[0121] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A fusion processing method based on multimodal oral image data, characterized in that: include: S1: Acquire multimodal oral image data, and obtain a structure mask and a grayscale image based on the multimodal oral image data; the multimodal oral image data includes 3D surface oral image data and CBCT oral image data; S2: Perform rigid registration on the structure mask to obtain alignment parameters, and perform non-rigid registration on the grayscale image to obtain an enhanced grayscale image; input the structure mask, alignment parameters and enhanced grayscale image into the CycleGAN model to obtain a multimodal registered fusion image; S3: Perform wavelet transform on the multimodal registered fusion image to obtain multi-scale frequency domain data, and annotate the multi-scale frequency domain data to obtain annotated areas; The multimodal oral imaging data collected at different times are used as time series imaging data; Perform spatiotemporal alignment based on the annotated regions and time-series image data to obtain a multimodal spatiotemporal registration dataset; S4: Feature decoupling based on multimodal spatiotemporal registration dataset to obtain hierarchical features; Based on hierarchical features, structural masks and grayscale images, regional adaptive fusion is performed to obtain a multimodal optimized fused image.
2. A fusion processing method based on multimodal oral image data according to claim 1, characterized in that: The generation of the structure mask includes: 3D Canny edge detection and topology restoration are used on 3D surface oral image data to obtain three-dimensional edge contours; dynamic threshold segmentation and connected domain analysis are used on CBCT oral image data to obtain bone structure segmentation results; the three-dimensional edge contours and bone structure segmentation results are superimposed to generate a structural mask.
3. The fusion processing method based on multimodal oral image data according to claim 2, characterized in that: The 3D Canny edge detection includes: Three-dimensional directional gradient detection is used on multimodal oral imaging data to calculate the brightness change intensity in the X, Y, and Z axis directions to obtain the gradient amplitude map and direction map; Based on the gradient magnitude map and direction map, the edge is refined to a single pixel width through non-maximum suppression to obtain a refined edge map; The gradient threshold is set based on the edge map to retain the boundary features between the crown edge and the gingiva and obtain the three-dimensional edge contour.
4. The fusion processing method based on multimodal oral image data according to claim 1, characterized in that: The rigid registration is to obtain the bone contour features based on the structure mask, and then calculate the rigid transformation parameters to obtain the alignment parameters; and obtain the enhanced grayscale image through non-rigid registration with the grayscale image according to the alignment parameters.
5. The fusion processing method based on multimodal oral image data according to claim 1, characterized in that: The CycleGAN model includes: The structure mask and alignment parameters are used as the structure channel data, and the enhanced grayscale image is used as the texture channel data to input into the CycleGAN model; Among them, the structure channel uses a convolutional network to extract anatomical contour features; the texture channel uses a residual network to extract multi-scale texture features; the anatomical contour features are fused with the multi-scale texture features to obtain a multimodal registration fusion image.
6. The fusion processing method based on multimodal oral image data according to claim 1, characterized in that: The wavelet transform comprises: Perform three-level wavelet decomposition on the multimodal registered fusion image to obtain low-frequency components and high-frequency sub-bands; Identify the artifact area in the high-frequency sub-band, determine the artifact interference level based on the artifact area, and then set the value range of the high-frequency weight coefficient and the value range of the normal area weight coefficient; Sharpen the low-frequency components, enhance the clarity of bone contours, and suppress interference from artifact areas; The weight-adjusted high-frequency subband is fused with the sharpened low-frequency component by inverse wavelet transform to obtain multi-scale frequency domain data.
7. The fusion processing method based on multimodal oral image data according to claim 1, characterized in that: The marked area includes: The multi-scale frequency domain data was divided into 8×8 pixel blocks. The deformation field residual within each block was calculated based on the alignment parameters. The area with a deformation error greater than 0.3 mm was identified as an uncorrected nonlinear deformation area. The sliding window standard deviation of the HU value of each block was calculated. The area with a standard deviation exceeding 200 was identified as a metal artifact interference area. The nonlinear deformation area and the metal artifact interference area together constituted the annotation area.
8. The fusion processing method based on multimodal oral image data according to claim 1, characterized in that: The feature decoupling step includes: The global structural features of the structure mask are extracted through a convolutional neural network, and the symmetry error with the gold standard is constrained to be less than or equal to 0.3mm. The local detail features of the grayscale image are extracted through the residual network, keeping the edge gradient amplitude no less than 90% of the original image; the spatial attention mechanism is used to dynamically fuse global structural features and local detail features to obtain hierarchical features; a unified spatial benchmark is provided for time series images to reduce the negative impact of artifacts on fusion quality.
9. The fusion processing method based on multimodal oral image data according to claim 1, characterized in that: The spatiotemporal alignment step comprises: The non-rigid deformation field of the soft tissue deformation is obtained by non-rigid registration calculation, and the composite alignment parameter is obtained according to the alignment parameter and the non-rigid deformation field; Based on composite alignment parameters, the time-series image data are registered to the static reference space to form a rigid registration grouping. A unified spatial coordinate system is established through the elastic deformation field, and the rigid registration error tolerance is less than or equal to 0.1mm. Based on the feature data extracted from the annotated area, a non-rigid compensation grouping is constructed, and a recurrent neural network is used to predict the local progressive deformation field, with a maximum deformation of less than or equal to 2.5mm. A hierarchical registration strategy is used to achieve precise alignment of bones and soft tissues across time points. The rigid registration grouping maintains the consistency of the spatial position of the implant, and the non-rigid compensation grouping corrects gingival deformation, outputting a multimodal spatiotemporal registration dataset.
10. The fusion processing method based on multimodal oral image data according to claim 1, characterized in that: The regional adaptive fusion step includes: Based on hierarchical features, the multimodal spatiotemporal registration dataset is segmented into [10, 15] functional regions; Compute spatial proximity edges based on structure masks; Calculate feature similarity edges based on multi-scale frequency domain data; According to the spatial proximity edge and feature similarity edge, a regional correlation matrix is established; Based on the regional correlation matrix, the attention network is used to dynamically allocate modal weights to obtain a multimodal optimized fusion image.
Citation Information
Patent Citations
A method for fusion and processing of multimodal oral image data
CN118537699B
Medical image registration method and device
CN113096166A
System and method for cardiac segmentation in mr-cine data using inverse consistent non-rigid registration
US20110081066A1