Fusion processing method based on multi-modal oral cavity image data

The method uses 3D Canny edge detection and non-rigid registration with a CycleGAN model to address complex deformations in multi-modal oral image data, achieving precise alignment and integration with enhanced clinical utility and artifact suppression.

CN120318093AActive Publication Date: 2025-07-15CENT SOUTH UNIV

Patent Information

Application Number
CN202510795584.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-15
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In the prior art, the local registration error and complex deformation adaptability of multimodal oral image data are insufficient, resulting in low multimodal fusion accuracy and efficiency.

Method used

A strategy of combining rigid registration and non-rigid registration is adopted, and a structural mask is generated by combining 3D Canny edge detection and dynamic threshold segmentation technology. Multimodal image fusion is used to perform multimodal image fusion, and the fusion process is optimized through wavelet transformation and feature decoupling technology.

Benefits of technology

It significantly improves the accuracy and efficiency of multimodal fusion images, effectively suppresses metal artifacts, maintains the spatial position accuracy of the implant and restores the biological characteristics of the gingival texture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318093A_ABST
    Figure CN120318093A_ABST
Patent Text Reader

Abstract

The invention discloses a fusion processing method based on multi-modal oral image data, and relates to the technical field of image data processing, and the method comprises the steps: carrying out the rigid registration of a structural mask, obtaining an alignment parameter, and carrying out the non-rigid registration of a gray image, and obtaining an enhanced gray image; inputting the structure mask, the alignment parameter and the enhanced gray level image into a CycleGAN model to obtain a multi-modal registration fusion image; performing wavelet transform on the multi-modal registration fusion image to obtain multi-scale frequency domain data; taking data acquired at different times as time sequence image data; performing space-time alignment to obtain a multi-modal space-time registration data set; performing feature decoupling based on the multi-modal space-time registration data set to obtain layered features; and based on the hierarchical features, the structure mask and the gray level image, carrying out regional adaptive fusion to obtain a multi-modal optimization fusion image. The technical effect of improving the multi-modal fusion precision and efficiency is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image data processing, and particularly relates to a fusion processing method for multi-modal oral image data. Background Art

[0002] In oral medical diagnosis, clinicians usually need to rely on various types of image data, such as X-rays, CT scans (such as CBCT), and oral endoscope images, etc., to comprehensively evaluate the oral health status of patients. However, different modalities of image data have their own advantages and disadvantages, and often cannot provide complete diagnostic information when used alone.

[0003] The Chinese invention patent with the publication number CN118537699B provides a method for fusing and processing multi-modal oral image data. Initial multi-modal oral image acquisition is performed through a 3D oral scanner device and a cone beam computed tomography imaging device to generate multi-modal oral image data; regional alignment of modal differences is performed on the multi-modal oral image data to generate aligned multi-modal oral image data; multi-modal oral image matching analysis is performed based on the aligned multi-modal oral image data to generate multi-modal oral image matching data; mapping fusion processing of multi-modal oral images is performed on the multi-modal oral image matching data based on a preset CycleGAN model to generate fused oral image data; fused oral image feature mining is performed based on the fused oral image data to generate fused oral image feature data.

[0004] Although affine transformation is used to achieve spatial alignment of multi-modal images, there are complex non-linear deformations in the internal structure of the oral cavity, and affine transformation cannot completely correct such deformations, which may lead to local registration errors and insufficient adaptability to complex deformations. Summary of the Invention

[0005] By providing a fusion processing method for multi-modal oral image data, the present application solves the problems of local registration errors and insufficient adaptability to complex deformations in the prior art, and achieves the technical effect of improving the accuracy and efficiency of multi-modal fusion.

[0006] The present application provides a fusion processing method for multi-modal oral image data, and the method includes: S1: Obtain multi-modal oral image data, and obtain a structure mask and a grayscale image based on the multi-modal oral image data; the multi-modal oral image data includes 3D surface oral image data and CBCT oral image data; S2: Perform rigid registration on the structure mask to obtain alignment parameters, and perform non-rigid registration on the grayscale image to obtain an enhanced grayscale image; input the structure mask, the alignment parameters, and the enhanced grayscale image into the CycleGAN model to obtain a multi-modal registered and fused image; S3: Perform wavelet transform on the multi-modal registration and fusion image to obtain multi-scale frequency domain data, and label the multi-scale frequency domain data to obtain labeled regions; Use the multi-modal oral image data collected at different times as time-series image data; Perform spatio-temporal alignment based on the labeled regions and the time-series image data to obtain a multi-modal spatio-temporal registration data set; S4: Decouple features based on the multi-modal spatio-temporal registration data set to obtain hierarchical features; perform region adaptive fusion based on the hierarchical features, the structure mask, and the grayscale image to obtain a multi-modal optimized fusion image.

[0007] Further, the generation of the structure mask includes: Perform 3D Canny edge detection and topological repair on the 3D surface oral image data to obtain a three-dimensional edge contour; perform dynamic threshold segmentation and connected component analysis on the CBCT oral image data to obtain a bone structure segmentation result; superimpose the three-dimensional edge contour and the bone structure segmentation result to generate a structure mask.

[0008] Further, the 3D Canny edge detection includes: Perform three-dimensional directional gradient detection on the multi-modal oral image data, calculate the intensity of brightness change in the X, Y, and Z axis directions to obtain a gradient magnitude map and a direction map; Based on the gradient magnitude map and the direction map, refine the edge to a single-pixel width through non-maximum suppression to obtain a refined edge map; Set a gradient threshold based on the edge map, retain the tooth crown edge and gingival junction features to obtain a three-dimensional edge contour.

[0009] Further, the rigid registration is to obtain bone contour features based on the structure mask, and then calculate the rigid transformation parameters to obtain alignment parameters; perform non-rigid registration based on the alignment parameters and the grayscale image to obtain an enhanced grayscale image.

[0010] Further, the CycleGAN model includes: Input the structure mask and the alignment parameters as structure channel data, and the enhanced grayscale image as texture channel data into the CycleGAN model; Among them, the structure channel uses a convolutional network to extract anatomical contour features; the texture channel uses a residual network to extract multi-scale texture features; fuse the anatomical contour features and the multi-scale texture features to obtain a multi-modal registration and fusion image.

[0011] Further, the wavelet transform optimization includes: Perform three-level wavelet decomposition on the multi-modal registration and fusion image to obtain a low-frequency component and high-frequency sub-bands; Identify the artifact regions in the high-frequency subbands, determine the artifact interference level based on the artifact regions, and then set the value range of the high-frequency weight coefficient and the value range of the normal region weight coefficient; Sharpen the low-frequency components to enhance the clarity of the bone contour and suppress the interference of the artifact regions; Perform inverse wavelet transform fusion on the high-frequency subband with adjusted weights and the sharpened low-frequency components to obtain multi-scale frequency domain data.

[0012] Further, the labeled regions include: Divide the multi-scale frequency domain data into 8×8 pixel blocks, calculate the deformation field residuals within each block based on the alignment parameters, and determine the regions with deformation errors greater than 0.3 mm as uncorrected non-linear deformation regions; at the same time, calculate the sliding window standard deviation of the HU values of each block, and identify the regions with a sudden increase in the standard deviation greater than 200 as metal artifact interference regions; the non-linear deformation regions and the metal artifact interference regions together constitute the labeled regions.

[0013] Further, the feature decoupling step includes: Extract the global structural features of the structural mask through a convolutional neural network, and constrain the symmetry error with the gold standard to be less than or equal to 0.3 mm; Extract the local detail features of the grayscale image through a residual network, and keep the edge gradient amplitude not less than 90% of the original image; adopt a spatial attention mechanism to dynamically fuse the global structural features and the local detail features to obtain hierarchical features; provide a unified spatial reference for the temporal images to reduce the negative impact of artifacts on the fusion quality.

[0014] Further, the spatio-temporal alignment step includes: Calculate the non-rigid deformation field of the soft tissue deformation through non-rigid registration, and obtain the composite alignment parameters according to the alignment parameters and the non-rigid deformation field; Based on the composite alignment parameters, register the temporal image data to the static reference space to form a rigid registration group, establish a unified spatial coordinate system through the elastic deformation field, and the rigid registration error tolerance is less than or equal to 0.1 mm; based on the feature data extracted from the labeled regions, construct a non-rigid compensation group, and use a recurrent neural network to predict the local progressive deformation field, with the maximum deformation amount less than or equal to 2.5 mm; achieve precise alignment of the bones and soft tissues of the images across time points through a hierarchical registration strategy, where the rigid registration group maintains the consistency of the spatial position of the implant, and the non-rigid compensation group corrects the gingival deformation, and outputs a multi-modal spatio-temporal registration data set.

[0015] Further, the region adaptive fusion step includes: Based on the hierarchical features, divide the multi-modal spatio-temporal registration data set into [10, 15] functional regions; Calculate the spatial proximity edges based on the structural mask; Calculate the feature similarity edges based on multi-scale frequency domain data; Establish a region association matrix according to the spatial proximity edges and the feature similarity edges; Based on the region association matrix, use an attention network to dynamically allocate modal weights to obtain a multi-modal optimized fusion image.

[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages: By adopting a strategy that combines rigid registration and non-rigid registration, and cooperating with 3D Canny edge detection and dynamic threshold segmentation technology, the problem of complex deformation correction between multi-modal images is effectively solved; by constructing a dual-channel CycleGAN model, while maintaining the spatial position accuracy of the implant, the biological characteristics of the gingival texture are restored, significantly improving the clinical usability of the multi-modal fusion image; by adopting a local wavelet optimization and feature decoupling fusion mechanism, metal artifacts are effectively suppressed, and at the same time, the calculation efficiency is increased through a region adaptive fusion strategy; by establishing a spatio-temporal alignment and hierarchical feature association model, the technical effects of improving the accuracy and efficiency of multi-modal fusion are achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of a fusion processing method for multi-modal oral image data in an embodiment of the present invention; Figure 2 It is a flowchart of feature decoupling in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To facilitate the understanding of the present invention, the present application will be described more comprehensively with reference to the relevant drawings; the preferred embodiments of the present invention are shown in the drawings, however, the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art belonging to the technical field of the present invention; the terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0020] Embodiment 1: As Figure 1 shown, a fusion processing method for multi-modal oral image data, the method includes: S1: Obtain multi-modal oral imaging data, and obtain a structure mask and a grayscale image based on the multi-modal oral imaging data; the multi-modal oral imaging data includes 3D surface oral imaging data and CBCT oral imaging data; Specifically, the 3D surface oral imaging data is obtained by a 3D oral scanner capturing three-dimensional geometric and texture information of the inner surface of the oral cavity through optical imaging technology; the cone beam CT device obtains CBCT (Cone Beam Computed Tomography) oral imaging data by the collaborative work of a cone X-ray beam and a two-dimensional flat panel detector to collect multi-angle projection data.

[0021] The generation of the structure mask includes: performing 3D Canny edge detection and topology repair on the 3D surface oral imaging data to obtain a three-dimensional edge contour; performing dynamic threshold segmentation and connected component analysis on the CBCT oral imaging data to obtain a bone structure segmentation result; superimposing the three-dimensional edge contour and the bone structure segmentation result to generate a structure mask.

[0022] The 3D Canny edge detection includes: performing three-dimensional directional gradient detection on the multi-modal oral imaging data, calculating the intensity of brightness change in the X, Y, and Z axis directions to obtain a gradient magnitude map and a direction map; based on the gradient magnitude map and the direction map, refining the edge to a single-pixel width through non-maximum suppression to obtain a refined edge map; setting a gradient threshold based on the edge map to retain the tooth crown edge and gingival junction features to obtain a three-dimensional edge contour.

[0023] Specifically, based on the 3D surface oral imaging data, use 3D Canny edge detection (for detecting structural boundaries in volume data) and topology repair to further process the data; Using 3D Canny edge detection, first remove the tiny noise points (noise reduction processing) in the scanned data, such as the fluctuations caused by device noise or patient movement, and perform overall smoothing on the three-dimensional data using a "blur filter"; this blur is not uniform, but assigns higher weights to the nearby areas to ensure that the edges are not overly blurred; then, calculate the gradient, detect the intensity of brightness change in the X, Y, and Z directions respectively, combine the slopes in the three directions to obtain a gradient magnitude map and a direction map, mark all possible edge positions, and find the steep areas in the data through the brightness change in the three-dimensional direction, such as the tooth edge and the gingival junction; secondly, refine through non-maximum suppression, check whether each point is the steepest position in that direction along the gradient direction. If so, retain it as an edge point; otherwise, eliminate it, so as to refine the thick-line edge into a single-pixel width, avoid overlap or blur, and obtain a refined edge map; finally, set the gradient threshold, and its formula is: , where, is the high / low threshold, is the global maximum value of the gradient amplitude of all data, is the proportional coefficient, the high threshold is [0.6, 0.8], and the low threshold is [0.2, 0.4]; According to the gradient threshold, the effective edges are screened. Areas with very obvious brightness changes, such as the junction between the crown and the gums, are directly retained; areas with slight brightness changes, such as the enamel surface texture, are retained if they are connected to strong edges, otherwise they are treated as noise and removed. Finally, a continuous, high-precision three-dimensional edge contour is output; To address the problems that may arise after 3D Canny edge detection, topology repair is used for optimization: first, the edge is expanded outward by 1-2 pixels to connect the broken areas, and then the expanded edge is contracted inward to restore the original width, thereby obtaining a complete and smooth anatomical structure outline. This can repair the problem of breakpoints that may occur after edge detection. Secondly, adjacent edge points are classified into the same connected domain, and the volume of each connected domain is calculated. If the volume is too small, such as being smaller than the size of a tooth tip, it is determined to be noise and deleted. This can repair the problem of misidentifying small noise points as edges.

[0024] Based on CBCT oral image data, the data is further processed using threshold segmentation and connected domain analysis: First, set the HU threshold (Hounsfield Unit, which is a standardized measurement unit used to quantify tissue density in CT images), such as bone tissue HU>400, cancellous bone HU=300-800, cortical bone HU>1500; in order to distinguish implants or fillings, dynamically adjust the threshold, such as HU>2000; classify voxels (volume elements, the smallest unit in a three-dimensional digital image), traverse each voxel, and if its HU value is higher than the threshold, it is marked as 1 (hard tissue), otherwise it is marked as 0 (soft tissue); for artifacts around implants or fillings, use local threshold enhancement, such as HU>500, to avoid misjudgment; for low-contrast areas such as gums, use gradient thresholds, such as edge gradient>10HU / mm, to assist segmentation; In order to remove isolated small areas caused by noise and retain continuous bone structures with anatomical significance, connected domain analysis was adopted: first, the connected areas were marked, and the voxels with adjacent faces, edges, and vertices in three-dimensional space were considered connected; starting from a seed point, all connected 1 (hard tissue) voxels were classified into the same area; secondly, the volume was filtered to eliminate areas with a volume <10 mm³; finally, the regional volume was obtained, and its formula was: Regional volume = number of voxels The volume of a single voxel.

[0025] The three-dimensional edge contour optimized from 3D surface oral image data is superimposed on the bone structure optimized from CBCT oral image data to generate a structure mask that combines surface details and internal anatomy (a binary template commonly used in image processing, medical image analysis, and computer vision for precisely marking or isolating specific target regions in an image).

[0026] S2: Rigid registration is performed on the structure mask to obtain alignment parameters, and non-rigid registration is performed on the grayscale image to obtain an enhanced grayscale image; the structure mask, alignment parameters, and enhanced grayscale image are input into the CycleGAN model to obtain a multi-modal registration and fusion image; Specifically, due to reasons such as uneven illumination and enamel reflection in 3D surface oral image data, histogram equalization is adopted to enhance the image contrast, and specular reflection suppression is used to eliminate highlight interference: The marked area includes: dividing the multi-scale frequency domain data into 8×8 pixel blocks, calculating the deformation field residual within each block based on the alignment parameters, and determining the area with a deformation error greater than 0.3 mm as an uncorrected non-linear deformation area; at the same time, calculating the sliding window standard deviation of the HU values of each block (a statistical index for detecting the fluctuation degree of HU values in the detection area), and identifying the area with a sudden increase in the standard deviation greater than 200 as a metal artifact interference area; the non-linear deformation area and the metal artifact interference area together constitute the marked area.

[0027] Specifically, an 8×8 pixel block corresponds to a physical size of 1.6×1.6 mm, which can cover the edge of a single anterior tooth or the anatomical unit of a posterior tooth cusp, ensuring that a single block contains complete anatomical landmarks. The sliding window standard deviation is a statistical index for detecting the fluctuation degree of HU values in the detection area.

[0028] The grayscale histogram (counting the frequency of each grayscale value from 0 - 255 in the image) and CDF (cumulative distribution function, which accumulates the histogram and represents the proportion of pixels with grayscale values less than or equal to a certain value to the total number of pixels) are calculated separately for each local area; to prevent local over-enhancement, the number of pixels with a certain grayscale value in the histogram of each local area is restricted not to exceed a threshold, such as Clip Limit = 2.0 (clipping limit, a parameter used in image processing to limit the intensity of histogram equalization); the pixels exceeding the threshold are clipped and evenly distributed to other gray levels; the CDF is recalculated using the restricted histogram, and the CDF is normalized to 0 - 255 to generate a grayscale mapping value, with the formula: , where, is the grayscale mapping value of grayscale value k, is the cumulative distribution function value of grayscale value k, with a range of [0,1], is the minimum valid CDF value in the current block, 255 is used to convert the [0, 1] interval to the actual image gray value [0, 255], and round ensures that the gray value is an integer; Perform bilinear interpolation on the mapping results of adjacent blocks to eliminate blocky artifacts and generate the final gray image.

[0029] There are deviations in HU values among different cone-beam CT devices. To map HU values to a unified standard space and ensure cross-device data consistency, an affine transformation is performed on each voxel's HU value: , where is the standardized HU value, representing the calibrated tissue density value that can be compared across devices; a is the slope of the linear transformation, used to adjust the proportional relationship of HU values; is the HU value in the original scan data, directly generated by the device; b is the intercept of the linear transformation, used to translate the overall offset of HU values; when the water phantom is placed beside the patient during device scanning, its true HU value is 0, and for the air region outside the scan, its true HU value is -1000. Substitute the measured HU value of the water phantom ( ) and the measured HU value of the air ( ) into the equation: , and the values of parameters a and b can be obtained.

[0030] The rigid registration is to obtain the bone contour features based on the structure mask, and then calculate the rigid transformation parameters to obtain the alignment parameters; according to the alignment parameters and the gray image, perform non-rigid registration to obtain the enhanced gray image.

[0031] Specifically, since rigid tissues such as bones and teeth are morphologically stable in a short period of time and do not deform, rigid registration is used to eliminate the global position differences through translation, rotation, and scaling operations to ensure the precise alignment of the bone contours of the 3D surface oral image data and the CBCT oral image data, and obtain the rigid transformation parameters (rotation matrix, translation vector, scaling factor); after rigid registration eliminates the global differences, soft tissues (gingiva, mucosa, tumor) will have local displacements due to body position, breathing, treatment deformation, or device resolution differences. Non-rigid registration aligns these dynamically changing structures through an elastic deformation field (used to describe the local displacement of each voxel in the image, so as to align images at different times, modalities, or states to the same anatomical space) to obtain the non-rigid deformation field; the alignment parameters are composed of the rigid transformation parameters and the non-rigid deformation field; fuse the gray images of the optimized image data, retaining the bone structure in the low frequency and the soft tissue texture in the high frequency, so as to obtain the enhanced gray image.

[0032] The CycleGAN model includes: inputting the structure mask and the alignment parameters as the structure channel data, and the enhanced gray image as the texture channel data into the CycleGAN model; Among them, the structural channel uses a convolutional network to extract anatomical contour features; the texture channel uses a residual network to extract multi-scale texture features; the anatomical contour features and multi-scale texture features are fused to obtain a multi-modal registration and fusion image.

[0033] Specifically, the structural mask and alignment parameters are used as the data of the structural channel, and the enhanced grayscale image is used as the data of the texture channel, and input into the CycleGAN model. The structural consistency constraint and texture fidelity constraint are added to achieve fine control of the generation quality. The convolutional layer is used in the structural channel to extract anatomical contour features, focusing on the generation of rigid structures such as bone boundaries and implant paths. The geometric information of the original mask is retained through skip connections to prevent deformation. The residual network is used in the texture channel to extract multi-scale texture features, capturing details from the macroscopic gingival morphology to the microscopic enamel cracks. In the deep network of the generator, the weights of the structural features and texture features are dynamically allocated through the channel attention module. For example, in the bone-soft tissue junction area, the model enhances the structural constraint; in the pure soft tissue area, it focuses on texture generation. The structural consistency constraint includes forcing the bone or tooth contours in the generated image to strictly match the input structural mask, avoiding problems such as implant offset or root misalignment; by comparing the differences in the alignment parameters of the input and generated data, it is ensured that the model does not introduce non-physical spatial distortions. The texture fidelity constraint includes using a pre-trained neural network to extract the high-level features of the generated image and comparing them with the target image to ensure the biological rationality of the texture; checking the coherence of the texture at different scales. For example, when observing the enamel microcracks under a magnifying glass, the generated result needs to be consistent with the microscopic structure of the real image. The generated image is remapped back to the original input domain to ensure that the generation process is reversible and logically self-consistent, ensuring the spatial accuracy of the anatomical boundaries and retaining the biological authenticity of the soft tissues, and finally obtaining a multi-modal registration and fusion image.

[0034] S3: Perform wavelet transform on the multi-modal registration and fusion image to obtain multi-scale frequency domain data, and label the multi-scale frequency domain data to obtain labeled regions; The wavelet transform includes: performing three-level wavelet decomposition on the multi-modal registration and fusion image to obtain a low-frequency component and high-frequency sub-bands; identifying the artifact regions in the high-frequency sub-bands, determining the artifact interference level based on the artifact regions, and then setting the value range of the high-frequency weight coefficient and the value range of the normal region weight coefficient; performing sharpening processing on the low-frequency component to enhance the clarity of the bone contour and suppress the interference of the artifact regions; performing inverse wavelet transform fusion on the high-frequency sub-bands with adjusted weights and the sharpened low-frequency component to obtain multi-scale frequency domain data.

[0035] Specifically, performing wavelet transform on the multi-modal registration and fusion image can effectively suppress artifacts without sacrificing the overall quality. First, the generated image is decomposed into sub-bands of different scales, including low-frequency components (overall structure, such as bone contours) and high-frequency components (details and noise, such as gum textures, artifacts); second, the abnormally high-response regions in the high-frequency sub-bands (such as bright spots corresponding to metal artifacts) are regarded as artifact regions; based on the artifact regions, the artifact interference levels are determined, including level 1 (2-5 mm, local artifacts caused by small metal objects), level 2 (5-10 mm, star-shaped artifacts generated by medium metal restorations or high-density materials), and level 3 (10-15 mm, extensive artifacts caused by large metal implants or dense metal fillings). According to the artifact interference levels, the high-frequency weights are reduced (the weights are used to specifically suppress artifacts or enhance effective information, which can be adaptively adjusted by the model or manually adjusted) to weaken the noise and artifacts. The value range for level 1 is [0.4, 0.5], for level 2 is [0.3, 0.4], and for level 3 is [0.3, 0.4]; for normal regions, the high-frequency weights are maintained or enhanced, with a value range of [0.8, 1.0], to retain the real details; then, the low-frequency sub-band is slightly sharpened to enhance the clarity of the anatomical structure; finally, the adjusted high-frequency and low-frequency components are fused through inverse wavelet transform to obtain multi-scale frequency-domain data.

[0036] Based on the multi-scale frequency-domain data, the labeled regions are obtained by manually annotating the key regions, which include: artifact interference regions, non-linear deformation regions, diagnostic sensitive regions, and feature conflict regions, so as to ensure the accuracy of the optimization target. The image is segmented into multiple blocks, and only the labeled regions are subjected to wavelet transform and weight adjustment, while other regions remain unchanged as the original image, reducing memory occupancy. Progressive optimization is adopted. First, the low-frequency correction is performed on the labeled regions to eliminate structural distortion, and then the high-frequency artifacts are fine-tuned to avoid detail loss caused by over-processing; while effectively suppressing artifacts, the global computational cost is reduced.

[0037] For example, a patient who needs to have multiple posterior tooth implants undergoes 3D intraoral scanning and CBCT scanning before surgery. Due to the presence of old metal restorations in the patient's mouth, metal artifacts are generated in the CBCT image; at the same time, during the scanning process, the patient's slight swallowing causes the tongue to shift, resulting in non-linear deformation differences between the gum contours of the 3D scan and the bone structure of the CBCT. Using the method of Example 1, through the synergistic effect of rigid registration and non-rigid registration, the complex deformation problem caused by patient movement is solved; the dual-channel CycleGAN reconstructs the key anatomical structures covered by artifacts using the deep bone information of the CBCT while retaining the surface accuracy of the 3D scan. In traditional methods, due to the lack of hierarchical registration and non-linear deformation correction, the risk of implant perforating the nerve canal in similar cases is as high as 15%, while Example 1 reduces this risk to less than 2%.

[0038] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages: The present application uses 3D Canny edge detection to extract contours, combines with CBCT dynamic thresholds to generate structure masks, and solves the problem of misalignment between 3D scan surface data and CBCT anatomical landmarks; optimizes the registration efficiency by using rigid registration and non-rigid registration; uses the CycleGAN model to improve the training efficiency and result accuracy through the structure channel and texture channel; uses wavelet transform to effectively suppress artifacts. In contrast, the prior art relies on a single affine transformation for global alignment and cannot handle the elastic deformation of the soft tissue around the implant.

[0039] Embodiment 2: In Embodiment 1, the structural consistency constraint of CycleGAN may lead to the loss of texture details at complex anatomical junctions (such as the tooth crown and gingiva), and the design does not explicitly separate global / local features, which may lead to structural distortion. To address this issue, this embodiment further improves Embodiment 1.

[0040] S4: Decouple features based on the multi-modal spatio-temporal registration dataset to obtain hierarchical features; The feature decoupling step includes: extracting the global structural features of the structure mask through a convolutional neural network, constraining the symmetry error with the gold standard to be less than or equal to 0.3 mm; extracting the local detail features of the grayscale image through a residual network, keeping the edge gradient amplitude not less than 90% of the original image; using a spatial attention mechanism to dynamically fuse the global structural features and local detail features to obtain hierarchical features; providing a unified spatial reference for the temporal image to reduce the negative impact of artifacts on the fusion quality.

[0041] Such as Figure 2As shown, the features are decoupled into local detail features and global structure features, collectively referred to as hierarchical features. Through independent optimization and dynamic fusion, the collaborative expression ability of multimodal data is enhanced. Based on the data of structural masks and rigid registration, a convolutional neural network (a deep learning model specialized for processing data with spatial structures) is used to extract global structure features. By means of an attention mechanism, key anatomical landmarks are focused on to obtain vectors of global structure features, such as bone density distribution and symmetry. Based on enhanced grayscale images and annotated regions, high-frequency details are extracted through a residual network, and noise is suppressed by adaptive filtering to obtain tensors of local detail features, such as gingival texture gradients and microcrack morphologies. The global structure features and local detail features are independently optimized. The Dice coefficient between the global features and the gold standard (such as CBCT bone segmentation) is forced to be ≥0.95 (measuring the overlap degree between the generated result and the gold standard) and the symmetry error of the left and right jawbone morphologies is ≤0.3 mm to maintain anatomical structure consistency, thereby completing the global structure optimization. By comparing the texture distributions of the generated image and the real image, it is ensured that the microstructures of soft tissues and hard tissues conform to anatomical laws. The differences between the generated image and the real image are quantified by integrating three indicators of brightness, contrast, and structural similarity. By maximizing the edge gradient amplitude of the generated image, the sharpness of key boundaries such as enamel-gingiva and bone-soft tissue is ensured. In the artifact region, the gradient loss can reduce edge diffusion and maintain the sharp contour of the anatomical structure, thereby completing the local detail optimization.

[0042] Generate a spatial attention map according to the global features to identify key anatomical regions; generate a channel attention vector according to the local features to strengthen the detail regions; then dynamically fuse the global features and the local features to obtain the final feature representation, and its formula is: , where, is the fused feature representation, is the global feature, is the local feature, is the weight coefficient, which is used to control the contribution ratio of the global feature and the local feature, and its value range is [0.6, 0.8]. The closer the value is to 0.8, the more it emphasizes the global feature, and vice versa, it emphasizes the local feature. For example, in the artifact region, α is reduced to reduce the amplification of noise by the global structure and adapt to the requirements of different clinical scenarios.

[0043] The technical solutions in the above embodiments of the present application at least have the following technical effects or advantages: This application uses feature decoupling to independently optimize the global structure and local details, and dynamically fuses weights to ensure that global interference is reduced in key areas such as artifacts, improving texture authenticity; combined with the edge gradient maximization constraint in local detail optimization, it specifically enhances effective high-frequency information. At the same time, through the spatial attention mechanism, the global feature weight is actively reduced in the artifact area to reduce noise amplification; a convolutional network is used to extract global features and a residual network is used to extract local features, and they are dynamically fused through spatial attention to solve the problem of deformation accumulation at the implant-soft tissue junction.

[0044] Embodiment 3: In Embodiment 2, feature decoupling only targets static data at a single time point and does not consider the temporal correlation of multiple examination images, resulting in the inability to capture dynamic changes such as implant displacement or bone resorption; non-rigid registration only processes the soft tissue deformation of a single scan and lacks a correction mechanism for the progressive anatomical structure changes during the treatment process. To address this problem, this embodiment further improves Embodiment 2.

[0045] Step S3 further includes: using multi-modal oral image data collected at different times as temporal image data; performing spatio-temporal alignment based on the labeled area and the temporal image data to obtain a multi-modal spatio-temporal registration data set; The spatio-temporal alignment step includes: calculating the non-rigid deformation field of soft tissue deformation through non-rigid registration, and obtaining the composite alignment parameter according to the alignment parameter and the non-rigid deformation field; based on the composite alignment parameter, registering the temporal image data to the static reference space to form a rigid registration group, establishing a unified space coordinate system through the elastic deformation field, and the rigid registration error tolerance is less than or equal to 0.1 mm; constructing a non-rigid compensation group based on the feature data extracted from the labeled area, using a recurrent neural network to predict the local progressive deformation field, and the maximum deformation amount is less than or equal to 2.5 mm; achieving precise alignment of bones and soft tissues of images across time points through a hierarchical registration strategy, where the rigid registration group maintains the consistency of the implant spatial position, and the non-rigid compensation group corrects the gingival deformation, and outputs a multi-modal spatio-temporal registration data set.

[0046] Specifically, use multi-modal oral image data collected at different times as temporal image data, and divide the data into a rigid registration group and a non-rigid compensation group according to the labeled area and the temporal image data by performing spatio-temporal alignment (synchronizing data or events in different time series to ensure they are aligned on the time axis); Obtain the non-rigid deformation field of soft tissue deformation through non-rigid registration; obtain the composite alignment parameter according to the alignment parameter and the non-rigid deformation field, and its formula is: , where, is the coordinate of a certain point, is the composite alignment parameter, is the alignment parameter, is a non-rigid deformation field.

[0047] Based on the composite alignment parameters, all temporal data are registered to the static reference space, and the data eliminating body position differences are used as the rigid registration group. A unified space coordinate system for all temporal data is established, with an error tolerance ≤ 0.1 mm to ensure the comparability of anatomical structures in the spatial dimension; the deformed data corrected by the deformation field are used as the non-rigid compensation group, allowing local deformations in the dynamically changing areas under the unified space reference, with a maximum deformation ≤ 2.5 mm; rigid registration eliminates global offsets, and non-rigid compensation retains local dynamics, forming a hierarchical group optimization.

[0048] Based on the registered data, the images of the patient's multiple examinations are used as dynamic temporal data to capture the dynamic changes of its structure; the single high-precision CT image is used as static structural data to provide a spatial reference to ensure the anatomical rationality of cross-time data alignment; a recurrent neural network (a neural network specialized in processing sequence data, whose core feature is the ability to capture the time dependence in data through "memory") is used to capture dynamic temporal features and obtain time-encoded vectors (such as bone resorption rate, implant displacement); a convolutional neural network is used to capture static structural features and obtain spatial-encoded vectors (such as bone density distribution, symmetry parameters). The dynamic temporal features and static structural features are fused and weighted to obtain the formula: , where, represents the fused feature, is the static structural feature, is the dynamic temporal feature, is the spatio-temporal attention, with a value range of [0, 1]. The closer the value is to 1, it indicates that the static structural feature is the dominant one, used to stabilize the anatomical area and strengthen the constraint of the reference structure; the closer the value is to 0, it indicates that the dynamic temporal feature is the dominant one, used for the significantly changing area, emphasizing the temporal evolution feature; The precise alignment of bones and soft tissues in images across time points is achieved through a hierarchical registration strategy. Among them, the rigid registration group maintains the consistency of the implant spatial position, and the non-rigid compensation group corrects the gingival deformation, outputting a multi-modal spatio-temporal registration dataset.

[0049] The technical solutions in the embodiments of the present application described above have at least the following technical effects or advantages: This application adopts spatio-temporal alignment, divides the temporal images into rigid groups and non-rigid groups, extracts dynamic temporal features through a recurrent neural network to achieve precise alignment of cross-time data; sets up a non-rigid compensation group, allows the local deformation field to be dynamically adjusted, and realizes precise compensation of progressive deformation through cumulative calculation of the deformation field; establishes a hierarchical registration strategy, where the rigid group enforces errors to ensure benchmark consistency, and the non-rigid group establishes regional deformation associations through a graph convolutional network to suppress error propagation.

[0050] Embodiment 4: In Embodiment 3, the spatio-temporal attention is a global parameter and cannot distinguish the sensitivity differences of regions such as bone tissue / soft tissue to CT and 3D scan data, resulting in the over-weakening of CT features in the implant region; the local deformation field of the non-rigid compensation group does not consider the motion consistency of adjacent regions, which may lead to conflicts in the deformation directions of the gingiva and the dental crown; the artifact suppression predicted by the recurrent neural network acts on the entire image, and there are still residual artifacts around the metal implant. To address this problem, this embodiment further improves Embodiment 3.

[0051] Step S4 further includes: performing region adaptive fusion based on hierarchical features, a structure mask, and a grayscale image to obtain a multi-modal optimized fusion image; The region adaptive fusion step includes: based on hierarchical features, dividing the multi-modal spatio-temporal registration data set into [10, 15] functional regions; calculating spatial proximity edges based on the structure mask; calculating feature similarity edges based on multi-scale frequency domain data; establishing a region association matrix according to the spatial proximity edges and the feature similarity edges; and using an attention network to dynamically assign modal weights based on the region association matrix to obtain a multi-modal optimized fusion image.

[0052] Specifically, based on hierarchical features, a structure mask, and a grayscale image, geometric features such as curvature and density distribution are extracted from the hierarchical features, and texture features such as gray level co-occurrence matrix and local binary pattern are also extracted. The K-means clustering algorithm is used to construct a graph structure based on feature similarity, and divided into K regions, such as k = 10, with a value range of [10, 15]; an output region label map is used to mark the region to which each voxel belongs; the effect of automatically dividing functional regions according to image features is achieved, realizing unsupervised region division.

[0053] Specifically, based on the divided regions, an association model of feature dependencies between regions is built; each region is regarded as a node, and feature vectors such as texture, volume, and spatial position are used as node features. The connection of edges is based on two independent conditions, and an edge is established as long as either condition is met; for the spatially adjacent edges, Euclidean distance is calculated with a threshold of 5 mm, and an edge is established if the distance between the geometric centers of two regions is ≤ 5 mm; for the feature similarity edges, cosine similarity is used with a threshold of 0.7, and an edge is established if the cosine similarity of the feature vectors of two regions is ≥ 0.7; the spatially adjacent edges and the feature similarity edges are input into a graph convolutional network, and an output region association matrix is obtained.

[0054] Specifically, based on the region association matrix, regional adaptive fusion is performed to implement a differential fusion strategy and optimize local accuracy. First, the importance weights of regions are calculated through an attention mechanism, and the formula is: , where is the weight of region ; is the set of neighbor regions connected to region ; is an element of the region association matrix, representing the association strength between regions i and j; is a fully connected layer that maps the feature vector of region j to a scalar; is a normalization function that maps the weights to [0, 1] and the sum is 1.

[0055] The modal weights are dynamically adjusted according to region characteristics. For example, the CT weight is high in the bone region, with a value range of [0.7, 0.9]; the 3D scan weight is high in the soft tissue region, with a value range of [0.6, 0.8]; the formula for the importance weights of the fusion regions is obtained as the formula: , where is the pixel value of the fused image at position x; M is the number of modalities; is the contribution weight of modality m in region i; is the pixel value of modality m at position x; by independently calculating the weights for each region, differential fusion is achieved, and a multi-modal optimized fused image is obtained.

[0056] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages: The present application uses K-means clustering to divide regions, enforces the CT weight in the bone tissue region, and accurately retains the internal trabecular bone structure; constructs a region association matrix, and through dual constraints of spatial proximity and feature similarity, avoids anatomical structure misalignment; automatically identifies the implant region with HU > 2000 during the region segmentation stage, and specifically reduces the weight of high-frequency components, resulting in an improvement in the artifact elimination rate.

[0057] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A fusion processing method based on multi-modal oral imaging data, characterized in that Including: S1: Obtain multi-modal oral image data, and obtain a structure mask and a grayscale image based on the multi-modal oral image data; the multi-modal oral image data includes 3D surface oral image data and CBCT oral image data; S2: Perform rigid registration on the structure mask to obtain alignment parameters, and perform non-rigid registration on the grayscale image to obtain an enhanced grayscale image; input the structure mask, alignment parameters and enhanced grayscale image into the CycleGAN model to obtain a multi-modal registration and fusion image; S3: Perform wavelet transform on the multi-modal registration and fusion image to obtain multi-scale frequency domain data, and perform annotation on the multi-scale frequency domain data to obtain an annotated area; Use the multi-modal oral image data collected at different times as time-series image data; Perform spatio-temporal alignment according to the annotated area and the time-series image data to obtain a multi-modal spatio-temporal registration data set; S4: Perform feature decoupling based on the multi-modal spatio-temporal registration data set to obtain hierarchical features; Perform region adaptive fusion based on the hierarchical features, structure mask and grayscale image to obtain a multi-modal optimized fusion image.

2. The fusion processing method based on multi-modal oral cavity image data according to claim 1, characterized in that, The generation of the structure mask includes: Perform 3D Canny edge detection and topological repair processing on the 3D surface oral image data to obtain a three-dimensional edge contour; perform dynamic threshold segmentation and connected component analysis on the CBCT oral image data to obtain a bone structure segmentation result; superimpose the three-dimensional edge contour and the bone structure segmentation result to generate a structure mask.

3. The fusion processing method based on multi-modal oral image data according to claim 2, wherein The 3D Canny edge detection includes: Perform three-dimensional direction gradient detection on the multi-modal oral image data, calculate the brightness change intensity in the X, Y, and Z axis directions to obtain a gradient amplitude map and a direction map; Based on the gradient amplitude map and the direction map, refine the edge to a single pixel width through non-maximum suppression to obtain a refined edge map; Set a gradient threshold based on the edge map, retain the feature of the junction between the tooth crown edge and the gingiva to obtain a three-dimensional edge contour.

4. The fusion processing method based on multimodal oral cavity image data according to claim 1, wherein, The rigid registration is to obtain bone contour features based on the structure mask, and then calculate the rigid transformation parameters to obtain alignment parameters; perform non-rigid registration on the grayscale image according to the alignment parameters to obtain an enhanced grayscale image.

5. A fusion processing method for multi-modal oral imaging data according to claim 1, characterized in that The CycleGAN model includes: Input the structure mask and alignment parameters as structure channel data, and the enhanced grayscale image as texture channel data into the CycleGAN model; Among them, the structure channel uses a convolutional network to extract anatomical contour features; the texture channel uses a residual network to extract multi-scale texture features; fuse the anatomical contour features and multi-scale texture features to obtain a multi-modal registration and fusion image.

6. The fusion processing method based on multi-modal oral cavity image data according to claim 1, characterized in that, The wavelet transform includes: Perform three-level wavelet decomposition on the multi-modal registration and fusion image to obtain a low-frequency component and high-frequency sub-bands; Identify the artifact areas in the high-frequency sub-bands, determine the artifact interference level based on the artifact areas, and then set the value range of the high-frequency weight coefficient and the value range of the normal area weight coefficient; Sharpen the low-frequency component to enhance the clarity of the bone contour and suppress the interference of the artifact areas; Perform inverse wavelet transform fusion on the high-frequency sub-bands with adjusted weights and the sharpened low-frequency component to obtain multi-scale frequency domain data.

7. The fusion processing method based on multi-modal oral cavity image data according to claim 1, wherein, The annotated area includes: The multi-scale frequency-domain data is segmented into 8×8 pixel blocks. The deformation field residuals within each block are calculated based on the alignment parameters, and the regions with deformation errors greater than 0.3 mm are determined as uncorrected non-linear deformation regions. At the same time, the sliding window standard deviation of the HU values of each block is calculated, and the regions with a sudden increase in the standard deviation greater than 200 are identified as metal artifact interference regions. The non-linear deformation regions and the metal artifact interference regions together constitute the labeled regions.

8. A fusion processing method based on multi-modal oral image data according to claim 1, characterized in that, The feature decoupling step includes: The global structural features of the structural mask are extracted through a convolutional neural network, and the symmetry error with the gold standard is constrained to be less than or equal to 0.3 mm. The local detail features of the grayscale image are extracted through a residual network, and the edge gradient amplitude is maintained at no less than 90% of the original image. A spatial attention mechanism is used to dynamically fuse the global structural features and the local detail features to obtain hierarchical features, providing a unified spatial reference for the temporal images and reducing the negative impact of artifacts on the fusion quality.

9. The fusion processing method of multi-modal oral cavity image data according to claim 1, wherein The spatio-temporal alignment step includes: The non-rigid deformation field of soft tissue deformation is calculated through non-rigid registration, and the composite alignment parameters are obtained based on the alignment parameters and the non-rigid deformation field. Based on the composite alignment parameters, the temporal image data is registered to the static reference space to form a rigid registration group. A unified spatial coordinate system is established through the elastic deformation field, and the rigid registration error tolerance is less than or equal to 0.1 mm. Based on the feature data extracted from the labeled regions, a non-rigid compensation group is constructed, and a recurrent neural network is used to predict the local progressive deformation field, with the maximum deformation amount less than or equal to 2.5 mm. The precise alignment of bones and soft tissues of images across time points is achieved through a hierarchical registration strategy, where the rigid registration group maintains the consistency of the spatial position of the implant, and the non-rigid compensation group corrects the gingival deformation, and a multi-modal spatio-temporal registration data set is output.

10. A fusion processing method for multi-modal oral imaging data according to claim 1, characterized in that The region adaptive fusion step includes: Based on the hierarchical features, the multi-modal spatio-temporal registration data set is segmented into [10, 15] functional regions. The spatial proximity edges are calculated based on the structural mask. The feature similarity edges are calculated based on the multi-scale frequency-domain data. Based on the spatial proximity edges and the feature similarity edges, a region association matrix is established. Based on the region association matrix, an attention network is used to dynamically allocate modal weights to obtain a multi-modal optimized fusion image.

Citation Information

Patent Citations

  • Medical image registration method and device

    CN113096166A

  • System and method for cardiac segmentation in mr-cine data using inverse consistent non-rigid registration

    US20110081066A1

  • Modeling and visualization of facial structure for dental treatment planning

    US20250200894A1

  • Modeling and visualization of facial structure for dental treatment planning

    WO2025097057A1

Cited By

  • Infrared analysis method for traditional Chinese medicine powder based on image recognition

    CN120581090A

  • Electrical impedance state evaluation method and system fused with CT (Computed Tomography) characteristics

    CN120827364A

  • Construction method of prostate cancer image recognition and classification model based on deep learning

    CN120894635A

  • PET image segmentation method and system, computer equipment and storage medium

    CN121147157A

  • A PET image segmentation method, system, computer device and storage medium

    CN121147157B