Multi-modal fusion oral cavity endoscopic image three-dimensional reconstruction method and system
Through an image processing method combining multispectral LED array probe and dental prior knowledge, a dual-channel feature extraction network and a non-rigid hierarchical fusion framework are configured to solve the problem of insufficient three-dimensional reconstruction accuracy of oral endoscopic images and realize high-precision three-dimensional reconstruction of oral cavity.
Patent Information
- Application Number
- CN202510258917.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-29
AI Technical Summary
In the prior art, the three-dimensional reconstruction accuracy of oral endoscopic images is insufficient, and the multi-dimensional information inside the oral cavity cannot be fully captured, resulting in insufficient reconstruction accuracy and details, which is difficult to meet the needs of diagnosis and treatment.
Multispectral LED array probes are used for imaging acquisition, combined with dental prior knowledge and adversarial generation network for image processing, a dual-channel feature extraction network is configured, and a three-dimensional reconstruction correction is performed through a non-rigid hierarchical fusion framework to generate corrected three-dimensional reconstruction results.
The accuracy and details of oral three-dimensional reconstruction are improved, and a more accurate and comprehensive oral three-dimensional model is generated, meeting the needs of diagnosis and treatment.
Smart Images

Figure CN120388129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional reconstruction, and particularly to a three-dimensional reconstruction method and system for oral endoscopy images with multi-modal fusion. Background Art
[0002] With the rapid development of oral medicine and digital technology, three-dimensional reconstruction technology of oral endoscopy images has been widely used in diagnosis, treatment, and surgical planning. However, with the increasing medical needs and the complexity of oral structures, the three-dimensional reconstruction accuracy of oral endoscopy images faces increasing challenges. Traditional single-modal imaging methods, such as only using visible light imaging or infrared light imaging, cannot comprehensively capture the multi-dimensional information inside the oral cavity, resulting in insufficient accuracy and details in three-dimensional reconstruction, and it is difficult to meet the requirements of accuracy, clarity, and comprehensiveness for oral images. Summary of the Invention
[0003] This application provides a three-dimensional reconstruction method and system for oral endoscopy images with multi-modal fusion, which is used to solve the technical problem of insufficient accuracy in three-dimensional reconstruction of oral endoscopy images in the prior art.
[0004] In view of the above problems, this application provides a three-dimensional reconstruction method and system for oral endoscopy images with multi-modal fusion.
[0005] In the first aspect of this application, a three-dimensional reconstruction method for oral endoscopy images with multi-modal fusion is provided. The method includes:
[0006] Configuring a multi-spectral LED array probe, performing oral imaging acquisition, and establishing an imaging data set. The spectrum of the multi-spectral LED array probe includes visible light, infrared light, and fluorescence imaging; introducing dental prior knowledge constraints for processing the imaging data set, performing adversarial reconstruction of the imaging data set, and generating an adversarial reconstruction result; configuring a dual-channel feature extraction network, performing dual-channel feature extraction of the adversarial reconstruction result, and establishing a dual-channel feature extraction result; obtaining the acquisition coordinates and acquisition parameters of the imaging data set, mapping the dual-channel feature extraction result to the same feature space according to the acquisition coordinates and acquisition parameters, and then performing pixel search alignment; performing three-dimensional reconstruction of the oral cavity based on the dual-channel feature extraction result according to the pixel search alignment result, and performing three-dimensional reconstruction correction through a non-rigid hierarchical fusion framework to generate a corrected three-dimensional reconstruction result.
[0007] In the second aspect of this application, a three-dimensional reconstruction system for oral endoscopy images with multi-modal fusion is provided. The system includes:
[0008] An imaging acquisition module, configured to configure a multi-spectral LED array probe, perform oral imaging acquisition, and establish an imaging data set. The spectrum of the multi-spectral LED array probe includes visible light, infrared light, and fluorescence imaging. An adversarial reconstruction module, configured to introduce dental prior knowledge constraints for processing the imaging data set, perform adversarial reconstruction of the imaging data set, and generate an adversarial reconstruction result. A feature extraction module, configured to configure a dual-channel feature extraction network, perform dual-channel feature extraction of the adversarial reconstruction result, and establish a dual-channel feature extraction result. A pixel search alignment module, configured to obtain the acquisition coordinates and acquisition parameters of the imaging data set, and perform pixel search alignment after mapping the dual-channel feature extraction result to the same feature space according to the acquisition coordinates and acquisition parameters. A three-dimensional reconstruction correction module, configured to perform three-dimensional reconstruction of the oral cavity based on the dual-channel feature extraction result according to the pixel search alignment result, and perform three-dimensional reconstruction correction through a non-rigid hierarchical fusion framework to generate a corrected three-dimensional reconstruction result.
[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0010] This application configures a multi-spectral LED array probe, performs oral imaging acquisition, and establishes an imaging data set. The spectrum of the multi-spectral LED array probe includes visible light, infrared light, and fluorescence imaging. It introduces dental prior knowledge constraints for processing the imaging data set, performs adversarial reconstruction of the imaging data set, and generates an adversarial reconstruction result. It configures a dual-channel feature extraction network, performs dual-channel feature extraction of the adversarial reconstruction result, and establishes a dual-channel feature extraction result. It obtains the acquisition coordinates and acquisition parameters of the imaging data set, performs pixel search alignment after mapping the dual-channel feature extraction result to the same feature space according to the acquisition coordinates and acquisition parameters. It performs three-dimensional reconstruction of the oral cavity based on the dual-channel feature extraction result according to the pixel search alignment result, and performs three-dimensional reconstruction correction through a non-rigid hierarchical fusion framework to generate a corrected three-dimensional reconstruction result. This invention solves the technical problem of insufficient three-dimensional reconstruction accuracy of oral endoscopy images in the prior art, and achieves the technical effect of improving the three-dimensional reconstruction accuracy of the oral cavity through a multi-spectral LED array probe, dual-channel feature extraction, pixel alignment, and non-rigid hierarchical fusion correction. Description of the Drawings
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic flowchart of a three-dimensional reconstruction method for multi-modal fusion of oral endoscopy images provided by an embodiment of this application;
[0013] Figure 2 This is a schematic structural diagram of a three-dimensional reconstruction system for multi-modal fusion of oral endoscopy images provided by an embodiment of the present application.
[0014] Explanation of reference numerals: Imaging acquisition module 11, adversarial reconstruction module 12, feature extraction module 13, pixel search and alignment module 14, three-dimensional reconstruction correction module 15. Specific implementation manners
[0015] By providing a method and system for three-dimensional reconstruction of multi-modal fusion of oral endoscopy images, the present application aims to solve the technical problem of insufficient accuracy of three-dimensional reconstruction of oral endoscopy images in the prior art. Through a multi-spectral LED array probe, dual-channel feature extraction, pixel alignment, and non-rigid hierarchical fusion correction, the technical effect of improving the accuracy of oral three-dimensional reconstruction is achieved.
[0016] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0017] It should be noted that any variations of the terms "including" and "having" are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0018] Embodiment 1, as Figure 1 shown, the present application provides a method for three-dimensional reconstruction of multi-modal fusion of oral endoscopy images, and the method includes:
[0019] Step S100: Configure a multi-spectral LED array probe, perform oral imaging acquisition, and establish an imaging data set. The spectrum of the multi-spectral LED array probe includes visible light, infrared light, and fluorescence imaging.
[0020] In the embodiments of the present application, a multi-spectral LED array probe is first configured. The multi-spectral LED array probe consists of multiple LED light sources and can emit light of different wavelengths. These spectra can be selected according to different imaging requirements. The spectra of the multi-spectral LED array probe include visible light, infrared light, and fluorescence imaging. Specifically, visible light (with wavelengths approximately between 400 - 700 nm) is used to obtain conventional images inside the oral cavity to help observe the surface morphology of teeth, gums, and other structures; infrared light (generally in the range of 700 nm to 1 mm) is used to obtain images of the internal structures of teeth and gums, which can reveal some deep information that cannot be identified by the naked eye, such as the mineralization of teeth or potential caries areas; and fluorescence imaging (usually by exciting a light source of a specific wavelength and then recording the fluorescence of different wavelengths emitted) can further help detect minor lesions or pathological changes inside the oral cavity, such as the initial stage of caries or early signs of oral cancer.
[0021] Through configuration by technical experts for the multi-spectral LED array probe, the configured multi-spectral LED array probe irradiates the inner wall of the oral cavity with light sources of different wavelengths, collects imaging data in multiple modes, and obtains an imaging data set.
[0022] Furthermore, in the method provided by the embodiments of the application, when performing oral imaging acquisition and establishing an imaging data set, it further includes:
[0023] Establishing a motion sensitivity map, performing motion artifact correction on the acquired images based on the motion sensitivity map. The motion sensitivity map is as follows:
[0024]
[0025] Among them, M sens (x, y) represents the motion sensitivity of the pixel point (x, y), I(x, y) represents the image intensity of the pixel point (x, y), t i is the sampling time, N represents the set of time points, i represents any one time point, and G σ is the Gaussian filter kernel; the imaging data set is established according to the motion artifact correction result.
[0026] In the embodiments of the present application, a motion sensitivity map is first established. In this process, first, continuous oral image data at multiple time points are obtained through the multi-spectral LED array probe. Each image corresponds to a time point, each image contains multiple pixel points, and the intensity I(x, y) of each pixel point in the image represents the brightness or color value of the pixel. Then, to capture the image changes caused by patient movement, the intensity changes of each pixel point at different time points are calculated. This is achieved by solving the intensity differences of each pixel point at each time point, and the specific calculation is where t iis the sampling time, representing the change amount of the pixel at time point t i At this point, the change amount is calculated. By this calculation, the temporal change of each pixel is quantified, and then the image changes caused by motion are identified. Subsequently, the Gaussian filter kernel G σ is used to perform weighted smoothing on these changes. Gaussian filtering is a low-pass filtering technique that smooths the image, removes high-frequency noise in the image, and only retains relatively stable low-frequency information. This filtering can remove the noise generated by short-term motion changes, highlight those continuous motion changes, and ensure that the effective information of the image is not weakened.
[0027] Through the above operations, the finally obtained motion-sensitive map reflects the sensitivity degree of each pixel point to motion artifacts.
[0028] Next, motion artifact correction is performed. The regions sensitive to motion are determined through the motion-sensitive map, and these regions usually appear as blurred or distorted artifacts in the image. In this process, Gaussian filtering is used to process these sensitive regions, aiming to smooth and correct the image distortion caused by patient motion. Gaussian filtering performs weighted average processing on each pixel in the image. Especially for those regions with large changes, these regions are smoothed more strongly through the Gaussian kernel. The filtering process helps to reduce the artifacts caused by motion and keep the image smooth and stable.
[0029] After this process, the motion artifacts in the image will be effectively removed, and the corrected image will be clearer and more accurate. Finally, an imaging data set corrected for motion artifacts is obtained.
[0030] Furthermore, in the method provided by the application embodiment, the execution of oral imaging acquisition and the establishment of the imaging data set further include:
[0031] Performing image self-verification on the oral imaging image to generate a self-verification result; when the self-verification result is a non-passing result, a reset acquisition instruction is generated; and according to the reset acquisition instruction, the image is reset and acquired.
[0032] In the embodiments of the present application, first, image self-verification is performed on oral imaging images to automatically evaluate the quality of the acquired images. The purpose of image self-verification is to detect the images by using image quality assessment algorithms to ensure that the images do not have problems such as blurring, noise, uneven exposure, or other common quality issues. To this end, common techniques include image quality assessment, such as PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index), edge detection, and exposure uniformity detection. PSNR is used to evaluate the clarity and noise level of the image. If the PSNR value of the image is lower than the set threshold, it indicates that the image has too much noise or distortion, which may lead to unqualified quality. SSIM is used to evaluate the structural similarity of the image. A low SSIM value indicates that the details of the image are lost or the contrast is insufficient, which may affect subsequent analysis. Edge detection, such as Canny edge detection, is used to detect the sharpness of the image. If the image is blurred or cannot clearly display the oral structure, it is judged as not passing. In addition, exposure uniformity detection ensures uniform brightness distribution of the image by checking the histogram distribution of the image, avoiding overexposure or underexposure. After these steps, a self-verification result is generated.
[0033] The self-verification result is divided into passing or not passing. When the image quality meets all the preset standards and passes all the quality assessment tests, a passing result will be returned, indicating that the image can be used for subsequent processing. If the image has quality problems and fails to meet the preset standards (such as blurred image, excessive noise, uneven exposure, or insufficient resolution, etc.), a not passing result will be returned, and the specific problem will be pointed out, such as blurring, excessive noise, or exposure problems.
[0034] When the self-verification result is not passing, a reset acquisition instruction is automatically generated. This instruction will trigger the image acquisition device to re-adjust the settings (such as refocusing, adjusting the exposure time, changing the shooting angle, etc.) to ensure that the images acquired next time meet the quality standards. Restart the image acquisition process according to the reset acquisition instruction to ensure that the finally obtained images can meet the high-quality requirements.
[0035] Through this process, the image quality is automatically detected, and corresponding processing is performed according to the self-verification result to ensure that the quality of each acquired image meets the requirements, providing reliable data support for subsequent image analysis and processing.
[0036] Step S200: Introduce dental prior knowledge constraints to process the imaging dataset, perform adversarial reconstruction of the imaging dataset, and generate an adversarial reconstruction result.
[0037] In the embodiments of the present application, first, the imaging data set is processed by introducing dental prior knowledge constraints. Dental prior knowledge constraints refer to using the known characteristics of oral anatomical structures and lesions, such as the shape, size, arrangement of teeth, and common lesion areas, etc., to provide additional guidance for image processing algorithms during the image reconstruction process. These prior knowledge help limit the space of the image processing results to conform to the actual anatomical structure and avoid errors caused by noise or artifact influence. Specifically, dental prior knowledge is embedded into the image processing algorithm and implemented through morphological constraints, regional constraints, and position constraints. For example, the prior knowledge can limit the shape and arrangement pattern of teeth in the image to ensure that the reconstructed image conforms to the normal oral structure and avoid generating incorrect images that do not conform to anatomical laws.
[0038] Next, the adversarial reconstruction step of the imaging data set is executed. In this process, a Generative Adversarial Network (GAN) is adopted. A GAN consists of two parts: a generator and a discriminator. The task of the generator is to generate as realistic a reconstructed image as possible from the input image data, while the discriminator evaluates the difference between these generated images and the real images. The training data set is the basis for training the generator and the discriminator, including a large number of labeled oral image data sets. These image data contain information such as the shape, size, arrangement of teeth, and lesion areas (such as dental caries, periodontal disease, etc.), and these data are pre-prepared. Through these labeled data, the generator learns how to extract key information from the images and generate accurate images, while the discriminator judges the authenticity of the generated images by comparing them with the real images. The generator and the discriminator are continuously optimized through adversarial training. The generator adjusts its parameters through backpropagation, making the generated images closer and closer to the real images. At the same time, the discriminator judges according to whether the images conform to the prior knowledge to ensure that the generated images not only conform to visual realism but also conform to the oral anatomical structure. In specific implementation, the generator processes the images through technologies such as Convolutional Neural Network (CNN), and uses prior knowledge for optimization to generate more accurate images. The discriminator evaluates the generated images according to the prior knowledge to judge whether the images conform to the dental anatomical standards.
[0039] In this process, the generator and the discriminator are alternately trained, constantly competing and optimizing with each other. After each training, the generator adjusts the generated images according to the feedback of the discriminator to make them more conform to the real oral structure and prior knowledge, while the discriminator optimizes its judgment ability according to the difference between the generated images and the real images. Through multiple rounds of adversarial training, the generator can generate more and more realistic and clear images. Especially when dealing with complex oral internal structures, it can better retain details and remove noise and artifacts in the images.
[0040] After training, the generator performs adversarial reconstruction of the imaging dataset and finally generates high-quality adversarial reconstruction results. These results not only remove the noise in the image but also accurately reconstruct the internal oral structures, such as the arrangement of teeth, the location of dental caries, and other lesion areas.
[0041] Step S300: Configure a dual-channel feature extraction network, perform dual-channel feature extraction of the adversarial reconstruction result, and establish a dual-channel feature extraction result.
[0042] In the embodiment of the present application, the process of configuring the dual-channel feature extraction network aims to extract key information from the adversarially reconstructed oral images. Specifically, first, two dedicated feature channels are set up, namely the geometric feature channel and the pathological feature channel. The geometric feature channel is used to identify the structural information in the image, such as the shape and arrangement of teeth, while the pathological feature channel focuses on extracting disease-related features, such as signs of lesions like dental caries and periodontal diseases.
[0043] In this process, the geometric feature channel and the pathological feature channel respectively extract the corresponding geometric features and pathological features from the adversarial reconstruction result, obtaining a geometric feature extraction result and a pathological feature extraction result. Then, these two feature results are synchronized to a cross-attention collaboration channel, which is responsible for fusing the information from the two feature channels, performing feature recombination, and effectively integrating the geometric and pathological features.
[0044] Finally, through the process of feature recombination, a dual-channel feature extraction result is generated.
[0045] Furthermore, in the method provided by the embodiment of the application, the configuring of the dual-channel feature extraction network, performing dual-channel feature extraction of the adversarial reconstruction result, and establishing a dual-channel feature extraction result further includes:
[0046] Configure a geometric feature channel, use the geometric feature channel to identify the geometric features of the adversarial reconstruction result, and establish a geometric feature extraction result; configure a pathological feature channel, use the pathological feature channel to identify the pathological features of the adversarial reconstruction result, and establish a pathological feature extraction result; synchronize the geometric feature extraction result and the pathological feature extraction result to a cross-attention collaboration channel, perform feature recombination, and establish a dual-channel feature extraction result according to the feature recombination result.
[0047] In the embodiments of the present application, a geometric feature channel is first configured. The geometric feature channel is specifically used to identify and extract structural features in images, such as the shape, arrangement, edges, and spatial positions of teeth. These geometric features provide strong support for the overall structure and interrelationships of teeth. This channel uses a convolutional neural network (CNN) to process the image through multiple convolutional layers and extract the geometric features of the teeth. During the training process, a large number of labeled oral image data are used, including information on the morphology and arrangement of teeth. Through these data, the network learns how to identify the geometric features in the image and finally generates the geometric feature extraction result.
[0048] Then, a pathological feature channel is configured, and the same convolutional neural network (CNN) is used to extract pathological features in the image. This network also specifically identifies the areas showing lesions in the image, such as dental caries, periodontitis, tooth cracks, etc. through convolutional operations. By using a pre-labeled image dataset for training, the network learns how to extract pathological features from the image, and these features can help us judge the health status of the teeth. The output of the pathological feature channel is the pathological feature extraction result.
[0049] Subsequently, the geometric feature extraction result and the pathological feature extraction result are synchronized to the cross-attention collaboration channel to perform feature recombination. After feature recombination, a dual-channel feature extraction result is obtained.
[0050] Furthermore, in the method provided by the embodiments of the application, the step of synchronizing the geometric feature extraction result and the pathological feature extraction result to the cross-attention collaboration channel to perform feature recombination further includes:
[0051] Establish an attention weight matrix as follows:
[0052]
[0053] where A represents the attention weight matrix, Q represents the query matrix, which is calculated by normalizing and linearly transforming the geometric feature extraction result, K represents the key matrix, which is calculated by normalizing and linearly transforming the pathological feature extraction result, T represents matrix transpose, d is the scaling factor, and M prior represents the prior mask;
[0054] Perform feature fusion on the attention weight matrix and the value matrix to construct a fused feature:
[0055] F fusion = GELU(A·V);
[0056] F fusion represents the fused feature, and V is the value matrix;
[0057] Perform gated attention calculation based on the fused feature to perform feature recombination:
[0058] F final = G ⊙ F g +(1 - G) ⊙ F fusion ;
[0059] Wherein, F final represents the recombination feature, G is the gating factor, G = σ(W g [F g ; F fusion + b g ), σ is the activation function, W g represents the weight matrix of the gating mechanism, F g represents the geometric feature, b g is the bias term of the gating mechanism.
[0060] In the embodiments of the present application, when synchronizing the geometric feature extraction result and the pathological feature extraction result to the cross-attention collaborative channel and performing feature recombination, first, an attention weight matrix is constructed, which is obtained by the interaction calculation of the query matrix Q and the key matrix K. The query matrix Q is obtained by normalizing and linearly transforming the geometric feature extraction result, representing the query of the image geometric information. The pathological feature extraction result is processed similarly to generate the key matrix K, representing the pathological features in the image. Specifically, the query matrix Q and the key matrix K perform a dot product operation, and the result is normalized by the Softmax function. In this way, a weighted attention weight matrix can be obtained. This process is calculated by the following formula:
[0061] Wherein, d is the scaling factor, which is the dimension of the matrix, M prior represents the prior mask to help adjust the priority of feature selection. For example, M prior can adjust the attention of the model to different regions based on the medical knowledge base, ensuring that the model can be more in line with anatomical laws or pathological characteristics when performing feature weighting.
[0062] After the attention weight matrix is calculated, the geometric feature and the pathological feature are fused through the cross-attention mechanism. Specifically, the value matrix V (representing the value of the features in the image) is multiplied by the attention matrix A, and a non-linear transformation is performed through the GELU activation function to generate the final fused feature F fusion . The role of this step is to aggregate different features through the cross-attention mechanism, so that the fused features can better express the key information in the image.
[0063] Next, gated attention calculation is performed based on the fused features to perform feature recombination. At this time, the gating factor G is introduced into the process of feature weighting. The gating factor is used to control the weighting ratio between the geometric feature and the fused feature. The gating factor G is calculated by the following formula G = σ(Wg [F g ; F fusion +b g ) is calculated, where σ is the sigmoid activation function that limits the output between 0 and 1, W g is the weight matrix representing the gating mechanism, F g represents the geometric features, b g is the bias term of the gating mechanism used to adjust the gating factor and can be preset by technical experts. Through this calculation, the value of the gating factor G determines the proportion of geometric features and fused features in the final fusion result. If G is close to 1, it means that the geometric features contribute more to the final fusion; if G is close to 0, it means that the fused features play a more important role in the final result.
[0064] Finally, the geometric features and the fused features are weighted and summed according to the gating factor, and through the formula F final =G⊙F g +(1 - G)⊙F fusion for calculation to obtain the reconstructed feature F final . Exemplarily, assume that in a specific oral image, the shape of the teeth (geometric features) is crucial for diagnosis, while the lesion area (such as dental caries or cracks) is relatively simple and has little impact. In this case, the gating factor may be close to 1, meaning that the reconstructed feature will rely more on the geometric features. Conversely, if the impact of the lesion area is greater, the gating factor will be close to 0, indicating that the reconstructed feature will rely more on the fused features.
[0065] Through the above process, geometric features and fused features are first extracted and weighted and fused through the cross - attention mechanism. Then, the gating mechanism is used to flexibly adjust the weighting ratio of geometric features and fused features, and finally the dual - channel feature extraction result is obtained.
[0066] Step S400: Obtain the acquisition coordinates and acquisition parameters of the imaging dataset, and after mapping the dual - channel feature extraction result to the same feature space according to the acquisition coordinates and acquisition parameters, perform pixel search alignment.
[0067] In the embodiments of the present application, first, the acquisition coordinates and acquisition parameters of the imaging dataset are obtained. The acquisition coordinates refer to the position and perspective of the imaging device during the image acquisition process. Through the built - in positioning system of the device, such as GPS or the positioning sensor of the camera, the acquisition coordinates record the spatial position and shooting angle of the device, and this information provides the basis for the subsequent spatial transformation of the image. In addition, the acquisition parameters include focal length, shooting angle, exposure time, resolution, and sampling frequency, etc., and these parameters have a direct impact on image quality, geometric morphology, and detail display.
[0068] After obtaining the acquisition coordinates and acquisition parameters, the next step is to map the geometric features and pathological features obtained by the dual-channel feature extraction network to the same feature space. Since the imaging conditions, acquisition perspectives, and focal lengths of different modality images are different, there may be spatial differences between the images. Therefore, the acquisition coordinates and parameters are used to align the geometric features and pathological features through geometric transformation techniques. This mapping process involves affine transformation, perspective transformation, or other transformation techniques, which adjust the spatial positions of the image features to a unified feature space, enabling features from different modalities to be compared and fused in the same coordinate system.
[0069] After completing the feature mapping, pixel search alignment is carried out next. In this process, first, a fixed feature is selected and a deviation search space is established based on the acquisition coordinates and acquisition parameters; then, within this space, alignment pixel points are randomly selected according to the coincidence range between the fixed feature and the feature to be aligned; next, the pixel features of the aligned pixel points are extracted, and the mapping coordinates of this point in the feature to be aligned are obtained; finally, using these mapping coordinates and pixel features, further pixel alignment is performed within the search space to obtain the optimal alignment result.
[0070] Furthermore, in the method provided by the application embodiment, the execution of pixel search alignment further includes:
[0071] After selecting the fixed feature, a deviation search space is established according to the acquisition coordinates and the acquisition parameters; the feature to be aligned with the fixed feature is obtained, and alignment pixel points are randomly selected based on the feature coincidence range between the fixed feature and the feature to be aligned; the pixel features of the aligned pixel points are extracted, and the mapping coordinates of the aligned pixel points within the feature to be aligned are obtained; centered on the mapping coordinates, search is carried out within the deviation search space using the pixel features to obtain the search alignment result; pixel search alignment is performed according to the search alignment result.
[0072] In the embodiment of the present application, first, fixed features are selected through a feature detection algorithm. These fixed features are regions in the image with high stability and distinctiveness, such as corner points, edges, textures, or other significant geometric structure features in the image. These features can be extracted from the image through a feature detection algorithm, such as SIFT (Scale-Invariant Feature Transform), etc. The selected fixed features provide a benchmark for subsequent image alignment. These features have a high matching degree in multiple image modalities, so they can be used as the starting point for alignment. Through this process, the fixed features are obtained.
[0073] After selecting the fixed features, a deviation search space is established based on the acquisition coordinates and acquisition parameters. The acquisition coordinates and acquisition parameters record factors affecting the geometric shape of the image, such as the spatial position, viewing angle, focal length, exposure time, etc. of the imaging device. Using this information, the possible range of feature deviations is inferred through geometric transformations (such as affine transformation, perspective transformation, or non-rigid transformation), and then a deviation search space containing all possible deformations is defined.
[0074] Then, the features to be aligned are obtained from the mapped feature space, which come from the image modality related to the fixed features. These features to be aligned may have a certain deviation from the fixed features in space, so adjustments are needed. For precise alignment, a set of alignment pixel points is randomly selected based on the feature coincidence range between the fixed features and the features to be aligned. These alignment pixel points are selected based on the similarity of the feature overlap region and the corresponding region in the image to ensure that the alignment points can correspond in the two modality images. Through this step, the alignment pixel points are obtained.
[0075] After selecting the alignment pixel points, the pixel features are extracted from the image. The pixel features include the color, brightness, texture information, etc. of the pixels. Common methods include calculating the gray-level co-occurrence matrix (GLCM), local binary pattern (LBP), or HOG features (histogram of oriented gradients). These pixel features contribute to the subsequent alignment accuracy, ensuring that the features can be more accurately matched during the search process. By extracting the features of the corresponding pixel points from the fixed features and the features to be aligned, the positions of the alignment pixel points in the image and their corresponding spatial information can be obtained.
[0076] Next, based on the pixel features of the alignment pixel points, their mapped coordinates in the features to be aligned are obtained. These mapped coordinates represent the spatial positions of the alignment pixel points in the other modality image. The mapped coordinates are usually inferred through a geometric transformation model. For example, rigid transformations (rotation, translation) or non-rigid transformations (such as B-spline transformation) are used to determine the correspondence of the alignment pixel points between different images. To enhance the mapping accuracy, optimization techniques such as the least squares method are used to further fine-tune the coordinates and reduce the spatial error.
[0077] Next, centered on the mapped coordinates, pixel search is performed within the deviation search space using the pixel features. The goal is to find the most suitable alignment position within the search space. The pixel search gradually adjusts the positions of the alignment pixel points through an optimization algorithm, such as gradient descent, to minimize the error between the images, thereby achieving the best alignment effect. Through this process, the search alignment result is obtained, which represents the optimal alignment position found within the deviation space.
[0078] Finally, based on the search alignment results, perform pixel search alignment. This step completes the final image alignment by integrating the search results of all aligned pixel points, ensuring that the features in the two images are perfectly aligned at the pixel level.
[0079] Step S500: Perform three-dimensional reconstruction of the oral cavity based on the results of the two-channel feature extraction according to the pixel search alignment results, and perform three-dimensional reconstruction correction through a non-rigid hierarchical fusion framework to generate a corrected three-dimensional reconstruction result.
[0080] In the embodiment of the present application, first, according to the pixel search alignment results, generate preliminary three-dimensional reconstruction data through three-dimensional reconstruction of the oral cavity based on the two-channel feature extraction results. The pixel search alignment results are obtained through the aforementioned pixel-level alignment step, which ensures that images from different modalities can be accurately aligned in space, thus providing a spatial basis for three-dimensional reconstruction. The geometric features (such as the shape and arrangement of teeth and other structural information) and pathological features (such as carious lesions and gingival lesions and other pathological regions) in the images will be constructed through three-dimensional reconstruction algorithms. In this process, the geometric features provide a structural framework for the reconstruction, and the pathological features provide the specific location and morphology of the lesion areas. Commonly used three-dimensional reconstruction methods include multi-view stereo matching and voxel reconstruction. Among them, the voxel reconstruction method converts the two-dimensional information of the image into three-dimensional volume data and uses the depth information of the reconstructed image to create a three-dimensional voxel model. Through this process, preliminary three-dimensional reconstruction data is generated.
[0081] Next, perform three-dimensional reconstruction correction through a non-rigid hierarchical fusion framework. Specifically, first, use the basic layer (hard tissue processing layer) to perform preliminary geometric correction on the hard tissue part to generate a first correction result. Then use the intermediate layer (connective tissue processing layer) to perform micro-deformation correction on the connection part between the hard tissue and the soft tissue to obtain a second correction result. Finally, the surface processing layer (soft tissue processing layer) performs deformation correction on the surface of the model to generate the final third correction result. After the gradual optimization of these three levels, a corrected three-dimensional reconstruction result is finally generated, which provides an accurate three-dimensional model of the oral cavity, covering the comprehensive optimization of hard tissue, connective tissue, and soft tissue, ensuring a high degree of accuracy in the structure and morphology of the model.
[0082] Furthermore, in the method provided by the embodiment of the application, the three-dimensional reconstruction correction through the non-rigid hierarchical fusion framework further includes:
[0083] Perform three-dimensional reconstruction correction using the base layer in the non-rigid hierarchical fusion framework to generate a first correction result, where the base layer is a hard tissue processing layer; perform micro-deformation correction on the basis of the first correction result using the intermediate layer in the non-rigid hierarchical fusion framework to establish a second correction result, where the intermediate layer is a connective tissue processing layer; perform deformation correction on the basis of the second correction result using the surface processing layer in the non-rigid hierarchical fusion framework to establish a third correction result, where the surface processing layer is a soft tissue processing layer.
[0084] In the embodiment of the present application, first, through the base layer in the non-rigid hierarchical fusion framework, that is, the hard tissue processing layer, perform preliminary geometric correction on the hard tissue part (such as teeth, bones, etc.) of the three-dimensional reconstruction. The goal of this layer is to roughly correct the hard tissue through geometric transformations (such as rotation, translation, scaling, etc.) to ensure that its position and shape conform to the actual anatomical structure. The hard tissue processing layer uses methods of rigid transformation and non-rigid transformation, allowing small-range deformation of the morphology of some regions to adapt to the changes in the actual anatomical structure. To achieve this, first compare the reconstructed hard tissue with a reference template (such as a standard oral anatomical model) through image registration technology, and use feature point-based matching methods (such as SIFT, SURF) to determine the spatial relationship between different parts. Then, use optimization algorithms such as the least squares method to adjust the image to generate a first correction result.
[0085] Next, in the intermediate layer, that is, the connective tissue processing layer, perform micro-deformation correction on the connection area between the hard tissue and the soft tissue using the first correction result. Connective tissue usually refers to the transitional area between hard and soft tissues such as teeth and gums, bones. The main task of the intermediate layer is to make fine adjustments to these connection areas to ensure a natural transition between hard and soft tissues and conform to the anatomical structure. For this purpose, non-rigid registration methods are adopted, especially using B-spline interpolation or thin plate splines, which can flexibly adjust the connection area between hard and soft tissues and ensure the naturalness of the transitional part. Through this fine-tuning, a second correction result is generated.
[0086] Finally, in the surface processing layer, that is, the soft tissue processing layer, perform deformation correction on the second correction result, focusing on the adjustment and optimization of the soft tissue (such as gums, oral inner wall, etc.) on the surface of the model. The morphological changes of soft tissues are relatively significant and require careful adjustment. To achieve this goal, shape matching and physical model-based deformation methods are adopted, such as the finite element method and shape optimization technology. These methods simulate the physical properties of soft tissues to adjust their morphology to ensure that the model surface can truly reflect the oral anatomical structure. Optimization methods such as the gradient descent method are used to minimize the deformation error and ensure smooth transition and natural display of the soft tissue area. After the adjustment of this layer, a third correction result is generated, which has completed the comprehensive correction of hard tissue, connective tissue, and soft tissue, ensuring that the three-dimensional reconstructed oral model is more accurate and conforms to the anatomical structure.
[0087] Through these three levels of calibration, the non-rigid hierarchical fusion framework ensures the accuracy of the oral three-dimensional reconstruction model, covering the comprehensive optimization of hard tissues, connective tissues, and soft tissues, and finally generating a calibrated three-dimensional reconstruction result.
[0088] In the embodiments of the present application, in summary, the embodiments of the present application at least have the following technical effects:
[0089] The present application configures a multi-spectral LED array probe to perform oral imaging acquisition and establish an imaging data set. The spectrum of the multi-spectral LED array probe includes visible light, infrared light, and fluorescence imaging. Dental prior knowledge constraints are introduced to process the imaging data set, and adversarial reconstruction of the imaging data set is performed to generate an adversarial reconstruction result. A dual-channel feature extraction network is configured to perform dual-channel feature extraction of the adversarial reconstruction result and establish a dual-channel feature extraction result. The acquisition coordinates and acquisition parameters of the imaging data set are obtained, and after mapping the dual-channel feature extraction result to the same feature space according to the acquisition coordinates and acquisition parameters, pixel search alignment is performed. Oral three-dimensional reconstruction is performed based on the dual-channel feature extraction result according to the pixel search alignment result, and three-dimensional reconstruction calibration is performed through a non-rigid hierarchical fusion framework to generate a calibrated three-dimensional reconstruction result. The present invention solves the technical problem of insufficient accuracy of three-dimensional reconstruction of oral endoscopy images in the prior art, and achieves the technical effect of improving the accuracy of oral three-dimensional reconstruction through a multi-spectral LED array probe, dual-channel feature extraction, pixel alignment, and non-rigid hierarchical fusion calibration.
[0090] Embodiment 2, based on the same inventive concept as the multi-modal fusion-based oral endoscopy image three-dimensional reconstruction method in the foregoing embodiment, as Figure 2 shown, the present application provides a multi-modal fusion-based oral endoscopy image three-dimensional reconstruction system. The system in the embodiments of the present application and the method embodiments are based on the same inventive concept. Among them, the system includes:
[0091] The imaging acquisition module 11 is used to configure a multi-spectral LED array probe, perform oral imaging acquisition, and establish an imaging data set. The spectrum of the multi-spectral LED array probe includes visible light, infrared light, and fluorescence imaging. The adversarial reconstruction module 12 is used to introduce dental prior knowledge constraints for processing the imaging data set, perform adversarial reconstruction of the imaging data set, and generate an adversarial reconstruction result. The feature extraction module 13 is used to configure a dual-channel feature extraction network, perform dual-channel feature extraction of the adversarial reconstruction result, and establish a dual-channel feature extraction result. The pixel search alignment module 14 is used to obtain the acquisition coordinates and acquisition parameters of the imaging data set, map the dual-channel feature extraction result to the same feature space according to the acquisition coordinates and acquisition parameters, and then perform pixel search alignment. The 3D reconstruction correction module 15 is used to perform oral 3D reconstruction based on the dual-channel feature extraction result according to the pixel search alignment result, and perform 3D reconstruction correction through a non-rigid hierarchical fusion framework to generate a corrected 3D reconstruction result.
[0092] Furthermore, the system is also used to implement the following functions:
[0093] Configure a geometric feature channel, use the geometric feature channel to perform geometric feature recognition of the adversarial reconstruction result, and establish a geometric feature extraction result; configure a pathological feature channel, use the pathological feature channel to perform pathological feature recognition of the adversarial reconstruction result, and establish a pathological feature extraction result; synchronize the geometric feature extraction result and the pathological feature extraction result to a cross-attention collaboration channel, perform feature recombination, and establish a dual-channel feature extraction result according to the feature recombination result.
[0094] Furthermore, the system is also used to implement the following functions:
[0095] Establish an attention weight matrix as follows:
[0096]
[0097] Among them, A represents the attention weight matrix, Q represents the query matrix, which is obtained by normalizing and linearly transforming the geometric feature extraction result, K represents the key matrix, which is obtained by normalizing and linearly transforming the pathological feature extraction result, T represents matrix transpose, d is a scaling factor, and M prior represents a prior mask;
[0098] Perform feature fusion on the attention weight matrix and the value matrix to construct a fusion feature:
[0099] F fusion = GELU(A·V);
[0100] F fusion represents the fusion feature, and V is the value matrix;
[0101] Perform gating attention calculation based on the fusion feature and execute feature recombination:
[0102] F final = G ⊙ F g + (1 - G) ⊙ F fusion ;
[0103] where F final represents the recombined feature, G is the gating factor, G = σ(W g [F g ; F fusion + b g ), σ is the activation function, W g represents the weight matrix of the gating mechanism, F g represents the geometric feature, and b g is the bias term of the gating mechanism.
[0104] Furthermore, the system is also used to implement the following functions:
[0105] Build a motion sensitivity map, perform motion artifact correction on the acquired image based on the motion sensitivity map. The motion sensitivity map is as follows:
[0106]
[0107] where M sens (x, y) represents the motion sensitivity of the pixel point (x, y), I(x, y) represents the image intensity of the pixel point (x, y), t i is the sampling time, N represents the set of time points, i represents any time point, and G σ is the Gaussian filter kernel; establish the imaging dataset according to the motion artifact correction result.
[0108] Furthermore, the system is also used to implement the following functions:
[0109] After selecting the fixed feature, establish a deviation search space according to the acquisition coordinates and the acquisition parameters; obtain the feature to be aligned with the fixed feature, randomly select an alignment pixel point based on the feature coincidence range of the fixed feature and the feature to be aligned; extract the pixel feature of the alignment pixel point and obtain the mapping coordinate of the alignment pixel point within the feature to be aligned; use the pixel feature to search within the deviation search space centered on the mapping coordinate to obtain the search alignment result; perform pixel search alignment according to the search alignment result.
[0110] Furthermore, the system is also used to implement the following functions:
[0111] Perform three-dimensional reconstruction correction using the base layer in the non-rigid hierarchical fusion framework to generate a first correction result, where the base layer is a hard tissue processing layer; perform micro-deformation correction on the basis of the first correction result using the intermediate layer in the non-rigid hierarchical fusion framework to establish a second correction result, where the intermediate layer is a connective tissue processing layer; perform deformation correction on the basis of the second correction result using the surface processing layer in the non-rigid hierarchical fusion framework to establish a third correction result, where the surface processing layer is a soft tissue processing layer.
[0112] Further, the system is also used to implement the following functions:
[0113] Perform image self-verification on the oral imaging image to generate a self-verification result; when the self-verification result is a non-passing result, generate a reset acquisition instruction; perform reset acquisition of the image according to the reset acquisition instruction.
[0114] It should be noted that the above order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above describes specific embodiments of this specification. The processes depicted in the drawings do not necessarily require the specific order and continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0115] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0116] This specification and the drawings are only exemplary descriptions of the present application and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.
Claims
1. A three-dimensional reconstruction method for oral endoscopic images with multi-modal fusion, characterized in that, The method includes: Configuring a multi-spectral LED array probe, performing oral imaging acquisition, and establishing an imaging data set. The spectrum of the multi-spectral LED array probe includes visible light, infrared light, and fluorescence imaging; Introducing dental prior knowledge constraints for processing the imaging data set, performing adversarial reconstruction of the imaging data set, and generating an adversarial reconstruction result; Configuring a dual-channel feature extraction network, performing dual-channel feature extraction of the adversarial reconstruction result, and establishing a dual-channel feature extraction result; Obtaining the acquisition coordinates and acquisition parameters of the imaging data set, mapping the dual-channel feature extraction result to the same feature space according to the acquisition coordinates and acquisition parameters, and then performing pixel search alignment; Performing oral three-dimensional reconstruction based on the dual-channel feature extraction result according to the pixel search alignment result, and performing three-dimensional reconstruction correction through a non-rigid hierarchical fusion framework to generate a corrected three-dimensional reconstruction result.
2. The three-dimensional reconstruction method of the multi-modal fusion oral endoscope image according to claim 1, characterized in that, The configuring a dual-channel feature extraction network, performing dual-channel feature extraction of the adversarial reconstruction result, and establishing a dual-channel feature extraction result includes: Configuring a geometric feature channel, using the geometric feature channel to identify the geometric features of the adversarial reconstruction result, and establishing a geometric feature extraction result; Configuring a pathological feature channel, using the pathological feature channel to identify the pathological features of the adversarial reconstruction result, and establishing a pathological feature extraction result; Synchronizing the geometric feature extraction result and the pathological feature extraction result to a cross-attention collaboration channel, performing feature recombination, and establishing a dual-channel feature extraction result according to the feature recombination result.
3. The three-dimensional reconstruction method of the multi-modal fusion oral endoscope image according to claim 2, wherein The synchronizing the geometric feature extraction result and the pathological feature extraction result to a cross-attention collaboration channel and performing feature recombination includes: Establishing an attention weight matrix as follows: Among them, A represents the attention weight matrix, Q represents the query matrix, which is calculated by normalizing and linearly transforming the geometric feature extraction result, K represents the key matrix, which is calculated by normalizing and linearly transforming the pathological feature extraction result, represents the matrix transpose, d is the scaling factor, M prior represents the prior mask; Performing feature fusion on the attention weight matrix and the value matrix to construct a fused feature: F fusion = GELU(A·V); F fusion characterizes the fusion feature, and V is a value matrix; Performing gated attention calculation according to the fused feature and performing feature recombination: F final = G ⊙ F g + (1 - G) ⊙ F fusion ; Among them, F final represents the recombination feature, G is the gating factor, G = σ(W g [F g ; F fusion +b g ), σ is the activation function, W g represents the weight matrix of the gating mechanism, F g represents the geometric feature, b g is the bias term of the gating mechanism.
4. The three-dimensional reconstruction method of the multi-modal fusion oral endoscope image according to claim 1, wherein, The performing oral imaging acquisition and establishing an imaging data set includes: Establishing a motion sensitivity map, performing motion artifact correction on the acquired image based on the motion sensitivity map. The motion sensitivity map is as follows: Among them, M sens (x, y) represents the motion sensitivity of the pixel point (x, y), I(x, y) represents the image intensity of the pixel point (x, y), t i is the sampling time, N represents the set of time points, i represents any time point, G σ is the Gaussian filter kernel; Establishing the imaging data set according to the motion artifact correction result.
5. The three-dimensional reconstruction method of multi-modal fusion oral endoscopy images according to claim 1, wherein, The performing pixel search alignment includes: After selecting a fixed feature, establishing a deviation search space according to the acquisition coordinates and the acquisition parameters; Obtaining the feature to be aligned with the fixed feature, and randomly selecting an alignment pixel point based on the feature coincidence range of the fixed feature and the feature to be aligned; Extracting the pixel feature of the alignment pixel point, and obtaining the mapping coordinates of the alignment pixel point within the feature to be aligned; Searching within the deviation search space using the pixel feature with the mapping coordinates as the center to obtain a search alignment result; Performing pixel search alignment according to the search alignment result.
6. The three-dimensional reconstruction method of the multi-modal fusion oral endoscope image according to claim 1, wherein The performing three-dimensional reconstruction correction through a non-rigid hierarchical fusion framework includes: Using the basic layer in the non-rigid hierarchical fusion framework to perform three-dimensional reconstruction correction to generate a first correction result. The basic layer is a hard tissue processing layer; Using the intermediate layer in the non-rigid hierarchical fusion framework to perform micro-deformation correction on the basis of the first correction result to establish a second correction result. The intermediate layer is a connective tissue processing layer; The surface treatment layer in the non-rigid hierarchical fusion framework is used to perform deformation correction on the basis of the second correction result to establish a third correction result, and the surface treatment layer is a soft tissue treatment layer.
7. The three-dimensional reconstruction method of the multi-modal fusion oral endoscope image according to claim 1, characterized in that Performing oral imaging acquisition to establish an imaging data set, including: Performing image self-verification on the oral imaging image to generate a self-verification result; When the self-verification result is a non-passing result, a reset acquisition instruction is generated; Performing reset acquisition of the image according to the reset acquisition instruction.
8. Three-dimensional reconstruction system for oral endoscopy images with multimodal fusion, characterized in that, The system includes: An imaging acquisition module, configured to configure a multi-spectral LED array probe, perform oral imaging acquisition, and establish an imaging data set. The spectrum of the multi-spectral LED array probe includes visible light, infrared light, and fluorescence imaging; An adversarial reconstruction module, configured to introduce dental prior knowledge constraints to process the imaging data set, perform adversarial reconstruction of the imaging data set, and generate an adversarial reconstruction result; A feature extraction module, configured to configure a dual-channel feature extraction network, perform dual-channel feature extraction on the adversarial reconstruction result, and establish a dual-channel feature extraction result; A pixel search alignment module, configured to obtain the acquisition coordinates and acquisition parameters of the imaging data set, and perform pixel search alignment after mapping the dual-channel feature extraction result to the same feature space according to the acquisition coordinates and acquisition parameters; A three-dimensional reconstruction correction module, configured to perform three-dimensional reconstruction of the oral cavity based on the dual-channel feature extraction result according to the pixel search alignment result, and perform three-dimensional reconstruction correction through a non-rigid hierarchical fusion framework to generate a corrected three-dimensional reconstruction result.