Oral disease analysis method and system based on AI intelligent recognition

By integrating oral scanning, endoscope and CBCT equipment into the dental chair, and combining it with a dynamic weight fusion mechanism and deep convolutional neural network, the problems of low diagnostic accuracy of single-modality data and scattered equipment are solved, and efficient and reliable diagnosis of multi-modal data and optimization of diagnosis and treatment processes are achieved.

CN120853872APending Publication Date: 2025-10-28FOSHAN CHUANGXIN MEDICAL APP CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510730860.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

The limitations of single-modality data diagnosis in existing oral disease diagnosis lead to low accuracy, the fixed weight strategy in multimodal data fusion cannot adapt to quality fluctuations, and the dispersion of diagnostic and treatment equipment leads to inefficient processes.

Method used

The dental chair integrates an intraoral scanning module to acquire three-dimensional surface data, while the endoscope module acquires soft tissue images. These images are then spatially registered using CBCT images. A dynamic weighted fusion mechanism is employed, and a deep convolutional neural network is used for multimodal data analysis to output a comprehensive diagnostic report.

Benefits of technology

It improves the early caries identification rate and periodontitis diagnosis compliance rate, shortens the diagnosis time, reduces the number of patient visits and equipment costs, and improves the reliability of analysis results and process efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853872A_ABST
    Figure CN120853872A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of oral medical treatment, in particular to an oral disease analysis method and system based on AI intelligent recognition. Comprising the following steps that oral cavity three-dimensional surface data are obtained through an oral scanning module integrated with a dental chair, and the oral scanning module adopts a structured light coding and multi-view fusion algorithm; an oral soft tissue high-definition image is collected through an endoscope module, and an AI analysis unit is arranged in an endoscope to recognize mucous membrane lesion features in real time; calling CBCT image data of the patient, and aligning a three-dimensional image with oral scanning data through a spatial registration algorithm; the mouth scanning data, the endoscope image and the CBCT data are subjected to multi-modal fusion by adopting a dynamic weight fusion mechanism. According to the scheme, the problem of accurate registration of cross-modal data, the lack of a dynamic quality evaluation mechanism of heterogeneous data and the real-time collaboration bottleneck of dispersed equipment can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of oral medical technology, specifically to an AI-based intelligent recognition method and system for analyzing oral diseases. Background Technology

[0002] The most prominent problem currently facing the field of oral disease diagnosis is the limitation of single-modal data diagnosis. Clinical practice shows that diagnostic methods relying solely on single data sources such as CBCT or intraoral scanning have significant drawbacks: CBCT has significant errors when observing fine structures such as the labial and buccal alveolar bone, and its spatial resolution limitations lead to blurred local margins and difficulties in quantitative observation; while using intraoral scanning data alone cannot comprehensively assess soft tissue conditions such as mucosal lesions. Current statistics show that this single-modal diagnostic approach maintains an overall accuracy rate of only about 78%, and the recognition rate for small lesions such as early caries is even lower at 72.5%. Specific manifestations of technical bottlenecks include: incomplete anatomical coverage: although CBCT can clearly display bone tissue, it lacks sufficient contrast for soft tissues; intraoral scanning excels at tooth surface reconstruction but struggles to assess intraosseous lesions. This blind spot effect leads to a clinical missed diagnosis rate of 15-23%. Error accumulation effect: independently used diagnostic equipment suffers from the accumulation of systematic errors, such as the superposition of measurement deviations caused by CBCT metal artifacts and motion artifacts from intraoral scanning, reducing the reliability of important decisions such as implant planning by 12-15%. The dispersed equipment system forces patients to move between different examination units, taking an average of 30-60 minutes, and there is a delay of about 200ms in the data transmission process, which seriously affects the continuity of the diagnosis and treatment process.

[0003] The industry urgently needs a technical solution that can integrate multi-source oral data to address the challenges of accurate registration of cross-modal data, the lack of dynamic quality assessment mechanisms for heterogeneous data, and the bottleneck of real-time collaboration among distributed devices. Summary of the Invention

[0004] This invention addresses the problems existing in the prior art by providing an AI-based intelligent recognition method for oral disease analysis, comprising the following steps: Three-dimensional oral surface data are acquired through an intraoral scanning module integrated into the dental chair. The intraoral scanning module employs a structured light coding and multi-view fusion algorithm. High-definition images of oral soft tissues are acquired through an endoscope module, and the endoscope has a built-in AI analysis unit to identify mucosal lesion characteristics in real time. Retrieve the patient's CBCT image data and align the 3D images with the intraoral scan data using a spatial registration algorithm; A dynamic weighted fusion mechanism is used to perform multimodal fusion of intraoral scan data, endoscopic images and CBCT data, where the weight of each data source is dynamically adjusted according to the acquisition quality index; By analyzing the fused multimodal features through deep convolutional neural networks, a comprehensive diagnostic report is output, including disease localization, severity assessment, and treatment recommendations. The system integrates an oral scanner, endoscope, and vital sign monitoring module through a dental chair carrier, enabling synchronous data collection and analysis throughout the entire diagnosis and treatment process.

[0005] Preferably, the endoscopic AI analysis unit adopts an improved YOLOv5s architecture, which realizes dynamic monitoring of mucosal lesions through a three-stage feature pyramid network. Specifically, the first-level pyramid extracts macroscopic morphological features, the second-level pyramid analyzes vascular distribution patterns, and the third-level pyramid identifies microscopic texture abnormalities. The three-stage features are weighted and fused through an attention mechanism.

[0006] More preferably, the spatial registration algorithm adopts an improved ICP registration technique, combined with M estimation to suppress the influence of outliers. The specific implementation method is as follows: first, an initial corresponding point set is established based on curvature features; then, the weight of outliers is reduced through the Huber weight function; and finally, the Levenberg-Marquardt optimization algorithm is used to complete the accurate registration.

[0007] Further preferably, the CBCT image preprocessing includes: temporal feature extraction based on spiking neuron networks, using the LIF neuron model to capture the dynamic changes of the image sequence; contrast enhancement of multi-scale convolution kernels, using three sets of parallel convolution kernels of 1×1, 3×3 and 5×5 to extract features of different granularities; missing region completion based on Poisson reconstruction, and flux domain optimization to achieve accurate repair of metal artifact regions.

[0008] More preferably, the calculation formula for the dynamic weight fusion mechanism is: ; in: The feature matrix after fusion; The feature vector representing the i-th data source: i=1 corresponds to oral scan data, i=2 corresponds to endoscopic data, i=3 corresponds to CBCT data; The dynamic weighting coefficient is calculated as follows: ; The quality score for the i-th data source includes three dimensions: resolution, signal-to-noise ratio, and motion artifact level. This is an adjustment factor, with a value range of 0.5-1.2.

[0009] Further preferably, the quality score The calculation formula is: ; in : Let be the standardized signal-to-noise ratio value of the i-th data source; This refers to spatial resolution metrics; Score the degree of motion artifacts; 、 、 These are the weighting coefficients for each dimension, with default values ​​of 0.4, 0.5, and 0.1.

[0010] More preferably, the output layer of the deep convolutional neural network employs an improved diagnostic scoring function: ; in: For the diagnostic score of the kth disease category; Sigmoid activation function ; The contribution of the m-th feature to the k-th disease category; The feature importance weights are obtained by optimizing them on multi-institutional data using a federated learning framework; This is a bias term.

[0011] An analysis system, applied to an AI-based intelligent recognition method for analyzing oral diseases as described in any of the above, comprising: Data acquisition module: integrates an oral scanning unit, an endoscope unit, and a CBCT interface. The oral scanning unit adopts blue light structured light projection technology with a sampling frequency of ≥15fps; the endoscope unit is equipped with a 4K CMOS sensor and a 470nm-940nm multispectral illumination system; the CBCT interface supports the DICOM3.0 standard protocol. The preprocessing module includes a point cloud filtering submodule, which uses a statistical outlier removal algorithm; an image registration submodule, which implements spatial alignment of multimodal data; and a feature normalization submodule, which performs Min-Max normalization and Z-score normalization. Dynamic fusion module: Configures a quality assessment engine, a weight calculation engine, and a feature fusion engine, supporting real-time weight adjustment and visualization of fusion results; AI Analysis Module: Deploys a multi-task deep learning network to process caries detection, periodontal disease assessment and mucosal lesion identification tasks in parallel. The network structure includes 121 convolutional layers and 8 attention modules. Report generation module: Generates interactive reports that include 3D reconstructed views, lesion heat maps, and treatment plan comparisons, supporting AR / VR display and electronic signature confirmation.

[0012] More preferably, the endoscope unit adopts a dual-channel optical system, with the visible light channel using an RGB three-band beam splitter and the near-infrared channel equipped with a 900nm bandpass filter. The dual-channel images are fused at the feature level through a pulse-coupled neural network.

[0013] More preferably, the system optimizes the hyperparameters of the neural network using a quantum evolution algorithm, specifically including: using qubit encoding to represent the number of network layers, convolution kernel size, and activation function type; and using a quantum rotation gate to update the parameters.

[0014] Technical effects: This invention proposes an AI-based intelligent recognition method for oral disease analysis. It acquires three-dimensional oral surface data through an intraoral scanning module integrated into the dental chair, while an endoscope module captures high-resolution soft tissue images and identifies mucosal lesions in real time. CBCT image data is retrieved and aligned using a spatial registration algorithm. A dynamic weighted fusion mechanism is employed for multimodal data fusion, and a comprehensive diagnostic report is output via a deep convolutional neural network. The system integrates an intraoral scanning module, endoscope, and vital sign monitoring module within the dental chair, enabling simultaneous data acquisition and analysis throughout the entire process. This method addresses the technical problems of traditional oral diagnosis, which relies on single-modal data, failing to comprehensively reflect the oral condition and thus limiting diagnostic accuracy; the use of a fixed weight strategy in multimodal data fusion, which cannot adapt to quality fluctuations from different data sources, affecting the reliability of analysis results; and the inefficiency caused by the dispersed and independent nature of diagnostic and treatment equipment, requiring multiple patient referrals for examination. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart of the oral disease analysis method based on AI intelligent recognition in this application; Figure 2 This is a flowchart illustrating the oral disease analysis method based on AI intelligent recognition proposed in this application. Detailed Implementation

[0017] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0018] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, operations, elements, components and / or collections thereof.

[0019] See also Figure 1 and Figure 2 Traditional technical solutions, for example, have the following technical problems: Traditional oral diagnosis relies on single-modal data, such as CBCT or oral scans alone, which cannot fully reflect the oral condition and thus limits the accuracy of diagnosis.

[0020] Using a fixed-weight strategy when fusing multimodal data cannot adapt to the quality fluctuations of different data sources, affecting the reliability of the analysis results.

[0021] The diagnostic and treatment equipment is scattered and independent, requiring patients to be transferred to multiple hospitals for examinations, resulting in low efficiency.

[0022] Based on this, this application provides an AI-based intelligent recognition method for analyzing oral diseases, characterized by the following steps: Three-dimensional oral surface data are acquired through an intraoral scanning module integrated into the dental chair. The intraoral scanning module employs a structured light coding and multi-view fusion algorithm. High-definition images of oral soft tissues are acquired through an endoscope module, and the endoscope has a built-in AI analysis unit to identify mucosal lesion characteristics in real time. Retrieve the patient's CBCT image data and align the 3D images with the intraoral scan data using a spatial registration algorithm; A dynamic weighted fusion mechanism is used to perform multimodal fusion of intraoral scan data, endoscopic images and CBCT data, where the weight of each data source is dynamically adjusted according to the acquisition quality index; By analyzing the fused multimodal features through deep convolutional neural networks, a comprehensive diagnostic report is output, including disease localization, severity assessment, and treatment recommendations. The system integrates an oral scanner, endoscope, and vital sign monitoring module through a dental chair carrier, enabling synchronous data collection and analysis throughout the entire diagnosis and treatment process.

[0023] It is worth mentioning that this embodiment proposes an AI-based intelligent recognition method for oral disease analysis. The method acquires three-dimensional oral surface data (10-50μm accuracy) through an intraoral scanning module integrated into the dental chair, while an endoscope module acquires high-resolution soft tissue images and identifies mucosal lesions in real time. CBCT image data is retrieved and aligned using a spatial registration algorithm. A dynamic weighted fusion mechanism is employed for multimodal data fusion, and finally, a comprehensive diagnostic report is output through a deep convolutional neural network. The system integrates an intraoral scanning module, an endoscope, and a vital signs monitoring module within the dental chair carrier, enabling simultaneous data acquisition and analysis throughout the entire process.

[0024] The technical effects of the above solution include: By fusing trimodal data—oral scan, endoscopy, and CBCT—the system can improve the early caries identification rate and increase the diagnostic accuracy of periodontitis. A dynamic weighting mechanism ensures that the system maintains high analytical accuracy even when the quality of single-modal data deteriorates.

[0025] Integrated dental chair solutions can shorten diagnosis time, reduce the number of patient visits, and lower equipment procurement costs.

[0026] Traditional technical solutions have the following technical problems: Conventional YOLO models suffer from high false negative rates in oral mucosal detection due to the large differences in lesion size, ranging from millimeter-scale leukoplakia to centimeter-scale erosion. Furthermore, interference from ambient light in the oral environment leads to unstable feature extraction, resulting in a high false positive rate for traditional single-stage networks.

[0027] Based on this, the endoscopic AI analysis unit adopts an improved YOLOv5s architecture and realizes dynamic monitoring of mucosal lesions through a three-stage feature pyramid network. Specifically, the first-level pyramid extracts macroscopic morphological features, the second-level pyramid analyzes vascular distribution patterns, and the third-level pyramid identifies microscopic texture abnormalities. The three-stage features are weighted and fused through an attention mechanism.

[0028] It is worth mentioning that this embodiment limits the endoscopic AI analysis unit to adopt a modified YOLOv5s architecture, which realizes mucosal lesion monitoring through a three-stage feature pyramid network. The first stage extracts macroscopic morphological features, such as ulcer area; the second stage analyzes vascular distribution patterns and capillary density / direction; and the third stage identifies microscopic texture abnormalities and the degree of epithelial keratinization. The three-stage features are weighted and fused through an attention mechanism, achieving a recognition accuracy of 90.7% and a false positive rate of less than 3.2%.

[0029] The technical effects of the above solution include: The three-stage pyramid network can improve the detection rate of lesions smaller than 1 mm, which is superior to traditional models. The attention-weighted fusion mechanism can reduce the misjudgment rate of blood vessel morphology, and under the same hardware conditions, the inference speed can be maintained at 25 FPS, meeting the requirements of real-time detection.

[0030] Traditional technical solutions have the following technical problems: Oral scans are susceptible to motion artifacts and saliva interference, resulting in high registration errors with traditional ICP, which affects implant planning accuracy. Traditional methods use feature point matching for pre-registration, but this does not address the local distortion caused by CBCT metal artifacts. While trimming ICP suppresses outliers, it results in the loss of effective features, leading to a high registration failure rate.

[0031] Based on this, the spatial registration algorithm adopts an improved ICP registration technique, combined with M estimation to suppress the influence of outliers. The specific implementation method is as follows: first, an initial corresponding point set is established based on curvature features; then, the weight of outliers is reduced through the Huber weight function; and finally, the Levenberg-Marquardt optimization algorithm is used to complete the accurate registration.

[0032] It is worth mentioning that the spatial registration algorithm in this embodiment adopts an improved ICP technique, combined with M-estimation to suppress outliers. Specifically, it includes: establishing an initial corresponding point set based on curvature features → reducing the influence of outliers using the Huber weight function → completing accurate registration through Levenberg-Marquardt optimization, achieving an error of <15μm and a processing speed of 30FPS.

[0033] The technical effects of the above solution include: M-estimation ensures the registration algorithm remains stable even with high outlier contamination, reducing the control failure rate. Multi-stage optimization strategies improve registration accuracy in the jaw region, meeting the precision requirements of implant surgery. GPU acceleration reduces the time required for full-arch registration.

[0034] Traditional technical solutions have the following technical problems: CBCT images exhibit beam hardening artifacts, and traditional filtering methods easily lead to loss of detail, resulting in blurred trabecular bone structures. Existing techniques employ CNN preprocessing, but fail to utilize the temporal feature extraction advantages of spiking neural networks. Star-shaped artifacts produced by metallic restorations affect the diagnostic accuracy of most implant cases.

[0035] Based on this, the CBCT image preprocessing includes: temporal feature extraction based on spiking neuron networks, using the LIF neuron model to capture the dynamic changes of the image sequence; contrast enhancement of multi-scale convolution kernels, using three sets of parallel convolution kernels of 1×1, 3×3 and 5×5 to extract features of different granularities; missing region completion based on Poisson reconstruction, and flux domain optimization to achieve accurate repair of metal artifact regions.

[0036] It is worth mentioning that the CBCT preprocessing in this embodiment includes: LIF neuron model temporal feature extraction → multi-scale contrast enhancement of three sets of parallel convolution kernels (1×1 / 3×3 / 5×5) → gradient domain Poisson reconstruction for metal artifact repair, which improves the PSNR of the reconstructed image by more than 12dB.

[0037] The technical effects of the above solution include: Dynamic features captured by spiking neural networks can reduce errors in bone mineral density measurements. Multi-scale convolutional combinations can effectively separate the root and alveolar bone boundaries, improving the Jaccard similarity coefficient. Poisson reconstruction can reduce the area of ​​artifacts around metal crowns, meeting the ISO 16142-2 image quality standard.

[0038] Traditional technical solutions have the following technical problems: Traditional technical solutions use fixed-weight fusion: for example, oral scan 0.6 + endoscopy 0.3 + CBCT 0.1, which is difficult to adapt to individual differences.

[0039] When multimodal conflicts occur, such as motion blurring in oral scans but clear CBCT, the signal-to-noise ratio of the fusion results obtained by traditional methods may decrease.

[0040] In addition, the existing dynamic weighting method does not have a quality assessment system designed for the characteristics of oral imaging.

[0041] Based on this, the calculation formula for the dynamic weight fusion mechanism is as follows: ; in: The feature matrix after fusion; The feature vector representing the i-th data source: i=1 corresponds to oral scan data, i=2 corresponds to endoscopic data, i=3 corresponds to CBCT data; The dynamic weighting coefficient is calculated as follows: ; The quality score for the i-th data source includes three dimensions: resolution, signal-to-noise ratio, and motion artifact level. This is an adjustment factor, with a value range of 0.5-1.2.

[0042] Feature fusion matrix F: Serves as a unified representation of multimodal data; its dimensionality depends on the sum of the dimensions of the input features. For example, when scanning data... ∈R(256×256×3), Endoscopic data ∈R(1024×1024×4), CBCT data When the modalities are in R(512×512×512), the system will automatically use a spatial transformation network (STN) to unify each modality into the feature space of R(512×512×32) and then perform weighted fusion.

[0043] Dynamic weighting coefficients Normalized weight allocation is implemented using the softmax function, with key parameters including: Quality rating The objective quality index reflecting each modal data is calculated from the formula in claim 6.

[0044] Regulatory factors : Controlling the steepness of the weight distribution, as verified by experiments, when A value of 0.8 achieves the best balance between discrimination and stability.

[0045] The summation term in the denominator ensures that the sum of all weights is 1, thus avoiding distortion of the characteristic amplitude.

[0046] The above scheme employs a real-time weight adjustment mechanism, recalculating the weights every 5ms. When a certain modality is detected... If the decrease exceeds the threshold, such as when motion artifacts increase due to patient movement during oral scan, the weights are redistributed within 3 sampling cycles; mixed precision calculation is achieved through Tensor Cores, with FP16 accumulation + FP32 weights, which can reduce the time consumption of a single fusion operation; when a certain modality completely fails, the system automatically switches to dual-modal fusion mode and triggers the redundancy check algorithm.

[0047] In periodontal disease diagnosis, when endoscopic bleeding leads to a decrease in image quality, the system automatically increases the intraoral scan weight from 0.4 to 0.7 to reduce fluctuations in diagnostic accuracy. Dynamic weighting improves decision confidence by 38% in multimodal conflict scenarios, such as when the difference between intraoral scan and CBCT measurement results is >15%. Compared to fixed-weight fusion, this approach increases the AUC value for early caries from 0.79 to 0.88.

[0048] The technical effects of the above solution include: When the SNR of the oral scan data is less than 15 dB, the system automatically increases the CBCT weight from 0.3 to 0.7 to ensure fusion stability. Dynamic adjustment achieves a diagnostic consistency Kappa value of 0.85 in multimodal conflict scenarios. The formula is highly interpretable and can improve the consistency between weight allocation and expert experience.

[0049] Traditional technical solutions have the following technical problems: Existing technical solutions use subjective scoring, which leads to inconsistent weighting; oral imaging quality assessment lacks standardized indicators, especially insufficient quantification of motion artifacts; and different modalities have inconsistent dimensions, such as μm for oral scans vs. mm for CBCT, which can easily lead to bias when directly weighted.

[0050] Based on this, the quality score The calculation formula is: ; in : Let be the standardized signal-to-noise ratio value of the i-th data source; This refers to spatial resolution metrics; Score the degree of motion artifacts; 、 、 These are the weighting coefficients for each dimension, with default values ​​of 0.4, 0.5, and 0.1.

[0051] Signal-to-noise ratio The unit used is decibel (dB), and the calculation method is as follows: ; The signal region is automatically selected from the Region of Interest (ROI), such as the enamel layer of teeth, while the noise region is selected from the uninhabited background area. The local variance is calculated using the sliding window method. resolution The quantitative standard is: Oral scan data: effective sampling point spacing; Endoscope: line pairs / mm (lp / mm), calibrated using the USAF 1951 resolution test chart; CBCT: isotropic voxel size (mm); artifact level Normalized score 0-1, based on: motion artifacts: inter-frame displacement calculated by optical flow method; metal artifacts: star-shaped stripe density detected by Radon transform; beam hardening: analysis of CT value linearity deviation.

[0052] Coefficient optimization process: =0.4: Emphasizes the fundamental role of signal-to-noise ratio, but does not over-rely on it, avoiding high SNR but excessively high scores for blurry images.

[0053] =0.5: Highlights the critical importance of resolution, especially its impact on the detection of tooth defect edges.

[0054] =0.1: Moderately penalize artifacts to prevent a single artifact from causing a sharp drop in the score.

[0055] The coefficient combination was optimized using orthogonal experimental design, resulting in a Spearman correlation coefficient of 0.91 between the scoring results and the subjective evaluations of the five experts.

[0056] The quality assessment process for the above formula includes: Preprocessing stage: All modal data are uniformly converted to DICOM format, and the acquisition parameters are extracted from the metadata; Feature extraction: SNR is calculated in parallel using frequency domain analysis, resolution is fitted using edge spread function, and artifacts are detected using convolutional neural network; Standardization: Normalize each indicator to a 0-100 point scale to eliminate the influence of dimensions; Weighted summation: Calculate the final result according to the formula. Write it to the "QualityScore" field of the DICOM label.

[0057] The technical effects of the above solution include: An objective scoring system can reduce the weight variation coefficient of data collected by different devices; prioritizing the resolution index β2 ensures a reasonable weighting of anatomical structure clarity; and automatically acquiring parameters through the DICOM standard interface can reduce evaluation time.

[0058] Traditional technical solutions have the following technical problems: The limited amount of data from a single medical institution results in poor model generalization ability; existing technical solutions employ centralized training, which requires the transmission of patient data, violating privacy regulations such as GDPR; and the fixed feature importance weight γ cannot adapt to regional differences in disease spectrum.

[0059] Based on this, the output layer of the deep convolutional neural network adopts an improved diagnostic scoring function: ; in: For the diagnostic score of the kth disease category; Sigmoid activation function ; The contribution of the m-th feature to the k-th disease category; The feature importance weights are obtained by optimizing them on multi-institutional data using a federated learning framework; This is a bias term.

[0060] Neural network architecture: Feature contribution Features from the bottleneck layer of a 121-layer CNN include: low-level features: edges / textures, convolutional layers 1-20; mid-level features: anatomical structures, convolutional layers 21-80; and high-level features: pathological patterns, convolutional layers 81-121.

[0061] Each feature is weighted by an attention gating mechanism and then input into the scoring function: Feature Importance. Dynamic updates are achieved through federated learning. The specific mechanism includes: local training: each medical institution calculates on its private data. Gradient; Secure aggregation: Transmit gradient mean using homomorphic encryption; Global update: Integrate cross-institutional parameters every 24 hours.

[0062] Protecting privacy while enabling Representativeness of the group: bias term Logit transformed values ​​of prior disease probabilities, initialized based on epidemiological data, such as dental caries. =ln(0.32 / 0.68).

[0063] Characteristics of the Sigmoid function: ; By constraining the score to the (0,1) interval, the probability of developing the disease can be intuitively reflected.

[0064] Decision threshold optimization: Determining the optimal threshold (e.g., for periodontitis) through ROC curve analysis. >0.63) Multi-label support: Simultaneously outputs K disease scores. In this system, K=12, which is achieved through a multi-head mechanism.

[0065] Visualization of enamel demineralization features contributed the most to the diagnosis of dental caries, with a weight of 0.38.

[0066] It is worth mentioning that the existing system does not integrate endoscopy and CBCT analysis functions, requiring third-party software bridging; and the existing hardware architecture does not support real-time multimodal data synchronization; traditional reports lack 3D visualization, resulting in low efficiency in doctor-patient communication.

[0067] Based on this, this embodiment provides an analysis system applied to an AI-based intelligent recognition method for analyzing oral diseases as described in any of the above embodiments, comprising: Data acquisition module: integrates an oral scanning unit, an endoscope unit, and a CBCT interface. The oral scanning unit adopts blue light structured light projection technology with a sampling frequency of ≥15fps; the endoscope unit is equipped with a 4K CMOS sensor and a 470nm-940nm multispectral illumination system; the CBCT interface supports the DICOM3.0 standard protocol. The preprocessing module includes a point cloud filtering submodule, which uses a statistical outlier removal algorithm; an image registration submodule, which implements spatial alignment of multimodal data; and a feature normalization submodule, which performs Min-Max normalization and Z-score normalization. Dynamic fusion module: Configures a quality assessment engine, a weight calculation engine, and a feature fusion engine, supporting real-time weight adjustment and visualization of fusion results; AI Analysis Module: Deploys a multi-task deep learning network to process caries detection, periodontal disease assessment and mucosal lesion identification tasks in parallel. The network structure includes 121 convolutional layers and 8 attention modules. Report generation module: Generates interactive reports that include 3D reconstructed views, lesion heat maps, and treatment plan comparisons, supporting AR / VR display and electronic signature confirmation.

[0068] It is worth mentioning that this embodiment limits the system that implements the aforementioned method to include: a blue light scanning unit, a 4K endoscope, and a data acquisition module with a DICOM 3.0 interface; a preprocessing module for point cloud filtering / image registration / feature standardization; a dynamic fusion module for quality assessment / weight calculation / feature fusion engine; an AI analysis module with 121 layers of CNN and 8 attention modules; and an interactive report generation module that supports AR / VR display.

[0069] The technical effects of the above solution include: Multispectral endoscopy can enhance the contrast of mucosal lesions and improve the clarity of vascular visualization; the integrated process controls data latency to <50ms, meeting the needs of real-time interaction; AR reports can improve patients' understanding of treatment plans.

[0070] Traditional technical solutions have the following technical problems: Existing dual-channel systems suffer from color differences, leading to high errors in tissue boundary positioning; conventional fusion methods, such as wavelet transform, result in decreased fusion quality in a moist oral environment; and the evaluation of near-infrared contrast agent (ICG) perfusion lacks quantitative standards and relies on subjective judgment.

[0071] Based on this, the endoscope unit adopts a dual-channel optical system. The visible light channel uses an RGB three-band beam splitter, and the near-infrared channel is equipped with a 900nm bandpass filter. The dual-channel images are fused at the feature level through a pulse-coupled neural network.

[0072] It is worth mentioning that: In this embodiment, the endoscope unit is limited to a dual-channel optical system. The visible light channel (RGB beam splitter) and the near-infrared channel (900nm bandpass filter) are fused at the feature level through a pulse-coupled neural network to solve the modal conflict problem.

[0073] The technical effects of the above solution include: The timing encoding of the spiking neural network reduces the alignment error of the two channels to <5μm; the near-infrared channel achieves a 100μm-level sensitivity for visualizing vascular networks, aiding in the early diagnosis of periodontitis. In fluorescence-guided surgery, it can improve the accuracy of tumor boundary identification.

[0074] Traditional technical solutions have the following technical problems: Traditional grid search hyperparameter optimization is time-consuming; classic evolutionary algorithms are prone to getting trapped in local optima, and the accuracy loss is relatively high after model compression.

[0075] Edge device deployment requires a lightweight model of less than 50MB, and existing technologies struggle to balance size and performance.

[0076] Based on this, the system optimizes the hyperparameters of the neural network through a quantum evolution algorithm, specifically including: using qubit encoding to represent the number of network layers, convolution kernel size, and activation function type; and using a quantum rotation gate to update parameters.

[0077] It is worth mentioning that: This embodiment limits the system to optimize the hyperparameters of the neural network through quantum evolution algorithm, uses qubit encoding to represent the number of network layers / convolution kernel size / activation function type, updates parameters through quantum rotation gate, and uses parallel quantum crossover operator to maintain population diversity, so that the model can improve the inference speed by 40% and reduce the size by 35% while maintaining 94% accuracy.

[0078] The technical effects of the above solution include: Quantum parallel computing will reduce hyperparameter search time. The optimal model parameter combination obtained through Pareto front optimization can further reduce computational cost. It also supports deployment on embedded platforms such as NVIDIA Jetson, reducing power consumption.

[0079] Unless otherwise specified, the equipment components involved in the above embodiments are all conventional equipment components, and the connection methods and control methods involved are all conventional connection methods and control methods unless otherwise specified.

[0080] The present invention has been described in detail above with reference to the embodiments. However, those skilled in the art will understand that, without departing from the spirit of the present invention, various specific parameters in the above embodiments can be changed to form multiple specific embodiments, all of which are common variations of the present invention, and will not be described in detail here.

Claims

1. A method for analyzing oral diseases based on AI intelligent recognition, characterized in that, Includes the following steps: Three-dimensional oral surface data are acquired through an intraoral scanning module integrated into the dental chair. The intraoral scanning module employs a structured light coding and multi-view fusion algorithm. High-definition images of oral soft tissues are acquired through an endoscope module, and the endoscope has a built-in AI analysis unit to identify mucosal lesion characteristics in real time. Retrieve the patient's CBCT image data and align the 3D images with the intraoral scan data using a spatial registration algorithm; A dynamic weighted fusion mechanism is used to perform multimodal fusion of intraoral scan data, endoscopic images and CBCT data, where the weight of each data source is dynamically adjusted according to the acquisition quality index; By analyzing the fused multimodal features through deep convolutional neural networks, a comprehensive diagnostic report is output, including disease localization, severity assessment, and treatment recommendations. The system integrates an oral scanner, endoscope, and vital sign monitoring module through a dental chair carrier, enabling synchronous data collection and analysis throughout the entire diagnosis and treatment process.

2. The method for analyzing oral diseases based on AI intelligent recognition according to claim 1, characterized in that, The endoscopic AI analysis unit adopts an improved YOLOv5s architecture and achieves dynamic monitoring of mucosal lesions through a three-stage feature pyramid network. Specifically, the first-level pyramid extracts macroscopic morphological features, the second-level pyramid analyzes vascular distribution patterns, and the third-level pyramid identifies microscopic texture abnormalities. The three-stage features are weighted and fused through an attention mechanism.

3. The method for analyzing oral diseases based on AI intelligent recognition according to claim 1, characterized in that, The spatial registration algorithm adopts an improved ICP registration technique, combined with M estimation to suppress the influence of outliers. The specific implementation method is as follows: first, an initial corresponding point set is established based on curvature features; then, the weight of outliers is reduced through the Huber weight function; and finally, the Levenberg-Marquardt optimization algorithm is used to complete the accurate registration.

4. The method for analyzing oral diseases based on AI intelligent recognition according to claim 1, characterized in that, The CBCT image preprocessing includes: temporal feature extraction based on spiking neuron networks, using the LIF neuron model to capture the dynamic changes of the image sequence; contrast enhancement using multi-scale convolution kernels, using three sets of parallel convolution kernels (1×1, 3×3, and 5×5) to extract features of different granularities; missing region completion based on Poisson reconstruction, and flux domain optimization to achieve accurate repair of metal artifact regions.

5. The method for analyzing oral diseases based on AI intelligent recognition according to claim 1, characterized in that, The calculation formula for the dynamic weight fusion mechanism is as follows: ; in: The feature matrix after fusion; The feature vector representing the i-th data source: i=1 corresponds to oral scan data, i=2 corresponds to endoscopic data, i=3 corresponds to CBCT data; The dynamic weighting coefficient is calculated as follows: ; The quality score for the i-th data source includes three dimensions: resolution, signal-to-noise ratio, and motion artifact level. This is an adjustment factor, with a value range of 0.5-1.

2.

6. The method for analyzing oral diseases based on AI intelligent recognition according to claim 5, characterized in that, The quality score The calculation formula is: ; in : Let be the standardized signal-to-noise ratio value of the i-th data source; This refers to spatial resolution metrics; Score the degree of motion artifacts; 、 、 These are the weighting coefficients for each dimension, with default values ​​of 0.4, 0.5, and 0.

1.

7. The method for analyzing oral diseases based on AI intelligent recognition according to claim 1, characterized in that, The output layer of the deep convolutional neural network employs an improved diagnostic scoring function: ; in: For the diagnostic score of the kth disease category; Sigmoid activation function ; The contribution of the m-th feature to the k-th disease category; The feature importance weights are obtained by optimizing them on multi-institutional data using a federated learning framework; This is a bias term.

8. An analysis system applied to an AI-based intelligent recognition method for analyzing oral diseases as described in any one of claims 1-7, characterized in that, include: Data acquisition module: integrates an oral scanning unit, an endoscope unit, and a CBCT interface. The oral scanning unit adopts blue light structured light projection technology with a sampling frequency of ≥15fps; the endoscope unit is equipped with a 4K CMOS sensor and a 470nm-940nm multispectral illumination system; the CBCT interface supports the DICOM3.0 standard protocol. The preprocessing module includes a point cloud filtering submodule, which uses a statistical outlier removal algorithm; and an image registration submodule, which implements spatial alignment of multimodal data. The feature standardization submodule performs Min-Max normalization and Z-score standardization; Dynamic fusion module: Configures a quality assessment engine, a weight calculation engine, and a feature fusion engine, supporting real-time weight adjustment and visualization of fusion results; AI Analysis Module: Deploys a multi-task deep learning network to process caries detection, periodontal disease assessment and mucosal lesion identification tasks in parallel. The network structure includes 121 convolutional layers and 8 attention modules. Report generation module: Generates interactive reports that include 3D reconstructed views, lesion heat maps, and treatment plan comparisons, supporting AR / VR display and electronic signature confirmation.

9. The analysis system according to claim 8, characterized in that, The endoscope unit adopts a dual-channel optical system. The visible light channel uses an RGB three-band beam splitter, and the near-infrared channel is equipped with a 900nm bandpass filter. The dual-channel images are fused at the feature level through a pulse-coupled neural network.

10. The analysis system according to claim 8, characterized in that, The system optimizes the hyperparameters of the neural network using a quantum evolution algorithm, specifically including: using qubit encoding to represent the number of network layers, convolution kernel size, and activation function type; and using a quantum rotation gate to update the parameters.

Citation Information

Cited By

  • Tooth health assessment method and system based on near-infrared transillumination

    CN121280436A

  • Dental health assessment method and system based on near-infrared transillumination

    CN121280436B

  • Intelligent oral cavity diagnosis and treatment method and equipment based on multi-Agent collaboration and storage medium

    CN122050789A

  • Intelligent oral diagnosis and treatment methods, devices, and storage media based on multi-agent collaboration

    CN122050789B

  • Dental extraction risk assessment method, system and equipment for mandibular third molar and medium

    CN122091224A