Weak fluorescence plaque detection method and device against toothpaste foam interference
Patent Information
- Application Number
- CN202610968621.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-08-18
AI Technical Summary
[0004]本发明的主要目的在于提供一种抗牙膏泡沫干扰的弱荧光牙菌斑检测方法及装置,旨在解决日常刷牙过程中,牙膏泡沫遮挡和环境光干扰导致普通摄像头无法稳定检测早期牙菌斑的微弱荧光信号,而传统显色剂方案操作繁琐、高光谱方案成本过高,均难以在消费级智能牙刷上实现实时、无感化的牙菌斑检测的技术问题
[0015]本发明通过激发光闪烁差分消除环境光,利用第一神经网络识别泡沫薄膜透光区域,引导第二神经网络从普通RGB图像中重建覆盖卟啉荧光特征峰的多光谱图像,再经第三神经网络融合空间纹理与光谱信息,建立薄膜区域与遮挡区域的空间关联以推理完整牙菌斑分布,最终在无需显色剂、无需真实高光谱相机的情况下,实现泡沫遮挡下早期弱荧光牙菌斑的实时检测与分割。
Smart Images

Figure CN122597808A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image detection technology, and in particular to a method and apparatus for detecting weak fluorescent dental plaque that resists interference from toothpaste foam. Background Technology
[0002] Dental plaque is a bacterial biofilm that adheres to the surface of teeth. Its metabolic products are the main pathogenic factors causing tooth decay, gingivitis, and periodontal disease. The maturity of dental plaque is positively correlated with the difficulty of its removal. Therefore, early detection and real-time removal of dental plaque during daily brushing has important clinical value for the prevention of oral diseases.
[0003] Currently, consumer-grade smart toothbrushes primarily rely on inertial sensors, pressure sensors, or cameras to monitor brushing areas, duration, and movements. These methods only indicate whether brushing was performed correctly, but cannot directly determine whether plaque actually exists or has been effectively removed. Clinically used plaque developers achieve visualization by binding chemical dyes to plaque. While offering good detection results, they require additional application, waiting, and rinsing, easily causing staining of teeth and oral mucosa. Furthermore, they necessitate continuous purchase of consumables, failing to meet the requirements of convenience, real-time performance, and long-term usability for consumer-grade products. Summary of the Invention
[0004] The main objective of this invention is to provide a method and device for detecting weak fluorescent dental plaque that resists interference from toothpaste foam. This invention aims to solve the technical problem that during daily brushing, toothpaste foam and ambient light interference prevent ordinary cameras from stably detecting the weak fluorescent signals of early dental plaque. Traditional chromogenic solutions are cumbersome to operate, and hyperspectral solutions are too expensive, making it difficult to achieve real-time, non-intrusive dental plaque detection on consumer-grade smart toothbrushes.
[0005] To achieve the above objectives, this invention proposes a method for detecting weakly fluorescent dental plaque that resists interference from toothpaste foam, comprising: Excitation light is used to irradiate the tooth surface, which stimulates porphyrin metabolites in dental plaque to produce red fluorescence; The long-pass filter suppresses the reflected light of the excitation light. The cutoff wavelength of the long-pass filter is located between the wavelength of the excitation light and the peak wavelength of dental plaque fluorescence, allowing only light with a wavelength greater than the cutoff wavelength to enter the camera. Eliminate ambient light entering the camera to obtain a weak fluorescence image after ambient light elimination; The weak fluorescence image is input into the first neural network to identify the foam film region in toothpaste foam that retains effective spectral information, and outputs a foam film region mask. Based on the foam film region mask, the image portion corresponding to the foam film region is extracted from the weak fluorescence image, and the image portion is input into the second neural network for hyperspectral reconstruction to obtain a multi-band multispectral reconstructed image. The band range of the multispectral reconstructed image at least covers the band where the porphyrin fluorescence characteristic peak of dental plaque is located. Weak fluorescence images and multispectral reconstructed images are input into a third neural network, which fuses spatial texture information and spectral information to output dental plaque segmentation results.
[0006] Furthermore, the excitation light is generated by a narrow-band 405nm LED with a full width at half maximum (FWHM) of less than 20nm; the flashing frequency of the excitation light is synchronously controlled with the exposure frame rate of the camera.
[0007] Furthermore, the first neural network processes the weak fluorescence image, distinguishes the foam film region from the foam edge region based on the different light transmittance characteristics of the toothpaste foam in the image, enhances the foam film region, and suppresses the foam edge region and background region to generate a foam film region mask.
[0008] Furthermore, during training, the first neural network uses a loss function that includes both a pixel-level classification loss function and a region overlap loss function.
[0009] Furthermore, the second neural network takes the image portion corresponding to the foam film region as input, restores the spectral response of teeth and dental plaque in multiple continuous bands through spectral reconstruction, and enhances the output response in the band where the porphyrin fluorescence characteristic peak of dental plaque is located, so as to obtain a multispectral reconstructed image.
[0010] Furthermore, the training of the second neural network employs a mask-constrained loss function, calculating the reconstruction error only within the foam film region. The mask-constrained loss function includes at least a spectral reconstruction loss term and a spectral morphology preservation loss term, with the spectral morphology preservation loss term constraining the curve variation trend of the band range where the dental plaque fluorescence characteristic peaks are located in the reconstructed spectrum.
[0011] Furthermore, the mask constraint loss function is a weighted combination loss function consisting of the spectral reconstruction loss term, the spectral contrast loss term, and the spectral slope loss term. The expression of the mask constraint combination loss function is: Total loss = 0.3 × mask-guided reconstruction loss + 0.5 × mask-guided spectral contrast loss + 0.2 × mask-guided spectral slope loss.
[0012] Furthermore, the third neural network extracts spatial texture and morphological features from the weak fluorescence image and extracts dental plaque fluorescence spectral features from the multispectral reconstructed image. It then fuses the two types of features, establishes the spatial relationship between the foam film region and the foam-covered region, infers the distribution of dental plaque in the foam-covered region, and outputs the dental plaque segmentation result.
[0013] Furthermore, the third neural network uses a dual-branch structure to extract spatial texture and morphological features and dental plaque fluorescence spectral features respectively. The deep features of the two branches are spliced together, and a self-attention mechanism is used to establish spatial association. Then, the spatial resolution is restored through the decoder to obtain the dental plaque segmentation result.
[0014] This invention also proposes a weakly fluorescent dental plaque detection device, comprising: An excitation light source is used to emit excitation light to illuminate the tooth surface, thereby exciting dental plaque to produce fluorescence; A filter is placed in the optical path of the excitation light source to suppress the reflected light of the excitation light and allow the fluorescence band of dental plaque to pass through; The image acquisition unit is used to acquire images of the on state and the off state respectively when the excitation light is alternately turned on and off; The processing unit is used to perform steps in the weak fluorescent dental plaque detection method that is resistant to toothpaste foam interference.
[0015] This invention eliminates ambient light by exciting light scintillation differentially, uses a first neural network to identify the light-transmitting area of the foam film, guides a second neural network to reconstruct a multispectral image covering the porphyrin fluorescence characteristic peak from a normal RGB image, and then uses a third neural network to fuse spatial texture and spectral information to establish a spatial correlation between the film area and the occluded area to infer the complete dental plaque distribution. Finally, it achieves real-time detection and segmentation of early weak fluorescent dental plaque under foam occlusion without the need for a colorimetric agent or a real hyperspectral camera. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic flowchart of the method for detecting weak fluorescent dental plaque that resists toothpaste foam interference according to the present invention; Figure 2 This is a flowchart illustrating the process of detecting toothpaste foam under a microscope in an embodiment of the method for detecting weak fluorescent dental plaque with resistance to toothpaste foam interference according to the present invention. Figure 3 This is a schematic diagram of the FilmMaskNet foam film region recognition network structure of the present invention; Figure 4This is a schematic diagram of the hyperspectral response characteristics under foam coverage conditions in an embodiment of the present invention; Figure 5 This is a schematic diagram of the MaskSRNet hyperspectral reconstruction network structure of the present invention; Figure 6 This is a schematic diagram of the PlaqueSegNet cross-modal plaque segmentation network structure of the present invention; Figure 7 This is a schematic diagram of the overall dental plaque detection method according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the weak fluorescent dental plaque detection device of the present invention.
[0019] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0020] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of the present invention and are not intended to limit the present invention.
[0021] To better understand the technical solution of the present invention, a detailed description will be provided below in conjunction with the accompanying drawings and specific embodiments.
[0022] like Figure 1 As shown, Figure 1 This is a schematic flowchart of the method for detecting weak fluorescent dental plaque that resists toothpaste foam interference according to the present invention.
[0023] This invention proposes a method and system for detecting weakly fluorescent dental plaque that resists interference from toothpaste foam. The method achieves real-time plaque detection during actual brushing by employing a cascaded deep learning process: "ambient light elimination—foam film region identification—region-guided hyperspectral reconstruction—cross-modal plaque inference and segmentation." The invention employs a three-stage cascaded network structure, including a foam film region identification network FilmMaskNet, a foam film region-guided hyperspectral reconstruction network MaskSRNet, and a cross-modal plaque segmentation network PlaqueSegNet.
[0024] S10, when excitation light is irradiated onto the tooth surface, excites porphyrin metabolites in dental plaque to produce red fluorescence.
[0025] In this embodiment, excitation light is used to irradiate the tooth surface, stimulating porphyrin metabolites in dental plaque to produce red fluorescence. In a preferred embodiment, the excitation light is generated by a narrow-band 405nm LED with a full width at half maximum (FWHM) of less than 20nm to improve the excitation efficiency of porphyrin metabolites in dental plaque. The excitation light source can be arranged in a ring array or symmetrically distributed around the RGB miniature camera to improve the excitation uniformity of the tooth surface and reduce local shadows and overexposure.
[0026] S20 uses a long-pass filter to suppress reflected excitation light. The cutoff wavelength of the long-pass filter is located between the wavelength of the excitation light and the peak wavelength of dental plaque fluorescence, allowing only light with wavelengths greater than the cutoff wavelength to enter the camera.
[0027] In this embodiment, a long-pass filter suppresses reflected excitation light. The cutoff wavelength of the long-pass filter is located between the wavelength of the excitation light and the peak wavelength of dental plaque fluorescence, allowing only light with wavelengths greater than the cutoff wavelength to enter the camera. In a preferred embodiment, a 490nm long-pass filter is placed in front of the camera to block 405nm excitation light and its specular reflection, allowing only wavelengths above 490nm to enter the camera, thereby preserving the autofluorescence of teeth and the weak fluorescence information of dental plaque.
[0028] S30, eliminates ambient light entering the camera to obtain a weak fluorescence image after ambient light elimination.
[0029] Specifically, real-time ambient light elimination is achieved using excitation light flashing and background subtraction: the excitation light is controlled to alternate between an on and off state, and images of the excitation light in both on and off states are captured by a camera; the on-state and off-state images are then differentially analyzed pixel by pixel to generate a weak fluorescence image after ambient light elimination. In a preferred embodiment, the LED operates using a high-frequency flashing mode, and the flashing frequency of the excitation light is synchronously controlled with the camera's exposure frame rate to achieve accurate frame synchronization differential.
[0030] By using the aforementioned optical acquisition structure and ambient light elimination processing, the effects of ambient light such as 405nm excitation light reflection interference, bathroom lights, sunlight, and vanity lights on the weak fluorescence detection of dental plaque can be effectively suppressed, resulting in a weak fluorescence image with a high signal-to-noise ratio.
[0031] To further reduce ambient light interference from bathroom lights, vanity lights, and natural light, this invention employs a combination of excitation light flashing and background subtraction for real-time ambient light elimination. Specifically, the camera captures images of the excitation light in both on and off states. The off-state image is used as the background image, and the excitation light is subtracted pixel-by-pixel from the on-state image to obtain a weak fluorescence image after ambient light elimination. This process can be represented as follows: ; in, This image represents the state of the excitation light being on. This image shows the image with the excitation light off. This represents a weak fluorescence image after ambient light removal. The above method effectively suppresses the influence of ambient light and background reflection on the weak fluorescence detection of dental plaque.
[0032] S40, input the weak fluorescence image into the first neural network, identify the foam film region in the toothpaste foam that retains effective spectral information, and output the foam film region mask; In this embodiment, a weak fluorescence image is input into a first neural network to identify the foam film region in toothpaste foam that retains effective spectral information, and outputs a foam film region mask.
[0033] like Figure 2 As shown, Figure 2 This is a flowchart illustrating the process of detecting toothpaste foam under a microscope in an embodiment of the method for detecting weak fluorescent dental plaque that resists toothpaste foam interference according to the present invention.
[0034] Toothpaste foam is not a uniform, transparent medium; it contains thin film areas and edge areas. Figure 2 This visually demonstrates the differences in the microstructure and optical properties of toothpaste foam. Figure 2 (a) is a macroscopic optical microscope image of toothpaste foam with a scale bar of 3 mm. It can be observed that the foam exhibits a typical porous honeycomb structure with uneven sizes. Bubbles ranging from tens of micrometers to several millimeters in diameter are stacked and distributed on each other, and the bubble walls are transparent or translucent liquid film morphology. Figure 2 Image (b) shows a magnified microscopic image of the foam, with a scale bar of 10 μm. It clearly marks two key structural regions of the foam: the Edge (foam edge region) and the Film (foam film region). The foam film region consists of an extremely thin liquid film between adjacent bubbles, exhibiting high optical transmittance. In contrast, the foam edge region is a liquid column (Platau boundary) formed by the convergence of three or more bubbles, significantly thicker than the film region, resulting in stronger light scattering and blocking effects.
[0035] Experiments showed that the foam film area had little effect on the autofluorescence of teeth and dental plaque, causing only slight and relatively uniform spectral attenuation, while the foam edge area caused stronger and more uneven spectral attenuation.
[0036] In this embodiment, after obtaining the RGB image after ambient light removal, the present invention first utilizes the FilmMaskNet foam film region recognition network to identify reliable regions in toothpaste foam. Research has found that toothpaste foam mainly consists of a foam film region and a foam edge region. The foam film region retains good light transmittance, while the foam edge region causes stronger and more uneven spectral attenuation. Therefore, the present invention does not directly "remove the foam," but first identifies the film region within the foam that still retains effective spectral information, and then performs hyperspectral recovery within that region.
[0037] like Figure 3 As shown, Figure 3 This is a schematic diagram of the FilmMaskNet foam film region recognition network structure of the present invention.
[0038] Figure 3 This paper showcases the encoder-decoder symmetric network architecture of FilmMaskNet, specifically designed for accurate identification of foam film regions. The network input is a 3×400×400 RGB image with foam. The encoding part consists of four convolutional blocks, each containing a 3×3 convolutional layer, a ReLU activation function, and a channel attention module. Simultaneously, max pooling with a stride of 2 progressively downsamples the feature map size from 400×400 to 50×50, and gradually increases the number of channels from 64 to 512, extracting deep semantic features of the foam. The decoding part correspondingly uses four upsampling blocks, progressively restoring the feature map size to 400×400 through bilinear interpolation, and gradually decreasing the number of channels from 512 to 1. The input of each upsampling block is fused with the output of the corresponding encoding block through skip connections to preserve the fine texture details of the foam. Finally, the network outputs a binarized foam film region mask with a size of 1×400×400 through a 1×1 convolutional layer and a sigmoid activation function, which is used for region selection in subsequent hyperspectral reconstruction.
[0039] In this embodiment, FilmMaskNet takes a 3×H×W RGB image with foam as input and outputs a 1×H×W foam film region mask. The network employs a U-Net encoder-decoder structure. The encoder extracts 64, 128, 256, and 512 channel features sequentially and progressively reduces the spatial resolution using max pooling. Each convolutional module includes... The system consists of convolutional layers, ReLU activation layers, and a spatial attention module. The spatial attention module generates a spatial response map through channel average pooling and max pooling, and then generates spatial attention weights through convolution and the sigmoid function, thereby enhancing the response of the foam film region and suppressing foam edges, high-reflectivity areas of teeth, and background regions.
[0040] The decoder recovers spatial resolution stepwise through bilinear upsampling and fuses shallow texture features with deep semantic features via skip connections. Finally, a foam film region mask is output through a 1×1 convolutional layer and a sigmoid function.
[0041] S50, based on the foam film region mask, extracts the image portion corresponding to the foam film region from the weak fluorescence image, and inputs the image portion into the second neural network for hyperspectral reconstruction to obtain a multi-band multispectral reconstructed image. The band range of the multispectral reconstructed image at least covers the band where the porphyrin fluorescence characteristic peak of dental plaque is located.
[0042] In this embodiment, based on the foam film region mask, the image portion corresponding to the foam film region is extracted from the weak fluorescence image, and the image portion is input into the second neural network for hyperspectral reconstruction to obtain a multi-band multispectral reconstructed image. The band range of the multispectral reconstructed image at least covers the band where the porphyrin fluorescence characteristic peak of dental plaque is located.
[0043] Hyperspectral experiments showed that healthy teeth have an autofluorescence peak around 500 nm, while early dental plaque shows a new weak fluorescence peak around 630 nm starting several hours after its formation.
[0044] like Figure 4 As shown, Figure 4 This is a schematic diagram of the hyperspectral response characteristics under foam coverage conditions in an embodiment of the present invention.
[0045] Figure 4 The experimental data verified the differences in spectral transmittance in different regions of the foam, providing a core basis for the technical feasibility of the present invention. Figure 4 (c) is a hyperspectral grayscale image of the tooth surface covered by foam. The image shows areas of uneven brightness on the tooth surface, corresponding to foam structures of different thicknesses. Two typical 2×2 pixel analysis areas are selected in the image: the brighter P1 area corresponds to the foam film area, and the lower-brightness P2 area corresponds to the foam edge area. Figure 4 Figure (d) shows the spectral response curves of different regions in the 500nm to 650nm wavelength range. The horizontal axis represents wavelength, and the vertical axis represents relative brightness. It includes four comparison curves: the solid line represents the spectrum of region P1 without foam coverage, the dotted line represents the spectrum of region P1 with foam film coverage, the short dashed line represents the spectrum of region P2 without foam coverage, and the dotted line represents the spectrum of region P2 at the foam edge. The curve comparison shows that the spectral curves of region P1 before and after foam coverage are highly consistent, with only a slight decrease in overall brightness, indicating that the foam film region still retains good spectral transmittance. In contrast, the spectral brightness of region P2 significantly decreases across the entire wavelength range after foam coverage, with the largest decrease occurring in the 500-550nm wavelength range, indicating stronger light scattering and spectral attenuation at the foam edge. These experimental results demonstrate that foam coverage does not completely block the weak fluorescence information of dental plaque; the foam film region still retains recoverable and effective spectral characteristics, which can serve as an important basis for hyperspectral reconstruction and dental plaque inference in this invention. The second neural network takes the image portion corresponding to the foam film region as input, and recovers the spectral response of teeth and dental plaque in multiple continuous bands through spectral reconstruction. It also enhances the output response in the band where the porphyrin fluorescence characteristic peak of dental plaque is located, so as to obtain a multispectral reconstructed image. Traditional hyperspectral cameras are bulky and expensive, making them difficult to integrate into toothbrushes. This invention recovers hyperspectral information from ordinary RGB images through a second neural network, eliminating the need to integrate a real hyperspectral camera into the toothbrush.
[0046] In this embodiment, the present invention utilizes the MaskSRNet hyperspectral reconstruction network, guided by a foam film region, to perform hyperspectral recovery of weak fluorescence information of teeth within the foam film region. Specifically, the foam film mask output by FilmMaskNet is multiplied pixel-by-pixel with the RGB image containing the foam film to obtain the RGB image of the foam film region. This process can be represented as follows: ; in, This indicates a mask for the foam film area. This indicates an RGB image with foam. This represents the RGB image of the foam film region.
[0047] Subsequently, the RGB image of the foam film region was input into MaskSRNet for hyperspectral reconstruction. The input to MaskSRNet is a data structure with a size of [size missing]. The RGB image is output as a size of The hyperspectral images correspond to 31 continuous spectral bands in the range of 500nm to 650nm.
[0048] like Figure 5 As shown, Figure 5 This is a schematic diagram of the MaskSRNet hyperspectral reconstruction network structure of the present invention.
[0049] Figure 5 This paper details the internal structure and data dimensionality variations of the MaskSRNet lightweight hyperspectral reconstruction network. The network input consists of two tensors: a 1×400×400 foam film mask and a 3×400×400 RGB image of the foam coverage. These are first multiplied pixel-wise to generate a 3×400×400 RGB mask image, retaining only information from the highly transparent foam film region, which serves as the input to the reconstruction network. The MaskSRNet core comprises five consecutive convolutional modules. Each module contains a 3×3 convolutional layer, a ReLU activation function, and a Squeeze-and-Excitation (SE) channel attention module (compression ratio 16) to adaptively enhance the response to weak fluorescence-related bands of dental plaque. The number of channels in each module is 64, 128, 256, 128, and 64, respectively. All convolutional layers use a stride of 1 and padding of 1 to ensure the feature map space size remains constant at 400×400. Finally, the network maps the number of channels to 31 through a 1×1 convolutional layer, and outputs a masked hyperspectral image with a size of 31×400×400, corresponding to 31 equally spaced bands in the wavelength range of 500nm to 650nm, accurately recovering the weak fluorescence characteristics of porphyrin near 630nm.
[0050] In this embodiment, the network employs a lightweight convolutional hyperspectral reconstruction structure with convolutional channels of 64, 128, 256, 128, 64, and 31, respectively. Each convolutional module includes... The system consists of convolutional layers, ReLU activation layers, and a channel attention module. The channel attention module employs a Squeeze-and-Excitation structure, generating channel description vectors through global average pooling, then compressing and restoring the channel dimensions through a fully connected layer, and finally generating channel weights through a sigmoid function to recalibrate the feature map channel by channel.
[0051] Through this mechanism, the network can enhance the characteristic channels associated with weak fluorescence in dental plaque, especially improving the recovery ability of weak fluorescence peaks of porphyrin in the 600nm to 650nm band and around 630nm.
[0052] MaskSRNet training employs a foam film region mask constraint loss function, calculating hyperspectral reconstruction error only within the foam film region to avoid interference from unreliable spectral information from the foam edge region. The loss function includes mask-guided reconstruction loss, mask-guided spectral contrast loss, and mask-guided spectral slope loss. Specifically, the mask-guided reconstruction loss constrains the pixel-level difference between the predicted and actual hyperspectral images; the spectral contrast loss enhances the spectral distinction between plaque and non-plaque regions; and the spectral slope loss constrains the spectral curve variation trend within the 610nm to 650nm band to maintain the morphological characteristics of the weak porphyrin fluorescence peak near 630nm.
[0053] The total loss function can be expressed as: ; in, To mask-guided reconstruction loss, To mask-guide spectral contrast loss, To guide the spectral slope loss for masking.
[0054] S60 inputs the weak fluorescence image and the multispectral reconstructed image into the third neural network, fuses spatial texture information and spectral information, and outputs the dental plaque segmentation result.
[0055] In this embodiment, weak fluorescence images and multispectral reconstructed images are input into a third neural network, spatial texture information and spectral information are fused, and dental plaque segmentation results are output.
[0056] The third neural network extracts spatial texture and morphological features from the weak fluorescence image and extracts plaque fluorescence spectral features from the multispectral reconstructed image. It then fuses these two types of features and establishes a spatial relationship between the foam film region and the foam-occluded region. This allows for the inference of plaque distribution in the foam-occluded region, outputting the plaque segmentation result. Since the foam film region retains reliable weak fluorescence cues, while the foam edge region causes occlusion, this invention utilizes a multi-head self-attention mechanism to establish a long-distance spatial relationship between the reliable cues in the film region and the edge-occluded region. This allows for the inference of plaque distribution obscured by foam based on the weak fluorescence information of locally visible plaque.
[0057] In this embodiment, after completing the hyperspectral reconstruction, the present invention further utilizes the cross-modal plaque segmentation network PlaqueSegNet to infer the complete plaque region under foam occlusion. The input to PlaqueSegNet includes the original RGB image with foam and the reconstructed hyperspectral image output by MaskSRNet. The RGB image is used to provide information on tooth structure, foam texture, and spatial morphology, while the hyperspectral image is used to provide weak fluorescence spectral information of plaque, especially the weak fluorescence features of porphyrins around 630 nm.
[0058] like Figure 6 As shown, Figure 6 This is a schematic diagram of the PlaqueSegNet cross-modal plaque segmentation network structure of the present invention.
[0059] Figure 6This paper demonstrates the complete structure of the PlaqueSegNet dual-branch cross-modal segmentation network, achieving accurate dental plaque detection under foam occlusion conditions. The network comprises two parallel encoder branches, each processing input data from a different modality. The hyperspectral branch takes a 31×400×400 masked hyperspectral image as input and extracts multi-scale weak fluorescence features through four convolutional blocks. Each convolutional block contains a 3×3 convolutional layer, a ReLU activation function, an SE channel attention module, and a max-pooling layer. The feature map size is progressively downsampled from 400×400 to 50×50, and the number of channels is increased from 64 to 512. The RGB branch takes a 3×400×400 foam-covered RGB image as input and employs the exact same encoder structure as the hyperspectral branch, extracting tooth morphology, foam texture, and spatial location features, outputting a 512×50×50 feature map. The feature maps output from the two branches are concatenated along the channel dimension to obtain a fused feature map of size 1024×50×50. This is then fed into a feature enhancement module containing a multi-head self-attention module (MHSA, 512 dimensions, 8 heads, 2 encoding layers) and a spatial attention module to model cross-regional contextual relationships and improve the segmentation accuracy of the plaque region. The decoding part consists of four upsampling blocks, which progressively restore the spatial resolution to 400×400 through bilinear interpolation. Each upsampling block is connected to the output of its corresponding encoder branch via skip connections to fuse shallow texture and deep semantic features. Finally, a 1×1 convolutional layer and a sigmoid activation function output a plaque segmentation result map of size 1×400×400.
[0060] In this embodiment, PlaqueSegNet employs a dual-branch encoder structure, including an RGB branch and a hyperspectral branch. Both branches utilize a multi-level convolutional coding structure to extract 64, 128, 256, and 512 channel features sequentially. Specifically, the hyperspectral branch incorporates a channel attention module to enhance the response in weak fluorescence-related bands; the RGB branch incorporates a spatial attention module to enhance the spatial morphological features of foam structures, tooth boundaries, and dental plaque.
[0061] Subsequently, the deep features of the two branches are concatenated along the channel dimension to form a joint feature map, which is then input into a multi-head self-attention module for cross-regional relationship modeling. In one implementation, the self-attention module adopts an MHSA structure with an embedding dimension of 512, 8 attention heads, and 2 encoding layers. This module is used to establish long-distance spatial associations between reliable weak fluorescence cues in the foam film region and the foam edge-occluded region, thereby inferring the distribution of dental plaque occluded by the foam edge based on the weak fluorescence information of locally visible dental plaque.
[0062] Subsequently, PlaqueSegNet restores the spatial resolution step by step through a progressive upsampling decoder and combines shallow texture features with deep semantic features using skip connections. Finally, a dental plaque segmentation map with a size of 1×H×W is output through a 1×1 convolutional layer and a sigmoid function.
[0063] like Figure 7 As shown, Figure 7 This is a schematic diagram of the overall dental plaque detection method according to an embodiment of the present invention.
[0064] Figure 7 This invention clearly demonstrates its three-stage end-to-end technical process based on "foam film region identification—mask-guided hyperspectral reconstruction—cross-modal plaque segmentation," with data flow proceeding sequentially from left to right. Stage 1: Foam film region identification. The input is a real-time RGB image of foam captured during brushing. Semantic segmentation is performed using a FilmMaskNet neural network, ultimately outputting a binarized foam film region mask. White areas represent foam film regions with high light transmittance and retaining effective weak fluorescence information, while black areas represent severely obscured foam edges and invalid background regions. Stage 2: Mask-guided hyperspectral reconstruction. The original RGB image with foam is multiplied pixel-by-pixel with the foam film region mask output from Stage 1, resulting in a mask RGB image that retains only foam film region information. This mask RGB image is then input into the MaskSRNet hyperspectral reconstruction network, outputting a reconstructed 31-band hyperspectral image covering the 500-650nm weak fluorescence characteristic band of dental plaque. Stage 3: Dental plaque segmentation. Simultaneously, the reconstructed 31-band hyperspectral image and the original RGB image with foam are input. Multimodal feature fusion and inference are performed through the PlaqueSegNet cross-modal segmentation network, and the final output is a dental plaque segmentation result map, where the white area is the detected dental plaque residue area and the black area is the healthy tooth and background area.
[0065] In this embodiment, the present invention outputs a plaque segmentation map, plaque coverage area, areas missed during brushing, and brushing suggestions. The detection results can be displayed on a smart mirror, mobile app, smart toothbrush display interface, or medical terminal to guide users to adjust their brushing position in real time and improve plaque removal effectiveness.
[0066] Experimental results show that the system of the present invention can achieve early detection of dental plaque during brushing, which is affected by toothpaste foam and ambient light, without the need for dental plaque developer or integrated real hyperspectral camera.
[0067] like Figure 8 As shown, Figure 8 This is a flowchart illustrating the control method for image detection according to Embodiment 1 of the present invention.
[0068] This invention also proposes a weakly fluorescent dental plaque detection device, comprising: An excitation light source 10 is used to emit excitation light to irradiate the tooth surface and excite dental plaque to produce fluorescence. In a preferred embodiment, the excitation light source is a narrow-band 405nm LED with a full width at half maximum (FWHM) of less than 20nm to improve the excitation efficiency for porphyrin metabolites in dental plaque. The excitation light source is arranged in a ring-shaped symmetrical layout around the image acquisition unit to improve the uniformity of excitation on the tooth surface.
[0069] Filter 20 is disposed in the optical path of the excitation light source to suppress reflected light from the excitation light and allow the fluorescence band of dental plaque to pass through. In a preferred embodiment, a 490nm long-pass filter is used to block 405nm excitation light and its specular reflection, allowing only wavelengths above 490nm to enter the image acquisition unit.
[0070] The image acquisition unit 30 is used to acquire images in the on state and the off state respectively when the excitation light is alternately turned on and off. In a preferred embodiment, the image acquisition unit is an RGB miniature camera, and its exposure frame rate is synchronously controlled with the flashing frequency of the excitation light.
[0071] The processing unit 40 is used to perform the steps in the aforementioned weak fluorescent dental plaque detection method that resists toothpaste foam interference. The processing unit can be deployed in a mobile terminal, smart mirror, or edge AI device to achieve real-time dental plaque detection during brushing. In a preferred embodiment, the overall system inference latency is controlled within 100ms, and the detection results can be displayed in real-time on the smart mirror, mobile app, or toothbrush display interface.
[0072] The aforementioned excitation light source, filter, and image acquisition unit are integrated into the head of the oral care device, while the processing unit is located in the main body of the oral care device or in an external smart terminal that communicates with the oral care device. The oral care device is either a smart toothbrush or an oral endoscope.
[0073] To further improve the stability of weak fluorescence detection of dental plaque, robustness against foam interference, and real-time detection performance, the present invention further proposes the following preferred technical solutions.
[0074] Regarding the optical structure for weak fluorescence acquisition, preferably, the 405nm excitation light source is arranged in a ring array or symmetrically distributed around the RGB miniature camera to improve the excitation uniformity of the tooth surface and reduce local shadows and overexposure. The excitation light source is preferably a narrow-band 405nm LED with a full width at half maximum (FWHM) of less than 20nm to improve the excitation efficiency for porphyrin metabolites in dental plaque. A 490nm long-pass filter is placed in front of the RGB camera to block the 405nm excitation light and its specular reflection, allowing only wavelengths above 490nm to enter the camera, thereby improving the signal-to-noise ratio of weak fluorescence in dental plaque. To improve ambient light suppression, the excitation light and camera exposure are synchronously controlled. The LED preferably operates in a high-frequency flicker mode. The camera separately acquires LED-on frames and LED-off frames, and background subtraction is used to eliminate the influence of ambient light. This method can effectively reduce the interference of bathroom lights, sunlight, and vanity lights on weak fluorescence detection.
[0075] For foam film region recognition networks, FilmMaskNet preferably employs a U-Net encoder-decoder structure and introduces a spatial attention module in the encoder convolutional module to improve foam film region recognition capabilities. The encoder preferably uses a four-level convolutional structure, with each level including a 3×3 convolutional layer, a ReLU activation layer, a spatial attention module, and a max-pooling layer. The spatial attention module generates a spatial attention map through channel average pooling and channel max pooling to enhance the response of the foam film region and suppress foam edges, high-reflectivity areas, and background areas. The decoder preferably uses bilinear interpolation for upsampling and fuses shallow texture features with deep semantic features through skip connections to improve the segmentation accuracy of foam edge regions. During training, FilmMaskNet preferably uses a combination of BCE loss and Dice loss to simultaneously ensure pixel-level classification stability and region overlap accuracy.
[0076] For the hyperspectral reconstruction network, MaskSRNet is preferably used to recover the hyperspectral information of teeth and dental plaque in the foam film region, outputting 31 continuous spectral bands corresponding to hyperspectral images in the range of 500nm to 650nm. The area around 500nm corresponds to the autofluorescence peak of teeth, and the area around 630nm corresponds to the weak porphyrin fluorescence peak of dental plaque. The network preferably adopts a lightweight convolutional hyperspectral reconstruction structure. Each convolutional module includes a 3×3 convolutional layer, a ReLU activation layer, and a channel attention module. The channel attention module preferably adopts a Squeeze-and-Excitation structure, assigning different weights to different channels to make the network pay more attention to the weak fluorescence-related features in the 600nm to 650nm band, thereby improving the recovery ability of the weak porphyrin fluorescence peak around 630nm. During training, a foam film region mask constraint is preferably used, calculating the hyperspectral reconstruction error only in the high-confidence film region to avoid unreliable spectral information from the foam edge region affecting model training. Furthermore, MaskSRNet preferably employs a combination of loss functions, including mask-guided reconstruction loss, mask-guided spectral contrast loss, and mask-guided spectral slope loss. The spectral slope loss focuses on constraining the spectral variation trend in the 610 nm to 650 nm band to maintain the morphological characteristics of the weak fluorescence peak of porphyrin near 630 nm.
[0077] For cross-modal plaque segmentation networks, PlaqueSegNet preferably employs a dual-branch cross-modal structure, including an RGB branch and a hyperspectral branch. The RGB branch is used to extract foam texture, tooth structure, and spatial morphological features, while the hyperspectral branch is used to extract weak fluorescence spectral features of plaque. Both branches use a four-level convolutional coding structure. The hyperspectral branch incorporates a channel attention module to enhance the response of weak fluorescence-related bands, while the RGB branch incorporates a spatial attention module to enhance the spatial morphological features of foam edges, tooth gaps, and plaque. The deep features of the two branches are concatenated along the channel dimension and then input into a multi-head self-attention module for cross-regional relationship modeling. This module can establish long-distance spatial associations between reliable weak fluorescence cues in the foam film region and the foam edge occlusion region, thereby enabling plaque inference in the foam occlusion region. The decoder preferably employs a progressive bilinear upsampling structure and fuses shallow texture information with deep semantic information through skip connections to improve plaque boundary recovery capabilities.
[0078] In terms of hyperspectral dataset acquisition and registration, preferably, a hyperspectral camera is used to acquire hyperspectral images of teeth without foam, and corresponding RGB images are acquired simultaneously to establish an RGB-hyperspectral correspondence dataset. To obtain data under realistic brushing foam scenarios, it is preferable to use a transparent ultra-thin dental crown to cover the teeth, allowing toothpaste foam to form on the crown surface while plaque remains on the tooth surface, thus avoiding direct removal of plaque during brushing. Spatial registration between RGB images with foam, hyperspectral images without foam, and plaque-annotated images is preferably performed using a homography matrix. Specifically, at least four corresponding feature points can be manually selected in different modal images, and the homography matrix can be calculated based on the corresponding feature points to achieve pixel-level spatial alignment. After registration, the hyperspectral images are preferably downsampled using nearest neighbor interpolation to make them consistent with the resolution of the RGB images, in order to establish a pixel-level RGB-hyperspectral correspondence.
[0079] Regarding real-time deployment and mobile implementation, preferably, the three-stage network is deployed in mobile terminals, smart mirrors, or edge AI devices to achieve real-time plaque detection during brushing. The overall system inference latency is preferably controlled within 100ms to enable real-time interactive feedback. The detection results preferably include plaque coverage area, brushing missed areas, brushing health score, and brushing suggestions. These results can be displayed in real-time on smart mirrors, mobile apps, or the toothbrush's display interface to guide users in continuing to brush areas with plaque residue.
[0080] To broaden the scope of protection of this invention, in addition to the embodiments described above, this invention also includes equivalent substitutions, combinations, or modifications to the optical structure, network structure, loss function, inference method, and deployment method. Modifications made by those skilled in the art without departing from the core concept of this invention are all within the scope of protection of this invention.
[0081] Regarding the excitation light source, this invention is not limited to using 405nm excitation light. In other embodiments, wavelengths capable of exciting weak fluorescence in dental plaque, such as 375nm, 385nm, 395nm, 410nm, and 420nm, can also be used. The type of excitation light source is not limited to LED light sources; lasers, VCSEL light sources, pulsed light sources, and multi-wavelength combination light sources can also be used. The arrangement of the excitation light sources is not limited to a ring layout; linear arrays, symmetrical arrays, distributed arrays, or embedded brush head structures can also be used.
[0082] Regarding the filter structure, the present invention is not limited to using a 490nm long-pass filter. In other embodiments, a bandpass filter, a narrow-band filter, a multilayer dielectric filter, a polarization filter structure, or a combination of multiple filters may also be used. The filter band range can also be adjusted according to different excitation wavelengths and fluorescence characteristic bands.
[0083] Regarding the method of eliminating ambient light, the present invention is not limited to background subtraction between LED on and off frames. In other embodiments, high-frequency modulation and demodulation, phase-locked detection, multi-frame background modeling, HDR fusion, polarization difference, frequency domain filtering, learning-based background estimation network or deep learning illumination compensation can also be used to further reduce ambient light interference.
[0084] In the foam region identification part, this invention is not limited to using the FilmMaskNet network based on the U-Net structure. In other embodiments, DeepLab, SegNet, HRNet, PSPNet, SwinTransformer, VisionTransformer, or other semantic segmentation network structures can also be used. The classification method of foam regions is not limited to the binary classification of "foam film region" and "non-film region", but can be further divided into multiple categories such as foam edge region, high-density foam region, bubble region, exposed tooth region, and background region.
[0085] In the hyperspectral reconstruction section, this invention is not limited to using a convolutional hyperspectral reconstruction network. In other embodiments, ResNet, DenseNet, Transformer, diffusion models, GAN generator networks, MLP-Mixer, multi-scale fusion networks, or implicit neural representation networks can also be used for RGB-to-hyperspectral mapping and recovery. The number of output bands is not limited to 31 bands and can be set to 16 bands, 24 bands, 64 bands, or other continuous spectral sampling methods according to actual needs. The hyperspectral range is not limited to 500nm to 650nm and can be extended to the ultraviolet band, near-infrared band, or a wider range of multispectral detection.
[0086] Regarding the attention mechanism, this invention is not limited to spatial attention and channel attention structures. In other embodiments, CBAM attention, non-local attention, multi-head self-attention, Transformer attention, frequency domain attention, or dynamic convolutional attention structures may also be used. The location of the attention module is not limited to the encoder part, but may also be set in the decoder, skip connections, feature fusion layer, or output layer.
[0087] Regarding the loss function, this invention is not limited to using reconstruction loss, spectral contrast loss, and spectral slope loss. In other embodiments, L1 loss, SSIM loss, perceptual loss, adversarial loss, Dice loss, focal loss, KL divergence loss, or frequency domain constraint loss may also be used. The weight relationship between different loss functions can also be adjusted using dynamic weights, adaptive weights, or a learning-based approach.
[0088] In the plaque segmentation part, this invention is not limited to using a two-branch cross-modal structure. In other embodiments, a single-branch fusion structure, a multi-branch fusion structure, a Transformer segmentation structure, a graph neural network structure, a temporal video segmentation network, or a diffusion model inference structure may also be used. The output results are not limited to pixel-level plaque segmentation maps, but may also include plaque heatmaps, plaque coverage area, plaque severity, brushing quality scores, brushing suggestions, and dynamic change trends.
[0089] Regarding data acquisition and spatial registration, this invention is not limited to using transparent dental braces to assist in constructing foam data. In other embodiments, transparent protective films, transparent resin layers, simulated tooth models, or artificial foam generation structures can also be used. Image registration methods are not limited to homography matrix transformation; affine transformation, optical flow registration, deep learning-based image registration, or 3D point cloud registration methods can also be used.
[0090] Regarding system deployment and application scenarios, this invention is not limited to smart toothbrushes, but can also be applied to oral endoscopes, smart mirrors, mobile phone peripherals, medical oral examination equipment, children's brushing guidance equipment, remote oral health monitoring systems, and AI-assisted dental diagnostic systems. System deployment methods are not limited to local inference, but can also employ edge AI deployment, cloud inference, mobile deployment, FPGA deployment, NPU accelerated deployment, and other implementation methods.
[0091] Furthermore, the weak fluorescence recovery and foam interference suppression method proposed in this invention can be extended to other weak fluorescence detection scenarios, including biofilm detection, bacterial fluorescence detection, medical fluorescence endoscopy, tissue weak fluorescence analysis, and skin fluorescence detection.
Claims
1. A method for detecting weakly fluorescent dental plaque that resists interference from toothpaste foam, characterized in that, include: Excitation light is used to irradiate the tooth surface, which stimulates porphyrin metabolites in dental plaque to produce red fluorescence; The reflected light of the excitation light is suppressed by a long-pass filter. The cutoff wavelength of the long-pass filter is located between the wavelength of the excitation light and the peak wavelength of the dental plaque fluorescence, allowing only light with a wavelength greater than the cutoff wavelength to enter the camera. Eliminate ambient light entering the camera to obtain a weak fluorescence image after ambient light elimination; The weak fluorescence image is input into the first neural network to identify the foam film region in the toothpaste foam that retains effective spectral information, and outputs a foam film region mask. Based on the foam film region mask, the image portion corresponding to the foam film region is extracted from the weak fluorescence image, and the image portion is input into the second neural network for hyperspectral reconstruction to obtain a multi-band multispectral reconstructed image. The band range of the multispectral reconstructed image at least covers the band where the porphyrin fluorescence characteristic peak of dental plaque is located. The weak fluorescence image and the multispectral reconstructed image are input into a third neural network, which fuses spatial texture information and spectral information to output dental plaque segmentation results.
2. The method for detecting weak fluorescent dental plaque with resistance to toothpaste foam interference according to claim 1, characterized in that, The excitation light is generated by a narrowband 405nm LED with a half-width of less than 20nm; the flashing frequency of the excitation light is synchronously controlled with the exposure frame rate of the camera.
3. The method for detecting weak fluorescent dental plaque with resistance to toothpaste foam interference according to claim 1, characterized in that, The first neural network processes the weak fluorescence image, distinguishes the foam film region from the foam edge region based on the different light transmittance characteristics of the toothpaste foam in the image, enhances the foam film region, and suppresses the foam edge region and the background region to generate the foam film region mask.
4. The method for detecting weak fluorescent dental plaque with resistance to toothpaste foam interference according to claim 3, characterized in that, During training, the first neural network uses a loss function that includes both a pixel-level classification loss function and a region overlap loss function.
5. The method for detecting weakly fluorescent dental plaque with resistance to toothpaste foam interference according to claim 1, characterized in that, The second neural network takes the image portion corresponding to the foam film region as input, restores the spectral response of teeth and dental plaque in multiple continuous bands through spectral reconstruction, and enhances the output response in the band where the porphyrin fluorescence characteristic peak of dental plaque is located, so as to obtain the multispectral reconstructed image.
6. The method for detecting weakly fluorescent dental plaque with resistance to toothpaste foam interference according to claim 5, characterized in that, The training of the second neural network uses a mask-constrained loss function, which calculates the reconstruction error only within the foam film region. The mask-constrained loss function includes at least a spectral reconstruction loss term and a spectral morphology preservation loss term. The spectral morphology preservation loss term is used to constrain the curve variation trend of the band range where the dental plaque fluorescence characteristic peak is located in the reconstructed spectrum.
7. The method for detecting weak fluorescent dental plaque with resistance to toothpaste foam interference according to claim 6, characterized in that, The mask constraint loss function is a weighted combination loss function consisting of a spectral reconstruction loss term, a spectral contrast loss term, and a spectral slope loss term. The expression of the mask constraint combination loss function is: Total loss = 0.3 × mask-guided reconstruction loss + 0.5 × mask-guided spectral contrast loss + 0.2 × mask-guided spectral slope loss.
8. The method for detecting weak fluorescent dental plaque with resistance to toothpaste foam interference according to claim 1, characterized in that, The third neural network extracts spatial texture and morphological features from the weak fluorescence image and extracts dental plaque fluorescence spectral features from the multispectral reconstructed image. It then fuses the two types of features, establishes a spatial correlation between the foam film region and the foam-covered region, infers the distribution of dental plaque in the foam-covered region, and outputs the dental plaque segmentation result.
9. The method for detecting weak fluorescent dental plaque with resistance to toothpaste foam interference according to claim 8, characterized in that, The third neural network uses a dual-branch structure to extract the spatial texture and morphological features and the dental plaque fluorescence spectral features, respectively. The deep features of the two branches are spliced together, and the spatial association is established using a self-attention mechanism. Then, the spatial resolution is restored through a decoder to obtain the dental plaque segmentation result.
10. A weakly fluorescent dental plaque detection device, characterized in that, include: An excitation light source is used to emit excitation light to illuminate the tooth surface, thereby exciting dental plaque to produce fluorescence; A filter is disposed in the optical path of the excitation light source to suppress the reflected light of the excitation light and allow the fluorescence band of the dental plaque to pass through; The image acquisition unit is used to acquire images of the on state and the off state respectively when the excitation light is alternately turned on and off; The processing unit is configured to perform the steps in the weak fluorescent dental plaque detection method with resistance to toothpaste foam interference as described in any one of claims 1 to 9.