Image processing method, device, medium, computer program product, and tooth recognition system
Through multispectral imaging technology and artificial intelligence algorithms, the problem of insufficient accuracy in early caries detection is solved, accurate positioning and classification of teeth are achieved, and an efficient and radiation-free caries detection method is provided.
Patent Information
- Application Number
- CN202510969064.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-14
AI Technical Summary
Existing caries detection technologies lack accuracy in early caries detection, have a single information dimension, and lack automated multispectral tooth image segmentation, positioning, and classification methods.
Multispectral imaging technology is combined with artificial intelligence algorithms to acquire visible light, near-infrared and fluorescence images. The segmentation network is used to embed a spatial attention module to segment the tooth area, extract multimodal features and classify them through a multi-branch convolutional neural network. Combined with FDI tooth position encoding, the precise positioning and classification of teeth can be achieved.
It achieves accurate monitoring and classification of early-stage caries, improves detection accuracy and efficiency, and provides a convenient detection method without ionizing radiation.
Smart Images

Figure CN120495283B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image technology, and in particular to an image processing method, device, medium, computer program product, and tooth recognition system. Background Art
[0002] Dental caries is a highly prevalent oral disease worldwide. Early detection and intervention are key to its prevention and treatment. In the early stages of dental caries, tooth enamel demineralization occurs. This process is not easily noticeable, but if identified promptly, it can be effectively reversed through non-invasive remineralization treatments, preventing the development of serious complications such as pulpitis, which require invasive treatment.
[0003] Existing caries detection technologies each have their limitations. Traditional X-ray examinations suffer from issues such as insufficient sensitivity and radiation risks, making them difficult to use for routine dynamic observation of early-stage caries. Recent advances in optical detection technologies, such as optical coherence tomography (OCT) and laser-induced fluorescence (LF), have improved the detection rate of early-stage caries. However, these devices are often expensive, bulky, or complex to operate, making them difficult to adopt for primary care and personal health monitoring.
[0004] Imaging using light of specific wavelengths is a promising approach for nondestructive testing. For example, near-infrared (NIR) light can penetrate tooth enamel, distinguishing healthy from demineralized areas based on changes in its scattering properties. Ultraviolet (UV) light of specific wavelengths can stimulate the autofluorescence of microorganisms and their metabolites on the tooth surface, thereby reflecting oral hygiene. However, a single optical imaging modality can only provide limited information.
[0005] Furthermore, with the development of artificial intelligence (AI) technology, the use of algorithms to automatically analyze medical images has become a trend. However, in the field of dentistry, effectively integrating information from multiple optical imaging modalities and building an intelligent process that automatically performs tooth location, numbering, feature extraction, and status classification remains a pressing technical challenge. Summary of the Invention
[0006] The purpose of the present invention is to provide an image processing method, device, medium, computer program product, and tooth recognition system to solve the technical problems of insufficient accuracy in early caries detection, single information dimension, and lack of intelligent analysis methods that can automatically and accurately segment, locate and classify multispectral tooth images.
[0007] A first aspect of the present invention discloses an image processing method for an electronic device, comprising:
[0008] Obtain visible light images, near-infrared images, and fluorescence images of the same tooth area to obtain multimodal images;
[0009] Inputting the multimodal image into a segmentation network, wherein a spatial attention module is embedded in the jump connection between the encoder and decoder of the segmentation network, and the segmentation network outputs a binary mask of the tooth region;
[0010] Extracting the outline of a single tooth from the binary mask, and determining the minimum circumscribed rectangle and geometric center of gravity of each of the outlines;
[0011] Assigning an FDI tooth position code to each of the contours according to the horizontal coordinate of the geometric center of gravity;
[0012] According to the minimum circumscribed rectangle, cropping a corresponding single tooth image block from the multimodal image;
[0013] The single tooth image block is input into a multi-branch convolutional neural network, where:
[0014] In the channel corresponding to the visible light image, the local binary pattern and gray-level co-occurrence matrix are extracted as the first set of features.
[0015] In the channel corresponding to the near-infrared image, the scattering intensity histogram of the pixel values is calculated as the second set of features.
[0016] In the channel corresponding to the fluorescence image, the fluorescence intensity information of the pixel value is extracted as the third set of features.
[0017] and performing channel splicing on the first set of features, the second set of features, and the third set of features to obtain a unified feature vector;
[0018] The unified feature vector is input into a fully connected classification network, and the fully connected classification network outputs probability values corresponding to different categories.
[0019] A second aspect of the present invention discloses a tooth recognition system, comprising:
[0020] A light source module, including white light LEDs, ultraviolet LEDs, and near-infrared LEDs distributed around the periphery;
[0021] Imaging module, including lens, CMOS image sensor and optical filter;
[0022] The control module is configured to control the light source module to switch spectral channels, so that in different spectral channels, the light source module lights up one of the white light LED, the ultraviolet LED and the near-infrared LED respectively; and control the CMOS image sensor to collect images of teeth,
[0023] The processing module is configured to recognize the image of the teeth according to the image processing method of the first aspect of the present invention.
[0024] A third aspect of the present invention discloses an electronic device comprising a memory storing computer-executable instructions and a processor. When the instructions are executed by the processor, the electronic device implements the image processing method according to the first aspect of the present invention.
[0025] A fourth aspect of the present invention discloses a computer storage medium having instructions stored thereon. When the instructions are executed on a computer, the computer is caused to execute the image processing method according to the first aspect of the present invention.
[0026] A fifth aspect of the present invention discloses a computer program product comprising computer executable instructions, which are executed by a processor to implement the image processing method according to the first aspect of the present invention.
[0027] Compared with the prior art, the main differences and effects of the embodiments of the present invention are:
[0028] The image processing method disclosed in the first aspect of the present invention can more accurately segment the tooth area from the multimodal image by embedding the spatial attention module in the jump connection of the segmentation network, output a high-quality binary mask, and provide a reliable basis for the subsequent single tooth analysis. Secondly, the method extracts and fuses the texture features of the visible light image, the scattering features of the near-infrared image, and the intensity information of the fluorescence image through a multi-branch network, and constructs a unified feature vector with complementary information. Compared with the method that relies on a single spectral information, it can more comprehensively and deeply characterize the health status of the teeth. Finally, by automatically extracting the tooth contour, determining its geometric center of gravity and minimum circumscribed rectangle, and allocating FDI tooth position codes and cropping image blocks accordingly, the accurate identification and positioning of a single tooth is achieved. Combined with the subsequent classification network, the full process of automated processing from the original image to the probability value of the specific tooth caries grade is completed, which significantly improves the accuracy, objectivity and efficiency of caries identification.
[0029] The tooth recognition system disclosed in the second aspect of the present invention integrates three LED light sources—white light, ultraviolet light, and near-infrared light—into a surround-distributed light source module. The module, controlled by a control module for time-sharing illumination, is combined with specific optical filters to efficiently capture three key diagnostic images reflecting the macroscopic morphology, surface microbial status, and internal structural changes of teeth on a single device, achieving comprehensive data collection. Because the system's processing module is configured to execute the image processing method of the first aspect of the present invention, the system is capable of automated, intelligent analysis and classification of the captured images, thereby providing a convenient, ionizing radiation-free, and reliable technical means for screening and dynamic monitoring of early-stage caries. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1A schematic diagram of the hardware structure of a tooth recognition system according to an embodiment of the present application is shown.
[0031] Figure 2 A three-channel tooth image according to an embodiment of the present application is shown.
[0032] Figure 3 A schematic structural diagram of a circuit board of a tooth recognition system according to an embodiment of the present application is shown.
[0033] Figure 4 FIG. 4 shows a visual annotation of a channel tooth image according to an embodiment of the present application.
[0034] Figure 5 A flowchart of an image processing method according to an embodiment of the present application is shown.
[0035] Figure 6 A multi-branch convolutional neural network according to an embodiment of the present application is shown.
[0036] Figure 7 A hardware structure block diagram of an electronic device according to an image processing method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0037] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0038] Example 1: Tooth Recognition System
[0039] An embodiment of the present invention provides a tooth recognition system that uses multispectral imaging technology combined with artificial intelligence algorithms to achieve accurate monitoring and classification of tooth conditions, especially early-stage caries.
[0040] See also Figure 1 This figure shows the hardware architecture of a tooth recognition system according to an embodiment of the present invention. The system can be designed as a portable intraoral camera, comprising a probe for capturing images and a handle housing the primary circuitry. A light source module 1 and an imaging module 2 are located at the tip of the probe. These modules are connected to a circuit board 4 located in the handle via control lines 3.1 and signal lines 3.2.
[0041] The light source module 1 is configured to provide multi-spectrum illumination. In a specific embodiment, the light source module 1 includes a white light LED 1.1, a near-infrared light LED 1.2, and an ultraviolet light LED 1.3. In order to achieve uniform illumination in a small oral environment and avoid shadows and reflections caused by a single light source position, preferably, multiple LEDs of each spectrum are provided, for example, three, and are arranged around the lens of the imaging module 2. For example, Figure 1The three white light LEDs 1.1 are spaced 120 degrees apart, the three near-infrared LEDs 1.2, and the three ultraviolet LEDs 1.3 are spaced 120 degrees apart around the lens of the imaging module 2. This layout ensures that the target tooth surface can be evenly and fully illuminated.
[0042] In this embodiment, the wavelength of the light source selected has a specific technical significance. The light emitted by the white light LED 1.1 is used to obtain a visible light image of the teeth, which can intuitively display information such as the shape, color, physical defects, and surface pigmentation of the teeth. Figure 2 The UV LED 1.3 preferably emits UV light with a wavelength of 405 nm. This wavelength of light can effectively stimulate microorganisms (such as dental plaque) and their metabolites (such as porphyrin) attached to the tooth surface to produce autofluorescence. By capturing this fluorescence signal, it is possible to generate Figure 2 The image shown in the "autofluorescence image" in the figure is used to evaluate the oral hygiene status and the distribution of harmful microorganisms, which is crucial for determining the risk of dental caries. The near-infrared LED 1.2 preferably emits near-infrared light with a wavelength of 780nm. Based on the principle of light scattering, the 780nm near-infrared LED can effectively identify early tooth demineralization, cracks and other early dental problems that are difficult to detect under white light. At the same time, this wavelength can also effectively distinguish light pigmentation on the tooth surface, which may be judged as a caries state under white light. Figure 2 As shown in the "Infrared Image" in the figure.
[0043] Imaging module 2 is used to capture optical signals reflected or emitted from the teeth after illumination by light source module 1. This module includes a lens, a CMOS image sensor, and an optical filter. The lens features a small field of view and small aperture design, matching the pixel size of the CMOS image sensor to improve imaging resolution and avoid information redundancy. The CMOS image sensor is responsible for capturing the signal and converting the optical signal into an electrical signal to enable multi-wavelength spectrum acquisition. The optical filter plays a key role in this system, being configured to block stray light that interferes with imaging. Specifically, the filter can effectively block ultraviolet light with wavelengths below 405nm, while filtering out infrared light with wavelengths greater than 795nm. This filter thus forms a bandpass window that allows light with wavelengths between 420nm and 790nm (including visible light, microbial autofluorescence, and near-infrared reflected light at 780nm) to pass through and be captured by the CMOS image sensor.
[0044] Please see the attached Figure 3The figure shows the structure of circuit board 4. Circuit board 4 integrates the system's core storage module, control module, and processing module. It includes a serial flash memory chip 4.1 (storage module) for temporarily storing processed image data; a capture button 4.2 and a mode switch button 4.4 (control module); an image processing chip 4.3 (processing module) for receiving and pre-processing raw image data from the CMOS image sensor; and a Wi-Fi module 4.5 for wirelessly transmitting the stored image data to an external smart terminal device, such as a mobile phone, tablet, or computer. Optionally, Wi-Fi module 4.5 can be replaced by other wireless or wired transmission modules.
[0045] The physical implementation of the control module on circuit board 4 includes the aforementioned capture button 4.2 and mode switch button 4.4, with control functions performed by firmware logic within a chip (e.g., image processing chip 4.3). This control module allows users to flexibly capture images. The system supports automatic and manual modes, which can be switched between by a first operation (e.g., a long press) on mode switch button 4.4. In automatic mode, a single operation (e.g., a single click) on capture button 4.2 automatically illuminates the white light LED, near-infrared LED, and ultraviolet LED, simultaneously triggering the CMOS image sensor to capture three images corresponding to different spectral channels. In manual mode, a single operation on capture button 4.2 captures only one image for the currently selected spectral channel. Alternatively, a second operation (e.g., a short press) on mode switch button 4.4 cycles through spectral channels (white light -> near-infrared -> ultraviolet -> white light, etc.), enabling independent capture of specific channels.
[0046] To ensure system stability and endurance in portable battery-powered applications, circuit board 4 also integrates a highly efficient power management module 4.6. This module provides independent, optimized power supplies for different chips. For example, it uses a low-dropout regulator (LDO) to provide precise and stable voltages for the image processing chip 4.3 and CMOS image sensor, which are sensitive to power supply noise. At the same time, it uses a more efficient DC-DC converter to power the relatively power-intensive serial flash memory chip 4.1 and Wi-Fi module 4.5, reducing overall energy consumption.
[0047] When the system is working, the user inserts the probe into the mouth and aligns it with the tooth to be examined. Through the above-mentioned key operations, the system collects three-channel images of the teeth, namely visible light, near-infrared and fluorescence. These image data are transmitted to the smart terminal via the WIFI module 4.5. After the application running on the smart terminal receives the image, it will call the image processing method detailed later to automatically analyze and classify the tooth status, and finally present the results and suggestions to the user in a visual way, such as Figure 4 Through the efficient integration and coordinated operation of the above modules, the system of this embodiment achieves the design requirements of miniaturization, high performance and low power consumption, and meets the needs of convenient and accurate detection of early caries in oral diagnosis and treatment.
[0048] Example 2: An image processing method
[0049] The embodiment of the present invention also provides a method for automatic tooth recognition and classification based on multimodal images. The method can be executed on an electronic device that receives the image collected by the above-mentioned tooth recognition system. The flowchart of the method is as follows: Figure 5 Shown, including:
[0050] S501, acquiring a visible light image, a near infrared image, and a fluorescence image of the same tooth area to obtain a multimodal image;
[0051] S502, inputting the multimodal image into a segmentation network, and the segmentation network outputting a binary mask of the tooth region;
[0052] S503, extracting the contour of a single tooth from the binary mask, and determining the minimum circumscribed rectangle and geometric center of gravity of each contour;
[0053] S504, assigning an FDI tooth position code to each of the contours according to the horizontal coordinate of the geometric center of gravity;
[0054] S505, cropping a corresponding single tooth image block from the multimodal image according to the minimum circumscribed rectangle;
[0055] S506, inputting the single tooth image block into a multi-branch convolutional neural network to obtain a unified feature vector output by the multi-branch convolutional neural network;
[0056] S507: Input the unified feature vector into a fully connected classification network, and the fully connected classification network outputs probability values corresponding to different categories.
[0057] The following will describe in detail the specific implementation methods of the key steps in the above process.
[0058] 1. Data Preprocessing and Registration (Corresponding to Step S501)
[0059] This step aims to convert the originally acquired visible light, near-infrared, and fluorescence images, with a resolution of at least 1280x960 pixels, into standardized input suitable for subsequent neural network analysis. First, all images are resized to a uniform size of 224x224 pixels. Second, to eliminate the effects of uneven illumination, each spectral channel is independently z-score normalized to achieve grayscale normalization. Finally, to suppress image noise, algorithms such as Gaussian filtering or bilateral filtering can be used for smoothing.
[0060] Multi-channel registration. Using the visible light image as the reference image, the SIFT (Scale Invariant Feature Transform) algorithm is used to detect and match feature points across the three images. The RANSAC (Random Sample Consensus) algorithm is then used to remove anomalous matching pairs. The visible light image, near-infrared image, and fluorescence image are then aligned at the pixel level. During this step, the pixel offset between each channel must be controlled within 0.5 pixels to ensure the accuracy of subsequent feature extraction.
[0061] 2. Tooth Region Segmentation (Corresponding to Step S502)
[0062] In order to accurately separate the teeth from the complex oral background, this embodiment uses an improved U-Net++ network. The encoder part of the network uses ResNet-34 as the backbone network, and its deep residual structure effectively extracts the hierarchical features of the image. The core improvement is that the spatial attention gate module (Spatial Attention Gate) is embedded in the skip connections between the encoder and the decoder. This module enables the network to adaptively learn and enhance the weights of the tooth area, especially the tooth edge area features, when fusing low-level detail features and high-level semantic features, while suppressing irrelevant features in the background area, thereby significantly improving the accuracy of segmentation. The output layer of the network consists of a 1x1 convolution layer and a Sigmoid activation function, which ultimately generates a binary mask of the tooth area with the same size as the input image.
[0063] During the training phase, the network uses mask images manually annotated by doctors as labels for supervised learning. Its loss function uses a composite loss function, which is a weighted sum of cross-entropy loss and Dice loss:
[0064] .
[0065] in, is the cross entropy loss, which is the standard loss function for classification tasks; and It is Dice loss, which is particularly effective in dealing with class imbalance problems (the tooth area is much smaller than the background) and optimizing the segmentation effect of small targets (such as the tooth gap area). Optimization can be performed through methods such as grid search.
[0066] After the segmentation is completed, post-processing operations can be performed. For example, small connected domains with an area of less than 100 pixels in the mask can be removed to eliminate noise artifacts, and morphological closing operations can be used to fill small holes that may exist inside the teeth to make the mask more complete.
[0067] 3. Single Tooth Positioning and Numbering (Corresponding to Steps S503 and S504)
[0068] After obtaining a binary mask of the tooth region, the Canny edge detection operator is used to extract edge point sets for all teeth in the mask. Considering that tooth-adjacent surfaces may produce contour breaks or gaps during segmentation, this embodiment constructs a convex hull for each detected set of edge points. This method effectively closes these breaks, forming a closed and smooth single tooth contour, overcoming the over-closing errors that can occur at tooth-adjacent interfaces with traditional methods.
[0069] For each generated complete contour, a geometric analysis method is used to calculate its minimum bounding rectangle. This rectangle defines the region of interest (ROI) of the individual tooth and is expressed as the upper left and lower right corner coordinates in the format [x1, y1, x2, y2]. This can be used for image cropping, number matching, and feature extraction. The geometric center of gravity of the contour is also calculated.
[0070] Next, all identified tooth contours were sorted in ascending order by the X-axis coordinate (i.e., horizontal position) of their geometric center of gravity. Using the internationally accepted FDI tooth position system, the maxillary left central incisor (FDI code 11) was used as the starting point for possible numbering, and the sorted teeth were assigned corresponding FDI tooth position codes. This method, combining anatomical structure and spatial position information, enables automatic determination of tooth position and standardized numbering.
[0071] 4. Multimodal Feature Extraction and Fusion (Corresponding to Steps S505 and S506)
[0072] This step is the core of this method. First, according to steps S503 and S505, a multimodal image block of a single tooth is obtained. Then, these image blocks are input into a specially constructed multi-branch convolutional neural network. Figure 6As shown, the backbone structure of the multi-branch convolutional neural network 61 uses the weight-sharing EfficientNet-B3 model to efficiently extract deep semantic features shared by all channels. On this basis, each branch performs customized feature extraction based on the physical characteristics of the spectral image it processes:
[0073] Visible light channel 611: This branch is used to extract the first set of features from visible light image block 621: local binary pattern (LBP) and gray-level co-occurrence matrix (GLCM) features. LBP captures the microtexture of the tooth surface, while GLCM describes the spatial correlation of pixel grayscale. The combination of the two effectively characterizes morphological changes on the tooth surface.
[0074] Near-infrared channel 612: This branch calculates the scattering intensity histogram of the pixel values based on near-infrared image block 622 as the second set of features. Because healthy tooth enamel is highly transparent to 780nm near-infrared light, while demineralized areas scatter it strongly, this histogram can quantitatively reflect characteristics such as internal tooth structural changes and the degree of mineral loss.
[0075] Fluorescence channel 613: This branch extracts the third set of features by analyzing the fluorescence intensity information excited by 405nm ultraviolet light from the fluorescence image block 623. Fluorescence intensity is directly related to the concentration of microorganisms and their metabolites on the tooth surface, so this feature reflects the microbial environment on the tooth surface.
[0076] The extracted multi-source features (the first, second, and third groups of features) are concatenated in the feature fusion module 63 to ultimately form a unified, highly information-concentrated 128-dimensional fused feature vector 631 (unified feature vector), which serves as the high-dimensional input for subsequent classification tasks.
[0077] 5. Classification and Result Output (Corresponding to Steps S507 and S508)
[0078] The 128-dimensional fused feature vector 631 is input into a three-layer fully connected neural network for classification. The network dimensions can be designed sequentially as 256 → 128 → 3. The network ultimately outputs a three-dimensional probability distribution vector [P0, P1, P2].
[0079] This three-dimensional probability distribution vector can be used to classify the dental caries grade. In this application, [P0, P1, P2] correspond to the confidence level of the corresponding tooth being "healthy," "early caries," and "obvious caries," respectively. Then, based on these three probability values and a preset classification confidence threshold T (T can range from 0.5 to 0.7 to control sensitivity to caries risk), the following logical comparison rules are applied to determine the final dental caries grade:
[0080] If P0>max(P1, P2), the classification result is "healthy".
[0081] If P1>P2 and P1>T, the classification result is "early caries".
[0082] If P2 ≥ P1 and P2>T, the classification result is "obvious caries".
[0083] This mechanism ensures high sensitivity of the model while avoiding over-judging minor anomalies. The final output can be presented in two forms: one is structured JSON data, such as {"FDI code":"37","bounding box":[x1,y1,x2,y2],"caries level":"early caries","confidence":0.92}; the other is visual annotation on the original image, such as Figure 4 As shown, the FDI code and caries grade icon are superimposed, and corresponding health advice is provided.
[0084] Figure 7 It is a hardware structure block diagram of an electronic device that implements the image processing method according to an embodiment of the present application.
[0085] like Figure 7 As shown, the electronic device 700 may include one or more processors 702, a system motherboard 708 connected to at least one of the processors 702, a system memory 704 connected to the system motherboard 708, a non-volatile memory (NVM) 706 connected to the system motherboard 708, and a network interface 710 connected to the system motherboard 708.
[0086] The processor 702 may include one or more single-core or multi-core processors. The processor 702 may include any combination of general-purpose processors and specialized processors (e.g., graphics processors, application processors, baseband processors, etc.). In embodiments of the present invention, the processor 702 may be configured to execute one or more embodiments according to various method embodiments of the present application.
[0087] In some embodiments, system board 708 may include any suitable interface controller to provide any suitable interface to at least one of processors 702 and / or any suitable device or component in communication with system board 708 .
[0088] In some embodiments, the system board 708 may include one or more memory controllers to provide an interface to the system memory 704. The system memory 704 may be used to load and store data and / or instructions. In some embodiments, the system memory 704 of the electronic device 700 may include any suitable volatile memory, such as a suitable dynamic random access memory (DRAM).
[0089] NVM 706 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, NVM 706 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as at least one of an HDD (Hard Disk Drive), a CD (Compact Disc) drive, and a DVD (Digital Versatile Disc) drive.
[0090] NVM 706 may include a portion of storage resources installed on a device of electronic device 700, or it may be accessible to the device but not necessarily part of the device. For example, NVM 706 may be accessed over a network via network interface 710.
[0091] In particular, system memory 704 and NVM 706 may respectively include a temporary copy and a permanent copy of instructions 720. Instructions 720 may include instructions that, when executed by at least one of processors 702, cause electronic device 700 to implement the methods of the present application. In some embodiments, instructions 720, hardware, firmware, and / or software components thereof may additionally or alternatively be located in system board 708, network interface 710, and / or processor 702.
[0092] The network interface 710 may include a transceiver for providing a radio interface for the electronic device 700, thereby communicating with any other suitable devices (e.g., a front-end module, an antenna, etc.) via one or more networks. In some embodiments, the network interface 710 may be integrated with other components of the electronic device 700. For example, the network interface 710 may be integrated with at least one of the processor 702, the system memory 704, the NVM 706, and a firmware device (not shown) having instructions. When at least one of the processors 702 executes the instructions, the electronic device 700 implements one or more of the various method embodiments of the present application.
[0093] The network interface 710 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 710 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.
[0094] In one embodiment, at least one of the processors 702 may be packaged together with one or more controllers for the system board 708 to form a system-in-package (SiP). In one embodiment, at least one of the processors 702 may be integrated on the same die with one or more controllers for the system board 708 to form a system-on-chip (SoC).
[0095] Electronic device 700 may further include an input / output (I / O) device 712 connected to system board 708. I / O device 712 may include a user interface to enable a user to interact with electronic device 700; peripheral component interfaces may also be designed to enable peripheral components to interact with electronic device 700. In some embodiments, electronic device 700 may also include a sensor for determining at least one of environmental conditions and location information related to electronic device 700.
[0096] In some embodiments, I / O device 712 may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), speakers, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., an LED flash), and a keyboard.
[0097] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.
[0098] In some embodiments, the sensors may include, but are not limited to, a gyroscope sensor, an accelerometer, a proximity sensor, an ambient light sensor, and a positioning unit. The positioning unit may also be part of or interact with the network interface 710 to communicate with components of a positioning network (e.g., Global Positioning System (GPS) satellites).
[0099] It should be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 700. In other embodiments of the present application, the electronic device 700 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0100] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a system for processing instructions including processor 702 includes any system having a processor such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0101] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in the present invention is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0102] One or more aspects of at least one embodiment may be implemented by instructions stored on a computer-readable storage medium. When the instructions are read and executed by a processor, the electronic device can implement the method of the embodiment described in the present invention.
[0103] According to some embodiments of the present application, a computer storage medium is disclosed. Instructions are stored on the computer storage medium. When the instructions are executed on a computer, the computer executes the image processing method according to the embodiments of the present application.
[0104] The method embodiments of the present application correspond to this embodiment, and this embodiment can be implemented in conjunction with the method embodiments of the present application. The relevant technical details mentioned in the method embodiments of the present application are still valid in this embodiment, and to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the method embodiments of the present application.
[0105] According to some embodiments of the present application, a computer program product is disclosed, including computer-executable instructions, where the instructions are executed by a processor to implement the image processing method according to the embodiment of the present application.
[0106] The method embodiments of the present application correspond to this embodiment, and this embodiment can be implemented in conjunction with the method embodiments of the present application. The relevant technical details mentioned in the method embodiments of the present application are still valid in this embodiment, and to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the method embodiments of the present application.
[0107] It should be understood that the specific embodiments described herein are intended only to illustrate the present application and are not intended to limit the present application. Furthermore, for ease of description, the accompanying drawings illustrate only some, but not all, structures or processes relevant to the present application. It should be noted that throughout this specification, similar reference numerals and letters denote similar items in the accompanying drawings.
[0108] It should be understood that although the terms "first," "second," and the like may be used herein to describe various features, these features should not be limited by these terms. These terms are used merely to distinguish and should not be understood as indicating or implying relative importance. For example, a first feature may be referred to as a second feature, and similarly, a second feature may be referred to as a first feature, without departing from the scope of the exemplary embodiments.
[0109] It should also be noted that, in the description of this application, unless otherwise expressly specified or limited, the terms "disposed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in this embodiment based on specific circumstances.
[0110] Illustrative embodiments of the present application include, but are not limited to, image processing methods, apparatus, media, computer program products, and tooth recognition systems.
[0111] The various aspects of the illustrative embodiments will be described using terms commonly used by those skilled in the art to convey the essence of their work to others skilled in the art. However, it will be apparent to those skilled in the art that some alternative embodiments may be implemented using some of the features described. For purposes of explanation, specific numbers and configurations have been set forth to provide a more thorough understanding of the illustrative embodiments. However, it will be apparent to those skilled in the art that alternative embodiments may be implemented without the specific details. In some other cases, well-known features have been omitted or simplified herein to avoid obscuring the illustrative embodiments of the present application.
[0112] Furthermore, various operations will be described as multiple, separate operations in a manner that is most helpful for understanding the illustrative embodiments; however, the order of description should not be construed to imply that the operations must be performed in order of description, and many of the operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when the described operations are completed, but can also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subprogram, and the like.
[0113] References in the specification to "one embodiment," "an embodiment," "an illustrative embodiment," etc., indicate that the described embodiment may include a particular feature, structure, or property, but that every embodiment may or may not necessarily include the particular feature, structure, or property. Furthermore, these phrases are not necessarily referring to the same embodiment. Furthermore, while particular features are described in conjunction with a specific embodiment, the knowledge of those skilled in the art can influence how those features can be combined with other embodiments, whether or not those embodiments are explicitly described.
[0114] Unless the context dictates otherwise, the terms "comprising," "having," and "including" are synonymous. The phrase "A and / or B" means "(A), (B), or (A and B)."
[0115] The terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings and are only used to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction, and therefore should not be understood as limiting the present invention.
[0116] As used herein, the term "module" may refer to, be part of, or include: memory (shared, dedicated, or group) for running one or more software or firmware programs, application specific integrated circuits (ASICs), electronic circuits and / or processors (shared, dedicated, or group), combinational logic circuits, and / or other suitable components that provide the functionality.
[0117] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order is not required. Rather, in some embodiments, these features may be illustrated in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular drawing does not mean that all embodiments need to include such features. In some embodiments, these features may not be included or may be combined with other features.
[0118] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented in the form of instructions or programs carried or stored on one or more transient or non-transitory machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors, etc. When the instructions or programs are executed by a machine, the machine can perform the various methods described above. For example, the instructions may be distributed over a network or other computer-readable medium. Therefore, machine-readable media may include, but is not limited to, any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), such as floppy disks, optical disks, compact disk read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electronically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, or flash memory or tangible machine-readable memory for transmitting network information via electrical, optical, acoustic, or other forms of signals (e.g., carrier waves, infrared signals, digital signals, etc.). Therefore, machine-readable media includes any form of machine-readable medium suitable for storing or transmitting electronic instructions or machine-readable information (e.g., a computer).
[0119] The embodiments of the present application are described in detail above with reference to the accompanying drawings. However, the application of the technical solution of the present application is not limited to the various applications mentioned in the embodiments of the present application. Various structures and modifications can be easily implemented with reference to the technical solution of the present application to achieve the various beneficial effects mentioned herein. Various changes made within the knowledge of ordinary technicians in this field without departing from the purpose of the present application should fall within the scope of the patent application.
Claims
1. An image processing method for an electronic device, characterized in that: include: Obtain visible light images, near-infrared images, and fluorescence images of the same tooth area to obtain multimodal images; Inputting the multimodal image into a segmentation network, wherein a spatial attention module is embedded in the jump connection between the encoder and decoder of the segmentation network, and the segmentation network outputs a binary mask of the tooth region; Extracting the outline of a single tooth from the binary mask, and determining the minimum circumscribed rectangle and geometric center of gravity of each of the outlines; Assigning an FDI tooth position code to each of the contours according to the horizontal coordinate of the geometric center of gravity; According to the minimum circumscribed rectangle, cropping a corresponding single tooth image block from the multimodal image; The single tooth image block is input into a multi-branch convolutional neural network, where: In the channel corresponding to the visible light image, the local binary pattern and gray-level co-occurrence matrix are extracted as the first set of features. In the channel corresponding to the near-infrared image, the scattering intensity histogram of the pixel values is calculated as the second set of features. In the channel corresponding to the fluorescence image, the fluorescence intensity information of the pixel value is extracted as the third set of features. and performing channel splicing on the first set of features, the second set of features, and the third set of features to obtain a unified feature vector; The unified feature vector is input into a fully connected classification network, and the fully connected classification network outputs probability values corresponding to different categories.
2. The method according to claim 1, characterized in that Performing pixel-level alignment and registration on the visible light image, the near-infrared image, and the fluorescence image to obtain the registered multimodal image; The pixel-level alignment and registration includes: using the SIFT algorithm to perform feature point matching, and using the RANSAC algorithm to eliminate abnormal matching points.
3. The method according to claim 1, characterized in that The segmentation network is a U-Net++ network structure.
4. The method according to claim 1, wherein The extraction of the outline of a single tooth includes: using the Canny operator to perform edge point detection on the binary mask, and constructing a convex hull for the detected edge point set.
5. A tooth recognition system, characterized in that: include: A light source module, including white light LEDs, ultraviolet LEDs, and near-infrared LEDs distributed around the periphery; Imaging module, including lens, CMOS image sensor and optical filter; The control module is configured to control the light source module to switch spectral channels, so that in different spectral channels, the light source module lights up one of the white light LED, the ultraviolet LED and the near-infrared LED respectively; and control the CMOS image sensor to collect images of teeth, The processing module is configured to recognize the image of the tooth according to the image processing method according to any one of claims 1 to 4.
6. The system according to claim 5, characterized in that In the light source module, there are three white light LEDs, three ultraviolet LEDs, and three near-infrared LEDs respectively, which are arranged around the lens at intervals of 120°.
7. The system according to claim 5 or 6, characterized in that The control module includes a capture button and a mode switching button, and the control module is configured to: switching between automatic mode and manual mode in response to a first operation on the mode switching button; In the automatic mode, in response to the operation of the capture button, the spectral channels are switched in sequence, and the CMOS image sensor is triggered to capture the image of the tooth; In the manual mode, in response to the operation of the capture button, the CMOS image sensor is triggered to capture the image of the teeth in the current spectral channel; in response to the second operation of the mode switching button, the spectral channel is switched.
8. The system according to claim 5, wherein: The system further comprises a power management module and a storage module. The power management module adopts a low voltage dropout regulator to supply power to the processing module and the CMOS image sensor, and adopts a DC-DC converter to supply power to the storage module.
9. The system according to claim 5, characterized in that The ultraviolet LED emits ultraviolet light with a wavelength of 405 nm, and the near-infrared LED emits near-infrared light with a wavelength of 780 nm.
10. The system according to claim 5, wherein: The optical filter is configured to block ultraviolet light with a wavelength of 405 nm or less and infrared light with a wavelength of 795 nm or more.
11. An electronic device, characterized in that: The electronic device includes a memory storing computer-executable instructions and a processor. When the instructions are executed by the processor, the electronic device implements the image processing method according to any one of claims 1 to 4.
12. A computer storage medium, characterized in that Instructions are stored on the computer storage medium. When the instructions are executed on a computer, the computer is enabled to execute the image processing method according to any one of claims 1 to 4.
13. A computer program product, characterized in that The method comprises computer executable instructions, wherein the instructions are executed by a processor to implement the image processing method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Eye vision function assessment device and eye vision function assessment method
CN109497925A
CBCT tooth segmentation method based on deep learning
CN112785609A