Interactive oral photography system and method for performing oral cancer screening via artificial intelligence image recognition using same
By combining an interactive oral photography system with image capture and artificial intelligence pattern recognition, the difficulty of oral cancer screening is solved, early identification and risk level warnings are achieved, and it is suitable for oral health monitoring and screening using smartphones.
Patent Information
- Application Number
- PCT/CN2025/085275
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-13
- Filing Date
- 2025-03-27
- Publication Date
- 2025-10-09
AI Technical Summary
Existing technologies have difficulty in effectively identifying oral cancer, especially in the early stages, and traditional medical imaging equipment is not suitable for the diagnosis of potential oral malignancies and oral cancer, resulting in delayed screening and long diagnosis time.
An interactive oral photography system has been developed that combines an image capture unit, an artificial intelligence pattern recognition module, and a storage module. It provides guidance lines and reference diagrams via a smartphone to assist users in capturing oral mucosal images. The system then uses a pattern recognition algorithm for screening and provides risk level warnings.
It achieves comprehensive monitoring and evaluation of oral health status, improves the recognition rate of early oral cancer, simplifies the screening process, and is suitable for oral cancer screening in telemedicine and remote areas.
Smart Images

Figure CN2025085275_09102025_PF_FP_ABST
Abstract
Description
Interactive oral photography system and oral cancer screening method using the system to perform artificial intelligence image recognition
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure claims priority to patent application 63 / 572,407 filed in the United States on April 1, 2024, and this disclosure claims priority to patent application 18 / 980,664 filed in the United States on December 13, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to the technical field of full-oral artificial intelligence image recognition, and more particularly to an interactive oral photography system and a method for performing oral cancer screening using the system using artificial intelligence image recognition. Background Art
[0004] According to the World Health Organization's (WHO) International Agency for Research on Cancer, global rates of lip and oral cancer are increasing, with 377,713 new cases and 177,757 deaths in 2020. This represents an increase from 2018 data and highlights the severity of this disease (Bray et al. 2018). Oral squamous cell carcinoma (OSCC) accounts for over 90% of all oral cancer cases (Ayaz et al. 2011) and is often used interchangeably with the term "oral cancer." It is a fatal disease, ranking 17th among the most common cancers worldwide, with a high prevalence in Central and South Asia and Melanesia. This may be due to the widespread use of betel quid in these regions (Cancer (IARC) 2022; Sung et al. 2021). Smoking and alcohol consumption are major risk factors for the development of oral lesions that can lead to oral cancer. We must increase awareness and education about these risk factors to prevent further spread of this deadly disease (Guha et al. 2014; Hashibe et al. 2007).
[0005] [Corrected 21.04.2025 according to rule 26] The National Institutes of Health Surveillance, Epidemiology, and End Results database shows that the average 5-year overall survival rate for patients with oral cavity and pharynx cancer is 68.0%, and further breakdown shows that the local survival rate for patients with oral cavity and pharynx cancer is 28% ("Cancer of the Oral Cavity and Pharynx-Cancer Stat Facts" 2022).
[0006] Similar to other types of cancer, early detection and prevention play a vital role in reducing oral cancer-related mortality and morbidity. Clinical oral examinations (COEs) are performed by dental specialists as part of routine screening to detect oral cancer. However, due to a lack of public awareness of oral cancer symptoms and a long interval between referrals to oral cancer specialists, most cases are diagnosed only after the disease has progressed. The diverse appearance of oral mucosal lesions makes it difficult for patients and even lay healthcare providers to recognize the subtle visual signs of oral squamous cell carcinoma (OSCC) (Gigliotti, Madathil, and Makhoul 2019; Liao et al. 2019). For example, the tumor may initially appear as a red, erythematous plaque or an ulcerated sore and is often asymptomatic, causing no discomfort until it has progressed.
[0007] Advances in computer vision, deep learning, and artificial intelligence have enabled the development of assistive technologies that can assist in oral screening and provide feedback to healthcare professionals during COEs or patient self-examinations. Among various deep learning models, single-stage detectors, such as You Only Look Once (YOLO) and Single Shot MultiBox Detector (SSD), learn class probabilities and bounding box coordinates from input images by treating object recognition as a simple regression problem. These models are significantly faster than two-stage object detectors (Soviany and Ionescu 2018). Early research on image-based automated diagnosis of oral cancer and other disease entities has focused on the use of imaging techniques such as hyperspectral imaging, autofluorescence imaging, optical coherence tomography (Khanagar et al. 2021), and endoscopic imaging (Heo et al. 2021). All of these medical imaging procedures must be performed using specialized machines, which may not be suitable for addressing the delayed diagnosis of oral potentially malignant disorders (OPMDs) and oral cancer. Since early detection and treatment remain the most effective way to improve oral cancer outcomes, the development of an objective standard visual profile of white-light macroscopic oral photographs that can be used to identify OPMDs and oral cancer provides a significant opportunity to facilitate the oral cancer screening process and facilitate telemedicine-based oral screening by healthcare providers in remote areas. Summary of the Invention
[0008] The present disclosure aims to provide an interactive oral photography system and a method for performing artificial intelligence image recognition for oral cancer screening using the system. The system combines image capture units / steps with application software operations to achieve comprehensive monitoring and evaluation of oral health status. During operation, the user or patient receives guidance from the application, including guide lines and reference schematic images, as well as the results generated by a pattern recognition algorithm. The guided steps performed assist the user or patient in capturing images of the entire oral mucosa (i.e., images of at least two different locations within the oral cavity), digitizing these images, and storing them in a storage module or uploading them to a server. The interactive oral photography system application not only has image capture and storage capabilities but also supports the simultaneous collection and organization of the patient's basic information, such as name, age, gender, and medical history. This data collection and organization provides valuable reference for subsequent diagnosis and treatment.
[0009] In accordance with the aforementioned objectives, the present disclosure provides an interactive oral photography system, comprising: a guidance module configured to provide a reference schematic image at at least two different locations within the oral cavity; an image capture unit communicatively connected to the guidance module and configured to capture at least two oral mucosa images at the at least two different locations within the patient's oral cavity based on a guide line corresponding to the reference schematic image, and digitize the at least two oral mucosa images; an artificial intelligence pattern recognition module communicatively connected to the image capture unit and configured to receive the at least two oral mucosa images and generate a result using a pattern recognition algorithm; and a storage module communicatively connected to the artificial intelligence pattern recognition module and configured to store the at least two oral mucosa images and the result.
[0010] In some embodiments, the at least two different locations include upper gum, upper palate, right cheek, right side of tongue, left side of tongue, right cheek, sublingual area, and lower gum.
[0011] In some embodiments, the guide line corresponds to an outline of the reference schematic image.
[0012] In some embodiments, the interactive oral photography system is a smart phone.
[0013] In some embodiments, the smartphone is installed with an application configured to execute the pattern recognition algorithm.
[0014] The present disclosure provides a method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system, comprising: a login step for logging in a user identity; a guidance step for providing a reference schematic image of at least two different locations within the oral cavity; an image capture step for capturing at least two oral mucosa images of the at least two different locations within the patient's oral cavity based on a guide line corresponding to the reference schematic image, and digitizing the at least two oral mucosa images; an artificial intelligence image recognition step for receiving the at least two oral mucosa images and generating a result using a pattern recognition algorithm; and a storage step for storing the at least two oral mucosa images and the result.
[0015] In some embodiments, the method further includes a risk warning step: providing different color light warnings based on the result to correspond to a risk level of oral cancer, the different color light signals including at least a green light, a yellow light, and a red light, the green light indicating that the risk level is low risk, the yellow light indicating that the risk level is medium risk, and the red light indicating that the risk level is high risk.
[0016] In some embodiments, the at least two different locations include upper gum, upper palate, right cheek, right side of tongue, left side of tongue, right cheek, sublingual area, and lower gum.
[0017] In some embodiments, in the guiding step, before providing the guide line and the reference schematic image at at least two different locations in the oral cavity, the step further includes creating a new folder to store the at least two oral mucosa images.
[0018] In some embodiments, the artificial intelligence image recognition step further includes uploading the at least two oral mucosa images to a server.
[0019] In order to make the above-mentioned objects, features and advantages of the present disclosure more obvious and easy to understand, the following is a detailed description of the specific embodiments listed in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] FIG1 is a schematic diagram of the structure of the interactive oral photography system of the present invention.
[0021] FIG. 2 is a flow chart illustrating a method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0022] FIG3 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0023] FIG. 4 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0024] FIG. 5 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0025] FIG6 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0026] FIG. 7 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0027] FIG8 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0028] FIG. 9 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0029] FIG. 10 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0030] FIG. 11 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0031] FIG. 12 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0032] FIG. 13 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0033] FIG. 14 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0034] FIG. 15 is a schematic diagram illustrating one process of the method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0035] FIG16 is a flow chart of the present disclosure for data creation and model training. DETAILED DESCRIPTION
[0036] The advantages, features, and technical methods achieved by the present disclosure will be described in more detail with reference to exemplary embodiments and the accompanying drawings so as to be more easily understood. The present disclosure can be implemented in different forms, and therefore should not be understood as being limited to the embodiments set forth herein. On the contrary, for those having ordinary knowledge in the relevant technical field, the provided embodiments will make the present disclosure more thorough and comprehensive and fully convey the scope of the present disclosure, and the present disclosure will only be defined by the scope of the attached patent applications.
[0037] In addition, the terms "include" and / or "comprising" refer to the presence of the stated features, regions, wholes, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, regions, wholes, steps, operations, elements, parts and / or their combinations.
[0038] To help you reviewers better understand the content of this disclosure and the effects that can be achieved, various specific embodiments listed in the drawings are described in detail below.
[0039] FIG1 is a schematic diagram of the structure of the interactive oral photography system of the present invention.
[0040] 1 , the interactive oral photography system 10 of the present disclosure includes a guidance module 100, an image capture unit 200, an artificial intelligence pattern recognition module 300, and a storage module 400. In some embodiments, the interactive oral photography system 10 of the present disclosure can be a smartphone, but is not limited thereto.
[0041] The guidance module 100 can be configured to provide a reference schematic image at at least two different locations within the oral cavity (see reference schematic image 120 in FIG. 5 ). In some embodiments, the at least two different locations include the upper gum, upper palate, right cheek, right side of the tongue, left side of the tongue, right cheek, sublingual area, and lower gum (see FIG. 6 to FIG. 13 ), but are not limited thereto.
[0042] The image capture unit 200 can be communicatively connected to the guidance module 100. The image capture unit 200 can be configured to capture at least two oral mucosal images (see images P1-P8 in Figures 6-13 ) at at least two different locations within the patient's oral cavity along a guide line (see guide line 110 in Figures 6-13 ) corresponding to the reference schematic image, and digitize the at least two oral mucosal images. The guide line can correspond to an outline of the reference schematic image.
[0043] The artificial intelligence pattern recognition module 300 can be communicatively connected to the image capture unit 200. The artificial intelligence pattern recognition module 300 can be configured to receive the at least two oral mucosa images and generate a result through a pattern recognition algorithm. In some embodiments, the interactive oral photography system 10 of the present disclosure, which can be a smartphone, can be installed with an application that is configured to execute the pattern recognition algorithm, that is, the application can perform artificial intelligence pattern recognition operations. Based on the result, the interactive oral photography system 10 of the present disclosure, which can be a smartphone, can provide different color light warnings through the application and its display screen to correspond to a risk level of oral cancer; the different color lights include at least a green light, a yellow light, and a red light, the green light indicating that the risk level is low risk, the yellow light indicating that the risk level is medium risk, and the red light indicating that the risk level is high risk.
[0044] The storage module 400 can be communicatively connected to the artificial intelligence pattern recognition module 300. The storage module 400 can be configured to store the at least two oral mucosa images and the results. Finally, the at least two oral mucosa images (see images P1-P8 in Figures 6 to 13) and the results can be uploaded to the server 500.
[0045] FIG. 2 is a flow chart illustrating a method for performing artificial intelligence image recognition for oral cancer screening using an interactive oral photography system according to the present disclosure.
[0046] 1 and 2 , the disclosed method S100 for performing oral cancer screening using an interactive oral photography system using artificial intelligence image recognition may include a login step S110 , a guidance step S120 , an image capture step S130 , an artificial intelligence image recognition step S140 , and a storage step S150 .
[0047] 2 and 3 , in the login step S110 , the user identity is logged in. In some embodiments, the user may enter a password or medical record number to confirm the identity of the login user.
[0048] Please refer to Figures 2 and 5 at the same time. In the guidance step S120, a reference schematic image of at least two different parts of the oral cavity is provided. In other words, the reference schematic image 120 of at least two different parts of the oral cavity is displayed on the display screen of the interactive oral photography system 10 of the present disclosure, which can be, for example, a smart phone. In some embodiments, the at least two different parts include the upper gum, the upper palate, the right cheek, the right side of the tongue, the left side of the tongue, the right cheek, the sublingual area, and the lower gum, but are not limited to this. Please refer to Figures 2, 4 and 5 at the same time. In the guidance step S120, before performing the providing of the guide line and the reference schematic image of at least two different parts of the oral cavity (as shown in Figure 5), it also includes performing the creation of a new folder to store the at least two oral mucosa images (as shown in Figure 4).
[0049] 2 and 6 to 13 , in image capture step S130 , at least two oral mucosa images at at least two different locations within the patient's oral cavity are captured based on the guide line and the reference schematic image, and the at least two oral mucosa images are digitized. Referring to FIG6 to FIG13 , at least two oral mucosa images at at least two different locations within the oral cavity are captured using guide lines 110 at different locations and corresponding to the reference schematic image 120 shown in FIG5 , such as upper gum image P1 (as shown in FIG6 ), upper jaw image P2 (as shown in FIG7 ), right cheek image P3 (as shown in FIG8 ), right tongue image P4 (as shown in FIG9 ), left tongue image P5 (as shown in FIG10 ), right cheek image P6 (as shown in FIG11 ), sublingual image P7 (as shown in FIG12 ), and lower gum image P8 (as shown in FIG13 ).
[0050] Referring to Figures 2 and 14 , in the artificial intelligence image recognition step S140 , the at least two oral mucosal images are received and a result is generated using a pattern recognition algorithm. The shadow recognition step S140 and the pattern recognition algorithm can be implemented by an application installed on the interactive oral photography system 10 of the present disclosure, which can be a smartphone, for example.
[0051] Referring to FIG. 2 and FIG. 14 , in the storage step S150 , the at least two oral mucosal images and the result are stored. In some embodiments, the at least two oral mucosal images and the result can be stored in the interactive oral photography system 10 (or application) of the present disclosure, such as a smartphone, or uploaded to a database for storage, but the present invention is not limited thereto.
[0052] In some embodiments, the present disclosure uses an interactive oral photography system to perform an artificial intelligence image recognition oral cancer screening method S100, which may also include a risk warning step S160. Please refer to Figures 2 and 15 at the same time. In the risk warning step S160, different color light warnings are provided according to the results to correspond to a risk level of oral cancer. In some embodiments, the different color lights may include at least a green light, a yellow light, and a red light. The green light indicates that the risk level is low risk, the yellow light indicates that the risk level is medium risk, and the red light indicates that the risk level is high risk. The risk levels of different color lights may correspond to different treatment suggestions. For example, a low risk with a green light may correspond to the treatment suggestion of "There are currently no obvious symptoms of oral precancerous lesions. Please continue to go to the dentist for teeth cleaning and oral examination every six months."
[0053] FIG16 is a flow chart of the present disclosure for data creation and model training.
[0054] 16 , in order to improve the judgment rate of the interactive oral photography system 10 and the results of executing the artificial intelligence image recognition oral cancer screening method S100 using the interactive oral photography system, the interactive oral photography system 10 may be deeply trained.
[0055] A total of 6,903 white-light macroscopic oral mucosal photographs with and without lesions, taken by senior oral and maxillofacial surgeon D1 using a digital SLR camera (Nikon D200 and D800, Nikon Inc., Tokyo, Japan) from 2006 to 2013, were retrospectively collected from different patients A0 (block B110). A native standard visual archive of white-light macroscopic oral photographs was established, and three-channel images representing different types of oral conditions were analyzed. The VGG Image Annotator (VIA) tool was used to assign the annotation task to two oral and maxillofacial dental residents, D2 and D3 (with 3 and 2 years of experience) (block B230). A detailed lesion polygon mask was generated for each photograph, or the photograph was classified as free of any lesion. These annotated photos were reviewed by D4, a senior oral and maxillofacial surgeon with over 30 years of experience. In biweekly meetings with medical imaging experts (including radiologists and neuroradiologists), the final masks were confirmed and assigned to one of 14 categories (Table 1) as ground truth (Block B240). Multiple data scientists D5 then trained neural network models with various backbones using YOLOv7 (Block B250). The polygonal annotation masks can be easily converted to bounding boxes (Block B260) for testing faster deep learning models for object detection. These masks then generate corresponding recommendation grades (or lesion risk levels), such as low risk (green), medium risk (yellow), and high risk (red), respectively (Block B270).
[0056] Table 1
[0057] In summary, the interactive oral photography system 10 and the method S100 for performing artificial intelligence image recognition for oral cancer screening using the interactive oral photography system 10 of the present disclosure combine the image capture unit 200 / image capture step S130 and application software operations to achieve comprehensive monitoring and evaluation of oral health status. During operation, the user or patient receives guidance from the application, including guide lines 110 and reference schematic images 120, as well as the results generated by the pattern recognition algorithm. The executed guidance step S120 assists the user or patient in capturing images of the entire oral mucosa (i.e., images of at least two different locations within the oral cavity, including upper gum image P1, upper palate image P2, right cheek image P3, right tongue image P4, left tongue image P5, right cheek image P6, sublingual image P7, and lower gum image P8), digitizing these images and storing them in the storage module 400 or uploading them to the server 500. The interactive oral photography system's application not only has image capture and storage functions, but also supports the simultaneous collection and organization of the patient's basic information, such as name, age, gender, medical history, etc. The collection and organization of this data provides valuable reference for subsequent diagnosis and treatment.
[0058] In summary, this case demonstrates significant differences from conventional technical features in terms of purpose, means, and effects. Furthermore, its first invention is practical and meets the patent requirements for inventions. We sincerely request that the Examining Committee carefully examine this matter and grant a patent as soon as possible, so that society can benefit greatly.
[0059] [Description of symbols] 10: Interactive oral photography system 100: Guidance module 110: Guide line 120: Reference schematic image 200: Image capture unit 300: Artificial intelligence pattern recognition module 400: Storage module 500: Server A0: Patient B210~B270: Block D1: Senior oral and maxillofacial surgeon D2: Oral and maxillofacial dental resident D3: Oral and maxillofacial dental resident D4: Senior oral and maxillofacial surgeon D5: Data scientist P1: Upper gum image P2: Upper jaw image P3: Right cheek image P4: Right side of tongue image P5: Left side of tongue image P6: Right cheek image P7: Sublingual image P8: Lower gum image S100: Method for oral cancer screening using artificial intelligence image recognition using an interactive oral photography system S110: Login step S120: Guidance step S130: Image capture step S140: Artificial intelligence image recognition step S150: Storage step S160: Risk warning step
Claims
1. An interactive oral photography system, comprising: a guide module configured to provide a reference graphical image at at least two different locations within the oral cavity; an image capture unit, communicatively connected to the guidance module, configured to capture at least two oral mucosal images at the at least two different locations in the patient's oral cavity along a guide line corresponding to the reference schematic image, and digitize the at least two oral mucosal images; an artificial intelligence pattern recognition module, communicatively connected to the image capture unit and configured to receive the at least two oral mucosa images and generate a result through a pattern recognition algorithm; as well as A storage module is communicatively connected to the artificial intelligence pattern recognition module and is configured to store the at least two oral mucosa images and the result.
2. The interactive oral photography system according to claim 1, wherein: The at least two different locations include upper gum, upper palate, right cheek, right side of tongue, left side of tongue, right cheek, sublingual area, and lower gum.
3. The interactive oral photography system according to claim 1, wherein: The guide line corresponds to a contour of the reference schematic image.
4. The interactive oral photography system according to claim 1, wherein: The interactive oral photography system is a smart phone.
5. The interactive oral photography system according to claim 4, wherein: The smart phone is installed with an application program configured to execute the pattern recognition algorithm.
6. A method for oral cancer screening using artificial intelligence image recognition using the interactive oral photography system according to any one of claims 1 to 5, comprising:
1. Login step: Log in as user; a guiding step: providing a reference schematic image at at least two different locations in the oral cavity; an image capturing step of capturing at least two oral mucosa images at the at least two different locations in the patient's oral cavity along a guide line corresponding to the reference schematic image, and digitizing the at least two oral mucosa images; An artificial intelligence image recognition step: receiving the at least two oral mucosa images and generating a result through a pattern recognition algorithm; as well as A storing step: storing the at least two oral mucosa images and the result.
7. The method of claim 6 further comprises a risk warning step of providing a warning light of different colors based on the result to indicate a risk level of oral cancer, wherein the different color lights include at least a green light, a yellow light, and a red light, wherein the green light indicates a low risk level, the yellow light indicates a medium risk level, and the red light indicates a high risk level.
8. The method of claim 6, wherein: The at least two different locations include upper gum, upper palate, right cheek, right side of tongue, left side of tongue, right cheek, sublingual area, and lower gum.
9. The method of claim 6, wherein: In the guiding step, before providing the guide line and the reference schematic image at at least two different positions in the oral cavity, the step further includes creating a new folder to store the at least two oral mucosa images.
10. The method of claim 6, wherein: The artificial intelligence image recognition step further includes uploading the at least two oral mucosa images to a server.
Citation Information
Patent Citations
Detection system and detection method thereof
CN109394172A
Oral state evaluation system, oral care recommendation system, and oral state notification system
CN116322469A
Imaging apparatus, imaging program, picture determination apparatus, picture determination program, and picture processing system
JP2019213652A
Portable terminal and imaging auxiliary program
JP2021053174A
Deduction device, learning model, learning model generation method, and computer program
WO2020153471A1