Child remote oral cavity detection and diagnosis report return system based on AI image processing

Through the children's remote oral examination system based on AI image processing, the problems of unbalanced children's oral examination resources and low examination efficiency are solved, efficient and accurate remote diagnosis is achieved, and professional diagnostic reports and personalized suggestions are provided, which improves the user experience.

CN120674022AInactive Publication Date: 2025-09-19BEIJING FUAN NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510737517.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing children's oral examination technology has problems such as resource imbalance, low examination efficiency, limited accuracy and imperfect remote examination solutions. It is especially difficult to obtain efficient and professional oral examination services in remote areas.

Method used

The children's remote oral examination system based on AI image processing is used, including a user-end device module, a data transmission module, an artificial intelligence analysis module and an expert review module. Through image quality judgment, encrypted transmission, deep learning algorithm feature recognition, expert review and other technical means, it generates diagnostic reports and provides personalized suggestions.

Benefits of technology

It improves detection efficiency and accuracy, reduces errors in manual detection, ensures the reliability of diagnosis, optimizes user experience, and saves time and labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120674022A_ABST
    Figure CN120674022A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical diagnosis, in particular to a child remote oral cavity detection and diagnosis report return system based on AI image processing. Comprising the steps that a user side equipment module collects a child oral cavity image, quality pre-analysis is carried out in the aspects of focusing / definition, illumination conditions, view finding, shielding object detection and the like, and a user is prompted to optimize shooting in real time; the data transmission module establishes a communication link between the user side equipment module and the server side, and performs privacy protection on data by adopting encryption transmission; the artificial intelligence analysis module analyzes the image data by using a deep learning algorithm and generates a diagnosis report in combination with a knowledge base; the expert auditing module is used for perfectly auditing the generated diagnosis report; and the result feedback module provides visual and personalized detection results and guidance for oral care, medical treatment, reexamination and the like for the user. By means of multi-module cooperation and artificial intelligence assistance, remote accurate detection of the oral cavity of the child is achieved, and the system has important application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical diagnosis technology, and specifically to a children's remote oral detection and diagnosis report return system based on AI image processing. Background Art

[0002] As people's health awareness continues to improve, children's oral health is receiving increasing attention. However, children's oral examinations currently face many challenges. On the one hand, the geographical distribution of professional pediatric dentists is extremely uneven, and children in many remote areas have difficulty obtaining timely and professional oral examination services. Even in cities, taking children to medical institutions for oral examinations often requires parents to spend a lot of time and energy, and children may not cooperate well, which makes the examination difficult. On the other hand, traditional oral examinations mainly rely on manual visual inspection combined with simple tools, which is not very efficient. Moreover, the test results are easily affected by the subjective factors of the testers, and the accuracy is limited.

[0003] Although telemedicine technology is constantly developing, the existing remote testing solutions for the specific field of children's oral examination are still imperfect. There is a lack of a system that can comprehensively consider the characteristics of children, efficiently collect high-quality oral data, use advanced artificial intelligence technology to accurately analyze and diagnose, and provide reliable expert review guarantees. Summary of the Invention

[0004] In response to the problems existing in the above-mentioned background technology, the present invention proposes a children's remote oral examination and diagnosis report feedback system based on AI image processing.

[0005] The AI-based image processing-based remote oral examination and diagnosis report transmission system for children includes: A user terminal device module is used to collect children's oral image data. The user terminal device module includes an image quality judgment unit. The image quality judgment unit performs a quality pre-analysis on the image data after data collection and before uploading to the server. When the data quality is qualified, the user is prompted to upload the data. When the data quality is unqualified, the user is prompted to re-collect the data. The data transmission module is used to establish a communication link between the user-end device module and the server-end to ensure the accurate transmission of oral image data to the server-end, and to transmit the processed results of the server-end back to the user-end device module. The data transmission module uses encryption to protect the privacy of the transmitted data; The artificial intelligence analysis module is set on the server side and includes: The image recognition unit receives the oral cavity image from the user terminal and identifies relevant features to determine whether there is any abnormality. If there is any abnormality, it locates the abnormal area; The intelligent diagnosis unit uses machine learning algorithms to generate a child-specific oral health diagnosis report based on the analysis results of the image recognition unit and the built-in children's oral health knowledge base; The expert review module is set up on the server side and is used to review and confirm the relevant data and preliminary diagnosis results of complex or suspected difficult cases output by the artificial intelligence analysis module, and to supplement and correct the diagnosis report; The result feedback module is used to feed back the final oral examination results to the user-end device module in an intuitive and easy-to-understand form, and provide parents with oral health care recommendations, medical guidance and regular review reminder services based on the test results.

[0006] Preferably, the performing quality pre-analysis on the image and video data specifically includes: Focus / clarity detection: detects whether the image is blurry. If the image is blurred, the user is prompted to retake the photo to obtain a clear image. Lighting condition assessment: evaluates the brightness, contrast, and shadows of the image, and based on the assessment results, prompts the user to improve the lighting conditions or use a flash as appropriate to ensure appropriate image lighting; Field of view / framing control ensures that all relevant structures are within the frame. If relevant structures are found to be incomplete in the frame, the user is prompted to zoom or reposition; Occlusion detection identifies whether fingers, tongues, or cheeks are blocking the field of view. If occlusion is detected, the user is prompted to correct it to ensure that the image fully presents the oral condition.

[0007] Preferably, the focus / clarity detection specifically includes: Image feature extraction: Preprocessing of the collected oral images includes grayscale processing to convert color images into grayscale images and Gaussian filtering to remove noise interference in the images; Apply the Laplace operator to perform convolution operation on the preprocessed image, calculate the second-order derivative of each pixel in the image, and generate a Laplace response image; Extract the gradient magnitude feature from the Laplace response image and characterize the edge sharpness of the image by calculating the gradient magnitude of each pixel; Calculation of clarity rating value: Count the gradient amplitudes of all pixels in the Laplace response image and calculate their variance as the clarity evaluation value; Fuzzy judgment and feedback: Comparing the calculated clarity evaluation value with a preset clarity threshold; If the clarity evaluation value is lower than the threshold, the image is judged to have a blur problem and a corresponding prompt message is generated; Through the display interface of the user-end device, the user is fed back the message "The image is blurry, please retake the photo" and operation instructions are provided.

[0008] Preferably, the focus / clarity detection further includes: Multi-region clarity evaluation: The oral image is divided into multiple regions of interest, including the tooth area, gum area, and occlusal area, and the clarity evaluation value of each area is calculated separately; Weighted comprehensive scoring: assign different weights to different areas and calculate the weighted comprehensive clarity evaluation value; Local blur recognition: If the clarity evaluation value of a certain area exceeds the preset clarity threshold, it is determined that the area is locally blurred, and a targeted prompt message is generated, such as "The maxillary tooth area is blurred, please adjust the angle and retake the photo."

[0009] Preferably, the lighting condition assessment specifically includes: Image preprocessing: The collected oral images were converted from RGB color space to HSV color space, and hue, saturation, and brightness channels were separated; Perform histogram equalization on the brightness channel to enhance the contrast of the image to highlight the details; Global illumination analysis: Brightness evaluation: Calculate the average grayscale value and standard deviation of the brightness channel. If the average grayscale value is lower than the first threshold, the image is considered dark; if it is higher than the second threshold, the image is considered overexposed. Contrast evaluation: Calculate the entropy value or grayscale range of the luminance channel. If the entropy value is lower than the preset entropy threshold or the grayscale range is less than 30% of the total grayscale, the image is judged to have insufficient contrast. Local illumination anomaly detection: Shadow detection: Divide the image into multiple sub-regions, calculate the average brightness and local contrast of each sub-region, and mark a region as a shadow region if its average brightness is lower than 50% of the global average and its local contrast is lower than the contrast threshold. Lighting optimization suggestion generation: Generate corresponding prompt information based on the test results: If the image is dark: "Insufficient lighting, please turn on the room lighting or use the device flash"; If there is a shadow: "Intraoral shadow detected, please adjust the light source angle or shooting position"; If the contrast is insufficient: "The image contrast is low, please try adjusting the shooting environment lighting"; A lighting adjustment diagram is displayed on the user-side device interface, including a flash control button, a brightness adjustment slider, and a shooting angle reference line.

[0010] Preferably, the lighting field of view / framing control specifically includes: Key structure identification: Apply deep learning-based object detection models to identify key anatomical structures in oral images; The model outputs the bounding box coordinates, category labels, and confidence scores of each structure, and filters out detection results with confidence scores below the confidence threshold; Visual field integrity assessment: Define the necessary structure set corresponding to each oral view; Calculate the ratio of the number of actually detected necessary structures to the total number of necessary structures. If the ratio is lower than the completeness threshold, the field of view is determined to be incomplete. Analysis of framing parameters: Scaling ratio assessment: Calculate the ratio of the average size of the detected teeth to the standard size. If the ratio is less than the scaling threshold, a prompt will appear: "Please zoom in for clearer details"; Centering assessment: Calculates the distance between the center points of all detected teeth and the image center. If the distance exceeds the centering threshold, a prompt "Please align the center of your mouth with the center of the viewfinder" will be displayed. Angle deviation assessment: Analyzes the angle between the main tooth arrangement direction and the horizontal / vertical axis of the image. If the angle deviation exceeds the angle threshold, a prompt "Please adjust the camera angle to horizontal / vertical" will be displayed; Intelligent framing guidance: A semi-transparent oral structure template is superimposed on the camera preview interface to show the complete tooth layout that should be included in the current view; Mark detected and missing structures with different colors and dynamically generate adjustment suggestions; Multi-regional collaborative assessment: For panoramic views that require stitching multiple images, plan the optimal shooting sequence and overlapping area; Check whether the overlapping areas between adjacent images meet the stitching requirements; The continuity between adjacent images is verified by the feature point matching algorithm. If the matching degree is lower than the matching threshold, a prompt is given to retake the photo.

[0011] Preferably, the obstruction detection specifically includes: Image preprocessing: Convert the collected oral images into HSV color space and separate the brightness channel and saturation channel; Gaussian blur was applied to remove image noise, and the kernel size was set to 5 × 5 pixels; The brightness channel is binarized using the Otsu threshold method to generate a preliminary segmentation of the foreground and background; Candidate occlusion region extraction: Identify skin areas in HSV space based on skin color model; Perform morphological operations on the binary image to merge adjacent skin areas and remove small noise points; Continuous skin area contours are extracted using a contour detection algorithm, and the area and perimeter of each contour are calculated; The contours with an area larger than 1% of the total image area and an aspect ratio greater than 2:1 are selected as candidate occluder areas; Classification and positioning of obstructions: Extract texture features and shape features of candidate regions; Input the features into the pre-trained support vector machine classifier to distinguish different types of occlusions; For each region classified as an occluder, calculate its spatial position relationship with the key structures of the oral cavity; If the overlapping area between the occluder area and the key structure exceeds the occlusion threshold, it is determined to be a valid occlusion; Occlusion hint generation: Generate prompt information based on the type and location of the occluder: If finger occlusion is detected: "Please move your finger to avoid covering your teeth"; If tongue occlusion is detected: "Please relax your tongue and do not lick your teeth"; If cheek occlusion is detected: "Please open your mouth to allow light to enter your mouth"; The occluded area is marked with a red bounding box on the user-side device interface, and an arrow is superimposed to indicate the removal direction.

[0012] AI remote oral examination for children and the method for transmitting diagnostic reports to parents include: Collect children's oral image data through user-end devices; The collected oral image data is transmitted to the server through an encrypted network communication link; On the server side, deep learning algorithms are used to process oral images, identify different features within the oral cavity, and determine whether there are any abnormalities. Based on the image recognition results and combined with the children's oral health knowledge base, a preliminary diagnosis report is generated; Professional pediatric dentists review the preliminary diagnostic reports generated by AI analysis and provide supplementary diagnoses and revisions for complex or difficult cases; The final oral examination results will be fed back to the user end, and personalized oral health care recommendations, medical guidance and regular check-up reminders will be provided based on the test results.

[0013] Preferably, the collecting of children's oral image data through a user terminal device further includes: During the acquisition process, the image is pre-analyzed in real time for quality, including: Focus / clarity detection: Calculate the image gradient variance using the Laplace operator to determine whether the image is blurred; Lighting condition assessment: Analyze image brightness histogram and color distribution to identify dark, overexposed, and shadow issues; Field of view / framing control: Identify key oral structures through object detection algorithms and assess whether they appear completely within the frame; Occluder detection: Identify occluders based on skin color model and contour analysis; If quality issues are detected, targeted prompt information will be fed back to the user and they will be guided to reshoot.

[0014] Compared with the prior art, the advantages of the present invention are: Improve detection efficiency and accuracy: Using artificial intelligence technology, especially the powerful capabilities of deep learning algorithms in image recognition and intelligent diagnosis, children's oral data can be analyzed quickly and accurately, reducing errors caused by subjective factors in manual detection, and efficiently generating diagnostic results and related recommendations, saving a lot of time and labor costs.

[0015] Guaranteeing diagnostic reliability: The configured expert review module involves a professional team of pediatric dentists to review, confirm, and supplement and correct complex or difficult cases, further ensuring the professionalism and reliability of the diagnostic results, allowing parents to rest assured to carry out subsequent oral care or medical arrangements based on the test results.

[0016] Optimize user experience: The image quality pre-analysis function of the user-end device module, as well as its friendly operation interface and guidance prompts, make it easier for parents to operate and use, help guide children to cooperate with data collection, and avoid repeated operations due to collected data not meeting requirements. This overall improves the user experience in using the system and is also conducive to improving the efficiency of the entire detection process. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is the system architecture diagram of the smart factory dynamic optimization management system based on digital twin and big data analysis proposed in this invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0019] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0020] Example 1: Reference Figure 1 This embodiment provides a remote pediatric oral imaging and diagnosis report transmission system based on AI image processing, including a user-end device module, a data transmission module, an AI analysis module, an expert review module, and a result feedback module. These modules collaborate to complete the entire process from pediatric oral image data acquisition, transmission, analysis and diagnosis, to result feedback. The following details the implementation of each module.

[0021] User terminal equipment module: Hardware equipment: User-end devices can be mobile devices such as smartphones and tablets with corresponding applications installed. These devices have cameras to capture images of children's oral cavity, built-in sensors such as gyroscopes and accelerometers, and are equipped with basic components such as display screens and audio output devices (for displaying prompts to users).

[0022] Image acquisition and preprocessing: Image Capture: After launching the app, the user enters the image capture interface, which displays easy-to-understand instructions. For example, animations demonstrate correct oral opening and head positioning, guiding parents to help their children prepare for the image capture. Furthermore, prompts for capturing oral images at specific angles based on different testing requirements are displayed, along with corresponding text prompts.

[0023] Image quality pre-analysis function realizes: Focus / clarity detection: First, the collected oral images are grayscaled and converted into grayscale images to simplify subsequent calculations and highlight the brightness information of the images.

[0024] Next, the Gaussian filtering algorithm is used to filter the grayscale image to remove noise interference in the image. The size of the Gaussian kernel can be set to 3×3 pixels with a standard deviation of 0.8 to ensure that detailed information such as image edges is retained as much as possible while removing noise.

[0025] The filtered image is then convolved with the Laplacian operator to calculate the second-order derivative of each pixel in the image, generating a Laplacian response image. The Laplacian operator uses a template with a central coefficient of 4 and a coefficient of -1 for the surrounding pixels. This template is convolved with the image pixels to determine the gradient change for each pixel.

[0026] Then extract the gradient amplitude feature in the Laplace response image, and characterize the edge sharpness of the image by calculating the gradient amplitude of each pixel. The calculation formula uses the common gradient amplitude calculation method (for example, for two-dimensional images, the gradient amplitude ,in and for the horizontal and vertical gradients, respectively).

[0027] Finally, the gradient amplitudes of all pixels in the Laplace response image are counted, and their variance is calculated as the clarity evaluation value. This clarity evaluation value is compared with a preset clarity threshold, which is determined through experimental analysis of a large number of clear and blurred oral image samples. For example, the threshold is set to 50 after multiple tests. If the clarity evaluation value is lower than the threshold, the image is considered blurry, and a prompt box pops up on the user's display interface to provide feedback to the user, saying "Image is blurry, please retake the picture." The user is also provided with operational instructions, such as "touch the screen to focus and retake the picture" or "adjust the shooting distance."

[0028] Lighting condition assessment: The collected oral images are first converted from the RGB color space to the HSV color space so that the hue (H), saturation (S), and brightness (V) channel information of the image can be analyzed more conveniently later.

[0029] For the luminance channel, calculate its average grayscale value and standard deviation to assess the overall brightness. If the average grayscale value is lower than the first threshold (e.g., 80 / 255), the image is considered dark; if it is higher than the second threshold (e.g., 180 / 255), the image is considered overexposed.

[0030] The contrast is detected by analyzing the histogram distribution of the luminance channel and calculating the entropy value or grayscale range (i.e., the difference between the maximum grayscale value and the minimum grayscale value) of the luminance channel. If the entropy value is lower than the preset threshold (e.g., 6.0 bits) or the grayscale range is less than 30% of the total grayscale, the image is judged to have insufficient contrast.

[0031] For color balance assessment, the average grayscale value distribution of the RGB channels is analyzed, and the average difference between any two channels is calculated. If the difference exceeds a threshold (such as 20 / 255), color cast is determined to exist.

[0032] For local illumination anomaly detection, the image is divided into multiple sub-regions (for example, using an 8×8 grid division method), and the average brightness and local contrast of each sub-region are calculated. If the average brightness of a region is lower than 50% of the global average and the local contrast is lower than the contrast threshold (such as 0.8), it is marked as a shadow area.

[0033] Targeted prompts are generated based on these detection results. For example, if the image is dark, a prompt will appear: "Insufficient lighting, please turn on room lighting or use the device flash." If shadows are present, a prompt will appear: "Intraoral shadows detected, please adjust the light source angle or shooting position." If the contrast is insufficient, a prompt will appear: "The image has low contrast, please try adjusting the shooting environment lighting." A lighting adjustment diagram is also displayed on the user-side device interface, including a flash control button, a brightness adjustment slider, and a shooting angle guide to facilitate user adjustments.

[0034] Field of view / framing control: A deep learning-based object detection model (such as the YOLOv5 model, pre-trained on a large dataset of images containing labeled images of various oral structures) is used to identify key anatomical structures in oral images. These include, but are not limited to, central incisors, lateral incisors, canines, premolars, and molars; soft tissue landmarks such as the gingival margin, alveolar crest, and palatal vault; and orthodontic devices (e.g., brackets, archwires) or restorations (e.g., fillings, crowns). The model outputs the bounding box coordinates, class label, and confidence score for each structure. Detection results with a confidence score below a threshold (e.g., 0.95) are filtered to remove less reliable detections.

[0035] A set of required structures for each oral view is defined (e.g., including at least bilateral first molars and central incisors). The ratio of the number of required structures actually detected to the total number of required structures is calculated. If the ratio falls below a completeness threshold (e.g., 80%), the field of view is considered incomplete. For specific pathological areas (e.g., suspected caries sites), their completeness within the image is checked. If truncation is detected, it is marked as a field of view defect.

[0036] In terms of framing parameter analysis, the ratio of the average size of the detected teeth to the standard size is calculated to determine whether the current zoom is appropriate. If the ratio is less than the zoom threshold (such as 0.7), a prompt will be displayed saying "Please zoom in to get clearer details"; the distance between the center point of all detected teeth and the center of the image is calculated. If it exceeds the centering threshold (such as 15% of the diagonal length of the image), a prompt will be displayed saying "Please align the center of the mouth with the center of the framing frame"; the angle between the main tooth arrangement direction and the horizontal / vertical axis of the image is analyzed. If the angle deviation exceeds the threshold (such as 10°), a prompt will be displayed saying "Please adjust the camera angle to horizontal / vertical".

[0037] In the intelligent framing guidance, a translucent oral structure template is superimposed on the camera preview interface, showing the complete tooth layout that should be included in the current view, marking the detected structures (such as green) and missing structures (such as red) with different colors, and displaying targeted prompts such as text prompts "Please include the right molars". At the same time, adjustment suggestions are dynamically generated, such as "Move the camera up to include the maxillary teeth" or "Shift to the left to show all front teeth". Augmented reality (AR) technology can also be used to display a 3D virtual tooth model on the preview interface to indicate the areas that need additional shooting, helping users obtain a complete oral image.

[0038] For panoramic views that require stitching multiple images, plan the optimal shooting sequence and overlapping area, check whether the overlapping areas between adjacent images meet the stitching requirements (such as an overlap rate of at least 30%), and verify the continuity between adjacent images through feature point matching (such as SIFT or ORB algorithms, extracting feature points in the image and then matching based on feature descriptors). If the matching degree is lower than the matching degree threshold (such as 85%), prompt to reshoot to ensure that the multiple images finally obtained can be accurately stitched into a complete panoramic oral image.

[0039] Occlusion detection: First, the collected oral images were converted into the HSV color space, and the brightness (V) channel and saturation (S) channel were separated. Gaussian blur was applied to eliminate image noise, and the kernel size was set to 5×5 pixels. The brightness channel was binarized using the Otsu thresholding method to generate a preliminary segmentation of the foreground (oral tissue) and background (possible occluders). The Otsu thresholding method can automatically find an optimal threshold based on the grayscale histogram of the image to maximize the variance between the two segments after segmentation, achieving a good preliminary segmentation effect.

[0040] Skin areas are identified in the HSV space based on the skin color model. The skin color range is set to H∈[0,25]∪[160,180], S∈[25,160], and V∈[40,240]. Morphological operations (dilation and erosion) are performed on the binary image (dilation and erosion, by defining a suitable structural element, such as a 3×3 square structural element, dilating the image to merge adjacent skin areas, and then eroding to remove small noise). Adjacent skin areas are merged and small noise is removed. Then, the contour detection algorithm is used to extract continuous skin area contours, and the area and perimeter of each contour are calculated.

[0041] Contours with an area greater than 1% of the total image area and an aspect ratio greater than 2:1 are selected as candidate occluder areas, and texture features (such as LBP and HOG, which describe the texture features of the image by dividing the image into small cells and counting the gradient direction histogram in each cell) and shape features (such as circularity and rectangularity, which can be obtained by calculating the relationship between the perimeter and area of ​​the area enclosed by the contour, and rectangularity measured by the ratio of the contour area to the area of ​​the circumscribed rectangle) are extracted from the candidate areas.

[0042] The features are input into a pre-trained support vector machine (SVM) classifier to distinguish different types of occluders, such as fingers, tongues, and cheeks. For each area classified as an occluder, its spatial position relationship with key oral structures (such as teeth and gums) is calculated. If the overlapping area of ​​the occluder area and the key structure exceeds a threshold (such as 15% of the area of ​​the key structure), it is judged as a valid occlusion.

[0043] Generate customized prompt information based on the type and location of the obstruction. For example: if a finger obstruction is detected, the prompt will say "Please move your fingers to avoid blocking your teeth"; if a tongue obstruction is detected, the prompt will say "Please relax your tongue and do not lick your teeth"; if a cheek obstruction is detected, the prompt will say "Please open your mouth to let light into your mouth". The obstructed area is marked with a red bounding box on the user-end device interface, and an arrow is superimposed to indicate the removal direction, so that users can clearly receive the prompt and take corresponding actions.

[0044] Data transmission: After the user-end device completes image acquisition and quality pre-analysis, the oral image data that meets the requirements is sent to the server through the data transmission module.

[0045] Encrypted transmission mechanism: The data transmission module uses SSL / TLS encryption protocol to encrypt the transmitted data to ensure the privacy and security of the data during network transmission.

[0046] Network transmission protocol selection: Based on the current network environment, the data transmission module will flexibly select an appropriate network transmission protocol, such as HTTP or MQTT, for data transmission. In scenarios where the network is stable and real-time requirements are not extremely high, HTTP is preferred. Its request-response model facilitates reliable data transmission. In scenarios where real-time requirements are high and frequent message push is required (such as real-time feedback on server-side processing progress), MQTT is used. Its lightweight, low-bandwidth nature and support for asynchronous communication make it better suited to data transmission needs in mobile network environments.

[0047] Artificial intelligence analysis module: Image recognition unit: Deep Learning Model Selection and Training: The image recognition unit utilizes a convolutional neural network (CNN) as its core algorithm. In this embodiment, a modified VGGNet architecture was selected (an appropriately lightweight modification of the original VGGNet to accommodate the computational resource constraints of mobile devices while ensuring sufficient recognition accuracy). During model training, tens of thousands of oral images of children of varying ages and oral health conditions were collected and annotated. These annotations included key features such as tooth morphology (e.g., presence of defects, deformities), color (normal color, discoloration due to caries), and gum condition (presence of redness, swelling, bleeding, etc.). Stochastic Gradient Descent (SGD) was used as the optimization algorithm, with an initial learning rate of 0.01 and a learning rate decay strategy that halved the learning rate every 10 training epochs. The batch size was set to 64, and the cross-entropy loss function was used as the loss function. Through extensive iterative training, the model was able to accurately extract relevant features from the input oral images and perform classification and recognition.

[0048] Image feature recognition process: After receiving an oral image from the user, the image is first preprocessed, including resizing (uniformly adjusting it to the input size required for model training, such as 224×224 pixels) and normalization (normalizing the image pixel values ​​to a specific range, such as [0,1]) to ensure that the data format of the input model meets the requirements. The preprocessed image is then input into the trained CNN model, which gradually extracts deep features of the image through multiple convolutional and pooling layers. For example, different levels of convolutional layers can capture characteristic information such as tooth edge contours, texture details, and the boundary between gums and teeth. Finally, the fully connected layer performs feature integration and classification, outputting judgment results on features such as tooth morphology, color, and gum condition, determining whether there are any abnormalities, and accurately locating abnormal areas. These results are then transmitted to the intelligent diagnosis unit for further comprehensive analysis.

[0049] Intelligent diagnostic unit: Knowledge Base Construction and Integration: The intelligent diagnostic unit relies on a rich pediatric oral health knowledge base, encompassing information on oral developmental characteristics of children of different ages, diagnostic criteria for common diseases, and correlations between diseases and oral features. This knowledge base is constructed by collecting authoritative oral medical literature and clinical case data, and by inviting professional pediatric dentists to organize and input knowledge. Furthermore, the knowledge base is regularly updated to incorporate the latest oral medical research findings and clinical diagnostic experience, ensuring the timeliness and accuracy of its knowledge.

[0050] Comprehensive diagnostic process: The intelligent diagnostic unit receives the analysis results from the image recognition unit and conducts a comprehensive analysis based on the built-in children's oral health knowledge base. For example, when the image recognition unit detects a suspicious area of ​​discoloration on a tooth, the intelligent diagnostic unit searches the knowledge base for matching disease information. By comparing the diagnostic criteria for common diseases (such as the morphology and color changes of caries), the intelligent diagnostic unit uses machine learning algorithms (such as rule-based reasoning algorithms combined with probabilistic statistical models to make comprehensive judgments based on the probability of occurrence of different features and the diagnostic rules set in the knowledge base) to generate an oral health diagnosis report for the child. The report includes information such as whether the disease exists, the type of disease, the severity, and possible development trends. The diagnosis report is then sent to the expert review module for further review and confirmation.

[0051] Expert review module: Workflow: After receiving an oral health diagnosis report from the intelligent diagnostic unit, the system assigns it to a specialist for review based on the case's complexity (for example, cases involving multiple coexisting conditions, atypical disease presentations, or suspected difficult conditions are considered complex). After logging into a dedicated review interface, the specialist can view detailed case information, including the oral image data uploaded by the user and the preliminary diagnostic report generated by the intelligent diagnostic unit.

[0052] Review and result feedback: The expert will carefully compare the oral features in the image with the conclusions in the preliminary diagnosis report, and use their professional knowledge and clinical experience to review and confirm. If the expert believes that the preliminary diagnosis is accurate, they will confirm it and enter the diagnosis report directly into the result feedback module. If they find that the diagnosis is inaccurate, important information is omitted, or further observation is required, the expert can supplement and correct the diagnosis report on the review interface, such as adding a more detailed description of the condition, adjusting the diagnosis type or severity of the disease, etc., and can also attach corresponding annotations to explain the reasons and basis for the corrections. The revised diagnosis report will be entered into the result feedback module as the final oral examination result.

[0053] Result feedback module: Results Feedback: The results feedback module is responsible for providing the final oral examination results to the user-end device module in an intuitive and easy-to-understand format. The diagnostic report is presented on the user's application interface with both pictures and text. For example, a clear image will be used to mark the abnormal oral area, accompanied by concise and clear text describing the disease name, severity, and other key information.

[0054] Personalized Service Provision: Based on the test results, parents are provided with personalized oral health care recommendations, medical guidance, and regular checkup reminders. Oral health care recommendations are customized based on factors such as the child's age, current oral health status (such as whether they have caries, gingivitis, and other conditions), and daily oral hygiene habits. For example, for children with caries, parents will be advised to supervise their children's daily brushing frequency, precautions for using fluoride toothpaste, and control their sweet intake. Regarding medical guidance, parents will be clearly informed whether further examination is necessary. If so, geolocation information will be used to recommend nearby professional pediatric dental clinics. Details such as the clinic's name, address, contact number, and specialized dental treatments will be provided, making it easier for parents to quickly choose the right clinic. Regular checkup reminders are set at appropriate times based on the nature of the condition and recovery period. Reminders can be sent to parents via in-app push notifications or text messages, ensuring they don't miss checkups and can stay informed of their children's oral health recovery.

[0055] Through the above specific implementation methods, the various modules work together to realize the complete process of the artificial intelligence-based remote children's oral detection system and method from data collection, analysis and diagnosis to result feedback, providing an efficient, accurate and convenient solution for the remote monitoring and diagnosis of children's oral health.

[0056] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0057] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. The remote oral examination and diagnosis report transmission system for children based on AI image processing is characterized by: include: A user terminal device module is used to collect children's oral image data. The user terminal device module includes an image quality judgment unit. The image quality judgment unit performs a quality pre-analysis on the image data after data collection and before uploading to the server. When the data quality is qualified, the user is prompted to upload the data. When the data quality is unqualified, the user is prompted to re-collect the data. The data transmission module is used to establish a communication link between the user-end device module and the server-end to ensure the accurate transmission of oral image data to the server-end, and to transmit the processed results of the server-end back to the user-end device module. The data transmission module uses encryption to protect the privacy of the transmitted data; The artificial intelligence analysis module is set on the server side and includes: The image recognition unit receives the oral cavity image from the user terminal and identifies relevant features to determine whether there is any abnormality. If there is any abnormality, it locates the abnormal area; The intelligent diagnosis unit uses machine learning algorithms to generate a child-specific oral health diagnosis report based on the analysis results of the image recognition unit and the built-in children's oral health knowledge base; The expert review module is set up on the server side and is used to review and confirm the relevant data and preliminary diagnosis results of complex or suspected difficult cases output by the artificial intelligence analysis module, and to supplement and correct the diagnosis report; The result feedback module is used to feed back the final oral examination results to the user-end device module in an intuitive and easy-to-understand form, and provide parents with oral health care recommendations, medical guidance and regular review reminder services based on the test results.

2. The AI ​​image processing-based children's remote oral examination and diagnosis report transmission system according to claim 1 is characterized in that: The performing quality pre-analysis on the image and video data specifically includes: Focus / clarity detection: detects whether the image is blurry. If the image is blurred, the user is prompted to retake the photo to obtain a clear image. Lighting condition assessment: evaluates the brightness, contrast, and shadows of the image, and based on the assessment results, prompts the user to improve the lighting conditions or use a flash as appropriate to ensure appropriate image lighting; Field of view / framing control ensures that all relevant structures are within the frame. If relevant structures are found to be incomplete in the frame, the user is prompted to zoom or reposition; Occlusion detection identifies whether fingers, tongues, or cheeks are blocking the field of view. If occlusion is detected, the user is prompted to correct it to ensure that the image fully presents the oral condition.

3. The AI ​​image processing-based children's remote oral examination and diagnosis report transmission system according to claim 2 is characterized in that: The focus / clarity detection specifically includes: Image feature extraction: Preprocessing of the collected oral images includes grayscale processing to convert color images into grayscale images and Gaussian filtering to remove noise interference in the images; Apply the Laplace operator to perform convolution operation on the preprocessed image, calculate the second-order derivative of each pixel in the image, and generate a Laplace response image; Extract the gradient magnitude feature from the Laplace response image and characterize the edge sharpness of the image by calculating the gradient magnitude of each pixel; Calculation of clarity rating value: Count the gradient amplitudes of all pixels in the Laplace response image and calculate their variance as the clarity evaluation value; Fuzzy judgment and feedback: Comparing the calculated clarity evaluation value with a preset clarity threshold; If the clarity evaluation value is lower than the threshold, the image is judged to have a blur problem and a corresponding prompt message is generated; Through the display interface of the user terminal device, the user is fed back the message "Image is blurry, please retake the photo" and provided with operation instructions.

4. The AI ​​image processing-based children's remote oral examination and diagnosis report transmission system according to claim 2 is characterized in that: The focus / clarity detection further includes: Multi-region clarity evaluation: The oral image is divided into multiple regions of interest, including the tooth area, gum area, and occlusal area, and the clarity evaluation value of each area is calculated separately; Weighted comprehensive scoring: assign different weights to different areas and calculate the weighted comprehensive clarity evaluation value; Local blur recognition: If the clarity evaluation value of a certain area exceeds the preset clarity threshold, it is determined that the area is locally blurred and a targeted prompt message is generated, such as "The maxillary tooth area is blurred, please adjust the angle and retake the photo." 5. The AI ​​image processing-based children's remote oral examination and diagnosis report transmission system according to claim 2 is characterized in that: The lighting condition assessment specifically includes: Image preprocessing: The collected oral images were converted from RGB color space to HSV color space, and hue, saturation, and brightness channels were separated; Perform histogram equalization on the brightness channel to enhance the contrast of the image to highlight the details; Global illumination analysis: Brightness evaluation: Calculate the average grayscale value and standard deviation of the brightness channel. If the average grayscale value is lower than the first threshold, the image is considered dark; if it is higher than the second threshold, the image is considered overexposed. Contrast evaluation: Calculate the entropy value or grayscale range of the luminance channel. If the entropy value is lower than the preset entropy threshold or the grayscale range is less than 30% of the total grayscale, the image is judged to have insufficient contrast. Local illumination anomaly detection: Shadow detection: Divide the image into multiple sub-regions, calculate the average brightness and local contrast of each sub-region, and mark a region as a shadow region if its average brightness is lower than 50% of the global average and its local contrast is lower than the contrast threshold. Lighting optimization suggestion generation: Generate corresponding prompt information based on the test results: If the image is dark: "Insufficient lighting, please turn on the room lighting or use the device flash"; If there is a shadow: "A shadow in the mouth has been detected. Please adjust the light source angle or shooting position"; If the contrast is insufficient: "The image contrast is low. Please try adjusting the shooting environment lighting"; A lighting adjustment diagram is displayed on the user-side device interface, including a flash control button, a brightness adjustment slider, and a shooting angle reference line.

6. The AI ​​image processing-based children's remote oral examination and diagnosis report transmission system according to claim 2 is characterized in that: The lighting field of view / framing control specifically includes: Key structure identification: Apply deep learning-based object detection models to identify key anatomical structures in oral images; The model outputs the bounding box coordinates, category labels, and confidence scores of each structure, and filters out detection results with confidence scores below the confidence threshold; Visual field integrity assessment: Define the necessary structure set corresponding to each oral view; Calculate the ratio of the number of actually detected necessary structures to the total number of necessary structures. If the ratio is lower than the completeness threshold, the field of view is determined to be incomplete. Analysis of framing parameters: Scaling ratio assessment: Calculate the ratio of the average size of the detected teeth to the standard size. If the ratio is less than the scaling threshold, a prompt "Please zoom in for clearer details" will be displayed; Centering assessment: Calculates the distance between the center points of all detected teeth and the image center. If the distance exceeds the centering threshold, a prompt "Please align the center of your mouth with the center of the viewfinder" will be displayed. Angle deviation assessment: Analyzes the angle between the main tooth arrangement direction and the horizontal / vertical axis of the image. If the angle deviation exceeds the angle threshold, a prompt "Please adjust the camera angle to horizontal / vertical" will be displayed; Intelligent framing guidance: A semi-transparent oral structure template is superimposed on the camera preview interface to show the complete tooth layout that should be included in the current view; Mark detected and missing structures with different colors and dynamically generate adjustment suggestions; Multi-regional collaborative assessment: For panoramic views that require stitching multiple images, plan the optimal shooting sequence and overlapping area; Check whether the overlapping areas between adjacent images meet the stitching requirements; The continuity between adjacent images is verified by the feature point matching algorithm. If the matching degree is lower than the matching threshold, a prompt is given to retake the photo.

7. The AI ​​image processing-based children's remote oral examination and diagnosis report transmission system according to claim 2 is characterized in that: The obstruction detection specifically includes: Image preprocessing: Convert the collected oral images into HSV color space and separate the brightness channel and saturation channel; Gaussian blur was applied to remove image noise, and the kernel size was set to 5 × 5 pixels; The brightness channel is binarized using the Otsu threshold method to generate a preliminary segmentation of the foreground and background; Candidate occlusion region extraction: Identify skin areas in HSV space based on skin color model; Perform morphological operations on the binary image to merge adjacent skin areas and remove small noise points; Continuous skin area contours are extracted using a contour detection algorithm, and the area and perimeter of each contour are calculated; The contours with an area larger than 1% of the total image area and an aspect ratio greater than 2:1 are selected as candidate occluder areas; Classification and positioning of obstructions: Extract texture features and shape features of candidate regions; Input the features into the pre-trained support vector machine classifier to distinguish different types of occlusions; For each region classified as an occluder, calculate its spatial position relationship with the key structures of the oral cavity; If the overlapping area between the occluder area and the key structure exceeds the occlusion threshold, it is determined to be a valid occlusion; Occlusion hint generation: Generate prompt information based on the type and location of the occluder: If finger occlusion is detected: "Please move your finger to avoid covering your teeth"; If tongue occlusion is detected: "Please relax your tongue and do not lick your teeth"; If cheek occlusion is detected: "Please open your mouth to allow light to enter your mouth"; The occluded area is marked with a red bounding box on the user-side device interface, and an arrow is superimposed to indicate the removal direction.

8. AI children's remote oral examination and parent-side diagnostic report transmission method, characterized by: include: Collect children's oral image data through user-end devices; The collected oral image data is transmitted to the server through an encrypted network communication link; On the server side, deep learning algorithms are used to process oral images, identify different features within the oral cavity, and determine whether there are any abnormalities. Based on the image recognition results and combined with the children's oral health knowledge base, a preliminary diagnosis report is generated; Professional pediatric dentists review the preliminary diagnostic reports generated by AI analysis and provide supplementary diagnoses and revisions for complex or difficult cases; The final oral examination results will be fed back to the user end, and personalized oral health care recommendations, medical guidance and regular check-up reminders will be provided based on the test results.

9. The method for remote oral examination of children by AI and transmission of diagnosis report to parents according to claim 8, characterized in that: The collecting of children's oral image data by the user terminal device also includes: During the acquisition process, the image is pre-analyzed in real time for quality, including: Focus / clarity detection: Calculate the image gradient variance using the Laplace operator to determine whether the image is blurred; Lighting condition assessment: Analyze image brightness histogram and color distribution to identify dark, overexposed, and shadow issues; Field of view / framing control: Identify key oral structures through object detection algorithms and assess whether they appear completely within the frame; Occluder detection: Identify occluders based on skin color model and contour analysis; If quality issues are detected, targeted prompt information will be fed back to the user and they will be guided to reshoot.