system
The system uses generative adversarial networks and variational autoencoders to predict future tooth alignment and facial changes, addressing uncertainty in orthodontic treatments by simulating scenarios and offering personalized predictions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional orthodontic treatments lack the ability to accurately predict future tooth alignment and facial appearance changes, leading to uncertainty in treatment choices and underutilization of past treatment data.
A system that utilizes facial and dental images to simulate future tooth alignment and facial features using generative adversarial networks and variational autoencoders, supported by an interactive interface for user interaction and prediction based on past treatment data.
Provides accurate and personalized predictions of future tooth alignment and facial changes, enabling informed treatment decisions by simulating various scenarios and providing real-time explanations.
Smart Images

Figure 2026073432000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional orthodontic treatment, it is difficult to predict how future tooth alignment and facial appearance will change depending on the presence or absence and method of treatment. For this reason, patients and their guardians may feel不安 when choosing a treatment method. In addition, with the conventional methods, past treatment data cannot be fully utilized, and there is a demand for improving the accuracy of diagnosis based on individual cases.
Means for Solving the Problems
[0005] This invention provides a system that receives facial photographs and dental images of children and uses them to make predictions. The system extracts features from the images and uses a generative adversarial network and a variational autoencoder to simulate future tooth alignment and facial features. It also supports the selection of treatment methods by predicting similar cases based on past treatment data. Furthermore, it has the function to display the simulation results on an interface and communicate with the user using an interactive interface.
[0006] "Image preprocessing" refers to processes such as resizing, noise reduction, color correction, and cropping of important parts of a received image in order to make it easier to analyze.
[0007] "Feature point extraction" is the process of identifying important locations and shapes on the face and teeth using a convolutional neural network.
[0008] Generative Adversarial Networks (GANs) are a machine learning technique that improves the accuracy of generating new data by having two neural networks compete against each other.
[0009] A variational autoencoder (VAE) is a model that uses latent variables in the data to generate new data samples, and is used to learn the distribution of the data.
[0010] "Simulation" is a process that virtually reproduces a certain situation or change to predict future changes in tooth alignment and facial appearance.
[0011] "Predicting similar cases" is the process of identifying similar cases that match the current symptoms based on accumulated past treatment data, and then predicting future changes.
[0012] A "user interface" refers to the screen layout and operating methods that allow a user to interact with a system and visually confirm information and results.
[0013] An "interactive interface" is a system design that enables real-time communication between the user and the system, and includes functions such as question answering. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when combined with an emotion engine.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] The system according to the present invention combines image recognition technology and generative AI technology to predict a child's future tooth alignment and facial features. It is implemented through the interaction of the user, server, and terminal as follows.
[0036] First, users upload a photo of their child's face and images of their teeth to the system from a mobile device or computer. These images are used as foundational data to simulate future conditions.
[0037] Next, the server receives the uploaded images and performs a series of preprocessing steps. These preprocessing steps include resizing, denoising, and color correction. This prepares the images for accurate extraction of facial and dental features.
[0038] Next, the server uses a convolutional neural network (CNN) to extract key feature points of the face and teeth from the image. At this stage, for example, the position and shape of the eyes, nose, mouth, and teeth are analyzed and used as foundational data for predicting future changes.
[0039] Based on the extracted feature data, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate multiple future scenarios depending on whether orthodontic treatment is performed and what method is used. These include various cases such as when no orthodontic treatment is performed and when a specific treatment method is used.
[0040] Next, the server retrieves information for analyzing past cases, particularly those that are similar. In this process, past treatment outcome data is searched using a transformer model to identify similarities with the current case.
[0041] The generated simulation results are then sent to the terminal. The terminal visually displays these images and information to the user. This display is organized in a format that allows the user to intuitively understand and compare treatment options.
[0042] Finally, when a user asks a question based on the simulation results, the terminal uses ChatGPT® to provide explanations in real time that correspond to the content of the question, helping the user understand. For example, if a user requests to see "how the face would change after partial orthodontic treatment," that scenario is immediately generated, and detailed images and explanations are provided.
[0043] Based on the above, the system of the present invention provides users with detailed predictive data to select the optimal orthodontic treatment method, and continuously improves its accuracy through the use of past data and expert feedback.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] Users upload photos of their child's face and teeth to the system. The image files must be submitted in a specified format, so users should prepare them according to the guidelines.
[0047] Step 2:
[0048] The server receives the uploaded images and performs preprocessing on them. Specifically, this includes resizing, denoising, and color correction. This ensures the quality necessary for analysis.
[0049] Step 3:
[0050] The server feeds the preprocessed images to a convolutional neural network (CNN) to extract facial and tooth feature points. This feature extraction process identifies facial landmarks and tooth contours.
[0051] Step 4:
[0052] Based on the extracted features, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future tooth alignment and facial features. Multiple treatment scenarios are included, such as no orthodontics, partial orthodontics, and full orthodontics.
[0053] Step 5:
[0054] The server references past treatment data and searches for similar cases. It uses a transformer model to analyze the database and find the treatment pattern that best matches the current case.
[0055] Step 6:
[0056] The generated simulation results are sent to the terminal, which displays them visually to the user. The display includes detailed images showing the future facial features and teeth alignment, along with explanations of treatment options.
[0057] Step 7:
[0058] When a user asks a question about the simulation results, the terminal uses ChatGPT to generate a response in real time, providing detailed explanations and additional information.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] It is difficult for parents and medical professionals to obtain accurate early predictions about a child's future oral structure and facial shape. In particular, information is limited to determine which treatment options are most effective when corrections or adjustments are needed. This increases the risk of incorrect treatment choices and unnecessary procedures. Furthermore, the lack of readily available methods to obtain predictions based on past cases makes it difficult to select the optimal treatment plan for each individual child.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] This invention includes a server that receives human facial and oral images and preprocesses them based on pre-set criteria, extracts facial and oral feature points from the preprocessed images using a machine learning algorithm, and simulates future oral structures and facial shapes using an advanced generative algorithm based on the extracted features. This makes it easy for parents and medical professionals to obtain accurate and diverse predictions based on various treatment scenarios.
[0064] A "human facial image" is a digital image that visually records the characteristics of an individual's face and includes information that represents the structure and shape of the face.
[0065] An "oral image" is an image that captures the inside of a person's mouth, particularly the condition of the teeth and gums, and is used for treatment and diagnosis.
[0066] "Preprocessing" refers to the process of converting image data into a format that is easy to analyze, and encompasses all processes including resizing, noise reduction, and color correction.
[0067] A "machine learning algorithm" is a technology that learns patterns and features from large amounts of data to make predictions and decisions, and is applied to image analysis and feature extraction.
[0068] "Feature point extraction" is the process of identifying and extracting points that represent important locations and shapes within an image, and these points are used as basic data for analysis and simulation.
[0069] An "advanced generative algorithm" is a complex computational method used to generate new information from data, and is particularly used when predicting future states.
[0070] "Oral structure" refers to the specific shape and arrangement of the inside of the mouth, such as the teeth, jawbone, and oral mucosa.
[0071] "Facial shape" refers to the specific form and contour of the external structure of a human face, and represents individual characteristics.
[0072] A "similar case" refers to a past case that shares many commonalities with the case currently being analyzed.
[0073] "Medical information" refers to information about a patient, including treatment and diagnostic results, and historical health data, which is used to determine and predict treatment plans.
[0074] "Diverse orthodontic treatment methods" refers to the various treatment options available in orthodontic treatment, including specific examples such as bracket orthodontics and clear aligner orthodontics.
[0075] A "human-to-human interaction device" is a device or interface for exchanging information with a human user, and is designed to provide information and answer questions through dialogue.
[0076] The embodiments for carrying out the present invention will now be described. This system involves the interaction of a user, a server, and a terminal to predict a child's future oral structure and facial shape. The roles of each are described in detail below.
[0077] First, users upload images of their child's face and mouth to the system using a mobile device or computer. These devices include typical smartphones and PCs with cameras. These images form the basis of the simulation data.
[0078] Next, these images are sent to a server. The server, equipped with high-performance computers and image processing software, performs preprocessing on the images. Specifically, it adjusts the image size, removes noise, and corrects color. Open-source image processing libraries are often used at this stage. Once preprocessing is complete, the images are prepared as data suitable for analysis by machine learning algorithms.
[0079] Next, the server runs a convolutional neural network (CNN) to extract facial and oral feature points. This identifies the location and shape of the eyes, nose, mouth, and teeth. Based on this data, generative adversarial networks (GANs) and variational autoencoders (VAEs) are used to simulate future oral structures and facial shapes. Here, multiple predictive patterns are generated for each treatment scenario.
[0080] The server then uses a transformer model to search a database of previously accumulated treatments and identify similar cases. This similarity information is used to improve the accuracy of the current simulation results.
[0081] Furthermore, the generated simulation results are sent to the terminal and displayed to the user in an easy-to-understand interface. The terminal features an intuitive user interface and is designed to allow the user to select and compare treatment scenarios.
[0082] Finally, users can ask questions about the simulation results, and the device uses ChatGPT to answer them in real time. This interactive process allows users to receive detailed explanations and deepen their understanding of treatment options.
[0083] As a concrete example, a possible prompt message could be, "Please show me the changes in the face after partial orthodontic treatment." In response to this prompt, the system quickly generates a corresponding scenario and provides the user with clear information.
[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0085] Step 1:
[0086] Users upload facial and oral images to the system from their mobile devices or computers. The input data consists of facial and oral images, which are high-resolution digital images in JPEG or PNG format. These image data are then sent to the server as output. Users can easily select images taken with their device's camera and upload them according to the system's specified interface.
[0087] Step 2:
[0088] The server preprocesses the received images. The input data consists of images submitted by the user. Image processing software is used to perform noise reduction, color correction, and resizing. This process prepares the images for facial and oral feature extraction. The output is a preprocessed, clear image. This includes specific actions such as filtering and resizing pixel data using an image analysis library.
[0089] Step 3:
[0090] The server uses a convolutional neural network (CNN) to extract facial and oral feature points from pre-processed images. The input is pre-processed images. The CNN algorithm identifies the position and shape of the eyes, nose, mouth, and teeth. This feature extraction outputs detailed facial and oral data. By leveraging GPUs to accelerate calculations and efficiently execute large-scale models, the system performs the operation of accurately identifying feature points.
[0091] Step 4:
[0092] The server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future oral structures and facial shapes. The input is extracted feature data. This generates multiple future prediction results based on different orthodontic treatment scenarios. The output is a simulated image of the changes corresponding to each scenario. GANs and VAEs automatically generate various possibilities and create evolving images.
[0093] Step 5:
[0094] The server analyzes past treatment data using a transformer model to reference similar cases. Inputs include generated simulation results and a database of past treatments. The analysis outputs cases most similar to the current case, improving the reliability of the simulation. This includes using a deep learning model to detect patterns and similarities from a vast amount of case data.
[0095] Step 6:
[0096] The terminal displays the generated simulation results to the user. The input is simulation result data sent from the server. The output is in a format that allows for intuitive selection and comparison of each treatment scenario through an interactive user interface. The terminal utilizes high-speed browser rendering technology to provide smooth visual display.
[0097] Step 7:
[0098] When a user asks a question about the simulation results, the terminal uses ChatGPT to generate a response in real time. The input is the question prompt entered by the user. Based on this, an answer including detailed explanations and suggestions is output. Specifically, natural language processing techniques are used to understand the user's intent and provide an appropriate answer.
[0099] (Application Example 1)
[0100] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0101] Conventional autonomous vehicles do not adequately analyze the driver's gaze movements and posture changes in real time to support safe driving. Preventing unforeseen incidents caused by changes in gaze is difficult, and there is potential room for improvement in safety. This invention aims to solve these problems and provide a safer and more efficient driving assistance system.
[0102] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0103] In this invention, the server includes means for receiving a child's facial photograph and tooth image and preprocessing them according to predetermined criteria; means for extracting facial and tooth feature points from the preprocessed images using a convolutional neural network; means for simulating future tooth alignment and facial features using a generative adversarial network and a variational autoencoder based on the extracted features; means for acquiring the driver's gaze and posture data with a camera device and extracting visual features from the preprocessed images; and means for predicting future gaze movements based on the extracted visual features. This makes it possible to predict the driver's gaze movements and provide driver assistance information in advance.
[0104] A "child's facial photograph" is an electronically recorded and stored image of a child's face.
[0105] A "dental image" is a visual record of the condition of a child's teeth.
[0106] "Preprocessing means" refers to methods for adjusting the size, removing noise, and color-correcting the received image.
[0107] A "convolutional neural network" is a deep learning technique used to extract feature points from images.
[0108] "Feature points" are important data points related to the shape and position of the face and teeth.
[0109] A "generative adversarial network" is a machine learning technique used to mimic and generate data distributions.
[0110] A variational autoencoder is a variant of an autoencoder that estimates the latent variable space of data and enables the generation of new data.
[0111] "Simulation" is the process of modeling a real-world situation and predicting future situations based on that model.
[0112] "Eye movement prediction means" refers to a technology that analyzes the driver's eye direction and movement to predict future eye movements.
[0113] "Driver assistance information" refers to information provided to enable drivers to drive more safely and efficiently.
[0114] The system implementing this invention analyzes the driver's gaze and posture within an autonomous vehicle to support safe driving. The system's program includes functions to monitor the driver, predict their gaze movements in real time, and display appropriate driving support information.
[0115] The system uses a camera module mounted on the vehicle (for example, a Raspberry Pi Camera) to acquire frame data of the driver's face and eyes. The server receives this data and performs preprocessing such as resizing, noise reduction, and color correction. The server also uses a convolutional neural network (CNN) to extract important visual features. After feature extraction, it uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to predict the driver's gaze movements.
[0116] Based on predicted eye movements, the server displays appropriate driving assistance information on the vehicle's dashboard and the driver's smart glasses. For example, if the driver's gaze shifts while driving on a highway, the system immediately notifies the driver of the danger and prompts appropriate evasive action.
[0117] Furthermore, to analyze past driving data, the server utilizes a transformer model to search for similar driver reaction patterns and provide optimal information. The resulting information helps improve driver safety and efficiency.
[0118] An example of a prompt message is: "Design a system that analyzes the driver's gaze data while driving, predicts the direction that requires attention next, and warns the driver." This would allow drivers to operate vehicles more safely.
[0119] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0120] Step 1:
[0121] The server uses a camera module mounted on the vehicle to acquire image data of the driver's face and eyes. Based on this data, preprocessing is performed, including image resizing, noise reduction, and color correction, to enable accurate extraction of visual features. The input is raw image data, and the output is preprocessed, clean image data.
[0122] Step 2:
[0123] The server uses a convolutional neural network (CNN) to extract the driver's visual features from pre-processed image data. The input is the pre-processed image data, and the output consists of features related to the driver's face, gaze direction, and posture. These features are then further analyzed by edge AI software.
[0124] Step 3:
[0125] The server uses extracted features to predict the driver's future gaze patterns using generative adversarial networks (GANs) and variational autoencoders (VAEs). The input is the driver's visual features, and the output is predicted information about the direction the driver will next look. This allows drivers to know in advance which directions they should focus their attention while driving.
[0126] Step 4:
[0127] The server generates driver assistance information based on predicted eye movements and displays it on the vehicle's dashboard and the driver's smart glasses. The input is predicted eye movement information, and the output is navigation instructions and warning information to support safe driving.
[0128] Step 5:
[0129] The user adjusts their driving behavior based on driver assistance information provided by the system. In this process, the user makes judgments about the displayed information and takes appropriate driving actions. The input is the displayed driver assistance information, and the output is the user's driving actions.
[0130] Step 6:
[0131] The server uses data collected during driving and employs a transformer model to search for similar past driving patterns. The input is accumulated driving data, and the output is the analysis results of similar patterns. This enables the provision of support information optimized for each driver.
[0132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0133] The system according to the present invention combines image recognition technology, generative AI technology, and emotion analysis technology to precisely predict a child's future tooth alignment and facial features, and further optimizes and presents the results according to the user's emotions. Specific embodiments are described below.
[0134] First, the user uploads a photo of their child's face and an image of their teeth to the system. This allows the system to obtain data to simulate future changes.
[0135] Next, the server receives these images and performs preprocessing. This preprocessing includes adjusting the image resolution and removing noise, preparing the images for image analysis.
[0136] Next, the server processes the pre-processed images using a convolutional neural network (CNN) to extract feature points of the face and teeth. At this stage, detailed landmarks are identified, enabling highly accurate predictions using a generative model.
[0137] Based on the extracted features, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future tooth alignment and facial appearance scenarios. The simulation provides multiple scenarios, including cases where orthodontic treatment is not performed and cases where different treatment methods are adopted.
[0138] Next, the server references accumulated historical treatment data and analyzes similar cases. By using a transformer model, insights gained from this data are applied to the current case to improve predictive accuracy.
[0139] The simulation results are then sent to the terminal and presented to the user. Crucially, this involves the use of an emotion analysis engine. The terminal collects emotional data from the user's camera and voice input, and analyzes it in real time. Based on the obtained emotional data, the presentation of the simulation results is optimized. For example, if the user is feeling anxious, more detailed explanations or additional support information may be displayed.
[0140] Furthermore, if a user has questions based on their response to the simulation results, the terminal generates a response using an interactive interface (e.g., ChatGPT) to address the question appropriately. This allows the user to feel secure through the system.
[0141] For example, if a user asks, "I want to know the difference between full and partial orthodontic treatment," the server will generate scenarios for each, and the device will provide a customized explanation that reflects the sentiment analysis results. Furthermore, the accuracy of predictions will be improved by continuously incorporating feedback from experts into the model.
[0142] Based on the above, the present invention provides an environment in which users can understand and select the optimal orthodontic treatment method, and further provides a more personalized experience by adjusting the method of presenting results according to their emotions.
[0143] The following describes the processing flow.
[0144] Step 1:
[0145] Users select a photo of their child's face and an image of their teeth and upload them to the system's online platform. The upload process is performed from the user's device, and the images are sent to the system's server.
[0146] Step 2:
[0147] The server stores the received images and begins preprocessing. Specifically, the images are resized, compressed, denoised, and retouched to ensure a clean dataset that can be input into machine learning models.
[0148] Step 3:
[0149] The server feeds the preprocessed images into a convolutional neural network (CNN) to extract facial and dental feature points. This identifies features such as jaw shape, tooth arrangement, and landmarks for the eyes and mouth.
[0150] Step 4:
[0151] The server uses feature point data to simulate future tooth alignment and facial appearance scenarios using generative adversarial networks (GANs) and variational autoencoders (VAEs). Each scenario is generated based on whether or not orthodontic treatment is performed and different treatment methods.
[0152] Step 5:
[0153] The server uses accumulated historical treatment data to search for similar cases and performs analysis using a transformer model. The results are then incorporated to improve the accuracy of the simulation.
[0154] Step 6:
[0155] The results of the generated scenarios are sent to the terminal, which visually presents them to the user through the user interface. The presentation includes images of the future facial appearance and tooth alignment related to each orthodontic treatment scenario.
[0156] Step 7:
[0157] The device activates an emotion analysis engine and collects emotional information from the user's facial expressions and voice. The emotional data is analyzed in real time to determine whether the user is experiencing anxiety or confusion.
[0158] Step 8:
[0159] The device optimizes the display of simulation results based on the user's emotional state, as needed. For example, if anxiety is detected, additional explanations or encouraging messages are provided.
[0160] Step 9:
[0161] When a user asks a question about the simulation results, the device responds in real time using an interactive interface. The AI model used here provides detailed information in response to the question.
[0162] Step 10:
[0163] Finally, the server receives feedback from experts and incorporates the newly acquired experience into the learning model. This allows the system to provide even more accurate simulations.
[0164] (Example 2)
[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0166] Given the current difficulty in comparing multiple orthodontic treatment methods and making the optimal choice for a child's future teeth alignment and facial appearance, there is a need for a method that provides clear information while considering the user's emotional state, allowing them to make informed decisions with confidence. Furthermore, improving predictive accuracy and creating an environment where users can interactively obtain information without feeling anxious are key challenges.
[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0168] In this invention, the server includes means for receiving a child's facial photograph and dental images and pre-processing them according to predetermined criteria; means for extracting facial and dental feature points from the pre-processed images using a convolutional neural network; means for simulating future tooth alignment and facial appearance using a generative adversarial network and a variational autoencoder based on the extracted features; means for making predictions based on similar cases using accumulated past treatment data; and means for displaying the simulation results generated based on multiple orthodontic treatment scenarios on a communication terminal, and optimizing the method of presenting the results while taking into account the user's emotional state. As a result, the user can more easily understand the future simulation results and can select an appropriate orthodontic treatment method without feeling anxious.
[0169] "Preprocessing" refers to the process of adjusting the resolution and removing noise in order to improve the quality of received image data.
[0170] A "convolutional neural network" is an artificial intelligence technology used to analyze images and features, and is a model that has particularly high discriminatory capabilities for visual data.
[0171] "Feature point extraction" is the process of deriving important identifying information about the shape of faces and teeth from images, and it forms the basis of predictive simulations.
[0172] A "generative adversarial network" is a type of machine learning model used to simulate images and data. It is a technique that generates highly accurate output by having a generative model and a discriminative model compete alternately.
[0173] A variational autoencoder is a type of autoencoder that has the ability to learn the latent structure of data and generate new data, making it possible to efficiently reproduce the distribution of data.
[0174] "Simulation" is a process of virtual experimentation that uses a model to predict future situations and conditions, and is a method of predicting future tooth alignment and facial appearance through different treatment scenarios.
[0175] "Emotion analysis" is a technology that evaluates a user's emotional state in real time from information collected through their camera and voice data, and then analyzes that data.
[0176] An "interactive interface" is a system designed to facilitate natural conversations and responses with users, and its role is to provide appropriate information and answers to user input.
[0177] The system of this invention combines image processing technology and machine learning technology to predict future changes in a child's teeth alignment and facial features, and supports the user in selecting an appropriate treatment method.
[0178] First, the user uploads a photo of their child's face and teeth to the system. The server then retrieves the data necessary for the simulation. This image data is preprocessed, including resolution adjustment and noise reduction, to be converted into a format suitable for analysis. Dedicated image editing software and algorithms are used for this preprocessing.
[0179] Next, the server uses a convolutional neural network (CNN) to extract facial and tooth feature points from the pre-processed images. This identifies landmarks of facial features and tooth shape. Based on these features, a generative adversarial network (GAN) and a variational autoencoder (VAE) are used to simulate future tooth alignment and facial appearance scenarios.
[0180] The system further analyzes past treatment data using a transformer model, compares it with similar cases, and improves the accuracy of predictions. As a result, multiple scenarios based on different orthodontic treatment options are generated, showing the expected outcomes associated with each treatment method.
[0181] These simulation results are presented to the user via a terminal. The terminal analyzes the user's emotional state in real time through the user's camera and voice input, and displays information in the most appropriate format for the user's state. For example, if the user is feeling anxious, the system will provide additional explanations or information to aid understanding.
[0182] If a user has questions based on the information presented, the device utilizes an interactive interface and uses natural language processing to respond to the user's questions. This includes using prompts such as "Please explain in detail the difference between full and partial orthodontic treatment" using a generative AI model.
[0183] In this way, the system provides users with clear and easy-to-understand information, creating an environment where they can confidently choose the optimal treatment method.
[0184] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0185] Step 1:
[0186] The user uploads a photo of their child's face and an image of their teeth to the system. By sending these images to the server as input data, the server obtains the basic data needed to simulate future changes. The specific action involved is for the user to select image files using a terminal and send them to the server via the upload function.
[0187] Step 2:
[0188] The server preprocesses the received image data. First, it applies resolution adjustment to the original input image, and then applies a noise reduction algorithm. This results in clear image data suitable for analysis as output. In this process, image editing software and specific algorithms are used to prepare the data to an optimal state.
[0189] Step 3:
[0190] The server inputs preprocessed images into a convolutional neural network (CNN) to extract feature points of the face and teeth. The CNN performs pixel-level analysis of the preprocessed input images to identify landmarks that indicate facial features and tooth shape. The output is the feature point data necessary for the simulation. This process utilizes a highly accurate model.
[0191] Step 4:
[0192] The server uses a generative adversarial network (GAN) and a variational autoencoder (VAE) to simulate future tooth alignment and facial features based on feature data. It constructs multiple scenarios using the feature data as input, making predictions that consider the impact of different orthodontic treatment methods. The final output is a visual future scenario for each treatment.
[0193] Step 5:
[0194] The server analyzes the simulation results using a transformer model, comparing them with accumulated historical treatment data. At this stage, historical data is input, and similar cases are analyzed to improve the accuracy of the simulation results and derive the optimal solution.
[0195] Step 6:
[0196] The terminal presents the final simulation results to the user. At this time, the user's emotional data, collected through a camera or voice input device, is used as input. The terminal performs real-time emotional analysis and optimizes the presentation method of the results according to the user's emotional state. The output is provided in a format that is easy for the user to understand.
[0197] Step 7:
[0198] If the user has questions based on the simulation results, the terminal responds with an interactive interface. It processes user input as prompts, uses a generative AI model to create natural language answers, and presents them to the user. Detailed explanations and additional information are provided as output to aid user understanding.
[0199] (Application Example 2)
[0200] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0201] The objective of this invention is to create an environment in which users can make appropriate decisions with confidence by accurately predicting future changes in the face and oral cavity, enabling the comparison of different orthodontic treatment scenarios, and providing optimal information presentation tailored to the user's emotional state.
[0202] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0203] In this invention, the server includes means for receiving images of the face and oral cavity and preprocessing them according to predetermined criteria, means for extracting facial and oral cavity features from the preprocessed images using deep learning technology, and means for predicting future oral and facial features using a generative model based on the extracted features. This enables highly accurate predictions and the presentation of information that responds to the user's emotional state.
[0204] A "facial image" is visual data that includes the entire face of the subject, and it forms the basis for feature extraction.
[0205] An "oral image" is visual data that captures the inside of a subject's mouth in detail, and is used to analyze the condition of their teeth and gums.
[0206] "Preprocessing" refers to processes that include adjusting the resolution of image data and removing noise, and is preparatory work to facilitate subsequent data analysis.
[0207] "Deep learning technology" is an information processing technology that uses multi-layered artificial neural networks to achieve advanced pattern recognition and feature extraction.
[0208] A "generative model" is a technique that generates new data using methods such as generative adversarial networks and variational autoencoders, and is used for prediction and simulation.
[0209] An "orthodontic treatment scenario" is a future treatment plan that assumes different orthodontic methods and treatment flows, and serves as a reference for users to make their choices.
[0210] "Emotional analysis technology" is a technology that detects and analyzes a user's emotional state from images and audio, and is used to provide personalized information.
[0211] The system for carrying out this invention uses a smart device to collect images of the face and oral cavity and transmit them to a cloud server. When a user uploads images through the device, the server preprocesses the received images. This preprocessing adjusts the image resolution and removes noise to prepare the images for analysis.
[0212] Next, the server utilizes deep learning techniques to extract facial and oral features from the preprocessed images. Specifically, it uses a convolutional neural network (CNN) to identify facial and oral feature points. This data then forms the basis for future predictions by subsequent generative models.
[0213] Based on the extracted features, the server predicts the future appearance of the oral cavity and face using generative adversarial networks (GANs) and variational autoencoders (VAEs). This makes it possible to provide users with data that allows them to compare and consider different orthodontic treatment scenarios.
[0214] Furthermore, the server references accumulated past treatment data and analyzes similar cases. By using a transformer model, past knowledge is reflected in current predictions, improving accuracy.
[0215] Furthermore, it utilizes emotion analysis technology to detect the user's emotional state in real time. Based on the obtained emotional data, the system optimizes how prediction results are presented, providing the user with the most reassuring information. For example, if the user is feeling anxious, it will provide more detailed explanations or additional support information.
[0216] For example, if a user asks a question such as, "I want to know the difference between full or partial orthodontic treatment," the system will generate scenarios for each, and the device will provide a customized explanation that reflects the sentiment analysis results.
[0217] An example of a prompt message for a generative AI model would be: "I have uploaded a photo. Please predict future changes in teeth alignment and facial features and create a scenario for orthodontic treatment. Also, since the user appears anxious, please include detailed explanations and reassuring information."
[0218] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0219] Step 1:
[0220] The user takes images of their face and mouth using a device and uploads them to the server. The input here is the captured image data, and the output is the transmission of the image data to the server.
[0221] Step 2:
[0222] The server performs preprocessing on the received image data. The input is uploaded facial and oral cavity image data, and the server processes the data, such as adjusting the resolution and removing noise, to output image data that is ready for analysis.
[0223] Step 3:
[0224] The server applies deep learning techniques to preprocessed image data to extract facial and oral features. Specifically, it uses a convolutional neural network (CNN) to extract feature points. The input is preprocessed image data, and the output is the extracted feature data.
[0225] Step 4:
[0226] The server uses a generative model to predict future oral and facial features based on extracted feature data. Generative adversarial networks (GANs) and variational autoencoders (VAEs) are used here. The input is feature data, and the server generates and outputs simulation data for future predictions.
[0227] Step 5:
[0228] The server references past treatment data and uses a transformer model to analyze similar cases. This provides insights to improve prediction accuracy. The input is past treatment data, and the output is enhanced prediction data.
[0229] Step 6:
[0230] The server uses emotion analysis technology to analyze the user's emotional state in real time. The input is data on the user's current emotional state (obtained from facial and voice), and the server outputs data to optimize the presentation method based on the emotional information.
[0231] Step 7:
[0232] The terminal displays predictive data received from the server through a user interface. Here, the information is presented using optimized data tailored to the user's emotional state. Input consists of predictive data and sentiment analysis data, and the terminal outputs customized information.
[0233] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0234] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0235] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0236] [Second Embodiment]
[0237] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0238] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0239] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0240] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0241] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0242] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0243] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0244] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0245] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0246] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0247] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0248] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0249] The system according to the present invention combines image recognition technology and generative AI technology to predict a child's future tooth alignment and facial features. It is implemented through the interaction of the user, server, and terminal as follows.
[0250] First, users upload a photo of their child's face and images of their teeth to the system from a mobile device or computer. These images are used as foundational data to simulate future conditions.
[0251] Next, the server receives the uploaded images and performs a series of preprocessing steps. These preprocessing steps include resizing, denoising, and color correction. This prepares the images for accurate extraction of facial and dental features.
[0252] Next, the server uses a convolutional neural network (CNN) to extract key feature points of the face and teeth from the image. At this stage, for example, the position and shape of the eyes, nose, mouth, and teeth are analyzed and used as foundational data for predicting future changes.
[0253] Based on the extracted feature data, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate multiple future scenarios depending on whether orthodontic treatment is performed and what method is used. These include various cases such as when no orthodontic treatment is performed and when a specific treatment method is used.
[0254] Next, the server retrieves information for analyzing past cases, particularly those that are similar. In this process, past treatment outcome data is searched using a transformer model to identify similarities with the current case.
[0255] The generated simulation results are then sent to the terminal. The terminal visually displays these images and information to the user. This display is organized in a format that allows the user to intuitively understand and compare treatment options.
[0256] Finally, when a user asks a question based on the simulation results, the device uses ChatGPT to provide a real-time explanation tailored to the question, aiding the user's understanding. For example, if a user requests to see "how the face would change after partial orthodontic treatment," that scenario is immediately generated, and detailed images and explanations are provided.
[0257] Based on the above, the system of the present invention provides users with detailed predictive data to select the optimal orthodontic treatment method, and continuously improves its accuracy through the use of past data and expert feedback.
[0258] The following describes the processing flow.
[0259] Step 1:
[0260] Users upload photos of their child's face and teeth to the system. The image files must be submitted in a specified format, so users should prepare them according to the guidelines.
[0261] Step 2:
[0262] The server receives the uploaded images and performs preprocessing on them. Specifically, this includes resizing, denoising, and color correction. This ensures the quality necessary for analysis.
[0263] Step 3:
[0264] The server feeds the preprocessed images to a convolutional neural network (CNN) to extract facial and tooth feature points. This feature extraction process identifies facial landmarks and tooth contours.
[0265] Step 4:
[0266] Based on the extracted features, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future tooth alignment and facial features. Multiple treatment scenarios are included, such as no orthodontics, partial orthodontics, and full orthodontics.
[0267] Step 5:
[0268] The server references past treatment data and searches for similar cases. It uses a transformer model to analyze the database and find the treatment pattern that best matches the current case.
[0269] Step 6:
[0270] The generated simulation results are sent to the terminal, which displays them visually to the user. The display includes detailed images showing the future facial features and teeth alignment, along with explanations of treatment options.
[0271] Step 7:
[0272] When a user asks a question about the simulation results, the terminal uses ChatGPT to generate a response in real time, providing detailed explanations and additional information.
[0273] (Example 1)
[0274] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0275] It is difficult for parents and medical professionals to obtain accurate early predictions about a child's future oral structure and facial shape. In particular, information is limited to determine which treatment options are most effective when corrections or adjustments are needed. This increases the risk of incorrect treatment choices and unnecessary procedures. Furthermore, the lack of readily available methods to obtain predictions based on past cases makes it difficult to select the optimal treatment plan for each individual child.
[0276] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0277] In this invention, the server includes means for receiving human facial images and oral images and performing preprocessing on them based on pre-set criteria, means for extracting facial and oral feature points from the preprocessed images using a machine learning algorithm, and means for simulating future oral structures and facial shapes using an advanced generation algorithm based on the extracted features. As a result, parents and medical experts can easily obtain predictions based on accurate and diverse treatment scenarios.
[0278] The "human facial image" is a digital image that visually records the features of an individual's face and includes information representing the structure and shape of the face.
[0279] The "oral image" is an image that captures the state of the human oral cavity, particularly the dentition and gums, and is used for treatment and diagnosis.
[0280] "Preprocessing" is a process of converting image data into a form that is easy for subsequent analysis, and generally refers to all processes including size adjustment, noise removal, color correction, etc.
[0281] The "machine learning algorithm" is a technology that learns patterns and features from a large amount of data and makes predictions and judgments, and is applied to image analysis and feature extraction.
[0282] "Extraction of feature points" is a process of identifying and extracting points that represent important positions and shapes in an image, and is used as basic data for analysis and simulation.
[0283] The "advanced generation algorithm" is a complex calculation method for generating new information from data, and is particularly used when predicting future states.
[0284] The "oral structure" refers to the specific shapes and arrangements within the oral cavity, such as the dentition, jawbone, oral mucosa, etc.
[0285] "Facial shape" refers to the specific form and contour of the external structure of a human face, representing individual characteristics.
[0286] "Similar cases" refer to cases that existed in the past and have many common points with the case currently being analyzed.
[0287] "Medical information" refers to information including treatment, diagnosis results, and historical health data related to patients, which is used for determining treatment policies and predictions.
[0288] "Diverse orthodontic treatment methods" refer to various treatment means that can be selected in orthodontics, including, as specific examples, bracket correction and mouthpiece correction.
[0289] "Device for interaction with humans" refers to devices or interfaces for exchanging information with human users, and is for providing information and answering questions through interaction.
[0290] The embodiments for implementing the present invention will be described. This system is one in which a user, a server, and a terminal interact with each other to predict the future oral structure and facial shape of a child. The respective roles will be described in detail below.
[0291] First, the user uploads facial images and oral images of the child to the system using a mobile device or a computer. Devices used for this include general smartphones with cameras and personal computers. These images serve as the base data for the simulation.
[0292] Next, these images are sent to the server. The server is equipped with a high-performance computer and image processing software, and performs preprocessing of the images. Specifically, it adjusts the size of the images, removes noise, and performs color correction. In this stage, open-source image processing libraries are often used. The images after preprocessing are prepared as data suitable for analysis by machine learning algorithms.
[0293] Next, the server runs a convolutional neural network (CNN) to extract facial and oral feature points. This identifies the location and shape of the eyes, nose, mouth, and teeth. Based on this data, generative adversarial networks (GANs) and variational autoencoders (VAEs) are used to simulate future oral structures and facial shapes. Here, multiple predictive patterns are generated for each treatment scenario.
[0294] The server then uses a transformer model to search a database of previously accumulated treatments and identify similar cases. This similarity information is used to improve the accuracy of the current simulation results.
[0295] Furthermore, the generated simulation results are sent to the terminal and displayed to the user in an easy-to-understand interface. The terminal features an intuitive user interface and is designed to allow the user to select and compare treatment scenarios.
[0296] Finally, users can ask questions about the simulation results, and the device uses ChatGPT to answer them in real time. This interactive process allows users to receive detailed explanations and deepen their understanding of treatment options.
[0297] As a concrete example, a possible prompt message could be, "Please show me the changes in the face after partial orthodontic treatment." In response to this prompt, the system quickly generates a corresponding scenario and provides the user with clear information.
[0298] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0299] Step 1:
[0300] Users upload facial and oral images to the system from their mobile devices or computers. The input data consists of facial and oral images, which are high-resolution digital images in JPEG or PNG format. These image data are then sent to the server as output. Users can easily select images taken with their device's camera and upload them according to the system's specified interface.
[0301] Step 2:
[0302] The server preprocesses the received images. The input data consists of images submitted by the user. Image processing software is used to perform noise reduction, color correction, and resizing. This process prepares the images for facial and oral feature extraction. The output is a preprocessed, clear image. This includes specific actions such as filtering and resizing pixel data using an image analysis library.
[0303] Step 3:
[0304] The server uses a convolutional neural network (CNN) to extract facial and oral feature points from pre-processed images. The input is pre-processed images. The CNN algorithm identifies the position and shape of the eyes, nose, mouth, and teeth. This feature extraction outputs detailed facial and oral data. By leveraging GPUs to accelerate calculations and efficiently execute large-scale models, the system performs the operation of accurately identifying feature points.
[0305] Step 4:
[0306] The server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future oral structures and facial shapes. The input is extracted feature data. This generates multiple future prediction results based on different orthodontic treatment scenarios. The output is a simulated image of the changes corresponding to each scenario. GANs and VAEs automatically generate various possibilities and create evolving images.
[0307] Step 5:
[0308] The server analyzes past treatment data using a transformer model to reference similar cases. The inputs are the generated simulation results and the past treatment database. Through the analysis, cases most similar to the current case are output, improving the reliability of the simulation. This includes the operation of using a deep learning model to detect patterns and similarities from a large amount of case data.
[0309] Step 6:
[0310] The terminal displays the generated simulation results to the user. The input is the simulation result data sent from the server. Through an interactive user interface, each treatment scenario can be intuitively selected and output in a comparable format. The terminal utilizes high-speed browser rendering technology to implement smooth visual display.
[0311] Step 7:
[0312] When the user asks questions about the simulation results, the terminal uses ChatGPT to generate responses in real time. The input is the question prompt entered by the user. Based on this, answers including detailed explanations and suggestions are output. As a specific operation, natural language processing technology is used to understand the user's intention and provide appropriate answers.
[0313] (Application Example 1)
[0314] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0315] Conventional autonomous vehicles do not adequately analyze the driver's gaze movements and posture changes in real time to support safe driving. Preventing unforeseen incidents caused by changes in gaze is difficult, and there is potential room for improvement in safety. This invention aims to solve these problems and provide a safer and more efficient driving assistance system.
[0316] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0317] In this invention, the server includes means for receiving a child's facial photograph and tooth image and preprocessing them according to predetermined criteria; means for extracting facial and tooth feature points from the preprocessed images using a convolutional neural network; means for simulating future tooth alignment and facial features using a generative adversarial network and a variational autoencoder based on the extracted features; means for acquiring the driver's gaze and posture data with a camera device and extracting visual features from the preprocessed images; and means for predicting future gaze movements based on the extracted visual features. This makes it possible to predict the driver's gaze movements and provide driver assistance information in advance.
[0318] A "child's facial photograph" is an electronically recorded and stored image of a child's face.
[0319] A "dental image" is a visual record of the condition of a child's teeth.
[0320] "Preprocessing means" refers to methods for adjusting the size, removing noise, and color-correcting the received image.
[0321] A "convolutional neural network" is a deep learning technique used to extract feature points from images.
[0322] "Feature points" are important data points related to the shape and position of the face and teeth.
[0323] A "generative adversarial network" is a machine learning technique used to mimic and generate data distributions.
[0324] A variational autoencoder is a variant of an autoencoder that estimates the latent variable space of data and enables the generation of new data.
[0325] "Simulation" is the process of modeling a real-world situation and predicting future situations based on that model.
[0326] "Eye movement prediction means" refers to a technology that analyzes the driver's eye direction and movement to predict future eye movements.
[0327] "Driver assistance information" refers to information provided to enable drivers to drive more safely and efficiently.
[0328] The system implementing this invention analyzes the driver's gaze and posture within an autonomous vehicle to support safe driving. The system's program includes functions to monitor the driver, predict their gaze movements in real time, and display appropriate driving support information.
[0329] The system uses a camera module mounted on the vehicle (for example, a Raspberry Pi Camera) to acquire frame data of the driver's face and eyes. The server receives this data and performs preprocessing such as resizing, noise reduction, and color correction. The server also uses a convolutional neural network (CNN) to extract important visual features. After feature extraction, it uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to predict the driver's gaze movements.
[0330] Based on predicted eye movements, the server displays appropriate driving assistance information on the vehicle's dashboard and the driver's smart glasses. For example, if the driver's gaze shifts while driving on a highway, the system immediately notifies the driver of the danger and prompts appropriate evasive action.
[0331] Furthermore, to analyze past driving data, the server utilizes a transformer model to search for similar driver reaction patterns and provide optimal information. The resulting information helps improve driver safety and efficiency.
[0332] An example of a prompt message is: "Design a system that analyzes the driver's gaze data while driving, predicts the direction that requires attention next, and warns the driver." This would allow drivers to operate vehicles more safely.
[0333] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0334] Step 1:
[0335] The server uses a camera module mounted on the vehicle to acquire image data of the driver's face and eyes. Based on this data, preprocessing is performed, including image resizing, noise reduction, and color correction, to enable accurate extraction of visual features. The input is raw image data, and the output is preprocessed, clean image data.
[0336] Step 2:
[0337] The server uses a convolutional neural network (CNN) to extract the driver's visual features from pre-processed image data. The input is the pre-processed image data, and the output consists of features related to the driver's face, gaze direction, and posture. These features are then further analyzed by edge AI software.
[0338] Step 3:
[0339] The server uses extracted features to predict the driver's future gaze patterns using generative adversarial networks (GANs) and variational autoencoders (VAEs). The input is the driver's visual features, and the output is predicted information about the direction the driver will next look. This allows drivers to know in advance which directions they should focus their attention while driving.
[0340] Step 4:
[0341] The server generates driver assistance information based on predicted eye movements and displays it on the vehicle's dashboard and the driver's smart glasses. The input is predicted eye movement information, and the output is navigation instructions and warning information to support safe driving.
[0342] Step 5:
[0343] The user adjusts their driving behavior based on driver assistance information provided by the system. In this process, the user makes judgments about the displayed information and takes appropriate driving actions. The input is the displayed driver assistance information, and the output is the user's driving actions.
[0344] Step 6:
[0345] The server uses data collected during driving and employs a transformer model to search for similar past driving patterns. The input is accumulated driving data, and the output is the analysis results of similar patterns. This enables the provision of support information optimized for each driver.
[0346] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0347] The system according to the present invention combines image recognition technology, generative AI technology, and emotion analysis technology to precisely predict a child's future tooth alignment and facial features, and further optimizes and presents the results according to the user's emotions. Specific embodiments are described below.
[0348] First, the user uploads a photo of their child's face and an image of their teeth to the system. This allows the system to obtain data to simulate future changes.
[0349] Next, the server receives these images and performs preprocessing. This preprocessing includes adjusting the image resolution and removing noise, preparing the images for image analysis.
[0350] Next, the server processes the pre-processed images using a convolutional neural network (CNN) to extract feature points of the face and teeth. At this stage, detailed landmarks are identified, enabling highly accurate predictions using a generative model.
[0351] Based on the extracted features, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future tooth alignment and facial appearance scenarios. The simulation provides multiple scenarios, including cases where orthodontic treatment is not performed and cases where different treatment methods are adopted.
[0352] Next, the server references accumulated historical treatment data and analyzes similar cases. By using a transformer model, insights gained from this data are applied to the current case to improve predictive accuracy.
[0353] The simulation results are then sent to the terminal and presented to the user. Crucially, this involves the use of an emotion analysis engine. The terminal collects emotional data from the user's camera and voice input, and analyzes it in real time. Based on the obtained emotional data, the presentation of the simulation results is optimized. For example, if the user is feeling anxious, more detailed explanations or additional support information may be displayed.
[0354] Furthermore, if a user has questions based on their response to the simulation results, the terminal generates a response using an interactive interface (e.g., ChatGPT) to address the question appropriately. This allows the user to feel secure through the system.
[0355] For example, if a user asks, "I want to know the difference between full and partial orthodontic treatment," the server will generate scenarios for each, and the device will provide a customized explanation that reflects the sentiment analysis results. Furthermore, the accuracy of predictions will be improved by continuously incorporating feedback from experts into the model.
[0356] Based on the above, the present invention provides an environment in which users can understand and select the optimal orthodontic treatment method, and further provides a more personalized experience by adjusting the method of presenting results according to their emotions.
[0357] The following describes the processing flow.
[0358] Step 1:
[0359] Users select a photo of their child's face and an image of their teeth and upload them to the system's online platform. The upload process is performed from the user's device, and the images are sent to the system's server.
[0360] Step 2:
[0361] The server stores the received images and begins preprocessing. Specifically, the images are resized, compressed, denoised, and retouched to ensure a clean dataset that can be input into machine learning models.
[0362] Step 3:
[0363] The server feeds the preprocessed images into a convolutional neural network (CNN) to extract facial and dental feature points. This identifies features such as jaw shape, tooth arrangement, and landmarks for the eyes and mouth.
[0364] Step 4:
[0365] The server uses feature point data to simulate future tooth alignment and facial appearance scenarios using generative adversarial networks (GANs) and variational autoencoders (VAEs). Each scenario is generated based on whether or not orthodontic treatment is performed and different treatment methods.
[0366] Step 5:
[0367] The server uses accumulated historical treatment data to search for similar cases and performs analysis using a transformer model. The results are then incorporated to improve the accuracy of the simulation.
[0368] Step 6:
[0369] The results of the generated scenarios are sent to the terminal, which visually presents them to the user through the user interface. The presentation includes images of the future facial appearance and tooth alignment related to each orthodontic treatment scenario.
[0370] Step 7:
[0371] The device activates an emotion analysis engine and collects emotional information from the user's facial expressions and voice. The emotional data is analyzed in real time to determine whether the user is experiencing anxiety or confusion.
[0372] Step 8:
[0373] The device optimizes the display of simulation results based on the user's emotional state, as needed. For example, if anxiety is detected, additional explanations or encouraging messages are provided.
[0374] Step 9:
[0375] When a user asks a question about the simulation results, the device responds in real time using an interactive interface. The AI model used here provides detailed information in response to the question.
[0376] Step 10:
[0377] Finally, the server receives feedback from experts and incorporates the newly acquired experience into the learning model. This allows the system to provide even more accurate simulations.
[0378] (Example 2)
[0379] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0380] Given the current difficulty in comparing multiple orthodontic treatment methods and making the optimal choice for a child's future teeth alignment and facial appearance, there is a need for a method that provides clear information while considering the user's emotional state, allowing them to make informed decisions with confidence. Furthermore, improving predictive accuracy and creating an environment where users can interactively obtain information without feeling anxious are key challenges.
[0381] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0382] In this invention, the server includes means for receiving a child's facial photograph and dental images and pre-processing them according to predetermined criteria; means for extracting facial and dental feature points from the pre-processed images using a convolutional neural network; means for simulating future tooth alignment and facial appearance using a generative adversarial network and a variational autoencoder based on the extracted features; means for making predictions based on similar cases using accumulated past treatment data; and means for displaying the simulation results generated based on multiple orthodontic treatment scenarios on a communication terminal, and optimizing the method of presenting the results while taking into account the user's emotional state. As a result, the user can more easily understand the future simulation results and can select an appropriate orthodontic treatment method without feeling anxious.
[0383] "Preprocessing" refers to the process of adjusting the resolution and removing noise in order to improve the quality of received image data.
[0384] A "convolutional neural network" is an artificial intelligence technology used to analyze images and features, and is a model that has particularly high discriminatory capabilities for visual data.
[0385] "Feature point extraction" is the process of deriving important identifying information about the shape of faces and teeth from images, and it forms the basis of predictive simulations.
[0386] A "generative adversarial network" is a type of machine learning model used to simulate images and data. It is a technique that generates highly accurate output by having a generative model and a discriminative model compete alternately.
[0387] A variational autoencoder is a type of autoencoder that has the ability to learn the latent structure of data and generate new data, making it possible to efficiently reproduce the distribution of data.
[0388] "Simulation" is a process of virtual experimentation that uses a model to predict future situations and conditions, and is a method of predicting future tooth alignment and facial appearance through different treatment scenarios.
[0389] "Emotion analysis" is a technology that evaluates a user's emotional state in real time from information collected through their camera and voice data, and then analyzes that data.
[0390] An "interactive interface" is a system designed to facilitate natural conversations and responses with users, and its role is to provide appropriate information and answers to user input.
[0391] The system of this invention combines image processing technology and machine learning technology to predict future changes in a child's teeth alignment and facial features, and supports the user in selecting an appropriate treatment method.
[0392] First, the user uploads a photo of their child's face and teeth to the system. The server then retrieves the data necessary for the simulation. This image data is preprocessed, including resolution adjustment and noise reduction, to be converted into a format suitable for analysis. Dedicated image editing software and algorithms are used for this preprocessing.
[0393] Next, the server uses a convolutional neural network (CNN) to extract facial and tooth feature points from the pre-processed images. This identifies landmarks of facial features and tooth shape. Based on these features, a generative adversarial network (GAN) and a variational autoencoder (VAE) are used to simulate future tooth alignment and facial appearance scenarios.
[0394] The system further analyzes past treatment data using a transformer model, compares it with similar cases, and improves the accuracy of predictions. As a result, multiple scenarios based on different orthodontic treatment options are generated, showing the expected outcomes associated with each treatment method.
[0395] These simulation results are presented to the user via a terminal. The terminal analyzes the user's emotional state in real time through the user's camera and voice input, and displays information in the most appropriate format for the user's state. For example, if the user is feeling anxious, the system will provide additional explanations or information to aid understanding.
[0396] If a user has questions based on the information presented, the device utilizes an interactive interface and uses natural language processing to respond to the user's questions. This includes using prompts such as "Please explain in detail the difference between full and partial orthodontic treatment" using a generative AI model.
[0397] In this way, the system provides users with clear and easy-to-understand information, creating an environment where they can confidently choose the optimal treatment method.
[0398] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0399] Step 1:
[0400] The user uploads a photo of their child's face and an image of their teeth to the system. By sending these images to the server as input data, the server obtains the basic data needed to simulate future changes. The specific action involved is for the user to select image files using a terminal and send them to the server via the upload function.
[0401] Step 2:
[0402] The server preprocesses the received image data. First, it applies resolution adjustment to the original input image, and then applies a noise reduction algorithm. This results in clear image data suitable for analysis as output. In this process, image editing software and specific algorithms are used to prepare the data to an optimal state.
[0403] Step 3:
[0404] The server inputs preprocessed images into a convolutional neural network (CNN) to extract feature points of the face and teeth. The CNN performs pixel-level analysis of the preprocessed input images to identify landmarks that indicate facial features and tooth shape. The output is the feature point data necessary for the simulation. This process utilizes a highly accurate model.
[0405] Step 4:
[0406] The server uses a generative adversarial network (GAN) and a variational autoencoder (VAE) to simulate future tooth alignment and facial features based on feature data. It constructs multiple scenarios using the feature data as input, making predictions that consider the impact of different orthodontic treatment methods. The final output is a visual future scenario for each treatment.
[0407] Step 5:
[0408] The server analyzes the simulation results using a transformer model, comparing them with accumulated historical treatment data. At this stage, historical data is input, and similar cases are analyzed to improve the accuracy of the simulation results and derive the optimal solution.
[0409] Step 6:
[0410] The terminal presents the final simulation results to the user. At this time, the user's emotional data, collected through a camera or voice input device, is used as input. The terminal performs real-time emotional analysis and optimizes the presentation method of the results according to the user's emotional state. The output is provided in a format that is easy for the user to understand.
[0411] Step 7:
[0412] If the user has questions based on the simulation results, the terminal responds with an interactive interface. It processes user input as prompts, uses a generative AI model to create natural language answers, and presents them to the user. Detailed explanations and additional information are provided as output to aid user understanding.
[0413] (Application Example 2)
[0414] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0415] The objective of this invention is to create an environment in which users can make appropriate decisions with confidence by accurately predicting future changes in the face and oral cavity, enabling the comparison of different orthodontic treatment scenarios, and providing optimal information presentation tailored to the user's emotional state.
[0416] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0417] In this invention, the server includes means for receiving images of the face and oral cavity and preprocessing them according to predetermined criteria, means for extracting facial and oral cavity features from the preprocessed images using deep learning technology, and means for predicting future oral and facial features using a generative model based on the extracted features. This enables highly accurate predictions and the presentation of information that responds to the user's emotional state.
[0418] A "facial image" is visual data that includes the entire face of the subject, and it forms the basis for feature extraction.
[0419] An "oral image" is visual data that captures the inside of a subject's mouth in detail, and is used to analyze the condition of their teeth and gums.
[0420] "Preprocessing" refers to processes that include adjusting the resolution of image data and removing noise, and is preparatory work to facilitate subsequent data analysis.
[0421] "Deep learning technology" is an information processing technology that uses multi-layered artificial neural networks to achieve advanced pattern recognition and feature extraction.
[0422] A "generative model" is a technique that generates new data using methods such as generative adversarial networks and variational autoencoders, and is used for prediction and simulation.
[0423] An "orthodontic treatment scenario" is a future treatment plan that assumes different orthodontic methods and treatment flows, and serves as a reference for users to make their choices.
[0424] "Emotional analysis technology" is a technology that detects and analyzes a user's emotional state from images and audio, and is used to provide personalized information.
[0425] The system for carrying out this invention uses a smart device to collect images of the face and oral cavity and transmit them to a cloud server. When a user uploads images through the device, the server preprocesses the received images. This preprocessing adjusts the image resolution and removes noise to prepare the images for analysis.
[0426] Next, the server utilizes deep learning techniques to extract facial and oral features from the preprocessed images. Specifically, it uses a convolutional neural network (CNN) to identify facial and oral feature points. This data then forms the basis for future predictions by subsequent generative models.
[0427] Based on the extracted features, the server predicts the future appearance of the oral cavity and face using generative adversarial networks (GANs) and variational autoencoders (VAEs). This makes it possible to provide users with data that allows them to compare and consider different orthodontic treatment scenarios.
[0428] Furthermore, the server references accumulated past treatment data and analyzes similar cases. By using a transformer model, past knowledge is reflected in current predictions, improving accuracy.
[0429] Furthermore, it utilizes emotion analysis technology to detect the user's emotional state in real time. Based on the obtained emotional data, the system optimizes how prediction results are presented, providing the user with the most reassuring information. For example, if the user is feeling anxious, it will provide more detailed explanations or additional support information.
[0430] For example, if a user asks a question such as, "I want to know the difference between full or partial orthodontic treatment," the system will generate scenarios for each, and the device will provide a customized explanation that reflects the sentiment analysis results.
[0431] An example of a prompt message for a generative AI model would be: "I have uploaded a photo. Please predict future changes in teeth alignment and facial features and create a scenario for orthodontic treatment. Also, since the user appears anxious, please include detailed explanations and reassuring information."
[0432] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0433] Step 1:
[0434] The user takes images of their face and mouth using a device and uploads them to the server. The input here is the captured image data, and the output is the transmission of the image data to the server.
[0435] Step 2:
[0436] The server performs preprocessing on the received image data. The input is uploaded facial and oral cavity image data, and the server processes the data, such as adjusting the resolution and removing noise, to output image data that is ready for analysis.
[0437] Step 3:
[0438] The server applies deep learning techniques to preprocessed image data to extract facial and oral features. Specifically, it uses a convolutional neural network (CNN) to extract feature points. The input is preprocessed image data, and the output is the extracted feature data.
[0439] Step 4:
[0440] The server uses a generative model to predict future oral and facial features based on extracted feature data. Generative adversarial networks (GANs) and variational autoencoders (VAEs) are used here. The input is feature data, and the server generates and outputs simulation data for future predictions.
[0441] Step 5:
[0442] The server references past treatment data and uses a transformer model to analyze similar cases. This provides insights to improve prediction accuracy. The input is past treatment data, and the output is enhanced prediction data.
[0443] Step 6:
[0444] The server uses emotion analysis technology to analyze the user's emotional state in real time. The input is data on the user's current emotional state (obtained from facial and voice), and the server outputs data to optimize the presentation method based on the emotional information.
[0445] Step 7:
[0446] The terminal displays predictive data received from the server through a user interface. Here, the information is presented using optimized data tailored to the user's emotional state. Input consists of predictive data and sentiment analysis data, and the terminal outputs customized information.
[0447] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0448] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0449] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0450] [Third Embodiment]
[0451] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0452] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0453] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0454] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0455] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0456] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0457] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0458] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0459] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0460] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0461] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0462] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0463] The system according to the present invention combines image recognition technology and generative AI technology to predict a child's future tooth alignment and facial features. It is implemented through the interaction of the user, server, and terminal as follows.
[0464] First, users upload a photo of their child's face and images of their teeth to the system from a mobile device or computer. These images are used as foundational data to simulate future conditions.
[0465] Next, the server receives the uploaded images and performs a series of preprocessing steps. These preprocessing steps include resizing, denoising, and color correction. This prepares the images for accurate extraction of facial and dental features.
[0466] Next, the server uses a convolutional neural network (CNN) to extract key feature points of the face and teeth from the image. At this stage, for example, the position and shape of the eyes, nose, mouth, and teeth are analyzed and used as foundational data for predicting future changes.
[0467] Based on the extracted feature data, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate multiple future scenarios depending on whether orthodontic treatment is performed and what method is used. These include various cases such as when no orthodontic treatment is performed and when a specific treatment method is used.
[0468] Next, the server retrieves information for analyzing past cases, particularly those that are similar. In this process, past treatment outcome data is searched using a transformer model to identify similarities with the current case.
[0469] The generated simulation results are then sent to the terminal. The terminal visually displays these images and information to the user. This display is organized in a format that allows the user to intuitively understand and compare treatment options.
[0470] Finally, when a user asks a question based on the simulation results, the device uses ChatGPT to provide a real-time explanation tailored to the question, aiding the user's understanding. For example, if a user requests to see "how the face would change after partial orthodontic treatment," that scenario is immediately generated, and detailed images and explanations are provided.
[0471] Based on the above, the system of the present invention provides users with detailed predictive data to select the optimal orthodontic treatment method, and continuously improves its accuracy through the use of past data and expert feedback.
[0472] The following describes the processing flow.
[0473] Step 1:
[0474] Users upload photos of their child's face and teeth to the system. The image files must be submitted in a specified format, so users should prepare them according to the guidelines.
[0475] Step 2:
[0476] The server receives the uploaded images and performs preprocessing on them. Specifically, this includes resizing, denoising, and color correction. This ensures the quality necessary for analysis.
[0477] Step 3:
[0478] The server feeds the preprocessed images to a convolutional neural network (CNN) to extract facial and tooth feature points. This feature extraction process identifies facial landmarks and tooth contours.
[0479] Step 4:
[0480] Based on the extracted features, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future tooth alignment and facial features. Multiple treatment scenarios are included, such as no orthodontics, partial orthodontics, and full orthodontics.
[0481] Step 5:
[0482] The server references past treatment data and searches for similar cases. It uses a transformer model to analyze the database and find the treatment pattern that best matches the current case.
[0483] Step 6:
[0484] The generated simulation results are sent to the terminal, which displays them visually to the user. The display includes detailed images showing the future facial features and teeth alignment, along with explanations of treatment options.
[0485] Step 7:
[0486] When a user asks a question about the simulation results, the terminal uses ChatGPT to generate a response in real time, providing detailed explanations and additional information.
[0487] (Example 1)
[0488] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0489] It is difficult for parents and medical professionals to obtain accurate early predictions about a child's future oral structure and facial shape. In particular, information is limited to determine which treatment options are most effective when corrections or adjustments are needed. This increases the risk of incorrect treatment choices and unnecessary procedures. Furthermore, the lack of readily available methods to obtain predictions based on past cases makes it difficult to select the optimal treatment plan for each individual child.
[0490] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0491] This invention includes a server that receives human facial and oral images and preprocesses them based on pre-set criteria, extracts facial and oral feature points from the preprocessed images using a machine learning algorithm, and simulates future oral structures and facial shapes using an advanced generative algorithm based on the extracted features. This makes it easy for parents and medical professionals to obtain accurate and diverse predictions based on various treatment scenarios.
[0492] A "human facial image" is a digital image that visually records the characteristics of an individual's face and includes information that represents the structure and shape of the face.
[0493] An "oral image" is an image that captures the inside of a person's mouth, particularly the condition of the teeth and gums, and is used for treatment and diagnosis.
[0494] "Preprocessing" refers to the process of converting image data into a format that is easy to analyze, and encompasses all processes including resizing, noise reduction, and color correction.
[0495] A "machine learning algorithm" is a technology that learns patterns and features from large amounts of data to make predictions and decisions, and is applied to image analysis and feature extraction.
[0496] "Feature point extraction" is the process of identifying and extracting points that represent important locations and shapes within an image, and these points are used as basic data for analysis and simulation.
[0497] An "advanced generative algorithm" is a complex computational method used to generate new information from data, and is particularly used when predicting future states.
[0498] "Oral structure" refers to the specific shape and arrangement of the inside of the mouth, such as the teeth, jawbone, and oral mucosa.
[0499] "Facial shape" refers to the specific form and contour of the external structure of a human face, and represents individual characteristics.
[0500] A "similar case" refers to a past case that shares many commonalities with the case currently being analyzed.
[0501] "Medical information" refers to information about a patient, including treatment and diagnostic results, and historical health data, which is used to determine and predict treatment plans.
[0502] "Diverse orthodontic treatment methods" refers to the various treatment options available in orthodontic treatment, including specific examples such as bracket orthodontics and clear aligner orthodontics.
[0503] A "human-to-human interaction device" is a device or interface for exchanging information with a human user, and is designed to provide information and answer questions through dialogue.
[0504] The embodiments for carrying out the present invention will now be described. This system involves the interaction of a user, a server, and a terminal to predict a child's future oral structure and facial shape. The roles of each are described in detail below.
[0505] First, users upload images of their child's face and mouth to the system using a mobile device or computer. These devices include typical smartphones and PCs with cameras. These images form the basis of the simulation data.
[0506] Next, these images are sent to a server. The server, equipped with high-performance computers and image processing software, performs preprocessing on the images. Specifically, it adjusts the image size, removes noise, and corrects color. Open-source image processing libraries are often used at this stage. Once preprocessing is complete, the images are prepared as data suitable for analysis by machine learning algorithms.
[0507] Next, the server runs a convolutional neural network (CNN) to extract facial and oral feature points. This identifies the location and shape of the eyes, nose, mouth, and teeth. Based on this data, generative adversarial networks (GANs) and variational autoencoders (VAEs) are used to simulate future oral structures and facial shapes. Here, multiple predictive patterns are generated for each treatment scenario.
[0508] The server then uses a transformer model to search a database of previously accumulated treatments and identify similar cases. This similarity information is used to improve the accuracy of the current simulation results.
[0509] Furthermore, the generated simulation results are sent to the terminal and displayed to the user in an easy-to-understand interface. The terminal features an intuitive user interface and is designed to allow the user to select and compare treatment scenarios.
[0510] Finally, users can ask questions about the simulation results, and the device uses ChatGPT to answer them in real time. This interactive process allows users to receive detailed explanations and deepen their understanding of treatment options.
[0511] As a concrete example, a possible prompt message could be, "Please show me the changes in the face after partial orthodontic treatment." In response to this prompt, the system quickly generates a corresponding scenario and provides the user with clear information.
[0512] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0513] Step 1:
[0514] Users upload facial and oral images to the system from their mobile devices or computers. The input data consists of facial and oral images, which are high-resolution digital images in JPEG or PNG format. These image data are then sent to the server as output. Users can easily select images taken with their device's camera and upload them according to the system's specified interface.
[0515] Step 2:
[0516] The server preprocesses the received images. The input data consists of images submitted by the user. Image processing software is used to perform noise reduction, color correction, and resizing. This process prepares the images for facial and oral feature extraction. The output is a preprocessed, clear image. This includes specific actions such as filtering and resizing pixel data using an image analysis library.
[0517] Step 3:
[0518] The server uses a convolutional neural network (CNN) to extract facial and oral feature points from pre-processed images. The input is pre-processed images. The CNN algorithm identifies the position and shape of the eyes, nose, mouth, and teeth. This feature extraction outputs detailed facial and oral data. By leveraging GPUs to accelerate calculations and efficiently execute large-scale models, the system performs the operation of accurately identifying feature points.
[0519] Step 4:
[0520] The server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future oral structures and facial shapes. The input is extracted feature data. This generates multiple future prediction results based on different orthodontic treatment scenarios. The output is a simulated image of the changes corresponding to each scenario. GANs and VAEs automatically generate various possibilities and create evolving images.
[0521] Step 5:
[0522] The server analyzes past treatment data using a transformer model to reference similar cases. Inputs include generated simulation results and a database of past treatments. The analysis outputs cases most similar to the current case, improving the reliability of the simulation. This includes using a deep learning model to detect patterns and similarities from a vast amount of case data.
[0523] Step 6:
[0524] The terminal displays the generated simulation results to the user. The input is simulation result data sent from the server. The output is in a format that allows for intuitive selection and comparison of each treatment scenario through an interactive user interface. The terminal utilizes high-speed browser rendering technology to provide smooth visual display.
[0525] Step 7:
[0526] When a user asks a question about the simulation results, the terminal uses ChatGPT to generate a response in real time. The input is the question prompt entered by the user. Based on this, an answer including detailed explanations and suggestions is output. Specifically, natural language processing techniques are used to understand the user's intent and provide an appropriate answer.
[0527] (Application Example 1)
[0528] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0529] Conventional autonomous vehicles do not adequately analyze the driver's gaze movements and posture changes in real time to support safe driving. Preventing unforeseen incidents caused by changes in gaze is difficult, and there is potential room for improvement in safety. This invention aims to solve these problems and provide a safer and more efficient driving assistance system.
[0530] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0531] In this invention, the server includes means for receiving a child's facial photograph and tooth image and preprocessing them according to predetermined criteria; means for extracting facial and tooth feature points from the preprocessed images using a convolutional neural network; means for simulating future tooth alignment and facial features using a generative adversarial network and a variational autoencoder based on the extracted features; means for acquiring the driver's gaze and posture data with a camera device and extracting visual features from the preprocessed images; and means for predicting future gaze movements based on the extracted visual features. This makes it possible to predict the driver's gaze movements and provide driver assistance information in advance.
[0532] A "child's facial photograph" is an electronically recorded and stored image of a child's face.
[0533] A "dental image" is a visual record of the condition of a child's teeth.
[0534] "Preprocessing means" refers to methods for adjusting the size, removing noise, and color-correcting the received image.
[0535] A "convolutional neural network" is a deep learning technique used to extract feature points from images.
[0536] "Feature points" are important data points related to the shape and position of the face and teeth.
[0537] A "generative adversarial network" is a machine learning technique used to mimic and generate data distributions.
[0538] A variational autoencoder is a variant of an autoencoder that estimates the latent variable space of data and enables the generation of new data.
[0539] "Simulation" is the process of modeling a real-world situation and predicting future situations based on that model.
[0540] "Eye movement prediction means" refers to a technology that analyzes the driver's eye direction and movement to predict future eye movements.
[0541] "Driver assistance information" refers to information provided to enable drivers to drive more safely and efficiently.
[0542] The system implementing this invention analyzes the driver's gaze and posture within an autonomous vehicle to support safe driving. The system's program includes functions to monitor the driver, predict their gaze movements in real time, and display appropriate driving support information.
[0543] The system uses a camera module mounted on the vehicle (for example, a Raspberry Pi Camera) to acquire frame data of the driver's face and eyes. The server receives this data and performs preprocessing such as resizing, noise reduction, and color correction. The server also uses a convolutional neural network (CNN) to extract important visual features. After feature extraction, it uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to predict the driver's gaze movements.
[0544] Based on predicted eye movements, the server displays appropriate driving assistance information on the vehicle's dashboard and the driver's smart glasses. For example, if the driver's gaze shifts while driving on a highway, the system immediately notifies the driver of the danger and prompts appropriate evasive action.
[0545] Furthermore, to analyze past driving data, the server utilizes a transformer model to search for similar driver reaction patterns and provide optimal information. The resulting information helps improve driver safety and efficiency.
[0546] An example of a prompt message is: "Design a system that analyzes the driver's gaze data while driving, predicts the direction that requires attention next, and warns the driver." This would allow drivers to operate vehicles more safely.
[0547] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0548] Step 1:
[0549] The server uses a camera module mounted on the vehicle to acquire image data of the driver's face and eyes. Based on this data, preprocessing is performed, including image resizing, noise reduction, and color correction, to enable accurate extraction of visual features. The input is raw image data, and the output is preprocessed, clean image data.
[0550] Step 2:
[0551] The server uses a convolutional neural network (CNN) to extract the driver's visual features from pre-processed image data. The input is the pre-processed image data, and the output consists of features related to the driver's face, gaze direction, and posture. These features are then further analyzed by edge AI software.
[0552] Step 3:
[0553] The server uses extracted features to predict the driver's future gaze patterns using generative adversarial networks (GANs) and variational autoencoders (VAEs). The input is the driver's visual features, and the output is predicted information about the direction the driver will next look. This allows drivers to know in advance which directions they should focus their attention while driving.
[0554] Step 4:
[0555] The server generates driver assistance information based on predicted eye movements and displays it on the vehicle's dashboard and the driver's smart glasses. The input is predicted eye movement information, and the output is navigation instructions and warning information to support safe driving.
[0556] Step 5:
[0557] The user adjusts their driving behavior based on driver assistance information provided by the system. In this process, the user makes judgments about the displayed information and takes appropriate driving actions. The input is the displayed driver assistance information, and the output is the user's driving actions.
[0558] Step 6:
[0559] The server uses data collected during driving and employs a transformer model to search for similar past driving patterns. The input is accumulated driving data, and the output is the analysis results of similar patterns. This enables the provision of support information optimized for each driver.
[0560] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0561] The system according to the present invention combines image recognition technology, generative AI technology, and emotion analysis technology to precisely predict a child's future tooth alignment and facial features, and further optimizes and presents the results according to the user's emotions. Specific embodiments are described below.
[0562] First, the user uploads a photo of their child's face and an image of their teeth to the system. This allows the system to obtain data to simulate future changes.
[0563] Next, the server receives these images and performs preprocessing. This preprocessing includes adjusting the image resolution and removing noise, preparing the images for image analysis.
[0564] Next, the server processes the pre-processed images using a convolutional neural network (CNN) to extract feature points of the face and teeth. At this stage, detailed landmarks are identified, enabling highly accurate predictions using a generative model.
[0565] Based on the extracted features, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future tooth alignment and facial appearance scenarios. The simulation provides multiple scenarios, including cases where orthodontic treatment is not performed and cases where different treatment methods are adopted.
[0566] Next, the server references accumulated historical treatment data and analyzes similar cases. By using a transformer model, insights gained from this data are applied to the current case to improve predictive accuracy.
[0567] The simulation results are then sent to the terminal and presented to the user. Crucially, this involves the use of an emotion analysis engine. The terminal collects emotional data from the user's camera and voice input, and analyzes it in real time. Based on the obtained emotional data, the presentation of the simulation results is optimized. For example, if the user is feeling anxious, more detailed explanations or additional support information may be displayed.
[0568] Furthermore, if a user has questions based on their response to the simulation results, the terminal generates a response using an interactive interface (e.g., ChatGPT) to address the question appropriately. This allows the user to feel secure through the system.
[0569] For example, if a user asks, "I want to know the difference between full and partial orthodontic treatment," the server will generate scenarios for each, and the device will provide a customized explanation that reflects the sentiment analysis results. Furthermore, the accuracy of predictions will be improved by continuously incorporating feedback from experts into the model.
[0570] Based on the above, the present invention provides an environment in which users can understand and select the optimal orthodontic treatment method, and further provides a more personalized experience by adjusting the method of presenting results according to their emotions.
[0571] The following describes the processing flow.
[0572] Step 1:
[0573] Users select a photo of their child's face and an image of their teeth and upload them to the system's online platform. The upload process is performed from the user's device, and the images are sent to the system's server.
[0574] Step 2:
[0575] The server stores the received images and begins preprocessing. Specifically, the images are resized, compressed, denoised, and retouched to ensure a clean dataset that can be input into machine learning models.
[0576] Step 3:
[0577] The server feeds the preprocessed images into a convolutional neural network (CNN) to extract facial and dental feature points. This identifies features such as jaw shape, tooth arrangement, and landmarks for the eyes and mouth.
[0578] Step 4:
[0579] The server uses feature point data to simulate future tooth alignment and facial appearance scenarios using generative adversarial networks (GANs) and variational autoencoders (VAEs). Each scenario is generated based on whether or not orthodontic treatment is performed and different treatment methods.
[0580] Step 5:
[0581] The server uses accumulated historical treatment data to search for similar cases and performs analysis using a transformer model. The results are then incorporated to improve the accuracy of the simulation.
[0582] Step 6:
[0583] The results of the generated scenarios are sent to the terminal, which visually presents them to the user through the user interface. The presentation includes images of the future facial appearance and tooth alignment related to each orthodontic treatment scenario.
[0584] Step 7:
[0585] The device activates an emotion analysis engine and collects emotional information from the user's facial expressions and voice. The emotional data is analyzed in real time to determine whether the user is experiencing anxiety or confusion.
[0586] Step 8:
[0587] The device optimizes the display of simulation results based on the user's emotional state, as needed. For example, if anxiety is detected, additional explanations or encouraging messages are provided.
[0588] Step 9:
[0589] When a user asks a question about the simulation results, the device responds in real time using an interactive interface. The AI model used here provides detailed information in response to the question.
[0590] Step 10:
[0591] Finally, the server receives feedback from experts and incorporates the newly acquired experience into the learning model. This allows the system to provide even more accurate simulations.
[0592] (Example 2)
[0593] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0594] Given the current difficulty in comparing multiple orthodontic treatment methods and making the optimal choice for a child's future teeth alignment and facial appearance, there is a need for a method that provides clear information while considering the user's emotional state, allowing them to make informed decisions with confidence. Furthermore, improving predictive accuracy and creating an environment where users can interactively obtain information without feeling anxious are key challenges.
[0595] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0596] In this invention, the server includes means for receiving a child's facial photograph and dental images and pre-processing them according to predetermined criteria; means for extracting facial and dental feature points from the pre-processed images using a convolutional neural network; means for simulating future tooth alignment and facial appearance using a generative adversarial network and a variational autoencoder based on the extracted features; means for making predictions based on similar cases using accumulated past treatment data; and means for displaying the simulation results generated based on multiple orthodontic treatment scenarios on a communication terminal, and optimizing the method of presenting the results while taking into account the user's emotional state. As a result, the user can more easily understand the future simulation results and can select an appropriate orthodontic treatment method without feeling anxious.
[0597] "Preprocessing" refers to the process of adjusting the resolution and removing noise in order to improve the quality of received image data.
[0598] A "convolutional neural network" is an artificial intelligence technology used to analyze images and features, and is a model that has particularly high discriminatory capabilities for visual data.
[0599] "Feature point extraction" is the process of deriving important identifying information about the shape of faces and teeth from images, and it forms the basis of predictive simulations.
[0600] A "generative adversarial network" is a type of machine learning model used to simulate images and data. It is a technique that generates highly accurate output by having a generative model and a discriminative model compete alternately.
[0601] A variational autoencoder is a type of autoencoder that has the ability to learn the latent structure of data and generate new data, making it possible to efficiently reproduce the distribution of data.
[0602] "Simulation" is a process of virtual experimentation that uses a model to predict future situations and conditions, and is a method of predicting future tooth alignment and facial appearance through different treatment scenarios.
[0603] "Emotion analysis" is a technology that evaluates a user's emotional state in real time from information collected through their camera and voice data, and then analyzes that data.
[0604] An "interactive interface" is a system designed to facilitate natural conversations and responses with users, and its role is to provide appropriate information and answers to user input.
[0605] The system of this invention combines image processing technology and machine learning technology to predict future changes in a child's teeth alignment and facial features, and supports the user in selecting an appropriate treatment method.
[0606] First, the user uploads a photo of their child's face and teeth to the system. The server then retrieves the data necessary for the simulation. This image data is preprocessed, including resolution adjustment and noise reduction, to be converted into a format suitable for analysis. Dedicated image editing software and algorithms are used for this preprocessing.
[0607] Next, the server uses a convolutional neural network (CNN) to extract facial and tooth feature points from the pre-processed images. This identifies landmarks of facial features and tooth shape. Based on these features, a generative adversarial network (GAN) and a variational autoencoder (VAE) are used to simulate future tooth alignment and facial appearance scenarios.
[0608] The system further analyzes past treatment data using a transformer model, compares it with similar cases, and improves the accuracy of predictions. As a result, multiple scenarios based on different orthodontic treatment options are generated, showing the expected outcomes associated with each treatment method.
[0609] These simulation results are presented to the user via a terminal. The terminal analyzes the user's emotional state in real time through the user's camera and voice input, and displays information in the most appropriate format for the user's state. For example, if the user is feeling anxious, the system will provide additional explanations or information to aid understanding.
[0610] If a user has questions based on the information presented, the device utilizes an interactive interface and uses natural language processing to respond to the user's questions. This includes using prompts such as "Please explain in detail the difference between full and partial orthodontic treatment" using a generative AI model.
[0611] In this way, the system provides users with clear and easy-to-understand information, creating an environment where they can confidently choose the optimal treatment method.
[0612] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0613] Step 1:
[0614] The user uploads a photo of their child's face and an image of their teeth to the system. By sending these images to the server as input data, the server obtains the basic data needed to simulate future changes. The specific action involved is for the user to select image files using a terminal and send them to the server via the upload function.
[0615] Step 2:
[0616] The server preprocesses the received image data. First, it applies resolution adjustment to the original input image, and then applies a noise reduction algorithm. This results in clear image data suitable for analysis as output. In this process, image editing software and specific algorithms are used to prepare the data to an optimal state.
[0617] Step 3:
[0618] The server inputs preprocessed images into a convolutional neural network (CNN) to extract feature points of the face and teeth. The CNN performs pixel-level analysis of the preprocessed input images to identify landmarks that indicate facial features and tooth shape. The output is the feature point data necessary for the simulation. This process utilizes a highly accurate model.
[0619] Step 4:
[0620] The server uses a generative adversarial network (GAN) and a variational autoencoder (VAE) to simulate future tooth alignment and facial features based on feature data. It constructs multiple scenarios using the feature data as input, making predictions that consider the impact of different orthodontic treatment methods. The final output is a visual future scenario for each treatment.
[0621] Step 5:
[0622] The server analyzes the simulation results using a transformer model, comparing them with accumulated historical treatment data. At this stage, historical data is input, and similar cases are analyzed to improve the accuracy of the simulation results and derive the optimal solution.
[0623] Step 6:
[0624] The terminal presents the final simulation results to the user. At this time, the user's emotional data, collected through a camera or voice input device, is used as input. The terminal performs real-time emotional analysis and optimizes the presentation method of the results according to the user's emotional state. The output is provided in a format that is easy for the user to understand.
[0625] Step 7:
[0626] If the user has questions based on the simulation results, the terminal responds with an interactive interface. It processes user input as prompts, uses a generative AI model to create natural language answers, and presents them to the user. Detailed explanations and additional information are provided as output to aid user understanding.
[0627] (Application Example 2)
[0628] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0629] The objective of this invention is to create an environment in which users can make appropriate decisions with confidence by accurately predicting future changes in the face and oral cavity, enabling the comparison of different orthodontic treatment scenarios, and providing optimal information presentation tailored to the user's emotional state.
[0630] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0631] In this invention, the server includes means for receiving images of the face and oral cavity and preprocessing them according to predetermined criteria, means for extracting facial and oral cavity features from the preprocessed images using deep learning technology, and means for predicting future oral and facial features using a generative model based on the extracted features. This enables highly accurate predictions and the presentation of information that responds to the user's emotional state.
[0632] A "facial image" is visual data that includes the entire face of the subject, and it forms the basis for feature extraction.
[0633] An "oral image" is visual data that captures the inside of a subject's mouth in detail, and is used to analyze the condition of their teeth and gums.
[0634] "Preprocessing" refers to processes that include adjusting the resolution of image data and removing noise, and is preparatory work to facilitate subsequent data analysis.
[0635] "Deep learning technology" is an information processing technology that uses multi-layered artificial neural networks to achieve advanced pattern recognition and feature extraction.
[0636] A "generative model" is a technique that generates new data using methods such as generative adversarial networks and variational autoencoders, and is used for prediction and simulation.
[0637] An "orthodontic treatment scenario" is a future treatment plan that assumes different orthodontic methods and treatment flows, and serves as a reference for users to make their choices.
[0638] "Emotional analysis technology" is a technology that detects and analyzes a user's emotional state from images and audio, and is used to provide personalized information.
[0639] The system for carrying out this invention uses a smart device to collect images of the face and oral cavity and transmit them to a cloud server. When a user uploads images through the device, the server preprocesses the received images. This preprocessing adjusts the image resolution and removes noise to prepare the images for analysis.
[0640] Next, the server utilizes deep learning techniques to extract facial and oral features from the preprocessed images. Specifically, it uses a convolutional neural network (CNN) to identify facial and oral feature points. This data then forms the basis for future predictions by subsequent generative models.
[0641] Based on the extracted features, the server predicts the future appearance of the oral cavity and face using generative adversarial networks (GANs) and variational autoencoders (VAEs). This makes it possible to provide users with data that allows them to compare and consider different orthodontic treatment scenarios.
[0642] Furthermore, the server references accumulated past treatment data and analyzes similar cases. By using a transformer model, past knowledge is reflected in current predictions, improving accuracy.
[0643] Furthermore, it utilizes emotion analysis technology to detect the user's emotional state in real time. Based on the obtained emotional data, the system optimizes how prediction results are presented, providing the user with the most reassuring information. For example, if the user is feeling anxious, it will provide more detailed explanations or additional support information.
[0644] For example, if a user asks a question such as, "I want to know the difference between full or partial orthodontic treatment," the system will generate scenarios for each, and the device will provide a customized explanation that reflects the sentiment analysis results.
[0645] An example of a prompt message for a generative AI model would be: "I have uploaded a photo. Please predict future changes in teeth alignment and facial features and create a scenario for orthodontic treatment. Also, since the user appears anxious, please include detailed explanations and reassuring information."
[0646] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0647] Step 1:
[0648] The user takes images of their face and mouth using a device and uploads them to the server. The input here is the captured image data, and the output is the transmission of the image data to the server.
[0649] Step 2:
[0650] The server performs preprocessing on the received image data. The input is uploaded facial and oral cavity image data, and the server processes the data, such as adjusting the resolution and removing noise, to output image data that is ready for analysis.
[0651] Step 3:
[0652] The server applies deep learning techniques to preprocessed image data to extract facial and oral features. Specifically, it uses a convolutional neural network (CNN) to extract feature points. The input is preprocessed image data, and the output is the extracted feature data.
[0653] Step 4:
[0654] The server uses a generative model to predict future oral and facial features based on extracted feature data. Generative adversarial networks (GANs) and variational autoencoders (VAEs) are used here. The input is feature data, and the server generates and outputs simulation data for future predictions.
[0655] Step 5:
[0656] The server references past treatment data and uses a transformer model to analyze similar cases. This provides insights to improve prediction accuracy. The input is past treatment data, and the output is enhanced prediction data.
[0657] Step 6:
[0658] The server uses emotion analysis technology to analyze the user's emotional state in real time. The input is data on the user's current emotional state (obtained from facial and voice), and the server outputs data to optimize the presentation method based on the emotional information.
[0659] Step 7:
[0660] The terminal displays predictive data received from the server through a user interface. Here, the information is presented using optimized data tailored to the user's emotional state. Input consists of predictive data and sentiment analysis data, and the terminal outputs customized information.
[0661] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0662] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0663] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0664] [Fourth Embodiment]
[0665] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0666] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0667] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0668] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0669] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0670] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0671] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0672] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0673] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0674] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0675] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0676] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0677] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0678] The system according to the present invention combines image recognition technology and generative AI technology to predict a child's future tooth alignment and facial features. It is implemented through the interaction of the user, server, and terminal as follows.
[0679] First, users upload a photo of their child's face and images of their teeth to the system from a mobile device or computer. These images are used as foundational data to simulate future conditions.
[0680] Next, the server receives the uploaded images and performs a series of preprocessing steps. These preprocessing steps include resizing, denoising, and color correction. This prepares the images for accurate extraction of facial and dental features.
[0681] Next, the server uses a convolutional neural network (CNN) to extract key feature points of the face and teeth from the image. At this stage, for example, the position and shape of the eyes, nose, mouth, and teeth are analyzed and used as foundational data for predicting future changes.
[0682] Based on the extracted feature data, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate multiple future scenarios depending on whether orthodontic treatment is performed and what method is used. These include various cases such as when no orthodontic treatment is performed and when a specific treatment method is used.
[0683] Next, the server retrieves information for analyzing past cases, particularly those that are similar. In this process, past treatment outcome data is searched using a transformer model to identify similarities with the current case.
[0684] The generated simulation results are then sent to the terminal. The terminal visually displays these images and information to the user. This display is organized in a format that allows the user to intuitively understand and compare treatment options.
[0685] Finally, when a user asks a question based on the simulation results, the device uses ChatGPT to provide a real-time explanation tailored to the question, aiding the user's understanding. For example, if a user requests to see "how the face would change after partial orthodontic treatment," that scenario is immediately generated, and detailed images and explanations are provided.
[0686] Based on the above, the system of the present invention provides users with detailed predictive data to select the optimal orthodontic treatment method, and continuously improves its accuracy through the use of past data and expert feedback.
[0687] The following describes the processing flow.
[0688] Step 1:
[0689] Users upload photos of their child's face and teeth to the system. The image files must be submitted in a specified format, so users should prepare them according to the guidelines.
[0690] Step 2:
[0691] The server receives the uploaded images and performs preprocessing on them. Specifically, this includes resizing, denoising, and color correction. This ensures the quality necessary for analysis.
[0692] Step 3:
[0693] The server feeds the preprocessed images to a convolutional neural network (CNN) to extract facial and tooth feature points. This feature extraction process identifies facial landmarks and tooth contours.
[0694] Step 4:
[0695] Based on the extracted features, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future tooth alignment and facial features. Multiple treatment scenarios are included, such as no orthodontics, partial orthodontics, and full orthodontics.
[0696] Step 5:
[0697] The server references past treatment data and searches for similar cases. It uses a transformer model to analyze the database and find the treatment pattern that best matches the current case.
[0698] Step 6:
[0699] The generated simulation results are sent to the terminal, which displays them visually to the user. The display includes detailed images showing the future facial features and teeth alignment, along with explanations of treatment options.
[0700] Step 7:
[0701] When a user asks a question about the simulation results, the terminal uses ChatGPT to generate a response in real time, providing detailed explanations and additional information.
[0702] (Example 1)
[0703] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0704] It is difficult for parents and medical professionals to obtain accurate early predictions about a child's future oral structure and facial shape. In particular, information is limited to determine which treatment options are most effective when corrections or adjustments are needed. This increases the risk of incorrect treatment choices and unnecessary procedures. Furthermore, the lack of readily available methods to obtain predictions based on past cases makes it difficult to select the optimal treatment plan for each individual child.
[0705] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0706] This invention includes a server that receives human facial and oral images and preprocesses them based on pre-set criteria, extracts facial and oral feature points from the preprocessed images using a machine learning algorithm, and simulates future oral structures and facial shapes using an advanced generative algorithm based on the extracted features. This makes it easy for parents and medical professionals to obtain accurate and diverse predictions based on various treatment scenarios.
[0707] A "human facial image" is a digital image that visually records the characteristics of an individual's face and includes information that represents the structure and shape of the face.
[0708] An "oral image" is an image that captures the inside of a person's mouth, particularly the condition of the teeth and gums, and is used for treatment and diagnosis.
[0709] "Preprocessing" refers to the process of converting image data into a format that is easy to analyze, and encompasses all processes including resizing, noise reduction, and color correction.
[0710] A "machine learning algorithm" is a technology that learns patterns and features from large amounts of data to make predictions and decisions, and is applied to image analysis and feature extraction.
[0711] "Feature point extraction" is the process of identifying and extracting points that represent important locations and shapes within an image, and these points are used as basic data for analysis and simulation.
[0712] An "advanced generative algorithm" is a complex computational method used to generate new information from data, and is particularly used when predicting future states.
[0713] "Oral structure" refers to the specific shape and arrangement of the inside of the mouth, such as the teeth, jawbone, and oral mucosa.
[0714] "Facial shape" refers to the specific form and contour of the external structure of a human face, and represents individual characteristics.
[0715] A "similar case" refers to a past case that shares many commonalities with the case currently being analyzed.
[0716] "Medical information" refers to information about a patient, including treatment and diagnostic results, and historical health data, which is used to determine and predict treatment plans.
[0717] "Diverse orthodontic treatment methods" refers to the various treatment options available in orthodontic treatment, including specific examples such as bracket orthodontics and clear aligner orthodontics.
[0718] A "human-to-human interaction device" is a device or interface for exchanging information with a human user, and is designed to provide information and answer questions through dialogue.
[0719] The embodiments for carrying out the present invention will now be described. This system involves the interaction of a user, a server, and a terminal to predict a child's future oral structure and facial shape. The roles of each are described in detail below.
[0720] First, users upload images of their child's face and mouth to the system using a mobile device or computer. These devices include typical smartphones and PCs with cameras. These images form the basis of the simulation data.
[0721] Next, these images are sent to a server. The server, equipped with high-performance computers and image processing software, performs preprocessing on the images. Specifically, it adjusts the image size, removes noise, and corrects color. Open-source image processing libraries are often used at this stage. Once preprocessing is complete, the images are prepared as data suitable for analysis by machine learning algorithms.
[0722] Next, the server runs a convolutional neural network (CNN) to extract facial and oral feature points. This identifies the location and shape of the eyes, nose, mouth, and teeth. Based on this data, generative adversarial networks (GANs) and variational autoencoders (VAEs) are used to simulate future oral structures and facial shapes. Here, multiple predictive patterns are generated for each treatment scenario.
[0723] The server then uses a transformer model to search a database of previously accumulated treatments and identify similar cases. This similarity information is used to improve the accuracy of the current simulation results.
[0724] Furthermore, the generated simulation results are sent to the terminal and displayed to the user in an easy-to-understand interface. The terminal features an intuitive user interface and is designed to allow the user to select and compare treatment scenarios.
[0725] Finally, users can ask questions about the simulation results, and the device uses ChatGPT to answer them in real time. This interactive process allows users to receive detailed explanations and deepen their understanding of treatment options.
[0726] As a concrete example, a possible prompt message could be, "Please show me the changes in the face after partial orthodontic treatment." In response to this prompt, the system quickly generates a corresponding scenario and provides the user with clear information.
[0727] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0728] Step 1:
[0729] Users upload facial and oral images to the system from their mobile devices or computers. The input data consists of facial and oral images, which are high-resolution digital images in JPEG or PNG format. These image data are then sent to the server as output. Users can easily select images taken with their device's camera and upload them according to the system's specified interface.
[0730] Step 2:
[0731] The server preprocesses the received images. The input data consists of images submitted by the user. Image processing software is used to perform noise reduction, color correction, and resizing. This process prepares the images for facial and oral feature extraction. The output is a preprocessed, clear image. This includes specific actions such as filtering and resizing pixel data using an image analysis library.
[0732] Step 3:
[0733] The server uses a convolutional neural network (CNN) to extract facial and oral feature points from pre-processed images. The input is pre-processed images. The CNN algorithm identifies the position and shape of the eyes, nose, mouth, and teeth. This feature extraction outputs detailed facial and oral data. By leveraging GPUs to accelerate calculations and efficiently execute large-scale models, the system performs the operation of accurately identifying feature points.
[0734] Step 4:
[0735] The server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future oral structures and facial shapes. The input is extracted feature data. This generates multiple future prediction results based on different orthodontic treatment scenarios. The output is a simulated image of the changes corresponding to each scenario. GANs and VAEs automatically generate various possibilities and create evolving images.
[0736] Step 5:
[0737] The server analyzes past treatment data using a transformer model to reference similar cases. Inputs include generated simulation results and a database of past treatments. The analysis outputs cases most similar to the current case, improving the reliability of the simulation. This includes using a deep learning model to detect patterns and similarities from a vast amount of case data.
[0738] Step 6:
[0739] The terminal displays the generated simulation results to the user. The input is simulation result data sent from the server. The output is in a format that allows for intuitive selection and comparison of each treatment scenario through an interactive user interface. The terminal utilizes high-speed browser rendering technology to provide smooth visual display.
[0740] Step 7:
[0741] When a user asks a question about the simulation results, the terminal uses ChatGPT to generate a response in real time. The input is the question prompt entered by the user. Based on this, an answer including detailed explanations and suggestions is output. Specifically, natural language processing techniques are used to understand the user's intent and provide an appropriate answer.
[0742] (Application Example 1)
[0743] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0744] Conventional autonomous vehicles do not adequately analyze the driver's gaze movements and posture changes in real time to support safe driving. Preventing unforeseen incidents caused by changes in gaze is difficult, and there is potential room for improvement in safety. This invention aims to solve these problems and provide a safer and more efficient driving assistance system.
[0745] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0746] In this invention, the server includes means for receiving a child's facial photograph and tooth image and preprocessing them according to predetermined criteria; means for extracting facial and tooth feature points from the preprocessed images using a convolutional neural network; means for simulating future tooth alignment and facial features using a generative adversarial network and a variational autoencoder based on the extracted features; means for acquiring the driver's gaze and posture data with a camera device and extracting visual features from the preprocessed images; and means for predicting future gaze movements based on the extracted visual features. This makes it possible to predict the driver's gaze movements and provide driver assistance information in advance.
[0747] A "child's facial photograph" is an electronically recorded and stored image of a child's face.
[0748] A "dental image" is a visual record of the condition of a child's teeth.
[0749] "Preprocessing means" refers to methods for adjusting the size, removing noise, and color-correcting the received image.
[0750] A "convolutional neural network" is a deep learning technique used to extract feature points from images.
[0751] "Feature points" are important data points related to the shape and position of the face and teeth.
[0752] A "generative adversarial network" is a machine learning technique used to mimic and generate data distributions.
[0753] A variational autoencoder is a variant of an autoencoder that estimates the latent variable space of data and enables the generation of new data.
[0754] "Simulation" is the process of modeling a real-world situation and predicting future situations based on that model.
[0755] "Eye movement prediction means" refers to a technology that analyzes the driver's eye direction and movement to predict future eye movements.
[0756] "Driver assistance information" refers to information provided to enable drivers to drive more safely and efficiently.
[0757] The system implementing this invention analyzes the driver's gaze and posture within an autonomous vehicle to support safe driving. The system's program includes functions to monitor the driver, predict their gaze movements in real time, and display appropriate driving support information.
[0758] The system uses a camera module mounted on the vehicle (for example, a Raspberry Pi Camera) to acquire frame data of the driver's face and eyes. The server receives this data and performs preprocessing such as resizing, noise reduction, and color correction. The server also uses a convolutional neural network (CNN) to extract important visual features. After feature extraction, it uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to predict the driver's gaze movements.
[0759] Based on predicted eye movements, the server displays appropriate driving assistance information on the vehicle's dashboard and the driver's smart glasses. For example, if the driver's gaze shifts while driving on a highway, the system immediately notifies the driver of the danger and prompts appropriate evasive action.
[0760] Furthermore, to analyze past driving data, the server utilizes a transformer model to search for similar driver reaction patterns and provide optimal information. The resulting information helps improve driver safety and efficiency.
[0761] An example of a prompt message is: "Design a system that analyzes the driver's gaze data while driving, predicts the direction that requires attention next, and warns the driver." This would allow drivers to operate vehicles more safely.
[0762] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0763] Step 1:
[0764] The server uses a camera module mounted on the vehicle to acquire image data of the driver's face and eyes. Based on this data, preprocessing is performed, including image resizing, noise reduction, and color correction, to enable accurate extraction of visual features. The input is raw image data, and the output is preprocessed, clean image data.
[0765] Step 2:
[0766] The server uses a convolutional neural network (CNN) to extract the driver's visual features from pre-processed image data. The input is the pre-processed image data, and the output consists of features related to the driver's face, gaze direction, and posture. These features are then further analyzed by edge AI software.
[0767] Step 3:
[0768] The server uses extracted features to predict the driver's future gaze patterns using generative adversarial networks (GANs) and variational autoencoders (VAEs). The input is the driver's visual features, and the output is predicted information about the direction the driver will next look. This allows drivers to know in advance which directions they should focus their attention while driving.
[0769] Step 4:
[0770] The server generates driver assistance information based on predicted eye movements and displays it on the vehicle's dashboard and the driver's smart glasses. The input is predicted eye movement information, and the output is navigation instructions and warning information to support safe driving.
[0771] Step 5:
[0772] The user adjusts their driving behavior based on driver assistance information provided by the system. In this process, the user makes judgments about the displayed information and takes appropriate driving actions. The input is the displayed driver assistance information, and the output is the user's driving actions.
[0773] Step 6:
[0774] The server uses data collected during driving and employs a transformer model to search for similar past driving patterns. The input is accumulated driving data, and the output is the analysis results of similar patterns. This enables the provision of support information optimized for each driver.
[0775] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0776] The system according to the present invention combines image recognition technology, generative AI technology, and emotion analysis technology to precisely predict a child's future tooth alignment and facial features, and further optimizes and presents the results according to the user's emotions. Specific embodiments are described below.
[0777] First, the user uploads a photo of their child's face and an image of their teeth to the system. This allows the system to obtain data to simulate future changes.
[0778] Next, the server receives these images and performs preprocessing. This preprocessing includes adjusting the image resolution and removing noise, preparing the images for image analysis.
[0779] Next, the server processes the pre-processed images using a convolutional neural network (CNN) to extract feature points of the face and teeth. At this stage, detailed landmarks are identified, enabling highly accurate predictions using a generative model.
[0780] Based on the extracted features, the server uses generative adversarial networks (GANs) and variational autoencoders (VAEs) to simulate future tooth alignment and facial appearance scenarios. The simulation provides multiple scenarios, including cases where orthodontic treatment is not performed and cases where different treatment methods are adopted.
[0781] Next, the server references accumulated historical treatment data and analyzes similar cases. By using a transformer model, insights gained from this data are applied to the current case to improve predictive accuracy.
[0782] The simulation results are then sent to the terminal and presented to the user. Crucially, this involves the use of an emotion analysis engine. The terminal collects emotional data from the user's camera and voice input, and analyzes it in real time. Based on the obtained emotional data, the presentation of the simulation results is optimized. For example, if the user is feeling anxious, more detailed explanations or additional support information may be displayed.
[0783] Furthermore, if a user has questions based on their response to the simulation results, the terminal generates a response using an interactive interface (e.g., ChatGPT) to address the question appropriately. This allows the user to feel secure through the system.
[0784] For example, if a user asks, "I want to know the difference between full and partial orthodontic treatment," the server will generate scenarios for each, and the device will provide a customized explanation that reflects the sentiment analysis results. Furthermore, the accuracy of predictions will be improved by continuously incorporating feedback from experts into the model.
[0785] Based on the above, the present invention provides an environment in which users can understand and select the optimal orthodontic treatment method, and further provides a more personalized experience by adjusting the method of presenting results according to their emotions.
[0786] The following describes the processing flow.
[0787] Step 1:
[0788] Users select a photo of their child's face and an image of their teeth and upload them to the system's online platform. The upload process is performed from the user's device, and the images are sent to the system's server.
[0789] Step 2:
[0790] The server stores the received images and begins preprocessing. Specifically, the images are resized, compressed, denoised, and retouched to ensure a clean dataset that can be input into machine learning models.
[0791] Step 3:
[0792] The server feeds the preprocessed images into a convolutional neural network (CNN) to extract facial and dental feature points. This identifies features such as jaw shape, tooth arrangement, and landmarks for the eyes and mouth.
[0793] Step 4:
[0794] The server uses feature point data to simulate future tooth alignment and facial appearance scenarios using generative adversarial networks (GANs) and variational autoencoders (VAEs). Each scenario is generated based on whether or not orthodontic treatment is performed and different treatment methods.
[0795] Step 5:
[0796] The server uses accumulated historical treatment data to search for similar cases and performs analysis using a transformer model. The results are then incorporated to improve the accuracy of the simulation.
[0797] Step 6:
[0798] The results of the generated scenarios are sent to the terminal, which visually presents them to the user through the user interface. The presentation includes images of the future facial appearance and tooth alignment related to each orthodontic treatment scenario.
[0799] Step 7:
[0800] The device activates an emotion analysis engine and collects emotional information from the user's facial expressions and voice. The emotional data is analyzed in real time to determine whether the user is experiencing anxiety or confusion.
[0801] Step 8:
[0802] The device optimizes the display of simulation results based on the user's emotional state, as needed. For example, if anxiety is detected, additional explanations or encouraging messages are provided.
[0803] Step 9:
[0804] When a user asks a question about the simulation results, the device responds in real time using an interactive interface. The AI model used here provides detailed information in response to the question.
[0805] Step 10:
[0806] Finally, the server receives feedback from experts and incorporates the newly acquired experience into the learning model. This allows the system to provide even more accurate simulations.
[0807] (Example 2)
[0808] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0809] Given the current difficulty in comparing multiple orthodontic treatment methods and making the optimal choice for a child's future teeth alignment and facial appearance, there is a need for a method that provides clear information while considering the user's emotional state, allowing them to make informed decisions with confidence. Furthermore, improving predictive accuracy and creating an environment where users can interactively obtain information without feeling anxious are key challenges.
[0810] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0811] In this invention, the server includes means for receiving a child's facial photograph and dental images and pre-processing them according to predetermined criteria; means for extracting facial and dental feature points from the pre-processed images using a convolutional neural network; means for simulating future tooth alignment and facial appearance using a generative adversarial network and a variational autoencoder based on the extracted features; means for making predictions based on similar cases using accumulated past treatment data; and means for displaying the simulation results generated based on multiple orthodontic treatment scenarios on a communication terminal, and optimizing the method of presenting the results while taking into account the user's emotional state. As a result, the user can more easily understand the future simulation results and can select an appropriate orthodontic treatment method without feeling anxious.
[0812] "Preprocessing" refers to the process of adjusting the resolution and removing noise in order to improve the quality of received image data.
[0813] A "convolutional neural network" is an artificial intelligence technology used to analyze images and features, and is a model that has particularly high discriminatory capabilities for visual data.
[0814] "Feature point extraction" is the process of deriving important identifying information about the shape of faces and teeth from images, and it forms the basis of predictive simulations.
[0815] A "generative adversarial network" is a type of machine learning model used to simulate images and data. It is a technique that generates highly accurate output by having a generative model and a discriminative model compete alternately.
[0816] A variational autoencoder is a type of autoencoder that has the ability to learn the latent structure of data and generate new data, making it possible to efficiently reproduce the distribution of data.
[0817] "Simulation" is a process of virtual experimentation that uses a model to predict future situations and conditions, and is a method of predicting future tooth alignment and facial appearance through different treatment scenarios.
[0818] "Emotion analysis" is a technology that evaluates a user's emotional state in real time from information collected through their camera and voice data, and then analyzes that data.
[0819] An "interactive interface" is a system designed to facilitate natural conversations and responses with users, and its role is to provide appropriate information and answers to user input.
[0820] The system of this invention combines image processing technology and machine learning technology to predict future changes in a child's teeth alignment and facial features, and supports the user in selecting an appropriate treatment method.
[0821] First, the user uploads a photo of their child's face and teeth to the system. The server then retrieves the data necessary for the simulation. This image data is preprocessed, including resolution adjustment and noise reduction, to be converted into a format suitable for analysis. Dedicated image editing software and algorithms are used for this preprocessing.
[0822] Next, the server uses a convolutional neural network (CNN) to extract facial and tooth feature points from the pre-processed images. This identifies landmarks of facial features and tooth shape. Based on these features, a generative adversarial network (GAN) and a variational autoencoder (VAE) are used to simulate future tooth alignment and facial appearance scenarios.
[0823] The system further analyzes past treatment data using a transformer model, compares it with similar cases, and improves the accuracy of predictions. As a result, multiple scenarios based on different orthodontic treatment options are generated, showing the expected outcomes associated with each treatment method.
[0824] These simulation results are presented to the user via a terminal. The terminal analyzes the user's emotional state in real time through the user's camera and voice input, and displays information in the most appropriate format for the user's state. For example, if the user is feeling anxious, the system will provide additional explanations or information to aid understanding.
[0825] If a user has questions based on the information presented, the device utilizes an interactive interface and uses natural language processing to respond to the user's questions. This includes using prompts such as "Please explain in detail the difference between full and partial orthodontic treatment" using a generative AI model.
[0826] In this way, the system provides users with clear and easy-to-understand information, creating an environment where they can confidently choose the optimal treatment method.
[0827] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0828] Step 1:
[0829] The user uploads a photo of their child's face and an image of their teeth to the system. By sending these images to the server as input data, the server obtains the basic data needed to simulate future changes. The specific action involved is for the user to select image files using a terminal and send them to the server via the upload function.
[0830] Step 2:
[0831] The server preprocesses the received image data. First, it applies resolution adjustment to the original input image, and then applies a noise reduction algorithm. This results in clear image data suitable for analysis as output. In this process, image editing software and specific algorithms are used to prepare the data to an optimal state.
[0832] Step 3:
[0833] The server inputs preprocessed images into a convolutional neural network (CNN) to extract feature points of the face and teeth. The CNN performs pixel-level analysis of the preprocessed input images to identify landmarks that indicate facial features and tooth shape. The output is the feature point data necessary for the simulation. This process utilizes a highly accurate model.
[0834] Step 4:
[0835] The server uses a generative adversarial network (GAN) and a variational autoencoder (VAE) to simulate future tooth alignment and facial features based on feature data. It constructs multiple scenarios using the feature data as input, making predictions that consider the impact of different orthodontic treatment methods. The final output is a visual future scenario for each treatment.
[0836] Step 5:
[0837] The server analyzes the simulation results using a transformer model, comparing them with accumulated historical treatment data. At this stage, historical data is input, and similar cases are analyzed to improve the accuracy of the simulation results and derive the optimal solution.
[0838] Step 6:
[0839] The terminal presents the final simulation results to the user. At this time, the user's emotional data, collected through a camera or voice input device, is used as input. The terminal performs real-time emotional analysis and optimizes the presentation method of the results according to the user's emotional state. The output is provided in a format that is easy for the user to understand.
[0840] Step 7:
[0841] If the user has questions based on the simulation results, the terminal responds with an interactive interface. It processes user input as prompts, uses a generative AI model to create natural language answers, and presents them to the user. Detailed explanations and additional information are provided as output to aid user understanding.
[0842] (Application Example 2)
[0843] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0844] The objective of this invention is to create an environment in which users can make appropriate decisions with confidence by accurately predicting future changes in the face and oral cavity, enabling the comparison of different orthodontic treatment scenarios, and providing optimal information presentation tailored to the user's emotional state.
[0845] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0846] In this invention, the server includes means for receiving images of the face and oral cavity and preprocessing them according to predetermined criteria, means for extracting facial and oral cavity features from the preprocessed images using deep learning technology, and means for predicting future oral and facial features using a generative model based on the extracted features. This enables highly accurate predictions and the presentation of information that responds to the user's emotional state.
[0847] A "facial image" is visual data that includes the entire face of the subject, and it forms the basis for feature extraction.
[0848] An "oral image" is visual data that captures the inside of a subject's mouth in detail, and is used to analyze the condition of their teeth and gums.
[0849] "Preprocessing" refers to processes that include adjusting the resolution of image data and removing noise, and is preparatory work to facilitate subsequent data analysis.
[0850] "Deep learning technology" is an information processing technology that uses multi-layered artificial neural networks to achieve advanced pattern recognition and feature extraction.
[0851] A "generative model" is a technique that generates new data using methods such as generative adversarial networks and variational autoencoders, and is used for prediction and simulation.
[0852] An "orthodontic treatment scenario" is a future treatment plan that assumes different orthodontic methods and treatment flows, and serves as a reference for users to make their choices.
[0853] "Emotional analysis technology" is a technology that detects and analyzes a user's emotional state from images and audio, and is used to provide personalized information.
[0854] The system for carrying out this invention uses a smart device to collect images of the face and oral cavity and transmit them to a cloud server. When a user uploads images through the device, the server preprocesses the received images. This preprocessing adjusts the image resolution and removes noise to prepare the images for analysis.
[0855] Next, the server utilizes deep learning techniques to extract facial and oral features from the preprocessed images. Specifically, it uses a convolutional neural network (CNN) to identify facial and oral feature points. This data then forms the basis for future predictions by subsequent generative models.
[0856] Based on the extracted features, the server predicts the future appearance of the oral cavity and face using generative adversarial networks (GANs) and variational autoencoders (VAEs). This makes it possible to provide users with data that allows them to compare and consider different orthodontic treatment scenarios.
[0857] Furthermore, the server references accumulated past treatment data and analyzes similar cases. By using a transformer model, past knowledge is reflected in current predictions, improving accuracy.
[0858] Furthermore, it utilizes emotion analysis technology to detect the user's emotional state in real time. Based on the obtained emotional data, the system optimizes how prediction results are presented, providing the user with the most reassuring information. For example, if the user is feeling anxious, it will provide more detailed explanations or additional support information.
[0859] For example, if a user asks a question such as, "I want to know the difference between full or partial orthodontic treatment," the system will generate scenarios for each, and the device will provide a customized explanation that reflects the sentiment analysis results.
[0860] An example of a prompt message for a generative AI model would be: "I have uploaded a photo. Please predict future changes in teeth alignment and facial features and create a scenario for orthodontic treatment. Also, since the user appears anxious, please include detailed explanations and reassuring information."
[0861] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0862] Step 1:
[0863] The user takes images of their face and mouth using a device and uploads them to the server. The input here is the captured image data, and the output is the transmission of the image data to the server.
[0864] Step 2:
[0865] The server performs preprocessing on the received image data. The input is uploaded facial and oral cavity image data, and the server processes the data, such as adjusting the resolution and removing noise, to output image data that is ready for analysis.
[0866] Step 3:
[0867] The server applies deep learning techniques to preprocessed image data to extract facial and oral features. Specifically, it uses a convolutional neural network (CNN) to extract feature points. The input is preprocessed image data, and the output is the extracted feature data.
[0868] Step 4:
[0869] The server uses a generative model to predict future oral and facial features based on extracted feature data. Generative adversarial networks (GANs) and variational autoencoders (VAEs) are used here. The input is feature data, and the server generates and outputs simulation data for future predictions.
[0870] Step 5:
[0871] The server references past treatment data and uses a transformer model to analyze similar cases. This provides insights to improve prediction accuracy. The input is past treatment data, and the output is enhanced prediction data.
[0872] Step 6:
[0873] The server uses emotion analysis technology to analyze the user's emotional state in real time. The input is data on the user's current emotional state (obtained from facial and voice), and the server outputs data to optimize the presentation method based on the emotional information.
[0874] Step 7:
[0875] The terminal displays predictive data received from the server through a user interface. Here, the information is presented using optimized data tailored to the user's emotional state. Input consists of predictive data and sentiment analysis data, and the terminal outputs customized information.
[0876] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0877] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0878] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0879] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0880] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0881] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0882] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0883] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0884] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0885] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0886] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0887] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0888] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0889] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0890] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0891] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0892] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0893] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0894] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0895] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0896] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0897] The following is further disclosed regarding the embodiments described above.
[0898] (Claim 1)
[0899] A means for receiving a child's facial photograph and dental images, and pre-processing them according to predetermined standards,
[0900] A means for extracting facial and dental feature points from preprocessed images using a convolutional neural network,
[0901] A means for simulating future tooth alignment and facial features using a generative adversarial network and a variational autoencoder based on extracted features,
[0902] A means of making predictions based on similar cases using accumulated past treatment data,
[0903] A means for displaying simulation results generated based on multiple orthodontic treatment scenarios on a user interface,
[0904] A system that includes this.
[0905] (Claim 2)
[0906] The system according to claim 1, which responds in real time to user input and questions through an interactive interface based on the provided simulation results.
[0907] (Claim 3)
[0908] The system according to claim 1, which updates a machine learning model based on feedback from experts to improve its accuracy.
[0909] "Example 1"
[0910] (Claim 1)
[0911] A device that receives human facial and oral images and performs preprocessing on them based on pre-set criteria,
[0912] A device that extracts facial and oral feature points from pre-processed images using a machine learning algorithm,
[0913] A device that simulates future oral structures and facial shapes using advanced generation algorithms based on extracted features,
[0914] A device that uses stored historical medical information to make predictions based on similar cases,
[0915] A device that displays simulation results generated based on various orthodontic treatment methods on a human-interactive device,
[0916] A system that includes this.
[0917] (Claim 2)
[0918] The system according to claim 1, which responds immediately to human input and questions via an interactive device based on the provided simulation results.
[0919] (Claim 3)
[0920] The system according to claim 1, which improves the algorithm and accuracy based on feedback from professionals.
[0921] "Application Example 1"
[0922] (Claim 1)
[0923] A means for receiving a child's facial photograph and dental images, and pre-processing them according to predetermined standards,
[0924] A means for extracting facial and dental feature points from preprocessed images using a convolutional neural network,
[0925] A means for simulating future tooth alignment and facial features using a generative adversarial network and a variational autoencoder based on extracted features,
[0926] A means of making predictions based on similar cases using accumulated past treatment data,
[0927] A means for displaying simulation results generated based on multiple orthodontic treatment scenarios on a user interface,
[0928] A means for acquiring the driver's gaze and posture data with a camera device and extracting visual features from the pre-processed images,
[0929] A means of predicting future gaze patterns based on extracted visual features,
[0930] A means of displaying driver assistance information based on predicted eye movements,
[0931] A system that includes this.
[0932] (Claim 2)
[0933] The system according to claim 1, which responds in real time to user input and questions through an interactive interface based on the provided simulation results.
[0934] (Claim 3)
[0935] The system according to claim 1, which updates a machine learning model based on feedback from experts to improve its accuracy.
[0936] "Example 2 of combining an emotion engine"
[0937] (Claim 1)
[0938] A means for receiving a child's facial photograph and dental images, and pre-processing them according to predetermined standards,
[0939] A means for extracting facial and dental feature points from preprocessed images using a convolutional neural network,
[0940] A means for simulating future tooth alignment and facial features using a generative adversarial network and a variational autoencoder based on extracted features,
[0941] A means of making predictions based on similar cases using accumulated past treatment data,
[0942] A means for displaying simulation results generated based on multiple orthodontic treatment scenarios on a communication terminal, and for optimizing the presentation method of the results while taking into account the user's emotional state,
[0943] A system that includes this.
[0944] (Claim 2)
[0945] The system according to claim 1, which, based on the provided simulation results, responds in real time to user input and questions through an interactive interface to support user understanding.
[0946] (Claim 3)
[0947] The system according to claim 1, which updates a machine learning model based on feedback from experts and evaluations from sentiment analysis to improve accuracy and user experience.
[0948] "Application example 2 when combining with an emotional engine"
[0949] (Claim 1)
[0950] A means for receiving images of the face and oral cavity and performing preprocessing thereon according to predetermined standards,
[0951] A means for extracting facial and oral features from pre-processed images using deep learning technology,
[0952] A means of predicting future oral and facial features using a generative model based on extracted features,
[0953] A means of making predictions based on similar cases using accumulated past treatment data,
[0954] A means for displaying prediction results generated based on multiple orthodontic treatment scenarios on a display device,
[0955] A means for detecting the user's emotional state and optimizing the presentation method based on emotion analysis technology,
[0956] A system that includes this.
[0957] (Claim 2)
[0958] The system according to claim 1, which responds in real time to user input and questions via an interactive platform based on the provided prediction results.
[0959] (Claim 3)
[0960] The system according to claim 1, which updates the machine learning device based on feedback from experts to improve accuracy. [Explanation of Symbols]
[0961] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for receiving a child's facial photograph and dental images, and pre-processing them according to predetermined standards, A means for extracting facial and dental feature points from preprocessed images using a convolutional neural network, A means for simulating future tooth alignment and facial features using a generative adversarial network and a variational autoencoder based on extracted features, A means of making predictions based on similar cases using accumulated past treatment data, A means for displaying simulation results generated based on multiple orthodontic treatment scenarios on a user interface, A system that includes this.
2. The system according to claim 1, which responds in real time to user input and questions through an interactive interface based on the provided simulation results.
3. The system according to claim 1, which updates a machine learning model based on feedback from experts to improve its accuracy.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A