System for monitoring and improving brushing technique of user during brushing operation
A sensor-free system using machine learning to analyze hand images and predict toothbrush position in the mouth addresses the limitations of sensor-based systems, enhancing brushing technique and oral hygiene.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- KONINKLIJKE PHILIPS NV
- Filing Date
- 2025-11-11
- Publication Date
- 2026-05-21
AI Technical Summary
Existing toothbrush monitoring systems using gyroscopes and accelerometers are cumbersome, costly, and prone to errors, making them undesirable for consumers.
A system utilizing machine learning models to analyze images of a user's hand holding a toothbrush, extracting key points, and predicting the toothbrush's position in the mouth to provide real-time feedback on brushing technique without the need for sensors.
Enables effective toothbrushing by providing detailed feedback on brushing coverage and duration, improving oral hygiene without the drawbacks of sensor-based systems.
Smart Images

Figure EP2025082548_21052026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR MONITORING AND IMPROVING BRUSHING TECHNIQUE OF USER DURING BRUSHING OPERATIONFIELD OF THE INVENTION
[0001] The invention relates to the field of dental hygiene, and more specifically to monitoring toothbrush position and brushing technique of a user and providing instructions to the user for improving the same.BACKGROUND OF THE INVENTION
[0002] Common methods for determining positions of an object involve use of position sensors, such as gyroscopes and accelerometers, which play a crucial role in capturing and measuring movement, orientation, and acceleration of the object. Generally, gyroscopes provide information about an object’s angular velocity and rotation, and accelerometers measure linear acceleration along various axes. Combining data from these sensors enables tracking of the object’s position and orientation in three-dimensional space.
[0003] However, such sensors do have limitations, such as accumulating errors over time and sensitivity to external disturbances. Also, use of such sensors in handheld personal care devices, in particular, such as toothbrushes, increases the overall weight of these devices, making them more cumbersome and difficult to maneuver. Further, the sensors increase overall cost, making the personal care devices less desirable to consumers.SUMMARY OF THE INVENTION
[0004] According to a representative embodiment, a method is provided for improving brushing technique of a user during a brushing operation. The method includes receiving multiple images of a hand of the user holding a toothbrush during the brushing operation; extracting multiple key points on the user’s hand from each image of the multiple images and converting the key points to position data using a trained first machine learning model; predicting a position of the toothbrush in a mouth of the user in real time using a trained second machine learning model based on the position data, where the predicted position of the toothbrush in the user’s mouthcorresponds to a predetermined section of multiple predetermined sections of the user’s mouth; and providing feedback to the user regarding brushing technique based on the predicted position of the toothbrush in the user’s mouth.
[0005] According to another representative embodiment, a system is provided for monitoring and improving brushing technique of a user during a brushing operation of teeth of the user. The system includes a camera configured to acquire a multiple images of a hand of the user holding a toothbrush during the brushing operation; at least one processor coupled to the camera; and a non-transitory memory storing instructions that, when executed by the at least one processor, cause the at least one processor to receive image data from the multiple images acquired by the camera during the brushing operation; extract multiple key points on the user’s hand from the image data for each image of the multiple images, and convert the multiple key points to position data using a trained first machine learning model; predict a position of the toothbrush in a mouth of the user in real time using a trained second machine learning model based on the position data, where the predicted position of the toothbrush in the user’s mouth corresponds to a predetermined section of multiple predetermined sections of the user’s mouth; and provide feedback to the user regarding brushing technique based on the predicted position of the toothbrush in the user’s mouth.
[0006] According to another representative embodiment, a non-transitory computer readable medium stores instructions for monitoring and improving brushing technique of a user during a brushing operation of teeth of the user. When executed by at least one processor, the instructions cause the at least one processor to receive image data from multiple images acquired by a camera of a hand of the user holding a toothbrush during the brushing operation; extract multiple key points on the user’ s hand from the image data for each image of the multiple images, and convert the multiple key points to position data using a trained first machine learning model; predict a position of the toothbrush in a mouth of the user in real time using a trained second machine learning model based on the position data, where the predicted position of the toothbrush in the user’s mouth corresponds to a predetermined section of multiple predetermined sections of the user’s mouth; and provide feedback to the user regarding brushing technique based on the predicted position of the toothbrush in the user’s mouth.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The example embodiments are best understood from the following detailed description when read with the accompanying drawing figures. It is emphasized that the various features are not necessarily drawn to scale. In fact, the dimensions may be arbitrarily increased or decreased for clarity of discussion. Wherever applicable and practical, like reference numerals refer to like elements.
[0008] FIG. l is a simplified block diagram of a system for improving brushing technique of a user during a brushing operation, according to a representative embodiment.
[0009] FIG. 2 is a flow diagram of a method for extracting key points in image data from images of the user’s hand, and converting the key points to position data at inference of a trained first machine learning model, according to a representative embodiment.
[0010] FIG. 3 is an image showing the key points on the hand of the user detected by the first machine learning model, according to a representative embodiment.
[0011] FIG. 4 is a simplified block diagram showing layers of the trained second machine learning model, according to a representative embodiment.
[0012] FIG. 5 is a flow diagram of a method for predicting a position of the toothbrush in the user’s mouth in real time at inference of a trained second machine learning model, according to a representative embodiment.
[0013] FIG. 6A is an image of a user showing position data of key points on the user’s hand as detected by the first machine learning model, where the position data indicates the toothbrush positioned in the right section of the user’s mouth as predicted by the second machine learning model, according to a representative embodiment.
[0014] FIG. 6B is an image of a user showing position data of key points on the user’s hand as detected by the first machine learning model, where the position data indicates the toothbrush positioned in the center section of the user’s mouth as predicted by the second machine learning model, according to a representative embodiment.
[0015] FIG. 6C is an image of a user showing position data of key points on the user’s hand as detected by the first machine learning model, where the position data indicates the toothbrush positioned in the left section of the user’s mouth as predicted by the second machine learning model, according to a representative embodiment.
[0016] FIG. 7 is a flow diagram of a method of monitoring and improving brushing technique of a user during a brushing operation, according to a representative embodiment.DETAILED DESCRIPTION OF EMBODIMENTS
[0017] Aspects of the disclosure may be supported by various information technology (IT) backends, including either or both local architectures, either as monoliths, networked, or a combination thereof, and hosted architectures, such as a software as a service (SaaS), platform as a service (PaaS), and / or infrastructure as a service (laaS), or the like. In an example, a supporting infrastructure includes multiple interconnected layers respectively hosting, as an abstraction, various IT processes, services, accounts, and other management components.
[0018] Any of the steps described in relation to examples and / or training described below can be performed by a specific-purpose computer system or general-purpose computer system, or a computer-readable medium, or data carrier system configured to carry out any of the steps described previously. The computer system can include a set of software instructions that can be executed to cause the computer system to perform any of the methods or computer-based functions disclosed herein. The computer system may operate as a standalone device or may be connected, for example using a network, to other computer systems or peripheral devices. As an example, a computer system performs logical processing based on digital signals received via an analogue-to-digital converter.
[0019] In the following detailed description, for the purposes of explanation and not limitation, representative embodiments disclosing specific details are set forth in order to provide a thorough understanding of an embodiment according to the present teachings. Descriptions of known systems, devices, materials, methods of operation and methods of manufacture may be omitted so as to avoid obscuring the description of the representative embodiments. Nonetheless, systems, devices, materials and methods that are within the purview of one of ordinary skill in the art are within the scope of the present teachings and may be used in accordance with the representative embodiments. It is to be understood that the terminology used herein is for purposes of describing particular embodiments only and is not intended to be limiting. The defined terms are in addition to the technical and scientific meanings of the defined terms as commonly understood and accepted in the technical field of the present teachings.
[0020] It will be understood that, although the terms first, second, third, etc. may be used herein to describe various elements or components, these elements or components should not be limited by these terms. These terms are only used to distinguish one element or component from another element or component. Thus, a first element or component discussed below could be termed a second element or component without departing from the teachings of the inventive concept.
[0021] The terminology used herein is for purposes of describing particular embodiments only and is not intended to be limiting. As used in the specification and appended claims, the singular forms of terms “a,” “an” and “the” are intended to include both singular and plural forms, unless the context clearly dictates otherwise. Additionally, the terms “comprises,” “comprising,” and / or similar terms specify the presence of stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0022] Unless otherwise noted, when an element or component is said to be “connected to,” “coupled to,” or “adjacent to” another element or component, it will be understood that the element or component can be directly connected or coupled to the other element or component, or intervening elements or components may be present. That is, these and similar terms encompass cases where one or more intermediate elements or components may be employed to connect two elements or components. However, when an element or component is said to be “directly connected” to another element or component, this encompasses only cases where the two elements or components are connected to each other without any intermediate or intervening elements or components.
[0023] The present disclosure, through one or more of its various aspects, embodiments and / or specific features or sub-components, is thus intended to bring out one or more of the advantages as specifically noted below. For purposes of explanation and not limitation, example embodiments disclosing specific details are set forth in order to provide a thorough understanding of an embodiment according to the present teachings. However, other embodiments consistent with the present disclosure that depart from specific details disclosed herein remain within the scope of the appended claims. Moreover, descriptions of well-known apparatuses and methods may be omitted so as to not obscure the description of the exampleembodiments. Such methods and apparatuses are within the scope of the present disclosure.
[0024] The various embodiments described herein provide a system, method and computer readable medium storing instructions for calculating brushing coverage during a brushing operation without the use of sensors, such as accelerometers and gyroscopes, and providing detailed feedback to the user regarding brushing technique to ensure complete brushing coverage of the user’s teeth. Rather, machine learning models have been incorporated to eliminate the need for such sensors, thereby providing a technical solution to the problem posed by the cumbersome and inefficient conventional means of monitoring brushing technique. The embodiments are user-friendly because they do not require any sensor attachments in the brush head of the toothbrush, otherwise required in conventional systems. Also, the product is less costly and easier to maneuver.
[0025] Generally, an image-based machine learning model is used to detect key points on the hand of a user holding a toothbrush during a brushing operation from images of the user’s hand, where the images are captured during the brushing operation using a camera. The image-based machine learning model further converts the key points into a comprehensive dataset, e.g., utilizing respective pixel positions of the key points in two-dimensional space. The comprehensive dataset is fed into a feed forward machine learning model, which accurately predicts the position of the toothbrush within the user’s mouth during the brushing operation based on the comprehensive dataset. This enables feedback regarding the user’s brushing technique in real time. The feedback may include indicating the amount of time spent brushing different sections of the user’s mouth and overall brushing duration, as well as specifically identifying sections of the user’s mouth in which the teeth have not been brushed enough. The various embodiments improve oral health and hygiene by helping users brush their teeth more effectively, and educating users about proper brushing techniques.
[0026] FIG. l is a simplified block diagram of a system for improving brushing technique of a user during a brushing operation, according to a representative embodiment.
[0027] Referring to FIG. 1, system 100 includes a processing unit 105 for implementing and / or managing the processes described herein with regard to improving brushing technique. The processing unit 105 includes one or more processors indicated by processor 120, one or more memories indicated by memory 130, a user interface 122 and a display 124. The system 100further includes a camera 140 configured to obtain images of a hand 150 of a user holding a toothbrush 155 during the brushing operation, and a speaker 145 configured to provide audible information and / or instructions to the user as feedback during the brushing operation. The camera 140 interfaces with the processing unit 105 through a known imaging interface (not shown) to provide the images to the processor 120 and the memory 130. The camera 140 may be an RGB or an RGB-D video camera, for example, that provides image data and depth information for objects within its field of view. The camera 140 may be a still frame camera without departing from the scope of the present teachings. The camera 140 must be arranged in a position and orientation that enables it to capture unobstructed images of the user’s hand 150 holding the toothbrush 155 while the user is performing the brushing operation.
[0028] In various embodiments, the system 100 may be implemented using a cell phone, for example, where the user has downloaded an application (or app) directed to improving brushing technique that is loaded in the memory of the cell phone (memory 130). The cell phone may include an iOS operating system or an Android operating system, for example, run by the processor 120. The downloaded application enables access to the cell phone camera as the camera 140. The process for improving brushing technique may be included as part of a broader suite of applications for improving oral hygiene, such as Sonicare® available from Philips Koninklijke Philips N.V., for example.
[0029] The memory 130 stores instructions executable by the processor 120. When executed, the instructions cause the processor 120 to implement one or more processes for performing a process for monitoring tooth brushing and providing feedback to the user in real time for improving the brushing technique. For purposes of illustration, the memory 130 is shown to include software modules, each of which includes a set of instructions, executable by the processor 120, corresponding to an associated capability of the system 100.
[0030] The processor 120 is representative of one or more processing devices, and may be implemented by a general purpose computer, a central processing unit (CPU), a digital signal processor (DSP), a graphical processing unit, a computer processor, a microprocessor, a state machine, programmable logic device, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), or combinations thereof, using any combination of hardware, software, firmware, hard-wired logic circuits, or combinations thereof. The processor120 may include a single processor, multiple processors, and / or parallel processors. Multiple processors may be included in, or coupled to, a single device or multiple devices. The term “processor” as used herein encompasses an electronic component able to execute a program or machine executable instruction. The processor 120 may also refer to a collection of processors within a single computer system or distributed among multiple computer systems, such as in a cloud-based or other multi-site application. Programs have software instructions performed by one or multiple processors that may be within the same computing device or which may be distributed across multiple computing devices.
[0031] The memory 130 may include main memory and / or static memory, where such memories may communicate with each other and the processor 120 via one or more buses. The memory 130 may be implemented by any number, type and combination of random access memory (RAM) and read-only memory (ROM), for example, and may store various types of information, such as software algorithms, artificial intelligence (Al) machine learning models, and computer programs, all of which are executable by the processor 120. The various types of ROM and RAM may include any number, type and combination of computer readable storage media, such as a disk drive, flash memory, an electrically programmable read-only memory (EPROM), an electrically erasable and programmable read only memory (EEPROM), registers, a hard disk, a removable disk, tape, compact disk read only memory (CD-ROM), digital versatile disk (DVD), floppy disk, Blu-ray disk, a universal serial bus (USB) drive, or any other form of storage medium. The memory 130 is a tangible storage medium for storing data and executable software instructions, and is non-transitory during the time software instructions are stored therein. As used herein, the term “non-transitory” is to be interpreted not as an eternal characteristic of a state, but as a characteristic of a state that will last for a period. The term “non-transitory” specifically disavows fleeting characteristics such as characteristics of a carrier wave or signal or other forms that exist only transitorily in any place at any time. The memory 130 may store software instructions and / or computer readable code that enable performance of various functions. The memory 130 may be secure and / or encrypted, or unsecure and / or unencrypted.
[0032] The system 100 may also include a database 112 for storing information that may be used by the various software modules of the memory 130. For example, the database 112 may include image data from previously obtained images of users performing brushing operations, whichmay be used for training artificial intelligence (Al) machine learning models, such as a neural network model, for example, as discussed below. The database 112 may be implemented by any number, type and combination of RAM and ROM, for example. The various types of ROM and RAM may include any number, type and combination of computer readable storage media, such as a disk drive, flash memory, EPROM, EEPROM, registers, a hard disk, a removable disk, tape, CD-ROM, DVD, floppy disk, Blu-ray disk, USB drive, or any other form of storage medium known in the art. The database 112 comprises a tangible storage medium for storing data and executable software instructions and is non-transitory during the time data and software instructions are stored therein. The database 112 may be secure and / or encrypted, or unsecure and / or unencrypted. For purposes of illustration, the database 112 is shown as a separate storage medium, although it is understood that it may be combined with and / or included in the memory 130, without departing from the scope of the present teachings.
[0033] The processor 120 may include or have access to an Al engine, which may be implemented as software to provide artificial intelligence (e.g., deep learning, neutral network models) and applies machine learning described herein. The Al engine may reside in any of various components in addition to or other than the processor 120, such as the memory 130, an external server, and / or the cloud, for example. When the Al engine is implemented in a cloud, such as at a data center, for example, the Al engine may be connected to the processor 120 via the internet or other communication network using one or more wired and / or wireless connection(s). In various embodiments, all or part of the processes provided by first and second machine learning models, discussed below, may be implemented by the Al engine, for example. The training and execution of the first and second machine learning models cannot practically be performed in the human mind.
[0034] The user interface 122 is configured to provide information and data output by the processor 120, the memory 130 and / or the camera 140 to the user and / or for receiving information and data input by the user. That is, the user interface 122 enables the user to enter data and to control or manipulate aspects of the processes described herein, and also enables the processor 120 to indicate the effects of the user’s input, which may include control or manipulation of the camera 140 and / or the speaker 145. All or a portion of the user interface 122 may be implemented by a graphical user interface (GUI), such as GUI 128 viewable on thedisplay 124, discussed below. The user interface 122 may include one or more interface devices, such as a touchpad, a touchscreen, a keypad, a keyboard, a mouse, a trackball, a joystick, a microphone, a video camera, or voice or gesture recognition captured by a microphone or video camera, for example.
[0035] The display 124 may be any type of compatible display, such as a cell phone screen, a computer monitor, a liquid crystal display (LCD), an organic light emitting diode (OLED), a flat panel display, or a solid-state display, for example. The display 124 includes a screen 126 for viewing images of the user’s hand 150 holing the toothbrush 155, as well as the GUI 128 to enable the user to interact with the displayed images and features.
[0036] Referring to the memory 130, the various modules store sets of data and instructions executable by the processor 120 to monitor tooth brushing and provide feedback to the user in for improving the brushing technique. The memory 130 includes imaging module 131, which is configured to receive and process images acquired by the camera 140 to provide corresponding image data, where the images may be image frames when the camera 140 is a video camera. The images show the user’s hand 150 holding a toothbrush 155 throughout a brushing operation by the user. The images are received in real time or near real time from the camera 140, e.g., during a contemporaneous brushing operation by the user. The images may be displayed on the display 124, if desired, although this is not required.
[0037] Key point extraction module 132 is configured to automatically extract multiple key points on the user’s hand 150 from the image data for each image, and to convert the key points to position data. The position data may be provided in a text formatted dataset, such as a comma separated values (csv) dataset, for example. Generally, the key point extraction module 132 provides an image-based, first machine learning model that receives the image data from the imaging module 131 as input, provides position data for a set of a predetermined number of key points and converts the key points to the text formatted dataset as output, as further discussed below. The position data may indicate the key points in two-dimensional space, which may be provided in a text formatted dataset, for example. The first machine learning model may be implemented as any suitable type of trainable machine learning model (e.g., deep learning model), such as a convolutional neural network (CNN), an artificial neural network (ANN), a recurrent neural network (RNN), a vision transformer, or a U-net model, for example. A deeplearning model refers to a neural network model with large numbers of layers and / or parameters, which directly enable tasks such as classification and regression, for example, as would be apparent to one skilled in the art.
[0038] In an embodiment, the first machine learning model may be implemented using MediaPipe, for example, available as an open-source framework, as would be apparent to one skilled in the art. MediaPipe is able to detect landmarks of various items in images, such as hands and faces. The MediaPipe hand detection in particular is able to handle various hand shapes, sizes, positions, and orientations, making it highly versatile and adaptable to diverse applications. In this embodiment, the predetermined number of key points on the user’s hand in one set may consist of 21 key points to be extracted from the image data. The 21 key points may be identified using a previously populated library by MediaPipe or other previously populated library. These key points are converted to a text formatted dataset using their pixel positions in the corresponding image. Accordingly, a total of 42 values (i.e., x, y coordinates for the 21 key points) will be used to identify one position of the user’s hand 150. The MediaPipe framework may be integrated with iOS and Android operating systems, as it is supported by these platforms.
[0039] The first machine learning model is previously trained in a supervised fashion using a first training dataset, which includes labelled image data from thousands of images of hands of different human beings in various positions relative to the camera. Each of the images is labeled to show all of the key points on the hand visible in that image, where the key points correspond to predetermined locations on the anatomy of the hand and are labeled accordingly. For example, the key points may correspond to the tips, joints and the bases of all five fingers, and the base of the palm, respectively, as shown in FIG. 3, discussed below. The labels may be added by experts in human anatomy, or may be labeled automatically through techniques such as ground truth automation, as would be apparent to one skilled in the art. The labeled first training data set may be referred to as ground truth data. The first training data set for training the first machine learning model may be stored in the database 112, for example.
[0040] After being trained on the first training data set, the first machine learning model is able to process new images, identify a predetermined number of key points based on the same, and convert the key points to position data. FIG. 2 is a flow diagram of a method for extracting key points in image data from images of the user’s hand 150, and converting the key points to a textformatted dataset at inference of the trained first machine learning model, according to a representative embodiment.
[0041] Referring to FIG. 2, image data from images of the user’s hand 150 acquired by the camera 140 is input to the trained first machine learning model in block S311. The images are acquired in real time during a brushing operation performed by the user.
[0042] In block S212, the trained first machine learning model searches the image data of each image for a set of key points on the user’ s hand. . Each set of key points includes a predetermined number of key points at corresponding predetermined locations on the anatomy of the user’s hand. When trained first machine learning model is implemented using MediaPipe, for example, the predetermined number of key points in each set of key points is 21, as discussed above.
[0043] In block S213, the trained first machine learning model converts the key points in each set of key points into position data in two-dimensional space. As mentioned above, the position data may be provided in a text formatted dataset, such as a csv dataset, for example. For example, the first machine learning model may convert the key points to the position data by identifying pixel positions of the key points, respectively, in the two-dimensional space of the camera 140, and translating the pixel positions into corresponding text formatted data.
[0044] In block S214, the trained first machine learning model outputs the position data in the text formatted dataset. The position data identifies the key points in the set of key points in each of the images of the user’s hand 150.
[0045] FIG. 3 is an image showing the key points on the user’s hand detected by the first machine learning model, according to a representative embodiment. Referring to FIG. 3, the image of the user’s hand 150 shows 21 key points labeled 0 to 20 in specific predetermined locations, including the tips, joints and bases of all five fingers, and the base of the palm. The 21 key points are converted by the first machine learning model to respective pixel positions 0’ to 20’ in two-dimensional space, as shown.
[0046] Referring again to FIG. 1, brush position prediction module 133 is configured to predict (estimate) the position of the toothbrush 155 in the user’s mouth in real time, based on the position data of the user’s hand. Generally, the brush position prediction module 133 provides a feed-forward, second machine learning model that receives the position data of key pointpositions from the key point extraction module 132 as input, and provides a predicted position of the toothbrush 155 in the user’s mouth as output, as discussed below. The predicted position of the toothbrush 155 corresponds to one section of multiple predetermined sections of the user’s mouth. For example, the user’s mouth may be divided into three predetermined sections, including a left section, a right section and a center section, as further discussed below. The second machine learning model may be implemented as any suitable type of trainable machine learning model (e.g., deep learning model), such as a CNN, an ANN, an RNN, a vision transformer, or a U-net model, for example.
[0047] The second machine learning model is previously trained in a supervised fashion using image data from thousands of training images of left and right hands of different individuals holding toothbrushes during brushing operations, respectively, in various positions relative to the camera. The individuals may have varying complexions, heights, and genders to provide breadth to the training data. The training images are converted into a second training dataset by the first machine learning model (e.g., implemented using MediaPipe), which outputs a text formatted dataset, such as a csv file, for example, for training the second machine learning model to find the position of the brush relative to our mouth. That is, the first machine leaning model is used to extract a set of key points on the user’ s hand in each of training images in a text formatted dataset. Each extracted set of key points is classified according to the section of multiple predetermined sections of the user’s mouth in which the tooth brush is located, and whether the user is using the right or left hand. Classifying the extracted sets of key points in the second dataset may be done by labeling the sets of key points accordingly for each image. The labels may be added by experts in machine learning technology and / or anatomy, or may be labeled automatically through techniques such as ground truth automation, as would be apparent to one skilled in the art. The labeled second training data set may be referred to as ground truth data. The second training data set for training the second machine learning model may be stored in the database 112, for example. In an embodiment, the second machine learning model may use mean squared error loss as the loss function, and rectified linear unit (ReLU) and SoftMax as the activation functions, for example, during training, as discussed below.
[0048] In an embodiment, different versions of the second machine learning model may be assembled, and the classified sets of key points may be applied to the versions of the secondmachine learning models, respectively. Performances of the different versions of the second machine learning model may then be evaluated, respectively, in response to the applied classified sets of key point. The different versions of the second machine learning model may be evaluated based on precision, recall and / or fl scores, for example, as would be apparent to one skilled in the art. One of the versions of the second trained machine learning model is then selected as the second trained machine learning model based on the evaluated performances of the same.
[0049] FIG. 4 is a simplified block diagram showing layers of the trained second machine learning model, according to a representative embodiment. In an embodiment, a TensorFlow platform, for example, available in PYTHON®, may be used to build the second machine learning model as shown in FIG. 4.
[0050] Referring to FIG. 4, second machine learning model 400 is a neural network, such as a CNN, an ANN, an RNN, a vision transformer, or a U-net model, for example, that includes multiple layers, indicated by representative input layer 411, first hidden layer 412, second hidden layer 413, and output layer 414. The input layer 411 is configured to take values from a text formatted dataset 410 output by the first machine learning model, where the text formatted dataset 410 provides position data of the predetermined number of key points extracted from each image. The first column in the text formatted dataset 410 includes identifiers (IDs) of 21 key points, labeled 0 to 20, as shown in FIG. 3, for example. The second column in the text formatted dataset 410 includes x and y pixel coordinates of the position data for each of the key points, where the x coordinate is the top value and the y coordinate is the bottom value associated with each of the key points. So, for example, the x and y pixel coordinates of the position data for key point 3 are -0.5162 and -0.5704, respectively.
[0051] In the depicted example, the first hidden layer 412 of the trained second machine learning model 420 has 20 nodes and the second hidden layer 413 has 10 nodes. Dropout of 0.2 is applied to the first hidden layer 412, and dropout of 0.4 is applied to the second hidden layer 413. Both the first and second hidden layers 412 and 413 may utilize the ReLU activation function, for example, mentioned above. ReLU is an open source non-linear mathematical activation function that outputs the input directly when the input is positive, and outputs zero when the input is negative, such that negative input values are not output. The output layer 414 has the same number of nodes as possible outcomes (i.e., different locations in the user’s mouth in which thetoothbrush may be positioned) by the second machine learning model 420. Assuming three possible classifications (e.g., left section, a right section and a center section of the user’s mouth), as discussed above, the output layer 414 likewise has three nodes. The output layer 414 may utilize the SoftMax activation function, mentioned above, for example. SoftMax is an open source non-linear mathematical function that takes vectors of arbitrary real values as input and transforms them into values between 0 and 1, which helps in multi class classification problems. The second machine learning model thus outputs the location of the toothbrush 155 from the output layer 414 as one of the left section, the right section or the center section of the user’s mouth from the corresponding one of the three nodes in the output layer 414. Of course, the second machine learning model may have more or fewer layers with more or fewer nodes, respectively, without departing from the scope of the present teachings.
[0052] After being trained on the second training data set, the second machine learning model is able to process new sets of key points extracted from the images and converted to position data by the first machine learning model and determine the predetermined section of the user’s mouth in which the toothbrush is located. FIG. 5 is a flow diagram of a method for predicting a position of the toothbrush in the user’ s mouth in real time at inference of the trained second machine learning model, according to a representative embodiment.
[0053] Referring to FIG. 5, the position data identifying the key points in the sets of key points in the images of the user’s hand 150, as acquired by the camera 140, are input to the trained second machine learning model in block S511. The position data may be provided in a text formatted dataset, such as a csv dataset, for example, as discussed above.
[0054] In block S512, for each image, the trained second machine learning model associates the position data of the set of key points with the most similar position data indicating a similar hand position.
[0055] In block S513, for each image, the trained second machine learning model predicts the position of the toothbrush in the user’s mouth in real time based on the associated position data indicating the most similar hand position. The predicted position of the toothbrush corresponds to one predetermined section of multiple predetermined sections of the user’s mouth. For example, the user’s mouth may be divided into three predetermined sections, including a left section, a right section and a center section, as discussed above. The center section includesupper and lower teeth in the center part (e.g., four front teeth or one third) of the user’s mouth, the left section includes upper and lower teeth to the left of the center section (e.g., left-side third) of the user’s mouth, and the right section includes upper and lower teeth to the right of the center section (e.g., right-side third) of the user’s mouth. The user’s mouth may be divided into more or fewer predetermined sections without departing from the scope of the present teachings. The relative positions of the key points on the user’s hand while brushing are assumed to be correlated while brushing in a specific section of the user’s mouth. Each section may include lingual, buccal and crowns of the corresponding teeth, for example.
[0056] In block S514, the trained second machine learning model outputs the predicted position of the toothbrush 155 in the user’s mouth. The predicted position identifies the section of the predetermined sections of the user’s mouth in which the toothbrush is located.
[0057] FIGs. 6A, 6B and 6C are images of a user showing position data of key points on the user’s hand as detected by the first machine learning model, where the position data indicates the toothbrush positioned in different sections of the user’s mouth as predicted by the second machine learning model, according to a representative embodiment.
[0058] In particular, FIG. 6 A shows an arrangement of first position data 611 (white dots) on the user’s hand 150. Based on the first position data, the second machine learning model predicts that the toothbrush is held in the left hand of the user and positioned in the right section of the user’s mouth. FIG. 6B shows an arrangement of second position data 612 (white dots) on the user’s hand 150. Based on the second position data, the second machine learning model predicts that the toothbrush is held in the left hand of the user and is positioned in the center section of the user’s mouth. FIG. 6C shows an arrangement of third position data 612 (white dots) on the user’s hand 150. Based on the third position data, the second machine learning model predicts that the toothbrush is held in the left hand of the user and is positioned in the left section of the user’s mouth.
[0059] User feedback module 134 is configured to determine feedback to be provided to the user regarding brushing technique based on the predicted positions of the toothbrush 155 in the predetermined sections of the user’s mouth, provided by the brush position prediction module 133. That is, the user feedback module 134 monitors the amount of time the toothbrush 155 is located in each section of the predetermined sections of the user’s mouth, and may also track thetotal amount of time of the user is engaged in the brushing operation. For example, the camera 140 may provide the processor 120 images of the brushing operation to be monitored at a certain frame frequency, and the processor 120 determines the amount of time the toothbrush 155 is in each section or position of the mouth by analyzing the frame frequency. For example, when the camera 140 is a video camera that captures live images at 30 frames per second, the processor 120 is able to determine time by counting frames and dividing by 30. So, if the camera 140 records the toothbrush 140 in the left section of the user’s mouth over 300 frames, the processor 120 calculates that the toothbrush 140 has been in the left section of the user’s mouth for 10 seconds and may notify the user accordingly. The user feedback module 134 is thus able to determine the amount of time spent brushing each predetermined section of the multiple predetermined sections of the user’s mouth and / or the overall brushing duration based on the total brushing time of all predetermined sections. The user feedback module 134 may output this information to the user via the display 124, the speaker 145, and / or other interface accessible by the user, such as haptic indications incorporated in in the toothbrush 155, for example.
[0060] The user feedback module 134 may provide additional feedback indicating insufficient brushing by the user for the predetermined sections of the user’s mouth, as well as instructions on how to correct any deficiencies. For example, the user feedback module 134 may compare the amount of time spent brushing each predetermined section of the user’s mouth with a predetermined minimum brushing time required for adequate cleaning. The minimum brushing time may be set by the user or stored in the memory 130, for example, prior to the brushing operation for use by the user feedback module 134. For example, the system 100 may store adequate minimum brushing times for brushing each of the predetermined sections of the user’s mouth based on relevant factors, such as age of the user and / or oral health conditions of the user. In this case, the user feedback module 134 is able to identify the appropriate minimum brushing time for the predetermined sections by accessing the relational database. When the user feedback module 134 detects that he user has not met the minimum brushing time for a particular predetermined section, it may initiate a notification or instructions to the user as feedback to inform the user of the insufficiencies and to direct the user to continue brushing in that predetermined section for an additional time to satisfy the minimum brushing time. For example, the instructions may include identifying at least one predetermined section of the user’s mouththat is not brushed enough, and / or directing the user specifically to return the toothbrush to a previous section of user’s mouth for a particular amount of time. Again, the feedback may be provided to the user via the display 124, the speaker 145, and / or other accessible interface.
[0061] FIG. 7 is a flow diagram of a method of monitoring and improving brushing technique of a user during a brushing operation, according to a representative embodiment. The method may be implemented at least in part using instructions stored in memory 130 and executable by the processor 120 in the system 100, for example.
[0062] Referring to FIG. 7, image data from multiple images of the user’s hand holding a toothbrush during the brushing operation are received in block S711. The images of the user’s hand are acquired in quick succession, e.g., by a video camera, in order to capture movement of the user’s hand throughout the brushing operation.
[0063] In block S712, a set of multiple key points on the user’s hand is extracted from the image data for each image of the multiple images, and the extracted key points are converted to position data using a trained first machine learning model. As discussed above with reference to FIG. 2, key points may be extracted using edge and shape detection, for example, to identify the key points in the anatomy of the user’s hand. The position data corresponding to the key points may be a text formatted dataset, such as a csv dataset, that provides the position data in two-dimensional space. In an embodiment, the two-dimensional space may be pixel positions of the key points in the image.
[0064] In block S713, a position of the toothbrush in the user’s mouth is predicted in real time using a trained second machine learning model based on the text formatted dataset. The predicted position of the toothbrush in the user’s mouth corresponds to a predetermined section of multiple predetermined sections of the user’s mouth. As discussed above with reference FIG.5, for each image, the position data of the key points are associated with the most similar position data in a second training data set, indicating a similar hand position. The position of the toothbrush in the user’s mouth is predicted based on the associated position data. The predicted position of the toothbrush corresponds to one predetermined section of multiple predetermined sections of the user’s mouth. For example, the user’s mouth may be divided into three predetermined sections, including a left section, a right section and a center section, as discussed above.
[0065] In block S714, feedback is provided to the user regarding the user’s brushing technique based on the predicted position of the toothbrush in the user’s mouth. The feedback includes indicating an amount of time spent brushing each predetermined section of the multiple predetermined sections of the user’s mouth and / or an overall brushing duration based on the total brushing time of all predetermined section. The feedback may further include indicating insufficient brushing by the user for the predetermined sections of the user’s mouth, notifying the user when an insufficient amount of time has been spent brushing one or more particular predetermined sections. The feedback may be based on monitoring the amount of time the toothbrush is predicted to be in each of the predetermined sections, as well as the total amount of time the toothbrush is in the user’s mouth, as discussed above. The monitoring is based on video images of the brushing operation and the output from the trained second machine learning model indicating when the toothbrush enters and leaves each predetermined section during the brushing operation. The feedback may be provided to the user during the brushing operation visually via a display and / or audibly via a speaker, for example, to enable the user to adjust and thereby improve their brushing technique in real time. Other forms of communicating the feedback may be incorporated without departing from the scope of the present teachings.
[0066] In an embodiment, the feedback may include comparing the predicted positions of the toothbrush in the user’s mouth to desired positions of the toothbrush and brushing durations in each section according to a predetermined brushing routine to determine whether the predicted positions of the toothbrush correspond to the desired positions and brushing durations, respectively. In this case, providing feedback to the user regarding the brushing technique may further include providing visual and / or audio instructions to the user to alter the current position of the toothbrush to the desired position of the toothbrush in the user’s mouth when the predicted position of the toothbrush is determined not to correspond to the desired position of the toothbrush in the user’s mouth, as discussed above. The visual and / or audio instructions may also include instructions to move to another section of the mouth when the time in the present section has been exceeded.
[0067] In accordance with various embodiments of the present disclosure, the methods described herein may be implemented using a hardware computer system that executes software programs stored on non-transitory storage mediums. Further, in an exemplary, non-limited embodiment,implementations can include distributed processing, component / object distributed processing, and parallel processing. Virtual computer system processing may implement one or more of the methods or functionalities as described herein, and a processor described herein may be used to support a virtual processing environment.
[0068] Although monitoring and improving brushing technique have been described with reference to exemplary embodiments, it is understood that the words that have been used are words of description and illustration, rather than words of limitation. Changes may be made within the purview of the appended claims, as presently stated and as amended, without departing from the scope and spirit of the embodiments. Also, although monitoring and improving brushing technique have been described with reference to particular means, materials and embodiments, it is not intended to be limited to the particulars disclosed; rather the description extends to all functionally equivalent structures, methods, and uses such as are within the scope of the appended claims.
[0069] The illustrations of the embodiments described herein are intended to provide a general understanding of the structure of the various embodiments. The illustrations are not intended to serve as a complete description of all of the elements and features of the disclosure described herein. Many other embodiments may be apparent to those of skill in the art upon reviewing the disclosure. Other embodiments may be utilized and derived from the disclosure, such that structural and logical substitutions and changes may be made without departing from the scope of the disclosure. Additionally, the illustrations are merely representational and may not be drawn to scale. Certain proportions within the illustrations may be exaggerated, while other proportions may be minimized. Accordingly, the disclosure and the figures are to be regarded as illustrative rather than restrictive.
[0070] One or more embodiments of the disclosure may be referred to herein, individually and / or collectively, by the term “invention” merely for convenience and without intending to voluntarily limit the scope of this application to any particular invention or inventive concept. Moreover, although specific embodiments have been illustrated and described herein, it should be appreciated that any subsequent arrangement designed to achieve the same or similar purpose may be substituted for the specific embodiments shown. This disclosure is intended to cover any and all subsequent adaptations or variations of various embodiments. Combinations of the aboveembodiments, and other embodiments not specifically described herein, will be apparent to those of skill in the art upon reviewing the description.
[0071] The Abstract of the Disclosure is provided to comply with 37 C.F.R. § 1.72(b) and is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, various features may be grouped together or described in a single embodiment for the purpose of streamlining the disclosure. This disclosure is not to be interpreted as reflecting an intention that the claimed embodiments require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter may be directed to less than all of the features of any of the disclosed embodiments. Thus, the following claims are incorporated into the Detailed Description, with each claim standing on its own as defining separately claimed subject matter.
[0072] The preceding description of the disclosed embodiments is provided to enable any person skilled in the art to practice the concepts described in the present disclosure. As such, the above disclosed subject matter is to be considered illustrative, and not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other embodiments which fall within the true spirit and scope of the present disclosure. Thus, to the maximum extent allowed by law, the scope of the present disclosure is to be determined by the broadest permissible interpretation of the following claims and their equivalents and shall not be restricted or limited by the foregoing detailed description.
Claims
CLAIMS:
1. A method of monitoring and improving brushing technique of a user during a brushing operation, the method comprising:receiving image data from a plurality of images of a hand of the user holding a toothbrush during the brushing operation (S711);extracting a plurality of key points on the user’ s hand from the image data for each image of the plurality of images, and converting the plurality of key points to position data using a trained first machine learning model (S712);predicting a position of the toothbrush in a mouth of the user in real time using a trained second machine learning model based on the position data, wherein the predicted position of the toothbrush in the user’s mouth corresponds to a predetermined section of a plurality of predetermined sections of the user’s mouth (S713); andproviding feedback to the user regarding brushing technique based on the predicted position of the toothbrush in the user’s mouth (S714).
2. The method of claim 1, wherein the feedback includes indicating one or more of (i) an amount of time spent brushing each predetermined section of the plurality of predetermined sections of the user’s mouth, (ii) an overall brushing duration, and (iii) insufficient brushing of each predetermined section of the plurality of predetermined sections of the user’s mouth.
3. The method of claim 1, wherein providing feedback comprises:comparing the predicted position of the toothbrush in the user’s mouth to a desired position of the toothbrush in the user’s mouth according to a predetermined brushing routine to determine whether the predicted position of the toothbrush in the user’s mouth corresponds to the desired position of the toothbrush in the user’s mouth, andproviding instructions to the user to alter a current position of the toothbrush to the desired position of the toothbrush in the user’s mouth when the predicted position of the toothbrush is determined to not correspond to the desired position of the toothbrush in the user’s mouth.
4. The method of claim 3, wherein the instructions to the user comprise at least one of returning the toothbrush to a previous position in the user’s mouth for a predetermined minimum amount of time, or identifying at least one predetermined section of the plurality of predetermined sections of the user’s mouth that is not brushed enough.
5. The method of claim 1, wherein the trained first machine learning model comprises a MediaPipe framework, and wherein the plurality of key points on the user’s hand is 21 key points extracted by the MediaPipe framework.
6. The method of claim 1, wherein the position data comprise a text formatted dataset.
7. The method of claim 6, wherein the text formatted dataset comprises a comma separated values (csv) dataset.
8. The method of claim 1, wherein converting the plurality of key points to the position data using the trained first machine learning model comprises identifying pixel positions of the key points in two-dimensional space, and translating the pixel positions into corresponding text.
9. The method of claim 1, wherein the plurality of predetermined sections of the user’s mouth comprise a left section, a right section, and a center section of the user’s mouth.
10. A system for monitoring and improving brushing technique of a user during a brushing operation of teeth of the user, the system comprising:a camera (1 0) configured to acquire a plurality of images of a hand of the user holding a toothbrush during the brushing operation;at least one processor (120) coupled to the camera; anda non-transitory memory (130) storing instructions that, when executed by the at least one processor, cause the at least one processor to:receive image data from the plurality of images acquired by the camera during the brushing operation (S711);extract a plurality of key points on the user’ s hand from the image data for each image of the plurality of images, and convert the plurality of key points to position data using a trained first machine learning model (S712);predict a position of the toothbrush in a mouth of the user in real time using a trained second machine learning model based on the position data, wherein the predicted position of the toothbrush in the user’s mouth corresponds to a predetermined section of a plurality of predetermined sections of the user’s mouth (S713); andprovide feedback to the user regarding brushing technique based on the predicted position of the toothbrush in the user’s mouth (S714).
11. The system of claim 10, further comprising:at least one of a display or a speaker coupled to the processor and configured to output the feedback to the user regarding the brushing technique.
12. The system of claim 11, wherein the feedback includes indicating via the at least one of the display or the speaker one or more of (i) an amount of time spent brushing each predetermined section of the plurality of predetermined sections of the user’s mouth, (ii) an overall brushing duration, and (iii) insufficient brushing of each predetermined section of the plurality of predetermined sections of the user’s mouth.
13. The system of claim 11, wherein the instructions cause the at least one processor to provide feedback by:comparing the predicted position of the toothbrush in the user’s mouth to a desired position of the toothbrush in the user’s mouth according to a predetermined brushing routine to determine whether the predicted position of the toothbrush in the user’ s mouth corresponds to the desired position of the toothbrush in the user’s mouth, andproviding instructions to the user via the at least one of the display or the speaker to alter a current position of the toothbrush to the desired position of the toothbrush in the user’s mouthwhen the predicted position of the toothbrush is determined to not correspond to the desired position of the toothbrush in the user’s mouth.
14. The system of claim 13, wherein the instructions to the user comprise at least one of returning the toothbrush to a previous position in the user’s mouth for a predetermined minimum amount of time, or identifying at least one predetermined section of the plurality of predetermined sections of the user’s mouth that is not brushed enough.
15. The system of claim 10, wherein the trained first machine learning model comprises a MediaPipe framework, and wherein the plurality of key points on the user’s hand is 21 key points extracted by the MediaPipe framework.
16. The system of claim 10, wherein the position data comprise a text formatted dataset.
17. The system of claim 16, wherein the text formatted dataset comprises a comma separated values (csv) dataset.
18. The system of claim 10, wherein converting the plurality of key points to the position data using the trained first machine learning model comprises identifying pixel positions of the key points in two-dimensional space, and translating the pixel positions into corresponding text.
19. The system of claim 10, wherein the plurality of predetermined sections of the user’s mouth comprise a left section, a right section, and a center section of the user’s mouth.
20. A non-transitory computer readable medium (130) storing instructions for monitoring and improving brushing technique of a user during a brushing operation of teeth of the user, wherein when executed by at least one processor, cause the at least one processor to:receive image data from a plurality of images acquired by a camera of a hand of the user holding a toothbrush during the brushing operation (S711);extract a plurality of key points on the user’ s hand from the image data for each image of the plurality of images, and convert the plurality of key points to position data using a trained first machine learning model (S712);predict a position of the toothbrush in a mouth of the user in real time using a trained second machine learning model based on the position data, wherein the predicted position of the toothbrush in the user’s mouth corresponds to a predetermined section of a plurality of predetermined sections of the user’s mouth (S713); andprovide feedback to the user regarding brushing technique based on the predicted position of the toothbrush in the user’s mouth (S714).26