Oral care
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KONINKLIJKE PHILIPS NV
- Filing Date
- 2023-06-12
- Publication Date
- 2026-04-17
AI Technical Summary
Existing oral care devices lack affordable sensors for monitoring and providing accurate feedback on brushing and flossing performance, limiting users' ability to improve their oral care routines effectively.
A computer vision-based method analyzes video data from conventional devices with cameras to track the movement of personal care devices and users, determining parameter values such as brushing time ratios and flossing completion using neural networks and machine learning algorithms, providing feedback without requiring expensive sensors.
Enables improved oral care by accurately assessing routine performance, reinforcing good habits and highlighting deficiencies, thus enhancing user awareness and quality of oral hygiene practices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of oral care routines, and more particularly to the field of assisting a user's oral care routine.
Background Art
[0002] Oral care devices such as electric toothbrushes or mouthpieces are used regularly (e.g., daily), and oral care routines involving these devices often include multiple complex steps. As a result, users performing oral care routines often cannot accurately judge their own performance of the routine. Particular problem examples are proper brushing and flossing of teeth.
[0003] Guided brushing is a feature for providing better oral care to toothbrush users. Users tend to be biased, such as brushing one side of the mouth more than the other, or concentrating on the front teeth at the expense of the back teeth. Therefore, better oral care can be achieved by providing users with important brushing information such as the total time spent brushing each area of the oral cavity, or whether basic oral hygiene procedures such as flossing have been performed. This information is particularly useful for parents tracking their children's tooth cleaning routines.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Currently, guided brushing is provided using sensors embedded in toothbrushes such as gyroscopes. However, these sensors are expensive and are therefore only available in premium and expensive toothbrushes. Most conventional devices for oral care do not include sensors suitable for monitoring oral care routines. Furthermore, there are few solutions for tracking flossing performance.
[0005] US2020 / 201272A1 describes a system and method for operating a personal grooming / home appliance, and includes providing a personal grooming / home appliance including at least one physical sensor.
[0006] "Method and System for Measuring Effectiveness of Tooth Brushing ED-Darl Kuhn" describes a method and system for measuring the effectiveness of a user's toothbrushing by dynamically evaluating the degree of plaque removal along with the practice of proper brushing techniques.
[0007] WO2021 / 197801A1 describes a method for tracking a user's toothbrushing activity, which includes receiving, for example, a video image of the user's face during a toothbrushing session.
[0008] US2018 / 132602A1 describes an oral care system that may have a toothbrush including physical characteristics and a programmable processor configured to receive physical characteristic data.
[0009] US2020 / 179089A1 describes an oral hygiene monitoring system for tracking the movement and orientation of an oral hygiene device. The control system can process data output from a motion sensor to determine the position and orientation of the oral hygiene device.
Means for Solving the Problems
[0010] The present invention is defined by the claims.
[0011] According to an example according to an aspect of the present invention, a method for assisting a user's oral care routine is provided.
[0012] The method includes obtaining video data from a captured video of a user performing an oral care routine using a personal care device, processing the video data to obtain motion data representing some of the user's movements while performing the oral care routine, and analyzing the motion data to determine at least one parameter value of the oral care routine, wherein the portion of the user includes the user's hand.
[0013] Accordingly, the proposed concept aims to provide a scheme, solution, concept, design, method, and system for assisting a user's oral care routine.
[0014] In particular, embodiments aim to determine at least one parameter value of an oral care routine based on the movement of a personal care device and / or a portion of the user during the execution of the oral care routine. Information regarding the movement of the personal care device and / or the portion of the user can be obtained from a video of the user and / or the personal care device captured while performing the oral care routine. That is, in the example where the user brushes their teeth, a video of the user brushing their teeth is captured and analyzed. From the analysis of the movement of the user and the toothbrush, insights can be obtained regarding how well the user brushed their teeth, i.e., the exact amount of time applied to each tooth. Guidance can be provided to the user to inform them of the quality of their execution. In another example where the oral care routine is flossing, the movement of one or both of the user's hands is analyzed to determine whether the user is flossing correctly or comprehensively, and guidance can be provided to the user to inform them of a way to floss more effectively in the future.
[0015] In other words, it is proposed that captured videos of the user and / or the personal care device during the execution of the oral care routine be analyzed in order to obtain motion data that can be utilized to determine at least one parameter value of the personal care routine. That is, based on the movement of the user's body and / or a personal care device such as a toothbrush or a flossing device, insights regarding the user's execution of the oral care routine can be obtained. Such videos can be obtained using an existing or conventional device that includes a camera already owned by the user.
[0016] By providing a computer vision-based method for analyzing the oral care routine, feedback similar to guided brushing can be provided to the user regardless of the type of toothbrush used. However, embodiments are not limited to electric or manual toothbrushes. The personal care device can also have a mouthpiece, a flossing device, or any other arbitrary personal oral care device. Accordingly, one or more of the proposed concepts can be adopted within the scope of various personal care devices. Thus, embodiments are widely applicable in the field of personal care devices and can be particularly relevant to dental proposals. For example, it can enable improved cleaning of the user's teeth, gums, tongue, etc., and can reduce unnecessary tissue damage. Therefore, embodiments can be used in relation to dental procedures to support dental practitioners when providing treatment to a subject.
[0017] By being integrated into the user's normal brushing method, embodiments can support improved dental care. Accordingly, the proposed concept can provide an improved oral care routine.
[0018] For example, one or more insights or statistics useful to the user can be determined by automatically analyzing the user's portion while the oral care routine is being executed. Exemplary parameter values are, for example, those related to the user's bias. For example, if the oral care routine is the user's toothbrushing, the parameter value can be the ratio of the time spent brushing the left side of the mouth to the time spent brushing the right side of the mouth, or alternatively, the percentage of the oral cavity that has been cleaned. Such insights can enable the provision of feedback to the user. As a result, favorable habits are reinforced and unfavorable actions are highlighted.
[0019] Using only video data to obtain insights regarding the execution of the user of the oral care routine enables assistance to the user in that routine without the need for a dedicated sensor in a personal care device that is costly and increases the complexity of the necessary personal care device. Thus, assistance can be provided to the user executing the oral care routine regardless of whether the oral care device has a sensor suitable for monitoring the oral care routine.
[0020] Ultimately, the proposed concept supports improved execution of the oral care routine by the user.
[0021] In some embodiments, at least one parameter value can have at least one of a user bias, a measure of completion of the oral care routine, a measure of completion of a sub-routine of the oral care routine, and a duration. An example of a user bias is the ratio of the time spent brushing the upper teeth to the time spent brushing the lower teeth. An example of a measure of completion of the oral care routine is the number of remaining teeth that will be properly cleaned. An example of a measure of completion of a sub-routine of the oral care routine is whether the user has flossed. An example of a duration is the total time spent brushing each area of the oral cavity. These parameter values enable the provision of feedback to the user. As a result, the user becomes aware of deficiencies in the execution of the oral care routine and conscious improvement is enabled.
[0022] In some embodiments, the obtained motion data further represents the operation of the personal care device. This can for example enable the tracking of the movement of a toothbrush or flossing device, which can enable further insight into the user's execution of the oral care routine.
[0023] In some embodiments, processing video data to obtain motion data representing the movement of a personal care device can include providing the video data as an input to a first convolutional neural network CNN, where the first CNN is trained to predict motion data indicating a series of positions of the personal care device for the video data associated with the personal care device.
[0024] Using a CNN instead of a plane object detection algorithm enables pixel-level segmentation of the personal care device, as compared to a simple bounding box. This facilitates more accurate tracking of the position and movement of the personal care device.
[0025] In some embodiments, a series of positions of a personal care device can represent an area of the personal care device. For example, the upper part of a toothbrush, such as the position and movement of the brushing head, can be particularly used during the analysis of motion data to determine parameter values of an oral care routine.
[0026] In some embodiments, the first CNN can be trained using a training algorithm configured to receive a training input and an array of individual known outputs, the training input having video data related to a personal care device, and the individual known outputs having motion data indicating a series of positions of the personal care device. Thus, the first CNN can be trained to output motion data indicating a series of positions of the personal care device when video data related to the personal care device during an oral care routine is provided.
[0027] In some embodiments, the first CNN is a pre-trained model further trained with videos of subjects using a manually annotated personal care device. This enables the CNN to be in a state of particular proficiency in identifying the personal care device.
[0028] In some embodiments, processing the video data to obtain motion data representing the movement of a portion of the user is providing the video data as an input to a second neural network, the second neural network being trained to predict motion data indicating a series of positions of the user's palm for the portion of the user associated with the video data, providing the video data as an input to a third neural network, the third neural network being trained to predict motion data indicating a series of positions of the user's hand landmarks for the portion of the user associated with the video data.
[0029] This segmentation model, which separately identifies the position of the user's palm and the position of the user's hand landmarks, enables fast and accurate hand tracking without using special hardware such as a depth perception camera. A third neural network can quickly identify the positions of the hand landmarks by simply searching within the bounding box specified by the second neural network.
[0030] In some embodiments, the second neural network is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, where the training inputs have video data associated with the user's portion, and the individual known outputs have motion data indicating a series of positions of the user's palm. This enables the second neural network to become proficient at identifying the user's palm.
[0031] Furthermore, the third neural network is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, where the training inputs have video data associated with the user's portion, and the individual known outputs have motion data indicating a series of positions of the user's hand landmarks. This enables the third neural network to become proficient at identifying the user's hand landmarks.
[0032] In some embodiments, the second neural network has a single-shot multibox detector architecture, and the third neural network has a feature pyramid network. The single-shot multibox detector architecture enables multiple objects present in an image to be detected in a single forward pass of the network. This enables multiple objects, such as two palms, to be quickly detected. The feature pyramid network is particularly adept at identifying small objects such as knuckles or finger joints and is thus suitable for detecting the user's hand landmarks.
[0033] In some embodiments, processing video data to obtain motion data representing the movement of a portion of the user comprises providing the video data as an input to a fourth neural network, the fourth neural network being trained to predict motion data indicating a series of positions of the user's face for the portion of the user associated with the video data, and providing the video data as an input to a fifth neural network, the fifth neural network being trained to predict motion data indicating a series of positions of landmarks of the user's face for the portion of the user associated with the video data.
[0034] This segmentation model that separately identifies the position of the user's face and the position of the landmarks of the user's face enables fast and accurate face tracking without using special hardware such as a depth perception camera. By first identifying the user's face and then simply searching within the bounding box detected by the fourth neural network, the landmarks on the face can be quickly identified by the fifth neural network.
[0035] In some embodiments, the fourth neural network is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, the training inputs having video data associated with a portion of the user, and the individual known outputs having motion data indicating a series of positions of the user's face. This enables the fourth neural network to become proficient at identifying the user's face.
[0036] Furthermore, the fifth neural network is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, where the training inputs have video data associated with a portion of the user, and the individual known outputs have motion data indicating a series of positions of landmarks of the user's face. This enables the fifth neural network to become proficient at identifying the landmarks of the user's face.
[0037] In some embodiments, the fourth neural network has a single-shot multibox detector architecture, and the fifth neural network has a feature pyramid network. The single-shot multibox detector architecture enables multiple objects present within an image to be detected in a single forward pass of the network. This enables single or multiple objects, such as one or more faces, to be detected rapidly. The feature pyramid network is particularly adept at identifying small objects such as nostrils and eyes, and is thus suitable for detecting the landmarks of the user's face.
[0038] In some embodiments, analyzing the motion data to determine at least one parameter value of an oral care routine comprises providing the motion data as an input to a machine learning algorithm, where the machine learning algorithm is trained to predict at least one parameter value of the oral care routine for the oral care routine associated with the motion data. From the position of the personal care device, the position of the hands and face landmarks, and pattern matching, the machine learning algorithm can infer useful information. For example, in the case where the user is brushing their teeth, the machine learning algorithm can predict the angle of the personal care device relative to the user's mouth.
[0039] In some embodiments, the machine learning algorithm has a supervised classifier model. This is particularly useful for classifying detected actions such as brushing a particular region of the user's mouth.
[0040] In some embodiments, the machine learning algorithm is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, where the training inputs have motion data related to an oral care routine and the individual known outputs have at least one parameter value of the oral care routine. This enables the machine learning algorithm to become particularly proficient at determining the parameter values of the oral care routine.
[0041] In some embodiments, analyzing the motion data to determine at least one parameter value of the oral care routine includes providing the motion data as an input to a rule - based algorithm designed to predict at least one parameter value of the oral care routine for the oral care routine associated with the motion data. This enables the parameter values to be inferred from the motion data without using a machine learning algorithm, which can reduce the computational load.
[0042] In some embodiments, analyzing the motion data to determine at least one parameter value of the oral care routine includes predicting the contact position between the personal care device and the user's surface. For example, when the user is brushing their teeth, the area where brushing is occurring can be inferred from the contact position between the personal care device and the user's surface. For example, it can be inferred that the user is currently brushing their left molars.
[0043] In some embodiments, analyzing the motion data to determine at least one parameter value of the oral care routine includes predicting the distance between the palm of each of the user's hands and the user's face, and optionally, the predicted distance between each palm of the user and the user's face is compared to a predetermined threshold.
[0044] This enables the detection that the user is flossing. This is because the user needs to bring both hands close to the mouth during flossing.
[0045] In some embodiments, when executed on a processing system, a computer program is provided that includes code means for implementing any of the above-described methods. According to another aspect of the present invention, a system for assisting a user's oral care routine is provided, and this system an input interface configured to obtain video data from captured videos of a user performing a personal health care routine using a personal care device, and a processor, wherein the processor processes the video data to obtain motion data representing the movement of a part of the user during the execution of the oral care routine, analyzes the motion data to determine at least one parameter value of the oral care routine, and the part of the user has the hand of the user.
[0046] In this way, a concept for assisting the user's oral care can be proposed, which can be based on the visually observed movement of the user during the execution of the oral care routine. Determining the parameter value of the oral care routine may help to inform the user of the quality of the execution of the oral care routine and enable the user to adjust the execution accordingly.
[0047] These and other aspects of the present invention will become apparent from the embodiments described below and will be described with reference to the embodiments.
Brief Description of the Drawings
[0048]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
[0049] For a better understanding of the present invention and to more clearly show the method of its implementation, reference is made to the accompanying drawings which are only illustrative.
[0050] The present invention will be described with reference to the drawings.
[0051] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, system and method, are for the purpose of illustration only and are not intended to limit the scope of the invention. These and other features, aspects and advantages of the apparatus, system and method of the present invention will be better understood from the following description, the appended claims and the accompanying drawings. It should be understood that the figures are merely schematic and are not drawn to scale. It should also be understood that the same reference numbers are used throughout the figures to indicate the same or similar parts.
[0052] Modifications to the disclosed embodiments can be understood and implemented by those skilled in the art of implementing the invention recited in the claims from a review of the figures, the disclosure and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.
[0053] It should be understood that the figures are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the figures to indicate the same or similar parts.
[0054] Implementations according to the present disclosure relate to various techniques, methods, schemes, and / or solutions related to assisting a user's oral care routine. According to the proposed concepts, a plurality of possible solutions can be implemented individually or jointly. That is, while these possible solutions are described individually below, two or more of these possible solutions can be implemented in one combination or another combination.
[0055] Embodiments of the present invention are directed to assisting a user's oral care routine by analyzing the movements (i.e., operations) of the user and a personal care device during the execution of the oral care routine. This can be achieved by acquiring video data from a captured video of a user performing an oral care routine using a personal care device. Next, the video data is processed to obtain motion data representing at least one movement of the personal care device and a part of the user during the execution of the oral care routine. Thereafter, the motion data can be analyzed to determine at least one parameter value of the oral care routine.
[0056] Accordingly, the proposed concept aims to determine at least one parameter value of an oral care routine based on the movement of a personal care device and / or parts of the user during the execution of the oral care routine. Information about the movement of the personal care device and / or parts of the user can be obtained from a video of the user and / or the personal care device captured while performing the oral care routine. That is, in the example of a user brushing their teeth, a video of the user brushing their teeth can be captured and analyzed. From the analysis of the movement of the user and / or the toothbrush, insights can be obtained about how well the user brushed their teeth, i.e., the exact amount of time spent on each tooth. Guidance can be provided to the user to inform them of the quality of the execution.
[0057] Referring now to FIG. 1, a flowchart of a method 100 for assisting a user's oral care routine according to a proposed embodiment is depicted.
[0058] The method begins with step 110 of acquiring video data from a captured video of a user performing an oral care routine using a personal care device. Such a video can be obtained using an existing or conventional device that includes a camera already owned by the user. For example, the conventional device can be any of a smartphone, tablet, laptop, or any other suitable device. By providing a computer vision-based method for analyzing the oral care routine, feedback similar to guided brushing can be provided to the user regardless of the type of toothbrush or personal care device used. Using only video data to obtain insights into the user's execution of the oral care routine enables assistance to the user in that routine without the need for a dedicated sensor in the personal care device, which is costly and increases the complexity of the required personal care device. Thus, assistance can be provided to the user performing the oral care routine regardless of whether the personal care device has a sensor suitable for monitoring the oral care routine.
[0059] In step 120, video data is processed, and motion data representing at least one movement of the personal care device and the user's part during the execution of the oral care routine can be obtained. Details of the processing of the video data will be described in detail later.
[0060] In step 130, the motion data is analyzed, and at least one parameter value of the oral care routine is determined. Exemplary parameter values are related to the user's bias, for example. For example, if the oral care routine is the user's toothbrushing, the parameter value can be the ratio of the time spent brushing the left side of the mouth to the time spent brushing the right side of the mouth, or alternatively, the percentage of the mouth that has been cleaned. Such insights can enable the provision of feedback to the user. As a result, preferred habits are reinforced and undesirable behaviors are highlighted.
[0061] In some embodiments, the at least one parameter value can have at least one of the user's bias, a measure of completion of the oral care routine, a measure of completion of a subroutine of the oral care routine, and a duration. An example of the user's bias is the ratio of the time spent brushing the upper teeth to the time spent brushing the lower teeth. An example of a measure of completion of the oral care routine is the number of remaining teeth that will be properly cleaned. An example of a measure of completion of a subroutine of the oral care routine is whether the user has flossed. An example of the duration is the total time spent brushing each area of the mouth. These parameter values enable the provision of feedback to the user, and as a result, the user becomes aware of deficiencies in the execution of the oral care routine and conscious improvement is made possible.
[0062] Next, referring to FIG. 2, a more detailed flowchart of a method 200 for assisting a user's oral care routine according to an exemplary embodiment is depicted. The method begins, as described above, with step 110 of obtaining video data from a captured video of a user performing an oral care routine using a personal care device.
[0063] In step 210, the video data is processed to predict motion data indicating a series of positions of the personal care device. The processing of the video data is facilitated by providing the video data as an input to a first convolutional neural network CNN, and the first CNN is trained to predict motion data indicating a series of positions of the personal care device for the personal care device. Thus, the movement of the personal care device captured during the execution of the oral care routine is converted / translated into motion data.
[0064] Using a CNN instead of a plane object detection algorithm enables pixel-level segmentation of the personal care device as compared to a simple bounding box. This facilitates more accurate tracking of the position and movement of the personal care device.
[0065] The structure of an artificial neural network (or simply a neural network) is inspired by the human brain. A neural network has layers, and each layer has a plurality of neurons. Each neuron has a mathematical operation. In particular, each neuron has a different weighted combination of a single type of transformation (e.g., the same type of transformation, sigmoid, etc., but with different weightings). In the process of processing the input data, the mathematical operations of each neuron are performed on the input data, a numerical output is generated, and the output of each layer in the neural network is sequentially supplied to the next layer. The final layer provides the output.
[0066] There are several types of neural networks, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs). This exemplary embodiment of the present invention employs a CNN-based learning algorithm. This is because CNNs have been particularly successful in the analysis of videos and have been proven to be able to classify frames within a video with a much lower error rate than other types of neural networks.
[0067] A CNN typically includes several layers, such as convolutional layers, pooling layers, and fully connected layers. A convolutional layer consists of a set of learnable filters and extracts features from the input. A pooling layer is a form of non-linear downsampling that reduces the data size by combining the outputs of multiple neurons in one layer into one neuron in the next layer. A fully connected layer connects each neuron in one layer to all neurons in the next layer.
[0068] Methods for training machine learning algorithms are well known. Typically, such methods involve obtaining a training data set that includes training input data entries and corresponding training output data entries. An initialized machine learning algorithm is applied to each input data entry, and a predicted output data entry is generated. The error between the predicted output data entry and the corresponding training output data entry is used to modify the machine learning algorithm. This process can be repeated until the error converges and the predicted output data entry is sufficiently similar (e.g., ±1%) to the training output data entry. This is generally known as supervised learning.
[0069] For example, the weighting of the mathematical operations of each neuron can be changed until the error converges. Known methods for modifying neural networks include the gradient descent method, the backpropagation algorithm, and the like.
[0070] The training input data entries for the first CNN used in method 200 correspond to exemplary video data related to a personal care device. The training output data entries correspond to motion data indicating a series of positions of the personal care device. Further, several preprocessing methods can be employed to improve the training samples. In other words, the first CNN can be trained using a training algorithm configured to receive an array of training inputs and individual known outputs, where the training inputs have video data related to a personal care device and the individual known outputs have motion data indicating a series of positions of the personal care device. Thus, the first CNN can be trained to output motion data indicating a series of positions of the personal care device when video data related to the personal care device during an oral care routine is provided.
[0071] In some embodiments, the first CNN is a pre-trained model further trained on videos of subjects using a manually annotated personal care device. This enables the CNN to become particularly proficient at identifying the personal care device. In some embodiments, the pre-trained model is trained on a COCO dataset based on a RESNET34, RESNET50, and / or RESNET101 backbone architecture.
[0072] In some embodiments, the first CNN is more specifically a Mask-RCNN deep neural-based model trained based on custom toothbrush annotations.
[0073] In some embodiments, the series of positions of the personal care device represents regions of the personal care device. For example, the upper edge of the toothbrush, such as the position and movement of the brushing head, can be particularly used during the analysis of motion data to determine parameter values for an oral care routine.
[0074] In step 220, video data is processed to predict a series of positions of a part of the user during the execution of the oral care routine. In step 220, the part of the user is the palm of the user's hand. This processing is facilitated by providing the video data as an input to a second neural network, and the second neural network is trained to predict motion data indicating a series of positions of the palm of the user's hand for the part of the user associated with the video data. In one embodiment, the second neural network has a single-shot multibox detector architecture. The single-shot multibox detector architecture enables multiple objects present in an image to be detected in a single forward pass of the network. This enables multiple objects, such as two palms, to be detected quickly.
[0075] The second neural network is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, the training inputs having video data associated with a part of the user, and the individual known outputs having motion data indicating a series of positions of the palm of the user's hand. This enables the second neural network to become proficient at identifying the palm of the user's hand and drawing a bounding box around it.
[0076] In step 230, video data is processed to predict a series of positions of a portion of the user while the oral care routine is being performed. In step 230, the portion of the user is the position of the landmarks of the user's hand. This processing is facilitated by providing the video data as an input to a third neural network, which is trained to predict a series of positions of the landmarks of the user's hand for the portion of the user associated with the video data. In one embodiment, the third neural network has a feature pyramid network. The feature pyramid network is particularly adept at identifying small objects such as knuckles or finger joints and is thus suitable for detecting the landmarks of the user's hand. In one embodiment, at least 21 landmarks of the hand are detected.
[0077] The third neural network is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, where the training inputs have video data associated with a portion of the user and the individual known outputs have motion data indicating a series of positions of the landmarks of the user's hand. This enables the third neural network to become proficient at identifying the landmarks of the user's hand. In an embodiment, the coordinates of the identified landmarks in the user's hand are stored in memory for further analysis.
[0078] This segmentation model that separately identifies the position of the palm of the user's hand and the position of the landmarks of the user's hand enables fast and accurate hand tracking, such as tracking of the hand skeleton, without using special hardware such as a depth perception camera. The third neural network can quickly identify the positions of the hand landmarks by simply searching within the bounding box specified by the second neural network. The landmarks can be identified using an encoder-decoder architecture.
[0079] In step 240, video data is processed to predict motion data indicating a series of positions of a part of the user during the execution of the oral care routine. In step 240, the part of the user is the user's face. This processing is facilitated by providing the video data as an input to a fourth neural network, and the fourth neural network is trained to predict motion data indicating a series of positions of the user's face for the part of the user associated with the video data. In an embodiment, the fourth neural network has a single-shot multibox detector architecture. The single-shot multibox detector architecture enables multiple objects present in an image to be detected in a single forward pass of the network. This enables a single or multiple objects, such as one or more faces, to be detected quickly.
[0080] The fourth neural network is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, the training inputs having video data related to a part of the user, and the individual known outputs having motion data indicating a series of positions of the user's face. This enables the fourth neural network to become proficient at identifying the user's face.
[0081] In step 250, video data is processed to predict motion data indicating a series of positions of a portion of the user during the execution of the oral care routine. In step 250, the portion of the user is the position of the landmarks of the user's face. This processing is facilitated by providing the video data as an input to a fifth neural network, which is trained to predict motion data indicating a series of positions of the landmarks of the user's face for the portion of the user associated with the video data. In one embodiment, the fifth neural network has a feature pyramid network. The feature pyramid network is particularly adept at identifying small objects such as nostrils and eyes, and is thus suitable for detecting the landmarks of the user's face. In one embodiment, at least 230 landmarks of the face are detected.
[0082] The fifth neural network is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, the training inputs having video data associated with a portion of the user, and the individual known outputs having motion data indicating a series of positions of the landmarks of the user's face. This enables the fifth neural network to become proficient at identifying the landmarks of the user's face. In an embodiment, the coordinates of the identified landmarks on the user's face are stored in memory for further analysis.
[0083] This segmentation model that separately identifies the position of the user's face and the position of the landmarks of the user's face enables fast and accurate face tracking without using special hardware such as a depth perception camera. By first identifying the user's face and then simply searching within the bounding box detected by the fourth neural network, the landmarks on the face can be more quickly identified by the fifth neural network. The landmarks can be identified using an encoder-decoder architecture.
[0084] In step 260, the motion data determined by steps 210, 220, 230, 240, and 250 is analyzed to determine at least one parameter value of the oral care routine. This can be performed in either of two ways.
[0085] i.) The motion data is provided as input to a machine learning algorithm, and the machine learning algorithm is trained to predict at least one parameter value of the oral care routine for the oral care routine associated with the motion data. From the position of the personal care device, the position of the landmarks of the hand and face, and pattern matching, the machine learning algorithm can infer useful information. For example, in the case where the user is brushing their teeth, the machine learning algorithm can predict the angle of the personal care device with respect to the user's mouth.
[0086] In one embodiment, the machine learning algorithm has a supervised classifier model. This is particularly useful for classifying detected actions such as brushing a particular area of the user's mouth. The area of brushing can be inferred based on the personal care device, the landmark coordinates of the hand and face, and pattern matching.
[0087] The machine learning algorithm is trained using a training algorithm configured to receive an array of training inputs and individual known outputs, where the training inputs have motion data related to the oral care routine and the individual known outputs have at least one parameter value of the oral care routine. This enables the machine learning algorithm to become particularly proficient at determining the parameter values of the oral care routine.
[0088] ii.) The movement data can be provided as an input to a rule - based algorithm designed to predict at least one parameter value of an oral care routine associated with the movement data. This enables the inference of parameter values from the movement data without using a machine - learning algorithm, which can reduce the computational load. For example, if the toothbrush head is closer to the right end than the left end of the mouth and the angle between the toothbrush and a line parallel to the plane of the face is 95 degrees, it can be determined that brushing is being performed on the right side of the mouth.
[0089] In one embodiment, analyzing the movement data to determine at least one parameter value of an oral care routine by any method involves predicting the contact position between the personal care device and the user's surface. For example, in the case where the user is brushing their teeth, the area where brushing is being performed can be inferred from the contact position between the personal care device and the user's surface. For example, it can be inferred that the user is currently brushing their left molars.
[0090] In one embodiment, analyzing the movement data to determine at least one parameter value of an oral care routine involves predicting the distance between each palm of the user and the user's face. The predicted distance between each palm of the user and the user's face can be compared to a predetermined threshold value. This enables, but is not limited to, the detection of the user flossing, since the user needs to bring both hands closer to the mouth during flossing.
[0091] Referring now to FIGS. 3a and 3b, a simplified diagram of the tracking of a personal care device according to the proposed embodiment is depicted. These figures depict a user performing an oral care routine such as brushing. In either figure, the personal care device 320 is lifted up to the user's face 310. The upper end 340 and the lower end 350 of the personal care device are used to determine the angle 360 of the personal care device with respect to a line 330 passing through the upper end and parallel to the plane of the face.
[0092] From the processing of video data to obtain motion data indicating a series of positions of the personal care device 320, the first CNN outputs a mask of the personal care device. In this example, the personal care device is a toothbrush. From the output mask, the upper end 340 of the toothbrush (i.e., the minimum value of the Y coordinate where the upper part of the video frame is Y = 0) and the lower end 350 (i.e., the maximum value of the Y coordinate) can be found.
[0093] In one embodiment, the first CNN is Mask-RCNN. However, in this embodiment, the mask of the personal care device 320 generated by Mask-RCNN can be composed of a plurality of sub-masks. To overcome this problem and determine the upper end 340 and the lower end 350 of the personal care device, the following algorithm is implemented.
[0094] First, the upper end 340 is extracted by finding the minimum Y coordinate of the mask (where the upper part of the video frame is Y = 0) and the corresponding X coordinate. Next, the lower end 350 is extracted by finding the maximum Y coordinate of the mask and the corresponding X coordinate. Thirdly, if there are two or more masks of the personal care device 320, the upper and lower ends of each sub-mask are repeatedly found, and the first two steps are repeated to reach the upper and lower ends of the personal care device.
[0095] After the upper end 340 and the lower end 350 of the personal care device 320 are found, the angle 360 of the personal care device with respect to the plane 330 of the face 310 can be determined. The angle between the toothbrush and the face is used to determine the brushing location. For example, when the user is brushing the side of the mouth, the toothbrush is approximately perpendicular to the plane of the face, while when brushing the front of the mouth, the toothbrush is approximately parallel to the face. To determine the angle, the upper and lower ends of the toothbrush are connected by a straight line. This line intersects the reference line 330 parallel to the plane of the face 310 at the upper end, which enables the angle between the two lines to be determined using the dot product. This angle is equal to the angle between the toothbrush and the face.
[0096] By treating two lines as two vectors, the angle between the two vectors can be calculated using the following equation (i). It can be calculated using TIFF2025523350000002.tif1060.
[0097] Here,[[]]END]] TIFF2025523350000003.tif97 is the inner product between the two vectors, TIFF2025523350000004.tif910 is the magnitude of the vector.
[0098] In an embodiment, the output coordinates of each of the first to fifth neural networks are stored in a memory. Thereafter, the stored coordinates are analyzed, and the brushing position is predicted using a rule-based or machine learning-based approach.
[0099] The memory can include any one or combination of volatile memory elements (e.g., random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), etc.) and non-volatile memory elements (e.g., ROM, erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), tape, compact disc read-only memory (CD-ROM), disk, floppy disk, cartridge, cassette, etc.). Further, the memory can incorporate electronic, magnetic, optical, and / or other types of storage media. Note that the memory can have a distributed architecture. In that case, various elements are located at separate locations but can be accessed by a processor.
[0100] In some embodiments, when executed on a processing system, a computer program is provided that includes code means for implementing any of the above-described methods.
[0101] Next, referring to FIG. 4, a simplified block diagram of a system 400 for assisting a user's oral care routine according to an embodiment is depicted. A system configured to assist a user's oral care routine has an input interface 410 and one or more processors 430.
[0102] System 400 is configured to analyze the movement of the user and the movement of the personal care device while the oral care routine is being performed. A video 415 capturing the user performing the oral care routine is acquired. The system outputs an output 450 having at least one parameter value of the oral care routine. The output 450 is generated based on the analysis of the motion data and is for providing feedback regarding the execution of the oral care routine to the user.
[0103] More specifically, the input interface 410 receives video data 415 from the captured video of the user while the oral care routine is being performed. The input interface 410 provides the video data 415 to one or more neural networks 420 to predict the position of one or more parts of the user within the video data and to predict the position of the personal care device within the video data. This facilitates the analysis of the operation pattern of the oral care routine.
[0104] Based on the identified positions, one or more processors 430 then process the video data with a motion detection algorithm to obtain motion data representing the movement of at least one of the personal care device and the parts of the user while the oral health care routine is being performed. For this purpose, a multi-person detection network can be used to avoid disturbances by focusing only on the user performing the oral care routine.
[0105] The resulting motion data is processed by the processor 430, and an output 450 having at least one parameter value of the oral care routine is generated.
[0106] FIG. 5 shows an example of a computer 500 in which one or more portions of the embodiments may be employed. The various operations described above can utilize the capabilities of the computer 500. In this regard, it should be understood that the system function blocks can be executed on a single computer or can be distributed among multiple computers and locations (e.g., connected via the Internet).
[0107] The computer 500 includes, but is not limited to, a PC, a workstation, a laptop, a PDA, a palm device, a server, storage, etc. Generally, from the perspective of hardware architecture, the computer 500 can include one or more processors 510, a memory 520, and one or more I / O devices 530 communicatively coupled via a local interface (not shown). The local interface can be, for example, but is not limited to, one or more buses or other wired or wireless connections as known in the art. The local interface can have additional elements such as a controller, a buffer (cache), a driver, a repeater, and a receiver to enable communication. Further, the local interface can include address, control, and / or data connections to enable proper communication among the aforementioned elements.
[0108] The processor 510 is a hardware device that executes software that can be stored in the memory 520. The processor 510 can be virtually any custom-made or can be a commercially available processor, a central processing unit (CPU), a digital signal processor (DSP), or an auxiliary processor among multiple processors related to the computer 500. The processor 510 can be a semiconductor-based microprocessor (in the form of a microchip) or a microprocessor.
[0109] Memory 520 can include any one or combination of volatile memory elements (e.g., random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), etc.) and non-volatile memory elements (e.g., ROM, erasable programmable read-only memory (EPROM), electronically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), tape, compact disc read-only memory (CD-ROM), disk, floppy disk, cartridge, cassette, etc.). Further, memory 520 can incorporate electronic, magnetic, optical, and / or other types of storage media. Note that memory 520 can have a distributed architecture. In that case, various elements are located at separate locations but can be accessed by processor 510.
[0110] The software in memory 520 can include one or more separate programs, and each program has an ordered list of executable instructions for implementing a logical function. According to an exemplary embodiment, the software in memory 520 includes a suitable operating system (O / S) 550, compiler 560, source code 570, and one or more applications 580. As illustrated, application 580 has a plurality of functional elements for implementing the features and processes of the exemplary embodiment. The application 580 of computer 500 can represent various applications, computing units, logics, functional units, processes, operations, virtual entities, and / or modules according to an exemplary embodiment, but application 580 is not meant to be limiting.
[0111] The operating system 550 controls the execution of other computer programs and provides scheduling, input / output control, file and data management, memory management, communication control, and related services. The inventors assume that the application 580 for implementing the exemplary embodiments is applicable on all commercially available operating systems.
[0112] The application 580 can be a source program, an executable program (object code), a script, or any other entity having a set of instructions to be executed. In the case of a source program, the program is typically translated through a compiler (such as compiler 560), an assembler, an interpreter, etc. so as to operate properly in connection with the O / S 550. These may or may not be included in the memory 520. Further, the application 580 can be described, for example, but not limited to, object-oriented programming languages having classes of data and methods such as C, C++, C#, Pascal, BASIC, API calls, HTML, XHTML, XML, ASP scripts, JavaScript, FORTRAN, COBOL, Perl, Java, ADA,.NET, or procedural programming languages having routines, subroutines, and / or functions.
[0113] The I / O device 530 can include, for example, but is not limited to, input devices such as a mouse, keyboard, scanner, microphone, and camera. Further, the I / O device 530 can include, for example, but is not limited to, output devices such as a printer and a display. Finally, the I / O device 530 can further include, for example, but is not limited to, devices that communicate both input and output, such as a NIC, a modem (for accessing remote devices, other files, devices, systems, or networks), a radio frequency (RF) or other transceiver, a telephone interface, a bridge, or a router. The I / O device 530 also includes elements that communicate via various networks such as the Internet or an intranet.
[0114] When the computer 500 is a PC, a workstation, an intelligent device, etc., the software in the memory 520 can further include a basic input / output system (BIOS) (omitted for simplicity). The BIOS is an essential set of software routines that initializes and tests the hardware at startup, starts the O / S 550, and supports the transfer of data between hardware devices. The BIOS is stored in a certain type of read-only memory such as ROM, PROM, EPROM, or EEPROM, and as a result, the BIOS can be executed when the computer 800 is started.
[0115] When the computer 500 is operating, the processor 510 is configured to execute the software stored in the memory 520, communicate data with the memory 520, and generally control the operation of the computer 500 according to the software. The application 580 and the O / S 550 are read, in whole or in part, by the processor 510 and perhaps buffered in the processor 510 before being executed.
[0116] When the application 580 is implemented in software, it should be noted that the application 580 can be stored in substantially any computer-readable medium for use by or in connection with any computer-related system or method. In the context of this document, a computer-readable medium can be an electronic, magnetic, optical, or other physical device or means that can store or hold a computer program for use by or in connection with a computer-related system or method.
[0117] The application 580 can be implemented, for example, in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems, that can fetch instructions from and execute the instructions. In the context of this document, a "computer-readable medium" can be any means that can store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable medium can be, by way of example and not limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium.
[0118] The methods of FIGS. 1-2 and the system of FIG. 4 can be implemented in hardware or software, or a combination of both (e.g., firmware executed on a hardware device). To the extent embodiments are implemented partially or wholly in software, the functional steps shown in the process flowcharts can be executed by a suitably programmed physical computing device such as one or more central processing units (CPUs) or graphics processing units (GPUs). Each process and individual element step shown in the flowchart can be executed by the same or different computing devices. According to an embodiment, a computer-readable storage medium stores a computer program including computer program code configured to cause one or more physical computing devices to perform the above-described encoding or decoding method when the program is executed on the one or more physical computing devices.
[0119] The storage medium includes volatile and non-volatile computer memories such as RAM, PROM, EPROM, EEPROM, optical disks (such as CD, DVD, BD), and magnetic storage media (such as hard disks and tapes). The various storage media may be fixed within the computing device or may be transportable such that one or more programs stored therein can be loaded into the processor.
[0120] To the extent that an embodiment is implemented partially or wholly in hardware, the blocks shown in the block diagram of FIG. 4 may be separate physical elements, may be a logical subdivision of a single physical element, or may be implemented in a manner integrated into a single physical element. The function of one block shown in the figure may be divided among multiple elements in implementation, or the functions of multiple blocks shown in the figure may be combined into one element in implementation. Hardware elements suitable for use in embodiments of the present invention include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field programmable gate arrays (FPGAs). One or more blocks can be implemented as a combination of dedicated hardware for performing some functions and one or more programmed microprocessors and associated circuitry for performing other functions.
[0121] Variations to the disclosed embodiments can be understood and effected by a person skilled in the art of carrying out the invention as claimed, from a consideration of the drawings, the disclosure and the appended claims. In the claims, the term "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. One processor or other unit may perform the functions of a plurality of items recited in the claims. The mere fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used advantageously. Where a computer program is discussed, it may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but it may also be distributed in other forms, such as via the Internet or other wired or wireless communication systems. It should be noted that, where the term "adapted to" is used in the claims or the specification, the term "adapted to" is intended to be equivalent to the term "configured to". Any reference signs in the claims should not be construed as limiting the scope of the invention.
[0122] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible realizations of systems, methods, and computer programs according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions, which has one or more executable instructions for implementing a specified logical function. In some alternative realizations, the functions described in the blocks may occur out of the order described in the figures. For example, two blocks shown in succession may actually be executed substantially simultaneously, or depending on the related functions, the blocks may be executed in the reverse order. It should also be noted that each block of the block diagrams and / or flowchart diagrams, as well as combinations of blocks in the block diagrams and / or flowchart diagrams, can be implemented by a special-purpose hardware-based system that performs the specified functions or acts, or executes a combination of special-purpose hardware and computer instructions.
Claims
1. In a method for supporting a user's oral care routine, The steps include: acquiring video data from a captured video of a user performing the oral care routine using a personal care device; The steps include processing the aforementioned video data to obtain motion data representing the movement of the user's body part while performing the oral care routine, The procedure includes the step of analyzing the aforementioned movement data and determining at least one parameter value of the oral care routine, A method wherein the user portion comprises the palms of the user, and the step of analyzing the motion data to determine at least one parameter value of the oral care routine includes the step of predicting the distance between the palms of the user and the user's face.
2. The method according to claim 1, wherein the at least one parameter value has at least one of user bias, a measure of completion of the oral care routine, a measure of completion of a subroutine of the oral care routine, and duration.
3. The method according to claim 1 or 2, wherein the obtained exercise data further represents the exercise of the personal care device.
4. The step of processing the aforementioned video data to obtain motion data representing the movement of the personal care device is: The method according to claim 3, comprising the step of providing the video data as input to a first convolutional neural network (CNN), wherein the first CNN is trained to predict motion data indicating a set of positions of a personal care device associated with the video data.
5. The first CNN is trained using a training algorithm configured to receive a training input and an array of individual known outputs, wherein the training input has video data related to a personal care device, and the individual known outputs have motion data indicating a sequence of positions of the personal care device. The method according to claim 4.
6. The method according to claim 4, wherein the first CNN is a pre-trained model further trained on manually annotated videos of subjects using the personal care device.
7. The step of processing the aforementioned video data to obtain motion data representing the movement of the user's body part is: A step of providing the video data as input to a second neural network, wherein the second neural network is trained to predict motion data indicating a series of positions of the user's palm for the portion of the user associated with the video data. The method includes the step of providing the video data as input to a third neural network, the third neural network being trained to predict motion data indicating a set of positions of landmarks of the user's hand for the portion of the user associated with the video data. The method according to any one of claims 1 to 6.
8. The second neural network is trained using a training algorithm configured to receive a training input and an array of individual known outputs, wherein the training input has video data associated with a part of the user, and the individual known outputs have motion data indicating a sequence of positions of the user's palm. The method according to claim 7, wherein the third neural network is trained using a training algorithm configured to receive a training input and an array of individual known outputs, the training input having video data associated with a portion of the user, and the individual known outputs having motion data indicating a sequence of positions of landmarks in the user's hand.
9. The method according to claim 7 or 8, wherein the second neural network has a single-shot multibox detector architecture, and the third neural network has a feature pyramid network.
10. The step of processing the aforementioned video data to obtain motion data representing the movement of the user's body part is: A step of providing the video data as input to a fourth neural network, wherein the fourth neural network is trained to predict motion data indicating a series of positions of the user's face for the portion of the user associated with the video data. The method according to any one of claims 1 to 9, comprising the step of providing the video data as input to a fifth neural network, wherein the fifth neural network is trained to predict motion data indicating the positions of a set of landmarks of the user's face for the portion of the user associated with the video data.
11. The fourth neural network is trained using a training algorithm configured to receive a training input and an array of known outputs, wherein the training input has video data relating to a portion of the user, and the individual known outputs have motion data indicating a sequence of positions of the user's face. The method according to claim 10, wherein the fifth neural network is trained using a training algorithm configured to receive a training input and an array of individual known outputs, the training input having video data associated with a portion of the user, and the individual known outputs having motion data indicating a sequence of locations of landmarks on the user's face.
12. The fourth neural network has a single-shot multibox detector architecture, and the fifth neural network has a feature pyramid network. The method according to claim 10 or 11.
13. The step of analyzing the aforementioned exercise data to determine at least one parameter value of the oral care routine is: The method according to any one of claims 1 to 12, further comprising the step of providing the motion data as input to a machine learning algorithm, wherein the machine learning algorithm is trained to predict at least one parameter value of an oral care routine associated with the motion data.
14. A computer program that, when executed on a processing system, includes code means for implementing the method according to any one of claims 1 to 13.
15. A system that supports the user's oral care routine, An input interface configured to acquire video data from captured video of a user performing a personal care routine using a personal care device, It has a processor, and the processor is The video data is processed to obtain motion data representing the movement of the user's body part while the oral care routine is being performed. The system is configured to analyze the aforementioned exercise data and determine at least one parameter value of the oral care routine. A system in which the user portion has each of the user's palms, analyzes the movement data, determines at least one parameter value of the oral care routine, and predicts the distance between each of the user's palms and the user's face.