Tooth brushing action detection method, apparatus, device and medium
Patent Information
- Application Number
- CN202511553578.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-10-28
AI Technical Summary
然而,由于缺乏有效的实时指导与反馈,用户往往难以自我纠正错误的刷牙行为
本公开实施例的一个有益效果在于,在用户刷牙过程中,可获取头戴设备的目标摄像头拍摄的连续多帧用户刷牙图像,并将连续多帧用户刷牙图像输入预先训练好的刷牙动作检测模型,以便获得用户的刷牙动作检测结果,以及在用户的刷牙动作检测结果为用户的刷牙动作为横刷动作的情况下,输出用于提示用户的刷牙动作为横刷动作的提示信息。也就是说,在用户刷牙过程中,其能够精准地识别用户的刷牙动作是否为横刷动作,并为用户提供及时的反馈,以帮助用户养成正确的刷牙习惯,提升牙齿健康水平。
Smart Images

Figure CN121545213B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of head-mounted device technology, and more specifically, to a method, apparatus, device, and medium for detecting brushing motions. Background Technology
[0002] Proper brushing habits, such as proper brushing techniques, are key to maintaining oral and dental health. Incorrect brushing methods are a major cause of tooth damage, with horizontal brushing (side brushing) being particularly harmful. However, due to a lack of effective real-time guidance and feedback, users often find it difficult to correct their incorrect brushing behaviors on their own. Summary of the Invention
[0003] The purpose of this disclosure is to provide a new technical solution for detecting brushing motions.
[0004] According to a first aspect of the present disclosure, a method for detecting brushing motions is provided, applied to a head-mounted device, the head-mounted device including a target camera, the method comprising: Acquire multiple consecutive frames of user brushing teeth images captured by the target camera; The continuous multi-frame images of the user brushing their teeth are input into a pre-trained brushing action detection model to obtain the user's brushing action detection results. If the user's brushing action detection result indicates that the user's brushing action is a horizontal brushing action, a prompt message is output; wherein, the prompt message is used to indicate that the user's brushing action is a horizontal brushing action.
[0005] Optionally, the brushing action detection model sequentially includes: a spatial feature extraction network, a temporal feature extraction network, a temporal pooling network, a feature enhancement network, and a classification network; The step of inputting the continuous multi-frame user brushing images into a pre-trained spatiotemporal fusion neural network model to obtain the user's brushing action detection results includes: The continuous multi-frame images are input into the spatial feature extraction network, and the spatial features of each frame of the user brushing teeth image in the continuous multi-frame user brushing teeth image are extracted by the spatial feature extraction network to obtain a spatial feature sequence. Temporal features are extracted from the spatial feature sequence using the temporal feature extraction network to obtain a temporal feature sequence. The temporal feature sequence is pooled using the temporal pooling network to obtain the aggregated temporal feature vector; The aggregated temporal feature vector is enhanced by the feature enhancement network to obtain the enhanced temporal feature vector. The enhanced temporal feature vector is classified using the classification network to obtain the user's brushing action detection results. Optionally, the brushing action detection model further includes a spatial downsampling network. Before inputting the consecutive multi-frame images into the spatial feature extraction network, and extracting the spatial features of each frame of the user brushing teeth image from the consecutive multi-frame user brushing teeth images through the spatial feature extraction network to obtain a spatial feature sequence, the method further includes: The spatial downsampling network is used to preprocess the consecutive multi-frame user brushing images; The preprocessing includes at least one of frame normalization, center cropping, and random horizontal flipping.
[0006] Optionally, the step of pooling the temporal feature sequence through the temporal pooling network to obtain the aggregated temporal feature vector includes: The temporal feature sequence is subjected to max pooling and average pooling through the temporal pooling network to obtain a first pooling result and a second pooling result. The first pooling result and the second pooling result are then concatenated to obtain an aggregated temporal feature vector.
[0007] Optionally, the feature enhancement network sequentially includes a first fully connected layer, a layer normalization and activation layer, a second fully connected layer, an attention gating unit, and a global context enhancement unit. The step of performing feature enhancement processing on the aggregated temporal feature vector through the feature enhancement network to obtain the enhanced temporal feature vector includes: The aggregated temporal feature vector is upgraded using the first fully connected layer. The time-series feature vectors after dimensionality increase are normalized and activated sequentially through the layer normalization and activation layers. The second fully connected layer performs dimensionality reduction on the normalized and activated temporal feature vectors. The attention gating unit filters out noise features from the time-series feature vector after dimensionality reduction and amplifies the motion features associated with the horizontal brushing action to obtain an intermediate feature vector. The intermediate feature vector is subjected to global context enhancement processing by the global context enhancement unit to obtain the enhanced feature vector.
[0008] Optionally, the head-mounted device is smart glasses, and the target camera is disposed on the side area of the front frame of the smart glasses.
[0009] Optionally, the method further includes: Receive control commands input by the user; In response to the control command, the head-mounted device is controlled to enter the brushing monitoring mode; When the head-mounted device enters the brushing monitoring mode, the target camera is controlled to continuously capture multiple frames of user brushing images at a set sampling frequency; The user is in a facing position, with the mirror in front of them.
[0010] According to a second aspect of the present disclosure, a brushing motion detection device is provided for use in smart glasses, the smart glasses including a target camera, the device comprising: The acquisition module is used to acquire multiple consecutive frames of user brushing teeth images captured by the target camera; The detection module is used to input the continuous multi-frame user brushing images into a pre-trained brushing action detection model to obtain the user's brushing action detection results; The output module is used to output a prompt message when the user's brushing action detection result indicates that the user's brushing action is a horizontal brushing action; wherein the prompt message is used to indicate that the user's brushing action is a horizontal brushing action. According to a third aspect of the present disclosure, a head-mounted device is provided, comprising: a memory for storing executable computer instructions; and a processor for executing the method described in accordance with the first aspect above, under the control of the executable computer instructions.
[0011] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, perform the method described in the first aspect above. One beneficial effect of this disclosure is that, during the user's brushing process, multiple consecutive frames of user brushing images captured by the target camera of the head-mounted device can be acquired. These multiple frames are then input into a pre-trained brushing motion detection model to obtain the user's brushing motion detection results. If the user's brushing motion detection results indicate that the brushing motion is a horizontal brushing motion, a prompt message is output to indicate that the brushing motion is a horizontal brushing motion. In other words, during the user's brushing process, it can accurately identify whether the user's brushing motion is a horizontal brushing motion and provide timely feedback to help the user develop correct brushing habits and improve dental health.
[0012] Other features and advantages of this specification will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of this specification and, together with their description, serve to explain the principles of this specification.
[0014] Figure 1 This is a schematic diagram of the hardware configuration of the head-mounted device provided in an embodiment of this disclosure; Figure 2 This is a schematic flowchart of the brushing action detection method provided in the embodiments of this disclosure; Figure 3a This is one of the structural schematic diagrams of the brushing action detection model provided in the embodiments of this disclosure; Figure 3b This is the second schematic diagram of the structure of the brushing action detection model provided in the embodiments of this disclosure; Figure 4 This is a block diagram of the brushing action detection device provided in the embodiments of this disclosure; Figure 5 This is a block diagram of a wearable device provided in an embodiment of this disclosure. Detailed Implementation
[0015] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the embodiments of the present disclosure.
[0016] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0017] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0018] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0019] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0020] <Hardware Configuration> Figure 1 This is a block diagram of the hardware configuration of a head-mounted device 1000 according to an embodiment of the present disclosure.
[0021] like Figure 1As shown, the head-mounted device 1000 can be, for example, smart glasses. The head-mounted device 1000 may include a processor 1100, a memory 1200, and a target camera 1300. The processor 1100 may include, but is not limited to, a central processing unit (CPU), a microprocessor (MCU), etc. The memory 1200 may include, for example, ROM (read-only memory), RAM (random access memory), or non-volatile memory such as a hard disk. The target camera 1300 may be, for example, a red-green-blue (RGB) camera, and is typically located on the side of the front frame of the head-mounted device for capturing images of the user brushing their teeth.
[0022] Taking the head-mounted device 1000 as smart glasses as an example, the processor 1100 can be located on the right temple of the smart glasses, and the target camera 1300 can be located on the right side of the front frame of the smart glasses. Of course, the processor 1100 can also be located on the left temple of the smart glasses, and the target camera 1300 can be located on the left side of the front frame of the smart glasses.
[0023] Those skilled in the art should understand that, although in Figure 1 The present specification shows a number of devices of the head-mounted device 1000; however, the head-mounted device 1000 of the embodiments described herein may involve only some of the devices, or may include other devices, which is not limited herein.
[0024] In this embodiment, the memory 1200 of the head-mounted device 1000 is used to store instructions that control the processor 1100 to operate in order to implement or support the implementation of the brushing action detection method according to any embodiment. Those skilled in the art can design instructions based on the scheme disclosed in this specification. How the instructions control the processor to operate is well known in the art and will not be described in detail here.
[0025] In the above description, those skilled in the art can design instructions based on the solutions provided in this disclosure. How the instructions control the processor to operate is well known in the art, and therefore will not be described in detail here.
[0026] Figure 1 The head-mounted device shown is illustrative only and is by no means intended to limit this disclosure, its application or use.
[0027] <Method Implementation> Figure 2 This disclosure illustrates a brushing motion detection method according to an embodiment of the present disclosure. The brushing motion detection method can be performed by… Figure 1 The head-mounted device 1000 shown is used for execution. The head-mounted device 1000 may include a target camera 1300, which may be, for example, an RGB camera. Figure 2As shown, the brushing action detection method of this embodiment may include the following steps S2100 to S2300: Step S2100: Acquire multiple consecutive frames of user brushing teeth images captured by the target camera.
[0028] In this embodiment, the user is typically in a facing-the-mirror posture while brushing their teeth. The RGB camera of the head-mounted device can continuously capture multiple frames of the user brushing their teeth at a set sampling frequency. The set sampling frequency can be preset according to the actual scenario and needs.
[0029] Taking smart glasses as an example, an RGB camera can be set on the right side of the front frame of the smart glasses. When a user is brushing their teeth, they are usually in a facing position in front of the mirror. The RGB camera of the smart glasses will continuously capture 10 frames of the user brushing their teeth at a sampling frequency of 10 frames per second.
[0030] After executing step S2100 to acquire multiple consecutive frames of user brushing images captured by the target camera, proceed to: Step S2200: Input the continuous multi-frame user brushing images into the pre-trained brushing action detection model to obtain the user's brushing action detection results.
[0031] The brushing action detection model can be pre-trained and stored in the head-mounted device's memory. Its input can be multiple consecutive frames of user brushing images, and its output can be the user's brushing action detection result. This brushing action detection model can be a spatiotemporal fusion neural network model.
[0032] In one example, the brushing action detection result can be either horizontal brushing or non-horizontal brushing. Horizontal brushing can include horizontal vibrations, i.e., the toothbrush moves laterally from side to side across the tooth surface. Vertical brushing can include vertical sweeping, i.e., the toothbrush moves from top to bottom (upper teeth) or bottom to top (lower teeth). Non-horizontal brushing can be, for example, vertical brushing, which is the core action of the Bass brushing technique and is generally applicable to all people who need to maintain oral health.
[0033] Continuing with the example of smart glasses as a head-mounted device, the smart glasses continuously capture 10 frames of user brushing images at a sampling frequency of 10 frames per second during the user's brushing process. These 10 consecutive frames of user brushing images are then input into a pre-trained brushing action detection model, which can then output the user's brushing action detection results.
[0034] After executing step S2200, which inputs the consecutive multi-frame user brushing images into the pre-trained brushing action detection model to obtain the user's brushing action detection results, the process proceeds to: Step S2300: If the user's brushing action detection result indicates that the user's brushing action is a horizontal brushing action, output a prompt message.
[0035] The prompt information is used to remind the user that the brushing motion is a horizontal brushing motion.
[0036] In this embodiment, if the user's brushing action detection result indicates that the user's brushing action is a horizontal brushing action, the head-mounted device can output a prompt message to remind the user that the brushing action is a horizontal brushing action. This prompt message can remind the user to develop healthy brushing habits.
[0037] In one example, if the user's brushing motion is detected as a horizontal brushing motion, the speaker of the head-mounted device is controlled to emit a first voice message to prompt the user that the brushing motion is a horizontal brushing motion.
[0038] In one example, if the user's brushing action detection result indicates that the user's brushing action is a horizontal brushing action, the terminal device connected to the head-mounted device is controlled to issue a second voice message, which is used to prompt the user that the brushing action is a horizontal brushing action. Through the embodiments of this disclosure, during a user's brushing process, multiple consecutive frames of the user's brushing images captured by a target camera on a wearable device can be acquired. These multiple frames are then input into a pre-trained brushing motion detection model to obtain the user's brushing motion detection results. If the user's brushing motion detection results indicate that the brushing motion is a horizontal brushing motion, a prompt message is output to indicate that the brushing motion is a horizontal brushing motion. In other words, during the user's brushing process, it can accurately identify whether the user's brushing motion is a horizontal brushing motion and provide timely feedback to help the user develop correct brushing habits and improve dental health.
[0039] In one embodiment of this disclosure, reference is made to Figure 3a The brushing action detection model sequentially includes: a spatial feature extraction network, a temporal feature extraction network, a temporal pooling network, a feature enhancement network, and a classification network. Step S2200, which inputs the continuous multi-frame user brushing images into the pre-trained brushing action detection model to obtain the user's brushing action detection results, can further include the following steps S2210 to S2250: Step S2210: Input the continuous multi-frame user brushing images into the spatial feature extraction network, and extract the spatial features of each frame of the user brushing images in the continuous multi-frame user brushing images through the spatial feature extraction network to obtain a spatial feature sequence.
[0040] In one example, multiple consecutive frames of user brushing images can be directly input into a spatial feature extraction network to extract the spatial features of each frame of the user brushing images, thus obtaining a spatial feature sequence.
[0041] In one example, refer to Figure 3b The brushing action detection model can also include a spatial downsampling network. Alternatively, the spatial downsampling network can be used to downsample multiple consecutive frames of user brushing images. Then, the downsampled multiple consecutive frames of user brushing images can be input into a spatial feature extraction network. The spatial feature extraction network can extract the spatial features of each frame of user brushing images in the multiple consecutive frames of user brushing images to obtain a spatial feature sequence.
[0042] Typically, the downsampling process described above can include at least one of frame normalization, center cropping, and random horizontal flipping.
[0043] In this example, downsampling of multiple consecutive frames of user brushing images can speed up model computation and improve model detection accuracy.
[0044] In one example, the spatial feature extraction network can be the lightweight convolutional neural network MobileNet V3, with the classification layer removed from the end of the original MobileNet V3 network structure and a fully connected layer connected at the end of the MobileNet V3 network. This allows the network to extract spatial features from each frame of a user brushing their teeth image and perform dimensionality reduction, for example, compressing the spatial features of each frame of a user brushing their teeth image from high dimension to 128 dimensions, ultimately generating a spatial feature sequence of dimension N×128, where N is the number of frames.
[0045] Continuing with the example of smart glasses as a head-mounted device, the smart glasses continuously capture 10 frames of the user's brushing image at a sampling frequency of 10 frames per second during the brushing process. Each frame of the user's brushing image has a resolution of 1008*756. These 10 frames of the user's brushing image are first input into a spatial downsampling network for processing such as frame normalization, center cropping, and random horizontal flipping, outputting 10*224*224*3. Next, the downsampled 10 frames of the user's brushing image are input into a spatial feature extraction network to extract spatial features and perform dimensionality reduction, outputting 10*128.
[0046] Step S2220: Extract temporal features from the spatial feature sequence using the temporal feature extraction network to obtain a temporal feature sequence.
[0047] In one example, refer to Figure 3b Temporal feature extraction networks can include bidirectional long short-term memory networks (BiLSTM) and one-dimensional temporal convolutional networks. The BiLSTM network captures the sequence and context of brushing actions, while the one-dimensional temporal convolutional network captures short-term, high-frequency motion patterns within the brushing action. In this example, the temporal feature extraction network captures subtle local motion patterns using the one-dimensional temporal convolutional network, while simultaneously using the BiLSTM network to model long-term temporal dependencies and the global action context. These two networks work together to form a comprehensive and robust feature representation of the brushing action in the temporal dimension, thus laying the foundation for accurately identifying horizontal brushing behavior.
[0048] In this example, the bidirectional long short-term memory network has 64 bidirectional hidden units each, totaling 128. To prevent overfitting, the random deactivation dropout can be set to 0.3, and peephole connections can be enabled. The one-dimensional temporal convolutional network can be set with a kernel size of 3 and a stride of 1, and batch normalization and Leaky ReLU can be used.
[0049] Continuing with the example of smart glasses as a head-mounted device, the spatial feature sequence is input into the temporal feature extraction network. This network then passes through the bidirectional long short-term memory network and the one-dimensional temporal convolutional network, outputting a temporal feature sequence. The outputs of both the bidirectional long short-term memory network and the one-dimensional temporal convolutional network are 10*128.
[0050] Step S2230: The temporal feature sequence is pooled using the temporal pooling network to obtain the aggregated temporal feature vector.
[0051] In one example, the time-series feature sequence is pooled by the time-pooling network to obtain the aggregated time-series feature vector. This can be done as follows: the time-series feature sequence is subjected to max pooling and average pooling respectively by the time-pooling network to obtain the first pooling result and the second pooling result, and the first pooling result and the second pooling result are concatenated to obtain the aggregated time-series feature vector.
[0052] In this example, refer to Figure 3bThe temporal feature sequence is input into the temporal pooling network, which performs global max pooling and global average pooling in parallel to extract salient action features and overall average patterns from the temporal feature sequence, respectively. The two pooling results are then concatenated along the feature dimension to finally output a fixed-dimensional global temporal feature, such as a 128-dimensional global temporal feature.
[0053] Step S2240: The aggregated temporal feature vector is subjected to feature enhancement processing through the feature enhancement network to obtain the enhanced temporal feature vector.
[0054] In one example, refer to Figure 3b The feature enhancement network sequentially includes a first fully connected layer, a layer normalization and activation layer, a second fully connected layer, an attention gating unit, and a global context enhancement unit. The feature enhancement network performs feature enhancement processing on the aggregated temporal feature vector to obtain the enhanced temporal feature vector. This can be achieved as follows: the first fully connected layer increases the dimensionality of the aggregated temporal feature vector; the layer normalization and activation layer sequentially normalizes and activates the increased dimensionality temporal feature vector; the second fully connected layer reduces the dimensionality of the normalized and activated temporal feature vector; the attention gating unit filters out noise features from the reduced dimensionality temporal feature vector and amplifies motion features associated with the horizontal brush action to obtain an intermediate feature vector; and the global context enhancement unit performs global context enhancement processing on the intermediate feature vector to obtain the enhanced feature vector.
[0055] In this example, to better capture brushing patterns, the 128-dimensional global temporal features are upscaled to 256 dimensions using a first fully connected layer to enhance data expressiveness. Layer normalization and activation layers are then applied to the upscaled features, followed by non-linear mapping using the Swish activation function. A Dropout layer with a dropout rate of 0.3 is also included. A second fully connected layer then reduces the dimensionality of the processed 256-dimensional features back to 128 dimensions. Next, an attention gating unit filters out noisy features from the processed 128-dimensional features and amplifies motion features associated with the horizontal brushing motion to obtain an intermediate feature vector. The attention gating unit can... Calculate the gate value and through Get output ,in, and These are learnable weights and biases. It's the Sigmoid function, used to compress the output to a range between 0 and 1. These are learnable weights. Furthermore, the intermediate feature vectors can be input into the global context enhancement unit (GCU), which performs global average pooling on the intermediate feature vectors to obtain global context information. This global context information is then fused into the local feature vectors at each time step in the original sequence.
[0056] Step S2250: Classify the enhanced temporal feature vector through the classification network to obtain the user's brushing action detection result.
[0057] In one example, refer to Figure 3b The classification network can include a third fully connected layer and a classification sub-network, which can use the Softmax function.
[0058] In this example, the third fully connected layer receives the 128-dimensional feature vector output by the global context augmentation unit. This third fully connected layer linearly transforms the input 128-dimensional feature vector, reducing its dimensionality to a single 2-dimensional feature vector. Next, the 2-dimensional feature vector output by the third fully connected layer is fed into the Softmax function for processing. After processing by the Softmax layer, a 2-dimensional probability vector is finally output, which can be in the form [P(swiping), P(non-swiping)]. Here, P(swiping) represents the confidence probability that the model determines the user is performing a swiping action. In practical applications, the category with the higher probability value is usually chosen as the final detection result.
[0059] In one embodiment of this disclosure, the brushing action detection method of this disclosure may further include the following steps S3100 to S3300: Step S3100: Receive control commands input by the user.
[0060] The control command can be a voice control command or other forms of command, and it is used to trigger the head-mounted device to enter the brushing monitoring mode.
[0061] Typically, control commands can be issued by the user while in an upright position in front of the face mirror.
[0062] Step S3200: In response to the control command, control the head-mounted device to enter the brushing monitoring mode.
[0063] In this embodiment, the head-mounted device can enter the brushing monitoring mode upon receiving a control command input by the user.
[0064] Step S3300: When the head-mounted device enters the brushing monitoring mode, the target camera is controlled to continuously acquire multiple frames of user brushing images at a set sampling frequency.
[0065] In this embodiment, when the head-mounted device enters the brushing monitoring mode, the head-mounted device can control the target camera to continuously acquire multiple frames of user brushing images at a set sampling frequency, so as to recognize the user's brushing action based on the continuous multiple frames of user brushing images.
[0066] In this embodiment, the head-mounted device only controls the target camera to continuously capture multiple frames of user brushing images at a set sampling frequency when it receives a control command input by the user in order to perform user brushing action recognition. This can improve the targeting accuracy of brushing action recognition and reduce the power consumption of the head-mounted device.
[0067] <Device Embodiment> Figure 4 This is a schematic diagram of a brushing motion detection device 400 according to one embodiment, which is applied to smart glasses. The smart glasses include a target camera. Figure 4 As shown, the brushing action detection device 400 includes an acquisition module 410, a detection module 420, and an output module 430.
[0068] The acquisition module 410 is used to acquire multiple consecutive frames of user brushing teeth images captured by the target camera; The detection module 420 is used to input the continuous multi-frame user brushing images into a pre-trained brushing action detection model to obtain the user's brushing action detection results. The output module 430 is used to output a prompt message when the user's brushing action detection result indicates that the user's brushing action is a horizontal brushing action; wherein the prompt message is used to indicate that the user's brushing action is a horizontal brushing action.
[0069] In one embodiment, the brushing action detection model sequentially includes: a spatial feature extraction network, a temporal feature extraction network, a temporal pooling network, a feature enhancement network, and a classification network.
[0070] The detection module 420 is specifically used to input the continuous multi-frame user brushing images into the spatial feature extraction network, extract spatial features from each frame of the user brushing images through the spatial feature extraction network to obtain a spatial feature sequence; extract temporal features from the spatial feature sequence through the temporal feature extraction network to obtain a temporal feature sequence; pool the temporal feature sequence through the temporal pooling network to obtain an aggregated temporal feature vector; perform feature enhancement processing on the aggregated temporal feature vector through the feature enhancement network to obtain an enhanced temporal feature vector; and classify the enhanced temporal feature vector through the classification network to obtain the user's brushing action detection result. In one embodiment, the brushing action detection model further includes a spatial downsampling network.
[0071] The detection module 420 is also used to downsample the continuous multi-frame user brushing images through the spatial downsampling network.
[0072] The downsampling process includes at least one of frame normalization, center cropping, and random horizontal flipping.
[0073] In one embodiment, the detection module 420 is specifically used to perform max pooling and average pooling on the temporal feature sequence through the temporal pooling network to obtain a first pooling result and a second pooling result, and to concatenate the first pooling result and the second pooling result to obtain an aggregated temporal feature vector.
[0074] In one embodiment, the feature enhancement network sequentially includes a first fully connected layer, a layer normalization and activation layer, a second fully connected layer, an attention gating unit, and a global context enhancement unit.
[0075] The detection module 420 is specifically used to perform dimensionality upscaling on the aggregated temporal feature vector through the first fully connected layer; to perform normalization and activation processing on the dimensionality upscaling temporal feature vector through the layer normalization and activation layers in sequence; to perform dimensionality reduction processing on the normalized and activated temporal feature vector through the second fully connected layer; to filter out noise features in the dimensionality reduction temporal feature vector through the attention gating unit and amplify motion features associated with the horizontal brushing action to obtain an intermediate feature vector; and to perform global context enhancement processing on the intermediate feature vector through the global context enhancement unit to obtain an enhanced feature vector.
[0076] In one embodiment, the target camera is disposed in the side area of the front frame of the head-mounted device.
[0077] In one embodiment, the brushing action detection device 400 further includes a control module (not shown in the figure).
[0078] A control module is used to receive control commands input by the user; in response to the control commands, control the head-mounted device to enter the brushing monitoring mode; when the head-mounted device enters the brushing monitoring mode, control the target camera to continuously acquire multiple frames of user brushing images at a set sampling frequency; wherein the user is in a positive posture facing the mirror.
[0079] According to embodiments of this disclosure, during a user's brushing process, multiple consecutive frames of the user's brushing images captured by a target camera on a wearable device can be acquired. These multiple frames are then input into a pre-trained brushing motion detection model to obtain the user's brushing motion detection results. If the user's brushing motion detection results indicate that the brushing motion is a horizontal brushing motion, a prompt message is output to indicate that the brushing motion is a horizontal brushing motion. In other words, during the user's brushing process, the model can accurately identify whether the user's brushing motion is a horizontal brushing motion and provide timely feedback to help the user develop correct brushing habits and improve dental health.
[0080] <Equipment Example> Figure 5 This is a schematic diagram of the hardware structure of a head-mounted device according to one embodiment. For example... Figure 5 As shown, the head-mounted device 1000 includes a processor 1100 and a memory 1200.
[0081] The memory 1200 can be used to store executable computer instructions.
[0082] The processor 1100 can be used to execute a brushing action detection method according to embodiments of the present disclosure, under the control of executable computer instructions.
[0083] The head-mounted device 1000 can be as follows: Figure 1 The head-mounted device 1000 shown can also be a device with other hardware structures, which is not limited here.
[0084] In another embodiment, the head-mounted device 1000 may include the brushing action detection device 400 described above. In one embodiment, each module of the brushing action detection device 400 can be implemented by the processor 1100 running computer instructions stored in the memory 1200.
[0085] Computer-readable storage media This disclosure also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, perform the brushing action detection method provided in this disclosure.
[0086] This disclosure can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of this disclosure.
[0087] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0088] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0089] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0090] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0091] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0092] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation in a combination of software and hardware are equivalent.
[0094] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of this disclosure is defined by the appended claims.
Claims
1. A method for detecting brushing motion, characterized in that, Applied to a head-mounted device, the head-mounted device including a target camera, the method includes: Acquire multiple consecutive frames of user brushing teeth images captured by the target camera; The continuous multi-frame images of the user brushing their teeth are input into a pre-trained brushing action detection model to obtain the user's brushing action detection results. If the user's brushing motion detection result indicates that the user's brushing motion is a horizontal brushing motion, a prompt message is output; wherein, the prompt message is used to indicate that the user's brushing motion is a horizontal brushing motion; The brushing action detection model includes, in sequence, a spatial feature extraction network, a temporal feature extraction network, a temporal pooling network, a feature enhancement network, and a classification network. The feature enhancement network includes, in sequence, a first fully connected layer, a layer normalization and activation layer, a second fully connected layer, an attention gating unit, and a global context enhancement unit. The first fully connected layer performs dimensionality upscaling on the aggregated temporal feature vector output by the temporal pooling network. The layer normalization and activation layers then perform normalization and activation processes on the dimensionality-upgraded temporal feature vector. The second fully connected layer performs dimensionality reduction on the normalized and activated temporal feature vector. The attention gating unit filters out noise features from the dimensionality-reduced temporal feature vector and amplifies motion features associated with the horizontal brushing action to obtain an intermediate feature vector. The global context enhancement unit then performs global context enhancement on the intermediate feature vector to obtain an enhanced feature vector.
2. The method according to claim 1, characterized in that, The step of inputting the continuous multi-frame user brushing images into a pre-trained spatiotemporal fusion neural network model to obtain the user's brushing action detection results includes: The continuous multi-frame user brushing images are input into the spatial feature extraction network, and the spatial feature extraction network extracts the spatial features of each frame of the user brushing images in the continuous multi-frame user brushing images to obtain a spatial feature sequence. Temporal features are extracted from the spatial feature sequence using the temporal feature extraction network to obtain a temporal feature sequence. The temporal feature sequence is pooled using the temporal pooling network to obtain the aggregated temporal feature vector; The aggregated temporal feature vector is enhanced by the feature enhancement network to obtain the enhanced temporal feature vector. The enhanced temporal feature vector is classified using the classification network to obtain the user's brushing action detection results.
3. The method according to claim 2, characterized in that, The brushing action detection model also includes a spatial downsampling network. Before inputting the consecutive multi-frame images into the spatial feature extraction network, and extracting the spatial features of each frame of the user brushing teeth image from the consecutive multi-frame user brushing teeth images through the spatial feature extraction network to obtain a spatial feature sequence, the method further includes: The spatial downsampling network is used to downsample the continuous multi-frame user brushing images. The downsampling process includes at least one of frame normalization, center cropping, and random horizontal flipping.
4. The method according to claim 2, characterized in that, The step of pooling the temporal feature sequence through the temporal pooling network to obtain the aggregated temporal feature vector includes: The temporal feature sequence is subjected to max pooling and average pooling through the temporal pooling network to obtain a first pooling result and a second pooling result. The first pooling result and the second pooling result are then concatenated to obtain an aggregated temporal feature vector.
5. The method according to any one of claims 1 to 4, characterized in that, The target camera is located in the side area of the front frame of the head-mounted device.
6. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Receive control commands input by the user; In response to the control command, the head-mounted device is controlled to enter the brushing monitoring mode; When the head-mounted device enters the brushing monitoring mode, the target camera is controlled to continuously capture multiple frames of user brushing images at a set sampling frequency; The user is in a facing position, with the mirror in front of them.
7. A toothbrushing action detection device, characterized in that, Applied to smart glasses, the smart glasses including a target camera, the device includes: The acquisition module is used to acquire multiple consecutive frames of user brushing teeth images captured by the target camera; The detection module is used to input the continuous multi-frame user brushing images into a pre-trained brushing action detection model to obtain the user's brushing action detection results; The output module is used to output a prompt message when the user's brushing action detection result indicates that the user's brushing action is a horizontal brushing action; wherein the prompt message is used to indicate that the user's brushing action is a horizontal brushing action; The brushing action detection model includes, in sequence, a spatial feature extraction network, a temporal feature extraction network, a temporal pooling network, a feature enhancement network, and a classification network. The feature enhancement network includes, in sequence, a first fully connected layer, a layer normalization and activation layer, a second fully connected layer, an attention gating unit, and a global context enhancement unit. The first fully connected layer performs dimensionality upscaling on the aggregated temporal feature vector output by the temporal pooling network. The layer normalization and activation layers then perform normalization and activation processes on the dimensionality-upgraded temporal feature vector. The second fully connected layer performs dimensionality reduction on the normalized and activated temporal feature vector. The attention gating unit filters out noise features from the dimensionality-reduced temporal feature vector and amplifies motion features associated with the horizontal brushing action to obtain an intermediate feature vector. The global context enhancement unit then performs global context enhancement on the intermediate feature vector to obtain an enhanced feature vector.
8. A head-mounted device, characterized in that, include: Memory is used to store executable computer instructions; A processor configured to execute the method according to any one of claims 1-6, under the control of the executable computer instructions.
9. A computer-readable storage medium, characterized in that, It stores computer instructions that, when executed by a processor, perform the method described in any one of claims 1-6.