Hand washing detection method, wearable device and storage medium
By combining the data processing of the inertial detection unit and image sensor in the wearable device, the spatial and temporal characteristics of the image sequence are extracted, and the accurate detection of the hand washing state is achieved, solving the problem of low detection accuracy under the influence of noise.
Patent Information
- Application Number
- CN202510329963.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
When a wearable device detects whether a user is washing hands, methods based on acceleration, angular velocity and audio data are susceptible to noise, resulting in low detection accuracy.
After the motion data is collected by an inertial detection unit (IMU) for preliminary detection, the image sequence is acquired through the image sensor, the spatial and temporal features are extracted using the pre-trained feature extraction model and the multi-layer perception model, and the secondary detection of the hand washing state is performed based on the corresponding relationship of the training samples.
It improves the accuracy of hand washing detection, avoids misidentifying hand rubbing and other actions as hand washing states, and enhances the reliability of the detection.
Smart Images

Figure CN120260125A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of status detection, and particularly to a handwashing detection method, a wearable device, and a storage medium. Background Art
[0002] Wearable devices are the general term for devices developed by intelligent design of daily wear, such as smart glasses, smart watches, etc. Due to the advantages of being lightweight and easy to carry, more and more users choose to use wearable devices for health detection, such as handwashing detection.
[0003] The wearable device analyzes the handwashing data such as the acceleration and angular velocity collected by the built-in IMU (Inertial Measurement Unit) and the audio collected by the microphone to detect whether the user is in a handwashing state. However, in scenarios where the user rubs hands or there is background noise, the handwashing detection based on the handwashing data such as acceleration, angular velocity, and audio may be affected by noise, resulting in low accuracy of the handwashing detection using the above detection method. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a handwashing detection method, a wearable device, and a storage medium to improve the accuracy of handwashing detection. The specific technical solutions are as follows:
[0005] In a first aspect, the embodiments of this application provide a handwashing detection method applied to a wearable device, where the wearable device includes a motion sensor and an image sensor, and the motion sensor includes an IMU; the method includes:
[0006] In response to the first handwashing detection result indicating that the user is in a handwashing state, obtain the image sequence collected by the image sensor, where the first handwashing detection result is obtained by performing handwashing detection using the motion data collected by the IMU;
[0007] Input the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, where the spatial features are used to characterize the picture features of the water flow passing through the hand, and the temporal features are used to characterize the motion state of the user's hand and the characteristics of the change of the motion state of the water flow over time;
[0008] Use a pre-trained multi-layer perception model to obtain the handwashing detection result of the user based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perception model, where the handwashing detection result of the user indicates whether the user is in a handwashing state, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the handwashing state.
[0009] Optionally, the image sequence is a plurality of consecutive sub-image sequences collected within a preset time period;
[0010] The step of obtaining the handwashing detection result of the user based on the spatial feature, the temporal feature, and the corresponding relationship learned during the training process of the pre-trained multi-layer perceptron model includes:
[0011] For each sub-image sequence, use the pre-trained multi-layer perceptron model to obtain the second handwashing detection result corresponding to the sub-image sequence based on the spatial feature, the temporal feature, and the corresponding relationship learned during the training process of the multi-layer perceptron model;
[0012] When there are no less than a preset number of target results among the multiple second handwashing detection results, determine that the handwashing detection result of the user is that the user is in a handwashing state, where the target result indicates that the user is in a handwashing state;
[0013] When the number of target results among the multiple second handwashing detection results is less than the preset number, determine that the handwashing detection result of the user is that the user is not in a handwashing state.
[0014] Optionally, the multi-layer perceptron model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer;
[0015] The step of obtaining the second handwashing detection result corresponding to the sub-image sequence based on the spatial feature, the temporal feature, and the corresponding relationship learned during the training process of the pre-trained multi-layer perceptron model includes:
[0016] Input the linear feature vector corresponding to the sub-image sequence into the input layer, where the linear feature vector corresponding to the sub-image sequence is obtained by compressing the spatial feature and the temporal feature corresponding to the sub-image sequence, or output by the feature extraction model;
[0017] The input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer;
[0018] The hidden layer receives the linear feature vector, activates the non-linear relationship in the linear feature vector by using multiple neurons and activation functions in the at least one fully connected layer, and inputs the activated linear feature vector into the output layer;
[0019] The output layer receives the activated linear feature vector, uses a normalization function to perform a handwashing state mapping on the activated linear feature vector, and obtains the second handwashing detection result.
[0020] Optionally, the training method of the multi-layer perceptron model includes:
[0021] Obtain the linear feature vector corresponding to each training sample and the calibration label of each training sample. Among them, the multiple training samples include the sample image sequence collected by the image sensor of the wearable device when the user is in the handwashing state, and the sample image sequence collected by the image sensor when the user is not in the handwashing state. The calibration label of each training sample represents the result of whether the training sample is in the handwashing state.
[0022] Input the linear feature vector of each training sample into the initial multi-layer perceptron model.
[0023] Obtain the predicted label output by the initial multi-layer perceptron model based on the current model parameters for processing the linear feature vector of the training sample.
[0024] Based on the difference between the predicted label and the corresponding calibration label, adjust the model parameters of the initial multi-layer perceptron model until the initial multi-layer perceptron model converges to obtain the trained multi-layer perceptron model.
[0025] Optionally, after the step of extracting the spatial features and temporal features of the image sequence, the method further includes:
[0026] Use the activation function to activate the non-linear relationship between the spatial features and temporal features corresponding to each sub-image sequence respectively.
[0027] Perform three-dimensional pooling on the activated spatial features and temporal features corresponding to each sub-image sequence respectively, and reduce the resolution of the feature map composed of the activated spatial features and temporal features corresponding to each sub-image sequence to a preset first resolution.
[0028] Based on the preset dimensionality reduction function, reduce the dimension of the feature map with the first resolution corresponding to each sub-image sequence to one dimension to obtain the linear feature vector corresponding to each sub-image sequence.
[0029] Optionally, the step of inputting the image sequence into the pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence includes:
[0030] Input the image sequence into the pre-trained feature extraction model.
[0031] Slide the 3D convolutional kernel in the feature extraction model in the spatial dimension of the image sequence, extract the spatial features of the image sequence based on the content of the frames included in the image sequence, and slide the 3D convolutional kernel in the time dimension to extract the temporal features of the image sequence based on the law of change of the content included in the image sequence over time.
[0032] Optionally, before the step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial and temporal features of the image sequence, the method further includes:
[0033] Use a bilateral filter to denoise each frame image included in the image sequence to obtain a denoised image sequence;
[0034] Perform downsampling on the denoised image sequence to reduce the image sequence to a preset second resolution to obtain a downsampled image sequence;
[0035] Perform normalization on the downsampled image sequence to obtain an image sequence for inputting into a pre-trained feature extraction model.
[0036] Optionally, the first handwashing detection result is the confidence level that the user is in a handwashing state;
[0037] The step of, in response to the first handwashing detection result indicating that the user is in a handwashing state, obtaining the image sequence collected by the image sensor includes:
[0038] When the confidence level is higher than a preset threshold, activate the image sensor so that the image sensor collects an image sequence of the environment where the user is located.
[0039] Optionally, after the step of obtaining the handwashing detection result of the user, the method further includes:
[0040] When the handwashing detection result indicates that the user is in a handwashing state, send a first prompt message to the user and start a handwashing timer, and when the timing duration reaches a preset duration, send a second prompt message to the user, where the first prompt message is used to prompt the user to start the handwashing timer, and the second prompt message is used to prompt the user to end the handwashing timer; and / or,
[0041] When the handwashing detection result indicates that the user is not in a handwashing state, turn off the image sensor.
[0042] In a second aspect, an embodiment of the present application provides a handwashing detection device applied to a wearable device. The wearable device includes a motion sensor and an image sensor, and the motion sensor includes an IMU; the device includes:
[0043] An image sequence acquisition module, configured to obtain an image sequence collected by the image sensor in response to a first handwashing detection result indicating that the user is in a handwashing state, wherein the first handwashing detection result is obtained by performing handwashing detection on the motion data collected by the IMU;
[0044] A feature extraction module, configured to input the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, wherein the spatial features are used to characterize the picture features of the water flow passing through the hand, and the temporal features are used to characterize the motion state of the user's hand and the characteristics of the change of the motion state of the water flow over time;
[0045] A result determination module, configured to use a pre-trained multi-layer perceptron model to obtain the handwashing detection result of the user based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model, wherein the handwashing detection result of the user indicates whether the user is in a handwashing state, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the handwashing state.
[0046] Optionally, the image sequence is a plurality of consecutive sub-image sequences collected within a preset time period;
[0047] The result determination module is specifically configured to:
[0048] For each sub-image sequence, use a pre-trained multi-layer perceptron model to obtain a second handwashing detection result corresponding to the sub-image sequence based on the spatial features, temporal features corresponding to the sub-image sequence, and the corresponding relationship learned during the training process of the multi-layer perceptron model;
[0049] In the case that there are no less than a preset number of target results among the multiple second handwashing detection results, determine that the handwashing detection result of the user is that the user is in a handwashing state, wherein the target result indicates that the user is in a handwashing state;
[0050] In the case that the number of target results among the multiple second handwashing detection results is less than the preset number, determine that the handwashing detection result of the user is that the user is not in a handwashing state.
[0051] Optionally, the multi-layer perceptron model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer;
[0052] The result determination module is specifically configured to:
[0053] Input the linear feature vector corresponding to the sub-image sequence into the input layer, where the linear feature vector corresponding to the sub-image sequence is obtained by compressing the spatial feature and the temporal feature corresponding to the sub-image sequence, or is output by the feature extraction model;
[0054] The input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer;
[0055] The hidden layer receives the linear feature vector, utilizes multiple neurons in the at least one fully connected layer and an activation function to activate the non-linear relationship in the linear feature vector, and inputs the activated linear feature vector into the output layer;
[0056] The output layer receives the activated linear feature vector, uses a normalization function to perform a handwashing state mapping on the activated linear feature vector, and obtains a second handwashing detection result.
[0057] Optionally, the training method of the multi-layer perceptron model includes:
[0058] Obtain the linear feature vector corresponding to each training sample and the calibration label of each training sample, where the multiple training samples include the sample image sequences collected by the image sensor of the wearable device when the user is in the handwashing state and the sample image sequences collected by the image sensor when the user is not in the handwashing state, and the calibration label of each training sample represents the result of whether the training sample is in the handwashing state;
[0059] Input the linear feature vector of each training sample into the initial multi-layer perceptron model;
[0060] Obtain the predicted label output by the initial multi-layer perceptron model based on the current model parameters for processing the linear feature vector of the training sample;
[0061] Based on the difference between the predicted label and the corresponding calibration label, adjust the model parameters of the initial multi-layer perceptron model until the initial multi-layer perceptron model converges, and obtain the trained multi-layer perceptron model.
[0062] Optionally, the device further includes:
[0063] A post-processing module, configured to, after the step of extracting the spatial features and temporal features of the image sequence, use an activation function to respectively activate the non-linear relationships of the spatial features and temporal features corresponding to each sub-image sequence; perform three-dimensional pooling on the activated spatial features and temporal features corresponding to each sub-image sequence respectively, and reduce the resolution of the feature map formed by the activated spatial features and temporal features corresponding to each sub-image sequence to a preset first resolution; based on a preset dimensionality reduction function, reduce the dimension of the feature map with the first resolution corresponding to each sub-image sequence to one dimension, and obtain a linear feature vector corresponding to each sub-image sequence.
[0064] Optionally, the feature extraction module is specifically configured to:
[0065] Input the image sequence into a pre-trained feature extraction model;
[0066] Use a three-dimensional convolution kernel in the feature extraction model to slide in the spatial dimension of the image sequence, and based on the picture content included in the image sequence, extract the spatial features of the image sequence, and use the three-dimensional convolution kernel to slide in the temporal dimension, and based on the law of the content included in the image sequence changing over time, extract the temporal features of the image sequence.
[0067] Optionally, the device further includes:
[0068] A pre-processing module, configured to, before the step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, use a bilateral filter to perform noise reduction on each frame image included in the image sequence to obtain a denoised image sequence; perform downsampling processing on the denoised image sequence, reduce the image sequence to a preset second resolution, and obtain a downsampled image sequence; perform normalization processing on the downsampled image sequence to obtain an image sequence for inputting into a pre-trained feature extraction model.
[0069] Optionally, the first handwashing detection result is the confidence level of the user being in a handwashing state;
[0070] The image sequence acquisition module is specifically configured to:
[0071] When the confidence level is higher than a preset threshold, activate the image sensor to enable the image sensor to collect an image sequence of the user's environment.
[0072] Optionally, the device further includes:
[0073] A timing module, configured to, after obtaining the handwashing detection result of the user, when the handwashing detection result indicates that the user is in a handwashing state, send a first prompt message to the user and start handwashing timing, and when the timing duration reaches a preset duration, send a second prompt message to the user, wherein the first prompt message is used to prompt the user to start handwashing timing, and the second prompt message is used to prompt the user to end handwashing timing.
[0074] Optionally, the device further includes:
[0075] An image sensor closing module, configured to close the image sensor when the handwashing detection result indicates that the user is not in a handwashing state.
[0076] In a third aspect, an embodiment of the present application provides a wearable device, including:
[0077] An image sensor, configured to collect an image sequence;
[0078] A motion sensor, the motion sensor includes an IMU, and the IMU is configured to collect motion data;
[0079] A memory, configured to store a computer program;
[0080] A processor, configured to, when executing the program stored on the memory, implement the method according to any one of the first aspects described above.
[0081] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method according to any one of the first aspects described above is implemented.
[0082] Beneficial effects of the embodiments of the present application:
[0083] In the technical solution provided by the embodiments of the present application, when the first handwashing detection result obtained based on the motion data collected by the IMU indicates that the user is in the handwashing state, the wearable device further acquires the image sequence collected by the image sensor; and then based on the image sequence, the pre-trained feature extraction model, and the pre-trained multi-layer perceptron model, the final handwashing detection result is obtained. Since the pre-trained feature extraction model learns how to extract the spatial features and temporal features in the image sequence through a large number of training samples during the training process, the wearable device can accurately extract the spatial features and temporal features in the image sequence by using the pre-trained feature extraction model. Similarly, the pre-trained multi-layer perceptron model learns the correspondence between the sample spatial features and sample temporal features and the handwashing state through a large number of training samples during the training process, and the wearable device can accurately predict whether the user is in the handwashing state based on the spatial features, temporal features of the image sequence, and the pre-trained multi-layer perceptron model.
[0084] It can be seen that based on the first handwashing detection result obtained by using the motion data collected by the IMU, the wearable device can further use the image sequence of the user's environment collected by the image sensor for the second handwashing detection to determine whether the user is actually in the handwashing state. In this way, by using the motion data collected by the IMU and the image sequence of the environment where the user is located for two handwashing detections respectively, the handwashing state of the user can be determined more accurately, avoiding misidentifying actions similar to handwashing actions such as the user rubbing hands as the user being in the handwashing state, and improving the accuracy of the handwashing detection result.
[0085] Moreover, since the spatial features and temporal features of the image sequence are respectively extracted when using the image sequence for handwashing detection, the characteristics of the water flow passing through the hand, the motion state of the user's hand, and the characteristics of the change of the motion state of the water flow over time are captured simultaneously, further improving the accuracy of the second handwashing detection result, and thus further improving the accuracy of the finally determined handwashing detection result, that is, improving the accuracy of handwashing detection.
[0086] Of course, it is not necessary for any product or method implementing the present application to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0088] Figure 1The first process schematic diagram of the handwashing detection method provided by the embodiment of the present application;
[0089] Figure 2 A refined process schematic diagram of step S103 provided by the embodiment of the present application;
[0090] Figure 3 A refined process schematic diagram of step S201 provided by the embodiment of the present application;
[0091] Figure 4 The process schematic diagram of a training method of the multi-layer perceptron model provided by the embodiment of the present application;
[0092] Figure 5 The second process schematic diagram of the handwashing detection method provided by the embodiment of the present application;
[0093] Figure 6 The third process schematic diagram of the handwashing detection method provided by the embodiment of the present application;
[0094] Figure 7 The process schematic diagram of a training method of the SVM classification model for obtaining the first handwashing detection result provided by the embodiment of the present application;
[0095] Figure 8 The process schematic diagram of an example of applying the handwashing detection method provided by the embodiment of the present application;
[0096] Figure 9 The structural schematic diagram of a handwashing detection device provided by the embodiment of the present application;
[0097] Figure 10 The structural schematic diagram of a wearable device provided by the embodiment of the present application. Detailed implementation manners
[0098] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0099] In order to improve the accuracy of handwashing detection, the embodiment of the present application provides a handwashing detection method, device, wearable device, computer-readable storage medium, and computer program product. First, a handwashing detection method provided by the embodiment of the present application will be introduced below.
[0100] As Figure 1As shown in the figure, a handwashing detection method is applied to a wearable device. The wearable device includes a motion sensor and an image sensor, and the motion sensor includes an IMU. The method includes:
[0101] S101, in response to the first handwashing detection result indicating that the user is in a handwashing state, obtaining an image sequence collected by the image sensor;
[0102] Wherein, the first handwashing detection result is obtained by performing handwashing detection on the motion data collected by the IMU;
[0103] S102, inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence;
[0104] Wherein, the spatial features are used to characterize the characteristics of the picture of the water flow passing through the hand, and the temporal features are used to characterize the motion state of the user's hand and the characteristics of the change of the motion state of the water flow over time;
[0105] S103, using a pre-trained multi-layer perceptron model, based on the spatial features, temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model, obtaining the handwashing detection result of the user;
[0106] Wherein, the handwashing detection result of the user indicates whether the user is in a handwashing state, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the handwashing state.
[0107] It can be seen that in the embodiment of the present application, when the first handwashing detection result obtained based on the motion data collected by the IMU indicates that the user is in a handwashing state, the wearable device further obtains the image sequence collected by the image sensor; and then based on the image sequence, the pre-trained feature extraction model, and the pre-trained multi-layer perceptron model, the final handwashing detection result is obtained.
[0108] Since the pre-trained feature extraction model learns how to extract the spatial features and temporal features in the image sequence through a large number of training samples during the training process, the wearable device can accurately extract the spatial features and temporal features in the image sequence by using the pre-trained feature extraction model. Similarly, the pre-trained multi-layer perceptron model learns the corresponding relationship between the sample spatial features and sample temporal features and the handwashing state through a large number of training samples during the training process, and the wearable device can accurately predict whether the user is in a handwashing state based on the spatial features, temporal features of the image sequence, and the pre-trained multi-layer perceptron model.
[0109] It can be seen that based on the first handwashing detection result obtained from the motion data collected by the IMU, the wearable device can further use the image sequence of the user's environment collected by the image sensor for the second handwashing detection to determine whether the user is actually in the handwashing state. In this way, by using the motion data collected by the IMU and the image sequence of the environment for the two handwashing detections respectively, the handwashing state of the user can be determined more accurately, avoiding misidentifying actions similar to handwashing, such as the user rubbing hands, as the user being in the handwashing state, and improving the accuracy of the handwashing detection result.
[0110] Moreover, when using the image sequence for handwashing detection, the spatial features and temporal features of the image sequence are extracted respectively, and the characteristics of the water flow passing through the hand, the motion state of the user's hand, and the change characteristics of the motion state of the water flow over time are captured simultaneously, further improving the accuracy of the second handwashing detection result, and thus further improving the accuracy of the finally determined handwashing detection result, that is, improving the accuracy of handwashing detection.
[0111] The wearable device can include various devices such as smart glasses and smart watches. The IMU in the wearable device can collect motion data such as ACC (Acceleration, accelerometer) data and GYRO (Gyroscope, gyroscope) data generated when the user moves. Among them, the ACC data is the acceleration data generated by the part of the body wearing the wearable device when the user moves, and the GYRO data is the angular velocity data generated by the part of the body wearing the wearable device when the user moves. The ACC data and GYRO data collected by the IMU can also be collectively referred to as six-axis signals, including the linear acceleration in 3 dimensions and the angular velocity in 3 dimensions of an object in 3D space.
[0112] The wearable device can use the motion data collected by the IMU to detect whether the user is in the handwashing state. Specifically, the wearable device can extract features from the six-axis signals to obtain features that can better characterize the handwashing action in the time domain and frequency domain, including 8 features such as the mean, variance, energy, peak value, root mean square, zero crossing rate in the time domain, and the maximum peak frequency and maximum peak amplitude in the frequency domain. The wearable device can put these features into a pre-trained SVM (Support Vector Machine) classification model for the first handwashing detection to obtain the first handwashing detection result.
[0113] However, when the user makes actions similar to handwashing, such as rubbing hands, the motion data such as the acceleration data and angular velocity data collected by the IMU are also similar. When the wearable device uses the motion data collected by the IMU for handwashing detection, it will be misdetected that the user is washing hands, resulting in an incorrect handwashing detection result.
[0114] Therefore, in order to improve the accuracy of handwashing detection, after performing a handwashing detection using the data collected by the IMU and obtaining a handwashing detection result indicating that the user is in a handwashing state, the wearable device can also perform another handwashing detection through the image sequence collected by the image sensor. The image sequence can be used to distinguish actions similar to handwashing, such as rubbing hands, made by the user. When it is determined that the user is in a handwashing state through the handwashing detection using the data collected by the IMU, the wearable device can execute step S101 to obtain the image sequence collected by the image sensor and perform a second handwashing detection.
[0115] The above-mentioned image sensor can be a camera. The image sensor can be installed at a position on the wearable device where it can collect the handwashing image sequence. In one example, the wearable device is a smartwatch, the image sensor is a camera, and the camera can be installed on the side of the smartwatch or at the position of the crown.
[0116] When the first handwashing detection result indicates that the user is in a handwashing state, the wearable device can obtain the image sequence collected by the image sensor.
[0117] After the wearable device obtains the image sequence, it executes step S102, that is, inputs the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence. Among them, the spatial features are used to characterize the characteristics of the picture of the water flow passing through the hand, and the temporal features are used to characterize the motion state of the user's hand and the characteristics of the change of the motion state of the water flow over time.
[0118] The image sequence includes multiple frames of images of the user's hand. The feature extraction model can be a three-dimensional convolutional neural network model. The feature extraction model can extract spatial features from the image sequence based on its own model parameters, and the spatial features can characterize the characteristics of the picture of the water flow passing through the hand corresponding to the water flow, the hand, and the positional relationship between the water flow and the hand. The feature extraction model can also extract temporal features from consecutive frames of the image sequence based on its own model parameters, and the temporal features can characterize the characteristics of the change of the water flow and the change of the user's hand corresponding to the motion state of the user's hand and the change of the motion state of the water flow over time. The model parameters of the feature extraction model are obtained through training with a large number of sample image sequences during the training process and can accurately extract spatial features and temporal features.
[0119] After extracting the spatial features and temporal features of the image sequence, the wearable device can further execute step S103, that is, use a pre-trained multi-layer perceptron model to obtain the handwashing detection result of the user based on the spatial features, temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model.
[0120] Among them, the handwashing detection result of the user represents whether the user is in the handwashing state, and the corresponding relationship is the corresponding relationship between the sample space feature and the sample time feature of the training sample and the handwashing state.
[0121] In one implementation, the handwashing detection result can be a result label. The result label can be a symbol or a number set according to actual needs. For example, the result label can include "1" and "0". The handwashing detection result being the result label "1" represents that the user is in the handwashing state; the handwashing detection result being the result label "0" represents that the user is not in the handwashing state correspondingly.
[0122] In another implementation, the handwashing detection result can also be a confidence level, that is, a probability value. For example, the handwashing detection result can be any probability value between 0 and 1. When the probability value in the handwashing detection result is less than 0.7, it represents that the user is not in the handwashing state; when the probability value in the handwashing detection result is greater than or equal to 0.7, it represents that the user is in the handwashing state.
[0123] The wearable device inputs the space feature and the time feature into a pre-trained multi-layer perceptron model. The multi-layer perceptron model can include multiple network layers, and these multiple network layers perform mapping calculations on the space feature and the time feature, and finally obtain the handwashing detection result of the user according to the mapping results of the multiple network layers in the multi-layer perceptron model.
[0124] As an implementation of the embodiment of the present application, the image sequence can be multiple consecutive sub-image sequences collected within a preset time period. In this case, as Figure 2 shown, the above step S103, that is, the step of using the pre-trained multi-layer perceptron model to obtain the handwashing detection result of the user based on the space feature, the time feature, and the corresponding relationship learned during the training process of the multi-layer perceptron model, can include the following steps S201 - S203.
[0125] S201, for each sub-image sequence, use the pre-trained multi-layer perceptron model to obtain the second handwashing detection result corresponding to the sub-image sequence based on the space feature, the time feature, and the corresponding relationship learned during the training process of the multi-layer perceptron model.
[0126] S202, when there are no less than a preset number of target results among the multiple second handwashing detection results, determine that the handwashing detection result of the user is that the user is in the handwashing state.
[0127] Among them, the target result represents that the user is in the handwashing state.
[0128] S203, when the number of target results among the multiple second handwashing detection results is less than the preset number, determine that the handwashing detection result of the user is that the user is not in the handwashing state.
[0129] Each sub - image sequence included in the image sequence may include a continuous plurality of frames of images. In one embodiment, the multiple consecutive sub - image sequences included in the image sequence may be sub - image sequences collected within the same - duration time period, and each sub - image sequence includes a continuous plurality of images with the same number of frames. For example, the image sequence includes 5 consecutive sub - image sequences collected within 5 seconds. Among them, the first sub - image sequence includes 12 consecutive frames of images collected within the first second, the second sub - image sequence includes 12 consecutive frames of images collected within the second second, and so on. Each sub - image sequence includes 12 consecutive frames of images collected within 1 second.
[0130] In order to further improve the accuracy of hand - washing detection, the wearable device can perform hand - washing detection on each of the multiple sub - image sequences included in the image sequence to obtain multiple second hand - washing detection results, that is, execute step S201. The wearable device can input the spatial features and temporal features corresponding to each sub - image sequence into a pre - trained multi - layer perceptron model respectively. Based on the spatial features and temporal features corresponding to the input sub - image sequence and the corresponding relationship learned during the training process of the multi - layer perceptron model, the multi - layer perceptron model obtains the second hand - washing detection result corresponding to the sub - image sequence.
[0131] After obtaining the second hand - washing detection result corresponding to each sub - image sequence, the wearable device can determine the number of second hand - washing detection results indicating that the user is in the hand - washing state and judge the magnitude relationship between the determined number and the preset number.
[0132] When the determined number is greater than or equal to the preset number, that is, when there are no less than the preset number of target results among the multiple second hand - washing detection results, the wearable device can determine that the hand - washing detection result of the user is that the user is in the hand - washing state, that is, execute step S202.
[0133] When the number of second hand - washing detection results indicating that the user is in the hand - washing state is lower than the preset number, the wearable device can determine that the hand - washing detection result of the user is that the user is not in the hand - washing state, that is, execute step S203.
[0134] In one embodiment, the above - mentioned preset number may be the total number of second hand - washing detection results. If each second hand - washing detection result indicates that the user is in the hand - washing state, the wearable device can determine that the hand - washing detection result of the user is that the user is in the hand - washing state; if there is a second hand - washing detection result indicating that the user is not in the hand - washing state, the wearable device determines that the hand - washing detection result of the user is that the user is not in the hand - washing state. By setting the preset number in this way, only when each second hand - washing detection result indicates that the user is in the hand - washing state, the wearable device will determine that the hand - washing detection result of the user is that the user is in the hand - washing state, which can improve the accuracy of hand - washing detection.
[0135] In another embodiment, the above preset quantity may also be a quantity less than the total quantity of the second handwashing detection results set according to actual needs. For example, if the total quantity of the second handwashing detection results is 5, the preset quantity may be set to 4. When there are 4 or 5 target results in the second handwashing detection results, the wearable device may determine that the user's handwashing detection result is that the user is in a handwashing state. When there are 1-3 target results in the second handwashing detection results, the wearable device may determine that the user's handwashing detection result is that the user is not in a handwashing state. Compared with directly setting the preset quantity to the total quantity of the second handwashing detection results, setting the preset quantity in this way can avoid the influence of individual abnormal second handwashing detection results on the user's handwashing detection result, and improve the fault tolerance rate of handwashing detection.
[0136] It can be seen that in the embodiments of the present application, the wearable device obtains multiple second handwashing detection results through multiple sub-image sequences included in the image sequence and the multi-layer perception model, and then determines the user's handwashing detection result based on these multiple second handwashing detection results. Compared with directly obtaining a handwashing detection result by using the image sequence and the multi-layer perception model, accidental errors can be reduced, and the influence of some abnormal data on the finally obtained handwashing detection result can be avoided, thereby improving the accuracy of handwashing detection.
[0137] As an embodiment of the embodiments of the present application, the multi-layer perception model may include an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer. In this case, as Figure 3 shown, the above step S201, that is, the step of using the pre-trained multi-layer perception model to obtain the second handwashing detection result corresponding to the sub-image sequence based on the spatial feature, time feature corresponding to the sub-image sequence, and the corresponding relationship learned during the training process of the multi-layer perception model, may include the following steps S301-S304.
[0138] S301, input the linear feature vector corresponding to the sub-image sequence into the input layer.
[0139] The linear feature vector corresponding to the sub-image sequence is obtained by compressing the spatial feature and time feature corresponding to the sub-image sequence, or output by the feature extraction model.
[0140] S302, the input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer.
[0141] S303, the hidden layer receives the linear feature vector, uses multiple neurons in at least one fully connected layer and the activation function to activate the non-linear relationship in the linear feature vector, and inputs the activated linear feature vector into the output layer.
[0142] In S304, the output layer receives the activated linear feature vector, and uses a normalization function to perform a handwashing state mapping on the activated linear feature vector to obtain a second handwashing detection result.
[0143] Since the features in the high-dimensional spatial features and temporal features are relatively scattered, and there are local arrangement positional relationships, when processing the high-dimensional spatial features and temporal features, it is also necessary to process the positional relationship information of these scattered features, resulting in low processing efficiency. To improve the processing efficiency, the high-dimensional spatial features and high-dimensional temporal features corresponding to the sub-image sequence can be compressed into a linear feature vector, integrating the high-dimensional scattered features into a compact global feature vector.
[0144] In one implementation, the wearable device can use a compression algorithm to compress the spatial features and temporal features to obtain a linear feature vector. In one implementation, the compression algorithm can be the flatten algorithm, and the wearable device can use the flatten algorithm to compress the spatial features and temporal features to obtain the linear feature vectors of the spatial features and temporal features.
[0145] In one implementation, the wearable device can also use a feature extraction model to compress the spatial features and temporal features. The feature extraction model can include a feature compression layer, such as the feature compression layer can be a flatten layer. After the wearable device inputs the image sequence into the feature extraction model, the feature extraction model first extracts features from the image sequence to obtain spatial features and temporal features. Then, the feature compression layer in the feature extraction model compresses the extracted spatial features and temporal features to obtain a linear feature vector.
[0146] After obtaining the linear feature vector, the wearable device can input the linear feature vector into the input layer of the multi-layer perceptron model, that is, perform step S301.
[0147] After the input layer receives the linear feature vector, it can preprocess the linear feature vector, convert the linear feature vector into the data format required by the hidden layer, and then input the processed data into the hidden layer, that is, perform step S302. The preprocessing can include normalization processing, dimension adjustment, etc.
[0148] The activation function can include at least one of the following: ReLU function, Sigmoid function, Tanh function. After the hidden layer receives the linear feature vector input by the input layer, based on the multiple neurons in at least one fully connected layer, it uses the activation function to process the linear feature vector, activates the non-linear relationship in the linear feature vector, and then the hidden layer inputs the activated linear feature vector into the output layer, that is, perform step S303.
[0149] The normalization function can be set according to actual needs. For example, the normalization function can be a Softmax function or a Sigmoid function, etc.
[0150] After the output layer receives the activated feature vector input from the hidden layer, it can use the normalization function to perform a handwashing state mapping on the activated linear feature vector to obtain the second handwashing detection result, that is, step S304 is executed.
[0151] In one implementation, the second handwashing detection result can be the confidence that the user is in the handwashing state. The normalization function performs a handwashing state mapping on the activated linear feature vector to obtain a confidence between 0 and 1.
[0152] In another implementation, the second handwashing detection result can also be a result label indicating whether the user is in the handwashing state. The result label can be a symbol or number set according to actual needs. For example, the result label can be "0" and "1". The normalization function performs a handwashing state mapping on the activated linear feature vector to obtain a result label of "0" or "1". Among them, a result label of "0" can indicate that the user is not in the handwashing state, and a result label of "1" can indicate that the user is in the handwashing state.
[0153] It can be seen that in the embodiments of the present application, the multi-layer perceptron model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer. The wearable device can output the linear feature vector to the input layer of the multi-layer perceptron model. The input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer. The hidden layer uses multiple neurons in at least one fully connected layer and an activation function to activate the non-linear relationship in the linear feature vector and inputs the activated linear feature vector into the output layer. The output layer uses the normalization function to perform a handwashing state mapping on the activated linear feature vector to obtain the second handwashing detection result. Each network layer of the multi-layer perceptron model processes the linear feature vector corresponding to the sub-image sequence in turn, which can fully obtain the information related to the handwashing state in the linear feature vector and improve the accuracy of the second handwashing detection result.
[0154] As an implementation of the present embodiment, as Figure 4 shown, the training method of the multi-layer perceptron model can include the following steps S401-S404.
[0155] The training method of the multi-layer perceptron model provided by the embodiments of the present application can be applied to a first device, and the first device can be a computer or a server, etc.
[0156] S401, obtain the linear feature vector corresponding to each training sample and the calibration label of each training sample.
[0157] Among them, the multiple training samples include a sequence of sample images collected by an image sensor of a wearable device when the user is in a handwashing state, and a sequence of sample images collected by the image sensor when the user is not in a handwashing state. The calibration label of each training sample represents the result of whether the training sample is in a handwashing state.
[0158] S402, Input the linear feature vector of each training sample into the initial multi-layer perceptron model.
[0159] S403, Obtain the predicted label output by the initial multi-layer perceptron model based on the current model parameters for processing the linear feature vector of the training sample.
[0160] S404, Adjust the model parameters of the initial multi-layer perceptron model based on the difference between the predicted label and the corresponding calibration label until the initial multi-layer perceptron model converges to obtain a trained multi-layer perceptron model.
[0161] The training samples may include sample video data of multiple users washing hands under different conditions and sample video data of hands not in a handwashing state. The sample video data can also be referred to as a sequence of sample images. Different conditions for washing hands can include: different water flow rates, whether to use hand sanitizer, or different handwashing durations, etc. The sample video data of hands not in a handwashing state can include: video data of the user's hand swinging in the air, video data of the user's hand when taking a vehicle, video data of the user's hand when cycling, and video data of the user's hand when walking, writing, running, and brushing teeth, etc. The calibration label can be a symbol or number set according to actual needs. For example, the calibration labels include "0" and "1", where "0" represents that the user is not in a handwashing state, and "1" represents that the user is in a handwashing state.
[0162] To train a multi-layer perceptron model for handwashing detection, the first device can obtain these training samples and the calibration label of each training sample, and obtain the linear feature vector corresponding to these training samples.
[0163] In one implementation, the first device can first use a feature extraction model to extract the spatial and temporal features of the sequence of sample images, and then use a compression algorithm to compress the extracted spatial and temporal features into a linear feature vector.
[0164] In another implementation, the first device can also input the sequence of sample images into a feature extraction model including a feature compression layer to directly obtain the linear feature vector output by the feature extraction model. Specifically, the method of obtaining the linear feature vector described in step S301 above can be referred to, and details are not described here.
[0165] After obtaining the linear feature vectors, the first device may input the linear feature vectors of each training sample obtained into the initial multi-layer perceptron model. The initial multi-layer perceptron model may process the linear feature vectors of the training samples based on the current model parameters and output predicted labels.
[0166] The first device may obtain the predicted labels output by the initial multi-layer perceptron model, and then adjust the model parameters of the initial multi-layer perceptron model based on the difference between the predicted labels and the corresponding calibration labels until the initial multi-layer perceptron model converges, obtaining a trained multi-layer perceptron model.
[0167] The first device may use a preset loss function to calculate a loss value based on the preset labels and the corresponding calibration labels, and determine whether the loss value is greater than a preset loss value threshold. When the loss value is greater than the preset loss value threshold, the first device adjusts the model parameters of the initial multi-layer perceptron model according to the loss value; when the loss value is less than or equal to the preset loss value threshold, the first device may determine that the model converges, obtaining a trained multi-layer perceptron model.
[0168] After the first device adjusts the model parameters of the initial multi-layer perceptron model to new model parameters, it may use the adjusted initial multi-layer perceptron model and new training samples to perform a new round of training according to steps S402 - S404 until the model converges.
[0169] It can be seen that in the embodiments of the present application, the first device uses a large number of sample image sequences collected by the image sensor of the wearable device to train the multi-layer perceptron model, and can fully learn the corresponding relationship between the sample space features, sample time features of the sample image sequences and the handwashing state. Subsequently, the trained multi-layer perceptron model can be used to more accurately detect the handwashing state corresponding to any image sequence, improving the accuracy of handwashing detection.
[0170] As an implementation manner of the embodiments of the present application, as Figure 5 shown, the embodiments of the present application further provide a handwashing detection method, which may further include the following steps S501 - S506.
[0171] S501, in response to the first handwashing detection result indicating that the user is in the handwashing state, obtain the image sequence collected by the image sensor, which is the same as step S101 above.
[0172] S502, input the image sequence into a pre-trained feature extraction model to extract the spatial features and time features of the image sequence, which is the same as step S102 above.
[0173] S503, use an activation function to activate the non-linear relationship between the spatial features and time features corresponding to each sub-image sequence respectively.
[0174] S504, perform three-dimensional pooling on the activated spatial features and temporal features corresponding to each sub-image sequence respectively, and reduce the resolution of the feature map composed of the activated spatial features and temporal features corresponding to each sub-image sequence to a preset first resolution.
[0175] S505, based on a preset dimensionality reduction function, reduce the dimension of the feature map with the first resolution corresponding to each sub-image sequence to one dimension, and obtain a linear feature vector corresponding to each sub-image sequence.
[0176] S506, using a pre-trained multi-layer perceptron model, based on the spatial features, temporal features, and the corresponding relationships learned during the training process of the multi-layer perceptron model, obtain the handwashing detection result of the user, which is the same as step S103 above.
[0177] The feature extraction model may include: an input layer, a convolutional layer, an activation layer, a pooling layer, and an output layer.
[0178] The wearable device can input the sub-image sequence into the input layer of the feature extraction model. The input layer can preprocess the sub-image sequence, such as normalization, standardization, etc. The convolutional layer can extract the spatial features and temporal features of the preprocessed sub-image sequence.
[0179] After obtaining the spatial features and temporal features of the sub-image sequence extracted by the convolutional layer, the convolutional layer can input the spatial features and temporal features corresponding to the sub-image sequence into the activation layer. The activation layer can use an activation function to activate the non-linear relationship of the spatial features and temporal features corresponding to the sub-image sequence, and input the activated spatial features and temporal features corresponding to the sub-image sequence into the pooling layer.
[0180] The pooling layer can perform three-dimensional pooling on the activated spatial features and temporal features corresponding to the sub-image sequence, so as to reduce the resolution of the feature map composed of the activated spatial features and temporal features corresponding to the sub-image sequence to a preset first resolution on the basis of enhancing the translation characteristics and deformation invariance of the features, and obtain a feature map with the first resolution. The preset first resolution can be set according to actual needs. For example, the first resolution can be 1 / 2 of the original resolution.
[0181] After obtaining the pooled feature map, the pooling layer can input the feature map with the first resolution corresponding to the sub-image sequence into the output layer, and the output layer reduces the dimension of the feature map with the first resolution corresponding to the sub-image sequence to one dimension based on a preset dimensionality reduction function, and obtains a linear feature vector corresponding to the sub-image sequence.
[0182] For each sub-image sequence, the wearable device can use the feature extraction model to obtain the linear feature vector corresponding to the sub-image sequence according to the above process.
[0183] It can be seen that in the embodiment of the present application, after the wearable device extracts the spatial features and temporal features of the image sequence, for each sub-image sequence, operations such as activation, pooling, and dimensionality reduction are performed on the spatial features and temporal features corresponding to the sub-image sequence. In this way, while fully retaining the complex non-linear relationships contained in the spatial features and temporal features corresponding to the sub-image sequence, the data volume can be reduced, thereby saving storage space and saving the processing time for obtaining the handwashing detection result using the linear feature vector subsequently, and thus improving the efficiency of handwashing detection.
[0184] As an implementation manner of the embodiment of the present application, the above step S102, that is, the step of inputting the image sequence into the pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, may include the following steps:
[0185] Input the image sequence into the pre-trained feature extraction model;
[0186] Use the three-dimensional convolution kernel in the feature extraction model to slide in the spatial dimension of the image sequence, and based on the content of the frames included in the image sequence, extract the spatial features of the image sequence. And use the three-dimensional convolution kernel to slide in the temporal dimension, and based on the law of change of the content included in the image sequence over time, extract the temporal features of the image sequence.
[0187] The pre-trained feature extraction model can be a three-dimensional convolutional neural network model. The feature extraction model can include multiple three-dimensional convolution kernels. By using the three-dimensional convolution kernels to slide in the spatial dimension and temporal dimension of the image sequence respectively, the spatial features and temporal features of the image sequence can be extracted respectively.
[0188] In one example, the feature extraction model can use multiple 3*3*3 three-dimensional convolution kernels, that is, a filter bank of 3*3*3, to slide in the spatial dimension and temporal dimension of the image sequence respectively to extract the spatial features and temporal features of the image sequence. Among them, each filter can extract a kind of feature, such as edge features, color block features in the spatial dimension, and dynamic features in the temporal dimension, etc.
[0189] Next, the training method of the feature extraction model will be described. The training method of the feature extraction model provided by the embodiment of the present application can be applied to a second device, and the second device can be a computer or a server, etc.
[0190] The second device can obtain multiple sample image sequences, as well as the calibrated spatial features and calibrated temporal features corresponding to each sample image sequence. Among them, the multiple sample image sequences include the sample image sequences collected by the image sensor of the wearable device when the user is in the handwashing state, and the sample image sequences collected by the image sensor when the user is not in the handwashing state. The calibrated temporal features and calibrated spatial features of each sample image sequence can be obtained by using a pre-trained feature extraction model.
[0191] The second device can input each sample image sequence into the initial feature extraction model to obtain the predicted spatial features and predicted temporal features output by the initial feature extraction model based on the current model parameters for processing the sample image sequence.
[0192] After that, the second device can adjust the model parameters of the initial feature extraction model based on the difference between the predicted spatial features and the corresponding calibrated spatial features, and the difference between the predicted temporal features and the corresponding calibrated temporal features, until the initial feature extraction model converges, and a trained feature extraction model is obtained.
[0193] It can be seen that in the embodiments of the present application, the feature extraction model uses multiple three-dimensional convolutional kernels, which can fully extract the spatial features and temporal features of the image sequence, thereby improving the accuracy of subsequent handwashing detection based on the spatial features and temporal features of the image sequence.
[0194] As an implementation manner of the embodiments of the present application, as Figure 6 shown, the embodiments of the present application further provide a handwashing detection method, which may further include the following steps S601-S606.
[0195] S601, in response to the first handwashing detection result indicating that the user is in the handwashing state, obtain the image sequence collected by the image sensor, which is the same as step S101 above.
[0196] S602, use a bilateral filter to denoise each frame of the image sequence included in the image sequence to obtain a denoised image sequence;
[0197] S603, perform downsampling processing on the denoised image sequence to reduce the image sequence to a preset second resolution to obtain a downsampled image sequence;
[0198] S604, perform normalization processing on the downsampled image sequence to obtain an image sequence for inputting into the pre-trained feature extraction model.
[0199] S605, input the image sequence into the pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, which is the same as step S102 above.
[0200] S606. Using the pre-trained multi-layer perception model, based on the spatial features, temporal features, and the corresponding relationships learned during the training process of the multi-layer perception model, obtain the handwashing detection result of the user, which is the same as step S103 above.
[0201] Due to environmental factors, the image sequence will contain noise data. The wearable device can use a bilateral filter to denoise each frame of the image sequence included in the image sequence to obtain a denoised image sequence.
[0202] To reduce the amount of data for subsequent processing, the wearable device can also perform downsampling on the denoised image sequence, reducing the image sequence to a preset second resolution to obtain a downsampled image sequence. For example, if the image sequence collected by the image sensor has a resolution of 224×224, the second resolution can be set to 112×112, and the wearable device can downsample the 224×224 denoised image sequence to a 112×112 image sequence.
[0203] The wearable device can further perform normalization on the downsampled image sequence to obtain an image sequence for inputting into the pre-trained feature extraction model, further reducing the amount of data for subsequent processing.
[0204] It can be seen that in the embodiments of the present application, before extracting the spatial features and temporal features of the image sequence, the wearable device performs denoising processing on the image sequence, which can reduce the impact of environmental noise on the image sequence, and then improve the accuracy of subsequent handwashing detection based on the image sequence. The wearable device also performs processing such as downsampling and normalization on the image sequence, which can reduce the amount of data for subsequent processing, thereby improving the efficiency of handwashing detection.
[0205] As an implementation manner of the embodiments of the present application, the first handwashing detection result can be the confidence level of the user being in the handwashing state. In this case, step S101 above, that is, the step of obtaining the image sequence collected by the image sensor in response to the first handwashing detection result indicating that the user is in the handwashing state, can include:
[0206] When the confidence level is higher than the preset threshold, activate the image sensor so that the image sensor collects the image sequence of the user's environment.
[0207] After the wearable device obtains the first handwashing detection result, it can determine whether the first handwashing detection result indicates that the user is in a handwashing state. When the first handwashing detection result indicates that the user is in a handwashing state, the wearable device can activate the image sensor so that the image sensor collects an image sequence of the user's environment. Then, the wearable device can execute step S101 to perform a second handwashing detection. When the first handwashing detection result indicates that the user is not in a handwashing state, the wearable device does not need to perform a second handwashing detection and thus does not need to activate the image sensor.
[0208] It can be seen that in the embodiments of the present application, only when the first handwashing detection result indicates that the user is in a handwashing state, the wearable device activates the image sensor, which can save the power consumption required by the image sensor, thereby reducing the power consumption required for handwashing detection and improving the user experience.
[0209] As an implementation manner of the embodiments of the present application, after the above step S103, that is, after the step of obtaining the user's handwashing detection result, the method further includes:
[0210] When the handwashing detection result indicates that the user is in a handwashing state, sending a first prompt message to the user and starting a handwashing timer. When the timing duration reaches a preset duration, sending a second prompt message to the user, where the first prompt message is used to prompt the user to start the handwashing timer, and the second prompt message is used to prompt the user to end the handwashing timer; and / or,
[0211] When the handwashing detection result indicates that the user is not in a handwashing state, turning off the image sensor.
[0212] When the handwashing detection result indicates that the user is in a handwashing state, the wearable device can send a first prompt message to the user in a way such as vibration, screen lighting, or ringing to prompt the user to start the handwashing timer. The preset duration can be the duration required for healthy handwashing. When the timing duration reaches the preset duration, the wearable device can send a second prompt message to the user in a way such as vibration, screen lighting, or ringing to prompt the user to end the handwashing timer.
[0213] In one implementation manner, in order to distinguish the first prompt message from the second prompt message, their prompt methods can be different. For example, the first prompt message is sent in a vibration manner, and the second prompt message is sent in a ringing manner; or, the first prompt message uses ringtone 1 for ringing prompt, and the second prompt message uses ringtone 2 for ringing prompt. Specifically, it can be set according to the actual needs of the user. Here, only examples are given and no limitation is intended.
[0214] When the handwashing detection result indicates that the user is not in a handwashing state, in order to save power consumption, the wearable device can turn off the image sensor.
[0215] It can be seen that in the embodiments of the present application, when the user is in a handwashing state, the wearable device can time the handwashing duration for the user to assist the user in health management. When the user is not in a handwashing state, the wearable device can timely turn off the image sensor to avoid unnecessary power consumption and improve the user experience.
[0216] As an implementation manner of the embodiments of the present application, as Figure 7 shown, the embodiments of the present application further provide a training method for an SVM classification model for obtaining a first handwashing detection result, which may include the following steps S701 - S704.
[0217] The training method of the multi - layer perception model provided by the embodiments of the present application can be applied to a third device, and the third device can be a computer or a server, etc.
[0218] S701, obtain the data collected by the IMU corresponding to each training sample and the calibration label of each training sample.
[0219] Among them, multiple training samples include the data collected by the IMU of the wearable device when the user is in a handwashing state, and the data collected by the IMU when the user is not in a handwashing state. The calibration label of each training sample represents the result of whether the training sample is in a handwashing state.
[0220] S702, input the data collected by the IMU corresponding to each training sample into the initial SVM classification model.
[0221] S703, obtain the predicted label output by the initial SVM classification model based on the current model parameters for processing the data collected by the IMU corresponding to the training sample.
[0222] S704, adjust the model parameters of the initial SVM classification model based on the difference between the predicted label and the corresponding calibration label until the initial SVM classification model converges to obtain the trained SVM classification model.
[0223] The training samples may include six - axis signals collected by the IMU of the wearable device when multiple users wash their hands under different conditions and six - axis signals collected by the IMU when the users are not in a handwashing state.
[0224] The different conditions for handwashing can include: different water flow rates, whether to use hand sanitizer, or different handwashing durations, etc. The six-axis signals collected when not in the handwashing state can include: the six-axis signals collected when the user waves their hand, rides a bicycle, takes a vehicle, walks, writes, runs, and brushes their teeth. The calibration label can be a symbol or number set according to actual needs. For example, the calibration labels include "0" and "1", where "0" represents that the user is not in the handwashing state, and "1" represents that the user is in the handwashing state.
[0225] To train the SVM classification model for handwashing detection, the third device can obtain the six-axis signals collected by the IMU corresponding to these training samples and the calibration label of each training sample.
[0226] The third device can input the obtained six-axis signals into the initial SVM classification model to obtain the predicted label output by the initial SVM classification model based on the current model parameters for processing the data collected by the IMU corresponding to the training samples.
[0227] The third device can use a preset loss function to calculate the loss value based on the predicted label and the corresponding calibration label, and determine whether the loss value is greater than the preset loss value threshold. When the loss value is greater than the preset loss value threshold, the third device adjusts the model parameters of the initial SVM classification model according to the loss value; when the loss value is less than or equal to the preset loss value threshold, the third device can determine that the model converges and obtain the trained SVM classification model.
[0228] After the third device adjusts the model parameters of the initial SVM classification model to new model parameters, it can use the adjusted initial SVM classification model and new training samples to perform a new round of training according to steps S702 - S704 until the model converges.
[0229] To more clearly illustrate the handwashing detection method provided by the embodiments of the present application, an example is also provided below, as Figure 8 shown, this example includes the following steps S801 - S810. Among them, steps S801 - S804 are executed by the IMU status recognition module in the wearable device, and steps S805 - S810 are executed by the video detection module in the wearable device.
[0230] S801, Input of raw data.
[0231] S802, Filtering of raw data.
[0232] S803, Feature extraction.
[0233] The original data is the six-axis signal data collected by the IMU, that is, the acceleration data in 3 dimensions and the angular velocity data in 3 dimensions. After the IMU status recognition module obtains the original data collected by the IMU, it can filter the original data and extract features, and then input the extracted features into a pre-trained SVM model for the first handwashing detection to obtain the first handwashing detection result. The first handwashing detection result represents the confidence level that the user is in the handwashing state.
[0234] S804, whether the confidence level is greater than the threshold; if so, execute step S805; if not, execute step S801.
[0235] When the confidence level is greater than the threshold, it indicates that the first handwashing detection result represents that the user is in the handwashing state. The video detection module activates the image sensor to collect the video data stream, that is, the image sequence. When the confidence level is less than or equal to the threshold, the wearable device executes step S801 to continue the handwashing detection based on the data collected by the IMU.
[0236] S805, input of the video data stream.
[0237] S806, bilateral filtering denoising of a single-frame image.
[0238] After activating the image sensor, the video detection module can obtain the video data stream collected by the image sensor. The image sensor can be a front-end camera with a resolution of 224×224, and the frame rate can be set to 12 frames per second. The video detection module can use a bilateral filter to denoise each frame image in the obtained video data stream, such as steps S601 - S602 above. The video detection module can further perform downsampling processing on the denoised video to obtain video data with a resolution of 112×112.
[0239] S807, spatio-temporal feature extraction of 3D-CNN (3Dimensional-Convolutional Neural Network).
[0240] After denoising each frame image in the video data stream, the video detection module can use a three-dimensional convolutional neural network to further extract the temporal features and spatial features of the video data stream, such as step S102 above. The video detection module can also compress the extracted spatio-temporal features to obtain a linear feature vector with a length of 4096.
[0241] S808, classification of MLP (Multi Layer Perceptron).
[0242] The video detection module uses a pre-trained MLP model to obtain the user's handwashing detection result based on the spatio-temporal features extracted in step S807 and the corresponding relationships learned during the training process of the multi-layer perception model, as in step S103 above. The MLP model adopted in the embodiments of the present application can be composed of three fully connected layers, which can fully learn the information related to the handwashing state in the spatio-temporal features.
[0243] S809, Handwashing scenario, start handwashing timing.
[0244] S810, Non-handwashing scenario, the video module is turned off.
[0245] According to the user's handwashing detection result obtained in step S808, in the handwashing scenario where the user is in the handwashing state, the video detection module starts handwashing timing; in the non-handwashing scenario where the user is not in the handwashing state, the video detection module is turned off.
[0246] In the embodiments of the present application, the data collected by the IMU is used for the first handwashing detection, and then the video data collected by the image sensor is used for the second handwashing detection. Compared with the handwashing detection that combines the data collected by the IMU and the data collected by the microphone, it will not be affected by noise in scenarios such as rubbing noise and background noise, improving the accuracy of handwashing detection.
[0247] In the technical solution of the present application, operations such as the acquisition, storage, use, processing, transmission, provision, and disclosure of the user's personal information are all carried out with the user's authorization.
[0248] As Figure 9 shown, a handwashing detection device is applied to a wearable device. The wearable device includes a motion sensor and an image sensor, and the motion sensor includes an IMU; the device includes:
[0249] An image sequence acquisition module 901, configured to acquire an image sequence collected by the image sensor in response to a first handwashing detection result indicating that the user is in a handwashing state, where the first handwashing detection result is obtained by performing handwashing detection using the motion data collected by the IMU;
[0250] A feature extraction module 902, configured to input the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, where the spatial features are used to characterize the picture features of the water flow passing through the hand, and the temporal features are used to characterize the motion state of the user's hand and the characteristics of the change of the motion state of the water flow over time;
[0251] A result determination module 903, configured to use a pre-trained multi-layer perception model to obtain a handwashing detection result of a user based on spatial features, temporal features, and a corresponding relationship learned during the training process of the multi-layer perception model, where the handwashing detection result of the user represents whether the user is in a handwashing state, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the handwashing state.
[0252] It can be seen that in the embodiment of the present application, when the first handwashing detection result obtained based on the motion data collected by the IMU indicates that the user is in a handwashing state, the wearable device further obtains an image sequence collected by the image sensor; and then based on the image sequence, the pre-trained feature extraction model, and the pre-trained multi-layer perception model, the final handwashing detection result is obtained.
[0253] Since the pre-trained feature extraction model learns how to extract spatial features and temporal features in the image sequence through a large number of training samples during the training process, the wearable device can accurately extract the spatial features and temporal features in the image sequence by using the pre-trained feature extraction model. Similarly, the pre-trained multi-layer perception model learns the corresponding relationship between the sample spatial features and sample temporal features and the handwashing state through a large number of training samples during the training process, and the wearable device can accurately predict whether the user is in a handwashing state based on the spatial features, temporal features of the image sequence, and the pre-trained multi-layer perception model.
[0254] It can be seen that based on the first handwashing detection result obtained by using the motion data collected by the IMU, the wearable device can further use the image sequence of the user's surrounding environment collected by the image sensor of the wearable device to perform a second handwashing detection to determine whether the user is actually in a handwashing state. In this way, by using the motion data collected by the IMU and the image sequence of the surrounding environment for two handwashing detections respectively, the handwashing state of the user can be determined more accurately, avoiding misidentifying actions similar to handwashing actions such as the user rubbing hands as the user being in a handwashing state, and improving the accuracy of the handwashing detection result.
[0255] Moreover, since when performing handwashing detection using the image sequence, the spatial features and temporal features of the image sequence are extracted respectively, and the characteristics of the water flow passing through the hand, the motion state of the user's hand, and the characteristics of the change of the motion state of the water flow over time are captured simultaneously, the accuracy of the second handwashing detection result is further improved, and thus the accuracy of the finally determined handwashing detection result is further improved, that is, the accuracy of handwashing detection is improved.
[0256] As an implementation manner of the embodiment of the present application, the image sequence is a plurality of consecutive sub-image sequences collected within a preset time period;
[0257] The above result determination module 903 can specifically be used for:
[0258] For each sub-image sequence, using a pre-trained multi-layer perceptron model, based on the spatial features, temporal features corresponding to the sub-image sequence, and the corresponding relationships learned during the training process of the multi-layer perceptron model, obtain the second handwashing detection result corresponding to the sub-image sequence;
[0259] In the case that there are no less than a preset number of target results among multiple second handwashing detection results, determine that the user's handwashing detection result is that the user is in a handwashing state, where the target result indicates that the user is in a handwashing state;
[0260] In the case that the number of target results among multiple second handwashing detection results is less than the preset number, determine that the user's handwashing detection result is that the user is not in a handwashing state.
[0261] As an implementation manner of an embodiment of the present application, the multi-layer perceptron model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer;
[0262] The above result determination module 903 can specifically be used for:
[0263] Input the linear feature vector corresponding to the sub-image sequence into the input layer, where the linear feature vector corresponding to the sub-image sequence is obtained by compressing the spatial features and temporal features corresponding to the sub-image sequence, or output by the feature extraction model;
[0264] The input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer;
[0265] The hidden layer receives the linear feature vector, uses multiple neurons in at least one fully connected layer and an activation function to activate the non-linear relationship in the linear feature vector, and inputs the activated linear feature vector into the output layer;
[0266] The output layer receives the activated linear feature vector, and uses a normalization function to perform a handwashing state mapping on the activated linear feature vector to obtain the second handwashing detection result.
[0267] As an implementation manner of an embodiment of the present application, the training method of the multi-layer perceptron model can include:
[0268] Obtain the linear feature vector corresponding to each training sample and the calibration label of each training sample, where multiple training samples include sample image sequences collected by the image sensor of the wearable device when the user is in a handwashing state, and sample image sequences collected by the image sensor when the user is not in a handwashing state, and the calibration label of each training sample represents the result of whether the training sample is in a handwashing state;
[0269] Input the linear feature vectors of each training sample into the initial multi-layer perceptron model;
[0270] Obtain the predicted labels output by the initial multi-layer perceptron model after processing the linear feature vectors of the training samples based on the current model parameters;
[0271] Based on the difference between the predicted labels and the corresponding calibration labels, adjust the model parameters of the initial multi-layer perceptron model until the initial multi-layer perceptron model converges to obtain the trained multi-layer perceptron model.
[0272] As an implementation manner of the embodiment of the present application, the above handwashing detection device may further include:
[0273] A post-processing module, configured to, after the step of extracting the spatial features and temporal features of the image sequence, use an activation function to respectively activate the non-linear relationships of the spatial features and temporal features corresponding to each sub-image sequence; perform three-dimensional pooling on the activated spatial features and temporal features corresponding to each sub-image sequence respectively, and reduce the resolution of the feature map composed of the activated spatial features and temporal features corresponding to each sub-image sequence to a preset first resolution; based on a preset dimensionality reduction function, reduce the dimension of the feature map with the first resolution corresponding to each sub-image sequence to one dimension to obtain the linear feature vector corresponding to each sub-image sequence.
[0274] As an implementation manner of the embodiment of the present application, the above feature extraction module 902 may specifically be configured to:
[0275] Input the image sequence into a pre-trained feature extraction model;
[0276] Use the three-dimensional convolution kernel in the feature extraction model to slide in the spatial dimension of the image sequence, and based on the content of the pictures included in the image sequence, extract the spatial features of the image sequence, and use the three-dimensional convolution kernel to slide in the temporal dimension, and based on the law of the change of the content included in the image sequence over time, extract the temporal features of the image sequence.
[0277] As an implementation manner of the embodiment of the present application, the above handwashing detection device may further include:
[0278] A pre-processing module, configured to, before the step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, use a bilateral filter to denoise each frame image included in the image sequence to obtain a denoised image sequence;
[0279] Perform downsampling processing on the denoised image sequence to reduce the image sequence to a preset second resolution to obtain a downsampled image sequence;
[0280] Normalize the downsampled image sequence to obtain an image sequence for inputting into a pre-trained feature extraction model.
[0281] As an implementation manner of an embodiment of the present application, the first handwashing detection result is the confidence level of the user being in the handwashing state; the above-mentioned image sequence acquisition module 901 can specifically be used for:
[0282] When the confidence level is higher than a preset threshold, activate the image sensor so that the image sensor acquires an image sequence of the user's environment.
[0283] As an implementation manner of an embodiment of the present application, the above-mentioned handwashing detection device may further include:
[0284] A timing module, configured to, after obtaining the user's handwashing detection result, when the handwashing detection result indicates that the user is in the handwashing state, send a first prompt message to the user and start handwashing timing, and when the timing duration reaches a preset duration, send a second prompt message to the user, where the first prompt message is used to prompt the user to start handwashing timing, and the second prompt message is used to prompt the user to end handwashing timing.
[0285] As an implementation manner of an embodiment of the present application, the above-mentioned handwashing detection device may further include:
[0286] An image sensor closing module, configured to close the image sensor when the handwashing detection result indicates that the user is not in the handwashing state.
[0287] An embodiment of the present application further provides a wearable device, as Figure 10 shown, including:
[0288] An image sensor 1001, configured to acquire an image sequence;
[0289] A motion sensor 1002, the motion sensor includes an IMU, and the IMU is configured to acquire motion data;
[0290] A memory 1003, configured to store a computer program;
[0291] A processor 1004, configured to implement any of the above-mentioned handwashing detection methods when executing the program stored on the memory 1003.
[0292] And the above-mentioned wearable device may further include a communication bus and / or a communication interface, and the processor 1004, the communication interface, and the memory 1003 complete mutual communication through the communication bus.
[0293] The communication bus mentioned in the above wearable device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0294] The communication interface is used for communication between the above wearable device and other devices.
[0295] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0296] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0297] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above handwashing detection methods are implemented.
[0298] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute any of the handwashing detection methods in the above embodiments.
[0299] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.
[0300] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0301] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, wearable device, computer-readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0302] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A handwashing detection method, characterized in that, Applied to a wearable device, the wearable device includes a motion sensor and an image sensor, and the motion sensor includes an inertial measurement unit (IMU); the method includes: In response to a first handwashing detection result indicating that the user is in a handwashing state, obtaining an image sequence collected by the image sensor, where the first handwashing detection result is obtained by performing handwashing detection using the motion data collected by the IMU; Inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, where the spatial features are used to characterize the characteristics of the picture of water flowing through the hand, and the temporal features are used to characterize the motion state of the user's hand and the characteristics of the change of the motion state of the water flow over time; Using a pre-trained multi-layer perceptron model, based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model, obtaining the handwashing detection result of the user, where the handwashing detection result of the user indicates whether the user is in a handwashing state, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the handwashing state.
2. The method according to claim 1, wherein The image sequence is a plurality of consecutive sub-image sequences collected within a preset time period; The step of using a pre-trained multi-layer perceptron model to obtain the handwashing detection result of the user based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model includes: For each sub-image sequence, using a pre-trained multi-layer perceptron model, based on the spatial features, temporal features corresponding to the sub-image sequence, and the corresponding relationship learned during the training process of the multi-layer perceptron model, obtaining a second handwashing detection result corresponding to the sub-image sequence; In the case where there are no less than a preset number of target results among the multiple second handwashing detection results, determining that the handwashing detection result of the user is that the user is in a handwashing state, where the target result indicates that the user is in a handwashing state; In the case where the number of target results among the multiple second handwashing detection results is less than the preset number, determining that the handwashing detection result of the user is that the user is not in a handwashing state.
3. The method according to claim 2, wherein The multi-layer perceptron model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer; The step of using a pre-trained multi-layer perceptron model to obtain a second handwashing detection result corresponding to the sub-image sequence based on the spatial features, temporal features corresponding to the sub-image sequence, and the corresponding relationship learned during the training process of the multi-layer perceptron model includes: Inputting the linear feature vector corresponding to the sub-image sequence into the input layer, where the linear feature vector corresponding to the sub-image sequence is obtained by compressing the spatial features and temporal features corresponding to the sub-image sequence, or is output by the feature extraction model; The input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer; The hidden layer receives the linear feature vector, activates the non-linear relationship in the linear feature vector by using multiple neurons in the at least one fully-connected layer and an activation function, and inputs the activated linear feature vector into the output layer; The output layer receives the activated linear feature vector, and uses a normalization function to perform a handwashing state mapping on the activated linear feature vector to obtain a second handwashing detection result.
4. The method according to claim 3, wherein The training method of the multi-layer perceptron model includes: Obtaining a linear feature vector corresponding to each training sample and a calibration label of each training sample, where the multiple training samples include a sample image sequence collected by an image sensor of the wearable device when the user is in a handwashing state, and a sample image sequence collected by the image sensor when the user is not in a handwashing state, and the calibration label of each training sample represents the result of whether the training sample is in a handwashing state; Inputting the linear feature vector of each training sample into an initial multi-layer perceptron model; Obtaining a predicted label output by the initial multi-layer perceptron model based on the current model parameters for processing the linear feature vector of the training sample; Adjusting the model parameters of the initial multi-layer perceptron model based on the difference between the predicted label and the corresponding calibration label until the initial multi-layer perceptron model converges to obtain a trained multi-layer perceptron model.
5. The method according to claim 3, characterized in that After the step of extracting the spatial features and temporal features of the image sequence, the method further includes: Using an activation function to activate the non-linear relationships of the spatial features and temporal features corresponding to each sub-image sequence respectively; Performing three-dimensional pooling on the activated spatial features and temporal features corresponding to each sub-image sequence respectively, and reducing the resolution of the feature map composed of the activated spatial features and temporal features corresponding to each sub-image sequence to a preset first resolution; Based on a preset dimensionality reduction function, reducing the dimension of the feature map with the first resolution corresponding to each sub-image sequence to one dimension to obtain a linear feature vector corresponding to each sub-image sequence.
6. The method according to claim 1, characterized in that, The step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence includes: Inputting the image sequence into a pre-trained feature extraction model; Using a three-dimensional convolution kernel in the feature extraction model to slide in the spatial dimension of the image sequence, extracting the spatial features of the image sequence based on the content of the frames included in the image sequence, and using the three-dimensional convolution kernel to slide in the temporal dimension, and extracting the temporal features of the image sequence based on the law of change of the content included in the image sequence over time.
7. The method according to any one of claims 1 to 6, characterized in that, Before the step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, the method further includes: Using a bilateral filter to perform noise reduction on each frame image included in the image sequence to obtain a noise-reduced image sequence; Performing downsampling processing on the noise-reduced image sequence to reduce the image sequence to a preset second resolution to obtain a downsampled image sequence; Normalize the downsampled image sequence to obtain an image sequence for inputting into a pre-trained feature extraction model.
8. The method according to any one of claims 1 to 6, characterized in that, The first handwashing detection result is the confidence level that the user is in the handwashing state. The step of obtaining the image sequence collected by the image sensor in response to the first handwashing detection result indicating that the user is in the handwashing state includes: When the confidence level is higher than a preset threshold, activate the image sensor so that the image sensor collects an image sequence of the environment where the user is located.
9. The method according to any one of claims 1-6, characterized in that, After the step of obtaining the handwashing detection result of the user, the method further includes: When the handwashing detection result indicates that the user is in the handwashing state, send a first prompt message to the user and start a handwashing timer. When the timing duration reaches a preset duration, send a second prompt message to the user, where the first prompt message is used to prompt the user to start the handwashing timer, and the second prompt message is used to prompt the user to end the handwashing timer; and / or, When the handwashing detection result indicates that the user is not in the handwashing state, turn off the image sensor.
10. A wearable device, characterized in that, including: An image sensor for collecting an image sequence; A motion sensor, the motion sensor includes an inertial detection unit IMU, and the IMU is used to collect motion data; A memory for storing computer programs; A processor, when executing the programs stored on the memory, implements the method according to any one of claims 1-9.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1-9.
Citation Information
Cited By
Tooth brushing action detection method, device, equipment and medium
CN121545213A