Travel state detection method, wearable device and storage medium
By using image sensors and multi-layer perception models to extract the spatial and timing characteristics of the image sequence in wearable devices, the problem of low travel status detection accuracy when riding a vehicle is solved, and higher travel status detection accuracy is achieved.
Patent Information
- Application Number
- CN202510331644.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-04
AI Technical Summary
In wearable devices, the acceleration signal collected by the IMU changes when the user rides a vehicle, resulting in a low accuracy in travel status detection.
Using the image sequences collected by the image sensor, spatial and timing features are extracted through pre-trained feature extraction models, combined with multi-layer perception models, the probability of users riding on vehicles is determined, and the accuracy of travel status detection is improved.
Through two travel status detection, it is possible to accurately distinguish whether the user is riding on a transportation tool, avoid misidentification, and improve the accuracy of travel status detection.
Smart Images

Figure CN120259614A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of status detection, and particularly to a method for detecting travel status, a wearable device, and a storage medium. Background Art
[0002] Wearable devices are a general term for electronic devices developed through intelligent design of daily wear, such as smart glasses, smart gloves, smart watches, etc. Due to the advantages of being lightweight and easy to carry, more and more users choose to use wearable devices for sports tracking.
[0003] Wearable devices can detect the travel status of users by analyzing the acceleration signals collected by the built-in IMU (Inertial Measurement Unit). However, when the user travels by car, the acceleration signals of the IMU will also change, and the accuracy of detecting the user's travel status using the acceleration signals collected by the IMU is low. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a method for detecting travel status, a wearable device, and a storage medium to improve the accuracy of travel status detection results. The specific technical solutions are as follows:
[0005] In a first aspect, the embodiments of this application provide a method for detecting travel status, which is applied to a wearable device. The wearable device includes a motion sensor and an image sensor, and the motion sensor includes an IMU; the method includes:
[0006] In response to the accuracy of the first travel status detection result being lower than a preset threshold, obtain the image sequence collected by the image sensor, where the first travel status detection result is obtained by detecting the travel status using the motion data collected by the IMU;
[0007] Input the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, where the spatial features are used to characterize the characteristics of the user's environment, and the temporal features are used to characterize the characteristics of the user's motion state changing over time;
[0008] Use a pre-trained multi-layer perceptron model to obtain a second travel status detection result based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model, where the second travel status detection result represents the probability that the user is taking a means of transportation, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the travel status;
[0009] Determine a target travel status detection result based on the probability of the user taking a means of transportation.
[0010] Optionally, the step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence includes:
[0011] Input the image sequence into a pre-trained feature extraction model;
[0012] Use a three-dimensional convolutional kernel in the feature extraction model to slide in the spatial dimension of the image sequence, and based on the environmental content included in the image sequence, extract the spatial features of the image sequence. Then use the three-dimensional convolutional kernel to slide in the time dimension, and based on the law of change of the content included in the image sequence over time, extract the temporal features of the image sequence.
[0013] Optionally, the multi-layer perceptron model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer;
[0014] The step of using a pre-trained multi-layer perceptron model to obtain a second travel status detection result based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model includes:
[0015] Input a linear feature vector into the input layer, where the linear feature vector is obtained by compressing the spatial features and the temporal features, or is output by the feature extraction model;
[0016] The input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer;
[0017] The hidden layer receives the linear feature vector, uses multiple neurons in the at least one fully connected layer and an activation function to activate the non-linear relationship in the linear feature vector, and inputs the activated linear feature vector into the output layer;
[0018] The output layer receives the activated linear feature vector, and uses a normalization function to perform a travel status mapping on the activated linear feature vector to obtain a second travel status detection result.
[0019] Optionally, the training method of the multi-layer perceptron model includes:
[0020] Obtain the linear feature vectors corresponding to each training sample and the calibration label of each training sample, where the multiple training samples include the image sequences collected by the image sensor of the wearable device when the user is taking a means of transportation, and the image sequences collected by the image sensor when the user is not taking a means of transportation and is exercising, and the calibration label of each training sample represents the travel state corresponding to this training sample;
[0021] Input the linear feature vectors of each training sample into the initial multi-layer perceptron model;
[0022] Obtain the predicted labels output by the initial multi-layer perceptron model based on the current model parameters for processing the linear feature vectors of the training samples;
[0023] Adjust the model parameters of the initial multi-layer perceptron model based on the difference between the predicted labels and the corresponding calibration labels until the initial multi-layer perceptron model converges to obtain the trained multi-layer perceptron model.
[0024] Optionally, after the step of extracting the temporal features of the image sequence, the method further includes:
[0025] Use an activation function to activate the non-linear relationship between the spatial features and the temporal features;
[0026] Perform three-dimensional pooling on the activated spatial features and temporal features to reduce the resolution of the feature map composed of the activated spatial features and temporal features to a preset resolution;
[0027] Based on a preset dimensionality reduction function, reduce the dimension of the feature map with the preset resolution to one dimension to obtain the linear feature vector.
[0028] Optionally, before the step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, the method further includes:
[0029] Perform noise reduction on each frame of the image included in the image sequence to obtain a noise-reduced image sequence;
[0030] Use the optical flow method to process the noise-reduced image sequence to obtain the motion relationship between each frame of the image in the noise-reduced image sequence;
[0031] Perform alignment processing on each frame of the image in the noise-reduced image sequence based on the motion relationship, and perform weighted average filtering on the aligned image sequence to obtain a filtered image sequence;
[0032] Perform normalization processing on the filtered image sequence to obtain an image sequence for inputting into the pre-trained feature extraction model.
[0033] Optionally, the step of determining the target travel status detection result based on the probability of the user taking a vehicle includes:
[0034] When the probability of the user taking a vehicle is greater than a preset probability threshold, determining that the target travel status detection result is taking a vehicle;
[0035] When the probability of the user taking a vehicle is not greater than the preset probability threshold, determining that the target travel status detection result is the travel status indicated by the first travel status detection result.
[0036] Optionally, the motion sensor further includes a photoelectric sensor; before the step of obtaining the image sequence collected by the image sensor in response to the accuracy of the first travel status detection result being lower than a preset threshold, the method further includes:
[0037] Obtaining the motion signal collected by the motion sensor, where the motion signal includes the acceleration signal collected by the IMU, that is, the motion data collected by the IMU, or the acceleration signal collected by the IMU and the heart rate signal collected by the photoelectric sensor;
[0038] Extracting the time-domain features of the motion signal and performing travel status detection based on the time-domain features to obtain the first travel status detection result.
[0039] Optionally, the first travel status detection result is the confidence level corresponding to the user's travel status being running or the confidence level corresponding to walking;
[0040] The step of obtaining the image sequence collected by the image sensor in response to the accuracy of the first travel status detection result being lower than a preset threshold includes:
[0041] When the confidence level is lower than the preset threshold, starting the image sensor so that the image sensor collects an image sequence of the user's environment.
[0042] In a second aspect, an embodiment of the present application provides a travel status detection device, which is applied to a wearable device. The wearable device includes a motion sensor and an image sensor, and the motion sensor includes an inertial detection unit IMU; the device includes:
[0043] An image sequence acquisition module, configured to obtain the image sequence collected by the image sensor in response to the accuracy of the first travel status detection result being lower than a preset threshold, where the first travel status detection result is obtained by performing travel status detection using the motion data collected by the IMU;
[0044] A feature extraction module, configured to input the image sequence into a pre-trained feature extraction model, and extract the spatial features and temporal features of the image sequence, where the spatial features are used to characterize the characteristics of the user's environment, and the temporal features are used to characterize the characteristics of the change of the user's motion state over time;
[0045] A first result determination module, configured to use a pre-trained multi-layer perception model to obtain a second travel state detection result based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perception model, where the second travel state detection result characterizes the probability that the user takes a means of transportation, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the travel state;
[0046] A second result determination module, configured to determine a target travel state detection result based on the probability that the user takes a means of transportation.
[0047] In a third aspect, an embodiment of the present application provides a wearable device, including:
[0048] An image sensor, configured to collect an image sequence;
[0049] A motion sensor, the motion sensor includes an inertial measurement unit (IMU); the IMU is configured to collect motion data;
[0050] A memory, configured to store a computer program;
[0051] A processor, when executing the program stored on the memory, implements the method steps of any one of the first aspects described above.
[0052] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method steps of any one of the first aspects described above are implemented.
[0053] In a fifth aspect, an embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute the method steps of any one of the first aspects described above.
[0054] Advantageous effects of the embodiments of the present application:
[0055] In the technical solution provided by the embodiments of the present application, the wearable device may obtain an image sequence collected by an image sensor in response to the accuracy of the first travel state detection result being lower than a preset threshold, where the first travel state detection result is obtained by performing travel state detection using the motion data collected by the IMU; input the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, where the spatial features are used to characterize the characteristics of the user's environment, and the temporal features are used to characterize the characteristics of the user's motion state changing over time; use a pre-trained multi-layer perceptron model to obtain a second travel state detection result based on the spatial features, temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model, where the second travel state detection result characterizes the probability that the user is taking a means of transportation, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the travel state; determine the target travel state detection result based on the probability that the user is taking a means of transportation.
[0056] It can be seen that when the first travel state detection result obtained using the motion data collected by the IMU is inaccurate, the image sequence of the user's environment collected by the image sensor of the wearable device can be used for further travel state detection to determine the probability that the user is taking a means of transportation. In this way, by performing two travel state detections using the acceleration signal and the image sequence of the environment respectively, it is possible to more accurately determine whether the user is traveling by means of transportation or walking / running, avoiding misidentifying the travel state of the user taking a means of transportation as the user walking / running, and improving the accuracy of the travel state detection result. Moreover, since the spatial features and temporal features of the image sequence are extracted respectively when performing travel state detection using the image sequence, the characteristics of the user's environment and the characteristics of the user's motion state changing over time are captured simultaneously, further improving the accuracy of the travel state detection result, and thus further improving the accuracy of the actual travel state detection result of the user. Of course, implementing any product or method of the present application does not necessarily require achieving all of the above advantages at the same time. Description of the Drawings
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.
[0058] Figure 1 It is a flowchart of a method for detecting a travel state provided by an embodiment of the present application;
[0059] Figure 2Schematic diagram of a specific implementation manner of step S102;
[0060] Figure 3 Schematic diagram of a specific implementation manner of step S103;
[0061] Figure 4 Flow schematic diagram of the training method of the multi-layer perception model provided by the embodiment of the present application;
[0062] Figure 5 Flow schematic diagram of a detection example of a travel state provided by the embodiment of the present application;
[0063] Figure 6 Structural schematic diagram of a detection device for a travel state provided by the embodiment of the present application;
[0064] Figure 7 Structural schematic diagram of a wearable device provided by the embodiment of the present application. Detailed implementation manner
[0065] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0066] In the technical solutions of the present application, operations such as obtaining, storing, using, processing, transmitting, providing, and disclosing user personal information are all carried out under the condition of obtaining user authorization.
[0067] In order to improve the accuracy of the travel state detection result, the embodiment of the present application provides a travel state detection method, device, wearable device, computer-readable storage medium, and computer program product. First, a travel state detection method provided by the embodiment of the present application will be introduced below.
[0068] The travel state detection method provided by the embodiment of the present application can be applied to any wearable device, such as a smart watch, smart glasses, etc.
[0069] As Figure 1 shown, the travel state detection method provided by the embodiment of the present application is applied to a wearable device, and the wearable device includes a motion sensor and an image sensor, and the motion sensor includes an inertial detection unit IMU; the method includes:
[0070] S101: In response to the accuracy of the first travel state detection result being lower than a preset threshold, obtain an image sequence collected by the image sensor.
[0071] Among them, the first travel state detection result is obtained by performing travel state detection on the motion data collected by the IMU.
[0072] S102: Input the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence.
[0073] Among them, the spatial features are used to characterize the characteristics of the user's environment, and the temporal features are used to characterize the characteristics of the user's motion state changing over time.
[0074] S103: Use a pre-trained multi-layer perceptron model to obtain a second travel state detection result based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model.
[0075] Among them, the second travel state detection result represents the probability that the user is taking a means of transportation, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the travel state.
[0076] S104: Determine the target travel state detection result based on the probability that the user is taking a means of transportation.
[0077] In the technical solution provided by the embodiments of the present application, the wearable device can, in response to the accuracy of the first travel state detection result being lower than a preset threshold, obtain an image sequence collected by an image sensor. Among them, the first travel state detection result is obtained by performing travel state detection on the motion data collected by the IMU; input the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence. Among them, the spatial features are used to characterize the characteristics of the user's environment, and the temporal features are used to characterize the characteristics of the user's motion state changing over time; use a pre-trained multi-layer perceptron model to obtain a second travel state detection result based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model. Among them, the second travel state detection result represents the probability that the user is taking a means of transportation, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the travel state; determine the target travel state detection result based on the probability that the user is taking a means of transportation.
[0078] It can be seen that when the first travel status detection result obtained from the motion data collected by the IMU is inaccurate, the image sequence of the user's environment collected by the image sensor of the wearable device can be used for further travel status detection to determine the probability that the user is taking a means of transportation. In this way, by using the acceleration signal and the image sequence of the environment for two travel status detections respectively, it is possible to accurately determine whether the user is traveling by means of transportation or walking / running, avoiding misidentifying the travel status of the user taking a means of transportation as the user walking / running, and improving the accuracy of the travel status detection result. Moreover, since the spatial features and temporal features of the image sequence are extracted respectively when using the image sequence for travel status detection, the characteristics of the user's environment and the characteristics of the user's motion state changing over time are captured simultaneously, further improving the accuracy of the travel status detection result, and thus further improving the accuracy of the actual travel status detection result of the user.
[0079] The wearable device can include various devices such as smart glasses and smart watches, and the wearable device includes a motion sensor and an image sensor. Among them, the image sensor can collect images of the user's environment; the motion sensor can collect the user's motion signals, and the motion sensor includes an inertial detection unit IMU, and the IMU can collect the user's motion data.
[0080] The IMU includes an accelerometer and a gyroscope, and can collect motion data such as ACC (Acceleration, accelerometer) data and GYRO (Gyroscope, gyroscope) data generated when the user moves. Among them, the ACC data is the acceleration data generated by the part of the body wearing the wearable device when the user moves, and the GYRO data is the angular velocity data generated by the part of the body wearing the wearable device when the user moves. The ACC data and GYRO data collected by the IMU can also be collectively referred to as six-axis signals, including the linear acceleration of an object in 3 dimensions and the angular velocity in 3 dimensions in 3D space.
[0081] The wearable device can obtain the motion data collected by the IMU and use the data collected by the IMU to detect the user's travel status and determine whether the user is walking or running.
[0082] However, when a user travels by a vehicle, the motion data recorded by the IMU will also change. In this case, using the motion data recorded by the IMU for travel status detection may misdetect that the user is walking or running, resulting in an incorrect travel status detection result. Since the motion data displayed by the wearable device should be the motion data generated when the user is walking or running, when misdetecting travel by vehicle as walking or running, the user's motion data recorded by the wearable device is inaccurate, which affects the user's understanding of their true motion situation and thus affects the user experience.
[0083] Based on this, in order to improve the accuracy of the travel status detection result, after using the motion data collected by the IMU for travel status detection to obtain the first travel status detection result, it is possible to determine whether the first travel status detection result is accurate based on a preset threshold. Wherein, the preset threshold can be set according to actual needs. For example, when evaluating the accuracy of the first travel status detection result using the confidence level corresponding to walking or the confidence level corresponding to running, the preset threshold can be 0.8, 0.9, etc., and no specific limitation is made here.
[0084] When the accuracy of the first travel status detection result is lower than the preset threshold, the accuracy of the first travel status detection result is relatively low, and it is impossible to accurately determine whether the user is traveling by vehicle or traveling by running or walking indicated by the first travel status detection result.
[0085] Furthermore, in step S101, the wearable device can, in response to the accuracy of the first travel status detection result being lower than the preset threshold, obtain the image sequence collected by the image sensor. Wherein, the first travel status detection result is obtained by using the motion data collected by the IMU for travel status detection.
[0086] The wearable device can obtain the motion data collected by the IMU of the wearable device and use the motion data collected by the IMU for travel status detection to obtain the first travel status detection result of the user. Wherein, the first travel status detection result represents the accuracy corresponding to the user traveling by walking or the accuracy corresponding to the user traveling by running.
[0087] When the accuracy of the first travel status detection result is lower than the preset threshold, the accuracy of the first travel status detection result is relatively low, and it is impossible to accurately distinguish whether the user is traveling by vehicle or traveling in the travel status indicated by the first travel status detection result.
[0088] To improve the accuracy of the travel status detection result, the wearable device can obtain the image sequence collected by the image sensor and use the image sequence for travel status detection to determine the probability that the user is traveling by vehicle.
[0089] When detecting the travel state using an image sequence, the wearable device can execute step S102, input the image sequence into a pre-trained feature extraction model, and extract the spatial features and temporal features of the image sequence.
[0090] Among them, the spatial features are used to characterize the characteristics of the user's environment, and the temporal features are used to characterize the characteristics of the change of the user's motion state over time.
[0091] Each frame image in the image sequence includes the environmental content of the user's environment, and the change of the environmental content between multiple consecutive images in the image sequence can reflect the characteristics of the change of the user's motion state over time.
[0092] The wearable device can input the image sequence into a pre-trained feature extraction model to extract the spatial features of the image sequence for characterizing the characteristics of the user's environment, such as the people sitting inside the vehicle the user is taking, seats, armrests, and windows, etc.; and the temporal features for characterizing the characteristics of the change of the user's motion state over time, such as the characteristics of the continuous jitter of the vehicle the user is taking over time, etc.
[0093] Among them, the feature extraction model is a three-dimensional (3D, 3-Dimensional) model obtained by training the initial feature extraction model using multiple sample image sequences, and the calibrated spatial features and calibrated temporal features corresponding to each sample image sequence. This three-dimensional model can extract the spatial features and temporal features of the image features. For the sake of clear writing, the training method of the feature extraction model will be described in detail below.
[0094] Furthermore, in the above step S103, the wearable device can use a pre-trained multi-layer perceptron model to obtain a second travel state detection result based on the spatial features, temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model.
[0095] Among them, the second travel state detection result represents the probability that the user is taking a means of transportation, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the travel state.
[0096] In one implementation, the wearable device can input the spatial features and temporal features into a pre-trained multi-layer perceptron model. The multi-layer perceptron model processes the spatial features and temporal features using the corresponding relationship learned during the training process to obtain a second travel state detection result representing the probability that the user is taking a means of transportation output by the multi-layer perceptron model.
[0097] In one implementation, the wearable device can input the one-dimensional feature vectors of the spatial feature and the temporal feature into a pre-trained multi-layer perceptron model. The multi-layer perceptron model uses the corresponding relationships learned during the training process for the one-dimensional feature vectors of the spatial feature and the temporal feature to obtain a second travel state detection result representing the probability that the user is taking a transportation vehicle.
[0098] Among them, the multi-layer perceptron model is pre-trained by using the sample spatial features and sample temporal features of multiple training samples, and the calibration label of each training sample for the initial multi-layer perceptron model. Among them, the calibration label is the travel state corresponding to the sample spatial feature and sample temporal feature of each training sample.
[0099] After obtaining the second travel state detection result, the wearable device can execute step S104 to determine the target travel state detection result based on the probability that the user is taking a transportation vehicle.
[0100] When the probability that the user is taking a transportation vehicle represents that the user travels by taking a transportation vehicle, the wearable device can determine that the target travel state detection result is taking a transportation vehicle;
[0101] When the probability that the user is taking a transportation vehicle represents that the user does not travel by taking a transportation vehicle, the wearable device can determine that the target travel state detection result is the travel state indicated by the first travel state detection result, that is, the user walks or runs.
[0102] Since the spatial feature of the image sequence is used to represent the characteristics of the user's environment, and the temporal feature represents the characteristics of the change of the user's motion state in the current environment over time, the essence of the second travel state detection by the multi-layer perceptron model is to classify whether the scene where the user is located is a scene of taking a transportation vehicle, and obtain the detection result of whether the user travels by taking a transportation vehicle or does not travel by taking a transportation vehicle through the scene classification of the scene where the user is located.
[0103] In this way, scene detection is performed on the wearable device side through the video signal, assisting the first travel state detection result of the data collected by the IMU, improving the accuracy of the state detection result, and avoiding problems such as mis-identification of the number of steps and the movement distance caused by inaccurate travel state detection results.
[0104] It can be seen that in the case where the first travel state detection result obtained from the motion data collected by the IMU is inaccurate, the image sequence of the user's environment collected by the image sensor of the wearable device can be used for further travel state detection to determine the probability that the user is taking a vehicle. In this way, by performing two travel state detections using the acceleration signal and the image sequence of the environment respectively, it is possible to more accurately determine whether the user is traveling by vehicle or walking / running, avoiding misidentifying the travel state of the user taking a vehicle as the user walking / running, and improving the accuracy of the travel state detection result. Moreover, since when performing travel state detection using the image sequence, the spatial features and temporal features of the image sequence are extracted respectively, capturing both the characteristics of the user's environment and the characteristics of the user's motion state changing over time, the accuracy of the travel state detection result is further improved, and thus the accuracy of the actual travel state detection result of the user is further improved.
[0105] As an implementation manner of the embodiment of the present application, as Figure 2 shown, the above step S102, that is, the step of inputting the image sequence into the pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, may include the following steps:
[0106] S201: Input the image sequence into the pre-trained feature extraction model;
[0107] S202: Use the three-dimensional convolution kernel in the feature extraction model to slide in the spatial dimension of the image sequence, and based on the environmental content included in the image sequence, extract the spatial features of the image sequence, and use the three-dimensional convolution kernel to slide in the time dimension, and based on the law of change of the content included in the image sequence over time, extract the temporal features of the image sequence.
[0108] The pre-trained feature extraction model is a three-dimensional model, including a plurality of three-dimensional convolution kernels. By sliding in the spatial dimension and the time dimension, the three-dimensional convolution kernel can capture both the spatial features and temporal features of the image features.
[0109] When performing travel state detection using the image sequence, the wearable device can input the image sequence into the pre-trained feature extraction model, and use the three-dimensional convolution kernel in the feature extraction model to slide in the spatial dimension of the image sequence to extract the spatial features of the image sequence based on the environmental content included in the image sequence, and obtain the spatial features for characterizing the characteristics of the user's environment.
[0110] Use the three-dimensional convolution kernel in the feature extraction model to slide in the time dimension, and based on the law of change of the content included in the image sequence over time, to extract the temporal features of the image sequence, and obtain the temporal features for characterizing the characteristics of the user's motion state changing over time.
[0111] In one implementation, the feature extraction model is a 3D-CNN (Convolutional Neural Network) model, and the 3D-CNN model includes multiple 3D convolutional layers, and each convolutional layer automatically learns image features at different levels and abstraction degrees of the image sequence.
[0112] Exemplarily, the feature extraction model includes a filter bank (kernels) of 3×3×3. After the image sequence is input into the feature extraction model, each 3×3×3 filter can slide in the spatial dimension and the temporal dimension to extract a feature respectively. For example, filter A extracts the edge feature of the image sequence, filter B extracts the color blocks of the image sequence, and filter C extracts the dynamic feature of the environmental content in the vehicle in the temporal dimension, etc.
[0113] The training method of the feature extraction model will be described below. Among them, the electronic device used to train the feature extraction model and the wearable device that executes a detection method for a travel state provided in this application may be the same device or different devices.
[0114] The wearable device can obtain multiple sample image sequences and the calibrated spatial features and calibrated temporal features of each sample image sequence. Among them, the multiple sample image sequences include the image sequences collected by the image sensor of the wearable device when the user is taking a means of transportation, and the image sequences collected by the image sensor when the user is not taking a means of transportation and is moving.
[0115] The wearable device can input each sample image sequence into the initial feature extraction model, and obtain the predicted spatial features and predicted temporal features output by the initial feature extraction model based on the current model parameters for processing the sample image sequence.
[0116] After that, the wearable device can adjust the model parameters of the initial feature extraction model based on the differences between the predicted spatial features and predicted temporal features and the calibrated spatial features and calibrated temporal features respectively until the initial feature extraction model converges, and obtain the trained feature extraction model.
[0117] In this embodiment, the wearable device can input an image sequence into a pre-trained feature extraction model. By using the 3D convolutional kernels in the feature extraction model to slide in the spatial dimension and the temporal dimension of the image sequence respectively, the spatial features and temporal features of the image sequence are extracted, the efficiency of extracting the image features of the image sequence is improved, and further the detection efficiency of the travel state is improved. Moreover, since the spatial features and temporal features of the image sequence are extracted respectively, the characteristics of the user's environment and the characteristics of the user's motion state changing over time can be captured simultaneously. Using the spatial features and temporal features for travel state detection, the obtained travel state detection result has high accuracy.
[0118] As an implementation manner of the embodiment of the present application, the multi-layer perception model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer;
[0119] As Figure 3 shown, the above step S103, that is, the step of using the pre-trained multi-layer perception model to obtain the second travel state detection result based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perception model, may include the following steps:
[0120] S301: Input the linear feature vector into the input layer, where the linear feature vector is obtained by compressing the spatial features and the temporal features, or output by the feature extraction model;
[0121] S302: The input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer;
[0122] S303: The hidden layer receives the linear feature vector, uses multiple neurons in the at least one fully connected layer and an activation function to activate the non-linear relationship in the linear feature vector, and inputs the activated linear feature vector into the output layer;
[0123] S304: The output layer receives the activated linear feature vector, uses a normalization function to perform travel state mapping on the activated linear feature vector, and obtains the second travel state detection result.
[0124] The pre-trained multi-layer perception model (MLP, Multi-Layer Perception) may include an input layer, a hidden layer, and an output layer, and the hidden layer may include at least one fully connected layer, and each fully connected layer includes multiple neurons.
[0125] Since the spatial features and temporal features are high-dimensional feature vectors, storing high-dimensional feature vectors requires a large amount of storage space, and processing high-dimensional feature vectors requires a large amount of processing time. Therefore, in order to save storage space and processing time, the high-dimensional feature vectors can be compressed, and the high-dimensional feature vectors can be transformed into low-dimensional feature vectors to compress and store the low-dimensional feature vectors.
[0126] Based on this, the wearable device can obtain the linear feature vectors of the spatial features and temporal features, where the length of the linear feature vectors can be set according to actual needs. For example, it can be 4K, or 8K, etc., and no specific limitation is made here.
[0127] In one implementation, the wearable device can use the flatten (compression) algorithm to compress the spatial features and temporal features to obtain the linear feature vectors of the spatial features and temporal features.
[0128] In one implementation, the feature extraction model can include a flatten layer. After extracting the spatial features and temporal features, the spatial features and temporal features can be sent to the flatten layer for compression, and the linear feature vectors of the spatial features and temporal features are output. The wearable device can obtain the linear feature vectors output by the feature extraction model.
[0129] Then the wearable device can input the linear feature vectors into the input layer of the multi-layer perceptron model. The input layer can receive the linear feature vectors and input the linear feature vectors into the hidden layer.
[0130] The hidden layer can receive the linear feature vectors and, respectively, based on multiple neurons in at least one fully connected layer, use an activation function to process the linear feature vectors to activate the non-linear relationships in the linear feature vectors. Among them, the activation function can include ReLU, Sigmoid, Tanh, etc. The complex non-linear relationships in the linear feature vectors can be captured through the combination of neurons and activation functions.
[0131] After that, the hidden layer inputs the activated linear feature vectors into the output layer. The output layer receives the activated linear feature vectors and uses a normalization function to perform travel state mapping on the activated linear feature vectors, mapping the activated linear feature vectors to the corresponding classification neurons to obtain the second travel state detection result.
[0132] Among them, the normalization function can be set according to actual needs. For example, it can be the Softmax function, etc.
[0133] The number of neurons in the output layer is equal to the number of categories of travel states. In this application, the multi-layer perception model is used to detect the probability that the user takes a means of transportation. That is, the multi-layer perception model is a binary classification model, and the categories in the output layer include taking a means of transportation and not taking a means of transportation.
[0134] Using the multi-layer perception model, spatial features, and temporal features for travel state detection actually realizes obtaining the scene category of the scene where the user is located through an image sequence, that is, whether the scene where the user is located belongs to the scene of taking a means of transportation or does not belong to the scene of taking a means of transportation.
[0135] In this embodiment, the multi-layer perception model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer. The wearable device can input the linear feature vector into the input layer. The input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer; the hidden layer receives the linear feature vector, uses multiple neurons in at least one fully connected layer and an activation function to activate the non-linear relationship in the linear feature vector, and inputs the activated linear feature vector into the output layer; the output layer receives the activated linear feature vector and uses a normalization function to perform travel state mapping on the activated linear feature vector. In this way, using this multi-layer perception model, the second travel state detection result corresponding to the linear feature vector of the image sequence can be obtained, improving the detection efficiency of the second travel state and the accuracy of the second travel state detection result.
[0136] As an implementation manner of the embodiment of this application, as Figure 4 shown, the training method of the multi-layer perception model may include the following steps:
[0137] S401: Obtain the linear feature vector corresponding to each training sample and the calibration label of each training sample. Among them, multiple training samples include the image sequence collected by the image sensor of the wearable device when the user takes a means of transportation, and the image sequence collected by the image sensor when the user does not take a means of transportation and is moving. The calibration label of each training sample represents the travel state corresponding to this training sample;
[0138] S402: Input the linear feature vector of each training sample into the initial multi-layer perception model;
[0139] S403: Obtain the predicted label output by the initial multi-layer perception model based on the current model parameters for processing the linear feature vector of the training sample;
[0140] S404: Based on the difference between the predicted label and the corresponding calibration label, adjust the model parameters of the initial multi-layer perception model until the initial multi-layer perception model converges to obtain the trained multi-layer perception model.
[0141] To train a multi-layer perceptron model for detecting the probability that a user is taking a vehicle, the wearable device can respectively obtain the linear feature vectors corresponding to each training sample and the calibration label of each training sample.
[0142] Among them, the multiple training samples include the image sequences collected by the image sensor of the wearable device when the user is taking a vehicle, and the image sequences collected by the image sensor when the user is not taking a vehicle and is moving. The calibration label of each training sample represents the travel state corresponding to the training sample.
[0143] For the image sequences collected by the image sensor of the wearable device when the user is taking a vehicle, for example, video samples when the user is taking a car, video samples when the user is taking the subway, etc., its label represents that the travel state corresponding to the image sequence is taking a vehicle; for the image sequences collected by the image sensor when the user is not taking a vehicle and is moving, for example, video samples when the user is walking, video samples when the user is running, etc., its calibration label represents that the travel state corresponding to the image sequence is traveling without taking a vehicle.
[0144] It should be noted that the linear feature vectors corresponding to the training samples can be extracted by the electronic device for training the multi-layer perceptron model, or can be extracted by other electronic devices.
[0145] Moreover, the wearable device for implementing a travel state detection method provided in this application and the electronic device for training the multi-layer perceptron model can be the same device or different electronic devices.
[0146] In one implementation, the wearable device can respectively obtain the image sequences collected by the image sensor when the user is taking a vehicle, and the image sequences collected by the image sensor when the user is not taking a vehicle and is moving, take each image sequence as a training sample, and obtain the calibration label of each image sequence, and this calibration label represents the travel state corresponding to the training sample. Then the wearable device can input each training sample into a pre-trained feature extraction model, extract the sample space feature and sample time series feature of each training sample, and compress the sample space feature and sample time series feature into a linear feature vector.
[0147] Among them, the method for obtaining the calibration label of each image sequence includes outputting each image sequence and obtaining the calibration label of each image sequence marked by artificial / background server; inputting each image sequence into a pre-trained classification model and obtaining the classification result input by the classification model as the calibration label of this image sequence, etc.
[0148] In one implementation, the wearable device can respectively obtain the image sequences collected by the image sensor when the user is taking a vehicle, and the image sequences collected by the image sensor when the user is not taking a vehicle and is exercising. Then, the wearable device sends the image sequences collected by the image sensor to the background server, and obtains the linear feature vectors obtained by the background server processing each image sequence, as well as the calibration labels of each image sequence, as the linear feature vectors and calibration labels corresponding to each training sample, so as to directly obtain the linear feature vectors of multiple training samples.
[0149] Next, the wearable device inputs the linear feature vectors corresponding to each training sample into the initial multi-layer perceptron model, and obtains the predicted labels output by the initial multi-layer perceptron model based on the current model parameters for processing the linear feature vectors of the training samples.
[0150] Based on the difference between the predicted labels and the calibration labels corresponding to each training sample, and using this difference to adjust the model parameters of the initial multi-layer perceptron model until the initial multi-layer perceptron model converges, a trained multi-layer perceptron model is obtained.
[0151] In this embodiment, the image sequences of the user's travel status collected on-site by the image sensor of the wearable device are obtained as training samples, and the image features of the image sequences are extracted to form linear feature vectors. In this way, the electronic device uses the linear feature vectors of the training samples and the calibration labels of each training sample to train a multi-layer perceptron model. The multiple perceptron models learn the corresponding relationship between the sample space features, sample time series features and travel status of the training samples through deep learning, and have the ability to distinguish between the user taking a vehicle and the user not taking a vehicle for travel, and can realize binary classification of whether the user takes a vehicle for travel. In this way, the trained multi-layer perceptron model can detect whether the travel status corresponding to the linear feature vector of any image sequence is that the user takes a vehicle.
[0152] As an implementation manner of the embodiment of the present application, after the step of extracting the time series features of the image sequence, the method further includes:
[0153] Using an activation function to activate the non-linear relationship between the spatial features and the time series features;
[0154] Performing three-dimensional pooling on the activated spatial features and time series features to reduce the resolution of the feature map composed of the activated spatial features and time series features to a preset resolution;
[0155] Based on a preset dimensionality reduction function, reducing the dimension of the feature map with the preset resolution to one dimension to obtain the linear feature vector.
[0156] The feature extraction model can also include a convolutional layer, an activation layer, a pooling layer and an output layer.
[0157] After the convolutional layer extracts the spatial features and temporal features of the image sequence, the spatial features and temporal features can be input into the activation layer. The activation layer can utilize an activation function to activate the non-linear relationship between the spatial features and temporal features, obtain the activated spatial features and temporal features, and input the activated spatial features and temporal features into the pooling layer.
[0158] The pooling layer can perform three-dimensional pooling (Pooling) on the activated spatial features and temporal features to reduce the resolution of the feature map composed of the activated spatial features and temporal features to a preset resolution on the basis of enhancing the translational characteristics and deformation invariance of the features, and obtain a feature map with the preset resolution.
[0159] Among them, the preset resolution can be set according to actual needs, and can be a fixed resolution, for example, 300 ppi, etc.; it can also be a multiple of the original resolution, where the multiple is less than 1, for example, 1 / 2 of the original resolution, etc., and no specific limitation is made here.
[0160] After that, the feature map with the preset resolution is input into the output layer, and the output layer can reduce the dimension of the feature map with the preset resolution to one dimension based on a preset dimensionality reduction function, and obtain a linear feature vector corresponding to the spatial features and temporal features. Among them, the preset dimensionality reduction function can be a flatten function or other functions, and no specific limitation is made here.
[0161] In this embodiment, the wearable device can utilize an activation function to activate the non-linear relationship between the spatial features and temporal features, perform three-dimensional pooling on the activated spatial features and temporal features, and reduce the resolution of the feature map composed of the activated spatial features and temporal features to a preset resolution; and based on a preset dimensionality reduction function, reduce the dimension of the feature map with the preset resolution to one dimension to obtain a linear feature vector. By activating, pooling, and reducing the dimension of the spatial features and temporal features, the obtained linear feature vector includes complex non-linear relationships in the image sequence and has a smaller data volume, thereby saving storage space and subsequent feature vector processing time.
[0162] As an implementation manner of the embodiment of the present application, before the above step S102, that is, the step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, a method for detecting a travel state provided by the embodiment of the present application may further include:
[0163] Perform noise reduction on each frame of the image included in the image sequence to obtain a denoised image sequence;
[0164] Use the optical flow method to process the denoised image sequence to obtain the motion relationship between each frame of the denoised image sequence;
[0165] Align each frame image in the denoised image sequence based on the motion relationship, and perform weighted average filtering on the aligned image sequence to obtain a filtered image sequence;
[0166] Perform normalization processing on the filtered image sequence to obtain an image sequence for inputting into a pre-trained feature extraction model.
[0167] Under the influence of various factors such as environmental factors, device noise, and device jitter, the image sequence may have problems such as noise and misalignment between consecutive frames. Then, directly performing feature extraction on the image sequence collected by the image sensor will result in large errors in the extracted spatial features and temporal features. Using the spatial features and temporal features for travel state detection will result in poor accuracy of the obtained travel state detection results.
[0168] Based on this, after obtaining the image sequence collected by the image sensor, a single-frame denoising algorithm can be used to eliminate the noise points in each frame image of the image sequence, realizing the denoising of each frame image included in the image sequence to obtain a denoised image sequence. Among them, the single-frame denoising algorithm can be set according to actual needs. For example, median filtering algorithm, bilateral filtering algorithm, etc., are not specifically limited here.
[0169] After that, use the optical flow method to process the denoised image sequence to obtain the motion relationship between each frame image in the denoised image sequence.
[0170] Among them, the optical flow method can be set according to actual needs. For example, Lucas-Kanade, Farneback, etc., are not limited here. Using the optical flow method to process the image sequence can estimate the motion direction and speed of the pixel points on the surface of the object in each frame image of the image sequence, that is, the motion path of the pixel points on the surface of each object in the image sequence can be determined. According to the motion path of the pixel points on the surface of each object, the motion relationship between each frame image in the image sequence can be determined.
[0171] Then, the wearable device can align each frame image in the denoised image sequence according to the determined motion relationship between each frame image. For example, perform affine transformation on each frame image according to the motion relationship between each frame image to obtain an aligned image sequence.
[0172] To further improve the image quality of the image sequence, weighted average filtering can be performed on the aligned image sequence to obtain a filtered image sequence.
[0173] And perform normalization processing on the filtered image sequence, standardize the pixel values of the image sequence, that is, limit the pixel values within the range of (0, 1), and obtain the image sequence for inputting into the pre-trained feature extraction model.
[0174] In this embodiment, the wearable device can perform noise reduction on the image sequence including each frame of the image to obtain a noise-reduced image sequence, use the optical flow method to process the noise-reduced image sequence to obtain the motion relationship between each frame of the image in the noise-reduced image sequence, perform alignment processing on each frame of the image in the noise-reduced image sequence based on the motion relationship, and perform weighted average filtering on the aligned image sequence to obtain a filtered image sequence; perform normalization processing on the filtered image sequence. In this way, the obtained image sequence has higher image instructions and smaller data volume. Inputting this image sequence into the feature extraction model can reduce the processing time of the feature extraction model and improve the accuracy of the output features.
[0175] As an implementation manner of the embodiment of the present application, the step of determining the target travel status detection result based on the probability that the user takes a transportation means includes:
[0176] When the probability that the user takes a transportation means is greater than a preset probability threshold, determine that the target travel status detection result is taking a transportation means;
[0177] When the probability that the user takes a transportation means is not greater than the preset probability threshold, determine that the target travel status detection result is the travel status indicated by the first travel status detection result.
[0178] The second travel status detection result represents the probability that the user takes a transportation means. When the probability that the user takes a transportation means is greater than the preset probability threshold, it can be determined that the user travels by taking a transportation means.
[0179] And when the probability that the user takes a transportation means is not greater than the preset probability threshold, it can be determined that the user does not travel by taking a transportation means. Since the second travel status detection result excludes the possibility that the user travels by taking a transportation means, then the user travels by the travel status indicated by the first travel status detection result. Therefore, it can be determined that the target travel status detection result is the travel status indicated by the first travel status detection result.
[0180] Among them, the preset probability threshold can be set according to actual needs. For example, 0.8, 0.9, etc., and no specific limitation is made here.
[0181] Exemplarily, the travel status indicated by the first travel status detection result is running, and the preset probability threshold is 0.85.
[0182] If the probability that the second travel status detection result represents the user taking a means of transportation is 0.9, which is greater than the preset probability threshold of 0.85, then it can be determined that the target travel status detection result of the user is taking a means of transportation, that is, the user is not running.
[0183] If the probability that the second travel status detection result represents the user taking a means of transportation is 0.5, which is not greater than the preset probability threshold of 0.85, it can be determined that the target travel status detection result of the user is running.
[0184] In this embodiment, the wearable device can obtain the probability that the user takes a means of transportation by performing a second travel status detection on the image sequence. Furthermore, when the probability that the user takes a means of transportation is greater than the preset probability threshold, it is determined that the target travel status detection result is taking a means of transportation; when the probability that the user takes a means of transportation is not greater than the preset probability threshold, it is determined that the target travel status detection result is the travel status indicated by the first travel status detection result. In this way, by using the image sequence to determine whether the user travels by taking a means of transportation, it can be accurately determined whether the user travels by taking a means of transportation or runs or walks as indicated by the first travel status detection result, avoiding the user misidentifying the operation data of taking a means of transportation as the motion data of the user's own movement, improving the accuracy of the motion data of the user displayed by the wearable device, and further improving the user experience.
[0185] As an implementation manner of the embodiment of the present application, the motion sensor further includes a photoelectric sensor;
[0186] Before the above step S101, that is, the step of acquiring the image sequence collected by the image sensor in response to the accuracy of the first travel status detection result being lower than the preset threshold, a method for detecting a travel status provided by the present application may further include:
[0187] Acquire the motion signal collected by the motion sensor, where the motion signal includes the acceleration signal collected by the IMU, that is, the motion data collected by the IMU, or the acceleration signal collected by the IMU and the heart rate signal collected by the photoelectric sensor;
[0188] Extract the time-domain features of the motion signal, and perform travel status detection based on the time-domain features to obtain the first travel status detection result.
[0189] To detect the travel status of the user, the wearable device can acquire the motion signal collected by the motion sensor.
[0190] The motion sensors of the wearable device include an IMU and an optoelectronic sensor. Among them, the IMU can collect the acceleration information of the user, and this acceleration signal is the motion data of the user collected by the IMU. The optoelectronic sensor can collect the heart rate signal (PPG signal, Photoplethysmographic Signal) of the user.
[0191] Since both the acceleration and heart rate of the user can change as the user travels, the corresponding acceleration signal collected by the IMU and the heart rate signal collected by the optoelectronic sensor can both reflect the travel state of the user.
[0192] Based on this, the motion signal for travel state detection collected by the motion sensors of the wearable device can be the acceleration signal collected by the IMU, or the acceleration signal collected by the IMU and the heart rate signal collected by the optoelectronic sensor.
[0193] After that, the wearable device can extract the time-domain features of the motion signal and perform travel state detection based on the extracted time-domain features to obtain the first travel state detection result.
[0194] Among them, the time-domain features can include features such as zero-crossing rate, energy, and correlation. The zero-crossing rate (ZCR) can reflect the fluctuation frequency of the signal, and a high zero-crossing rate usually indicates strenuous exercise; Energy can reflect the total intensity of the signal, and a high energy usually indicates a large movement amplitude; Correlation can reflect the linear relationship between the three-axis signals, and a high correlation indicates a consistent movement direction.
[0195] By extracting the time-domain features of the motion signal, rich motion information can be obtained from the motion signal, providing support for subsequent travel state detection.
[0196] In one implementation, the wearable device can input the time-domain features into a pre-trained travel state detection model and obtain the first travel state detection result output by the travel state detection model. Among them, the travel state detection model can be an XGBoost (eXtreme Gradient Boosting) model, which is a machine learning model that has been greatly optimized based on the traditional gradient boosting algorithm.
[0197] In one implementation, the motion signal collected by the wearable device is the acceleration signal collected by the IMU. The IMU can continuously collect the user's acceleration signal at a first preset frequency and store the acceleration signal collected for the first preset duration in the buffer. Among them, the first preset frequency and the first preset duration can be set according to actual needs. For example, the first preset frequency can be 15Hz, 25Hz, etc., and the first preset duration can be 2 seconds, 8 seconds, etc., which are not specifically limited here.
[0198] The wearable device can obtain the acceleration signal of the first preset duration stored in the buffer, extract the time-domain features of the acceleration signal, and perform travel state detection based on the time-domain features to obtain the first travel state detection result.
[0199] Exemplarily, the IMU can collect the acceleration signal at a frequency of 25Hz and store 200 acceleration signals collected every 8 seconds in the buffer. The wearable device can obtain the acceleration signal of 8 seconds stored in the buffer, extract the time-domain features of the acceleration signal, and perform travel state detection based on the time-domain features to obtain the first travel state detection result.
[0200] In one implementation, the motion signal collected by the wearable device is the acceleration signal collected by the IMU and the heart rate signal collected by the photoelectric sensor. The IMU can continuously collect the user's acceleration signal at a first preset frequency and store the acceleration signal collected for the first preset duration in the buffer; the photoelectric sensor can continuously collect the user's heart rate signal at a second preset frequency and store the heart rate signal collected for the second preset duration in the buffer.
[0201] Among them, the first preset frequency and the first preset duration can be set according to actual needs. For example, the first preset frequency can be 15Hz, 25Hz, etc., and the first preset duration can be 2 seconds, 8 seconds, etc.; the second preset frequency and the second preset duration can also be set according to actual needs. For example, the second preset frequency can be 20Hz, 25Hz, 30Hz, etc., and the second preset duration can be 4 seconds, 8 seconds, etc.; and the first preset frequency and the second preset frequency can be the same or different; the first preset duration and the second preset duration can be the same or different, which are not specifically limited here.
[0202] The wearable device can obtain the acceleration signal of the first preset duration and the heart rate signal of the second preset duration stored in the buffer, and use a preset filter to filter the heart rate signal based on a preset filter bandwidth to obtain the filtered heart rate signal. Among them, the preset filter can be a Butterworth filter or other filters; the preset filter bandwidth can be set according to actual needs, for example, 0.5Hz - 4Hz, which is not specifically limited here.
[0203] After that, the time-domain features of the acceleration signal and the filtered heart rate signal are extracted, and the time-domain features are input into the XGBoost model to obtain the travel state output by the XGBoost model and the confidence corresponding to the travel state, so as to obtain the first travel state detection result.
[0204] Exemplarily, the IMU can collect acceleration signals at a frequency of 25 Hz and store 200 acceleration signals collected every 8 seconds in the buffer; the optoelectronic sensor can collect heart rate signals at a frequency of 10 Hz and store 80 heart rate signals collected every 8 seconds in the buffer.
[0205] The wearable device can obtain the cached acceleration signals for 8 seconds and 80 heart rate signals collected every 8 seconds. Using a Butterworth filter, the heart rate signal is filtered based on a filtering bandwidth of 0.5 Hz - 4 Hz to obtain the filtered heart rate signal. After that, the wearable device can extract the time-domain features of the acceleration signal and the filtered heart rate signal, and perform travel state detection based on the time-domain features to obtain the first travel state detection result.
[0206] In this embodiment, the motion sensor of the wearable device further includes an optoelectronic sensor. The wearable device can obtain the acceleration signal collected by the IMU, or the acceleration signal collected by the IMU and the heart rate signal collected by the optoelectronic sensor as the motion signal collected by the motion sensor, extract the time-domain features of the motion signal, and perform travel state detection based on the time-domain features to obtain the first travel state detection result. In this way, using the time-domain features of the motion signal collected by the motion sensor to detect the first travel state detection result of the user can improve the accuracy of the first travel state detection result.
[0207] As an implementation manner of the embodiment of the present application, the first travel state detection result is the confidence corresponding to the user's travel state being running or the confidence corresponding to walking.
[0208] The above step S101, that is, the step of obtaining the image sequence collected by the image sensor in response to the accuracy of the first travel state detection result being lower than the preset threshold, may include:
[0209] In the case where the confidence is lower than the preset threshold, the image sensor is activated so that the image sensor collects an image sequence of the user's environment.
[0210] The wearable device can use the data collected by the IMU to detect the user's travel mode. Among them, the travel modes that can be detected using the data collected by the IMU include running or walking. Furthermore, the first travel state detection result is the confidence corresponding to the user's travel state being running, or the first travel state detection result is the confidence corresponding to the user's travel state being walking.
[0211] In the case that the above confidence level is lower than a preset threshold, the wearable device can activate the image sensor. The image sensor can collect consecutive image frames of the environment where the user is located to obtain an image sequence, and send the image sequence to the electronic device.
[0212] The wearable device receives the image sequence collected by the image sensor, extracts the spatial features and temporal features of the image sequence, and performs travel state detection.
[0213] In one implementation, when the confidence level corresponding to the first travel state detection result is lower than the preset threshold, the wearable device activates the image sensor so that the image sensor collects an image sequence of the environment where the user is located and sends the image sequence to the electronic device. The wearable device can receive the image sequence and perform subsequent travel state detection based on the image sequence.
[0214] In one implementation, the image acquisition device collects an image sequence of the environment where the user is located in real time. When the confidence level corresponding to the first travel state detection result is lower than the preset threshold, the wearable device performs subsequent travel state detection using the image sequence corresponding to the IMU data for obtaining the first travel state detection result.
[0215] Exemplarily, the wearable device is a smart watch and the preset threshold is 0.7. When the first travel state detection result indicates that the user's travel state is walking and the confidence level corresponding to walking is 0.6, since the confidence level of walking 0.6 is less than the preset threshold 0.7, the image sensor located at the watch end is activated, and multiple consecutive images of the user's current location environment collected by the image sensor at a third preset frequency are obtained as the image sequence, and the travel state is detected using the image sequence.
[0216] In one implementation, when the accuracy of the first travel state detection result is not lower than the preset threshold, the accuracy of the first travel state detection result is relatively high, and it can be determined that the target travel state detection result of the user is the travel state indicated by the first travel state detection result.
[0217] In one implementation, when the accuracy of the first travel state detection result is not lower than the preset threshold, the wearable device can also obtain the image sequence collected by the image sensor, and use the image sequence to determine the probability that the user is taking a vehicle, and further verify the first travel state detection result.
[0218] In this embodiment, when the confidence level corresponding to the running travel state or the walking travel state of the user is the first travel state detection result, the wearable device may start the image sensor when the confidence level corresponding to the running travel state or the walking travel state of the user is lower than a preset threshold, so that the image sensor collects an image sequence of the environment where the user is located. After that, an image sequence of the environment where the user is located is obtained, and the second travel state detection is performed using the image sequence to further determine whether the user travels by means of transportation or walks or runs as detected by the first travel state detection result, so as to verify the first travel state detection result.
[0219] The following Figure 5 The flowchart of an example of travel state detection provided below is used to detail the travel state detection process provided in the embodiments of the present application. Among them, the wearable device includes an IMU state recognition model and a video detection module. As shown in the following figure, the IMU state recognition model executes the following steps S501 - S504, and the video detection module executes steps S505 - S512.
[0220] S501: Input of raw data.
[0221] The IMU state recognition model can obtain the acceleration signal (ACC signal) and PPG signal collected by the IMU, and store the acceleration signal and PPG signal in the cache.
[0222] S502: Filtering of raw data.
[0223] The IMU state recognition model filters the PPG signal using a Butterworth filter, and the filtering bandwidth is selected as 0.5Hz - 4Hz.
[0224] S503: Feature extraction.
[0225] Feature extraction is performed on the acceleration signal and PPG signal in the cache to obtain various time-domain features such as the zero-crossing rate, energy, and correlation of the acceleration signal and PPG signal.
[0226] S504: Whether the confidence level is less than the threshold; if so, execute step S505; otherwise, execute step S501.
[0227] S505: Input of video data stream.
[0228] The extracted time-domain features are input into a pre-trained XGBoost model to obtain the motion state (the first travel state result) and confidence level predicted and output by the XGBoost model for the time-domain features, and it is determined whether the confidence level is less than the threshold 0.7.
[0229] If the confidence level is less than 0.7, the image sensor is activated so that the image sensor acquires the video data stream of the user's current environment. The video detection module obtains the video data stream of the user's current environment acquired by the image sensor and obtains an image sequence (112*112*10).
[0230] If the confidence level is not less than 0.7, the user is currently walking / running. Continue to obtain the original data input and detect the user's travel status at the next detection time.
[0231] S506: Denoising of a single-frame image.
[0232] The video detection module uses a median filter or a bilateral filter to eliminate the noise points in the single-frame image.
[0233] S507: Filtering of multiple-frame images.
[0234] The video detection module uses the optical flow method to estimate the motion paths of each pixel point between consecutive frames, then aligns the sequence images according to the estimated motion, and performs weighted average filtering on the aligned multiple-frame images.
[0235] S508: 3D-CNN spatio-temporal feature extraction.
[0236] The video detection module normalizes the image sequence and sends the normalized image sequence into a 3D-CNN model (feature extraction model). A filter bank of 3×3×3 slides on the multi-dimensional data to extract the spatial features of the image sequence representing the environment inside the vehicle, such as sitting people, armrests, and windows, etc., and the temporal features of the image sequence representing the characteristics of the user's motion state changing over time, such as the motion feature of the vehicle continuously jittering.
[0237] S509: MLP classification.
[0238] The video detection module activates, performs 3D pooling, and dimensionality reduction and compression on the spatial features and temporal features of the image sequence to obtain a linear feature vector of length 4096, and sends the set of linear feature vectors into the MLP model for classification to obtain a classification result (the second travel status result) representing the probability that the user is taking a vehicle.
[0239] S510: Whether the probability is less than the threshold; if so, execute step S511; if not, execute step S512.
[0240] S511: The user is walking / running.
[0241] S512: The user is not walking / running.
[0242] If the probability that the user takes a means of transportation is less than the threshold of 0.85, then the user does not take a means of transportation and walks / runs to travel; if the probability that the user takes a means of transportation is not less than the threshold of 0.85, then the user takes a means of transportation and does not walk / run to travel.
[0243] Among them, the true travel state of the user obtained by detecting the travel state using IMU data and image sequences can be represented in Table 1 below:
[0244] Table 1
[0245]
[0246] Corresponding to the detection method of the travel state, the present application embodiment also provides a detection device for the travel state.
[0247] As Figure 6 shown, a detection device for travel state is applied to a wearable device. The wearable device includes a motion sensor and an image sensor, and the motion sensor includes an inertial detection unit IMU; the device includes:
[0248] An image sequence acquisition module 601, configured to acquire an image sequence collected by the image sensor in response to the accuracy of the first travel state detection result being lower than a preset threshold, where the first travel state detection result is obtained by detecting the travel state using the motion data collected by the IMU;
[0249] A feature extraction module 602, configured to input the image sequence into a pre-trained feature extraction model, and extract spatial features and temporal features of the image sequence, where the spatial features are used to characterize the characteristics of the user's environment, and the temporal features are used to characterize the characteristics of the user's motion state changing over time;
[0250] A first result determination module 603, configured to use a pre-trained multi-layer perception model to obtain a second travel state detection result based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perception model, where the second travel state detection result represents the probability that the user takes a means of transportation, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the travel state;
[0251] A second result determination module 604, configured to determine a target travel state detection result based on the probability that the user takes a means of transportation.
[0252] In the technical solution provided by the embodiment of the present application, the wearable device may obtain an image sequence collected by an image sensor in response to the accuracy of the first travel state detection result being lower than a preset threshold, where the first travel state detection result is obtained by performing travel state detection using the motion data collected by the IMU; input the image sequence into a pre-trained feature extraction model to extract the spatial feature and the temporal feature of the image sequence, where the spatial feature is used to characterize the characteristics of the user's environment, and the temporal feature is used to characterize the characteristics of the user's motion state changing over time; use a pre-trained multi-layer perceptron model to obtain a second travel state detection result based on the spatial feature, the temporal feature, and the corresponding relationship learned during the training process of the multi-layer perceptron model, where the second travel state detection result characterizes the probability that the user is taking a means of transportation, and the corresponding relationship is the corresponding relationship between the sample spatial feature and the sample temporal feature of the training sample and the travel state; determine the target travel state detection result based on the probability that the user is taking a means of transportation.
[0253] It can be seen that when the first travel state detection result obtained by using the motion data collected by the IMU is inaccurate, the image sequence of the user's environment collected by the image sensor of the wearable device can be used for further travel state detection to determine the probability that the user is taking a means of transportation. In this way, by performing two travel state detections using the acceleration signal and the image sequence of the environment respectively, it is possible to more accurately determine whether the user is traveling by means of transportation or walking / running, avoiding misidentifying the travel state of the user taking a means of transportation as the user walking / running, and improving the accuracy of the travel state detection result. Moreover, since the spatial feature and the temporal feature of the image sequence are extracted respectively when performing travel state detection using the image sequence, the characteristics of the user's environment and the characteristics of the user's motion state changing over time are captured simultaneously, further improving the accuracy of the travel state detection result, and thus further improving the accuracy of the actual travel state detection result of the user.
[0254] As an implementation manner of the embodiment of the present application, the feature extraction module 602 includes:
[0255] A first input sub-module, configured to input the image sequence into a pre-trained feature extraction model;
[0256] A feature extraction sub-module, configured to slide a three-dimensional convolution kernel in the spatial dimension of the image sequence using the feature extraction model, extract the spatial feature of the image sequence based on the environmental content included in the image sequence, and slide the three-dimensional convolution kernel in the time dimension, and extract the temporal feature of the image sequence based on the law of change of the content included in the image sequence over time.
[0257] As an implementation manner of an embodiment of the present application, the multi-layer perceptron model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer;
[0258] The first result determination module includes:
[0259] A second input sub-module, configured to input the linear feature vector into the input layer, where the linear feature vector is obtained by compressing the spatial feature and the temporal feature, or output by the feature extraction model;
[0260] A third input sub-module, configured to receive the linear feature vector by the input layer and input the linear feature vector into the hidden layer;
[0261] A fourth input sub-module, configured to receive the linear feature vector by the hidden layer, activate the non-linear relationship in the linear feature vector by using a plurality of neurons in the at least one fully connected layer and an activation function, and input the activated linear feature vector into the output layer;
[0262] A mapping sub-module, configured to receive the activated linear feature vector by the output layer, perform a travel state mapping on the activated linear feature vector by using a normalization function, and obtain a second travel state detection result.
[0263] As an implementation manner of an embodiment of the present application, the device further includes a training module, and the training module includes:
[0264] A calibration label acquisition sub-module, configured to acquire a linear feature vector corresponding to each training sample and a calibration label of each training sample, where the plurality of training samples include an image sequence collected by an image sensor of the wearable device when the user takes a vehicle, and an image sequence collected by the image sensor when the user does not take a vehicle and is exercising, and the calibration label of each training sample represents the travel state corresponding to the training sample;
[0265] A fifth input sub-module, configured to input the linear feature vector of each training sample into the initial multi-layer perceptron model;
[0266] A prediction label acquisition sub-module, configured to acquire a prediction label output by the initial multi-layer perceptron model based on the current model parameters for processing the linear feature vector of the training sample;
[0267] A parameter adjustment sub-module, configured to adjust the model parameters of the initial multi-layer perceptron model based on the difference between the prediction label and the corresponding calibration label until the initial multi-layer perceptron model converges, and obtain a trained multi-layer perceptron model.
[0268] As an implementation manner of an embodiment of the present application, the device further includes:
[0269] An activation module, configured to, after the step of extracting the temporal features of the image sequence, utilize an activation function to activate the non-linear relationship between the spatial features and the temporal features;
[0270] A pooling module, configured to perform three-dimensional pooling on the activated spatial features and temporal features, and reduce the resolution of the feature map formed by the activated spatial features and temporal features to a preset resolution;
[0271] A dimensionality reduction module, configured to reduce the dimension of the feature map with the preset resolution to one dimension based on a preset dimensionality reduction function, so as to obtain the linear feature vector.
[0272] As an implementation manner of an embodiment of the present application, the device further includes:
[0273] A noise reduction module, configured to perform noise reduction on each frame of the image included in the image sequence before the step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, so as to obtain a noise-reduced image sequence;
[0274] A processing module, configured to process the noise-reduced image sequence by using an optical flow method to obtain the motion relationship between each frame of the image in the noise-reduced image sequence;
[0275] A filtering module, configured to perform alignment processing on each frame of the image in the noise-reduced image sequence based on the motion relationship, and perform weighted average filtering on the aligned image sequence to obtain a filtered image sequence;
[0276] A normalization module, configured to perform normalization processing on the filtered image sequence to obtain an image sequence for inputting into a pre-trained feature extraction model.
[0277] As an implementation manner of an embodiment of the present application, the second result determination module includes:
[0278] A first determination sub-module, configured to determine that the target travel state detection result is traveling by vehicle when the probability that the user travels by vehicle is greater than a preset probability threshold;
[0279] A second determination sub-module, configured to determine that the target travel state detection result is the travel state indicated by the first travel state detection result when the probability that the user travels by vehicle is not greater than the preset probability threshold.
[0280] As an implementation manner of an embodiment of the present application, the motion sensor further includes a photoelectric sensor; the device further includes:
[0281] A motion signal acquisition module, configured to acquire a motion signal collected by the motion sensor before the step of acquiring an image sequence collected by the image sensor when the accuracy of the response to the first travel state detection result is lower than a preset threshold, where the motion signal includes an acceleration signal collected by the IMU, i.e., motion data collected by the IMU, or an acceleration signal collected by the IMU and a heart rate signal collected by the optoelectronic sensor;
[0282] A time-domain feature extraction module, configured to extract the time-domain features of the motion signal and perform travel state detection based on the time-domain features to obtain the first travel state detection result.
[0283] As an implementation manner of an embodiment of the present application, the first travel state detection result is the confidence level corresponding to the user's running travel state or the confidence level corresponding to walking;
[0284] The image sequence acquisition module includes:
[0285] An image sequence acquisition sub-module, configured to start the image sensor when the confidence level is lower than the preset threshold, so that the image sensor acquires an image sequence of the user's environment.
[0286] An embodiment of the present application further provides a wearable device, as Figure 7 shown, including:
[0287] An image sensor 701, configured to acquire an image sequence;
[0288] A motion sensor 702, where the motion sensor 702 includes an inertial detection unit IMU; the IMU is configured to collect motion data;
[0289] A memory 703, configured to store a computer program;
[0290] A processor 704, configured to implement the travel state detection method provided by the embodiment of the present application when executing the program stored in the memory 703.
[0291] And the above electronic device may further include a communication bus and / or a communication interface, and the processor 704, the communication interface, and the memory 703 complete communication with each other through the communication bus.
[0292] In the technical solution provided by the embodiment of the present application, the wearable device can obtain an image sequence collected by an image sensor in response to the accuracy of the first travel state detection result being lower than a preset threshold, where the first travel state detection result is obtained by performing travel state detection on the motion data collected by the IMU; input the image sequence into a pre-trained feature extraction model to extract the spatial feature and temporal feature of the image sequence, where the spatial feature is used to characterize the characteristics of the user's environment, and the temporal feature is used to characterize the characteristics of the user's motion state changing over time; use a pre-trained multi-layer perceptron model to obtain a second travel state detection result based on the spatial feature, temporal feature, and the corresponding relationship learned during the training process of the multi-layer perceptron model, where the second travel state detection result represents the probability that the user is taking a means of transportation, and the corresponding relationship is the corresponding relationship between the sample spatial feature and sample temporal feature of the training sample and the travel state; determine the target travel state detection result based on the probability that the user is taking a means of transportation.
[0293] It can be seen that when the first travel state detection result obtained from the motion data collected by the IMU is inaccurate, the image sequence of the user's environment collected by the image sensor of the wearable device can be used for further travel state detection to determine the probability that the user is taking a means of transportation. In this way, by performing two travel state detections using the acceleration signal and the image sequence of the environment respectively, it is possible to more accurately determine whether the user is traveling by means of transportation or walking / running, avoiding misidentifying the travel state of the user taking a means of transportation as the user walking / running, and improving the accuracy of the travel state detection result. Moreover, since when performing travel state detection using the image sequence, the spatial feature and temporal feature of the image sequence are extracted respectively, capturing both the characteristics of the user's environment and the characteristics of the user's motion state changing over time, the accuracy of the travel state detection result is further improved, and thus the accuracy of the actual travel state detection result of the user is further improved.
[0294] The communication bus mentioned in the above wearable device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0295] The communication interface is used for communication between the above wearable device and other devices.
[0296] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0297] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0298] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the detection method for any of the above travel states are implemented.
[0299] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, which when running on a computer, causes the computer to execute the detection method for any of the above travel states in the above embodiments.
[0300] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server, data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a solid-state disk (SSD), etc.
[0301] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0302] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the embodiments of the device, wearable device, computer-readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments.
[0303] The foregoing are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A method for detecting travel status, characterized in that, Applied to a wearable device, the wearable device includes a motion sensor and an image sensor, and the motion sensor includes an inertial measurement unit (IMU); the method includes: In response to the accuracy of the first travel state detection result being lower than a preset threshold, obtaining an image sequence collected by the image sensor, where the first travel state detection result is obtained by performing travel state detection using the motion data collected by the IMU; Inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, where the spatial features are used to characterize the characteristics of the user's environment, and the temporal features are used to characterize the characteristics of the user's motion state changing over time; Using a pre-trained multi-layer perceptron model, based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model, obtaining a second travel state detection result, where the second travel state detection result represents the probability that the user is taking a vehicle, and the corresponding relationship is the corresponding relationship between the sample spatial features and sample temporal features of the training samples and the travel state; Based on the probability that the user is taking a vehicle, determining a target travel state detection result.
2. The method according to claim 1, characterized in that, The step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence includes: Inputting the image sequence into a pre-trained feature extraction model; Using a three-dimensional convolution kernel in the feature extraction model to slide in the spatial dimension of the image sequence, and based on the environmental content included in the image sequence, extracting the spatial features of the image sequence, and using the three-dimensional convolution kernel to slide in the time dimension, and based on the law of change of the content included in the image sequence over time, extracting the temporal features of the image sequence.
3. The method according to claim 1, wherein The multi-layer perceptron model includes an input layer, a hidden layer, and an output layer, and the hidden layer includes at least one fully connected layer; The step of using a pre-trained multi-layer perceptron model, based on the spatial features, the temporal features, and the corresponding relationship learned during the training process of the multi-layer perceptron model, obtaining a second travel state detection result includes: Inputting a linear feature vector into the input layer, where the linear feature vector is obtained by compressing the spatial features and the temporal features, or output by the feature extraction model; The input layer receives the linear feature vector and inputs the linear feature vector into the hidden layer; The hidden layer receives the linear feature vector, uses multiple neurons in the at least one fully connected layer and an activation function to activate the non-linear relationship in the linear feature vector, and inputs the activated linear feature vector into the output layer; The output layer receives the activated linear feature vector, and uses a normalization function to perform travel state mapping on the activated linear feature vector to obtain a second travel state detection result.
4. The method according to claim 3, characterized in that, The training method of the multi-layer perceptron model includes: Obtain the linear feature vector corresponding to each training sample and the calibration label of each training sample, where the multiple training samples include an image sequence collected by the image sensor of the wearable device when the user is taking a transportation means, and an image sequence collected by the image sensor when the user is not taking a transportation means and is exercising. The calibration label of each training sample represents the travel state corresponding to this training sample; Input the linear feature vector of each training sample into the initial multi-layer perceptron model; Obtain the predicted label output by the initial multi-layer perceptron model based on the current model parameters for processing the linear feature vector of the training sample; Adjust the model parameters of the initial multi-layer perceptron model based on the difference between the predicted label and the corresponding calibration label until the initial multi-layer perceptron model converges to obtain the trained multi-layer perceptron model.
5. The method according to claim 3, characterized in that, After the step of extracting the temporal features of the image sequence, the method further includes: Use an activation function to activate the non-linear relationship between the spatial features and the temporal features; Perform three-dimensional pooling on the activated spatial features and temporal features to reduce the resolution of the feature map composed of the activated spatial features and temporal features to a preset resolution; Based on a preset dimensionality reduction function, reduce the dimension of the feature map with the preset resolution to one dimension to obtain the linear feature vector.
6. The method according to any one of claims 1-5, characterized in that, Before the step of inputting the image sequence into a pre-trained feature extraction model to extract the spatial features and temporal features of the image sequence, the method further includes: Denoise each frame of the image included in the image sequence to obtain a denoised image sequence; Use the optical flow method to process the denoised image sequence to obtain the motion relationship between each frame of the image in the denoised image sequence; Perform alignment processing on each frame of the image in the denoised image sequence based on the motion relationship, and perform weighted average filtering on the aligned image sequence to obtain a filtered image sequence; Perform normalization processing on the filtered image sequence to obtain an image sequence for inputting into the pre-trained feature extraction model.
7. The method according to any one of claims 1-5, characterized in that, The step of determining the target travel state detection result based on the probability that the user is taking a transportation means includes: When the probability that the user is taking a transportation means is greater than a preset probability threshold, determine that the target travel state detection result is taking a transportation means; When the probability that the user is taking a transportation means is not greater than the preset probability threshold, determine that the target travel state detection result is the travel state indicated by the first travel state detection result.
8. The method according to any one of claims 1-5, characterized in that, The motion sensor further includes a photoelectric sensor; before the step of obtaining the image sequence collected by the image sensor in response to the accuracy of the first travel state detection result being lower than a preset threshold, the method further includes: Obtain the motion signal collected by the motion sensor, where the motion signal includes the acceleration signal collected by the IMU, that is, the motion data collected by the IMU, or the acceleration signal collected by the IMU and the heart rate signal collected by the photoelectric sensor; Extract the time-domain features of the motion signal, and perform travel state detection based on the time-domain features to obtain the first travel state detection result.
9. The method according to any one of claims 1-5, characterized in that The first travel state detection result is the confidence level corresponding to the user's travel state being running or walking. The step of obtaining the image sequence collected by the image sensor in response to the accuracy of the first travel state detection result being lower than a preset threshold includes: When the confidence level is lower than the preset threshold, activate the image sensor so that the image sensor collects an image sequence of the user's environment.
10. A wearable device, characterized in that, Includes: An image sensor for collecting an image sequence; A motion sensor, the motion sensor includes an inertial detection unit IMU, and the IMU is used to collect motion data; A memory for storing computer programs; A processor, when executing the program stored on the memory, implements the method according to any one of claims 1-9.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1-9.
Citation Information
Cited By
Tooth brushing action detection method, device, equipment and medium
CN121545213A