Action prediction model training method based on smart watch data
By collecting multiple types of data through smartwatches, performing multimodal feature fusion, and conducting scientific training and evaluation, the problem of single data and imperfect evaluation in existing motion prediction models is solved, resulting in more efficient motion prediction performance.
Patent Information
- Application Number
- CN202511103937.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-28
AI Technical Summary
Existing methods for training motion prediction models suffer from limitations such as a single data acquisition method, difficulty in covering a wide variety of motion scenarios, lack of multimodal feature fusion mechanisms, and inadequate model evaluation, resulting in poor model generalization ability and low prediction accuracy.
By acquiring multiple types of data collected by smartwatches as positive and negative samples, multimodal feature extraction and fusion are performed, label classification vectors are configured, training and evaluation sets are constructed, and model parameters are optimized using loss and evaluation functions to ensure that the model approaches the optimal solution during training.
This improves the model's generalization ability and prediction accuracy, ensures continuous optimization during training, avoids overfitting or underfitting, and enhances the model's reliability and stability.
Smart Images

Figure CN121034535A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a motion prediction model training method based on smart watch data. BACKGROUND
[0002] At present, smart wearables are becoming increasingly popular, and smart watches have become common devices in people's lives due to their convenience and multifunctionality. Among them, motion recognition and prediction using data collected by smart watches have broad application prospects. For example, in the field of health monitoring, the physical activity status of a user can be evaluated by recognizing the user's daily motions; in the field of sports and fitness, whether the motion is standardized can be accurately judged to provide scientific guidance for users.
[0003] However, the existing motion prediction model training methods have many shortcomings. On the one hand, the data acquisition method is relatively single, and is mostly focused on limited motion data collection in specific scenarios, which is difficult to cover a wide variety of actual motion situations, resulting in poor model generalization ability and inability to accurately predict various complex motions. On the other hand, in the data processing and model training process, there is a lack of effective multi-modal feature fusion mechanism, which cannot fully exploit the internal relationship between the multi-type data collected by the smart watch, so that the extraction of motion features by the model is not comprehensive and accurate, thereby affecting the accuracy of motion prediction. In addition, the existing training methods are not perfect in the model evaluation link, lacking scientific and reasonable evaluation indicators and processes, making it difficult to accurately determine whether the model meets the requirements of practical application. Therefore, it is of great practical significance to develop an efficient and accurate motion prediction model training method based on smart watch data. SUMMARY
[0004] The purpose of the present application is to solve the problems existing in the prior art by providing a motion prediction model training method based on smart watch data.
[0005] To achieve the above-mentioned purpose, in a first aspect, the present application provides a motion prediction model training method based on smart watch data, which comprises:
[0006] obtaining positive sample data and corresponding negative sample data; the positive sample data is seven types of original collected data sequences collected when a prescribed motion is completed; the negative sample data is seven types of original collected data sequences collected when a non-prescribed motion is completed;
[0007] processing the positive sample data and the negative sample data respectively to obtain positive sample multi-modal feature maps and negative sample multi-modal feature maps;
[0008] configuring a label classification vector for each positive sample multi-modal feature map and a label classification vector for each negative sample multi-modal feature map;
[0009] The positive sample multi-modal feature map and the negative sample multi-modal feature map are taken as a first training feature map, the label classification vector is taken as a corresponding first label classification vector, the first training feature map and the first label classification vector form a corresponding first data record, and all the first data records form a first data set;
[0010] The first data set is processed into a first training set and a first evaluation set, the first training feature map of each first data record of the first training set is input into the action prediction model for processing, a prediction classification vector output by the model at this time is taken as a corresponding first prediction classification vector, and each first prediction classification vector and the corresponding first label classification vector form a first prediction-label pair;
[0011] The first prediction-label pair is substituted into a first model loss function to obtain a first loss value;
[0012] It is identified whether the first loss value meets a preset first loss value range, if yes, the first training feature map of each first data record of the first evaluation set is input into the action prediction model for processing, a prediction classification vector output by the model at this time is taken as a corresponding second prediction classification vector, and each second prediction classification vector and the corresponding first label classification vector form a corresponding second prediction-label pair;
[0013] All the second prediction-label pairs are substituted into a first model evaluation function to obtain a first evaluation value;
[0014] It is identified whether the first evaluation value meets a preset first evaluation value range, if yes, it is confirmed that the model training is completed.
[0015] In a possible implementation, the obtaining of the positive sample data and the corresponding negative sample data specifically includes:
[0016] Seven types of original collected data sequences of 10 types of specified actions of a user are collected for multiple times by a smart watch, and a group of seven types of original collected data sequences collected each time form a corresponding positive sample data;
[0017] Seven types of original collected data sequences of any non-specified action of the user are collected for multiple times by the smart watch, and a group of seven types of original collected data sequences collected each time form a corresponding negative sample data.
[0018] In a possible implementation, the smart watch comprises an inertial measurement unit, a GPS / Beidou positioning unit, a UWB / Bluetooth positioning unit, a heart rate monitoring unit, a body temperature monitoring unit, an ambient humidity sensor, and a recording unit, and the seven types of original collected data sequences comprise:
[0019] six-dimensional attitude data sequences collected by the inertial measurement unit, first three-dimensional positioning coordinate data sequences collected by the GPS / Beidou positioning unit, second three-dimensional positioning coordinate data sequences collected by the UWB / Bluetooth positioning unit, heart rate data sequences output by the heart rate monitoring unit, body temperature data sequences output by the body temperature monitoring unit, humidity sequence data output by the ambient humidity sensor, and sound wave sampling data sequences output by the recording unit; the six-dimensional attitude data sequences comprise X / Y / Z axis acceleration and rotation angle; the first three-dimensional positioning coordinate data sequences comprise longitude, latitude, and altitude; and the second three-dimensional positioning coordinate data sequences comprise three-axis coordinates in a current indoor pre-labeled relative coordinate system.
[0020] In a possible implementation, the specified actions comprise horizontal surface scrubbing, vertical surface scrubbing, O-ring surface scrubbing, faucet scrubbing, toilet scrubbing, dry floor mopping, wet floor mopping, spraying, garbage dumping, and garbage bag replacement.
[0021] In a possible implementation, the processing of the positive sample data and the negative sample data respectively to obtain positive sample multi-modal feature maps and negative sample multi-modal feature maps comprises:
[0022] obtaining seven types of original collected data sequences of the same time period T;
[0023] based on a unified sampling frequency f sam sampling points are set for the time period to obtain N sampling points; the sampling frequency f sam is an integer multiple of the seven types of original collected data;
[0024] the length of the time period T is taken as the sequence length, and the sampling frequency f sam is used to resample the seven types of collected data sequences to obtain seven types of aligned data sequences;
[0025] the seven types of aligned data sequences are respectively filtered and smoothed to obtain seven types of smoothed data sequences;
[0026] multi-modal data sequences are created based on the seven types of smoothed data sequences; each multi-modal data in the multi-modal data sequences comprises 16-dimensional data: X / Y / Z axis acceleration and rotation angle, longitude, latitude, altitude, X' / Y' / Z' three-axis coordinates, heart rate, body temperature, ambient humidity, and sound wave amplitude;
[0027] Normalizing heart rate, body temperature, ambient humidity, sound wave amplitude in the multi-modal data sequence;
[0028] The multi-modal data sequence is regarded as a multi-modal feature map P with a shape of HxW.
[0029] In a possible implementation, the processing of the first data set into the first training set and the first evaluation set, the input of the first training feature map of each first data record of the first training set into the action prediction model for processing, and the prediction classification vector output by the model at this time as a corresponding first prediction classification vector specifically include:
[0030] The multi-modal feature map is input into a feature extraction network for local feature extraction to obtain a corresponding first feature tensor;
[0031] The first feature vectors of the first feature tensor are sorted in ascending order according to the index j to form a first feature vector sequence, which is input into an LSTM model for time series feature extraction to obtain a corresponding second feature vector sequence, and the second feature vector sequence is converted into a second feature tensor;
[0032] The second feature tensor is input into an MLP model for feature vector mapping to obtain a third feature vector;
[0033] The third feature vector is input into a linear layer for full connection calculation to obtain a fourth feature vector;
[0034] The fourth feature vector is input into a Sigmoid function for calculation to obtain a corresponding prediction classification vector.
[0035] In a possible implementation, the action prediction model includes a feature extraction network, an LSTM model, an MLP model, a linear layer, and a Sigmoid function layer.
[0036] In a possible implementation, the prediction classification vector includes 10 action confidence levels; the 10 action confidence levels correspond to 10 action types one by one; each action execution degree is the confidence score of the corresponding action; and the maximum value in the 10 action confidence levels is the most likely predicted action type.
[0037] In a possible implementation, the processing of the first data set into the first training set and the first evaluation set specifically includes:
[0038] According to a preset first segmentation ratio, the first data set is segmented into two sub-data sets, denoted as a first training set and a first evaluation set; the proportions of positive sample data and negative sample data in the first training set and the first evaluation set are consistent.
[0039] In a second aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the action prediction model training method based on smart watch data according to any one of the first aspect.
[0040] The action prediction model training method based on smart watch data provided by the present application has the following technical effects:
[0041] The data is comprehensive, seven types of original collected data sequences when completing the specified action and the non-specified action are obtained as positive and negative sample data, the types and range of the training data are enriched, the model can learn more diversified action features, the generalization ability of the model is effectively improved, and the action prediction demand in different scenes can be adapted.
[0042] The feature fusion is accurate, the multi-modal feature maps are obtained by processing the positive and negative sample data respectively, and the label classification vector is configured to form the first data set, the multi-class data collected by the smart watch is fully utilized, the action features are more comprehensively and accurately extracted through the multi-modal feature fusion, high-quality feature input is provided for the model training, and the prediction accuracy of the model is improved.
[0043] The training process is scientific, the first data set is divided into a training set and an evaluation set, in the training process, the first prediction-label pair of the prediction classification vector output by the model and the label classification vector is calculated, and the loss value is obtained by substituting the loss function, the model parameters are adjusted according to whether the loss value meets the preset range, and the model is gradually optimized. This scientific training method can ensure that the model continuously approaches the optimal solution in the training process, and improves the performance of the model.
[0044] The evaluation mechanism is perfect, when the training loss value meets the requirement, the model is evaluated by using the evaluation set, the second prediction-label pair is calculated, the evaluation value is obtained by substituting the evaluation function, and whether the model is trained is judged according to whether the evaluation value meets the preset range. The perfect evaluation mechanism can accurately judge whether the performance of the model meets the actual application standard, avoid overfitting or underfitting of the model, and ensure the reliability and stability of the model. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 The action prediction model training method based on smart watch data provided by the present application is provided with a flowchart;
[0046] Figure 2 The action prediction model training method based on smart watch data provided by the present application is provided with a flowchart; Figure 1 The flowchart of step 120;
[0047] Figure 3 The action prediction model is provided with a schematic diagram;
[0048] Figure 4 The action prediction model is provided with a schematic diagram;Figure 1 Flowchart of step 150;
[0049] Figure 5 The computer readable storage medium provided for the embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0051] The technical solutions of the present application will be further described in detail below with reference to the drawings and embodiments.
[0052] Embodiment one
[0053] Figure 1 The flowchart of the action prediction model training method based on smart watch data provided for the embodiment of the present application is schematically shown. The method can be applied in multiple fields, such as fitness field and cleaning field. The present application will be described by applying the method in the cleaning field. When applied in the cleaning field, the method is realized by relying on the smart watch carried by the cleaning personnel. As shown in the figure, the method comprises the following steps: Figure 1
[0054] Step 110, obtaining positive sample data and corresponding negative sample data; the positive sample data is seven types of original collected data sequences collected when completing a specified action; the negative sample data is seven types of original collected data sequences collected when completing a non-specified action;
[0055] Specifically, the smart watch is carried on the arm of the cleaning personnel and is used for data collection. The present application collects seven types of original collected data sequences of 10 types of specified actions of the user by the smart watch for multiple times, and a corresponding positive sample data is composed of a group of seven types of original collected data sequences collected each time; seven types of original collected data sequences of any non-specified action of the user are collected by the smart watch for multiple times, and a corresponding negative sample data is composed of a group of seven types of original collected data sequences collected each time.
[0056] The ten types of specified actions are a series of cleaning classification actions that must be completed by the cleaning personnel when completing a task, including horizontal surface scrubbing, vertical surface scrubbing, O-ring surface scrubbing, faucet scrubbing, toilet scrubbing, dry ground mopping, wet ground mopping, spraying, dumping garbage and replacing garbage bag.
[0057] The intelligent watch comprises an inertial measurement unit, a GPS / Beidou positioning unit, a UWB / Bluetooth positioning unit, a heart rate monitoring unit, a body temperature monitoring unit, an environmental humidity sensor and a recording unit, and seven types of original collected data sequences comprise:
[0058] The six-dimensional attitude data sequence collected by the inertial measurement unit, the first three-dimensional positioning coordinate data sequence collected by the GPS / Beidou positioning unit, the second three-dimensional positioning coordinate data sequence collected by the UWB / Bluetooth positioning unit, the heart rate data sequence output by the heart rate monitoring unit, the heart rate data sequence output by the body temperature monitoring unit, the humidity sequence data output by the environmental humidity sensor and the sound wave sampling data sequence output by the recording unit; the six-dimensional attitude data sequence comprises X / Y / Z axis acceleration and rotation angle; the first three-dimensional positioning coordinate data sequence comprises longitude, latitude and altitude; and the second three-dimensional positioning coordinate data sequence comprises three-axis coordinates in a current indoor pre-marked relative coordinate system.
[0059] The inertial measurement unit is used for collecting the six-dimensional attitude data sequence The coordinate system thereof is a northeast celestial coordinate system XYZ; the six-dimensional attitude data comprises X / Y / Z axis acceleration and rotation angle, and the rotation angle comprises a pitch angle around the X axis, a roll angle around the Y axis and a yaw angle around the Z axis.
[0060] The GPS / Beidou positioning unit is used for collecting the three-dimensional positioning coordinate data sequence The three-dimensional positioning coordinate data comprise longitude, latitude and altitude.
[0061] The UWB / Bluetooth positioning unit is used for collecting the three-dimensional positioning coordinate data sequence The coordinate system thereof is a current indoor pre-marked relative coordinate system X'Y'Z'; the origin of the relative coordinate system X'Y'Z' coincides with a specified point in the northeast celestial coordinate system XYZ; and the coordinate conversion relationship between the relative coordinate system X'Y'Z' and the northeast celestial coordinate system XYZ is fixed. The three-dimensional positioning coordinate data comprise X ’ / Y ’ / Z ’ three-axis coordinates.
[0062] The heart rate monitoring unit is used for outputting the heart rate data sequence The heart rate data is a heart rate, and the unit is beats per minute.
[0063] The body temperature monitoring unit is used for outputting the body temperature data sequence The body temperature data is a body temperature, and the unit is degrees Celsius.
[0064] The environmental humidity sensor is used for outputting the humidity data sequence humidity data is an environmental humidity.
[0065] The audio recording unit is used to record the environment and output the corresponding sound sample data sequence sound sample data is the sound amplitude at a sampling point.
[0066] The correlation between the cleaning classification action and the collected data sequence is shown in Table 1, where Y represents correlation and N represents no correlation.
[0067] Table 1
[0068] ID Action Type [00000D 1 ]] [00000D 2 ]] [00000D 3 ]] D 4 ]]> [002D 5 ]]> D 6 ]]> [002D 7 ]]> [S1] Horizontal surface wiping Y Y Y Y Y Y N [S2] Longitudinal surface wiping Y Y Y Y Y Y N [S3] O-ring surface wiping Y Y Y Y Y Y N [S4] Faucet wiping Y Y Y Y Y Y N [S5] Toilet scrubbing Y Y Y Y Y Y N [S6] Dry floor mopping Y Y Y Y Y Y N [S7] Wet floor mopping Y Y Y Y Y Y N [S8] Spraying Y Y Y Y Y Y Y [S9] Dumping trash Y Y Y Y Y Y Y [SA 10 ]]> Changing trash bag Y Y Y Y Y Y Y
[0069] Step 120, processing the positive sample data and the negative sample data respectively to obtain the positive sample multi-modal feature map and the negative sample multi-modal feature map;
[0070] Wherein, when collecting positive and negative sample data, the total amount of the two types of data is required to be balanced, and the ratio of 1:1 is as far as possible.
[0071] Specifically, as shown in Figure 2 , step 120 includes the following steps:
[0072] Step 1201, obtaining seven types of original collected data sequences D 1-7 of the same period T
[0073] Step 1202, based on the uniform sampling frequency f sam , the sampling points of the period are set to obtain N sampling points; wherein, the sampling frequency f sam can be evenly divided by the seven types of original collected f 1-7 ;
[0074] Specifically, in order to meet the requirements of subsequent data processing and analysis, and considering the sampling ability of different sensor devices, the sampling frequency f sam can ensure data quality and will not cause excessive data redundancy. The reason why the sampling frequency f sam can be evenly divided by the seven types of original collected frequency f 1-7 is to facilitate subsequent data alignment and fusion processing, and to avoid data misalignment caused by mismatching of sampling frequency.
[0075] Step 1203, based on the linear interpolation method, the sampling frequency f sam is used to resample the seven types of collected data sequences D 1-7 with the length of the period T as the sequence length, so as to achieve time alignment, and obtain seven types of aligned data sequences
[0076] Step 1204, filtering and smoothing the 7 types of alignment data sequences respectively to obtain 7 types of smoothed data sequences
[0077] Specifically, for data such as acceleration and rotation angle that may contain high-frequency noise, a low-pass filter can be selected to remove high-frequency noise; for relatively stable data such as heart rate and body temperature, a moving average method can be used for smoothing.
[0078] Step 1205, based on the 7 types of smoothed data sequences create a multimodal data sequence Each multimodal data in the multimodal data sequence includes 16-dimensional data: X / Y / Z-axis acceleration and rotation angle, longitude, latitude, altitude, X’ / Y’ / Z’ three-axis coordinates, heart rate, body temperature, ambient humidity, sound wave amplitude; 1≤ sampling point index j≤N.
[0079] Step 1206, normalizing the heart rate, body temperature, ambient humidity, and sound wave amplitude in the multimodal data sequence .
[0080] Specifically, the normalization method is selected: the selected normalization method and the reason are explained. For example, the minimum-maximum normalization method is used to linearly map the data to the [0, 1] interval. The reason for selecting this method is that it is simple and easy to implement, and can maintain the relative relationship of the data.
[0081] Normalization process: detailed description of the specific steps of normalizing the four types of data of heart rate, body temperature, ambient humidity, and sound wave amplitude. For example, first traverse the multimodal data sequence find the minimum and maximum values of each type of data, and then normalize the data of each sampling point according to the normalization formula.
[0082] Step 1207, regarding the multimodal data sequence as a multimodal feature map P with a shape of HxW, or a feature tensor, H=16, W=N.
[0083] Step 130, configuring a label classification vector for each positive sample multimodal feature map, and configuring a label classification vector for each negative sample multimodal feature map
[0084] Specifically, for each positive sample multimodal feature map P * a corresponding label classification vector S * is configured, and the configuration rule is:
[0085] The label classification vector S * includes 10 types of label confidence The label confidence degree matching the action type of the specified action corresponding to the current positive sample The remaining 9 unmatched label confidence degrees
[0086] For each negative sample multi-modal feature map P * Configure a corresponding label classification vector S * The configuration rule is:
[0087] Label classification vector S * Including 10 label confidence degrees And the 10 label confidence degrees All are 0.
[0088] Step 140, taking the positive sample multi-modal feature map and the negative sample multi-modal feature map as a first training feature map, taking the label classification vector as a corresponding first label classification vector, and taking the first training feature map and the first label classification vector to form a corresponding first data record, and taking all the obtained first data records to form a first data set;
[0089] Specifically, taking each positive and negative sample multi-modal feature map P * As a corresponding first training feature map, and taking each label classification vector S * As a corresponding first label classification vector; and taking each first training feature map and its corresponding first label classification vector to form a corresponding first data record; and taking all the obtained first data records to form a first data set.
[0090] The obtained first data set includes a plurality of first data records; each first data record includes a first training feature map and a first label classification vector; the first training feature map is a multi-modal feature map P * , and the first label classification vector is a label classification vector S * .
[0091] Step 150, processing the first data set into a first training set and a first evaluation set, inputting the first training feature map of each first data record of the first training set into the action prediction model for processing, and taking the prediction classification vector output by the model at this time as a corresponding first prediction classification vector, and taking each first prediction classification vector and the corresponding first label classification vector to form a first prediction-label pair;
[0092] Specifically, the first dataset is divided into two subsets, denoted as the first training set and the first evaluation set, according to a preset first segmentation ratio. The first segmentation ratio is a preset ratio parameter, such as 8:2. Both the first training set and the first evaluation set consist of multiple first data records. The ratio of the total number of records in the first training set to the first evaluation set satisfies the first segmentation ratio. The ratio of positive to negative samples in the first training set and the first evaluation set is consistent with the ratio of positive to negative samples in the first dataset.
[0093] Step 150 involves processing the first dataset into a first training set and a first evaluation set. The first training feature maps of each first data record in the first training set are input into the action prediction model for processing. The predicted classification vector output by the model in this step is used as a corresponding first predicted classification vector. Specifically, this includes:
[0094] Step 1501: Input the multimodal feature map into the feature extraction network for local feature extraction to obtain the corresponding first feature tensor;
[0095] Specifically, the multimodal feature map P is input into the feature extraction network to perform local feature extraction and obtain the corresponding first feature tensor X1.
[0096] The shape of the first feature tensor X1 is H1×W1, where H1>H and W1=W;
[0097] The first characteristic tensor X1 consists of W1 vectors, each with a length of H1, representing the first characteristic vector. composition.
[0098] Step 1502: Sort the first feature vectors of the first feature tensor in ascending order of index j to form a first feature vector sequence, input it into the LSTM model for temporal feature extraction, obtain the corresponding second feature vector sequence, and convert the second feature vector sequence into a second feature tensor.
[0099] Specifically, the first eigenvector of the first feature tensor X1 The first feature vector sequence is formed by sorting the elements by index j in ascending order. This sequence is then input into the LSTM model for temporal feature extraction to obtain the corresponding second feature vector sequence. The second feature vector sequence is then converted into a second feature tensor X2 of shape H2×W2.
[0100] The second feature vector sequence consists of W2 second feature vectors of length H2. Composition; W2 = W1.
[0101] Step 1503: Input the second feature tensor into the MLP model for feature vector mapping and process to obtain the third feature vector;
[0102] Specifically, the MLP model is sequentially connected by two or more hidden layers, and each hidden layer is sequentially connected by a fully connected layer and a nonlinear activation function. The nonlinear activation function uses a ReLU activation function by default.
[0103] The second feature tensor X2 is input into the MLP model for feature vector mapping processing to obtain a third feature vector X3 with a vector length of W3. W3≠ W2.
[0104] In step 1504, the third feature vector is input into a linear layer for full connection calculation to obtain a fourth feature vector.
[0105] Specifically, the linear layer is implemented based on a full connection layer.
[0106] The third feature vector X3 is input into the linear layer for full connection calculation to obtain a fourth feature vector X4 with a vector length of W4. W4 = 10.
[0107] In step 1505, the fourth feature vector is input into a Sigmoid function for calculation to obtain a corresponding predicted classification vector.
[0108] Specifically, the Sigmoid function layer is used to input the fourth feature vector X4 into the Sigmoid function for calculation to obtain a corresponding predicted classification vector S and output.
[0109] Subsequently, the first training feature map of each first data record in the first training set is input into the action prediction model for processing, and the predicted classification vector S output by the model at this time is taken as a corresponding first predicted classification vector. Each first predicted classification vector and the corresponding first label classification vector form a corresponding first prediction-label pair.
[0110] In step 160, the first prediction-label pair is substituted into the first model loss function to obtain a first loss value.
[0111] Specifically, all first prediction-label pairs are substituted into the first model loss function to obtain a corresponding first loss value. The first model loss function is implemented based on a cross-entropy loss function.
[0112] In step 170, it is identified whether the first loss value meets a preset first loss value range. If yes, the first training feature map of each first data record in the first evaluation set is input into the action prediction model for processing, and the predicted classification vector output by the model at this time is taken as a corresponding second predicted classification vector. Each second predicted classification vector and the corresponding first label classification vector form a corresponding second prediction-label pair.
[0113] Specifically, when the first evaluation value meets the preset first evaluation value range, the first evaluation value is confirmed as the final evaluation value, and the model training is ended.
[0114] In an optional implementation, if the first evaluation value does not meet the preset first evaluation value range, the model parameters of the action prediction model are modulated based on a preset first model optimizer in a direction of minimizing the first model loss function, and the step 150 is returned at the end of the modulation.
[0115] At step 180, all the second prediction-label pairs are substituted into the first model evaluation function to obtain a first evaluation value.
[0116] The first model evaluation function is implemented based on a MAE function, a MSE function or a RMSE function.
[0117] At step 190, whether the first evaluation value meets the preset first evaluation value range is identified.
[0118] Specifically, whether the first evaluation value meets the preset first evaluation value range is identified.
[0119] The action prediction model training method based on smart watch data provided by the application has the following technical effects:
[0120] The data is comprehensive, seven types of original collected data sequences when completing the specified action and the non-specified action are obtained as positive and negative sample data, the types and range of the training data are enriched, the model can learn more diversified action features, the generalization ability of the model is effectively improved, and the action prediction demand in different scenes can be adapted.
[0121] The feature fusion is accurate, the positive and negative sample data are processed to obtain multi-modal feature maps, and the label classification vectors are configured to form a first data set, the multi-type data collected by the smart watch is fully utilized, the action features are more comprehensively and accurately extracted through the multi-modal feature fusion, high-quality feature input is provided for the model training, and the prediction accuracy of the model is improved.
[0122] The training process is scientific, the first data set is divided into a training set and an evaluation set, in the training process, a first prediction-label pair of a prediction classification vector output by the model and a label classification vector is continuously calculated, and the loss value is obtained by substituting the first prediction-label pair into the loss function, and the model parameters are adjusted according to whether the loss value meets a preset range, so that the model is gradually optimized. This scientific training method can ensure that the model continuously approaches the optimal solution in the training process, and improves the performance of the model.
[0123] The evaluation mechanism is perfect, when the training loss value meets the requirement, the model is evaluated by using the evaluation set, the second prediction-label pair is calculated, and the evaluation value is obtained by substituting the second prediction-label pair into the evaluation function, and whether the model is trained is judged according to whether the evaluation value meets a preset range. The perfect evaluation mechanism can accurately judge whether the performance of the model meets the actual application standard, avoid overfitting or underfitting of the model, and ensure the reliability and stability of the model.
[0124] Embodiment two
[0125] The embodiment two of the present application provides a computer server, which comprises a memory, a processor and a transceiver.
[0126] The processor is coupled with the memory, reads and executes instructions in the memory, to realize any one of the action prediction model training methods based on smart watch data provided in the above-mentioned embodiment one.
[0127] The transceiver is coupled with the processor, and the transceiver is controlled by the processor to perform message transmission and reception.
[0128] Embodiment three
[0129] The embodiment three of the present application provides a chip system, which comprises a processor, the processor is coupled with a memory, and the memory stores program instructions, when the program instructions stored in the memory are executed by the processor, any one of the action prediction model training methods based on smart watch data provided in the embodiment one is realized.
[0130] Embodiment four
[0131] The embodiment four of the present application provides a computer readable storage medium, as shown in the figure, comprising a program or instructions, when the program or instructions run on the computer, any one of the action prediction model training methods based on smart watch data provided in the embodiment one is realized. Figure 5
[0132] Those skilled in the art should further appreciate that the elements and algorithms described in connection with the examples disclosed herein can be embodied in electronic hardware, computer software, or in combinations of both. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein in terms of their general functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0133] The steps of a method or algorithm described in connection with the examples disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0134] The specific implementation described above is further to the purposes, technical solutions, and beneficial effects of the present application. It should be understood that the above description is merely a specific implementation of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for training a motion prediction model based on smartwatch data, characterized in that, The method includes: Acquire positive sample data and corresponding negative sample data; the positive sample data consists of seven types of raw data sequences collected when a specified action is performed; the negative sample data consists of seven types of raw data sequences collected when a non-specified action is performed. The positive sample data and the negative sample data are processed separately to obtain positive sample multimodal feature maps and negative sample multimodal feature maps; Configure a label classification vector for each positive sample multimodal feature map and configure a label classification vector for each negative sample multimodal feature map; The positive sample multimodal feature map and the negative sample multimodal feature map are used as a first training feature map, and the label classification vector is used as the corresponding first label classification vector. The first training feature map and the first label classification vector form a corresponding first data record, and all the obtained first data records form a first dataset. The first dataset is processed into a first training set and a first evaluation set. The first training feature map of each first data record in the first training set is input into the action prediction model for processing. The predicted classification vector output by the model at that time is used as a corresponding first predicted classification vector. Each first predicted classification vector and the corresponding first label classification vector form a first prediction-label pair. The first prediction-label pair is substituted into the first model loss function to calculate the first loss value; The system identifies whether the first loss value meets the preset first loss value range. If it does, the first training feature map of each first data record in the first evaluation set is input into the action prediction model for processing. The predicted classification vector output by the model at this time is used as a corresponding second predicted classification vector. Each second predicted classification vector and the corresponding first label classification vector form a corresponding second prediction-label pair. Substitute all the second prediction-label pairs into the first model evaluation function to calculate the first evaluation value; The system identifies whether the first evaluation value meets the preset range of the first evaluation value. If it does, the system confirms that the model training has ended.
2. The method according to claim 1, characterized in that, The acquisition of positive sample data and corresponding negative sample data specifically includes: The user's 10 prescribed actions are collected multiple times using a smartwatch, and each set of seven original data sequences is combined to form a corresponding positive sample. The smartwatch collects data multiple times from seven types of raw data sequences representing any non-prescribed action of the user, and each set of seven raw data sequences collected at one time forms a corresponding negative sample data.
3. The method according to claim 1, characterized in that, The smartwatch includes an inertial measurement unit, a GPS / BeiDou positioning unit, a UWB / Bluetooth positioning unit, a heart rate monitoring unit, a body temperature monitoring unit, an ambient humidity sensor, and a recording unit. The seven types of raw data sequences include: The system includes a six-dimensional attitude data sequence acquired by the inertial measurement unit, a first three-dimensional positioning coordinate data sequence acquired by the GPS / BeiDou positioning unit, a second three-dimensional positioning coordinate data sequence acquired by the UWB / Bluetooth positioning unit, a heart rate data sequence output by the heart rate monitoring unit, a heart rate data sequence output by the body temperature monitoring unit, a humidity sequence data output by the environmental humidity sensor, and a sound wave sampling data sequence output by the recording unit. The six-dimensional attitude data sequence includes acceleration and rotation angle along the X / Y / Z axes. The first three-dimensional positioning coordinate data sequence includes longitude, latitude, and altitude. The second three-dimensional positioning coordinate data sequence includes three-axis coordinates in a pre-calibrated relative coordinate system within the current indoor environment.
4. The method according to claim 1, characterized in that, The prescribed actions include horizontal surface cleaning, vertical surface cleaning, O-ring surface cleaning, faucet cleaning, toilet brush cleaning, dry mopping, wet mopping, spraying, dumping garbage, and changing garbage bags.
5. The method according to claim 1, characterized in that, The specific steps of processing the positive sample data and the negative sample data to obtain the positive sample multimodal feature map and the negative sample multimodal feature map include: Obtain seven types of raw data sequences collected within the same time period T; Based on a unified sampling frequency f sam Sampling points are set for the time period to obtain N sampling points; where the sampling frequency f sam Divisible by seven types of original data collection; Using the duration T as the sequence length, and based on linear interpolation at the sampling frequency f sam The seven types of collected data sequences were resampled to obtain seven types of aligned data sequences. Seven types of aligned data sequences were filtered and smoothed to obtain seven types of smoothed data sequences. Multimodal data sequences are created based on 7 types of smooth data sequences; each multimodal data sequence includes 16 dimensions of data: acceleration and rotation angle along the X / Y / Z axes, longitude, latitude, altitude, X' / Y' / Z' coordinates, heart rate, body temperature, ambient humidity, and sound wave amplitude; Normalize the heart rate, body temperature, ambient humidity, and sound wave amplitude in the multimodal data sequence; The multimodal data sequence is considered as a multimodal feature map P with shape H×W.
6. The method according to claim 1, characterized in that, The step of processing the first dataset into a first training set and a first evaluation set, inputting the first training feature map of each first data record in the first training set into the action prediction model for processing, and taking the predicted classification vector output by the model in this step as a corresponding first predicted classification vector specifically includes: The multimodal feature map is input into a feature extraction network for local feature extraction to obtain the corresponding first feature tensor; The first feature vectors of the first feature tensor are sorted in ascending order of index j to form the first feature vector sequence. This sequence is then input into the LSTM model for temporal feature extraction to obtain the corresponding second feature vector sequence. Finally, the second feature vector sequence is converted into the second feature tensor. The second feature tensor is input into the MLP model for feature vector mapping, and then processed to obtain the third feature vector. The third feature vector is input into a fully connected linear layer for computation to obtain the fourth feature vector; The fourth feature vector is input into the Sigmoid function for calculation to obtain the corresponding predicted classification vector.
7. The method according to claim 6, characterized in that, The action prediction model includes a feature extraction network, an LSTM model, an MLP model, a linear layer, and a Sigmoid function layer.
8. The method according to claim 1, characterized in that, The predicted classification vector includes 10 action confidence scores; each of the 10 action confidence scores corresponds one-to-one with one of the 10 action types; the execution score of each action is the confidence score of the corresponding action; the maximum value among the 10 action confidence scores is taken as the most likely predicted action type.
9. The method according to claim 1, characterized in that, The process of dividing the first dataset into a first training set and a first evaluation set specifically includes: According to the preset first segmentation ratio, the first dataset is divided into two sub-datasets, denoted as the first training set and the first evaluation set; the ratio of positive sample data to negative sample data is the same in the first training set and the first evaluation set.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is executed by a processor according to any one of claims 1-9, the motion prediction model training method based on smartwatch data.