Action recognition method and device, electronic equipment and storage medium

By using low-frequency sensors and simple models in motion recognition, and combining first and second data to perform type analysis of motion content, the problems of false triggering and high power consumption in motion recognition are solved, achieving higher accuracy and longer device standby time.

CN121959091APending Publication Date: 2026-05-01BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XIAOMI MOBILE SOFTWARE CO LTD
Filing Date
2024-10-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In existing technologies, motion recognition algorithms are prone to misidentification, which can lead to false triggering of motion control commands, affecting user experience. Furthermore, the use of high-frequency sensors increases device power consumption and shortens standby time.

Method used

By acquiring the first data and determining the action content and time period based on the preset action recognition model, and combining the second data and the preset action classification model to determine the type of action content, the action recognition is performed using a low-frequency sensor and a simple machine learning model, reducing the requirements for sensor frequency and algorithm accuracy.

Benefits of technology

It improves the accuracy of motion recognition, avoids accidental triggering of motion control commands, reduces device power consumption, increases device standby time, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121959091A_ABST
    Figure CN121959091A_ABST
Patent Text Reader

Abstract

The invention relates to an action recognition method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring first data, wherein the first data is user action data acquired by at least one sensor in a first time period; based on a preset action recognition model, action content associated with the first data and an action time period of the action content are determined, and the action time period represents a time period when the user executes the action content; based on second data and a preset action classification model, the type of the action content is determined, the second data is user action data collected by at least one sensor in a second time period, and the second time period is at least one of a first preset time period before the action time period and a second preset time period after the action time period; and determining an action recognition result based on the action content and the type of the action content. According to the method, false triggering of the action control instruction can be avoided, and equipment power consumption can be reduced while the accuracy of the action recognition result is ensured, so that the use experience of a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of terminal technology, and in particular to a motion recognition method, device, electronic device, and storage medium. Background Technology

[0002] Motion control is a type of device control. When a user performs an action, sensors on the device detect the action, and a motion recognition algorithm identifies the detected action, converting it into a corresponding motion control command to control the device. However, in related technologies, when a motion is detected, regardless of whether it is a specific action or an action intended to control the device, it is often identified as one of several pre-stored specific actions. This can easily lead to misidentification, causing the motion control command to be triggered incorrectly and affecting the user experience. Summary of the Invention

[0003] To overcome the problems existing in related technologies, this disclosure provides an action recognition method, device, electronic device, and storage medium.

[0004] According to a first aspect of the present disclosure, an action recognition method is provided, the method comprising:

[0005] Acquire first data, which is user action data collected by at least one sensor during a first time period, wherein the first time period represents the time period during which the user performs an action;

[0006] Based on a preset action recognition model, the action content associated with the first data and the action time period of the action content are determined. The preset action recognition model is used to identify the action content based on the first data. The action time period represents the time period during which the user performs the action content. The duration of the action time period is less than or equal to the duration of the first time period.

[0007] Based on the second data and the preset action classification model, the type of the action content is determined, wherein the second data is user action data collected by the at least one sensor during the second time period, the second time period is at least one of the first preset time period before the action time period and the second preset time period after the action time period, and the preset action classification model is used to classify the action content based on the second data;

[0008] The action recognition result is determined based on the action content and the type of the action content.

[0009] In an exemplary embodiment, determining the type of the action content based on the second data and a preset action classification model includes:

[0010] Based on the second data and the preset action classification model, the probability that the action content is non-target action content is determined;

[0011] The type of the action content is determined based on the probability that the action content is a non-target action content.

[0012] In an exemplary embodiment, determining the type of the action content based on the probability that the action content is non-target action content includes:

[0013] If the probability that the action content is a non-target action content is less than a preset threshold, the type of the action content is determined to be a first type, and the first type is used to indicate that the action content is a target action content;

[0014] If the probability that the action content is a non-target action content is greater than or equal to the preset threshold, the type of the action content is determined to be a second type, and the second type is used to indicate that the action content is a non-target action content.

[0015] In an exemplary embodiment, determining the action recognition result based on the action content and the type of the action content includes:

[0016] If the type of the action content is the first type, the action content is used as the action recognition result, and the first type is used to indicate that the action content is the target action content;

[0017] If the type of the action content is the second type, the action recognition result is determined to be no action, and the second type is used to indicate that the action content is a non-target action content.

[0018] In an exemplary embodiment, determining the type of the action content based on the second data and a preset action classification model includes:

[0019] Extract the associated action features of the action content from the second data. The associated action features are the action features before and / or after the user performs the action content. The associated action features include at least one of signal features, statistical features and morphological features.

[0020] Based on the associated action features and the preset action classification model, the type of the action content is determined.

[0021] In an exemplary embodiment, the preset action recognition model is determined in the following manner:

[0022] Multiple first training sample pairs are acquired, each first training sample pair including reference action content and first training data, the first training data being user action data collected by the at least one sensor within a first reference time period, the first reference time period representing the time period during which the user performs the reference action content;

[0023] Based on the multiple first training sample pairs, the initial action recognition model is trained to obtain the preset action recognition model.

[0024] In an exemplary embodiment, the preset action classification model is determined in the following manner:

[0025] Multiple second training sample pairs are acquired. Each second training sample pair includes second training data and a reference type of reference action content. The second training data is user action data collected by the at least one sensor during a second reference time period. The second reference time period is at least one of a first preset time period before the first reference time period and a second preset time period after the first reference time period. The first reference time period represents the time period during which the user performs the reference action content.

[0026] Based on the multiple second training sample pairs, the initial action classification model is trained to obtain the preset action classification model.

[0027] In one exemplary embodiment, the at least one sensor includes at least one of the following: an accelerometer, a gyroscope, a magnetometer, and a biosensor.

[0028] According to a second aspect of the present disclosure, an action recognition device is provided, the device comprising:

[0029] The acquisition module is configured to acquire first data, which is user action data collected by at least one sensor during a first time period, wherein the first time period represents the time period during which the user performs an action;

[0030] The first determining module is configured to determine the action content associated with the first data and the action time period of the action content based on a preset action recognition model, wherein the preset action recognition model is used to identify the action content based on the first data, the action time period represents the time period during which the user performs the action content, and the duration of the action time period is less than or equal to the duration of the first time period.

[0031] The second determining module is configured to determine the type of the action content based on the second data and a preset action classification model, wherein the second data is user action data collected by the at least one sensor during a second time period, the second time period is at least one of a first preset time period before the action time period and a second preset time period after the action time period, and the preset action classification model is used to classify the action content based on the second data;

[0032] The third determining module is configured to determine the action recognition result based on the action content and the type of the action content.

[0033] According to a third aspect of the present disclosure, an electronic device is provided, comprising:

[0034] processor;

[0035] Memory used to store processor-executable instructions;

[0036] The processor is configured to perform the method described in the first aspect of the embodiments of this disclosure.

[0037] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method described in the first aspect of the present disclosure.

[0038] The method disclosed herein has the following advantages: after identifying the action content based on the first data, the method further analyzes the action content based on the second data to determine the type of the action content, which can avoid the problem of accidental triggering of action control commands. While ensuring the accuracy of the action recognition results, it can also reduce device power consumption and improve the standby time of the device, thereby improving the user experience.

[0039] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0040] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0041] Figure 1 This is a flowchart illustrating an action recognition method according to an exemplary embodiment;

[0042] Figure 2 This is a schematic diagram illustrating an action recognition method according to an exemplary embodiment;

[0043] Figure 3 This is a schematic diagram illustrating a model training process according to an exemplary embodiment;

[0044] Figure 4 This is a flowchart illustrating an action recognition method according to an exemplary embodiment;

[0045] Figure 5 This is a flowchart illustrating an action recognition method according to an exemplary embodiment;

[0046] Figure 6 This is a block diagram illustrating an action recognition device according to an exemplary embodiment;

[0047] Figure 7 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation

[0048] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0049] In related technologies, action recognition algorithms are typically used to distinguish different actions, that is, to identify which of a number of pre-stored specific actions a detected action belongs to. However, sensors do not filter actions when detecting them; regardless of whether it is a specific action or an action used to control the device, it will be identified as one of the pre-stored specific actions. Therefore, misidentification is prone to occur, leading to the false triggering of action control commands and affecting the user experience. For example, the pre-stored specific actions of a device may include waving up and down, but not waving left and right. If a left or right wave is detected, it will also be identified as one of the pre-stored specific actions, resulting in the false triggering of action commands. Furthermore, besides controlling the device, there are many scenarios in daily life where people wave up and down. If the corresponding action control command is triggered every time this action is detected, the corresponding command will also be falsely triggered, affecting the user experience.

[0050] In addition, in order to ensure the accuracy of motion recognition, especially for small-amplitude motion recognition, high-frequency sensors of 208Hz and above are usually required to collect more detailed motion data and to identify the motion through more complex motion recognition algorithms. However, using high-frequency sensors to collect motion data and using heavy-duty motion recognition algorithms will greatly increase the power consumption of the device. For wearable devices such as watches or bracelets, the increased power consumption will lead to a shorter standby time due to the small battery capacity, which will affect the user experience.

[0051] In an exemplary embodiment of this disclosure, to overcome the problems of high power consumption and high false recognition rate in related technologies, an action recognition method is provided, comprising: acquiring first data, wherein the first data is user action data collected by at least one sensor within a first time period, the first time period representing the time period during which the user performs an action; determining action content associated with the first data and the action time period of the action content based on a preset action recognition model, wherein the action time period represents the time period during which the user performs the action content, and the duration of the action time period is less than or equal to the duration of the first time period; determining the type of action content based on second data and a preset action classification model, wherein the second data is user action data collected by at least one sensor within a second time period, the second time period being at least one of a first preset time period before the action time period and a second preset time period after the action time period; and determining an action recognition result based on the action content and the type of action content. This method can avoid the problem of false triggering of action control commands, and while ensuring the accuracy of the action recognition result, it can also reduce device power consumption and improve the standby time of the device, thereby improving the user experience.

[0052] In the exemplary embodiments of this disclosure, an action recognition method is provided, which is applied to electronic devices, including smartphones, tablets, personal computers, smart wearable devices, smart IoT devices, smart vehicle systems, etc. The application scenarios of this method include, but are not limited to, the following scenarios: users control smart wearable devices or applications on wearable devices such as smartwatches or smart bracelets through actions, such as controlling to ignore message pop-ups, controlling volume, etc.; users control mobile phones, tablets, computers, smart home devices, smart vehicle systems, etc. connected to smart wearable devices through actions, such as controlling video playback, controlling music playback, controlling PPT page turning, controlling lights on and off, controlling curtains on and off, controlling the opening or closing of car sunroofs, controlling the in-car air conditioning, etc.

[0053] In the exemplary embodiments of this disclosure, Figure 1 This is a flowchart illustrating an action recognition method according to an exemplary embodiment, such as... Figure 1 As shown, the process includes the following steps S101-S104:

[0054] Step S101: Obtain first data, which is user action data collected by at least one sensor during a first time period, and the first time period represents the time period during which the user performs an action.

[0055] At least one sensor is a sensor in a smart wearable device. The number and type of sensors are determined according to the hardware configuration of the smart wearable device. The frequency of the sensors is not limited, but can be as low as 50Hz. The user action data collected by the sensors is related to the type of sensor. In one example, the user action data collected by the accelerometer is an acceleration signal, and the user action data collected by the biosensor is a biological signal, such as a photoplethysmography (PPG) signal.

[0056] The first time period represents the period during which a user performs an action, and the first data is the user action data within the first time period. The first time period can be determined based on the collected user action data. For example, by performing preliminary identification on the user action data, the approximate time period during which the user performs the action can be identified as the first time period. The user action data within this time period is the first data. For instance, if the user action data before time t1 and after time t2 is relatively stable, while the user action data during the time period t1-t2 fluctuates, then the first time period is the time period t1-t2, and the user action data within the time period t1-t2 is the first data. The first time period can also be any time period determined according to a preset duration. Each preset duration is a first time period, and the user action data within each preset duration is a first data. In some implementations, the first data is obtained from the user action data through a time sliding window with a duration of the preset duration. The preset duration is the maximum value among multiple execution durations corresponding to multiple preset actions in the electronic device. For example, if there are three preset actions in the electronic device, namely action 1, action 2 and action 3, the execution duration of action 1 is 1 second, the execution duration of action 2 is 3 seconds, and the execution duration of action 3 is 2 seconds, then the preset duration is 3 seconds.

[0057] If the electronic device is a smartwatch, smart bracelet, smart ring, smart glasses, or other smart wearable device, it collects user motion data and obtains first data from the user motion data. If the electronic device is another device networked with the smart wearable device, after the smart wearable device collects user motion data, it sends the user motion data to the electronic device, which receives the user motion data. The smart wearable device can send all collected user motion data to the electronic device, which then obtains the first data from all the collected user motion data; alternatively, the smart wearable device can obtain the first data from the collected user motion data and then send the first data to the electronic device, which receives the first data.

[0058] In some embodiments, at least one sensor includes at least one of the following: an accelerometer, a gyroscope, a magnetic sensor, or a biosensor.

[0059] Step S102: Based on the preset action recognition model, determine the action content and the action time period associated with the first data.

[0060] The preset action recognition model is used to recognize action content based on the first data. The action time period represents the time period during which the user performs the action content. The duration of the action time period is less than or equal to the duration of the first time period.

[0061] The preset action recognition model is a neural network model. The model structure of the neural network model includes, but is not limited to, recurrent neural networks, autoencoders, residual networks, multilayer perceptrons, and reinforcement learning networks. In some implementations, the preset action recognition model adopts a structure of two Long Short-Term Memory (LSTM) layers and two convolutional layers. The electronic device stores multiple preset action contents, each corresponding to a control command or multiple control commands for different application scenarios. For example, the preset action content might be a single fist clench or two fist clenches. The preset action recognition model is used to recognize action content based on first data. The first data is input into the preset action recognition model, which identifies which of the multiple preset action contents in the electronic device the user's action corresponds to—that is, it identifies the action content associated with the first data. Simultaneously, after identifying the action content associated with the first data, it can also identify the specific time period during which the user performs the action content—that is, the action time period. The output of the preset action recognition model is the action content associated with the first data and the action time period of the action content. Since the action content associated with the first data is identified from the user action data within the first time period, the time period in which the user performs the action content is also identified within the first time period. Therefore, the action time period is located within the first time period, and the duration of the action time period is less than or equal to the duration of the first time period.

[0062] Step S103: Based on the second data and the preset action classification model, determine the type of action content.

[0063] The second data refers to user action data collected by at least one sensor during the second time period. The second time period is at least one of a first preset time period before the action period and a second preset time period after the action period. The preset action classification model is used to classify the action content based on the second data.

[0064] After determining the action time period of the action content, a first preset time period before the action time period and a second preset time period after the action time period are selected. The end time of the first preset time period is the start time of the action time period, and the start time of the second preset time period is the end time of the action time period. The duration of the first preset time period and the duration of the second preset time period can be the same or different. Both the duration of the first preset time period and the duration of the second preset time period can be multiples of the duration of the action time period. For example, the duration of the first preset time period is 0.5 times the duration of the action time period, and the duration of the second preset time period is 1 time the duration of the action time period. Optionally, the duration of either the first or second preset time period can be between 0.5 and 2 times the duration of the action time period. The second time period can be either the first or the second preset time period. To ensure the accuracy of the recognition results, the second time period can also be either the first or the second preset time period.

[0065] The second data is user behavior data within the second time period. If the electronic device is a smartwatch, smart bracelet, smart ring, smart glasses, or other smart wearable device, the second data is obtained from the collected user action data through the smart wearable device. If the electronic device is another device networked with the smart wearable device, the smart wearable device sends all collected user action data to the electronic device, which receives all user action data and, after determining the second time period, obtains the second data from it; alternatively, the electronic device can determine the second time period and send it to the smart wearable device, which, upon receiving the second time period, obtains the second data from the collected user action data and sends the second data to the electronic device, which then receives the second data.

[0066] The preset action classification model is a machine learning model, employing classification models including, but not limited to, Lightweight Gradient Boosting (LightGBM), Support Vector Machines, eXtreme Gradient Boosting (XGBoost), and Random Forest. The electronic device stores multiple preset action types, the number of which can be set according to actual needs. The preset action classification model is used to classify action content based on second data. The second data is input into the preset action classification model, which classifies the action content according to the second data, determining which of the multiple preset action types the action content belongs to, and outputting the type of the action content. Specifically, if there are two preset action types, the preset action classification model is a binary classification model; if there are multiple preset action types, the preset action classification model is a multi-class classification model.

[0067] Since the second data represents user action data before and / or after the user performs an action, it allows for the determination of the user's pre-action and / or post-action motion data. For similar or identical actions, the second data enables more specific differentiation to determine the type of action. In one example, the type of action indicates whether the actual action is one of the preset actions. If the identified action is an up-and-down hand gesture, since the pre-action and / or post-action motions of similar actions can differ, further analysis based on the user's actions before or after performing the action can be conducted to verify its accuracy. If the second data differs from the pre-action and / or post-action motion data of the up-and-down hand gesture, it indicates an error in action identification; the actual action is not one of the preset actions, and other similar actions have been identified as an up-and-down hand gesture. If the second data is the same as the pre-action and / or post-action motion data of the up-and-down hand gesture, it indicates correct action identification; the actual action is one of the preset actions. In another example, the type of action content indicates whether it is an action content used for device control. If the identified action content is waving up and down, we can analyze whether the user is using the action content to control the device or an unintentional action in other life scenarios based on the actions before or after the user performs the action content. For example, if there are no other actions before or after performing the action content, it means that the user is likely using the action content to control the device. If there are continuous waving up and down before or after performing the action content, it means that the user is likely using the action content unintentionally in other life scenarios, such as waving up and down while wiping a table.

[0068] Step S104: Determine the action recognition result based on the action content and the type of action content.

[0069] If the type of the action content indicates that the action content is one of the preset action content or an action content used for motion control of the device, then the action recognition result is determined to be the recognized action content, and the corresponding action control command is triggered; if the type of the action content indicates that the action content is not one of the preset action content or an action content used for motion control of the device, then the action recognition result is determined to be that the corresponding action control command is not triggered.

[0070] In one example, Figure 2 This is a schematic diagram illustrating an action recognition method according to an exemplary embodiment, such as... Figure 2As shown, user action data is collected by at least one sensor in a smart wearable device. First data is obtained from the user action data by sliding a time window. The first data is input into a preset action recognition model, which outputs the action content and the action time period of the action content. Second data is obtained from the user action data based on the action time period. The second data is input into a preset action classification model, which outputs the type of action content. The action recognition result is determined according to the action content and the type of action content.

[0071] In the exemplary embodiments of this disclosure, user action data collected by at least one sensor during a first time period is input into a preset action recognition model to determine the action content and the action time period of the action content. At least one of the first preset time period before the action time period and the second preset time period after the action time period is taken as the second time period. User action data collected by at least one sensor during the second time period is input into a preset action classification model to determine the type of action content. Based on the action content and the type of action content, the action recognition result is determined. This method, after recognizing the action content based on the first data, further analyzes the action content based on the second data to determine the type of action content. This can improve the accuracy of the action recognition result and avoid the problem of erroneous triggering of action control commands due to the action content not being one of the preset action content or not being the action content used for action control of the device, thereby improving the user experience. In addition, since at least one sensor is used to collect action data, and the type of action content is further determined based on the second data after the action content is recognized, the requirements for sensor frequency and action recognition algorithm accuracy are reduced. Even using low-frequency sensors and relatively simple algorithm models, the accuracy of action recognition results can be ensured, thereby reducing device power consumption and increasing device standby time.

[0072] In some embodiments, the preset action recognition model is determined in the following ways:

[0073] Multiple first training sample pairs are acquired, each first training sample pair including reference action content and first training data. The first training data is data obtained by at least one sensor during a first reference time period, and the first reference time period represents the time period during which the user performs the reference action content.

[0074] Based on multiple first training sample pairs, the initial action recognition model is trained to obtain the preset action recognition model.

[0075] The reference action content can be any action content. The more types of reference action content there are, the better the performance of the preset action recognition model will be. After determining the reference action content, first training data corresponding to each reference action content is collected, and the correspondence between the first training data and the reference action content is labeled. Multiple different first training data can be collected for each reference action content. For example, different first training data can be collected when different users perform the reference action content, or different first training data can be collected when the same user performs the reference action content in different scenarios, in order to adapt to the differences between different users or different scenarios. Each first training data and the reference action content corresponding to the first training data form a first training sample pair. The first training data from all the first training sample pairs are input into the initial action recognition model for iterative training. The initial parameters of the initial action recognition model are randomly initialized. In each iteration, the predicted action content and the predicted time period of the predicted action content corresponding to the first training data in each first training sample pair are output. It is determined whether the predicted action content corresponding to the first training data in each first training sample pair is the same as its corresponding reference action content, and whether the predicted time period corresponding to the first training data in each first training sample pair is the same as its corresponding first reference time period. If they are the same, it means that the action content recognized by the first training sample pair is correct; if they are different, it means that the action content recognized by the first training sample pair is incorrect. The ratio of the number of correctly recognized first training sample pairs to the total number of first training sample pairs is calculated. The parameters of the initial action recognition model are adjusted according to the output results. The adjusted model parameters are used for the next iteration of training until the ratio of the number of correctly recognized first training sample pairs to the total number of first training sample pairs is greater than a first preset ratio. The training ends when the preset action recognition model is obtained.

[0076] In some embodiments, the preset action classification model is determined in the following ways:

[0077] Multiple second training sample pairs are acquired. Each second training sample pair includes second training data and a reference type of reference action content. The second training data is user action data collected by at least one sensor during a second reference time period. The second reference time period is at least one of a first preset time period before the first reference time period and a second preset time period after the first reference time period. The first reference time period represents the time period during which the user performs the reference action content.

[0078] The initial action classification model is trained based on multiple second training sample pairs to obtain the preset action classification model.

[0079] After determining the reference action content, the reference type of each reference action content is labeled. When collecting the first training data corresponding to each reference action content, the second training data is collected simultaneously, and the correspondence between the second training data and the reference type is labeled. Multiple different second training data can be collected for each reference type. For example, different second training data can be collected before and after different users execute the reference action content, or different second training data can be collected before and after the same user executes the reference action content in different scenarios, in order to adapt to the differences between different users or different scenarios. Each second training data and the reference type corresponding to the second training data form a second training sample pair. All second training sample pairs are divided into training and test sets. An initial action classification model is trained using data from the training set. Second training data from all second training sample pairs in the training set are input into the initial action classification model for iterative training. The initial parameters of the initial action recognition model are randomly initialized. In each iteration, the predicted type corresponding to the second training data in each second training sample pair is output. It is determined whether the predicted type of the second training data in each second training sample pair is the same as its corresponding reference type. If they are the same, the classification is correct; if they are different, the classification is incorrect. The ratio of the number of correctly classified second training sample pairs to the total number of second training sample pairs in the training set is calculated. The parameters of the initial action classification model are adjusted based on the output results. The adjusted model parameters are used for the next iteration of training until the ratio of the number of correctly classified second training sample pairs to the total number of second training sample pairs in the training set is greater than a set value, at which point training ends. The trained initial action classification model is tested using data from the test set. The ratio of the number of correctly classified second training sample pairs using the model to the total number of second training sample pairs in the test set is determined. If this ratio is greater than a second preset ratio, the trained initial action classification model is used as the preset action classification model.

[0080] In some implementations, an initial action recognition model and an initial action classification model are trained simultaneously based on multiple third training sample pairs. After determining the reference action content and the reference type for each reference action content, first training data is collected when each reference action content is executed, as well as second training data before and after the execution of each reference action content. The correspondence between the first and second training data and the reference action content and its reference type is labeled, with each pair of correspondences serving as a third training sample pair. In one example, Figure 3 This is a schematic diagram illustrating a model training process according to an exemplary embodiment, such as... Figure 3As shown, user action data is collected by at least one sensor in a smart wearable device. First training data and second training data are obtained from the user action data by sliding a time window. The first training data is input into an initial action recognition model, which outputs the predicted action content and the predicted time period of the predicted action content. The second training data is input into an initial action classification model, which outputs the predicted type of the predicted action content. The initial action recognition model is trained based on the predicted action content and the reference action content. The initial action classification model is trained based on the predicted type and the reference type, thus obtaining a preset action recognition model and a preset action classification model.

[0081] It should be noted that when using the model for action recognition, it is necessary to ensure that the duration of the second time segment is the same as the duration of the second reference time segment when training the model.

[0082] In the exemplary embodiments of this disclosure, Figure 4 This is a flowchart illustrating an action recognition method according to an exemplary embodiment, such as... Figure 4 As shown, the process includes the following steps S401-S405:

[0083] Step S401: Obtain first data, which is user action data collected by at least one sensor during a first time period, and the first time period represents the time period during which the user performs an action.

[0084] Step S402: Based on the preset action recognition model, determine the action content and the action time period associated with the first data.

[0085] The preset action recognition model is used to recognize action content based on the first data. The action time period represents the time period during which the user performs the action content. The duration of the action time period is less than or equal to the duration of the first time period.

[0086] For specific implementation methods of steps S401 and S402, please refer to steps S101-S102, which will not be repeated here.

[0087] Step S403: Extract the associated action features of the action content from the second data. The associated action features are the action features before and / or after the user executes the action content. The associated action features include at least one of signal features, statistical features and morphological features.

[0088] The associated action features of the action content are extracted from the second data using a preset feature extraction structure. If the second time period is the first preset time period before the user executes the action content, the associated action features are the action features before the user executes the action content, i.e., the pre-action action features of the action content; if the second time period is the second preset time period after the user executes the action content, the associated action features are the action features after the user executes the action content, i.e., the post-action action features of the action content; if the second time period is both the first and second preset time periods, the associated action features are the action features before and after the user executes the action content. The associated action features include at least one of signal features, statistical features, and morphological features. For example, signal features include time domain features, frequency domain features, etc.; statistical features include deviation, variance, mean, median, percentage, etc.; and morphological features include peaks, troughs, slopes, etc.

[0089] Step S404: Determine the type of action content based on associated action features and a preset action classification model.

[0090] The preset feature extraction structure used to extract associated action features can be a part of the preset action classification model or a separate structure. If the preset feature extraction is a separate structure, the extracted associated action features are input into the preset action classification model, and the type of action content is output. If the preset feature extraction is a part of the preset action classification model, the extracted associated action features are input into the remaining parts of the preset action classification model, and the type of action content is output.

[0091] Step S405: Determine the action recognition result based on the action content and the type of action content.

[0092] For a detailed implementation of step S405, please refer to step S104, which will not be repeated here.

[0093] In this embodiment, by extracting various associated action features from the second data and classifying the action content, the accuracy of the determined action content type can be improved.

[0094] In the exemplary embodiments of this disclosure, Figure 5 This is a flowchart illustrating an action recognition method according to an exemplary embodiment, such as... Figure 5 As shown, the steps S501-S505 are included:

[0095] Step S501: Obtain first data, which is user action data collected by at least one sensor during a first time period, and the first time period represents the time period during which the user performs an action.

[0096] Step S502: Based on the preset action recognition model, determine the action content and the action time period associated with the first data.

[0097] The preset action recognition model is used to recognize action content based on the first data. The action time period represents the time period during which the user performs the action content. The duration of the action time period is less than or equal to the duration of the first time period.

[0098] For the specific implementation of steps S501 and S502, please refer to steps S101-S102, which will not be repeated here.

[0099] Step S503: Based on the second data and the preset action classification model, determine the probability that the action content is non-target action content.

[0100] The second data is user action data collected by at least one sensor during the second time period, and the second time period is at least one of a first preset time period before the action period and a second preset time period after the action period.

[0101] The type of action content indicates whether the action content is a target action content. Action content that is a target action content indicates that the action content is one of the preset action content or that the action content is used for motion control of the device. Action content that is a non-target action content indicates that the action content is not one of the preset action content or that the action content is not used for motion control of the device. The second data is input into the preset action classification model. The preset action classification model determines the probability that the action content is a non-target action content based on the second data and outputs the probability that the action content is a non-target action content. It should be noted that when training the preset action classification model, the output of the preset action classification model is also the probability that the action content is a non-target action content.

[0102] Step S504: Determine the type of action content based on the probability that the action content is a non-target action content.

[0103] Based on the probability that the action content is a non-target action content, the action content is divided into different types, which are used to indicate whether the action content is a target action content or a non-target action content.

[0104] In some implementations, if the probability that the action content is a non-target action content is less than a preset threshold, the type of the action content is determined to be a first type, which is used to indicate that the action content is a target action content; if the probability that the action content is a non-target action content is greater than or equal to the preset threshold, the type of the action content is determined to be a second type, which is used to indicate that the action content is a non-target action content.

[0105] If the probability that the action content is a non-target action content is less than a preset threshold, it means that the action content is the target action content, and the type of the action content is determined to be the first type; if the probability that the action content is a non-target action content is greater than or equal to the preset threshold, it means that the action content is a non-target action content, and the type of the action content is determined to be the second type.

[0106] Step S505: Determine the action recognition result based on the action content and the type of action content.

[0107] In some implementations, if the type of the action content is a first type, the action content is used as the action recognition result, and the first type is used to indicate that the action content is the target action content; if the type of the action content is a second type, the action recognition result is determined to be no action, and the second type is used to indicate that the action content is not the target action content.

[0108] If the type of the action content is the first type, it means that the action content is the target action content, that is, the action content is one of the preset action content or the action content used to control the device. In this case, the action recognition result is determined to be the recognized action content, and the corresponding action control command is triggered. If the type of the action content is the second type, it means that the action content is not the target action content, that is, the action content is not one of the preset action content or the action content used to control the device. In this case, the action recognition result is determined to be no action, that is, the corresponding action control command is not triggered.

[0109] In some embodiments, step S504 can be omitted, and step S505 can be replaced by: determining the action recognition result based on the action content and the probability that the action content is non-target action content.

[0110] In some implementations, if the probability that the action content is a non-target action content is less than a preset threshold, the action content is taken as the action recognition result; if the probability that the action content is a non-target action content is greater than or equal to the preset threshold, the action recognition result is determined to be no action.

[0111] In an exemplary embodiment of this disclosure, an action recognition device is provided. Figure 6 This is a block diagram illustrating an action recognition device according to an exemplary embodiment, such as... Figure 6 As shown, the motion recognition device includes:

[0112] The acquisition module 601 is configured to acquire first data, which is user action data collected by at least one sensor during a first time period, and the first time period represents the time period during which the user performs an action;

[0113] The first determining module 602 is configured to determine the action content and the action time period of the action content associated with the first data based on a preset action recognition model. The preset action recognition model is used to identify the action content based on the first data. The action time period represents the time period during which the user performs the action content. The duration of the action time period is less than or equal to the duration of the first time period.

[0114] The second determining module 603 is configured to determine the type of action content based on the second data and the preset action classification model. The second data is user action data collected by at least one sensor during the second time period. The second time period is at least one of the first preset time period before the action period and the second preset time period after the action period. The preset action classification model is used to classify the action content based on the second data.

[0115] The third determining module 604 is configured to determine the action recognition result based on the action content and the type of the action content.

[0116] In one exemplary embodiment, the second determining module 603 is further configured to:

[0117] Based on the second data and the preset action classification model, determine the probability that the action content is non-target action content;

[0118] The type of action content is determined based on the probability that the action content is a non-target action content.

[0119] In one exemplary embodiment, the second determining module 603 is further configured to:

[0120] If the probability that the action content is a non-target action content is less than a preset threshold, the type of the action content is determined to be the first type. The first type is used to indicate that the action content is the target action content.

[0121] If the probability that the action content is a non-target action content is greater than or equal to a preset threshold, the type of the action content is determined to be the second type. The second type is used to indicate that the action content is a non-target action content.

[0122] In one exemplary embodiment, the third determining module 604 is further configured to:

[0123] If the type of the action content is the first type, the action content will be used as the action recognition result. The first type is used to indicate that the action content is the target action content.

[0124] If the type of the action content is the second type, the action recognition result is determined to be no action. The second type is used to indicate that the action content is not the target action content.

[0125] In one exemplary embodiment, the second determining module 603 is further configured to:

[0126] Extract the associated action features of the action content from the second data. The associated action features are the action features before and / or after the user performs the action content. The associated action features include at least one of signal features, statistical features and morphological features.

[0127] Based on associated action features and a pre-defined action classification model, the type of action content is determined.

[0128] In one exemplary embodiment, the action recognition device further includes a training module 605, configured to:

[0129] Multiple first training sample pairs are acquired. Each first training sample pair includes reference action content and first training data. The first training data is user action data collected by at least one sensor during a first reference time period. The first reference time period represents the time period during which the user performs the reference action content.

[0130] Based on multiple first training sample pairs, the initial action recognition model is trained to obtain the preset action recognition model.

[0131] In one exemplary embodiment, the training module 605 is further configured to:

[0132] Multiple second training sample pairs are acquired. Each second training sample pair includes second training data and a reference type of reference action content. The second training data is user action data collected by at least one sensor during a second reference time period. The second reference time period is at least one of a first preset time period before the first reference time period and a second preset time period after the first reference time period. The first reference time period represents the time period during which the user performs the reference action content.

[0133] The initial action classification model is trained based on multiple second training sample pairs to obtain the preset action classification model.

[0134] In one exemplary embodiment, at least one sensor includes at least one of the following: an accelerometer, a gyroscope, a magnetic sensor, and a biosensor.

[0135] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0136] Figure 7 This is a block diagram illustrating an electronic device 700 according to an exemplary embodiment.

[0137] Reference Figure 7The electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.

[0138] Processing component 702 typically controls the overall operation of electronic device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.

[0139] Memory 704 is configured to store various types of data to support the operation of electronic device 700. Examples of this data include instructions for any application or method operating on electronic device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0140] Power supply component 706 provides power to various components of electronic device 700. Power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 700.

[0141] Multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0142] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when electronic device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.

[0143] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0144] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of electronic device 700. For example, sensor assembly 714 can detect the on / off state of electronic device 700, the relative positioning of components such as the display and keypad of electronic device 700, changes in position of electronic device 700 or a component of electronic device 700, the presence or absence of user contact with electronic device 700, orientation or acceleration / deceleration of electronic device 700, and temperature changes of electronic device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0145] Communication component 716 is configured to facilitate wired or wireless communication between electronic device 700 and other devices. Electronic device 700 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0146] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0147] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by a processor 720 of an electronic device 700 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0148] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform an action recognition method, the action recognition method including any of the methods described above.

[0149] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.

[0150] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. An action recognition method, characterized in that, The method includes: Acquire first data, which is user action data collected by at least one sensor during a first time period, wherein the first time period represents the time period during which the user performs an action; Based on a preset action recognition model, the action content associated with the first data and the action time period of the action content are determined. The preset action recognition model is used to identify the action content based on the first data. The action time period represents the time period during which the user performs the action content. The duration of the action time period is less than or equal to the duration of the first time period. Based on the second data and the preset action classification model, the type of the action content is determined, wherein the second data is user action data collected by the at least one sensor during the second time period, the second time period is at least one of the first preset time period before the action time period and the second preset time period after the action time period, and the preset action classification model is used to classify the action content based on the second data; The action recognition result is determined based on the action content and the type of the action content.

2. The action recognition method according to claim 1, characterized in that, The determination of the type of action content based on the second data and a preset action classification model includes: Based on the second data and the preset action classification model, the probability that the action content is non-target action content is determined; The type of the action content is determined based on the probability that the action content is a non-target action content.

3. The action recognition method according to claim 2, characterized in that, Determining the type of the action content based on the probability that the action content is a non-target action content includes: If the probability that the action content is a non-target action content is less than a preset threshold, the type of the action content is determined to be a first type, and the first type is used to indicate that the action content is a target action content; If the probability that the action content is a non-target action content is greater than or equal to the preset threshold, the type of the action content is determined to be a second type, and the second type is used to indicate that the action content is a non-target action content.

4. The action recognition method according to claim 1 or 3, characterized in that, The determination of the action recognition result based on the action content and the type of the action content includes: If the type of the action content is the first type, the action content is used as the action recognition result, and the first type is used to indicate that the action content is the target action content; If the type of the action content is the second type, the action recognition result is determined to be no action, and the second type is used to indicate that the action content is a non-target action content.

5. The action recognition method according to claim 1, characterized in that, The determination of the type of action content based on the second data and a preset action classification model includes: Extract the associated action features of the action content from the second data. The associated action features are the action features before and / or after the user performs the action content. The associated action features include at least one of signal features, statistical features and morphological features. Based on the associated action features and the preset action classification model, the type of the action content is determined.

6. The action recognition method according to claim 1, characterized in that, The preset action recognition model is determined in the following way: Multiple first training sample pairs are acquired, each first training sample pair including reference action content and first training data, the first training data being user action data collected by the at least one sensor within a first reference time period, the first reference time period representing the time period during which the user performs the reference action content; Based on the multiple first training sample pairs, the initial action recognition model is trained to obtain the preset action recognition model.

7. The action recognition method according to claim 1, characterized in that, The preset action classification model is determined in the following way: Multiple second training sample pairs are acquired. Each second training sample pair includes second training data and a reference type of reference action content. The second training data is user action data collected by the at least one sensor during a second reference time period. The second reference time period is at least one of a first preset time period before the first reference time period and a second preset time period after the first reference time period. The first reference time period represents the time period during which the user performs the reference action content. Based on the multiple second training sample pairs, the initial action classification model is trained to obtain the preset action classification model.

8. The action recognition method according to claim 1, characterized in that, The at least one sensor includes at least one of the following: an accelerometer, a gyroscope, a magnetic sensor, or a biosensor.

9. A motion recognition device, characterized in that, The device includes: The acquisition module is configured to acquire first data, which is user action data collected by at least one sensor during a first time period, wherein the first time period represents the time period during which the user performs an action; The first determining module is configured to determine the action content associated with the first data and the action time period of the action content based on a preset action recognition model, wherein the preset action recognition model is used to identify the action content based on the first data, the action time period represents the time period during which the user performs the action content, and the duration of the action time period is less than or equal to the duration of the first time period. The second determining module is configured to determine the type of the action content based on the second data and a preset action classification model, wherein the second data is user action data collected by the at least one sensor during a second time period, the second time period is at least one of a first preset time period before the action time period and a second preset time period after the action time period, and the preset action classification model is used to classify the action content based on the second data; The third determining module is configured to determine the action recognition result based on the action content and the type of the action content.

10. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to perform the method as described in any one of claims 1-8.

11. A non-transitory computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-8.