Action recognition prediction method, device, server and storage medium

By mapping the feature sequence of time series data to high-dimensional space and training of preset models, the problem of difficulty in accurately identifying action types in the prior art is solved, and accurate recognition of actions and real-time feedback are achieved.

CN114510981BActive Publication Date: 2025-05-13CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011281245.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-16
Publication Date
2025-05-13
Estimated Expiration
2040-11-16

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the types of actions in time series data of the user's daily behavior activities collected by the wearable device.

Method used

By obtaining the original time series data, extracting the feature sequence, and mapping it into high-dimensional space, and using a preset model for training, a prediction model for identifying actions is obtained.

Benefits of technology

It realizes accurate identification of actions in the target time series data, and can be fed back to the terminal in real time, improving the accuracy and efficiency of action recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114510981B_ABST
    Figure CN114510981B_ABST
Patent Text Reader

Abstract

The present invention discloses a prediction method, device, server and storage medium for action recognition. The method comprises: obtaining original time series data; extracting feature sequences corresponding to the original time series data; mapping the feature sequences to high-dimensional space to obtain feature sequences in the high-dimensional space; using the feature sequences in the high-dimensional space as sample training data to train a preset model to obtain a prediction model for action recognition; wherein the prediction model is used to recognize actions in target time series data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a prediction method, device, server and storage medium for action recognition. Background Art

[0002] With the rapid development of machine learning and Internet of Things technology, the fields of intelligent monitoring, human-computer interaction, medical health, etc. are developing more and more rapidly. At present, in the case of busy human social behavior life, the time series classification technology based on daily activities can not only allow users to have a clear understanding of their own activity patterns, but also provide users with more personalized recording support. Therefore, how to perceive the behavioral time series information of users or other applications, and perform timely actions or services based on this, has become a new challenge. In related technologies, the sensors equipped with wearable devices can be used to collect the time series data of users' daily behavior activities, and the time series data can be recorded through installed applications. However, it is impossible to accurately identify the action type of the recorded time series data. Summary of the invention

[0003] In view of this, embodiments of the present invention are intended to provide a prediction method, device, server, and storage medium for action recognition.

[0004] The technical solution of the embodiment of the present invention is achieved as follows:

[0005] At least one embodiment of the present invention provides a prediction method for action recognition, the method comprising:

[0006] Get the original time series data;

[0007] Extracting a feature sequence corresponding to the original time series data;

[0008] Mapping the feature sequence into a high-dimensional space to obtain a feature sequence in the high-dimensional space;

[0009] Using the characteristic sequence in the high-dimensional space as sample training data, training a preset model to obtain a prediction model for identifying the action;

[0010] The prediction model is used to identify actions in target time series data.

[0011] In addition, according to at least one embodiment of the present invention, mapping the feature sequence into a high-dimensional space comprises:

[0012] intercepting the feature sequence to obtain a plurality of one-dimensional feature sequences;

[0013] The multiple one-dimensional feature sequences are mapped into a high-dimensional space using a first preset parameter and a second preset parameter; the first parameter represents the dimension of the high-dimensional space, and the second parameter represents the flow shape formed by the feature sequences in the high-dimensional space.

[0014] In addition, according to at least one embodiment of the present invention, the using the characteristic sequence in the high-dimensional space as sample training data to train a preset model includes:

[0015] Determine a plurality of hyperplanes by using the characteristic sequence in the high-dimensional space;

[0016] Using the multiple hyperplanes, dividing the high-dimensional space into multiple subspaces;

[0017] The feature sequences of each subspace are used as sample training data to train the preset model.

[0018] In addition, according to at least one embodiment of the present invention, the determining of multiple hyperplanes by using the characteristic sequence in the high-dimensional space includes:

[0019] Grouping the feature sequences in the high-dimensional space to obtain multiple groups of feature sequences;

[0020] For each group of feature sequences in the multiple groups of feature sequences, a hyperplane corresponding to the corresponding group of feature sequences is determined to obtain multiple hyperplanes.

[0021] In addition, according to at least one embodiment of the present invention, the method of using the feature sequences of each subspace as sample training data to train the preset model includes:

[0022] For each feature sequence of each subspace in each subspace, count the number of feature points that meet the preset conditions in the corresponding feature sequence; and determine the total number of feature points included in the corresponding feature sequence;

[0023] Calculate the ratio of the number of statistical feature points to the total number of feature points contained in the corresponding feature sequence;

[0024] When the ratio is greater than or equal to a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the first action, and is used as the training result of the preset model.

[0025] In addition, according to at least one embodiment of the present invention, the method further includes:

[0026] When the ratio is less than a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the second action, and is used as the training result of the preset model.

[0027] In addition, according to at least one embodiment of the present invention, the counting of the number of feature points in the corresponding feature sequence that meet a preset condition includes:

[0028] Determine the coordinates of each feature point in the corresponding feature sequence;

[0029] For each feature point in the corresponding feature sequence, determine whether the coordinates of the corresponding feature point meet the preset coordinate conditions;

[0030] When the coordinates of the corresponding feature points meet the preset coordinate conditions, the corresponding feature points are used as coordinate points to be processed, and a plurality of coordinate points to be processed are obtained;

[0031] Count the number of the plurality of feature points to be processed.

[0032] In addition, according to at least one embodiment of the present invention, when performing one training of the preset model, the method further includes:

[0033] For each subspace in each subspace, determine whether the corresponding subspace has been trained to obtain the corresponding training result;

[0034] When it is determined that the corresponding subspace has been trained to obtain the corresponding training result, the corresponding subspace is excluded;

[0035] The feature sequences in the remaining subspace are used as sample training data to train the prediction model.

[0036] At least one embodiment of the present invention provides a prediction device for action recognition, comprising:

[0037] An acquisition unit, used to acquire original time series data;

[0038] A first processing unit is used to extract a feature sequence corresponding to the original time series data; map the feature sequence to a high-dimensional space to obtain a feature sequence in the high-dimensional space;

[0039] A second processing unit is used to train a preset model using the characteristic sequence in the high-dimensional space as sample training data to obtain a prediction model for identifying the action;

[0040] The prediction model is used to identify actions in target time series data.

[0041] In addition, according to at least one embodiment of the present invention, the first processing unit is specifically configured to:

[0042] intercepting the feature sequence to obtain a plurality of one-dimensional feature sequences;

[0043] The multiple one-dimensional feature sequences are mapped into a high-dimensional space using a first preset parameter and a second preset parameter; the first parameter represents the dimension of the high-dimensional space, and the second parameter represents the flow shape formed by the feature sequences in the high-dimensional space.

[0044] In addition, according to at least one embodiment of the present invention, the second processing unit is specifically configured to:

[0045] Determine a plurality of hyperplanes by using the characteristic sequence in the high-dimensional space;

[0046] Using the multiple hyperplanes, dividing the high-dimensional space into multiple subspaces;

[0047] The feature sequences of each subspace are used as sample training data to train the preset model.

[0048] In addition, according to at least one embodiment of the present invention, the determining of multiple hyperplanes by using the characteristic sequence in the high-dimensional space includes:

[0049] Grouping the feature sequences in the high-dimensional space to obtain multiple groups of feature sequences;

[0050] For each group of feature sequences in the multiple groups of feature sequences, a hyperplane corresponding to the corresponding group of feature sequences is determined to obtain multiple hyperplanes.

[0051] In addition, according to at least one embodiment of the present invention, the second processing unit is specifically used to:

[0052] For each feature sequence of each subspace in each subspace, count the number of feature points that meet the preset conditions in the corresponding feature sequence; and determine the total number of feature points included in the corresponding feature sequence;

[0053] Calculate the ratio of the number of statistical feature points to the total number of feature points contained in the corresponding feature sequence;

[0054] When the ratio is greater than or equal to a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the first action, and is used as the training result of the preset model.

[0055] In addition, according to at least one embodiment of the present invention, the second processing unit is further configured to:

[0056] When the ratio is less than a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the second action, and is used as the training result of the preset model.

[0057] In addition, according to at least one embodiment of the present invention, the second processing unit is specifically configured to:

[0058] Determine the coordinates of each feature point in the corresponding feature sequence; for each feature point in the corresponding feature sequence, determine whether the coordinates of the corresponding feature point meet the preset coordinate conditions; when the coordinates of the corresponding feature point meet the preset coordinate conditions, use the corresponding feature point as a coordinate point to be processed to obtain multiple coordinate points to be processed; and count the number of the multiple feature points to be processed.

[0059] In addition, according to at least one embodiment of the present invention, the second processing unit is specifically configured to:

[0060] When performing a training of the preset model, for each subspace in each subspace, it is determined whether the corresponding subspace has been trained to obtain the corresponding training result;

[0061] When it is determined that the corresponding subspace has been trained to obtain the corresponding training result, the corresponding subspace is excluded;

[0062] The feature sequences in the remaining subspace are used as sample training data to train the prediction model.

[0063] At least one embodiment of the present invention provides a server, comprising:

[0064] Communication interface, used to obtain raw time series data;

[0065] A processor is used to extract a feature sequence corresponding to the original time series data; map the feature sequence to a high-dimensional space to obtain a feature sequence in the high-dimensional space; and use the feature sequence in the high-dimensional space as sample training data to train a preset model to obtain a prediction model for identifying actions; wherein the prediction model is used to identify actions in the target time series data.

[0066] At least one embodiment of the present invention provides a server, comprising a processor and a memory for storing a computer program that can be run on the processor.

[0067] Wherein, the processor is used to execute the steps of any of the above methods when running the computer program.

[0068] At least one embodiment of the present invention provides a storage medium having a computer program stored thereon, wherein the computer program implements the steps of any of the above methods when executed by a processor.

[0069] The prediction method, device, equipment and storage medium for action recognition provided by the embodiment of the present invention obtains original time series data; extracts the feature sequence corresponding to the original time series data; maps the feature sequence to a high-dimensional space to obtain the feature sequence in the high-dimensional space; uses the feature sequence in the high-dimensional space as sample training data to train a preset model to obtain a prediction model for action recognition; wherein the prediction model is used to recognize actions in target time series data. Using the technical solution of the embodiment of the present invention, the feature sequence corresponding to the original time series data is mapped to a high-dimensional space, and the feature sequence in the high-dimensional space is used as sample training data to train a preset model, so that the trained model can be used to recognize the action category of the target time series and provide real-time feedback to the terminal. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 It is a schematic diagram of the system architecture of the application of the prediction method for action recognition according to an embodiment of the present invention;

[0071] Figure 2 It is a schematic diagram of an implementation flow of a prediction method for action recognition provided by an embodiment of the present invention;

[0072] Figure 3a and Figure 3b Schematic diagram of two artificial time series provided by the embodiment of the present invention Figure 1 ;

[0073] Figure 4a and Figure 4b Schematic diagram of two artificial time series provided by the embodiment of the present invention Figure 2 ;

[0074] Figure 5a and Figure 5b It is a schematic diagram of an implementation flow of mapping a feature sequence corresponding to an original time series into a high-dimensional space in an embodiment of the present invention;

[0075] Figure 6 It is a schematic diagram of an implementation flow of mapping a feature sequence corresponding to an original time series into a high-dimensional space in an embodiment of the present invention;

[0076] Figure 7 is a schematic diagram of dividing a high-dimensional space into two subspaces according to an embodiment of the present invention;

[0077] Figure 8 is a schematic diagram of a hyperplane H1 according to an embodiment of the present invention;

[0078] Fig. 9 It is a schematic diagram of the implementation process of training the prediction model of the embodiment of the present invention and using the trained prediction model to perform action recognition on the target time series;

[0079] Fig.10 It is a schematic diagram of the composition structure of a prediction device for action recognition provided by an embodiment of the present invention;

[0080] Fig.11 It is a schematic diagram of the composition structure of the server provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0081] Before introducing the technical solution of the embodiment of the present invention, the related technology is first described.

[0082] In the related technologies, with the rapid development of machine learning and Internet of Things technology, the fields of intelligent monitoring, human-computer interaction, medical health, etc. are developing more and more rapidly. At present, in the case of busy human social behavior life, the time series classification technology based on daily activities can not only allow users to have a clear understanding of their own activity patterns, but also provide users with more personalized recording support. Therefore, how to perceive the behavioral time series information of users or other applications and perform timely actions or services accordingly has become a new challenge. In the related technologies, the data of daily behavior activities can be collected by sensors equipped with wearable devices such as mobile phones, watches, and glasses, and the data of daily behavior activities of users can be recorded through installed applications such as WeChat, Keep, Mint Health, etc., to obtain time series data, and the collected time series data can be classified using traditional time series classification methods; among them, the WeChat application can record the number of steps a user walks in a day, and the Keep application can record the duration of a user's running, etc. Traditional time series classification methods can be divided into four categories: classification methods based on integration, classification methods based on models, classification methods based on local features, and classification methods based on global features.

[0083] In the related art, the use of traditional time series classification methods to classify collected time series data has the following defects: First, the characteristics of the collected time series are not taken into account, that is, the time series data corresponding to the daily behavior activities of users collected by multi-functional sensors have the characteristics of "massive and high-dimensional" big data and "real-time updated" "streaming data". Second, when cutting the collected time series, the activity type corresponding to the time series is not considered, and the cutting length and cutting start and end points of each time slice are determined solely by human experience, resulting in more than one action category corresponding to a time slice, and poor scalability and reproducibility. Since the position and time of each slice or sequence of time stream data in actual daily collected data are uncertain, different cutting methods can generate different data samples. Using the same cutting length for different human activity cycles is prone to lose important features of some categories. For example, the traditional small window cutting method is used to cut the collected time series, and each small window is labeled with the category. However, since the relative position of the window cutting point is not fixed in a complete cycle, it is impossible to ensure that the accuracy remains stable under different cutting methods. For another example, the traditional large window cutting method is used to cut the collected time series. Although it can accommodate multiple activity cycles so that the impact of the cutting point on the sample is low, the large window can contain multi-category sequence information, which greatly increases the difficulty of sequence classification. Third, there can be multiple activity types in the time series in continuous time, so it is more difficult to feedback the category corresponding to the time series in real time. In addition, human daily life behaviors are not limited to two types (walking and running), which makes it impossible to meet the application requirements of different types of activities. For example, different types of activities have different calorie consumption per unit time, which greatly increases the difficulty of calculating the calorie consumption throughout the day. Fourth, using segmented and annotated public activity datasets, such as MSR Action3D, PAMAP, etc., is only applicable to laboratory environments and cannot be applied to various applications.

[0084] Based on this, in various embodiments of the present invention, original time series data are obtained; feature sequences corresponding to the original time series data are extracted; the feature sequences are mapped to a high-dimensional space to obtain feature sequences in the high-dimensional space; the feature sequences in the high-dimensional space are used as sample training data to train a preset model to obtain a prediction model for identifying actions; wherein the prediction model is used to identify actions in target time series data.

[0085] The present invention will be described in further detail below in conjunction with the accompanying drawings and embodiments.

[0086] Figure 1 is a schematic diagram of a system architecture for an application of a prediction method for action recognition provided by an embodiment of the present invention. Figure 1 As shown, the system includes:

[0087] The data transmission module is used to obtain the original time series collected by the terminal.

[0088] The feature extraction module is used to extract the feature sequence corresponding to the original time series data.

[0089] The model training module is used to map the feature sequence into a high-dimensional space to obtain a feature sequence in the high-dimensional space; the feature sequence in the high-dimensional space is used as sample training data to train a preset model to obtain a prediction model for identifying actions.

[0090] The online testing module is used to identify actions in the target time series data using the prediction model.

[0091] Figure 2 is a schematic diagram of an implementation flow of a prediction method for action recognition provided by an embodiment of the present invention, combined with Figure 1 The schematic diagram shown illustrates Figure 2 The implementation process, such as Figure 2 As shown, the method includes:

[0092] Step 201: Obtain original time series data;

[0093] Step 202: extracting a feature sequence corresponding to the original time series data;

[0094] Step 203: Mapping the feature sequence into a high-dimensional space to obtain a feature sequence in the high-dimensional space;

[0095] Step 204: using the characteristic sequence in the high-dimensional space as sample training data to train a preset model to obtain a prediction model for identifying the action;

[0096] The prediction model is used to identify actions in target time series data.

[0097] Here, in step 201, in actual application, the original time series data may refer to the data recorded by the terminal application at different times, such as WeChat, Keep, Mint Health and other applications. Specifically, the terminal can use an acceleration sensor such as a MEMS inertial sensor device to collect the three-axis acceleration data of the user's daily activities, and record it through the terminal application to form an original time series, and finally output the recorded original time series to the data transmission module in the system in real time through a transmission method such as TCP / IP, HTTP and Bluetooth.

[0098] Here, in step 202, before extracting the feature sequence corresponding to the original time series data, the original time series data may be preprocessed, such as filtering, etc. Since the original time series data is collected by using the acceleration sensor of the terminal, considering that the acceleration sensor is susceptible to noise, in order to prevent the noise deviation between the collected data and the real data from affecting the feature extraction, that is, in order to eliminate the noise deviation between the original time series and the real data, the original time series may be filtered.

[0099] Here, in step 203, before mapping the feature sequence to a high-dimensional space, the filtered original time series may be cut using overlapping sliding windows, non-overlapping sliding windows, etc. to obtain multiple one-dimensional time segments; each time segment corresponds to complete or partial information of a single action category.

[0100] Here, in step 204, in order to improve the training speed and accuracy, the high-dimensional space can be divided into multiple subspaces, and parallel recognition is performed on each of the multiple subspaces. After the action type recognition of the feature sequence of a certain subspace is completed, the feature sequence of the subspace is excluded from the total sample training data, and model training is performed based on the feature sequences of the remaining subspaces to shorten the training time and improve the training efficiency.

[0101] It should be noted that after the prediction model is used to identify the action in the target time series data, the server can interact with the terminal, that is, the server can feed back the identified action type to the terminal in real time, so as to realize end-to-end real-time feedback. In addition, the prediction model can be used to perform offline and online recognition on different types of continuous time series, wherein offline recognition refers to the recognition of the action type of time series data within a period of time, for example, the recognition of the action type of time series data recorded by the terminal application within a month. Online recognition refers to the recognition of the action type of time series data generated at the current moment, for example, the recognition of the action type of time series data recorded by the terminal application at the current moment.

[0102] The following is a detailed description of the process of how to extract the feature sequence corresponding to the original time series.

[0103] Specifically, Kalman filtering is first performed on the original time series, and then feature extraction is performed on the filtered original time series.

[0104] In practical applications, since the original time series data is collected by the accelerometer of the terminal, and the accelerometer is susceptible to noise and electromagnetic interference in the surrounding environment when collecting data, resulting in a certain gap between the collected data and the real data, in the preprocessing of the stream time series, it is necessary to denoise the collected original time series data to reduce the error caused by interference. Among them, noise can be divided into random noise with uniform frequency distribution, white noise, and frequency noise introduced by other algorithms or external devices.

[0105] Based on this, in one embodiment, extracting the feature sequence corresponding to the original time series data includes:

[0106] The original time series is subjected to Kalman filtering to obtain the original time series after filtering; and feature extraction is performed on the original time series after filtering to obtain a corresponding feature sequence.

[0107] Here, Kalman filtering refers to an optimized autoregressive data processing algorithm. This method breaks through the limitations of the classic Wiener filtering method and does not require a large amount of historical data. Therefore, it is suitable for situations where real-time feedback is required. Kalman filtering introduces a state space model, takes the minimum mean square estimate as a criterion, recursively estimates the value at the current moment, and then continues to iterate. It can not only achieve a good filtering effect, but also ensure a high data output efficiency and less calculation amount to meet the needs of real-time feedback. In the activity recognition system based on the acceleration sensor, due to the real-time requirements, the filtering algorithm should be able to occupy as little calculation amount and storage space as possible while ensuring the filtering effect.

[0108] The following is a detailed description of the process of mapping the feature sequence corresponding to the original time series into a high-dimensional space.

[0109] In practical applications, considering that in the related art, feature sequences are intercepted according to the length of time to obtain multiple time segments, due to the different lengths of time, each time segment may correspond to multiple action categories, resulting in the inability to accurately identify the action type. In this way, in the embodiment of the present invention, considering the fine-grained monitoring requirements and the transient characteristics of certain feature sequences, the switching time will be very short, and overlapping sliding windows, non-overlapping sliding windows, etc. can be used to cut the filtered feature sequence, and the action category corresponding to the feature sequence can be fed back in real time later.

[0110] In order to identify multiple time segments at the same time, multiple time segments are mapped into a high-dimensional space.

[0111] Based on this, in one embodiment, mapping the feature sequence to a high-dimensional space includes:

[0112] intercepting the feature sequence to obtain a plurality of one-dimensional feature sequences;

[0113] The multiple one-dimensional feature sequences are mapped into a high-dimensional space using a first preset parameter and a second preset parameter; wherein the first parameter represents the dimension of the high-dimensional space, and the second parameter represents the flow shape formed by the feature sequences in the high-dimensional space, that is, the time interval of the feature sequences embedded in the high-dimensional space is selected from the one-dimensional feature sequences.

[0114] Here, the process of intercepting the feature sequence may include:

[0115] Determine the length of the sliding window;

[0116] The feature sequence is truncated according to the determined length of the sliding window.

[0117] Specifically, considering that the sliding window is too long, it is necessary to use multiple labels to mark the window. In this way, the feature sequence can be analyzed to determine the length of the sliding window to ensure that the feature sequence in a sliding window only contains complete information or partial information of one action category, and should not contain relevant information of multiple action categories. For example, the length of the sliding window can be set to 100ms, which is lower than the duration of most users performing a single activity category.

[0118] For example, assuming that the original time series represents the sports activities completed by the user within 2 minutes, the length of the sliding window is determined. Assuming that the determined sliding window length is 30 seconds, the continuous feature sequence corresponding to the original time series is intercepted to obtain 4 feature sequences. The action category corresponding to the first feature sequence is running action, the action category corresponding to the second feature sequence is running action, the action category corresponding to the third feature sequence is walking action, and the action category corresponding to the fourth feature sequence is walking action.

[0119] Here, the process of mapping the multiple one-dimensional feature sequences into a high-dimensional space may include:

[0120] Using the first preset parameters, establishing a coordinate system;

[0121] According to the second preset parameter, selecting a plurality of feature sequences to be processed from the feature sequences;

[0122] Determining feature points corresponding to the plurality of feature sequences to be processed respectively in the established coordinate system to obtain a plurality of feature points;

[0123] Based on the multiple feature points, a feature sequence in a high-dimensional space is formed.

[0124] Specifically, based on the delayed embedding theory proposed by Taken in Detecting strange attractors in turbulence, the multiple one-dimensional feature sequences obtained by interception can be embedded into a high-dimensional space.

[0125] For example, d is used to represent the first preset parameter, and τ is used to represent the second preset parameter. Assuming d = 2, τ = 1, the one-dimensional feature sequence T = {y t ,y t+1 , y t+2 , …}, then a two-dimensional coordinate system is established, and {y t ,y t+1 ,y t+2 , …} as a feature sequence in high-dimensional space, as shown in Figure 3.

[0126] Here, the set of feature sequences selected from the one-dimensional feature sequence T can be expressed by formula (1), as follows:

[0127] SW d,τ (T) = {(t i ,t i+τ ,...,t i+(d-1)τ )li∈N}=M (1)

[0128] Among them, SW d,τ (T) can transform an ordered sequence of real values ​​into a set M of geometric spaces, which can be called a set of manifolds. Each element of M can be considered as a probe SW on the entire series. d,τ (T) The detection results. d and τ together determine the sensing window of the probe, where τ determines the granularity of the sensing, and d is related to the number of samples involved in the window.

[0129] Here, the process of determining the first preset parameter and the second preset parameter may include:

[0130] Determine the distribution map corresponding to the original time series;

[0131] Based on the determined distribution graph, a parameter value of the first preset parameter and a parameter value of the second preset parameter are determined.

[0132] Specifically, the first preset parameter and the second preset parameter may be determined by one of cross-validation, expert knowledge, and a control variable method.

[0133] For example, assuming that the first preset parameter is represented by d and the second preset parameter is represented by τ, Figure 4a and Figure 4b are two artificial time series, Figure 4a The artificial time series in is distributed locally, then d τ The value of is set to a smaller value, such as d = 2, τ = 2. After mapping the feature sequence corresponding to the artificial time series to the high-dimensional space, two manifolds of the same shape are obtained, that is, the artificial time series corresponds to an action category. Figure 4b The artificial time series in is distributed in a larger range, then d τ The value of is set to a larger value, such as d = 2, τ = 40. After mapping the feature sequence corresponding to the artificial time series to the high-dimensional space, two manifolds of different shapes are obtained, that is, the artificial time series corresponds to two action categories.

[0134] For another example, Figure 5a The artificial time series in is distributed locally, then d τ The value of is set to a small value, such as d = 2, τ = 40. After mapping the feature sequence corresponding to the artificial time series to the high-dimensional space, two manifolds of the same shape are obtained, that is, the artificial time series corresponds to an action category. Figure 5b The artificial time series in is distributed locally, then d τ The value of is set to a smaller value, such as d = 3, τ = 40. After mapping the feature sequence corresponding to the artificial time series to the high-dimensional space, two manifolds of different shapes are obtained, that is, the artificial time series corresponds to two action categories.

[0135] In summary, when the distribution of the original artificial time series is local, d τ The value of is set to a smaller value; when the distribution of the original artificial time series is a hash distribution in a large range, d τ to a larger value.

[0136] In one example, if Figure 6 As shown in Figure 2, the process of mapping the feature sequence corresponding to the original time series into a high-dimensional space is described, including:

[0137] Step 601: obtaining an original time series; performing filtering processing on the original time series, performing feature extraction on the filtered original time series, and obtaining a corresponding feature series.

[0138] Step 602: determine the length of the sliding window; truncate the feature sequence according to the determined length of the sliding window to obtain a truncated feature sequence.

[0139] Step 603: using the first preset parameter, establishing a coordinate system; selecting a feature sequence to be processed from the intercepted feature sequence according to the second preset parameter, to obtain a plurality of feature sequences to be processed;

[0140] Step 604: determining feature points respectively corresponding to the plurality of feature sequences to be processed in the established coordinate system to obtain a plurality of feature points; and forming a feature sequence in a high-dimensional space based on the plurality of feature points.

[0141] Here, the feature sequence corresponding to the original time series is mapped to a high-dimensional space, which has the following advantages:

[0142] (1) Based on the delayed embedding theory, it can efficiently represent the geometric space features of "massive, high-dimensional, and real-time updated" raw data, so that the global change trend and local data characteristics of the original data are retained; and it has high representation accuracy. The feature sequence data in the subsequent high-dimensional space can be more conveniently and accurately recognized by the classifier, which can improve the robustness of training and is not affected by the sequence cutting length and the first and last positions.

[0143] (2) A sliding window is used to truncate the feature sequence, that is, a cutting method is adopted that is insensitive to the cutting point and retains all the information of the sequence category as much as possible. It is not affected by the sequence cutting length and the first and last positions, and each sample sequence can be accurately labeled later.

[0144] (3) Mapping the feature sequence corresponding to the original time series into a high-dimensional space can sensitively capture the characteristics of different types of sequences. At the same time, it can consider the different forms of expression of sequences of the same category and the distinction between different categories, that is, the diversity and comprehensiveness of the data.

[0145] (4) By filtering and denoising the original time series, extracting the sequence, and performing geometric representation of the features, a data foundation is provided for building a prediction model.

[0146] (5) Preprocessing the acquired raw stream data can eliminate noise and redundancy, thereby avoiding incomplete data. At the same time, it can also promote data aggregation and standardization.

[0147] (6) Extract features (such as spatiotemporal information) based on the collected raw data, denoise the input data through appropriate filtering methods, then intercept the feature sequence, and finally convert it into a feature sequence in a high-dimensional space so that it can be subsequently input into the model to train the prediction model for action recognition.

[0148] The following is a detailed description of how to use feature sequences in high-dimensional space to train a prediction model for action recognition.

[0149] Specifically, the hyperplane of the high-dimensional space is first determined; then the determined hyperplane is used to divide the feature sequence in the high-dimensional space into multiple batches, and the feature sequence of each batch forms a feature subspace; the feature sequence of each feature subspace is trained as a prediction model respectively until the training is completed.

[0150] In practical applications, considering that a large number of training samples will result in a long training time, after mapping the feature sequence to a high-dimensional space, the sample training data can be trained using a "batch learning" method. That is, multiple hyperplanes are used to divide the feature sequence in the high-dimensional space into multiple batches. The feature sequence of each batch corresponds to a feature subspace, and the feature sequences of each subspace are used as sample training data to train the preset model. In this way, the training time can be greatly shortened without affecting the training accuracy.

[0151] Based on this, in one embodiment, the characteristic sequence in the high-dimensional space is used as sample training data to train a preset model, including:

[0152] Determine a plurality of hyperplanes by using the characteristic sequence in the high-dimensional space;

[0153] Using the multiple hyperplanes, dividing the high-dimensional space into multiple subspaces;

[0154] The feature sequences of each subspace are used as sample training data to train the preset model.

[0155] Here, determining a plurality of hyperplanes by using the characteristic sequence in the high-dimensional space may include:

[0156] Grouping the feature sequences in the high-dimensional space to obtain multiple groups of feature sequences;

[0157] For each group of feature sequences in the multiple groups of feature sequences, a hyperplane corresponding to the corresponding group of feature sequences is determined to obtain multiple hyperplanes.

[0158] Specifically, a feature sequence set is generated by using the feature sequence in the high-dimensional space; the feature sequence set is divided into a plurality of subsets; and a plurality of hyperplanes are determined by using the plurality of subsets.

[0159] In practical applications, support vector machines can be used to find the embedding space R d Divided into S - and S + The hyperplane H. Among them, the support vector machine is a model that classifies data into two categories, which is a linear classifier that maximizes the classification interval in the feature space. Its core lies in converting the problem of finding a classification hyperplane into a convex quadratic programming problem.

[0160] For example, suppose the set formed by the feature sequence in the high-dimensional space is represented by M, and the feature sequence in the high-dimensional space is divided into p groups, and the subset corresponding to each group is represented by M 1 ,M 2 ,...,M PIndicates that each subset M i is to use the hyperplane H relative to M i Determined. Taking the subset M i For example, the corresponding feature sequence in the subset is used for model training. The training result can be that the feature subspace corresponding to the subset contains two categories of feature sequences, that is, the feature sequence M formed by +1 category. i + The characteristic sequence M formed by the -1 class i - .

[0161] Figure 7 It is a schematic diagram of dividing the high-dimensional space into two subspaces, such as Figure 7 As shown in , the set formed by the feature sequence in the high-dimensional space is represented by M, and the feature sequence in the high-dimensional space is divided into 2 groups. The corresponding subsets of each group are represented by M1 and M2. M1 and M2 are determined by the hyperplane H1 relative to M. Among them, since the training samples closest to the optimal hyperplane are called support vectors, the greater the distance between the support vector and the optimal hyperplane, the higher the ability of the hyperplane to distinguish between the two categories and the classification accuracy, therefore, the maximum distance between M1 and M2 is used to determine the hyperplane H1, as shown in Figure 8 shown.

[0162] It should be noted that since batch learning allows the prediction model to be trained in parallel, it can further improve the training efficiency of the prediction model. At the same time, batch learning does not incur additional overhead for the classification time complexity due to the use of multiple hyperplanes. In addition, since the time complexity of batch learning is O(p(ml / p) 2 ), so the training efficiency of the prediction model can be greatly improved.

[0163] For example, when p = ml / 2 (i.e., batch training is performed in pairs), according to O(p(ml / p) 2 ) The time complexity of the prediction model calculated is O(2ml). Obviously, O(2ml) is linear with the original relative ml complexity O(ml 2 ), it can be seen that the linear complexity of O(2ml) is significantly lower than that of ml. For another example, given m t For a time series with a test length of l, according to the formula, the classification time complexity is O(mm t l), which is relative to the training sample size m t It is linear.

[0164] In practical applications, the “threshold learning” method can be used to identify the action category of the feature sequence in each subspace.

[0165] Based on this, in one embodiment, the feature sequences of each subspace are used as sample training data to train the preset model, which may include:

[0166] For each feature sequence of each subspace in each subspace, count the number of feature points that meet the preset conditions in the corresponding feature sequence; and determine the total number of feature points included in the corresponding feature sequence;

[0167] Calculate the ratio of the number of statistical feature points to the total number of feature points contained in the corresponding feature sequence;

[0168] When the ratio is greater than or equal to a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the first action, and is used as the training result of the preset model.

[0169] Here, when the ratio is less than a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the second action, and is used as the training result of the preset model.

[0170] Here, counting the number of feature points that meet a preset condition in the corresponding feature sequence may include:

[0171] Determine the coordinates of each feature point in the corresponding feature sequence; for each feature point in the corresponding feature sequence, determine whether the coordinates of the corresponding feature point meet the preset coordinate conditions; when the coordinates of the corresponding feature point meet the preset coordinate conditions, use the corresponding feature point as a coordinate point to be processed to obtain multiple coordinate points to be processed; and count the number of the multiple feature points to be processed.

[0172] Here, the coordinates of the corresponding feature point satisfying the preset coordinate conditions may mean that the value of each coordinate of the corresponding feature point falls within a preset value range. For example, the feature point is a two-dimensional coordinate point, i.e., an x-coordinate point and a y-coordinate point. When the x-coordinate point falls within a first preset value range and the y-coordinate point falls within a second preset value range, it indicates that the coordinates of the feature point satisfy the preset coordinate conditions.

[0173] For example, suppose that subspace 1 contains three feature sequences, represented by T1, T2 and T3 respectively, where T1 = {(0,0), (0,1), (0,2)}, T2 = {(1,0), (2,0), (3,0)}, T3 = {(1,1), (2,2), (3,3)}. Assuming that the first preset value range is [1,2] and the second preset value range is [1,3], the feature points (1,1), (2,2) and (3,3) in T3 are feature points that meet the preset coordinate conditions.

[0174] In practical applications, the “boost learning” method can be used to parallelly identify the action categories of the feature sequences in each subspace to speed up the training of the prediction model.

[0175] Based on this, in one embodiment, when performing a training of the preset model, the method further includes:

[0176] For each subspace in each subspace, determine whether the corresponding subspace has been trained to obtain the corresponding training result; when it is determined that the corresponding subspace has been trained to obtain the corresponding training result, exclude the corresponding subspace; use the feature sequences in the remaining subspaces as sample training data to train the prediction model.

[0177] Here, in actual application, in order to reduce the total amount of training data and thus improve training efficiency, during a model training, for the subspace corresponding to each batch, if the category of the feature sequence in the subspace has been identified, the sample training data corresponding to the batch can be deleted from the total training data, where the total training data can refer to the feature sequence data corresponding to the entire high-dimensional space.

[0178] Table 1 is a code diagram of the boosting learning algorithm. As shown in Table 1, line 7 describes the training results. Line 8 describes the use of the trained prediction model to delete correctly classified +1-level feature points in the total training data. Among them, the proportion of +1 or -1 feature points in each subspace remains the same as the proportion of +1 or -1 feature points in the total training data in the entire high-dimensional space.

[0179]

[0180] Table 1

[0181] The following describes in detail the process of using the prediction model to perform action recognition on the target time series with a specific embodiment.

[0182] Taking binary classification as an example, Table 1 shows the code of the algorithm for action recognition of the target time series. The specific analysis is as follows:

[0183] 1) Lines 1 to 6 describe how to map the feature sequence corresponding to the original time series data into a high-dimensional space and divide the high-dimensional space into subspaces, that is, how to perform partition learning.

[0184] First, all training time series are embedded into the same space as the set M sampled from the manifold from line 1 to line 5; then, the hyperplane H is determined to divide the embedding space into S relative to M. + and S - Among them, M i The label of each embedded point in y i Keep the same, i.e. the labels of the corresponding time series.

[0185] 2) From line 7 to line 20, it describes how to use the "threshold learning" method to identify the categories of feature sequences in each subspace.

[0186] From lines 7 to 10, calculate each sequence T i Embed S i The proportion of embedded points p i If p i ≥th, where th is the preset threshold, then the corresponding time series T i is classified as -1; otherwise, it is classified as +1.

[0187] Starting from line 10, the optimal threshold with the minimum number of misclassified time series in the training data set is searched within (0,1) to verify whether the above classification results are wrong according to the optimal threshold.

[0188] 3) From line 21 onwards, describe how to predict the unlabeled target time series.

[0189] The target time series Tx is embedded into the same space as Mx, where the training data is in line 21. Then, the proportion Px of Mx is calculated to learn the feature subspace S - According to "threshold learning", the predicted label y x .

[0190] Here, it should be noted that the target time series can be embedded into the feature subspace S corresponding to class-1 - For classification, we can also embed the target time series into the feature subspace S corresponding to class 1 + The results of the two implementations can vary depending on the problem. For example, in the problem of classifying a specific ECG dataset, - The learning seems to be better than S + ,Therefore, in practice, the two classification results can be compared, and the classification result with the smaller training error can be selected as the classification result obtained by the final classifier.

[0191]

[0192]

[0193] Table 2

[0194] In one example, if Fig. 9 As shown in FIG. 1 , the process of training the prediction model and using the trained prediction model to perform action recognition on the target time series is described, including:

[0195] Step 901: Acquire original time series data collected by the sensor.

[0196] Step 902: Perform Kalman filtering on the original time series, perform feature extraction on the filtered original time series, and obtain a corresponding feature sequence.

[0197] Step 903: determine the length of the sliding window; and intercept the feature sequence according to the determined length of the sliding window to obtain multiple one-dimensional feature sequences.

[0198] Step 904: Map the multiple one-dimensional feature sequences into a high-dimensional space to obtain a feature sequence in the high-dimensional space.

[0199] Step 905: using the characteristic sequence in the high-dimensional space to determine multiple hyperplanes; using the multiple hyperplanes to divide the high-dimensional space into multiple subspaces; using the characteristic sequence of each subspace as sample training data to train the classifier model.

[0200] Step 906: using the trained classifier model to identify the action category of the time series corresponding to the continuous action of the target human body.

[0201] Here, after using the trained classifier model to identify the action category of the time series corresponding to the target human body's continuous actions, the result of identifying the type of human body action can be fed back to the application on the terminal in real time. The application on the terminal can calculate the calories consumed by the user based on the result of identifying the type of human body action, for example, counting the calories consumed by the user per week.

[0202] Table 3 is a comparative diagram of the accuracy of identifying the action category of the target time series using the prediction model provided by the embodiment of the present invention and the accuracy of identifying the action category of the target time series using the DDE-MGM model in the related art. As shown in Table 3, the accuracy of identifying the action category of the target time series using the prediction model provided by the embodiment of the present invention is higher than the accuracy of identifying the action category of the target time series using the DDE-MGM model in the related art.

[0203]

[0204] Table 3

[0205] Here, the training of the prediction model and the use of the trained prediction model to perform action recognition on the target time series have the following advantages:

[0206] (1) Using the entire training data set and adopting batch learning methods can reduce training time. Compared with the method of selecting appropriate artificial subsets for training in related technologies, it can avoid the problem of inaccurate classifiers caused by small amounts of training data.

[0207] (2) When the training scale is very large, the size of training data can be reduced by improving the learning method. Compared with the method of using fast algorithms such as the sequential minimal optimization algorithm (SMO) for model training in related technologies, it can reduce the training cost.

[0208] (3) It can solve the following technical problems: how to achieve the goal of not significantly affecting the sequence recognition accuracy due to the length of time slice segmentation; how to select feature subspaces that can vary significantly with different activity categories; how to implement an offline and online recognition algorithm for multiple types of continuous sequences; how to make the system easy to identify continuous time series, provide a friendly interactive interface, and achieve end-to-end real-time feedback.

[0209] (4) As a discriminant model algorithm of supervised learning, the support vector machine is based on the principle of minimizing structured risk. It improves generalization ability by minimizing empirical risk and confidence range. Even when there are not many training samples, it can still achieve good classification results. In the support vector machine, the mapping can map low-dimensional linear inseparable problems to high-dimensional linear separable problems. At the same time, in order to avoid the dimensionality disaster and reduce the amount of calculation, high-dimensional discrimination is performed through the kernel function.

[0210] (5) After the time series is feature processed, a new geometric feature (which can be called a high-dimensional space) is extracted through manifold partitioning learning. Through partitioning learning and threshold learning, all training time series are embedded in the same space as the set sampled on the specific manifold and then divided to obtain multiple feature subspaces. The feature sequence of each feature subspace is used in combination with support vector machine (SVM) and boosting strategy to train the classifier model. Subsequently, the trained classifier model can be used to recognize specific action categories.

[0211] By adopting the technical solution provided in the embodiment of the present invention, the feature sequence corresponding to the original time series data is mapped to a high-dimensional space, and the feature sequence in the high-dimensional space is used as sample training data to train the preset model. In this way, the trained model can be used to identify the action category of the target time series and provide real-time feedback to the terminal.

[0212] In order to implement the prediction method of action recognition of the embodiment of the present invention, the embodiment of the present invention also provides a prediction device for action recognition. Fig.10 FIG. 1 is a schematic diagram of the structure of a prediction device for action recognition according to an embodiment of the present invention; Fig.10 As shown, the device comprises:

[0213] An acquisition unit 101 is used to acquire original time series data;

[0214] The first processing unit 102 is used to extract the feature sequence corresponding to the original time series data; map the feature sequence to a high-dimensional space to obtain the feature sequence in the high-dimensional space;

[0215] The second processing unit 103 is used to train a preset model using the characteristic sequence in the high-dimensional space as sample training data to obtain a prediction model for identifying the action;

[0216] The prediction model is used to identify actions in target time series data.

[0217] In one embodiment, the first processing unit 102 is specifically configured to:

[0218] intercepting the feature sequence to obtain a plurality of one-dimensional feature sequences;

[0219] The multiple one-dimensional feature sequences are mapped into a high-dimensional space using a first preset parameter and a second preset parameter; the first parameter represents the dimension of the high-dimensional space, and the second parameter represents the flow shape formed by the feature sequences in the high-dimensional space.

[0220] In one embodiment, the second processing unit 103 is specifically configured to:

[0221] Determine a plurality of hyperplanes by using the characteristic sequence in the high-dimensional space;

[0222] Using the multiple hyperplanes, dividing the high-dimensional space into multiple subspaces;

[0223] The feature sequences of each subspace are used as sample training data to train the preset model.

[0224] In one embodiment, the second processing unit 103 is specifically configured to:

[0225] Grouping the feature sequences in the high-dimensional space to obtain multiple groups of feature sequences;

[0226] For each group of feature sequences in the multiple groups of feature sequences, a hyperplane corresponding to the corresponding group of feature sequences is determined to obtain multiple hyperplanes.

[0227] In one embodiment, the second processing unit 103 is specifically used to:

[0228] For each feature sequence of each subspace in each subspace, count the number of feature points that meet the preset conditions in the corresponding feature sequence; and determine the total number of feature points included in the corresponding feature sequence;

[0229] Calculate the ratio of the number of statistical feature points to the total number of feature points contained in the corresponding feature sequence;

[0230] When the ratio is greater than or equal to a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the first action, and is used as the training result of the preset model.

[0231] In one embodiment, the second processing unit 103 is further configured to:

[0232] When the ratio is less than a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the second action, and is used as the training result of the preset model.

[0233] In one embodiment, the second processing unit 103 is specifically configured to:

[0234] Determine the coordinates of each feature point in the corresponding feature sequence;

[0235] For each feature point in the corresponding feature sequence, determine whether the coordinates of the corresponding feature point meet the preset coordinate conditions;

[0236] When the coordinates of the corresponding feature points meet the preset coordinate conditions, the corresponding feature points are used as coordinate points to be processed, and a plurality of coordinate points to be processed are obtained;

[0237] Count the number of the plurality of feature points to be processed.

[0238] In one embodiment, the second processing unit 103 is specifically configured to:

[0239] When performing a training of the preset model, for each subspace in each subspace, it is determined whether the corresponding subspace has been trained to obtain the corresponding training result;

[0240] When it is determined that the corresponding subspace has been trained to obtain the corresponding training result, the corresponding subspace is excluded;

[0241] The feature sequences in the remaining subspace are used as sample training data to train the prediction model.

[0242] In actual application, the acquisition unit 101 can be implemented by a communication interface in the motion recognition prediction device; the first processing unit 102 and the second processing unit 103 can be implemented by a processor in the motion recognition prediction device.

[0243] It should be noted that: the prediction device for action recognition provided in the above embodiment only uses the division of the above program modules as an example when performing prediction based on transmission power. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the prediction device for action recognition provided in the above embodiment and the prediction method embodiment based on transmission power belong to the same concept. The specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0244] The embodiment of the present invention also provides a server, such as Fig.11 As shown, including:

[0245] Communication interface 111, capable of exchanging information with other devices;

[0246] The processor 112 is connected to the communication interface 111 and is used to execute the method provided by one or more technical solutions on the smart device side when running the computer program. The computer program is stored in the memory 113.

[0247] It should be noted that the specific processing process of the processor 112 and the communication interface 111 is detailed in the method embodiment and will not be repeated here.

[0248] Of course, in actual application, the various components in the server 110 are coupled together through the bus system 114. It can be understood that the bus system 114 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 114 also includes a power bus, a prediction bus for action recognition, and a status signal bus. However, for the sake of clarity, Fig.11 Various buses are labeled as bus system 114 .

[0249] The memory 113 in the embodiment of the present application is used to store various types of data to support the operation of the server 110. Examples of such data include: any computer program used to operate on the server 110.

[0250] The method disclosed in the above embodiment of the present application can be applied to the processor 112, or implemented by the processor 112. The processor 112 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 112 or an instruction in the form of software. The above-mentioned processor 112 may be a general-purpose processor, a digital data processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 112 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 113, and the processor 112 reads the information in the memory 113 and completes the steps of the above method in combination with its hardware.

[0251] In an exemplary embodiment, the server 110 may be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general-purpose processor, motion recognition predictor, micro-motion recognition predictor (MCU), micro controller unit (Microprocessor), or other electronic components to execute the aforementioned method.

[0252] It can be understood that the memory (memory 113) of the embodiment of the present application can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a ferromagnetic random access memory, a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.

[0253] In an exemplary embodiment, the embodiment of the present invention further provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a memory 113 storing a computer program, and the computer program can be executed by the processor 112 of the server 110 to complete the steps of the aforementioned action recognition prediction server-side method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.

[0254] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0255] In addition, the technical solutions described in the embodiments of the present invention can be arbitrarily combined without conflict.

[0256] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.

Claims

1. A prediction method for action recognition, characterized in that: The method comprises: Get the original time series data; Extracting a feature sequence corresponding to the original time series data; Mapping the feature sequence into a high-dimensional space to obtain a feature sequence in the high-dimensional space; Using the feature sequence in the high-dimensional space as sample training data to train the model, and obtaining a prediction model for identifying the action; Wherein, the prediction model is used to identify actions in the target time series data; Mapping the feature sequence to a high-dimensional space comprises: intercepting the feature sequence to obtain a plurality of one-dimensional feature sequences; Mapping the multiple one-dimensional feature sequences into a high-dimensional space using a first preset parameter and a second preset parameter; the first preset parameter represents the dimension of the high-dimensional space, and the second preset parameter represents the flow shape formed by the feature sequences in the high-dimensional space; The method of using the feature sequence in the high-dimensional space as sample training data to train the model includes: Determine a plurality of hyperplanes by using the feature sequence in the high-dimensional space; Using the multiple hyperplanes, dividing the high-dimensional space into multiple subspaces; The feature sequences of each subspace are used as sample training data to train the model.

2. The method according to claim 1, characterized in that The method of using the feature sequence in the high-dimensional space to determine a plurality of hyperplanes includes: Grouping the feature sequences in the high-dimensional space to obtain multiple groups of feature sequences; For each group of feature sequences in the multiple groups of feature sequences, a hyperplane corresponding to the corresponding group of feature sequences is determined to obtain multiple hyperplanes.

3. The method according to claim 1, characterized in that The feature sequences of each subspace are used as sample training data to train the model, including: For each feature sequence of each subspace in each subspace, count the number of feature points that meet the preset conditions in the corresponding feature sequence; and determine the total number of feature points included in the corresponding feature sequence; Calculate the ratio of the number of statistical feature points to the total number of feature points contained in the corresponding feature sequence; When the ratio is greater than or equal to a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the first action, and is used as the training result of the model.

4. The method according to claim 3, characterized in that The method further comprises: When the ratio is less than a preset threshold, the action type corresponding to the corresponding feature sequence is determined to be the second action, and is used as the training result of the model.

5. The method according to claim 3, characterized in that: The counting of the number of feature points in the corresponding feature sequence that meet the preset conditions includes: Determine the coordinates of each feature point in the corresponding feature sequence; For each feature point in the corresponding feature sequence, determine whether the coordinates of the corresponding feature point meet the preset coordinate conditions; When the coordinates of the corresponding feature points meet the preset coordinate conditions, the corresponding feature points are used as coordinate points to be processed, and a plurality of coordinate points to be processed are obtained; Count the number of the plurality of feature points to be processed.

6. The method according to claim 3, characterized in that When performing a training of the model, the method further includes: For each subspace in each subspace, determine whether the corresponding subspace has been trained to obtain the corresponding training result; When it is determined that the corresponding subspace has been trained to obtain the corresponding training result, the corresponding subspace is excluded; The feature sequences in the remaining subspace are used as sample training data to train the prediction model.

7. A prediction device for action recognition, characterized in that: include: An acquisition unit, used to acquire original time series data; A first processing unit, used for extracting a feature sequence corresponding to the original time series data; Mapping the feature sequence into a high-dimensional space to obtain a feature sequence in the high-dimensional space; A second processing unit is used to train a model by using the feature sequence in the high-dimensional space as sample training data to obtain a prediction model for identifying the action; Wherein, the prediction model is used to identify actions in the target time series data; The first processing unit is specifically configured to: The feature sequence is intercepted to obtain a plurality of one-dimensional feature sequences; the plurality of one-dimensional feature sequences are mapped to a high-dimensional space using a first preset parameter and a second preset parameter; the first preset parameter represents the dimension of the high-dimensional space, and the second preset parameter represents the flow shape formed by the feature sequence in the high-dimensional space; The second processing unit is specifically configured to: Using the feature sequence in the high-dimensional space, multiple hyperplanes are determined; using the multiple hyperplanes, the high-dimensional space is divided into multiple subspaces; using the feature sequence of each subspace as sample training data to train the model.

8. A server, characterized in that: include: Communication interface, used to obtain raw time series data; A processor is used to extract a feature sequence corresponding to the original time series data; map the feature sequence to a high-dimensional space to obtain a feature sequence in the high-dimensional space; and use the feature sequence in the high-dimensional space as sample training data to train a model to obtain a prediction model for identifying actions; wherein the prediction model is used to identify actions in the target time series data; The processor is specifically used for: The feature sequence is intercepted to obtain a plurality of one-dimensional feature sequences; the plurality of one-dimensional feature sequences are mapped to a high-dimensional space using a first preset parameter and a second preset parameter; the first preset parameter represents the dimension of the high-dimensional space, and the second preset parameter represents the flow shape formed by the feature sequence in the high-dimensional space; Using the feature sequence in the high-dimensional space, multiple hyperplanes are determined; using the multiple hyperplanes, the high-dimensional space is divided into multiple subspaces; using the feature sequence of each subspace as sample training data to train the model.

9. A server, characterized in that: comprising a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the steps of the method described in any one of claims 1 to 6.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • IMU action recognition based on machine learning

    AU2019101148A4

  • Time sequence anomaly detection method and device, server and storage medium

    CN110276409A