Motion recognition model construction method and device, computer equipment and storage medium

By performing tensor decomposition and pooling on the weight matrix of the recursive neural network model and combining it with attention weight adjustment, an action recognition model is constructed, which solves the problem of large number of parameters and achieves efficient action recognition.

CN120708025APending Publication Date: 2025-09-26HAINAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510789477.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing action recognition models have a large number of parameters, which makes the models difficult to train and takes up huge computing resources.

Method used

By performing tensor decomposition on the first and second weight matrices in the recursive neural network model, the first tensor chain and the second tensor chain are obtained. The training tensor samples and hidden states are transformed through these tensor chains. Combined with pooling processing and attention weight adjustment, an action recognition model is constructed.

Benefits of technology

It reduces the number of model parameters, improves training efficiency and recognition efficiency, reduces computing resource usage, and improves recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708025A_ABST
    Figure CN120708025A_ABST
Patent Text Reader

Abstract

The invention discloses an action recognition model construction method and device, computer equipment and a storage medium, and relates to the field of neural networks. The method comprises the following steps: determining a first weight matrix, a second weight matrix and a hidden state of a first to-be-trained model, and respectively performing tensor decomposition on the first weight matrix and the second weight matrix to obtain a first tensor chain and a second tensor chain; and obtaining a target hidden state according to the first tensor chain, the second tensor chain, the hidden state and the training tensor sample, performing pooling processing on the target hidden state to obtain a pooling result, and adjusting the first to-be-trained model according to the pooling result and the label of the training tensor sample to obtain a first target model. According to the method, the parameter quantity is reduced, so that the computing resource occupation of the first target model during use is reduced, and the recognition efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of neural network technology, and in particular to a method, device, computer equipment and storage medium for constructing an action recognition model. Background Art

[0002] In contemporary society, human-computer interaction has become an integral part of our work and lives. With the continuous development of artificial intelligence, particularly deep learning, motion recognition in human-computer interaction is becoming increasingly intelligent and efficient. Motion recognition enables human body detection and attribute recognition, gesture recognition, and crowd counting.

[0003] Existing action recognition models are usually built using neural networks. The neural networks are trained using action videos or radar data to obtain action recognition models. However, existing model training has a large number of parameters, which makes the model difficult to train and consumes huge computing resources. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, device, computer equipment and storage medium for constructing an action recognition model to address the above technical problems, so as to solve the problem that the existing technology has a large number of parameters, which makes the model difficult to train and occupies huge computing resources.

[0005] In a first aspect, the present application provides a method for constructing an action recognition model. The method comprises:

[0006] Determine a first weight matrix, a second weight matrix, and a hidden state of a first to-be-trained model according to the training tensor sample;

[0007] Performing tensor decomposition on the first weight matrix and the second weight matrix respectively to obtain a first tensor chain and a second tensor chain;

[0008] Obtain the target hidden state according to the first tensor chain, the second tensor chain, the hidden state, and the training tensor sample;

[0009] Perform pooling on the target hidden state to obtain the pooling result;

[0010] Adjust the first to-be-trained model according to the pooling result and the label of the training tensor sample to obtain a first target model;

[0011] The training tensor samples include high-order tensor data arranged in time series, and the high-order tensor data includes frame data corresponding to preset actions.

[0012] In one embodiment, obtaining a target hidden state based on the first tensor chain, the second tensor chain, the hidden state, and the training tensor sample includes:

[0013] Obtain first transformed data according to the first tensor chain and the training tensor sample;

[0014] Obtain second transformed data according to the second tensor chain and the hidden state;

[0015] A target hidden state is obtained according to the first transformed data and the second transformed data.

[0016] In one embodiment, the step of obtaining training tensor samples includes:

[0017] Acquire video data of a preset action, frame data in the video data, and a time sequence of the frame data;

[0018] Determine width data, height data, and channel information of the frame data according to the frame data;

[0019] Construct high-order tensor data based on width data, height data, and channel information;

[0020] Sort the high-order tensor data according to the time series to obtain training tensor samples.

[0021] In one embodiment, determining a first weight matrix, a second weight matrix, and a hidden state of the first to-be-trained model based on the training tensor sample and the first to-be-trained model includes:

[0022] Determining a first weight matrix and a second weight matrix of a first to-be-trained model;

[0023] Determine a hidden state according to the first weight matrix and the second weight matrix of the first to-be-trained model.

[0024] In one embodiment, obtaining a target hidden state based on the first tensor chain, the second tensor chain, the hidden state, and the training tensor sample includes:

[0025] Obtaining first transformed data according to the first tensor chain and the high-order tensor data at the first target moment;

[0026] Obtaining second transformed data according to the second tensor chain and the hidden state of the second target moment, wherein the first target moment is a next target moment of the second target moment;

[0027] A target hidden state is obtained according to the first transformed data and the second transformed data.

[0028] In one embodiment, after adjusting the first to-be-trained model based on the pooling results and the labels of the training tensor samples to obtain the first target model, the method further includes:

[0029] Training the second model to be trained according to the training tensor sample to obtain the attention weight of the second model to be trained;

[0030] According to the pooling results and attention weights, the enhanced feature output is obtained;

[0031] According to the enhanced feature output and the preset action, the first model to be trained and the second model to be trained are adjusted to obtain the target model.

[0032] In one embodiment, training the second model to be trained based on the training tensor sample to obtain the attention weight of the second model to be trained includes:

[0033] Perform tensor locality-preserving projection on the training tensor samples to obtain a projection matrix corresponding to the dimension of the high-order tensor data;

[0034] According to the training tensor sample and the projection matrix, a low-dimensional embedding tensor of the training tensor sample is obtained;

[0035] The second model to be trained is trained according to the low-dimensional embedding tensor to obtain the attention weight of the second model to be trained.

[0036] In one embodiment, a tensor locality preserving projection is performed on the training tensor samples to obtain a projection matrix corresponding to the dimension of the high-order tensor data, including:

[0037] According to the training tensor samples, a neighbor graph corresponding to the training tensor samples is obtained;

[0038] According to the nearest neighbor graph, the correlation matrix is ​​obtained, wherein the correlation matrix is ​​obtained by the heat kernel method according to the nearest neighbor graph;

[0039] According to the correlation matrix, eigendecomposition is performed to obtain the projection matrix corresponding to the dimension of the high-order tensor data.

[0040] In a second aspect, the present application also provides a device for constructing an action recognition model. The device includes:

[0041] A determination module, configured to determine a first weight matrix, a second weight matrix, and a hidden state of a first to-be-trained model based on the training tensor sample;

[0042] a decomposition module, configured to perform tensor decomposition on the first weight matrix and the second weight matrix respectively to obtain a first tensor chain and a second tensor chain;

[0043] A get module, configured to get a target hidden state based on the first tensor chain, the second tensor chain, the hidden state, and the training tensor sample;

[0044] The pooling module is used to perform pooling processing on the target hidden state to obtain the pooling result;

[0045] An adjustment module, configured to adjust the first to-be-trained model according to the pooling result and the label of the training tensor sample to obtain a first target model;

[0046] The training tensor samples include high-order tensor data arranged in time series, and the high-order tensor data includes frame data corresponding to preset actions.

[0047] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned method for constructing an action recognition model when executing the computer program.

[0048] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-mentioned method for constructing an action recognition model.

[0049] In a fifth aspect, the present application further provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps in the above-mentioned method for constructing an action recognition model.

[0050] The above-mentioned action recognition model construction method, device, computer equipment and storage medium construct training tensor samples from tensor data, so that the samples contain more information and improve the training effect. By performing tensor decomposition on the first weight matrix and the second weight matrix in the first to-be-trained model, the first tensor chain and the second tensor chain are obtained and replaced by the first weight matrix and the second weight matrix, the number of parameters of the first to-be-trained model is reduced, the training efficiency of the first to-be-trained model is improved, and the number of parameters of the obtained first target model is reduced, thereby reducing the computing resource usage and improving the recognition efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A schematic diagram of a scenario for constructing an action recognition model provided in one embodiment;

[0052] Figure 2 A schematic diagram of a method for building an action recognition model in one embodiment Figure 1 ;

[0053] Figure 3 A schematic diagram of a method for building an action recognition model in one embodiment Figure 2 ;

[0054] Figure 4 The figure is a structural block diagram of a device for building a motion recognition model in one embodiment. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0056] The present application embodiment provides a method for constructing an action recognition model, such as Figure 1 As shown, the following steps are included:

[0057] In order to clearly understand the technical solution of the present application, the solution of the prior art is first introduced in detail.

[0058] In contemporary society, human-computer interaction has gradually become an integral part of our work and lives. With the continuous development of artificial intelligence, particularly deep learning, human-computer interaction is becoming increasingly intelligent and efficient. In recent years, thanks to collaboration and research between scholars from diverse fields, a variety of human-computer interaction methods have been developed to facilitate our lives and work. As an important form of body language, gestures are an integral part of our lives. For example, traffic police use gestures to direct traffic, and the deaf and mute use sign language to communicate with others. Consequently, gesture interaction has garnered increasing attention from both academia and industry, becoming a key method in the field of human-computer interaction.

[0059] Existing gesture recognition models are typically built using neural networks, which are trained using gesture videos or radar data to obtain gesture recognition models. However, existing model training has a large number of parameters, making the model difficult to train and occupying huge computing resources.

[0060] In response to the above problems, the inventors discovered in their research that the first weight matrix used to transform the training tensor samples and the second weight matrix used to transform the hidden state in the existing recursive neural network model can be determined, and the first tensor chain and the second tensor chain can be obtained by performing tensor decomposition on the first weight matrix and the second weight matrix. The training tensor samples are then transformed by the first tensor chain, and the hidden state is transformed by the second tensor chain to obtain the target hidden state at the last moment. The pooling result is obtained by pooling the target hidden state. Finally, a correspondence is established between the pooling result and the label of the training tensor sample, and the first model to be trained is adjusted to obtain the first target model, thereby achieving the purpose of reducing the number of parameters.

[0061] The following introduces the application scenarios of the action recognition model construction method provided in the embodiment of the present application.

[0062] Figure 1 A schematic diagram of a scene for constructing an action recognition model provided in an embodiment of the present application, such as Figure 1As shown, the scene includes video data and a computer. The video data is video data of a preset action. The video data is converted into training tensor samples by the computer, and then the first model to be trained is trained by the training tensor samples to obtain a first target model.

[0063] The preset action may refer to a pre-set action that needs to be recognized by the model, such as a hand gesture, a human posture, an animal posture, or an expression.

[0064] The video data may refer to a video file of a preset action shot by a camera, and the video data includes frame data sorted in a time series.

[0065] Frame data may refer to image data of each frame in the video data, and each image data includes width data, height data, and channel information.

[0066] Width data can refer to the width of the image, and height data can refer to the height of the image. For example, if the image size is 10 cm * 20 cm, it can be understood that the width is 10 cm and the height is 20 cm. Channel information can refer to the data of the three color channels of red (R), green (G), and blue (B) of the image, which is used to determine the color of the image.

[0067] A tensor can refer to a multidimensional array, which is a generalization of vectors and matrices. In machine learning and deep learning, tensors are the basic structure for representing data. They can be scalars (0-dimensional tensors), vectors (1-dimensional tensors), matrices (2-dimensional tensors), or higher-dimensional arrays.

[0068] High-order tensor data may refer to a high-order array, which is data that extends the concepts of vectors and matrices to multi-order and multi-dimensional spaces and can retain more data information. In this embodiment, high-order tensor data may refer to a third-order tensor constructed based on width data, height data, and channel information in the video data.

[0069] The Long-Short Term Memory (LSTM) neural network model is a special type of recurrent neural network (RNN) designed for processing and predicting time series data. Compared to standard RNNs, LSTMs have a stronger memory capacity and are better able to handle long-term dependencies. They use special structural units, including input gates, forget gates, and output gates, to control the flow of information and memory updates. This enables LSTMs to more effectively learn and remember long-term dependencies, resulting in better performance when processing time series data.

[0070] Figure 2A flow chart of a method for constructing an action recognition model provided in an embodiment of the present application is shown as follows: Figure 2 As shown, the method includes:

[0071] S201: Determine a first weight matrix, a second weight matrix, and a hidden state of a first to-be-trained model.

[0072] Among them, the first model to be trained can refer to a neural network model. For example, in an embodiment of the present application, the neural network model can be a long short-term memory neural network model.

[0073] The weight matrix may refer to a matrix composed of weight parameters of each neuron connection in a neural network model. In an embodiment of the present application, when the first model to be trained is a long short-term memory neural network model, the first weight matrix may refer to a weight matrix composed of parameters between the input layer and the hidden layer, and the second weight matrix may refer to a weight matrix composed of parameters between the hidden layer and the output layer. The first weight matrix may include a first weight matrix of the input gate, a first weight matrix of the forget gate, a first weight matrix of the output gate, and a first weight matrix of the candidate cell state. The second weight matrix may include a second weight matrix of the input gate, a second weight matrix of the forget gate, a second weight matrix of the output gate, and a second weight matrix of the candidate cell state.

[0074] The hidden state refers to the state within the network at each time step. It contains information about the network's input at the current time step and the hidden state of the previous time step. The hidden state can be seen as an internal representation of the network's past information. It is continuously updated and passed during training to capture the temporal correlation and contextual information in the input sequence.

[0075] In an embodiment of the present application, the training process of the first model to be trained can be: at the beginning of training, the initial first weight matrix, the initial second weight matrix, and the initial hidden state are initialized, and the high-order tensor data of the first moment in the training tensor sample is input into the first model to be trained to obtain the updated first weight matrix, the updated second weight matrix, and the hidden state at the first moment. When the high-order tensor data of the second moment is input, the updated first weight matrix of the first moment is multiplied by the high-order tensor data of the second moment, and the updated second weight matrix of the first moment is multiplied by the hidden state of the first moment to obtain the updated first weight matrix, the updated second weight matrix, and the hidden state at the second moment. After the high-order tensor data of the last moment is input, the updated first weight matrix, the updated second weight matrix, and the hidden state at the last moment are obtained. At this point, the training of the first model to be trained is completed. Wherein, the first weight matrix is ​​the updated first weight matrix of the last moment, the second weight matrix is ​​the updated second weight matrix of the last moment, and the hidden state includes the hidden state of each moment.

[0076] In the embodiment of the present application, determining the first weight matrix, the second weight matrix, and the hidden state of the first to-be-trained model includes:

[0077] Get the training tensor sample corresponding to the preset action;

[0078] According to the training tensor sample and the first model to be trained, a first weight matrix, a second weight matrix, and a hidden state of the first model to be trained are determined.

[0079] In this embodiment of the present application, obtaining a training tensor sample corresponding to a preset action includes:

[0080] Acquire video data of a preset action, frame data in the video data, and a time sequence of the frame data;

[0081] Determine width data, height data, and channel information of the frame data according to the frame data;

[0082] Construct high-order tensor data based on width data, height data, and channel information;

[0083] Sort the high-order tensor data according to the time series to obtain training tensor samples.

[0084] Among them, the method for obtaining training tensor samples may include: dividing the video data into frame data in time series by a computer, then performing data recognition on each frame data, obtaining width data, height data, and channel data in each frame data, and finally converting the width data, height data, and channel data into high-order tensor data to retain the complete information in the frame data. In some embodiments, when processing an image of size 10 cm * 20 cm, the width and height of the image can be first converted into pixel representation, thereby making it correspond to the representation method of the channel information, and then determining the color brightness representation of the red channel, blue channel, and yellow channel on each pixel, so as to obtain channel information, and finally representing the width, height and channel information in the above information with tensors to obtain third-order tensor data.

[0085] In the embodiment of the present application, determining the first weight matrix, the second weight matrix, and the hidden state of the first to-be-trained model according to the training tensor sample and the first to-be-trained model includes:

[0086] Determining a first weight matrix and a second weight matrix of a first to-be-trained model;

[0087] Determine a hidden state according to a first weight matrix and a second weight matrix of the first to-be-trained model, wherein the hidden state satisfies:

[0088]

[0089] Among them, σ is the sigmoid function, tanh is the hyperbolic tangent function, W f is the first weight matrix of the input gate in the first model to be trained, U f is the second weight matrix of the input gate in the first model to be trained, W i is the first weight matrix of the forget gate in the first model to be trained, U i is the second weight matrix of the forget gate in the first model to be trained, W o is the first weight matrix of the output gate in the first model to be trained, U o is the second weight matrix of the output gate in the first model to be trained, W c is the first weight matrix in the candidate cell state in the first model to be trained, U c is the second weight matrix in the candidate cell state in the first model to be trained, x t is the target high-order tensor data at time t in the training tensor sample, b f 、b i 、b o and b c is the bias value, h t-1 is the hidden state at time t-1, h t is the hidden state at time t.

[0090] Among them, ⊙ can refer to the element-based product operation, c t Can be cell state, used to convey information in a long-term context.

[0091] The input gate may refer to a network layer in the first model to be trained, which is used to add new training tensor samples to the hidden state, thereby updating the hidden state until all training tensor samples have been input.

[0092] The forget gate may refer to a network layer in the first model to be trained, which is used to determine whether the information in the memory cells needs to be forgotten.

[0093] The output gate may refer to a network layer in the first model to be trained, which is used to decide whether to use the information of the memory cell as the current output.

[0094] The control of the input gate, forget gate, and output gate can effectively capture important long-term dependencies in the training tensor sample sequence and solve the gradient problem.

[0095] A candidate cell state may refer to a candidate value of a cell state.

[0096] The target high-order tensor data may refer to the high-order tensor data that needs to be input and is currently selected from the training tensor samples.

[0097] In this embodiment, training tensor samples are constructed through tensor data so that the samples contain more information, thereby improving the training effect.

[0098] S202 , performing tensor decomposition on the first weight matrix and the second weight matrix respectively to obtain a first tensor chain and a second tensor chain.

[0099] Among them, tensor decomposition can refer to the process of decomposing a high-order tensor into the product of multiple low-order tensors. In an embodiment of the present application, tensor decomposition of the first weight matrix can obtain a first tensor chain represented by a series of lightweight tensor chains, and tensor decomposition of the second weight matrix can obtain a second tensor chain represented by a series of lightweight tensor chains.

[0100] For example, assuming that the high-order tensor data is an N-order tensor The output hidden state is a D-order tensor Then the first tensor chain and the second tensor chain can be expressed as:

[0101]

[0102] Among them, QTT(W l ) is the first tensor chain, W l is the first weight matrix, QTT(U l ) is the second tensor chain, U l is the second weight matrix, is the strong Kronecker product.

[0103] In this embodiment, by decomposing the first weight matrix and the second weight matrix into a first tensor chain and a second tensor chain respectively through tensor decomposition, the number of parameters of the first target model can be reduced and the recognition efficiency of the model can be increased.

[0104] S203. Obtain a target hidden state based on the first tensor chain, the second tensor chain, the hidden state, and the training tensor sample, wherein the training tensor sample includes high-order tensor data arranged in a time series, the high-order tensor data includes frame data corresponding to a preset action, and the target hidden state is obtained based on the first transformation data and the second transformation data, the first transformation data is obtained based on the first tensor chain and the training tensor sample, and the second transformation data is obtained based on the second tensor chain and the hidden state.

[0105] Among them, the first tensor chain is used to transform the training tensor samples, and the second tensor chain is used to transform the hidden state. The first weight matrix is ​​replaced by the first tensor chain, and the second weight matrix is ​​replaced by the second tensor chain, and then the target hidden state is obtained by the formula satisfied by the target hidden state.

[0106] Transformed data may refer to data that is pre-processed or converted to better suit the needs of the model, wherein raw data may refer to data that has not been processed or converted, that is, the initial data obtained from the data source. In an embodiment of the present application, raw data may refer to training tensor samples and hidden states. The first transformed data may be data after the training tensor samples are transformed according to the first tensor chain, and the second transformed data may be data after the hidden state is transformed according to the second tensor chain. For example, at the current moment, the first transformed data may refer to the data after the training tensor samples at the current moment are transformed with the first tensor chain, and the second transformed data may refer to the data after the hidden state at the previous moment is transformed with the second tensor chain.

[0107] The target hidden state can refer to the hidden state that the model needs to learn and predict during the training process. For example, in time series data, the target hidden state can be the state of the next time step that needs to be predicted.

[0108] In the embodiment of the present application, the target hidden state is obtained according to the first tensor chain, the second tensor chain, the hidden state and the training tensor sample, including:

[0109] Obtaining first transformed data according to the first tensor chain and the high-order tensor data at the first target moment;

[0110] Obtaining second transformed data according to the second tensor chain and the hidden state of the second target moment, wherein the first target moment is a next target moment of the second target moment;

[0111] A target hidden state is obtained according to the first transformed data and the second transformed data.

[0112] The first transformed data may be obtained by multiplying the first tensor chain by the training tensor sample, and the second transformed data may be obtained by multiplying the second tensor chain by the hidden state.

[0113] The first target time can refer to any time of the training tensor sample, such as the time corresponding to the last frame of the video file. The second target time is the time before the first target time.

[0114] When the first weight matrix and the second weight matrix are obtained, the hidden state at each moment is also obtained. Since the first weight matrix is ​​updated to the first tensor chain and the second weight matrix is ​​updated to the second tensor chain, the updated model output is obtained, thereby obtaining the target hidden state.

[0115] In this embodiment of the present application, the target hidden state satisfies:

[0116]

[0117] Among them, QTTL (Wi , x t ) is the first transformation data of the input gate in the first model to be trained, QTTL(W f , x t ) is the first transformation data of the forget gate in the first model to be trained, QTTL(W o , x t ) is the first transformation data of the output gate in the first model to be trained, QTTL(W c , x t ) is the first transformation data in the candidate cell state in the first to-be-trained model, QTTL(U i , h t-1 ) is the second transformation data of the input gate in the first to-be-trained model, QTTL(U f , h t-1 ) is the second transformation data of the forget gate in the first to-be-trained model, QTTL(U o , h t-1 ) is the second transformed data of the output gate in the first to-be-trained model and QTTL(U c , h t-1 ) is the second transformed data of the candidate cell state in the first model to be trained, and t is the first target time.

[0118] S204: Perform pooling processing on the target hidden state to obtain a pooling result.

[0119] Among them, pooling processing can refer to reducing the target hidden state, achieving downsampling, while retaining the main information in the target hidden state to prevent overfitting. For example, when identifying whether an image is a face, we need to know that there is an eye on the left and an eye on the right of the face, without knowing the exact location of the eyes. Therefore, we can obtain the overall statistical characteristics by pooling the pixels of a certain area, highlighting the recognition of the eyes.

[0120] S205: Adjust the first model to be trained according to the pooling result and the label of the training tensor sample to obtain a first target model.

[0121] The label of the training tensor sample may refer to a preset action corresponding to the training tensor sample.

[0122] Adjusting the first model to be trained may refer to establishing a correspondence between the pooling result and the preset action by processing the pooling result and comparing the processed result with the preset action.

[0123] When the first training model is trained using multiple preset actions, the obtained first target model can recognize multiple preset actions. By inputting the video of the action to be recognized into the first target model, the first target model outputs the recognition pooling result, and the recognition pooling result is processed and mapped to each preset action to obtain the probability of being similar to each preset action. The preset action with the highest probability is selected as the recognition result of the first target model.

[0124] In the embodiment of the present application, after adjusting the first to-be-trained model according to the pooling result and the label of the training tensor sample to obtain the first target model, the method further includes:

[0125] Training the second model to be trained according to the training tensor sample to obtain the attention weight of the second model to be trained;

[0126] According to the pooling results and attention weights, the enhanced feature output is obtained;

[0127] According to the enhanced feature output and the preset action, the first model to be trained and the second model to be trained are adjusted to obtain the target model.

[0128] Among them, the second model to be trained can refer to a computer model that simulates the human attention mechanism, which is used to simulate the selective process of humans paying attention to and focusing on different parts of information processing. In the second model to be trained, after the training tensor samples are encoded, a set of attention weights are calculated to determine the degree of attention to different parts of the input. These weights represent the importance of each input part in the output.

[0129] The attention weight may refer to the attention weight corresponding to the high-order tensor data output by the second model to be trained. Since the number of parameters of the first target model is reduced by replacing the first weight matrix with the first tensor chain and the second weight matrix with the second tensor chain, although the model calculation efficiency is improved, the output accuracy of the first target model is reduced. Therefore, the output accuracy of the first target model is improved by obtaining the attention weight of the training tensor sample.

[0130] The enhanced feature output can be calculated by combining the pooling result and the attention weight.

[0131] Adjusting the first model to be trained and the second model to be trained may refer to adjusting the parameters in the first model to be trained and the second model to be trained according to the expected corresponding effect of the enhanced feature output and the preset action, thereby improving the recognition accuracy of the first model to be trained and the second model to be trained, and thus obtaining the target model.

[0132] In this embodiment of the present application, the second model to be trained is trained according to the training tensor sample to obtain the attention weight of the second model to be trained, including:

[0133] Perform tensor locality-preserving projection on the training tensor samples to obtain a projection matrix corresponding to the dimension of the high-order tensor data;

[0134] According to the training tensor sample and the projection matrix, a low-dimensional embedding tensor of the training tensor sample is obtained;

[0135] The second model to be trained is trained according to the low-dimensional embedding tensor to obtain the attention weight of the second model to be trained.

[0136] Among them, tensor local preservation projection can refer to the extraction and dimensionality reduction of high-order tensor data in training tensor samples, which can effectively retain the information in high-order tensor data to provide support for subsequent work.

[0137] The projection matrix may refer to the result of tensor locality-preserving projection of the high-order tensor data in the training tensor sample. The number of projection matrices corresponds to the dimension of the high-order tensor data. For example, if the dimension of the high-order tensor data is 3-dimensional, there is one projection matrix for each dimension of the high-order tensor data, for a total of 3 projection matrices.

[0138] A low-dimensional embedding tensor can refer to data with a lower dimension than the high-order tensor data obtained by reducing the dimension of the high-order tensor data. For example, the dimension of the high-order tensor data A is 3. In order to make it trainable by the second to-be-trained model, the high-order tensor data A is reduced in dimension to obtain a low-dimensional embedding tensor B with 2 dimensions.

[0139] In the embodiment of the present application, a tensor local preservation projection is performed on the training tensor sample to obtain a projection matrix corresponding to the dimension of the high-order tensor data, including:

[0140] According to the training tensor samples, a neighbor graph corresponding to the training tensor samples is obtained;

[0141] According to the nearest neighbor graph, the correlation matrix is ​​obtained, wherein the correlation matrix is ​​obtained by the heat kernel method according to the nearest neighbor graph;

[0142] According to the correlation matrix, eigendecomposition is performed to obtain the projection matrix corresponding to the dimension of the high-order tensor data.

[0143] The neighborhood graph can refer to the local geometric organization of the training tensor samples in each dimension.

[0144] The heat kernel method can refer to the process of nonlinear dimensionality reduction of the neighbor graph. By presetting the initial heat for the data points in the neighbor graph and then calculating the residual heat after changing with time, the heat kernel characteristics of each point are obtained. The heat kernel characteristics are represented by a matrix to obtain the correlation matrix.

[0145] Eigendecomposition can refer to the decomposition of an incidence matrix into a set of products of eigenvalues ​​and eigenvectors.

[0146] The projection matrix can be obtained by solving the generalized eigenvalue problem on the correlation matrix after eigendecomposition, and obtaining the eigenvector corresponding to the corresponding minimum eigenvalue, thereby obtaining the projection matrix.

[0147] In the embodiment of the present application, the calculation process of the low-dimensional embedding tensor can be:

[0148] Assume that the given sample X T , where X T Including n sample points in I k is the k-mode dimension of the high-order tensor data. Then a neighbor graph is constructed to represent the local geometric structure of M. Then, based on the constructed neighbor graph, the corresponding correlation matrix s = [S ij ] n*n , where s is defined based on the heat kernel method, i = 1…n, j = 1…n. Assume that is the projection matrix. According to the neighbor graph and the correlation matrix s, the optimization problem based on tensor local preservation projection can be expressed as:

[0149]

[0150] in, It can refer to the F norm of A. The F norm of A can refer to the square root of the sum of the squares of all elements in A. In the embodiment of the present application,

[0151] d ii are diagonal elements.

[0152] In this way, the projection matrix can be obtained by calculating the above formula. According to the training tensor sample and the projection matrix, the low-dimensional embedding tensor of the training tensor sample can be expressed as:

[0153] F TLPP =X T ×1U1…× k U k .

[0154] Then, the calculation process of obtaining the attention weight can be:

[0155] According to the attention mechanism method, F TLPP The non-normalized tensor attention feature map α is then normalized by the Softmax function to obtain the normalized tensor attention weight The above normalization calculation process can be expressed as:

[0156]

[0157] Among them, f(F TLPP ) is a function expression for finding α. In this embodiment of the present application, the enhanced feature output satisfies:

[0158]

[0159] in: To enhance feature output, F is the pooling result, is the attention weight.

[0160] The enhanced feature output is obtained by performing an element-wise multiplication operation on the attention weight and the pooling result, and then adding the pooling result.

[0161] In some embodiments, the enhanced feature output can be flattened by a flattening layer to compress the enhanced feature output into a vector, and then the vector is classified by a fully connected layer to output the action category to achieve the purpose of action recognition.

[0162] The embodiment of the present application provides a method for constructing a motion recognition model, which constructs a training tensor sample through tensor data so that the sample contains more information and improves the training effect. By replacing the first weight matrix and the second weight matrix with the first tensor chain and the second tensor chain, the number of parameters of the first model to be trained is reduced, the training efficiency of the first model to be trained is improved, and the computing resource occupancy is reduced. In order to make up for the lack of accuracy of the first target model, the model accuracy of the target model is improved by obtaining the attention weight of the training tensor sample and combining the attention weight with the output of the first target model, so that the target model finally obtained has faster computing efficiency, lower resource occupancy, and higher model accuracy, which ultimately brings better motion recognition effect.

[0163] Figure 3 A flow chart of another method for constructing an action recognition model provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the method includes:

[0164] S301. Obtain a first weight matrix and a second weight matrix in the long short-term memory neural network model according to the training tensor sample and the long short-term memory neural network model.

[0165] Among them, the first weight matrix is ​​used to multiply the input of the long short-term memory neural network model, and the second weight matrix is ​​used to multiply the hidden state of the long short-term memory neural network model.

[0166] The formula of the long short-term memory neural network model is:

[0167]

[0168] Among them, σ is the sigmoid function, tanh is the hyperbolic tangent function, W f is the first weight matrix of the input gate in the first model to be trained, U f is the second weight matrix of the input gate in the first model to be trained, W i is the first weight matrix of the forget gate in the first model to be trained, U i is the second weight matrix of the forget gate in the first model to be trained, W o is the first weight matrix of the output gate in the first model to be trained, U o is the second weight matrix of the output gate in the first model to be trained, W c is the first weight matrix in the candidate cell state in the first model to be trained, U c is the second weight matrix in the candidate cell state in the first model to be trained, x t is the target high-order tensor data at time t in the training tensor sample, b f 、b i 、b o and b c is the bias value, h t-1 is the hidden state at time t-1, h t is the hidden state at time t.

[0169] S302 : Perform tensor decomposition on the first weight matrix and the second weight matrix to obtain a first tensor chain and a second tensor chain.

[0170] The tensor decomposition process is: Assume that the high-order tensor data is an N-order tensor The output hidden state is a D-order tensor Then the first tensor chain and the second tensor chain can be expressed as:

[0171]

[0172] Among them, QTT(W l ) is the first tensor chain, W l is the first weight matrix, QTT(U l ) is the second tensor chain, U l is the second weight matrix, is the strong Kronecker product.

[0173] S303. Replace the first weight matrix and the second weight matrix in the long short-term memory neural network model with the first tensor chain and the second tensor chain respectively to obtain the target hidden state of the long short-term memory neural network model.

[0174] Among them, the target hidden state satisfies:

[0175]

[0176] Among them, QTTL (W i , x t ) is the first transformation data of the input gate in the first model to be trained, QTTL(W f , x t ) is the first transformation data of the forget gate in the first model to be trained, QTTL(W o , x t ) is the first transformation data of the output gate in the first model to be trained, QTTL(W c , x t ) is the first transformation data in the candidate cell state in the first to-be-trained model, QTTL(U i , h t-1 ) is the second transformation data of the input gate in the first to-be-trained model, QTTL(U f , h t-1 ) is the second transformation data of the forget gate in the first to-be-trained model, QTTL(U o , h t-1 ) is the second transformed data of the output gate in the first to-be-trained model and QTTL(U c , h t-1 ) is the second transformed data of the candidate cell state in the first model to be trained, and t is the first target time.

[0177] S304: Perform tensor local preservation projection on the training tensor samples to obtain a projection matrix.

[0178] The process of tensor local preservation projection can be expressed as follows: Assume that n sample points A1,…A are given n , where A i ∈M(i=1…n), I k is the k-mode dimension of the high-order tensor data. First, we construct a neighbor graph G to represent the local geometric structure of M. Then, based on the constructed neighbor graph G, we can obtain its corresponding correlation matrix s = [S ij ] n*n , where s is defined based on the heat kernel method. Then, the eigendecomposition of the training sample set is calculated and the corresponding projection matrix is ​​obtained. Finally, the low-dimensional embedding tensor of the training sample is calculated. Assume is the projection matrix. According to the neighbor graph G and the incidence matrix s, the optimization problem based on tensor local preservation projection can be expressed as:

[0179]

[0180] in, It can refer to the F norm of A. The F norm of A can refer to the square root of the sum of the squares of all elements in A. In the embodiment of the present application,

[0181] d ii are diagonal elements.

[0182] By calculating the above formula, we can get the projection matrix. According to the training tensor sample and the projection matrix, the low-dimensional embedding tensor of the training tensor sample can be expressed as:

[0183] F TLPP =X T ×1U1…× k U k .

[0184] S305. According to the projection matrix and the attention mechanism model, the attention weight of the attention mechanism model is obtained.

[0185] The main process of obtaining the attention weight is as follows: the attention mechanism model includes three trainable parameter matrices, through which the projection matrix is ​​transformed into Q, K, and V of the attention mechanism model so that the attention weight satisfies:

[0186]

[0187] Among them, Attention(Q,K,V) is the attention weight, d k is the dimension size of K.

[0188] S306: Obtain enhanced feature output according to the attention weight and the target hidden state.

[0189] Among them, the enhanced feature output satisfies:

[0190]

[0191] in, To enhance feature output, F is the pooling result, is the attention weight.

[0192] S307, flattening the enhanced feature output according to the enhanced feature output and the flattening layer to obtain an output vector;

[0193] S308. According to the labels of the output vector and the training tensor samples, the long short-term memory neural network model and the attention mechanism model are adjusted to obtain the target model.

[0194] Another method for constructing an action recognition model provided by an embodiment of the present application reduces the number of parameters of the long short-term memory neural network model, improves the training efficiency of the long short-term memory neural network model, and reduces computing resource usage by replacing the first weight matrix and the second weight matrix in the long short-term memory neural network model with the first tensor chain and the second tensor chain. In order to make up for the lack of accuracy of the long short-term memory neural network model, the attention weight of the training tensor sample is obtained through the attention mechanism model, and the attention weight is combined with the output of the long short-term memory neural network model to improve the model accuracy of the target model, so that the target model finally obtained has faster computing efficiency, lower resource usage, and higher model accuracy, which ultimately brings better action recognition effect.

[0195] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0196] Based on the same inventive concept, the present application also provides an apparatus for constructing an action recognition model for implementing the aforementioned method for constructing an action recognition model. The solution provided by this apparatus is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more embodiments of the apparatus for constructing an action recognition model provided below can be found in the aforementioned limitations on the method for constructing an action recognition model, and will not be further elaborated here.

[0197] Figure 4 A schematic diagram of a structure of a motion recognition model building device provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the apparatus includes: a determination module 401, a decomposition module 402, an acquisition module 403, a pooling module 404, and an adjustment module 405, wherein:

[0198] Determination module 401, used to determine a first weight matrix, a second weight matrix and a hidden state of a first to-be-trained model;

[0199] A decomposition module 402 is configured to perform tensor decomposition on the first weight matrix and the second weight matrix to obtain a first tensor chain and a second tensor chain;

[0200] an obtaining module 403 for obtaining a target hidden state based on the first tensor chain, the second tensor chain, the hidden state, and the training tensor sample, wherein the training tensor sample includes high-order tensor data arranged in a time series, the high-order tensor data includes frame data corresponding to a preset action, and the target hidden state is obtained based on the first transformation data and the second transformation data, the first transformation data is obtained based on the first tensor chain and the training tensor sample, and the second transformation data is obtained based on the second tensor chain and the hidden state;

[0201] Pooling module 404, used to perform pooling processing on the target hidden state to obtain a pooling result;

[0202] The adjustment module 405 is used to adjust the first to-be-trained model according to the pooling result and the label of the training tensor sample to obtain a first target model.

[0203] In this embodiment of the present application, the determining module 401 is further configured to:

[0204] Get the training tensor sample corresponding to the preset action;

[0205] According to the training tensor sample and the first model to be trained, a first weight matrix, a second weight matrix, and a hidden state of the first model to be trained are determined.

[0206] In this embodiment of the present application, the determining module 401 is further configured to:

[0207] Acquire video data of a preset action, frame data in the video data, and a time sequence of the frame data;

[0208] Determine width data, height data, and channel information of the frame data according to the frame data;

[0209] Construct high-order tensor data based on width data, height data, and channel information;

[0210] Sort the high-order tensor data according to the time series to obtain training tensor samples.

[0211] In this embodiment of the present application, the determining module 401 is further configured to:

[0212] Determining a first weight matrix and a second weight matrix of a first to-be-trained model;

[0213] Determine a hidden state according to a first weight matrix and a second weight matrix of the first to-be-trained model, wherein the hidden state satisfies:

[0214]

[0215] Among them, σ is the sigmoid function, tanh is the hyperbolic tangent function, W f is the first weight matrix of the input gate in the first model to be trained, U f is the second weight matrix of the input gate in the first model to be trained, W i is the first weight matrix of the forget gate in the first model to be trained, U i is the second weight matrix of the forget gate in the first model to be trained, W o is the first weight matrix of the output gate in the first model to be trained, U o is the second weight matrix of the output gate in the first model to be trained, W c is the first weight matrix in the candidate cell state in the first model to be trained, U c is the second weight matrix in the candidate cell state in the first model to be trained, x t is the target high-order tensor data at time t in the training tensor sample, b f 、b i 、b o and b c is the bias value, h t-1 is the hidden state at time t-1, h t is the hidden state at time t.

[0216] In this embodiment of the present application, the obtaining module 403 is further used to:

[0217] Obtaining first transformed data according to the first tensor chain and the high-order tensor data at the first target moment;

[0218] Obtaining second transformed data according to the second tensor chain and the hidden state of the second target moment, wherein the first target moment is a next target moment of the second target moment;

[0219] A target hidden state is obtained according to the first transformed data and the second transformed data.

[0220] In this embodiment of the present application, the obtaining module 403 is further used to:

[0221] The target hidden state satisfies:

[0222]

[0223] Among them, QTTL (W i , x t ) is the first transformation data of the input gate in the first model to be trained, QTTL(W f , x t ) is the first transformation data of the forget gate in the first model to be trained, QTTL(Wo , x t ) is the first transformation data of the output gate in the first model to be trained, QTTL(W c , x t ) is the first transformation data in the candidate cell state in the first to-be-trained model, QTTL(U i , h t-1 ) is the second transformation data of the input gate in the first to-be-trained model, QTTL(U f , h t-1 ) is the second transformation data of the forget gate in the first to-be-trained model, QTTL(U o , h t-1 ) is the second transformed data of the output gate in the first to-be-trained model and QTTL(U c , h t-1 ) is the second transformed data of the candidate cell state in the first model to be trained, and t is the first target time.

[0224] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in all the above method embodiments when executing the computer program.

[0225] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in all the above method embodiments are implemented.

[0226] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in all the above method embodiments when executed by a processor.

[0227] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0228] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.

[0229] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0230] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for constructing an action recognition model, characterized in that: The method comprises: Determine a first weight matrix, a second weight matrix, and a hidden state of a first to-be-trained model according to the training tensor sample; Performing tensor decomposition on the first weight matrix and the second weight matrix respectively to obtain a first tensor chain and a second tensor chain; Obtaining a target hidden state according to the first tensor chain, the second tensor chain, the hidden state, and a training tensor sample; Performing pooling processing on the target hidden state to obtain a pooling result; Adjusting the first to-be-trained model according to the pooling result and the label of the training tensor sample to obtain a first target model; The training tensor samples include high-order tensor data arranged in a time series, and the high-order tensor data includes frame data corresponding to a preset action.

2. The method according to claim 1, characterized in that Obtaining a target hidden state according to the first tensor chain, the second tensor chain, the hidden state, and a training tensor sample includes: Obtaining first transformed data according to the first tensor chain and the training tensor sample; Obtaining second transformed data based on the second tensor chain and the hidden state; The target hidden state is obtained according to the first transformed data and the second transformed data.

3. The method according to claim 1, characterized in that The step of obtaining the training tensor sample includes: Acquire video data of a preset action, frame data in the video data, and a time sequence of the frame data; Determining width data, height data, and channel information of the frame data according to the frame data; Constructing high-order tensor data according to the width data, the height data, and the channel information; The high-order tensor data is sorted according to the time series to obtain training tensor samples.

4. The method according to claim 1, wherein The determining, based on the training tensor sample and the first model to be trained, a first weight matrix, a second weight matrix, and a hidden state of the first model to be trained includes: Determining a first weight matrix and a second weight matrix of the first to-be-trained model; Determine a hidden state according to the first weight matrix and the second weight matrix of the first to-be-trained model.

5. The method according to claim 1, wherein Obtaining a target hidden state according to the first tensor chain, the second tensor chain, the hidden state, and a training tensor sample includes: Obtaining first transformed data according to the first tensor chain and the high-order tensor data at a first target time; obtaining second transformed data according to the second tensor chain and the hidden state at a second target moment, the first target moment being a target moment next to the second target moment; A target hidden state is obtained according to the first transformed data and the second transformed data.

6. The method according to claim 1, characterized in that After adjusting the first to-be-trained model according to the pooling result and the label of the training tensor sample to obtain a first target model, the method further includes: Training a second model to be trained according to the training tensor sample to obtain an attention weight of the second model to be trained; Obtaining enhanced feature output according to the pooling result and the attention weight; According to the enhanced feature output and the preset action, the first model to be trained and the second model to be trained are adjusted to obtain a target model.

7. The method according to claim 6, characterized in that The step of training the second model to be trained according to the training tensor sample to obtain the attention weight of the second model to be trained includes: Performing tensor locality-preserving projection on the training tensor samples to obtain a projection matrix corresponding to the dimension of the high-order tensor data; Obtaining a low-dimensional embedding tensor of the training tensor sample according to the training tensor sample and the projection matrix; A second model to be trained is trained according to the low-dimensional embedding tensor to obtain an attention weight of the second model to be trained.

8. The method according to claim 7, characterized in that The performing tensor local preservation projection on the training tensor sample to obtain a projection matrix corresponding to the dimension of the high-order tensor data includes: Obtaining a neighbor graph corresponding to the training tensor sample according to the training tensor sample; Obtaining an association matrix according to the neighbor graph, wherein the association matrix is ​​obtained according to the neighbor graph by using a heat kernel method; Perform eigendecomposition according to the incidence matrix to obtain a projection matrix corresponding to the dimension of the high-order tensor data.

9. A motion recognition model construction device, characterized in that: The device comprises: A determination module, configured to determine a first weight matrix, a second weight matrix, and a hidden state of a first to-be-trained model based on the training tensor sample; a decomposition module, configured to perform tensor decomposition on the first weight matrix and the second weight matrix respectively to obtain a first tensor chain and a second tensor chain; an obtaining module, configured to obtain a target hidden state according to the first tensor chain, the second tensor chain, the hidden state, and a training tensor sample; A pooling module, configured to perform pooling processing on the target hidden state to obtain a pooling result; an adjustment module, configured to adjust the first to-be-trained model according to the pooling result and the label of the training tensor sample to obtain a first target model; The training tensor samples include high-order tensor data arranged in a time series, and the high-order tensor data includes frame data corresponding to a preset action.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.