Memory enhancement action recognition method and system on Riemannian manifold and storage medium
By introducing memory enhancement methods on Riemann manifolds in human body movement recognition, the shortcomings of the prior art in spatial and temporal feature fusion, dynamic sensitivity and robustness are solved, and higher recognition accuracy and a wide range of application scenarios are achieved.
Patent Information
- Application Number
- CN202510699609.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing human action recognition methods have many challenges in dealing with complex spatiotemporal relationships, multi-view data fusion, keyframe capture and long-time series modeling, including time information loss, spatial and temporal feature modeling separation, and insufficient modeling of nonlinear dynamic feature of high-dimensional data.
A memory-enhanced action recognition method on Riemann manifold is proposed. By collecting human motion data, calculating first-order difference information, expanding into a third-order tensor, the weight matrix is calculated using the human short-term memory mechanism, performing principal component analysis, mapping to unit hypersphere, and optimizing weight parameters using Monte Carlo Markov chain algorithm, and finally using K nearest neighbor classifier for action recognition.
It improves recognition accuracy, enhances modeling ability for complex actions, reduces computational complexity, improves model robustness, and is suitable for multiple application scenarios.
Smart Images

Figure CN120220252A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular, to a memory-enhanced action recognition method and system on a Riemannian manifold. Background Art
[0002] Human Action Recognition (HAR) is an important research direction in the field of computer vision, and its main goal is to accurately and quickly identify human behavior patterns from video or sensor data. In recent years, with the rapid development of deep learning, human action recognition technology has been widely applied in fields such as healthcare monitoring, intelligent security, sports training, and virtual reality. However, existing methods still face many challenges in dealing with complex spatio-temporal relationships, multi-view data fusion, key frame capture, and long-time sequence modeling.
[0003] Traditional methods mainly rely on handcrafted feature extraction, such as optical flow, spatio-temporal interest points, etc. However, these methods have poor adaptability in complex scenarios and limited performance for high-dynamic actions. In recent years, deep learning-based action recognition methods have made significant progress. Among them, convolutional neural networks can effectively extract spatial features but cannot fully model temporal information; recurrent neural networks and their variants, such as long short-term memory gated recurrent units, can capture short-term dependencies but have the problem of gradient vanishing when modeling long-time sequence information; spatio-temporal graph convolutional networks use graph neural networks to model human skeleton information, but their dependence on skeleton data limits the generalization ability of application scenarios. In recent years, Transformer-based methods such as ST-TransNet capture long-time sequence dependencies through self-attention mechanisms, but they have high computational complexity and are difficult to apply to real-time scenarios.
[0004] Existing methods still have certain limitations in modeling fine-grained dynamic information of human actions, optimizing spatio-temporal feature fusion, and improving cross-class recognition ability. For long-time sequences, existing models are difficult to maintain long-term time dependencies, thus affecting recognition accuracy. In addition, most methods adopt a separation strategy when modeling spatial and temporal features, resulting in insufficient feature fusion and reducing the expression ability of action patterns. Specifically, existing methods have the following three main technical problems: First, the traditional covariance matrix method has a problem of time information loss when modeling spatio-temporal features. When quantifying the correlation between modalities, the averaging effect of the covariance matrix will dilute or mask important time information, including action rhythm and acceleration changes. Especially in long sequences, this limitation reduces the model's ability to detect and respond to key time features.
[0005] Second, existing methods perform poorly in handling fast actions (such as throwing or running) and long continuous action sequences. These methods often model space and time as independent processes, limiting their ability to effectively integrate spatio-temporal features and resulting in insufficient representation of fine-grained dynamics and complex spatio-temporal fusion patterns.
[0006] Third, existing dynamic weighting strategies mainly focus on the time dimension, while the dynamic modeling of spatial saliency features is still insufficient. Most methods adopt static or random weight assignment strategies, unable to balance the importance of spatio-temporal features. Especially in high-dimensional data such as human action sequences, the modeling of non-linear dynamic features is still insufficient.
[0007] Therefore, it is of great significance to propose a method that can simultaneously optimize spatio-temporal feature fusion, improve short-term dynamic sensitivity, and enhance model robustness for the accuracy and stability of human action recognition. Summary of the Invention
[0008] The technical problem to be solved by the embodiments of the present invention is to provide a memory-enhanced action recognition method and system on a Riemannian manifold to improve the recognition accuracy in view of the limitations of existing human action recognition methods in dealing with complex dynamic changes, long-term time dependencies, and spatio-temporal feature interactions. To solve the above technical problem, the embodiments of the present invention propose a memory-enhanced action recognition method on a Riemannian manifold, including: Step 1: Collect human action data, process the data according to the first-order difference information of the data, and represent the data as a third-order tensor X i ; Step 2: Unfold the third-order tensor X i to obtain three matrices, each matrix corresponding to a different subspace; Step 3: Calculate the corresponding memory-enhanced weight matrix using the human short-term memory mechanism; Step 4: Decompose the weight matrix by the principal component analysis method to obtain the weighted basis vectors; Step 5: Recombine the basis vectors into a matrix and normalize each row vector of the matrix so that it is mapped onto the unit hypersphere, where each row vector of the matrix corresponds to a point on the sphere; Step 6: Use the Monte Carlo Markov chain algorithm to learn the optimal modal weight parameters, calculate the angular distance between points on the hypersphere, and obtain the geometric differences between points; Step 7: Adopt a K-nearest neighbor classifier to classify human actions according to the geometric differences and output the classification results.
[0009] Correspondingly, the embodiments of the present invention also provide a memory-enhanced action recognition system on a Riemannian manifold, including: Data acquisition and processing module: It acquires human motion data, processes the data according to the first-order difference information of the data, represents the data as a third-order tensor; unfolds the third-order tensor to obtain three matrices, and each matrix corresponds to a different subspace; Dynamic weighting module: It calculates the corresponding memory-enhanced weight matrix by using the human short-term memory mechanism; decomposes the weight matrix by the principal component analysis method to obtain the weighted basis vectors; Riemannian metric module: It reorganizes the basis vectors into a matrix, normalizes each row vector of the matrix so that it is mapped onto the unit hypersphere, where each row vector of the matrix corresponds to a point on the sphere; calculates the angular distance between the points on the hypersphere to obtain the geometric difference between the points; Parameter optimization module: It uses the Monte Carlo Markov chain algorithm to optimize the weight parameter w of each modality j , and adaptively balances the contributions of different modalities to the overall geometric difference; Action recognition module: It adopts a K-nearest neighbor classifier to classify human actions according to the geometric difference and outputs the classification result.
[0010] Correspondingly, an embodiment of the present invention also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the memory-enhanced action recognition method on the Riemannian manifold.
[0011] The beneficial effects of the present invention are as follows: 1. Improved recognition accuracy: By mapping the tensor onto the hypersphere and using the angular relationship to quantify the changes between modalities, the present invention effectively preserves the geometric structure of spatial and temporal information, reduces information loss, and achieves state-of-the-art recognition results on datasets such as CMB and NW.
[0012] 2. Enhanced modeling ability for complex actions: The dynamic weighting mechanism of the present invention can adaptively capture rapid motion changes and long-term time dependencies, especially the sensitivity to subtle action changes in complex scenarios.
[0013] 3. Reduced computational complexity: Compared with deep learning methods, the method based on tensor geometry of the present invention has higher computational efficiency, especially when dealing with long time series, the computational time is significantly reduced.
[0014] 4. Improved model robustness: Through the normalization process of hypersphere mapping, the present invention has stronger resistance to external factors such as illumination changes, perspective changes, and background noise, and reduces the impact of environmental interference on the recognition result.
[0015] 5. Wide range of application scenarios: The present invention realizes accurate action recognition, laying a foundation for multiple application fields, including healthcare monitoring, virtual reality interaction, sports training analysis, intelligent security monitoring, and other fields. Brief Description of the Drawings
[0016] Figure 1 is a schematic structural diagram of the memory-enhanced action recognition method on the Riemannian manifold according to an embodiment of the present invention.
[0017] Figure 2 is a schematic diagram of tensor unfolding according to an embodiment of the present invention, showing the process of unfolding the tensor along different modalities.
[0018] Figure 3 is a schematic diagram of mapping according to an embodiment of the present invention, showing the distribution of the normalized points on the unit hypersphere. Detailed Embodiment
[0019] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0020] In the embodiments of the present invention, if there are directional indications (such as up, down, left, right, front, back...), they are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0021] In addition, in the present invention, the descriptions such as "first" and "second" are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature.
[0022] Please refer to Figure 1 , the memory-enhanced action recognition method on the Riemannian manifold according to an embodiment of the present invention includes steps 1 to 7.
[0023] Step 1: Collect human action data (as shown in (a) in Figure 1 , the original data of human actions includes original information and differential information). Process the data according to the first-order differential information of the data, and represent the data as a third-order tensor X i . The human action data contains n samples, denoted as . Each sample represents a video sequence of an action process, is the corresponding label, where is the total number of action categories. The data is organized as a third-order tensor, where and represent the spatial dimensions of the video frames respectively, represents the temporal dimension, i.e., the number of frames in the sequence. This tensor representation effectively captures the spatio-temporal structure of the action data.
[0024] Human action data is inherently complex and is often affected by external factors such as lighting changes, noise, and individual differences, etc. These factors can obscure key spatio-temporal patterns and pose challenges to accurate classification. To mitigate these effects, the present invention introduces first-order difference information to emphasize the motion dynamics along the temporal dimension, thereby reducing the influence of external static features and noise. The first-order difference tensor represents the discrete-time change of the temporal pattern along , and is defined as: ; where represents the discrete difference of the tensor along the temporal pattern, and is defined as: ; Circular padding is applied at the boundaries to ensure that the dimension of the first-order difference tensor is the same as that of the original tensor , thus ensuring compatibility with subsequent processing steps.
[0025] The first-order difference tensor extracts temporal features such as action onset, intensity change, and speed fluctuation. It suppresses redundant information such as static background and lighting changes, thereby highlighting key spatio-temporal features and improving the accuracy and robustness of the model in a dynamic environment.
[0026] Step 2: Unfold the third-order tensor X i to obtain three matrices (as shown in (b) of Figure 1 ), and each matrix corresponds to a different subspace.
[0027] The spatio-temporal tensor representation provides a structured geometric framework for video data and is analyzed by unfolding the tensor into subspaces. These representation methods have been widely used in action recognition because they can model the intrinsic relationship between spatial and temporal modalities. For a given video tensor that contains width , height and time dimensions, the tensor is unfolded into three mode-specific matrices: ; , , They represent matrices unfolded along the width dimension, the height dimension, and the time dimension, respectively. Each unfolded matrix corresponds to a different subspace, providing an in-depth understanding of the spatial and temporal features of the action data. This decomposition makes the geometric interpretation of the tensor more comprehensive by focusing on the inherent multimodal structure of the tensor.
[0028] A commonly used method for generating subspace representations is to calculate the covariance matrix of the unfolded tensor. The covariance matrix is usually defined as: ; where the tensor represents the tensor unfolded along the mode. Covariance-based methods are commonly used to generate subspace representations. They can effectively capture global correlations but also have significant limitations. These methods result in the loss of key spatial or temporal information through cross-modal variance averaging. Specifically, for the time mode , the covariance operation suppresses rapid changes and masks dynamic features such as speed changes and action boundaries.
[0029] To overcome the limitations of covariance-based methods, the present invention represents the video tensor as a spatio-temporal subspace. For an action video tensor that captures the width, height, and time dimensions, the tensor is unfolded along three modes to extract structured subspaces (see Figure 2, Figure 2 Modes 1 and 2: Visual attention mechanisms in the spatial dimension, from left to right / from right to left; from top to bottom / from bottom to top; Mode 3: In the time dimension, different viewing orders correspond to different dynamic memory patterns, emphasizing the start / end phases of the sequence). This method results in the loss of key spatial or temporal information, especially for tensors with time dependencies (such as videos).
[0030] In the time mode, the tensor can be unfolded as: ; where the mode matrix represents the frames along the time dimension. Each column vector corresponds to the time frame in the sequence. The time information is essentially encoded by the order of the columns in the matrix , reflecting the order of the frames.
[0031] However, constructing the subspace of the time modality requires performing a singular value decomposition (SVD) on the autocorrelation matrix . SVD is insensitive to the column order of the matrix, which means that regardless of how the frame sequence is arranged, it produces the same eigenvectors and eigenvalues. Therefore, the original matrix The temporal information encoded by the column order is not preserved in the subspace representation.
[0032] The present invention adopts a multi-layer feature fusion strategy, combining temporal information and spatial structure, so that the model can capture motion characteristics at different time scales and spatial scales, and improve the modeling ability of high-order spatiotemporal relationships.
[0033] Step 3: Use the human short-term memory mechanism to calculate the corresponding memory enhancement weight matrix (such as Figure 1 As shown in (c), the corresponding forward matrix and backward matrix are obtained according to the initial matrix). The present invention adopts a memory-enhanced subspace modeling strategy, calculates the covariance matrix by combining motion features with memory weights, and extracts the main motion mode to avoid the problem of time information loss during the calculation process, so as to enhance the time-dependent modeling capability. Since the time information is reduced (eliminated) by matrix calculation during the decomposition process when the covariance matrix is decomposed, the present invention solves this problem through STM weights.
[0034] Step 4: Decompose the weight matrix by principal component analysis to obtain basis vectors with weights.
[0035] The human short-term memory (STM) mechanism is able to retain historical information for a short period of time, allowing the encoding of sequential information. A variety of computational models simulate STM. For example, the biologically inspired neuron model encodes temporal information by calculating the difference between the input sequence and the neuron state by simulating the dynamic changes of membrane potential through delay. The response vector of the neuron response The calculation formula is: ; in, is the current input, It is The reference point of each neuron, is the forgetting parameter. This parameter controls the balance between past information and current information. Specifically, when , the system will retain all past information equally; and when It only keeps the latest information.
[0036] In order to better process sequence data, a time adaptive self-organizing map (TASOM) is proposed. This model captures the temporal characteristics of sequence signals through a dynamic update mechanism. In this model, the input pattern is defined as a weighted combination of the current input and the previous input: ; in, Indicates the current input mode, is the memory depth parameter, which dynamically controls the weighting between historical information and current information. A smaller emphasizes long-term historical information, while a larger prioritizes changes in the current input.
[0037] The above equation can be re-expressed as: ; can be expressed in the form of matrix multiplication: ; is the sum of the column terms of the weight matrix AX under the memory mechanism.
[0038] The traditional covariance matrix is calculated as , substituting into the memory mode, we can get , and obtain ; Since A is a diagonal matrix, , finally we get: ; At this time, the structure of the subspace is only determined by the covariance matrix, and has nothing to do with the memory matrix A, resulting in the loss of temporal information.
[0039] To solve this problem, the present invention specifically integrates the temporal structure and the memory mechanism into the subspace representation of the data. By centralizing the data and applying PCA (Principal Component Analysis) projection (as shown in (d) of Figure 1 ), a new subspace representation is obtained, and the PCA projection is: .
[0040] For the memory-enabled mode , its PCA projection expression is: ; In this way, the projection result includes both the influence of standard PCA and the offset of the memory mechanism on the data structure, avoiding the loss of temporal information.
[0041] Step 5: Recombine the basis vectors into a matrix, and normalize each row vector of the matrix so that it is mapped onto the unit hypersphere (as shown in (e) of Figure 1 ), where each row vector of the matrix corresponds to a point on the sphere. The present invention reduces the mean drift through normalization and improves the stability of the data distribution.
[0042] Step 6: Use the Monte Carlo Markov Chain algorithm to learn the optimal modal weight parameters, calculate the angular distance between points on the hypersphere, and obtain the geometric differences between samples (as shown in (e) of Figure 1 ).
[0043] To compare subspaces from different samples, a row normalization operation maps the mode matrix onto the hypersphere. The normalization is performed row by row, and the formula is as follows: ; where represents the Euclidean norm of the th row. The normalized matrix represents a set of points located on the hypersphere, and each point corresponds to a row vector of the original matrix (see Figure 3 ). This mapping preserves the geometric structure of the data while facilitating the calculation of the distance between subspaces.
[0044] Since the hypersphere is a Riemannian manifold with a constant positive curvature, the present invention uses the Riemannian metric to measure the geometric differences between points on the manifold. The metric on the hypersphere must consider its curvature characteristics. To this end, the present invention introduces the geodesic distance, which is the shortest distance between two points on the manifold. On the hypersphere, the geodesic distance is calculated through the angular distance between two points, and this angle measures the included angle between them on the hypersphere and effectively reflects their geometric differences.
[0045] The calculation of geometric differences includes: First, calculate the standard angular distance between two normalized mode matrices and : ; where the standard angle θ l is calculated through , represents the singular value of the matrix , and then use the Monte Carlo Markov Chain algorithm to learn the optimal modal weight parameter w j , and summarize the geometric relationships of all tensor modes: .
[0046] where is the mode-specific weight parameter used to adjust the contribution of each mode to the overall geometric difference measure. This weighting mechanism ensures that key modes (such as time or space dimensions) are given priority. Therefore, the model can focus more effectively on features highly relevant to the task.
[0047] The Monte Carlo Markov Chain algorithm includes the following steps: (1) Initialize the weight parameter w j ; (2) For each iteration t = 1 to T: a) Generate candidate weights from the proposal distribution: ; b) Calculate the acceptance probability ; where represents the posterior probability distribution; c) Generate a uniform random number ; d) If the condition is satisfied, accept the proposal and update ; otherwise retain the original value: ; (3) Return the optimized weight parameter .
[0048] Step 7: Use a K-nearest neighbor classifier (as shown in (e) in Figure 1 ) to classify human body movements according to geometric differences and output the classification results. To identify human body movement data, the present invention calculates the Riemannian distance between all data points, selects the K closest points through the KNN method (k-nearest neighbor algorithm, a commonly used supervised learning method), and determines the label of the human body movement data by voting according to the labels of these K points. To optimize the KNN classification results, the present invention optimizes the neighborhood relationship on the hypersphere so that the data distribution can adapt to the geometric structures of different categories and improve the cross-category recognition ability.
[0049] The present invention is applicable to different categories of human body movements, including basic movements (walking, running, jumping) and complex activities (throwing, dancing, martial arts, etc.), enhancing the adaptability of the model. The present invention supports the processing of video data with different resolutions to ensure that the recognition performance can be maintained in both high-resolution and low-resolution environments. The present invention adopts an adaptive weight adjustment strategy to optimize the sensitivity of the model to different motion categories, improve the cross-category generalization ability, and make it applicable to a variety of application scenarios. The present invention can be integrated into an online learning framework, enabling the model to be dynamically updated as new data is input, improving the adaptability to environmental changes and individual differences, and enhancing the recognition accuracy in long-term usage scenarios.
[0050] The memory-enhanced action recognition system on the Riemannian manifold of the present invention includes: Data acquisition and processing module: Collect human body movement data, and process the data according to the first-order difference information i of the data X to process the data, and the third-order tensor X i and They are respectively expanded to obtain three matrices, and each matrix corresponds to a different subspace; Dynamic weighting module: Calculate the corresponding memory-enhanced weight matrix by using the human short-term memory mechanism; Decompose the weight matrix by using the principal component analysis method to obtain the basis vectors with weights; Riemannian metric module: Recombine the basis vectors into a matrix, and normalize each row vector of the matrix to map it onto the unit hypersphere, where each row vector of the matrix corresponds to a point on the sphere; Calculate the angular distance between points on the hypersphere to obtain the geometric differences between points; Parameter optimization module: Use the MCMC algorithm to optimize the weight parameter w of each modality j , and adaptively balance the contributions of different modalities to the overall geometric differences; Action recognition module: Adopt a K-nearest neighbor classifier to classify human actions according to the geometric differences and output the classification results.
[0051] As an implementation, the data acquisition and processing module includes: Image acquisition unit: Used to acquire human action sequence data from the video source array; First-order difference calculation unit: Used to calculate the first-order difference in the time dimension of the acquired video data to emphasize the motion dynamic features; Tensor construction unit: Used to organize the differenced data into a third-order tensor structure ; Tensor expansion unit: Used to expand the third-order tensor into the corresponding matrix representation along different modalities 、 、 。
[0052] As an implementation, the dynamic weighting module includes: Memory parameter calculation unit: Used to calculate the memory depth parameter a∈(0,1) based on the human short-term memory mechanism; Weight matrix generation unit: Used to generate the corresponding weight matrix A according to the memory depth parameter; Principal component analysis unit: Used to perform principal component analysis on the weighted data to extract the main eigenvectors V with weights; Among them, the memory parameter calculation unit realizes the adaptive fusion of information on different time scales by dynamically adjusting the value of the parameter a.
[0053] As an implementation, the Riemannian metric module includes: Hypersphere mapping unit: Used to normalize the matrix row vectors and map them onto the unit hypersphere; Angle distance calculation unit: used to calculate the geodesic distance between points on the hypersphere and quantify the geometric differences between subspaces; wherein, the angle distance calculation unit calculates the standard angle θ through the singular value decomposition method l , ensuring the geometric preservation of spatial and temporal features.
[0054] As an implementation, the parameter optimization module includes: Objective function construction unit: used to construct an optimization objective function based on the recognition accuracy; Monte Carlo Markov chain sampling unit: used to execute the Monte Carlo Markov chain algorithm to generate a Markov chain of weight parameters; Parameter convergence judgment unit: used to evaluate the convergence of weight parameters and determine the final optimization result; wherein, the Monte Carlo Markov chain sampling unit guides the parameter search process through the posterior probability distribution P(·), ensuring that the weight assignment matches the action category discrimination ability.
[0055] In addition, the present invention also proposes a storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the memory-enhanced action recognition method on the Riemannian manifold are realized.
[0056] Inspired by the human cognitive mechanism, the present invention introduces a dynamic weighted allocation strategy that can adapt to changes in time and space. Specifically, the present invention maps the unfolded tensor pattern to the unit hypersphere to ensure the preservation of the global and local geometric relationships in spatio-temporal features. By utilizing the angular relationship on the hypersphere, the present invention can quantify the relative changes between tensor patterns, thus alleviating the common problem of time information loss in traditional covariance-based methods. In addition, the present invention prioritizes key time moments and spatial regions through dynamic weighting, effectively coping with the challenges brought by fast motion and subtle spatial changes. This dynamic weighting mechanism enhances the fusion and modeling of spatio-temporal features, thereby improving the model's ability to capture the inherent time dynamics and spatial dependencies in human actions. The method of the present invention has been evaluated on the CMB, NW, UTK, KTH, and MHAD datasets, and has achieved state-of-the-art results on the CMB and NW datasets. These results demonstrate the effectiveness and robustness of the method of the present invention in action recognition scenarios.
[0057] The main innovation points of the embodiments of the present invention include: 1. Propose an adaptive dynamic weighted allocation strategy inspired by the human memory pattern, which dynamically adjusts the weight distribution according to time changes and the spatial observation order, enabling the model to efficiently focus on key time points and important spatial regions.
[0058] 2. A dynamic feature modeling method based on hypersphere embedding is designed, which maps the matrix rows of the unfolded tensor onto the unit hypersphere, alleviating the common time information loss problem in traditional covariance matrix methods.
[0059] 3. The properties of the constant curvature of the hypersphere and manifold geometry tools are used to jointly analyze spatio-temporal features, adjust the spatial weight distribution, enhance the sensitivity of the model to key spatial features, and effectively fuse spatio-temporal information.
[0060] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A memory-enhanced action recognition method on Riemannian manifolds, characterized in that, Including: Step 1: Collect human motion data, process the data according to the first-order difference information of the data, and represent the data as a third-order tensor X i ; Step 2: Unfold the third-order tensor X i to obtain three matrices, each corresponding to a different subspace; Step 3: Calculate the corresponding memory-enhanced weight matrix by using the human short-term memory mechanism; Step 4: Decompose the weight matrix by using the principal component analysis method to obtain the basis vectors with weights; Step 5: Recombine the basis vectors into a matrix, and normalize each row vector of the matrix so that it is mapped onto the unit hypersphere, where each row vector of the matrix corresponds to a point on the sphere; Step 6: Use the Monte Carlo Markov chain algorithm to learn the optimal modal weight parameters, calculate the angular distance between the points on the hypersphere, and obtain the geometric differences between the samples; Step 7: Adopt a K-nearest neighbor classifier to classify human body actions according to the geometric differences and output the classification results.
2. The memory-enhanced action recognition method on a Riemannian manifold according to claim 1, wherein In step 1, the human motion data is represented as a third-order tensor , where and represent the width and height in the spatial dimension respectively; represents the time dimension and contains the number of frames in the video sequence.
3. The memory-enhanced action recognition method on a Riemannian manifold according to claim 1, characterized in that In step 1, the first-order difference tensor represents the discrete-time change in the time pattern along , which is defined as: ; Among them represents a tensor The discrete difference along the time mode is defined as: ; Among them, cyclic padding is adopted at the time dimension boundary to maintain the consistency of data dimensions.
4. The memory-enhanced action recognition method on a Riemannian manifold according to claim 2, wherein In Step 2, the three matrices are respectively expressed as: ; Among them, , , respectively represent matrices expanded along the width dimension, the height dimension, and the time dimension.
5. The memory-enhanced action recognition method on a Riemannian manifold according to claim 1, characterized in that, In Step 3, the time information is encoded by calculating the difference between the input sequence and the neuron state: ; Among them, is the current input, is the reference point of the th neuron, is the memory depth parameter, representing the response vector of the th neuron's response; And capture the temporal characteristics in the sequence signal through a dynamic update mechanism, and define the input pattern as a weighted combination of the current input and the previous input: 。 6. The memory-enhanced action recognition method on Riemannian manifolds according to claim 2, wherein In Step 5, the row normalization is performed according to the following formula: ; Among them, represents the Euclidean norm of the th row, and the normalized matrix represents a set of points located on the hypersphere, where each point corresponds to a row vector in the matrix.
7. The memory-enhanced action recognition method on a Riemannian manifold according to claim 1, wherein In Step 6, the calculation of the geometric differences includes: First, calculate the standard angular distance between two normalized pattern matrices and : ; where θ l Through It can be calculated that denotes the singular value of the matrix , and then the Monte Carlo Markov chain algorithm is used to learn the optimal modal weight parameter w j , and summarize the geometric relationships of all tensor modes: 。 8. The memory-enhanced action recognition method on a Riemannian manifold according to claim 7, wherein, The Monte Carlo Markov chain algorithm includes the following steps: (1) Initialize the weight parameter w j ; (2) For each iteration from t = 1 to T: a) Generate candidate weights from the proposal distribution: ; b) Calculate the acceptance probability ; wherein represents the posterior probability distribution; c) Generate uniform random numbers ; d) If the condition is satisfied then accept the proposal and update ; otherwise retain the original value: ; (3) Return the optimized weight parameters .
9. A memory-enhanced action recognition system on a Riemannian manifold, characterized in that, Including: Data acquisition and processing module: Acquire human body action data, process the data according to the first-order difference information of the data, and represent the data as a third-order tensor; Expand the third-order tensor to obtain three matrices, and each matrix corresponds to a different subspace; Dynamic weighting module: Calculate the corresponding memory-enhanced weight matrix by using the human short-term memory mechanism; Decompose the weight matrix by using the principal component analysis method to obtain the basis vectors with weights; Riemannian metric module: Recombine the basis vectors into a matrix, and normalize each row vector of the matrix so that it is mapped onto the unit hypersphere, where each row vector of the matrix corresponds to a point on the sphere; Calculate the angular distance between the points on the hypersphere to obtain the geometric differences between the points; Parameter optimization module: Use the Monte Carlo Markov Chain algorithm to optimize the weight parameter w of each modality j , and adaptively balance the contributions of different modalities to the overall geometric difference; Action recognition module: Adopt a K-nearest neighbor classifier to classify human body actions according to the geometric differences and output the classification results.
10. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the memory-enhanced action recognition method on the Riemannian manifold according to any one of claims 1 to 8.
Citation Information
Patent Citations
An identification method for movement by human bodies irrelevant with the viewpoint based on stencil matching
CN101216896A
Human body motion classification method based on compression perception
CN106056135A
Cited By
A data processing method for digital representation of sports intangible cultural heritage
CN122634021A