Skeleton data processing method based on graph Fourier transform and space-time synchronous graph convolutional network

By combining graph Fourier transform and spatiotemporal synchronous graph convolutional network, the problems of spatiotemporal relationship capture and noise filtering in skeleton data processing are solved, achieving higher recognition accuracy and generalization ability, which is suitable for scenarios such as motion prediction, human-computer interaction and intelligent monitoring.

CN120726337APending Publication Date: 2025-09-30WUXI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510729511.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively capturing the high-dimensional spatiotemporal relationships of skeleton data, ignore frequency domain information, lack noise filtering mechanisms, have poor generalization capabilities, and are prone to forgetting important features when processing long sequence data.

Method used

Graph Fourier transform is used to extract frequency domain features, which are then combined with spatiotemporal synchronous graph convolutional network for feature fusion. Kalman filtering and memory replay mechanism are used to optimize the prediction results, and a spatiotemporal synchronous graph convolutional network is constructed for feature extraction and prediction.

Benefits of technology

It improves the generalization ability and recognition accuracy of skeleton data, enhances the robustness of the model and its ability to process long sequence data, and adapts to the application needs of different scenarios and individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726337A_ABST
    Figure CN120726337A_ABST
Patent Text Reader

Abstract

The invention provides a skeleton data processing method based on graph Fourier transform and a space-time synchronization graph convolutional network, which comprises the following steps: converting skeleton data to a frequency domain through graph Fourier transform, then extracting space-time characteristics by using the space-time synchronization graph convolutional network, and finally carrying out data smoothing processing through memory playback and Kalman filtering. Therefore, the generalization ability and the recognition precision of the skeleton data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of skeleton data processing, and more specifically to a skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network. Background Art

[0002] This method predicts future motion sequences based on spatiotemporal feature analysis and pattern recognition of dynamic skeleton data. Dynamic skeleton data prediction is widely used in scenarios such as motion prediction, human-computer interaction, virtual reality, and intelligent monitoring. It has important applications in computer vision and artificial intelligence.

[0003] Although significant progress has been made in dynamic skeleton data processing technology, it still faces many challenges in practical applications: (1) Skeleton data has complex spatiotemporal dependencies, and traditional convolutional neural networks and recurrent neural networks are difficult to effectively capture this high-dimensional spatiotemporal relationship; (2) Existing methods often ignore the global frequency domain information of skeleton data during feature extraction, resulting in limited ability to express dynamic features; (3) Skeleton data is easily affected by noise during the acquisition process, affecting the robustness and accuracy of the model; (4) Existing methods have poor generalization capabilities when processing skeleton data from different scenarios and individuals, and are difficult to adapt to diverse application needs.

[0004] In response to the above problems, existing technologies have the following deficiencies: (1) Traditional methods usually process temporal and spatial features separately and cannot fully model the spatiotemporal correlation in skeleton data; (2) The frequency domain features of skeleton data contain rich global information, but existing methods mainly rely on time domain features and fail to fully utilize frequency domain information; (3) Existing methods lack an effective noise filtering mechanism, resulting in the performance degradation of the model under noise interference; (4) Existing models are prone to forgetting important features when processing long sequence data, and lack an effective feedback mechanism to optimize model performance. Summary of the Invention

[0005] In order to solve the problems of low generalization ability and low recognition accuracy of skeleton graph data in the existing technology, the present invention proposes a skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network.

[0006] In order to achieve the above technical effects, the technical solutions of the present invention are as follows: Obtaining initial image data, extracting skeleton data from the initial image data, the skeleton data including joint point coordinates; performing graph Fourier transform on the skeleton data to extract frequency domain features; normalizing the joint point coordinates and converting them into a tensor format; A spatiotemporal synchronous graph convolutional network is constructed and trained using the processed tensors to extract joint information from both the temporal and spatial domains. The skeleton data is then subjected to feature extraction using a multi-layer spatiotemporal synchronous graph convolutional layer constructed using the spatiotemporal synchronous graph convolutional network to obtain joint features in the skeleton data. The extracted frequency domain features, joint information, and joint features in the skeleton data are then fused to obtain fused features. The joint point coordinate information in the fusion feature is converted into tensor form and input into the trained spatiotemporal synchronous graph convolutional network to obtain the prediction results of the joint point coordinates; The prediction results of the joint point coordinates are smoothed and optimized by Kalman filtering after extracting the features of the joint point coordinates in the memory playback to obtain the processed joint point coordinate data.

[0007] Compared with the prior art, the beneficial effects of the technical solution of the present invention are: The present invention proposes a skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network. The present invention converts the skeleton data into the frequency domain through graph Fourier transform, then uses the spatiotemporal synchronous graph convolutional network to extract spatiotemporal features and fuses the features. Finally, the spatiotemporal synchronous graph convolutional network is input to obtain the skeleton data prediction result, thereby improving the generalization ability and recognition accuracy of the skeleton data. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 The flowchart of the skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network is shown in an embodiment of the present invention.

[0009] Figure 2 This is an overall flow chart showing an embodiment of the present invention.

[0010] Figure 3 This is a structural diagram of the spatiotemporal synchronization graph convolution module shown in an embodiment of the present invention. DETAILED DESCRIPTION

[0011] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0012] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0013] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0014] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0015] Example 1 This embodiment proposes a skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network, the flow chart of which is as follows: Figure 1 shown.

[0016] Obtaining initial image data, extracting skeleton data from the initial image data, the skeleton data including joint point coordinates; performing graph Fourier transform on the skeleton data to extract frequency domain features; normalizing the joint point coordinates and converting them into a tensor format; A spatiotemporal synchronous graph convolutional network is constructed and trained using the processed tensors to extract joint information from both the temporal and spatial domains. The skeleton data is then subjected to feature extraction using a multi-layer spatiotemporal synchronous graph convolutional layer constructed using the spatiotemporal synchronous graph convolutional network to obtain joint features in the skeleton data. The extracted frequency domain features, joint information, and joint features in the skeleton data are then fused to obtain fused features. The joint point coordinate information in the fusion feature is converted into tensor form and input into the trained spatiotemporal synchronous graph convolutional network to obtain the prediction results of the joint point coordinates; The prediction results of the joint point coordinates are smoothed and optimized by Kalman filtering after extracting the features of the joint point coordinates in the memory playback to obtain the processed joint point coordinate data.

[0017] This embodiment converts the skeleton data into the frequency domain through graph Fourier transform, then uses a spatiotemporal synchronous graph convolutional network to extract spatiotemporal features, and finally performs data smoothing through memory playback and Kalman filtering, thereby improving the generalization ability and recognition accuracy of the skeleton data.

[0018] Example 2 This embodiment further explains the scheme in detail based on embodiment 1. The overall flow chart is as follows: Figure 2 shown.

[0019] In an optional embodiment, obtaining initial image data, extracting skeleton data of the initial image data, the skeleton data including joint point coordinates and labels, and performing a graph Fourier transform on the skeleton data specifically comprises the following steps: Use OpenPose to preprocess the initial image to obtain skeleton data; The skeleton data is converted into a grayscale image and then Fourier transformed to obtain frequency domain features; The frequency domain features are centralized and then filtered to obtain the filtered frequency domain features; The filtered frequency domain features are inverse Fourier transformed to obtain the inverse transformed skeleton data.

[0020] Furthermore, the graph Fourier transform extracts the global features of the skeleton data through frequency domain analysis and filters out high-frequency noise, making full use of the frequency domain information and enhancing the robustness and feature expression ability of the model.

[0021] First, the skeleton data is read and converted into a grayscale image. Grayscale processing can reduce computational complexity while retaining the main structural information of the image. Secondly, the grayscale image is subjected to Fourier transform to convert the image from the spatial domain to the frequency domain. The Fourier transform can decompose the image into components of different frequencies. The low-frequency components usually correspond to the overall structure of the image, such as the main motion trend of the skeleton, while the high-frequency components correspond to details and noise. Through the Fourier transform, the module can extract the global features of the skeleton data from the frequency domain, especially the low-frequency components, which usually correspond to the overall motion trend of the skeleton. Compared with traditional time domain analysis methods, frequency domain analysis can better capture the global dynamic patterns of skeleton data.

[0022] Secondly, the Fourier-transformed spectrum is centered, moving the zero-frequency component (DC component) to the center of the spectrum. This step facilitates subsequent filtering operations, as low-frequency components are typically concentrated in the center of the spectrum. A low-pass filter is also used to filter out high-frequency components. This involves creating a two-dimensional mask with the center region set to 1 (allowing transmission) and the rest of the region set to 0 (blocking). This mask is then applied to the Fourier-transformed image. The low-pass filter effectively filters out high-frequency noise, improving data quality. Skeleton data is susceptible to noise during acquisition, such as sensor jitter and changes in ambient lighting. Through frequency-domain filtering, the graph Fourier transform can significantly reduce the impact of this noise on the model, improving its robustness.

[0023] Finally, the filtered spectrum is shifted back to its original position and an inverse Fourier transform is performed to convert the image from the frequency domain back to the spatial domain. Finally, the image is normalized to ensure that the data values ​​are within a reasonable range for subsequent model processing. Through the inverse transform, the module can smooth the skeleton data, reducing outliers and jitter in the data, making subsequent model processing more stable and reliable.

[0024] The computational complexity of Fourier transforms and low-pass filtering is relatively low, especially when the image size is small. They can quickly complete frequency domain analysis and filtering operations to extract the frequency domain information of the skeleton data, fully utilizing global features and being suitable for real-time applications. Therefore, this module can adapt to different skeleton data sources, such as data from different devices and different perspectives. Through normalization, it eliminates scale differences between different data sources and enhances the model's generalization capabilities.

[0025] In an optional embodiment, the frequency domain features are centered and then filtered to obtain filtered frequency domain features, and a low-pass filter is used to filter high-frequency noise; the mathematical expression of the filtered frequency domain features is: filtered =X⊙M in, filtered Represents the frequency domain features after filtering, M represents the mask matrix of the low-pass filter, and ⊙ represents element-by-element multiplication.

[0026] The spatiotemporal synchronous graph convolutional network includes a fully connected layer, a spatiotemporal synchronous graph convolutional layer and a fully connected layer; wherein, the spatiotemporal synchronous graph convolutional layer includes a local spatiotemporal graph and a spatiotemporal synchronous graph convolutional module, the local spatiotemporal graph is constructed according to the preprocessed skeleton graph, the spatiotemporal synchronous graph convolutional module is constructed according to the local spatiotemporal graph, and the spatiotemporal synchronous graph convolutional layer is constructed according to the spatiotemporal synchronous graph convolutional module.

[0027] In an optional embodiment, the structure diagram of the spatiotemporal synchronization graph convolution module is as follows: Figure 3 As shown, it includes a first local spatiotemporal graph signal matrix submodule, a first graph convolutional layer submodule, a second graph convolutional layer submodule, a first aggregation layer submodule and a first local spatiotemporal neighborhood feature representation submodule connected in sequence, and the output end of the first graph convolutional layer submodule is also connected to the input end of the first aggregation layer submodule.

[0028] In an optional embodiment, the skeleton data captures the temporal and spatial relationships between skeleton joints by performing convolution operations on skeleton node features and adjacency matrices in a spatiotemporal synchronous graph convolution module to obtain joint point information in the time domain and the space domain.

[0029] In an optional embodiment, the joint point coordinate information in the fusion feature is converted into a tensor form and input into the trained spatiotemporal synchronous graph convolutional network to obtain the prediction result of the joint point coordinates; the time step i The prediction result calculation expression is:

[0030] in, (i) represents the time step i prediction, W1(i)∈R TC×C′ , b1(i)∈R C′ , W2(i)∈R C′×1 , b2(i)∈R represents a learnable parameter; The prediction results of all time steps are concatenated; the expression is: =[ (1), (2),…, (T)]∈R N×T in, A matrix representing the concatenation of prediction results for all time steps.

[0031] Furthermore, the spatiotemporal synchronized graph convolutional network (GCN) combines information from both spatial and temporal dimensions to capture the complex spatiotemporal dependencies in skeleton data, enabling accurate prediction of action sequences. The GCN design leverages the strengths of GCNs, effectively processing non-Euclidean data structures such as graphs. By combining the temporal and spatial dimensions through a spatiotemporal synchronization mechanism, it further enhances the model's ability to capture high-dimensional spatiotemporal relationships.

[0032] The features output after the graph Fourier transform include the node feature matrix X and the adjacency matrix A of the dynamic skeleton data, which serve as input to the spatiotemporal synchronous graph convolution module. A graph convolution operation is performed on the input node feature matrix X to extract local spatial features. Joint coordinates are normalized to the range [0, 1] through normalization and padding. The processed data is then converted into a tensor for input into the spatiotemporal synchronous graph convolution module for training and testing.

[0033] The graph convolutional layer is responsible for extracting spatial features from the skeleton data. The skeleton graph consists of nodes (joints) and edges (connections between joints). The graph convolutional layer captures the spatial relationships between joints by convolving the node features with the adjacency matrix. Specifically, a 1×1 convolution is performed to output a new feature map with the same number of channels as the shallow-layer features. The spatiotemporal synchronized graph convolutional layer performs the following operations: First, a learnable weight matrix is ​​introduced and bitwise multiplied with the adjacency matrix to assign greater weights to important edges in the adjacency matrix while suppressing the weights of less important edges. This learnable weight matrix allows the module to dynamically adjust the edge weights in the adjacency matrix, focusing on important joint connections while suppressing less important ones, thereby improving the model's expressiveness. The weighted adjacency matrix and the input data are then fed into the graph convolutional network for computation. Furthermore, a residual structure is introduced to calculate the residual, which is bitwise added to the output of the graph convolutional network to aggregate spatial information. This residual structure enables the model to preserve shallow-layer features, avoid losing important information in deeper networks, and further enhance model performance. The model integrates features by performing graph convolution operations on a local spatiotemporal graph. This approach uses a spatiotemporal synchronization mechanism to simultaneously consider both spatial and temporal dimensions, specifically the influence of the previous time node (joint positions in the previous frame) on the next time node (joint positions in the current frame). This spatiotemporal synchronization enables the model to effectively capture long-range spatiotemporal correlations within action sequences, enabling accurate action recognition and analysis. This spatiotemporal synchronization mechanism enables the spatiotemporally synchronized graph convolutional network to effectively capture dynamic information within action sequences, enabling the modeling and analysis of spatiotemporal features. The C×H×W feature layer undergoes global average pooling and global max pooling along the channel dimension, resulting in four feature vectors of size C×1×1. H and W are the height and width of the feature, respectively, and C is the number of channels. Through global average pooling and global max pooling, the module extracts key information from spatiotemporal features, enhancing the model's ability to recognize action categories and improve its generalization.

[0034] Furthermore, the following steps are included in the spatiotemporal synchronous graph convolutional network: The preprocessed joint point coordinates are converted into tensors for input into the spatiotemporal synchronized graph convolution module for training and testing.

[0035] The processed multidimensional array is converted into multiple C×1×1 feature vectors and concatenated along the channel dimension into a one-dimensional feature vector of size NC×1×1. This one-dimensional feature vector is fed into the fully connected layer, and its mathematical expression is as follows: z(l)=σ(ω(l)z(l-1)+b(l)) Among them, l represents the number of layers in the current fully connected layer, z(l) represents the output of the current layer of the fully connected layer, z(l-1) represents the output of the previous layer, σ(i) represents the activation function, ω represents the weight coefficient, and b represents the bias coefficient.

[0036] The extracted feature data is input into the spatiotemporal graph convolution module for high-level feature extraction and action prediction.

[0037] Compute mean and error estimates: Among them, the mean value of the predicted coordinate point is estimated: for each coordinate axis direction (x, y), the mean of the true coordinate point and the predicted coordinate point is calculated:

[0038] Among them, i represents the coordinate axis direction (x, y), j represents the time point, and N represents the total number of time points.

[0039] Error estimation: Calculate the error between the true coordinate point and the predicted coordinate point at each time point:

[0040] For each axis direction, calculate the mean and standard deviation of the error:

[0041] in, represents the mean of the errors, Represents the standard deviation of the error.

[0042] In an optional embodiment, the prediction result of the joint point coordinates is extracted by memory playback, and the feature of the joint point coordinates is further included in dynamically updating the memory library by a sliding window mechanism; the calculation expression is: M=update(M, (t-1)) (t)=predict(X(t),M) in, (t) represents the prediction result at time step t, M represents the memory bank, update represents the operation of updating the memory bank, and predict represents the operation of fusing the current input and historical information for prediction.

[0043] Furthermore, the memory replay module enhances the model's ability to capture long-term dependencies in action sequences by introducing historical prediction results as additional input to the current time step. In action sequence prediction tasks, the action in the current frame is often closely related to the action in previous frames. Therefore, leveraging historical information can help the model better understand the dynamic changes in the action.

[0044] The memory replay module optimizes the memory replay process by introducing a dynamic error threshold mechanism, further improving the accuracy of action sequence prediction. After predicting the skeleton coordinates, the model calculates the Euclidean distance error between the predicted value and the true value and sets an adjustable threshold to classify the results: predicted points with errors exceeding the threshold are marked as inaccurate, while those with errors below the threshold are classified as accurate. The memory replay module focuses on extracting the spatiotemporal features of inaccurate parts of historical predictions and uses them as supplementary information for the current time step input. At the same time, it dynamically updates the memory bank through a sliding window mechanism, retaining predictions for recent keyframes and gradually eliminating earlier redundant data. This strategy enables the model to focus on action segments with large errors, strengthening its ability to model complex spatiotemporal relationships during training. At the same time, it smoothes and corrects historical predictions in the memory bank through Kalman filtering, forming a closed-loop "prediction-evaluation-feedback" optimization process. This effectively addresses the problem of feature forgetting in long-sequence tasks and significantly improves the model's robustness to noise interference and cross-scenario generalization.

[0045] The memory replay module significantly enhances the model's ability to capture long-term dependencies in action sequences by introducing historical prediction results as additional input to the current time step. By incorporating historical prediction results, the memory replay module can effectively capture long-term dependencies in action sequences. This is particularly important when processing long sequences of data, as changes in actions often exhibit temporal continuity, requiring the model to understand the correlations between previous and subsequent frames. Furthermore, by combining information from the current and historical frames, the memory pool enables the model to generate more stable and consistent predictions. This stability is crucial when processing noisy data or complex action sequences, reducing jitter and outliers in the predictions.

[0046] In an optional embodiment, the prediction result is smoothed and optimized by Kalman filtering; the calculation expression is: X t =F t X t-1 +B t u t P t =F t P t-1 +Q t Among them, X t represents the state vector, F t represents the state transition matrix, X t-1 represents the state vector, B t represents the control input matrix, u t represents the control input vector, Q trepresents the process noise covariance matrix.

[0047] In an optional embodiment, the prediction results are smoothed and optimized by Kalman filtering to obtain processed joint point coordinate data; the expression is: smoothed (t)=x t Among them, Xt represents the smoothed state estimate.

[0048] Furthermore, the Kalman filter reduces noise and outliers in the prediction results through state estimation and iterative updates, thereby improving the model's robustness and prediction accuracy. The Kalman filter is a recursive algorithm that dynamically adjusts predictions based on the system's state transition model and observations, bringing them closer to the true value. Specifically, the Kalman filter module achieves data smoothing and optimization through the following steps: First, state prediction is performed. Based on the system's state transition model, the Kalman filter predicts the state Xt+1 and covariance matrix Pt+1 for the next time step. The state transition matrix Ft describes the temporal continuity of joint motion. The control input matrix Bt and input vector ut are used to adjust the state prediction. Next, the observation update is performed. The Kalman filter updates the state estimate using the observation Zt, i.e., the predicted coordinates output by the spatiotemporally synchronized graph convolutional network. The observation matrix Ht maps the state vector to the observation space, and the Kalman gain Kt balances the weight between the predicted and observed values. Through continuous iterative optimization, the Kalman filter gradually updates the state estimate Xt and covariance matrix Pt, reducing prediction error. This iterative optimization process forms a closed-loop feedback mechanism, enabling the model to make more accurate predictions in subsequent time steps. Ultimately, the Kalman filter generates the smoothed output Xt, the denoised two-dimensional coordinates of the joint points, as the final prediction. This smoothing process effectively reduces the impact of noise and outliers on the prediction results, improving the robustness of the model.

[0049] Kalman filtering can effectively reduce noise and outliers in prediction results, improving model robustness. Through smoothing, Kalman filtering significantly reduces the impact of these disturbances on model performance. Furthermore, by combining a state transition model with observations, Kalman filtering dynamically adjusts predictions to bring them closer to the true value. This dynamic adjustment mechanism enables the model to adapt to varying data changes and improve prediction accuracy. Kalman filtering, combined with a feedback mechanism, forms a closed-loop optimization process. The model iteratively updates the state estimate and covariance matrix, gradually reducing prediction error and enabling more accurate predictions in subsequent time steps. Kalman filtering is particularly well-suited for action recognition tasks with time dependencies. By fully leveraging information in time series, the model can better understand the dynamics of actions and improve its prediction capabilities for complex action sequences. Furthermore, Kalman filtering enhances the model's generalization across diverse scenarios, enabling it to maintain stable performance across diverse data sources and environmental conditions. Furthermore, through iterative updates and state prediction, Kalman filtering effectively processes long sequences of data, avoiding the loss of important features. This design is particularly suitable for action recognition tasks that require capturing long-term dependencies.

[0050] In this embodiment, a spatiotemporal synchronous graph convolutional network effectively extracts spatiotemporal correlations from skeleton data by combining information from both spatial and temporal dimensions, dynamically adjusting weights, introducing a residual structure, and performing global pooling operations, significantly improving the model's motion prediction capabilities. It boasts the advantages of efficient joint modeling of spatiotemporal features, dynamic weight adjustment, information retention, strong generalization, adaptability, and excellent noise suppression, making it suitable for processing long-sequence data. This embodiment reduces noise interference through Kalman filtering, improving the robustness of the network model. A memory replay and feedback mechanism enhances the processing capability of long-sequence data, preventing the forgetting of important features and optimizing network model performance.

[0051] Example 3 This embodiment is described by way of example based on Embodiment 1 and Embodiment 2.

[0052] Obtain the Human3.6M dataset. During training, use the Human3.6M dataset to train the model. The Human3.6M dataset is a widely used dataset for 3D human pose estimation. It contains a large number of human action sequences performed by professional actors in a controlled environment. This dataset was released by the University of Pennsylvania to promote research in 3D human pose estimation and action recognition.

[0053] We then extracted skeleton images from the Human3.6M dataset by extracting frames using a video processing library. This process involves extracting frames from the videos, dividing them into individual frames at a frame rate of 30 frames per second (fps). The extracted frames are then saved as an image sequence for subsequent processing.

[0054] Each frame is processed using OpenPose, an open-source tool developed by Carnegie Mellon University, or another skeleton extraction tool to detect and extract human joints. OpenPose is a popular open-source tool that can detect key points in real time from a single image or video sequence. Model parameters are adjusted to optimize joint detection accuracy. During processing, OpenPose outputs the coordinates of each joint, which are then used to generate a skeleton graph.

[0055] In practical application scenarios, first, video data is collected in real time using a camera or other sensor, typically at a frame rate of 30 frames per second. Next, a video processing library is used to extract frames from the live video, segmenting the video into individual frames. Each frame is processed using OpenPose or other skeleton extraction tools to detect and extract human joints within the image, generating a skeleton graph. Finally, the extracted skeleton graph is normalized, and the joint coordinates are converted to tensor form, which is then fed into a trained spatiotemporal graph convolutional network for prediction.

[0056] This example performs a graph Fourier transform on the skeleton graph. First, a normalized Laplace matrix is ​​constructed for the skeleton graph, transforming the complex joint connections into a frequency-domain parseable graph signal. Frequency-domain basis functions are then derived by eigendecomposing the matrix, and node features are projected into the frequency domain. This process not only extracts low-frequency components that reflect limb movement trends but also suppresses high-frequency noise through a dynamic low-pass filter, significantly improving the model's robustness in complex environments.

[0057] Furthermore, to address the lack of generalization due to cross-scene data discrepancies, this method incorporates adaptive normalization and quality enhancement techniques during the frequency domain feature reconstruction phase. Spatial features after the inverse Fourier transform are normalized using joint coordinates to eliminate scale deviations caused by different devices or viewpoints. Image resolution is also adjusted to enhance detail clarity. This frequency-spatial domain collaborative processing mechanism enables the model to remove individual-specific interference from macroscopic motion patterns (such as walking and waving), enhancing the learning of common features across different scenes.

[0058] Furthermore, the Laplacian matrix of the skeleton graph is constructed. The skeleton graph consists of nodes (joints) and edges (connections between joints). The adjacency matrix A∈R N×N Represents the connection relationship between nodes, where N is the number of joints.

[0059] Calculate the degree matrix D, D∈R N×N , whose diagonal elements are the degrees of each node. Construct the Laplacian matrix, whose mathematical expression is as follows: L = DA And normalized, its mathematical expression is as follows: L=ID -1 / 2 AD -1 / 2 Where I represents the identity matrix.

[0060] Perform eigendecomposition on the normalized Laplace matrix Lnorm to obtain the eigenvector matrix U∈R N×N and the eigenvalue matrix Λ∈R N×N , its mathematical expression is as follows: Lnorm = UΛU T Where each column of U is an eigenvector, Λ is a diagonal matrix, and the diagonal elements are eigenvalues.

[0061] Project the node feature matrix of the skeleton graph into the frequency domain: X = U T X Where X represents the frequency domain feature matrix.

[0062] In the frequency domain, a low-pass filter is applied to filter out high-frequency noise and retain important low-frequency features. filtered The mathematical expression is as follows: filtered =X⊙M Where M represents the mask matrix of the low-pass filter and ⊙ represents element-by-element multiplication.

[0063] The filtered frequency domain features filtered Converted back to the spatial domain through the inverse Fourier transform, its mathematical expression is as follows: X filtered =U filtered Among them, X filtered ∈R N×C Represents the smoothed spatial domain feature matrix.

[0064] The skeleton image after inverse transformation is further preprocessed to normalize the joint point coordinates, adjust the image size, and enhance the image quality. Normalizing the joint point coordinates helps to eliminate the differences between different videos and make the model training more stable.

[0065] Ultimately, this technology has demonstrated significant advantages in practical applications. In industrial safety monitoring, the model uses global frequency domain features to accurately identify whether workers are properly wearing safety equipment, maintaining high accuracy even in complex environments with multiple occluders or fluctuating lighting. In sports training analysis, multi-band fusion features can simultaneously capture an athlete's overall posture stability and joint micro-movement details, providing a quantitative basis for technical improvements.

[0066] Construct a spatiotemporal synchronized graph convolutional network. The spatiotemporal synchronized graph convolutional network is a structure of fully connected layer-spatiotemporal synchronized graph convolutional layer-fully connected layer, in which the spatiotemporal synchronized graph convolutional layer is mainly composed of multiple spatiotemporal synchronized graph convolutional layers.

[0067] In the process of human motion prediction, the extraction of spatiotemporal feature information is very important. The present invention improves on the graph convolutional network to perform feature extraction operations. The present invention uses a spatiotemporal synchronized graph convolutional network for spatiotemporal network data prediction. The model can effectively capture complex local spatiotemporal correlations through a carefully designed spatiotemporal synchronized modeling mechanism. At the same time, multiple modules for different time periods are designed in the model to effectively capture the heterogeneity in the local spatiotemporal graph. In order to capture the long-range spatiotemporal correlation of the entire network sequence, we use multiple spatiotemporal synchronized graph convolution modules to simulate different periods instead of sharing one for all periods. Using a sliding window to cut out different periods, this design enables the spatiotemporal synchronized graph convolutional network to simultaneously consider the spatial and temporal dependencies of the skeleton data, thereby improving the accuracy of motion recognition. By deploying multiple spatiotemporal synchronized graph convolutional modules in different time periods, the spatiotemporal synchronized graph convolutional network can better capture the spatiotemporal features within different periods, thereby enhancing the model's understanding and prediction capabilities of motion sequences.

[0068] The spatiotemporal graph convolution module consists of a set of graph convolution operations. The input of the graph convolution operation is the graph signal matrix of the local spatiotemporal graph. Each node aggregates its own features and those of its neighbors at adjacent time steps. The aggregation function is a linear combination whose weight is equal to the weight of the edge between the node and its neighbors. A fully connected layer with an activation function is then added to transform the node features into a new space. The graph convolution operation can be expressed as follows:

[0069] in, represents the adjacency matrix of the local space-time graph, represents the input of the l-th layer graph convolution, Represents a learnable parameter and σ represents an activation function. In addition, a self-loop is added to each node in the local spatiotemporal graph to allow the graph convolution operation to consider its own features when aggregating features. The steps are as follows: Introducing the weight matrix, first introduce a learnable weight matrix and multiply it with the adjacency matrix bit by bit to give the important edges in the adjacency matrix a larger weight while suppressing the weight of the unimportant edges. Its mathematical expression is as follows:

[0070] Here, ⊙ represents bitwise multiplication, A represents the adjacency matrix, and W represents the learnable weight matrix.

[0071] The weighted adjacency matrix and input data are fed into the graph convolutional network for computation. At the same time, a residual structure is introduced to calculate the residual, which is then bitwise added to the output of the graph convolutional network to aggregate spatial dimension information.

[0072] By integrating features through graph convolution operations on the local spatiotemporal graph, it simultaneously considers both spatial and temporal dimensions through a spatiotemporal synchronization mechanism, namely the influence of the previous time node (the joint position in the previous frame) on the next time node (the joint position in the current frame). This spatiotemporal synchronization allows the model to effectively capture long-range spatiotemporal correlations in action sequences, thereby achieving accurate action recognition and analysis. This spatiotemporal synchronization mechanism enables the spatiotemporally synchronized graph convolutional network to effectively capture dynamic information in action sequences, thereby enabling the modeling and analysis of spatiotemporal features.

[0073] The feature layer of size C×H×W performs global maximum pooling along the channel dimension, obtaining four feature vectors of size C×1×1. Among them, H and W are the height and width of the feature, respectively, and C is the number of channels. The calculation formula for global maximum pooling is: GMPc=maxc(X(i,j,c)) Among them, H and W represent the height and width of the feature layer, GMPc represents the global maximum pooling on channel c, X(i,j,c) represents the value of the pixel coordinate position (i,j) on channel c, and maxc represents the maximum value at all positions of channel X of the object.

[0074] The input of the above spatiotemporal synchronization graph convolutional layer is X∈R T×N×C , first convert it to ∈R N×TC , and then generate the prediction results through two layers of fully connected layers:

[0075] in, (i) represents the time step i prediction, W1(i)∈R TC×C′ , b1(i)∈R C′ , W2(i)∈R C′×1, b2(i)∈R is a learnable parameter. Finally, the prediction results of all time steps are concatenated into a matrix.

[0076] In this process, the ReLU function is used to introduce nonlinear characteristics, and the fully connected layer is used to map the features to the final output space. In this way, the model can learn action categories or regression targets from spatiotemporal features to achieve accurate action prediction.

[0077] The network model was trained using the dataset using supervised training. Data augmentation was first performed on the images in the dataset. The original images and their corresponding labels were then converted into tensors and fed into the model for training. The optimizer used was the Adaptive Moment Estimator (Adam), and the learning rate strategy adopted the dynamically adjusted "Poly" strategy. The initial learning rate was set to 0.0001, with a decay coefficient of 0.98. The learning rate was updated every three training runs, for a total of 150 runs, using a batch size of 16.

[0078] Use the trained network model to predict actions. After the training is completed, the model weight will be obtained, and then enter the prediction stage of the model. When predicting, the present invention uses the trained spatiotemporal synchronous graph convolutional network for prediction.

[0079] By leveraging the temporal characteristics of the memory replay model when processing sequential data, the memory replay process is optimized by introducing a dynamic error threshold mechanism, further improving the accuracy of action sequence prediction. After predicting the skeleton graph coordinates, the model calculates the Euclidean distance error between the predicted value and the true value, and sets an adjustable threshold to classify the results: predicted points with errors exceeding the threshold are marked as inaccurate, while those with errors below the threshold are classified as accurate. The memory replay module focuses on extracting the spatiotemporal features of the inaccurate parts of historical predictions and uses them as supplementary information for the current time step input. At the same time, it dynamically updates the memory bank through a sliding window mechanism, retaining the prediction results of recent key frames and gradually eliminating earlier redundant data. This strategy enables the model to focus on action segments with large errors, strengthening its ability to model complex spatiotemporal relationships during training. At the same time, the Kalman filter is used to smooth and correct historical predictions in the memory bank, forming a closed-loop optimization process of "prediction-evaluation-feedback", which effectively solves the problem of feature forgetting in long sequence tasks and significantly improves the model's robustness to noise interference and cross-scenario generalization ability. Its mathematical representation is as follows: M=update(M, (t-1)) (t)=predict(X(t),M) in, ( t ) represents the time step tThe prediction result, M represents the memory bank, update represents the operation of updating the memory bank, and predict represents the operation of fusing the current input and historical information for prediction.

[0080] Furthermore, the predictions for the previous time step of the model are stored in memory. This can be done with a simple array or list, or using more complex data structures like queues to store sequential data.

[0081] At the current time step, the prediction for the previous time step is retrieved from the storage structure. This can be a single prediction or all predictions for a period of time. After completing the prediction for the current time step, the storage structure is updated, adding the current prediction to the storage and removing the prediction for the oldest time step to limit the storage size.

[0082] The processed data is smoothed and optimized using Kalman filtering, which can effectively reduce noise and outliers in the prediction results. This paper uses Kalman filtering to iteratively optimize the output of the spatiotemporal synchronous graph convolutional network, forming a closed-loop feedback mechanism. The specific mathematical formula is as follows: X t =F t X t-1 +B t u t P t =F t P t-1 +Q t The input includes: state vector: X t X t-1 ; Observation value: Zt (with X t X t-1 The dimensions are consistent), that is, the predicted coordinates output by the spatiotemporal synchronous graph convolutional network at the current time step; the state transfer matrix: Ft (dimension 2N×2N), used to describe the temporal continuity of joint motion; the observation matrix: Ht (dimension 2N×2N), maps the state vector to the observation space; Among them, F t represents the state transition matrix, B t represents the control input matrix, u t represents the control input vector, Q t represents the process noise covariance matrix.

[0083] According to the observation value Zt, the Kalman gain Kt is calculated, which is mathematically expressed as follows: K t =Pt (H t P t +R t ) -1 Among them, H t represents the observation matrix, R t represents the observation noise covariance matrix, Xt represents the updated state estimate and Pt represents the covariance matrix;

[0084] Applying the Kalman filter to the model's prediction results generates the following mathematical expression for the smoothed output: smoothed (t)=x t The output is the smoothed state estimate Xt (dimension N×2), that is, the two-dimensional coordinates of the joint point after denoising, as the final prediction result.

[0085] The feedback operation in the above method helps the model make more accurate predictions in subsequent time steps. By applying Kalman filtering and feedback operations, the spatiotemporal synchronized graph convolution model can effectively process action sequence data and improve the accuracy and robustness of action recognition. This method is particularly suitable for handling time-dependent action recognition tasks, fully utilizing the information in the time series to achieve more accurate predictions.

[0086] Each embodiment of the present invention is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment. The device embodiment described above is merely exemplary. The modules described as separate components may or may not be physically separated. When implementing the scheme of the present invention, the functions of each module can be implemented in the same one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the scheme of this embodiment.

[0087] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network, characterized in that: The following steps are involved: Obtaining initial image data, extracting skeleton data of the initial image data, the skeleton data including the coordinates of joint points; performing graph Fourier transform on the skeleton data to extract frequency domain features; Normalize the joint point coordinates and convert them into tensor format; A spatiotemporal synchronous graph convolutional network is constructed and trained using the processed tensors to extract joint information from both the temporal and spatial domains. The skeleton data is then subjected to feature extraction using the spatiotemporal synchronous graph convolutional layers constructed using a multi-layer spatiotemporal synchronous graph convolutional network to obtain joint features in the skeleton data. The extracted frequency domain features, joint information, and joint features in the skeleton data are then fused to obtain fused features. The joint point coordinate information in the fusion feature is converted into tensor form and input into the trained spatiotemporal synchronous graph convolutional network to obtain the prediction results of the joint point coordinates; The prediction results of the joint point coordinates are smoothed and optimized by Kalman filtering after extracting the features of the joint point coordinates in the memory playback to obtain the processed joint point coordinate data.

2. The skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network according to claim 1 is characterized in that: Obtaining initial image data, extracting skeleton data from the initial image data, and performing graph Fourier transform on the skeleton data specifically include the following steps: Use OpenPose to preprocess the initial image to obtain skeleton data; The skeleton data is converted into a grayscale image and then Fourier transformed to obtain frequency domain features; The frequency domain features are centralized and then filtered to obtain the filtered frequency domain features; The filtered frequency domain features are inverse Fourier transformed to obtain the inverse transformed skeleton data.

3. The skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network according to claim 1, characterized in that: The frequency domain features are centralized and then filtered to obtain the filtered frequency domain features. A low-pass filter is used to filter high-frequency noise. The mathematical expression of the filtered frequency domain features is: filtered =X⊙M in, filtered Represents the frequency domain features after filtering, M represents the mask matrix of the low-pass filter, and ⊙ represents element-by-element multiplication.

4. The skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network according to claim 1, characterized in that: The spatiotemporal synchronized graph convolutional network includes a local spatiotemporal graph, a spatiotemporal synchronized graph convolutional module and a spatiotemporal synchronized graph convolutional layer.

5. The skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network according to claim 4 is characterized in that: The spatiotemporal synchronous graph convolution module includes a first local spatiotemporal graph signal matrix submodule, a first graph convolution layer submodule, a second graph convolution layer submodule, a first aggregation layer submodule and a first local spatiotemporal neighborhood feature representation submodule, which are connected in sequence. The output end of the first graph convolution layer submodule is also connected to the input end of the first aggregation layer submodule.

6. The skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network according to claim 4, characterized in that: The skeleton data is convolved with skeleton node features and adjacency matrices in a spatiotemporal synchronous graph convolution module to capture the temporal and spatial relationships between skeleton joints, thereby obtaining joint point information in the temporal and spatial domains.

7. The skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network according to claim 6, characterized in that: The joint point coordinate information in the fusion feature is converted into tensor form and input into the trained spatiotemporal synchronous graph convolutional network to obtain the prediction result of the joint point coordinates; time step i The prediction result calculation expression is: in, (i) represents the time step i prediction, W1(i)∈R TC×C′ , b1(i)∈R C′ , W2(i)∈R C′×1 , b2(i)∈R represents a learnable parameter; The prediction results of all time steps are concatenated; the expression is: =[ (1), (2),…, (T)]∈R N×T in, A matrix representing the concatenation of prediction results for all time steps.

8. The skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network according to claim 1, characterized in that: The prediction results of the joint point coordinates are extracted through memory playback. The features of the joint point coordinates also include dynamically updating the memory library through a sliding window mechanism; its calculation expression is: M=update(M, (t-1)) (t)=predict(X(t),M) in, (t) represents the prediction result at time step t, M represents the memory bank, update represents the operation of updating the memory bank, and predict represents the operation of fusing the current input and historical information for prediction.

9. The skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network according to claim 1, characterized in that: The prediction results are smoothed and optimized through Kalman filtering; its calculation expression is: X t =F t X t-1 +B t u t P t =F t P t-1 +Q t Among them, X t represents the state vector, F t represents the state transition matrix, X t-1 represents the state vector, B t represents the control input matrix, u t represents the control input vector, Q t represents the process noise covariance matrix.

10. The skeleton data processing method based on graph Fourier transform and spatiotemporal synchronous graph convolutional network according to claim 9, characterized in that: The prediction results are smoothed and optimized through Kalman filtering to obtain the processed joint point coordinate data; Its expression is: smoothed (t)=x t Among them, Xt represents the smoothed state estimate.

Citation Information

Cited By

  • Graph neural network machine learning model anomaly detection method based on graph Fourier transform enhancement

    CN121117805A

  • Anomaly Detection Method Based on Graph Fourier Transform Enhanced Graph Neural Network Machine Learning Model

    CN121117805B