A Pose Detection Method for Video Style Transfer Based on Graph Neural Network
By constructing a dynamic spatiotemporal correlation network and generating a dynamic mask, the problem of insufficient capture of dynamic changes in joint relationships in style transfer is solved, and the effects of high-precision pose recognition and motion abnormality determination are achieved.
Patent Information
- Application Number
- CN202510413754.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing graph convolution method adopts a fixed adjacency matrix, which cannot adaptively capture the dynamic changes in joint relationships caused by style transfer, and is not robust enough in occlusion or fast motion scenarios.
By constructing a dynamic spatiotemporal correlation network, the pixel map of the video frame is dynamically calculated as the connection weight of the spatiotemporal graph node, and local feature fusion is performed on the corresponding areas of adjacent frames based on the mask coverage range to constrain the geometric consistency of the joint structure in the stylized output.
It realizes the geometric consistency of joint structure after style transfer, improves the accuracy and robustness of posture recognition, can accurately determine movement abnormalities, and improves the reliability of detection results through multimodal fusion and pre-processing optimization mechanisms.
Smart Images

Figure CN119919458B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a method for video style transfer pose detection based on a graph neural network. Background Art
[0002] At present, with the booming development of the Internet, pictures and videos have been deeply integrated into people's lives. Creative artists and designers actively explore and incorporate artistic elements of different styles into their works, giving rise to the highly regarded field of style transfer, which shows broad prospects in aspects such as film and television special effects and virtual reality. Style transfer can transfer the style features of an image, such as color, texture, composition, etc., to another image, making the latter seem to be re-rendered in the new style of the former.
[0003] However, video data has complex spatio-temporal characteristics and highly non-linear mapping relationships, and traditional machine learning algorithms are difficult to effectively capture and process these characteristics. Moreover, most of the existing style transfer methods only focus on optimizing the visual quality of single frames and lack the modeling of the consistency of continuous actions, resulting in problems such as distorted human poses and sudden joint mutations after style transfer. Traditional solutions rely on optical flow methods or temporal smoothing constraints, but it is difficult to handle the topological structure distortion in complex motion scenes.
[0004] In recent years, graph neural networks (GNNs) have made certain progress in the field of human pose estimation due to their advantages in modeling non-Euclidean data. However, most of the existing graph convolution methods use fixed adjacency matrices and cannot adaptively capture the dynamic changes in joint relationships caused by style transfer, and their robustness is also insufficient in occlusion or fast motion scenes. In addition, pure convolution-based spatio-temporal models are difficult to simultaneously take into account long-range temporal dependencies and local structural features. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for video style transfer pose detection based on a graph neural network, and solve the following technical problems:
[0006] Most of the existing graph convolution methods use fixed adjacency matrices and cannot adaptively capture the dynamic changes in joint relationships caused by style transfer, and their robustness is also insufficient in occlusion or fast motion scenes.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] A method for video style transfer pose detection based on a graph neural network includes the following steps:
[0009] Parse the original video frames frame by frame, extract a continuous video sequence, dynamically adjust the resolution according to the scene and divide it into local regions to generate a preprocessed video data stream;
[0010] Construct a dynamic spatio-temporal correlation network, map the pixels of video frames to spatio-temporal graph nodes, and dynamically calculate the connection weights between nodes according to the motion trends and regional content complexities of adjacent frames. By aggregating the spatio-temporal features of nodes layer by layer, convert the original video frames into target artistic styles;
[0011] During the style transfer process, generate a dynamic mask for the regions where human joint points are located, perform local feature fusion on the corresponding regions of adjacent frames based on the mask coverage, and constrain the geometric consistency of joint structures in the stylized output;
[0012] Extract the joint point coordinates of each frame from the stylized video, construct a spatio-temporal trajectory sequence, learn the motion patterns and spatial topological relationships of joint points through the dynamic spatio-temporal correlation network, and generate feature vectors containing spatio-temporal dependencies;
[0013] Perform multi-dimensional comparison of the feature vectors with a preset action pattern library, calculate the trajectory offset and pattern matching degree, and determine action anomalies according to the threshold and output the detection results.
[0014] As a further solution of the present invention: The construction of the dynamic spatio-temporal correlation network specifically includes:
[0015] Divide the pixels of video frames into several local regions, map each region to a graph node, and the node attributes include color gradient, motion vector, and local texture entropy value, where the motion vector is generated by calculating the optical flow field, and the texture entropy value reflects the detail complexity within the region;
[0016] In the time dimension, establish cross-frame connections for corresponding nodes of adjacent frames, and the connection weights are dynamically calculated by the spatio-temporal correlation of node attributes: If the motion vector directions of adjacent frame nodes are the same and the difference in texture entropy values is less than the adaptive threshold, the weight is increased, otherwise it is decreased;
[0017] In the space dimension, construct hierarchical connections for human joint points within the same frame: Define first-level connections based on anatomical priors, define second-level connections based on motion coordination, and the connection weights between levels are dynamically adjusted according to the joint motion amplitude;
[0018] Through spatio-temporal graph convolution operations, layer by layer fuse the dynamic connection features across frames and across levels to generate spatio-temporal correlation features with both style transfer and action perception.
[0019] As a further solution of the present invention: The generation of the dynamic mask includes:
[0020] Based on the preprocessed video frames, use a lightweight pose estimation network to real-time predict human joint points, generate an initial mask centered on the joint points, and the mask radius linearly expands with the joint motion speed;
[0021] Combined with the optical flow information and the joint trajectory prediction model, infer the deformation parameters of the next-frame mask. The deformation direction is consistent with the joint movement direction, and the deformation amount is proportional to the motion acceleration.
[0022] During the style transfer process, apply double constraints to the mask-covered area. The double constraints include geometric constraints and semantic constraints. The geometric constraints enhance the contour feature extraction of the joints within the mask through a dynamic spatio-temporal correlation network, suppressing the edge blurring caused by stylization. The semantic constraints compare the semantic segmentation results of the mask area before and after style transfer. If key anatomical structures are lost, trigger a local style transfer intensity attenuation.
[0023] When it is detected that the joint point movement speed exceeds the threshold or is occluded, generate a temporary mask by interpolating the mask trajectory of the historical frames, and automatically expand the mask range to cover the motion blur area.
[0024] As a further solution of the present invention: The construction of the spatio-temporal trajectory sequence includes:
[0025] Extract the joint point coordinates from each frame of the stylized video, construct the original trajectory sequence in chronological order, and perform segmented noise reduction processing on the sequence. For the low-speed motion segment, use a moving average filter to smooth the jitter noise. For the high-speed motion segment, retain the original data and only remove the outliers.
[0026] Assign an adaptive time window to each joint point. The window length is dynamically adjusted according to the motion acceleration: the greater the acceleration, the shorter the window length.
[0027] Within the spatio-temporal correlation network, perform multi-scale feature extraction on the trajectory data within the window: fuse the multi-scale features with the hierarchical connection weights to generate a spatio-temporal feature vector containing local details and global patterns.
[0028] As a further solution of the present invention: The construction of the spatio-temporal trajectory sequence further includes:
[0029] Extract visual modality features from the stylized video, including color distribution, edge sharpness, and texture density, and generate a first feature vector through normalization processing.
[0030] Extract motion modality features from the spatio-temporal trajectory sequence, including joint point displacement, acceleration, and motion direction consistency, and generate a second feature vector through dynamic time warping.
[0031] Construct a cross-modal correlation matrix, align the dimensions of the first feature vector and the second feature vector, and fuse the two types of features through a dynamic weight allocation strategy. The weight value is adjusted in real time according to the scene illumination intensity, motion complexity, and style transfer interference degree.
[0032] During the fusion process, if it is detected that the visual modality features are distorted due to style transfer, local feature compensation is performed based on the spatio-temporal continuity of the motion modality features to ensure that the fused feature vector reflects both the visual style and the essence of the action;
[0033] When performing anomaly detection on the fused feature vector, a conflict resolution mechanism is introduced: when the determination results of the two modalities are inconsistent, the detection result of the motion modality is preferentially adopted, and the result is verified by backtracking the trajectory consistency of adjacent frames.
[0034] As a further solution of the present invention: extract the joint point coordinates of each frame and the color distribution of the corresponding pixel region from the stylized video, calculate the motion trajectory offset of the joint points and the color distribution difference of the stylized region, and generate correction parameters;
[0035] Optimize the node connection weights of the dynamic spatio-temporal correlation network according to the correction parameters by gradient descent. For nodes with a trajectory offset exceeding the threshold, reduce their connection weights with adjacent nodes;
[0036] For pixel regions with a color distribution difference exceeding the set range, recalculate their spatio-temporal correlation with the corresponding regions of adjacent frames, and update the feature aggregation rules of local nodes;
[0037] Feed the optimized node connection weights and feature aggregation rules back to the dynamic spatio-temporal correlation network to generate updated style transfer results and spatio-temporal trajectory sequences, and iteratively execute until the trajectory offset and color difference converge to the preset range.
[0038] As a further solution of the present invention: the determination of abnormal actions specifically includes:
[0039] The preset action mode library contains spatio-temporal feature vectors, trajectory offset tolerances, and motion direction constraint conditions of standard actions;
[0040] Align the spatio-temporal feature vectors generated in real time with the features in the mode library dimension by dimension, and calculate the trajectory offset, direction deviation angle, and motion acceleration difference;
[0041] If the trajectory offset exceeds the tolerance, the direction deviation angle is greater than the set angle, or the acceleration difference breaks through the threshold, it is determined as an abnormal action;
[0042] Perform multi-frame continuity verification on the determination results, and only output the final result when anomalies are detected in multiple consecutive frames.
[0043] As a further solution of the present invention: the preprocessing includes:
[0044] Perform dynamic region segmentation on the video frames, divide the inside of the video frames into high-detail regions and low-detail regions according to the content complexity threshold, and perform local super-resolution enhancement on the high-detail regions;
[0045] Extract the optical flow information of video frames, generate a motion vector map, and mark the high-motion regions and static regions; dynamically adjust the preprocessing parameters according to the motion vector map, perform noise reduction and blur suppression on the high-motion regions, and perform color balance optimization on the static regions;
[0046] Synchronously input the preprocessed video data stream and the motion vector map into the dynamic spatio-temporal correlation network to guide the network to focus on and optimize the key regions.
[0047] Advantages of the present invention:
[0048] Through key technical features such as frame-by-frame analysis of the original video frames, construction of a dynamic spatio-temporal correlation network, generation of a dynamic mask, construction of a spatio-temporal trajectory sequence, and action anomaly determination, the present invention effectively solves the problems that traditional methods are difficult to process complex spatio-temporal features of videos, lack modeling of the consistency of continuous actions, cannot adaptively capture dynamic changes in joint relationships, and have insufficient robustness in complex scenarios. It realizes high-precision pose recognition, strong stability and reliability, can maintain the geometric consistency of joint structures after style transfer, accurately determine action anomalies, and further improves the robustness of detection results and algorithm performance through multi-modal fusion, preprocessing, and evaluation optimization mechanisms. It can iteratively improve the collaborative performance of style transfer and anomaly detection and achieve model self-evolution. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The present invention will be further described below with reference to the accompanying drawings.
[0050] Figure 1 is a schematic structural diagram of the ST-GCN, which is the basic network of the present invention for processing human pose detection tasks;
[0051] Figure 2 is a schematic structural diagram of the overall network of the style transfer pose detection algorithm of the present invention;
[0052] Figure 3 is a schematic flowchart of the style transfer network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] Embodiment 1, please refer to Figures 1 - 3 As shown, the present invention is a method for video style transfer pose detection based on a graph neural network, including the following steps:
[0055] 1.1 Parsing of original video frames and extraction of continuous video sequences:
[0056] The original video is parsed frame by frame. Through precise timestamp and frame number identification, each frame of the video is arranged in order according to time. On this basis, a specific algorithm is used to detect scene change points in the video, thereby accurately extracting continuous video sequences. This algorithm can effectively identify significant changes at the moment of scene change by analyzing multi-dimensional information such as color histograms, brightness changes, and feature point matching between adjacent frames, ensuring the integrity and accuracy of continuous video sequences.
[0057] 1.2 Dynamic region segmentation and super-resolution enhancement of video frames:
[0058] For each frame of the video, a method based on content complexity calculation is used for dynamic region segmentation. Specifically, by calculating features such as the gradient magnitude and texture complexity of image blocks and comparing them with a pre-set content complexity threshold. Regions with complexity higher than the threshold are determined as high-detail regions; otherwise, they are low-detail regions.
[0059] For high-detail regions, local super-resolution enhancement technology is applied. This technology is based on deep learning algorithms. By constructing a neural network model containing multiple convolutional layers and deconvolutional layers, feature extraction and reconstruction are performed on image blocks in high-detail regions. The model learns the mapping relationship between a large number of high-resolution images and corresponding low-resolution images during the training process, so that in actual applications, low-resolution image blocks in high-detail regions can be restored to high resolution, significantly improving the detail clarity.
[0060] 1.3 Extraction of optical flow information and marking of motion regions:
[0061] Classic optical flow estimation algorithms, such as the Lucas-Kanade algorithm or the Farneback algorithm, are used to extract the optical flow information of video frames. These algorithms calculate the motion vectors of each pixel point by tracking the motion trajectories of feature points between adjacent frames. Based on the obtained motion vectors, a motion vector map is generated.
[0062] Based on the motion vector map, a motion amplitude threshold is set to divide the video frames into high-motion regions and static regions. Regions with motion vector amplitudes greater than the threshold are marked as high-motion regions, while regions with motion vector amplitudes close to zero are static regions.
[0063] 1.4 Dynamically adjusting preprocessing parameters according to the motion vector map:
[0064] For high-motion regions, due to the easy introduction of noise and blurring during fast movement, targeted noise reduction and blurring suppression algorithms are adopted. In terms of noise reduction, a noise reduction method based on wavelet transform is used. By performing wavelet decomposition on the images in high-motion regions, noise components are removed in different frequency sub-bands, and then wavelet reconstruction is carried out. In terms of blurring suppression, deconvolution technology is utilized. According to the direction and magnitude of the motion vector, inverse operations are performed on the blurred images to restore the clear edges of the images.
[0065] For static regions, the focus is on color balance optimization. By calculating the color histogram of the static region images, the histogram equalization algorithm is used to adjust the color distribution, making the colors of the images more uniform and rich, and enhancing the visual effect.
[0066] 1.5 Synchronously input the preprocessed video data stream and the motion vector map:
[0067] Synchronize the video data stream and the motion vector map obtained through the above series of preprocessing steps to ensure their one-to-one correspondence in the time dimension. Then, input the synchronized information into the dynamic spatio-temporal correlation network. This network can perform in-depth analysis and processing of video content in the spatio-temporal dimension based on the key region information provided by the motion vector map, achieve focused optimization of key regions, and thus improve the accuracy and efficiency of video processing.
[0068] 2.1 Video frame region division and node mapping:
[0069] The pixels of the video frame are finely divided and segmented into numerous appropriately sized local regions. The size of these regions needs to be determined comprehensively considering the complexity of the video content and subsequent calculation efficiency. Usually, the optimal size can be determined through experiments. For example, the video frame can be divided into several small regions of 16x16 pixels. Each local region is mapped to a graph node to build the basis of the graph structure.
[0070] Rich attributes are assigned to each node. The color gradient is obtained by calculating the change rate of pixel colors in the horizontal and vertical directions within the region, which can reflect the changing trend of colors within the region. The motion vector is generated by calculating the optical flow field. Common optical flow estimation algorithms such as the Lucas-Kanade algorithm or the Farneback algorithm can accurately calculate the motion direction and displacement of each pixel point by tracking the motion trajectories of feature points between adjacent frames, and then obtain the motion vector of the region. The texture entropy value is used to reflect the complexity of details within the region and is obtained by calculating the information entropy of the distribution of pixel gray values within the region. The higher the entropy value, the richer the texture details and the more complex the changes within the region.
[0071] 2.2 Cross-frame connection construction and weight calculation:
[0072] In the time dimension, cross-frame connections are established for corresponding nodes in adjacent frames to capture the dynamic changes of the video in the time series. For the corresponding nodes in each pair of adjacent frames, the connection weights are not fixed, but are dynamically calculated from the spatio-temporal correlation of the node attributes.
[0073] In the specific calculation process, first compare the motion vector directions of the nodes in adjacent frames. If the directions are the same, it indicates that the region has a relatively stable motion trend in time. At the same time, compare the texture entropy values of the nodes. If the difference in texture entropy values is less than a threshold adaptively set according to the video content (this threshold can be dynamically adjusted by statistical analysis of the overall texture entropy value of the video combined with the average texture entropy value of the local region), it means that the texture details in this region change less in time and have a high spatio-temporal correlation. In this case, the corresponding cross-frame connection weights will be increased, otherwise decreased. Through this dynamic weight adjustment mechanism, the key information transmission of the video in the time dimension can be more accurately reflected.
[0074] 2.3 Hierarchical connection construction and weight adjustment of human joint points within the same frame:
[0075] In the spatial dimension, hierarchical connections are constructed for human joint points within the same frame. Based on anatomical prior knowledge, the first-level connections are defined as connections between adjacent joints, such as shoulder-elbow, elbow-wrist, etc. These connections can directly reflect the basic composition and motion relationships of the human limb structure.
[0076] The second-level connections are defined based on motion synergy, that is, connections across limb joints, such as left wrist-right ankle, etc. This connection method takes into account the coordinated motion relationships between different limb parts during human movement and can more comprehensively describe the overall motion pattern of the human body.
[0077] The connection weights between levels are dynamically adjusted according to the joint motion amplitude. By tracking the position changes of joint points in the video sequence, the motion amplitude of the joints is calculated. For joint connections with a larger motion amplitude, higher weights are assigned to highlight the important roles and dynamic changes of these joints in human movement. For joint connections with a smaller motion amplitude, the weights are correspondingly reduced, so as to reasonably allocate the importance of different joint connections in the constructed hierarchical connection structure.
[0078] 2.4 Spatio-temporal graph convolution operation and feature generation:
[0079] Using spatio-temporal graph convolution operations, the constructed dynamic spatio-temporal graph structure is processed layer by layer. Spatio-temporal graph convolution can effectively fuse the dynamic connection features across frames and across levels. In each layer of convolution operation, through a specific convolution kernel and the graph nodes and their connection weights for calculation, the information of neighboring nodes is aggregated to the central node, thus realizing the propagation and fusion of features.
[0080] During the convolution process, the cross-frame connection information in the time dimension and the hierarchical connection information in the spatial dimension are fully considered, so that the generated features can not only capture the dynamic changes of human body movements in the video, but also fuse the style features of different regions. After multiple layers of spatio-temporal graph convolution operations, spatio-temporal correlation features with both style transfer and action perception are finally generated. These features can comprehensively and accurately describe the key information of video content in the spatio-temporal dimension, providing a solid data foundation for subsequent applications such as video analysis, action recognition, and style conversion.
[0081] 3.1 Initial mask generation:
[0082] At the initial stage of style transfer, relying on the preprocessed video frame data, a lightweight pose estimation network is used to perform real-time operations. This network has been trained with a large amount of human body pose data and can quickly and accurately predict the positions of human body joint points in the video frame. These joint points serve as the key reference points for subsequent operations.
[0083] Taking each predicted joint point as the center, an initial mask is generated. The shape of the mask is usually set to be circular, and the determination of its radius is closely related to the joint movement speed. Specifically, by tracking and analyzing the positions of joint points in the current frame and previous frames, the displacement of the joint within a unit time is calculated, thereby obtaining the joint movement speed. The mask radius is linearly extended according to the joint movement speed, that is, the faster the joint movement speed, the larger the mask radius expands. For example, a basic radius value r0 is preset. When the joint movement speed is v, the mask radius r = r0 + k * v, where k is a proportionality coefficient that can be determined through experimental optimization to ensure that the mask can reasonably cover the area related to the joint movement around the joint.
[0084] 3.2 Mask deformation parameter inference:
[0085] Combined with the optical flow information and the joint trajectory prediction model, the deformation parameters of the next-frame mask are accurately inferred. The optical flow information is obtained through the optical flow estimation algorithm, which reflects the movement of pixel points in the video frame. The joint trajectory prediction model uses the position information of joint points in historical frames and adopts time series prediction algorithms such as Kalman filtering to predict the positions of joints in future frames.
[0086] The deformation direction of the mask is consistent with the joint movement direction, which is based on the actual situation that joint movement drives changes in the surrounding area. The magnitude of the deformation is proportional to the motion acceleration. By calculating the change in the joint movement speed between adjacent frames, the motion acceleration a is obtained. Let the deformation of the mask in the current frame be Δx1 and the deformation in the next frame be Δx2. If the motion acceleration is a, then Δx2 = Δx1 + m * a, where m is a proportionality factor related to the acceleration and also needs to be optimized through experiments. Such a setting enables the mask to adaptively deform according to the dynamic changes of joint movement, better fitting the real movement state of the joints in the video sequence.
[0087] 3.3 Dual Constraint Mechanism:
[0088] Geometric Constraint: During the style transfer process, a geometric constraint is imposed on the mask-covered area by means of a dynamic spatio-temporal correlation network. The dynamic spatio-temporal correlation network can effectively capture the key information of the video in the spatio-temporal dimension. For the joint area within the mask, the network enhances the clarity of the joint contour during the stylization process by strengthening contour feature extraction. For example, the convolutional layer in the network specifically extracts and amplifies the edge features of the joint contour, suppressing the edge blur phenomenon that may be caused by the stylization operation. Through this geometric constraint, it is ensured that the joints still maintain clear and accurate geometric contours in the visual presentation after style transfer.
[0089] Semantic Constraint: By comparing the semantic segmentation results of the mask area before and after style transfer, semantic constraint is achieved. Semantic segmentation can divide different objects and regions in the video frame according to semantic categories. For the mask area related to human joints, key anatomical structures are focused on, such as the five-finger contours of the hand. If key anatomical structures are found to be missing in the semantic segmentation result after style transfer, the system will automatically trigger a local style intensity attenuation mechanism. The specific approach is to reduce the intensity of the style algorithm within this mask area, so that this area retains the semantic features of the original video frame to a certain extent, avoiding the loss of important human structure information due to over-stylization.
[0090] 3.4 Temporary Mask Generation and Range Adjustment:
[0091] When the detected movement speed of the joint points exceeds a pre-set threshold, it indicates that the joint movement is relatively intense, which may cause problems such as motion blur. Or when the joint points are occluded by other objects, to ensure the effective coverage of the mask for the relevant area, interpolation operations are performed based on the mask trajectory of the historical frames to generate a temporary mask. Interpolation methods can use common algorithms such as linear interpolation or spline interpolation. According to the position and shape information of the historical mask, the shape of the mask that should be in the current frame is deduced. At the same time, the mask range is automatically expanded so that it can cover the areas where motion blur may occur, ensuring that in these special cases, the style transfer operation can still accurately target the joint and its surrounding areas, guaranteeing the coherence and accuracy of the style transfer effect.
[0092] 3.5 Local Feature Fusion and Geometric Consistency Constraint:
[0093] Based on the mask coverage range, local feature fusion is performed on the corresponding areas of adjacent frames. By extracting features such as color, texture, and structure within the corresponding mask areas of adjacent frames, and using feature fusion algorithms, such as weighted average fusion or deep learning-based feature fusion networks, these features are organically combined. During the fusion process, the geometric consistency of the joint structure in the stylized output is strictly constrained. For example, by comparing the positions of joint points and the shapes of joint contours in adjacent frames, the weights and methods of feature fusion are adjusted to ensure that in the stylized video sequence, the transition of the joint structure between adjacent frames is natural and smooth, without sudden changes or unreasonable deformations in geometric shapes, thereby improving the visual quality and realism of the video after style transfer.
[0094] 4.1 Construction and Denoising Processing of Joint Point Trajectory Sequence:
[0095] From each frame of the stylized video, using a high-precision joint point detection algorithm, the coordinates of the joint points are accurately extracted. These algorithms are based on deep learning models and are trained on a large amount of video data with human joint annotations, enabling accurate identification of joint point positions in different scenarios and postures. The extracted joint point coordinates are arranged in chronological order to construct an original trajectory sequence. This sequence completely records the movement path of the joint on the video time axis.
[0096] To improve the quality of the trajectory sequence, segmental denoising processing is performed on it. By analyzing the displacement of joint points within a unit time, the joint movement speed is calculated. For low-speed movement segments, since they may be affected by noise interference and cause jitter in the trajectory, a moving average filtering algorithm is used. This algorithm calculates the average of the data within a fixed-length sliding window to smooth out the jitter noise. For high-speed movement segments, to avoid losing key movement information due to excessive smoothing, only outliers are removed. By calculating the mean and standard deviation of the data, points that deviate from the mean by more than a certain multiple of the standard deviation are regarded as outliers and removed.
[0097] 4.2 Adaptive Time Window Allocation and Multi-scale Feature Extraction:
[0098] An adaptive time window is allocated for each joint point to better capture its features in different motion states. The window length is dynamically adjusted according to the motion acceleration. The greater the acceleration, the faster the change in the joint motion state, and the shorter the required time window to more timely reflect the instantaneous changes in joint motion. The acceleration is obtained by calculating the change in the joint motion speed between adjacent frames. For example, let the joint motion speed in the current frame be v i , and the previous frame be v i−1 , then the acceleration a = v i −v i−1 . Through the pre-set mapping relationship between acceleration and window length, the time window length of each joint point in each frame is determined.
[0099] Within the spatio-temporal correlation network, multi-scale feature extraction is performed on the trajectory data within the time window. Convolution kernels or pooling windows of different sizes are used to perform feature extraction at different levels on the trajectory data. Smaller scales can capture local detailed features, such as the minute movement changes of joints in a short period of time; larger scales focus on global patterns, such as the motion trends of joints over a longer period of time. The multi-scale features are fused with hierarchical connection weights, and the hierarchical connection weights reflect the correlation strength between joint points based on anatomy and motion synergy. Through methods such as weighted summation, the multi-scale features and connection weights are combined to generate a spatio-temporal feature vector containing local details and global patterns, providing rich motion information for subsequent analysis and processing.
[0100] 4.3 Multi-modal Feature Extraction and Fusion:
[0101] Visual Modal Feature Extraction: Visual modal features are extracted from the stylized video, covering color distribution, edge sharpness, and texture density. The color distribution is described by calculating the pixel ratio and histogram of different colors in the video frame, reflecting the overall color characteristics of the video. Edge sharpness uses edge detection algorithms, such as the Sobel operator or Canny operator, to calculate the intensity and clarity of the image edges. Texture density is measured by counting the number and distribution of texture elements in the image. These features are normalized, mapping their values to the range of [0,1] or [-1,1] to generate the first feature vector, ensuring the weight balance of different features in the subsequent fusion process.
[0102] Motion modality feature extraction: Extract motion modality features from the spatio-temporal trajectory sequence, including joint displacement, acceleration, and motion direction consistency. The joint displacement is directly calculated by the difference in joint coordinates between adjacent frames. The acceleration is calculated by the change in velocity as described above. The motion direction consistency is judged by comparing the included angle of joint motion directions between adjacent frames. The smaller the included angle, the higher the consistency. The dynamic time warping algorithm is used to align the motion trajectory sequences of different lengths on the time axis, making the motion features of different video segments comparable, and then generating the second feature vector.
[0103] Feature fusion: Construct a cross-modal correlation matrix, align the dimensions of the first feature vector and the second feature vector, and fuse the two types of features through a dynamic weight allocation strategy. The weight values are adjusted in real time according to the scene illumination intensity, motion complexity, and stylization interference degree. For example, when the scene illumination intensity changes greatly, the weight of the visual modality features is appropriately reduced because illumination may affect the accuracy of visual features such as color and edges. When the motion complexity is high, the weight of the motion modality features is increased to highlight the importance of motion information. During the fusion process, if it is detected that the visual modality features are distorted due to style transfer, such as serious color deviation or excessive edge blurring, local feature compensation is performed based on the spatio-temporal continuity of the motion modality features. By analyzing the change rules of the motion modality features between adjacent frames, reasonable visual features of the affected area are inferred, so as to ensure that the fused feature vector reflects both the visual style and the essence of the action.
[0104] 4.4 Anomaly detection and conflict resolution:
[0105] When performing anomaly detection on the fused feature vector, a conflict resolution mechanism is introduced. When the judgment results based on the visual modality features and the motion modality features are inconsistent, the detection results of the motion modality are preferred. This is because the motion modality features can better reflect the intrinsic nature of human actions and are relatively less affected by external factors such as style transfer. At the same time, the result is verified by backtracking the trajectory consistency of adjacent frames. Check whether the joint trajectories of adjacent frames conform to the normal motion logic and continuity. If there are abnormal jumps or situations that do not conform to the motion law, the reasons are further analyzed and the detection results are corrected to improve the accuracy and reliability of anomaly detection.
[0106] 4.5 Correction parameter generation and network optimization:
[0107] Extract the joint point coordinates of each frame and the color distribution of the corresponding pixel regions from the stylized video, and calculate the offset of the joint point movement trajectory and the color distribution difference of the stylized region. The offset of the movement trajectory is obtained by comparing the joint point trajectory of the current frame with the expected trajectory (such as predicted by the historical trajectory or deduced by the normal movement model). The color distribution difference is measured by calculating the distance between the color histograms of the corresponding pixel regions before and after stylization, such as using the Bhattacharyya distance or the chi-square distance. Generate correction parameters from these calculation results to guide the optimization of the dynamic spatio-temporal correlation network.
[0108] Optimize the node connection weights of the dynamic spatio-temporal correlation network by gradient descent according to the correction parameters. For nodes with trajectory offsets exceeding the threshold, reduce their connection weights with adjacent nodes. This is because a trajectory offset may indicate an abnormal association between the node and its adjacent nodes. By reducing the weight, the impact of abnormal associations on the overall performance of the network is reduced. For pixel regions with color distribution differences exceeding the set range, recalculate their spatio-temporal correlation with the corresponding regions in adjacent frames. Use spatio-temporal correlation algorithms, such as methods based on optical flow and feature matching, to update the feature aggregation rules of local nodes, enabling the network to more accurately handle the impact of color changes during the stylization process.
[0109] Feed the optimized node connection weights and feature aggregation rules back to the dynamic spatio-temporal correlation network to generate updated style transfer results and spatio-temporal trajectory sequences. By continuously iterating the above process, that is, re-extracting features, calculating correction parameters, and optimizing the network, until the trajectory offset and color difference converge to the preset range. At this time, the dynamic spatio-temporal correlation network can more accurately reflect the real situation of human joint movement while ensuring the style transfer effect, improving the quality and reliability of the stylized video.
[0110] 5.1 Construction of the preset action pattern library:
[0111] The construction of the preset action pattern library is the basis of the entire comparison process. A large amount of information related to standard actions is included in this library, and the spatio-temporal feature vectors are extracted through in-depth analysis of videos of various standard actions in different scenarios. Using advanced motion capture devices and professional data analysis software, accurately record the spatio-temporal change information of joint points when the human body performs standard actions, and then through complex algorithm processing, generate spatio-temporal feature vectors that can accurately represent the characteristics of standard actions.
[0112] The trajectory offset tolerance is the allowable deviation range set for the joint point trajectories of each standard action. It is obtained based on the statistical analysis of a large number of standard action samples, taking into account the possible minor differences in actual execution. For example, for the common waving action, after statistically analyzing the trajectory data of 1000 standard waving actions, the trajectory offset tolerance of the arm joint points in a specific direction is determined to be ±5 pixels.
[0113] The motion direction constraints specify the rules that the joint motion direction should follow during the execution of the standard action. By analyzing the standard action video frame by frame, the ideal motion direction of each joint at each stage is determined and expressed as an angle range. For example, in a standard squat, the motion direction of the knee joint during the descending phase should be within an angle of ±15° from the vertical.
[0114] 5.2 Feature vector dimension-by-dimension alignment and difference calculation:
[0115] The spatiotemporal feature vector generated in real time is aligned dimension by dimension with the features in the preset action pattern library. Since different spatiotemporal feature vectors may differ in dimensional order and representation, specific alignment algorithms are required to achieve accurate comparison. These algorithms match and adjust each dimension based on the physical meaning of the feature vector and the action representation logic. For example, for the joint point coordinate dimension, ensure that the order of the joint points in the real-time feature vector is consistent with the order of the joint points of the standard action in the preset library, and the coordinate units and reference system are the same.
[0116] After completing the dimension-by-dimension alignment, the trajectory offset, direction deviation angle, and motion acceleration difference are calculated. The trajectory offset is obtained by calculating the spatial distance between the corresponding joint points of the real-time joint point trajectory and the standard motion trajectory at the same time point.
[0117] The direction deviation angle is obtained by comparing the angle between the real-time joint movement direction and the joint movement direction in the standard action. First, calculate the real-time movement direction vector v real and the standard motion direction vector v std , and then use the vector angle formula to calculate the angle, which is the direction deviation angle.
[0118] The motion acceleration difference is measured by comparing the sum of the absolute values of the differences between the real-time motion acceleration and the standard motion acceleration in each dimension.
[0119] 5.3 Abnormal action determination:
[0120] The calculated trajectory offset, direction deviation angle, and motion acceleration difference are compared with the corresponding tolerance, set angle, and threshold in the preset motion pattern library to determine whether the action is abnormal. If the trajectory offset exceeds the tolerance, it means that the joint motion trajectory deviates from the normal range of the standard action, which may indicate that the action is not performed in a standardized manner or an abnormal situation has occurred. For example, in a standard running action, if the trajectory offset of the leg joint exceeds the preset tolerance of ±10 cm, it may mean that the running posture is incorrect or there is a deformation of the movement caused by physical discomfort.
[0121] When the direction deviation angle is greater than the set angle, it indicates that the joint movement direction does not meet the requirements of the standard action. For example, in the standard punching action of boxing, if the movement direction of the arm deviates from the set angle by more than 30°, it may be that the punching direction is incorrect or the force application method is improper.
[0122] If the acceleration difference breaks through the threshold, it means that there is a significant difference in the acceleration change of the joint movement compared with the standard action. This may reflect problems in the control of the force and speed of the action. For example, in the standard weightlifting action, if the acceleration difference during the lifting of the barbell exceeds the preset threshold, it may be that the athlete applies force unevenly or makes a technical mistake. As long as any one of the above three conditions is met, it is determined as an abnormal action.
[0123] 5.4 Multi-frame Continuity Verification and Final Result Output:
[0124] To improve the accuracy and reliability of abnormal action determination and avoid misjudgment caused by noise or accidental factors in single-frame data, multi-frame continuity verification is performed on the determination results. The system continuously monitors the action determination situation of multiple consecutive frames. Only when abnormalities are detected in multiple consecutive frames (for example, 5 consecutive frames), the final abnormal action result is output. This is because in actual scenarios, single-frame data may be affected by various interferences, such as light changes and occlusions, resulting in short-term feature abnormalities, but it does not necessarily represent a real action abnormality. Through multi-frame continuity verification, these interference factors can be effectively excluded to ensure that the output abnormal action result is more credible. Once abnormalities are detected in multiple consecutive frames, the system will trigger the corresponding alarm or recording mechanism to prompt relevant personnel to pay attention to the abnormal action situation for further analysis and processing.
[0125] The above has described an embodiment of the present invention in detail, but the content described is only a preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the present invention application should still fall within the patent coverage scope of the present invention.
Claims
1. A video style transfer posture detection method based on graph neural network, characterized in that: The following steps are involved: Parse the original video frames frame by frame, extract continuous video sequences, dynamically adjust the resolution according to the scene and divide it into local areas to generate pre-processed video data streams; Construct a dynamic spatiotemporal association network to map the pixels of video frames into spatiotemporal graph nodes. The connection weights between nodes are dynamically calculated based on the motion trends of adjacent frames and the complexity of regional content. The original video frames are converted into the target artistic style by aggregating the spatiotemporal features of nodes layer by layer. In the process of style transfer, a dynamic mask is generated for the area where the human body joints are located, and local features of the corresponding areas of adjacent frames are fused based on the mask coverage to constrain the geometric consistency of the joint structure in the stylized output; Extract the coordinates of the joint points of each frame from the stylized video, construct a spatiotemporal trajectory sequence, learn the motion pattern and spatial topological relationship of the joint points through a dynamic spatiotemporal association network, and generate a feature vector containing spatiotemporal dependencies; The feature vector is compared with the preset action pattern library in multiple dimensions, the trajectory offset and pattern matching degree are calculated, the action abnormality is determined according to the threshold and the detection result is output.
2. According to the method of video style transfer posture detection based on graph neural network in claim 1, it is characterized in that: The construction of the dynamic spatiotemporal association network specifically includes: The pixels of the video frame are divided into several local areas, each of which is mapped to a graph node. The node attributes include color gradient, motion vector and local texture entropy value. The motion vector is generated by optical flow field calculation, and the texture entropy value reflects the complexity of details in the area. In the time dimension, cross-frame connections are established for corresponding nodes in adjacent frames, and the connection weight is dynamically calculated based on the spatiotemporal correlation of node attributes: if the motion vector directions of nodes in adjacent frames are consistent and the texture entropy value difference is less than the adaptive threshold, the weight is increased, otherwise it is decreased; In the spatial dimension, hierarchical connections are constructed for human joints in the same frame: the first-level connection is defined based on anatomical priors, the second-level connection is defined based on motion synergy, and the weights of the connections between the levels are dynamically adjusted according to the range of joint motion. Through the spatiotemporal graph convolution operation, the dynamic connection features across frames and levels are fused layer by layer to generate spatiotemporal correlation features that combine style transfer and motion perception.
3. According to the method of video style transfer posture detection based on graph neural network in claim 1, it is characterized in that: The generation of the dynamic mask includes: Based on the preprocessed video frames, a lightweight posture estimation network is used to predict the human joints in real time, and an initial mask is generated with the joints as the center. The mask radius expands linearly with the joint movement speed. Combining the optical flow information with the joint trajectory prediction model, the deformation parameters of the next frame mask are inferred. The deformation direction is consistent with the joint movement direction, and the deformation amount is proportional to the movement acceleration. In the process of style transfer, the mask coverage area is subject to dual constraints, including geometric constraints and semantic constraints. The geometric constraints strengthen the contour feature extraction of the joints in the mask through a dynamic spatiotemporal association network to suppress the edge blurring caused by stylization; the semantic constraints compare the semantic segmentation results of the mask area before and after style transfer. If the key anatomical structure is lost, the local stylization strength attenuation is triggered. When it is detected that the motion speed of a joint exceeds a threshold or is occluded, a temporary mask is generated based on the mask trajectory interpolation of the historical frames, and the mask range is automatically expanded to cover the motion blur area.
4. According to the method of video style transfer posture detection based on graph neural network in claim 1, it is characterized in that: The construction of the spatiotemporal trajectory sequence includes: Extract the coordinates of joint points from each frame of the stylized video, construct the original trajectory sequence in chronological order, and perform segmented denoising on the sequence. Use sliding average filtering to smooth jitter noise in low-speed motion segments; retain the original data for high-speed motion segments and only remove outliers. An adaptive time window is allocated to each joint point, and the window length is dynamically adjusted according to the motion acceleration: the greater the acceleration, the shorter the window length; In the spatiotemporal correlation network, multi-scale feature extraction is performed on the trajectory data in the window: the multi-scale features are fused with the hierarchical connection weights to generate a spatiotemporal feature vector containing local details and global patterns.
5. According to the method of video style transfer posture detection based on graph neural network in claim 4, it is characterized in that: The construction of the spatiotemporal trajectory sequence also includes: Extracting visual modality features from the stylized video, including color distribution, edge sharpness, and texture density, and generating a first feature vector through normalization; Extract motion modal features from the spatiotemporal trajectory sequence, including joint displacement, acceleration, and motion direction consistency, and generate a second eigenvector through dynamic time warping; Construct a cross-modal association matrix, align the dimensions of the first eigenvector and the second eigenvector, and fuse the two types of features through a dynamic weight allocation strategy. The weight value is adjusted in real time according to the scene lighting intensity, motion complexity, and stylized interference degree; During the fusion process, if it is detected that the visual modality features are distorted due to style transfer, local feature compensation is performed based on the spatiotemporal continuity of the motion modality features to ensure that the fused feature vector reflects both the visual style and the essence of the action. When performing anomaly detection on the fused feature vector, a conflict resolution mechanism is introduced: when the judgment results of the two modes are inconsistent, the detection results of the motion mode are adopted first, and the results are verified by tracing the consistency of the trajectories of adjacent frames.
6. The method for detecting postures in video style transfer based on graph neural network according to claim 1, characterized in that: Extract the coordinates of the joint points of each frame and the color distribution of the corresponding pixel area from the stylized video, calculate the difference between the motion trajectory offset of the joint points and the color distribution of the stylized area, and generate correction parameters; The node connection weights of the dynamic spatiotemporal association network are optimized by gradient descent according to the correction parameters, and the connection weights of the nodes with trajectory offsets exceeding the threshold value and the adjacent nodes are reduced; For pixel areas where the color distribution difference exceeds the set range, recalculate its spatiotemporal correlation with the corresponding areas of adjacent frames and update the feature aggregation rules of local nodes; The optimized node connection weights and feature aggregation rules are fed back to the dynamic spatiotemporal association network to generate updated style migration results and spatiotemporal trajectory sequences, which are iterated until the trajectory offset and color difference converge to the preset range.
7. The method for detecting postures in video style transfer based on graph neural network according to claim 1, characterized in that: The determination of abnormal action specifically includes: The preset action pattern library contains the spatiotemporal feature vectors, trajectory deviation tolerances, and motion direction constraints of standard actions; Align the spatiotemporal feature vector generated in real time with the features in the pattern library dimension by dimension, and calculate the trajectory offset, direction deviation angle, and motion acceleration difference; If the trajectory deviation exceeds the tolerance, the direction deviation angle is greater than the set angle, or the acceleration difference exceeds the threshold, it is judged as an abnormal action; The judgment result is verified for multi-frame continuity, and the final result is output only when anomalies are detected in multiple consecutive frames.
8. The method for detecting postures in video style transfer based on graph neural network according to claim 1, characterized in that: The pre-processing comprises: Perform dynamic region segmentation on the video frame, divide the interior of the video frame into high-detail area and low-detail area according to the content complexity threshold, and perform local super-resolution enhancement on the high-detail area; Extract the optical flow information of the video frame, generate a motion vector, and mark the high-motion area and the static area; dynamically adjust the preprocessing parameters according to the motion vector, perform noise reduction and blur suppression on the high-motion area, and perform color balance optimization on the static area; The preprocessed video data stream and the motion vector graph are synchronously input into the dynamic spatiotemporal association network to guide the network to optimize the focus on key areas.
Citation Information
Patent Citations
Face pose migration method and device based on video driving
CN113808005A
Real-time video quality optimization and enhancement method based on deep learning
CN119418254A