Hierarchical feature modeling method and system for motion gesture recognition
By analyzing human motion image sequences, detecting dynamic changes in interference and adjusting the perspective, identifying node timing and calculating trajectory compensation paths, the problem of adaptability and accuracy in complex scenarios for motion recognition in existing technologies is solved, achieving higher recognition accuracy and environmental adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to distinguish between environmental changes and target motion-related feature interference in dynamic scenarios such as background disturbances and perspective changes. This results in insufficient response at different action stages, difficulty in effectively compensating for spatial trajectories, and limited feature grouping and hierarchical representation capabilities, thus reducing the accuracy and adaptability of recognition in complex scenarios.
By analyzing human motion image sequences, detecting dynamic changes in interference, adjusting the perspective of the acquisition device, identifying node temporal trigger sequences, calculating spatial trajectory compensation paths, and constructing a hierarchical feature expression structure, including disturbance fluctuations, edge stability indicators, posture and perspective adjustment parameters, node temporal trigger sequences, and spatial trajectory compensation results.
It improves the hierarchical representation capability of motion capture, enhances adaptability and recognition accuracy in complex environments, can actively filter background disturbances, dynamically adapt to motion direction, separate core motion nodes, and improve node response association and grouped angle and temporal structure.
Smart Images

Figure CN121305683B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of posture recognition technology, and in particular to a hierarchical feature modeling method and system for motion posture recognition. Background Technology
[0002] The field of posture recognition involves perceiving, analyzing, and judging the posture of humans, animals, or objects in space to identify and track their actions. This technical field includes image acquisition, feature extraction, motion modeling, and classification. Its research objects include the changes in static postures and dynamic movements, and it is mainly applied in various scenarios such as human-computer interaction, intelligent monitoring, virtual reality, sports training, and medical rehabilitation. Traditional hierarchical feature modeling methods for motion posture recognition involve constructing a deep neural network model, using acquired temporal image data to detect and extract key human points, then encoding spatiotemporal features using a convolutional neural network, and finally using a recurrent neural network or long short-term memory network to model and classify the action sequence, thereby achieving the recognition of various actions during movement.
[0003] Existing technologies struggle to distinguish between environmental changes and feature interference caused by target motion in dynamic scenarios such as background disturbances and perspective changes. During data acquisition, the accuracy of the input is easily affected by the fixed perspective, the response of nodes in different stages of the action is insufficient, the spatial trajectory is difficult to compensate effectively, and the feature grouping is too simple and lacks hierarchical expression. In practical applications, this can easily lead to problems such as incomplete recognition of complex action details, disordered node response order, and limited feature extraction granularity, thereby reducing the overall system's adaptability and ability to distinguish complex scenarios. Summary of the Invention
[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide a hierarchical feature modeling method and system for motion posture recognition. The technical solution is as follows:
[0005] On the one hand, a hierarchical feature modeling method for motion pose recognition is provided, including the following steps:
[0006] S1: Based on human motion image sequences, analyze the gray-scale change area outside the subject boundary, determine the gradient direction of edge pixels frame by frame, compare the temporal changes of the edge and the center, and detect frames with reversed direction to obtain the dynamic change characteristics of interference.
[0007] S2: Based on the dynamic change characteristics of the interference, analyze the spatial coordinates of the skeleton center within the associated time period, determine the direction of the skeleton trajectory in consecutive frames, compare the trend of the angle between the motion direction and the axis of the acquisition device, adjust the viewing angle of the acquisition device, and obtain the attitude viewing angle adjustment parameters.
[0008] S3: Based on the posture and viewpoint adjustment parameters, determine the changes in spatial coordinates and color attributes of skeleton node pixels in consecutive frames, analyze the differences in node position and color, identify nodes with change amplitude, and obtain the node timing trigger sequence.
[0009] S4: Based on the node timing trigger sequence, determine the spatial trajectory relationship between the delayed response node and the activated neighboring node, compare the trajectory direction and trend, calculate the dynamic compensation path, and obtain the spatial trajectory compensation result.
[0010] S5: Based on the spatial trajectory compensation results, analyze the change sequence of joint angle pairs between consecutive frames, determine the angle change trend, identify hierarchical change feature segments, and obtain the hierarchical feature expression structure.
[0011] On the other hand, the dynamic change features of the interference include disturbance fluctuation amount, edge stability index, and inter-frame change level; the attitude and view adjustment parameters include deflection angle parameters, target field of view interval, and acquisition consistency reference; the node temporal trigger sequence includes activation node index, stage marker information, and response temporal data; the spatial trajectory compensation result includes spatial position correction amount, compensation trajectory parameters, and node dynamic mapping; and the hierarchical feature expression structure includes hierarchical angle features, intra-group statistical parameters, and temporal grouping identifier.
[0012] On the other hand, the specific steps for obtaining the dynamic change characteristics of the interference are as follows:
[0013] S101: Based on the human motion image sequence, analyze the subject boundary of each frame, identify the gray-level change area outside the boundary, determine the gray-level gradient direction of each edge pixel in space, summarize the changes in the gradient direction of edge pixels in consecutive frames, aggregate the gradient direction trends between frames, and obtain the contour gradient sequence.
[0014] S102: Based on the contour gradient sequence, compare the corresponding pixel density state of the central region, analyze the changing characteristics of the pixel distribution in the central region of each frame, integrate the pixel density trend between consecutive frames of the central region, and obtain the regional dynamic sequence.
[0015] S103: Based on the dynamic sequence of the region, determine the direction of change between adjacent frames, identify the key frame number where the direction of change changes, and obtain the dynamic change characteristics of the interference.
[0016] On the other hand, the steps for obtaining the attitude view adjustment parameters are as follows:
[0017] S201: Based on the dynamic change characteristics of the interference, analyze the spatial coordinates of the skeleton center of each frame in the corresponding time period, determine the motion direction of the skeleton center point of adjacent frames, integrate the continuity relationship of motion paths between frames, and obtain motion trajectory data.
[0018] S202: Based on the motion trajectory data, calculate the spatial angle between the motion direction of each frame and the main axis direction of the acquisition device, compare the change process of the angle of each frame with the image sequence, identify the frame number whose angle fluctuation range is higher than the stable segment of the motion trajectory, and obtain the viewpoint offset segment.
[0019] S203: Based on the aforementioned viewpoint offset segment, determine the spatial correspondence between the skeleton center trajectory direction and the main axis of the acquisition device, summarize the offset change pattern of each frame within the interval, aggregate the orientation adjustment parameters of each interval, and obtain the attitude viewpoint adjustment parameters.
[0020] On the other hand, the steps for obtaining the node timing trigger sequence are as follows:
[0021] S301: Based on the posture and viewpoint adjustment parameters, analyze the spatial coordinates of the skeleton node pixels in consecutive frames, determine the trend of spatial coordinate changes of each node in the image sequence, and combine the changes in the color attributes of the corresponding pixels of each node to identify the skeleton nodes with spatial or color changes and obtain the node response group.
[0022] S302: Based on the node response group, determine the response time of each skeleton node in the time series, compare the order in which the response states of each skeleton node occur, and organize the response timing of the skeleton nodes in consecutive frames to obtain the node timing chain.
[0023] S303: Based on the node timing chain, compare the correspondence with the skeleton action stage division, determine the motion stage to which each node response time belongs, establish stage-by-stage triggering logic, and obtain the node timing trigger sequence.
[0024] On the other hand, the steps for obtaining the spatial trajectory compensation result are as follows:
[0025] S401: Based on the node timing trigger sequence, analyze the skeleton nodes activated in each sub-stage, determine the spatial trajectory connection between the delayed response node and its activated neighboring nodes, compare the change characteristics of the trajectory direction of adjacent nodes, and summarize the change trend of the spatial trajectory in each action stage to obtain the stage trajectory feature group.
[0026] S402: Based on the stage trajectory feature group, filter out the skeleton nodes with spatial position deviations in each motion stage, determine their offset magnitude in the spatial trajectory, and combine the trajectory changes of activated neighboring nodes to determine the spatial distribution and trajectory of abnormal nodes, thereby obtaining an abnormal node mapping set.
[0027] S403: Based on the abnormal node mapping set, determine the spatial relationship between the motion trajectory of each offset node and its neighboring nodes, analyze the dynamic characteristics of trajectory connection and action sequence, calculate the dynamic compensation trajectory of the offset node in the phased motion process, and obtain the spatial trajectory compensation result.
[0028] On the other hand, the specific steps for obtaining the hierarchical feature representation structure are as follows:
[0029] S501: Based on the spatial trajectory compensation results, analyze the change process of the angle pairs of skeleton joints between consecutive frames in the time series, determine the amplitude trend of the angle pairs in each time period, identify the segments with fluctuation characteristics of angle changes, and obtain the set of angle fluctuation segments.
[0030] S502: Based on the set of angle fluctuation segments, filter feature segments with hierarchical change characteristics, compare the differences in angle distribution and change trend of each level of segments, determine the angle feature performance of each level, and obtain grouped angle feature groups.
[0031] S503: Based on the grouped angle feature groups, optimize the grouping method and expression content of the angle features, adjust the arrangement and encoding of the angle sequences within each group, aggregate the statistical parameters and time series labels associated with the hierarchical features, and obtain the hierarchical feature expression structure.
[0032] On the other hand, the subject boundary refers to the outer edge contour of the human body, hand, and head targets detected by deep learning or image segmentation methods in a motion image sequence, and the grayscale change region refers to the part of the image where the brightness of a pixel changes in space or time.
[0033] On the other hand, the spatial coordinates of the skeleton center refer to the three-dimensional or two-dimensional coordinate sequence of key points of the human body center point and torso center point obtained by OpenPose in continuous frames, and the angle trend refers to the curve of the change of the angle between the skeleton movement direction and the main axis direction of the acquisition device over time.
[0034] On the other hand, a hierarchical feature modeling system for motion pose recognition is provided, including:
[0035] The feature perturbation extraction module is based on human motion image sequences. It analyzes the gray-scale change area outside the subject boundary, judges the gradient direction of edge pixels frame by frame, compares the temporal changes of the edge and the center, and detects frames with reversed direction to obtain the dynamic change features of the perturbation.
[0036] Based on the dynamic change characteristics of the interference, the perspective adaptive module analyzes the spatial coordinates of the skeleton center within the associated time period, determines the direction of the skeleton trajectory in consecutive frames, compares the trend of the angle between the motion direction and the axis of the acquisition device, adjusts the perspective of the acquisition device, and obtains the attitude perspective adjustment parameters.
[0037] The node dynamic detection module determines the changes in spatial coordinates and color attributes of skeleton node pixels in consecutive frames based on the posture and viewpoint adjustment parameters, analyzes the differences in node position and color, identifies nodes with changes in amplitude, and obtains the node temporal trigger sequence.
[0038] The trajectory compensation inference module determines the spatial trajectory relationship between the delayed response node and the activated neighboring node based on the node timing trigger sequence, compares the trajectory direction and trend, calculates the dynamic compensation path, and obtains the spatial trajectory compensation result.
[0039] Based on the spatial trajectory compensation results, the hierarchical feature modeling module analyzes the change sequence of joint angle pairs between consecutive frames, judges the angle change trend, identifies hierarchical change feature segments, and obtains the hierarchical feature expression structure.
[0040] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0041] By detecting dynamic changes in interference, and combining temporal perspective adjustments and phased node responses, the compensation reasoning of spatial trajectories is strengthened, enhancing the hierarchical representation capability of motion capture. This enables input features to proactively filter background disturbances, dynamically adapt to motion direction, separate core action nodes, improve node response correlation, and group angle temporal structure. This forms an adaptive processing capability for changing scenes and complex node states during motion posture recognition, ensuring the integrity and recognizability of action feature structure representation and hierarchical modeling, resulting in higher recognition accuracy and adaptability to complex environments. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This is a flowchart of the main steps of the present invention;
[0044] Figure 2 This is a flowchart of steps S1 of the present invention;
[0045] Figure 3 This is a flowchart of steps S2 of the present invention;
[0046] Figure 4 This is a flowchart of steps S3 of the present invention;
[0047] Figure 5 This is a flowchart of step S4 of the present invention;
[0048] Figure 6 This is a flowchart of steps S5 of the present invention;
[0049] Figure 7 This is a system block diagram of the present invention. Detailed Implementation
[0050] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0051] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0052] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0053] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0054] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0055] This invention provides a hierarchical feature modeling method for motion pose recognition, such as... Figure 1 As shown, it includes the following steps:
[0056] S1: Based on human motion image sequences, analyze the gray-scale change area outside the subject boundary, determine the gradient direction of edge pixels frame by frame, combine the pixel density trend of the central region, compare the changes of the edge and central regions in the time series, and simultaneously detect the direction reversal frames to obtain the dynamic change characteristics of interference.
[0057] S2: Based on the dynamic change characteristics of interference, analyze the spatial coordinate sequence of the skeleton center within the time period. By judging the trajectory direction of the skeleton center point in consecutive frames, compare the trend of the angle between the skeleton movement direction and the axis of the acquisition device, identify the acquisition angle that is inconsistent with the current movement direction, adjust the acquisition device angle, and obtain the attitude angle adjustment parameters.
[0058] S3: Based on the attitude viewpoint adjustment parameters, determine the changes in the spatial coordinates and color attributes of skeleton node pixels in consecutive frames, analyze the spatial position and color differences of skeleton nodes in the time series, identify skeleton nodes with large response change amplitude, and establish motion phased triggering logic by comparing the order of skeleton node response times to obtain node time sequence triggering sequence.
[0059] S4: Based on the node timing trigger sequence, analyze the determined stage trigger nodes, determine the spatial trajectory relationship between the delayed response node and its activated neighboring nodes, identify nodes with large spatial position deviations within the motion segment by comparing trajectory directions and trends, and calculate the dynamic compensation path of the node by combining the trajectories of neighboring nodes to obtain the spatial trajectory compensation result.
[0060] S5: Based on the spatial trajectory compensation results, analyze the time change sequence of joint angle pairs between consecutive frames, determine the angle change trend, select feature segments with hierarchical changes, compare the performance of angle features at each level, optimize feature grouping and expression form, and obtain the hierarchical feature expression structure.
[0061] The dynamic characteristics of interference include disturbance fluctuation, edge stability index, and inter-frame change classification. The attitude and view adjustment parameters include deflection angle parameters, target field of view interval, and acquisition consistency reference. The node temporal trigger sequence includes activation node index, stage marker information, and response temporal data. The spatial trajectory compensation results include spatial position correction amount, compensation trajectory parameters, and node dynamic mapping. The hierarchical feature expression structure includes hierarchical angle features, intra-group statistical parameters, and temporal grouping identifier.
[0062] In S1, the subject boundary refers to the external edge contour of targets such as the human body, hands, and head, detected by deep learning or image segmentation methods in a sequence of moving images. It is used to distinguish the human body from the background. The grayscale change region refers to the part of the image where the brightness (grayscale value) of a pixel changes significantly in space or time. It is often used to identify object contours or environmental changes. The edge pixel gradient direction refers to the direction of change of pixel grayscale in space in pixels at or near the boundary. It can be obtained using operators such as Sobel and Canny, reflecting the trend of the edge. The pixel density trend refers to the trajectory of the change of the number of pixels, brightness, or density of pixels in each frame over time in the central region. It is used to describe regional dynamics and environmental changes. The direction reversal frame refers to the key frame where the change direction (e.g., from rising to falling) is detected when analyzing pixel gradient or density trends. It is often used to capture dynamic turning points or environmental transition moments.
[0063] In S2, the time period refers to the human motion segment or time range that needs further analysis, determined by background disturbance trend detection. The spatial coordinate sequence of the skeleton center refers to the three-dimensional or two-dimensional coordinate sequence of key points such as the human body center point and torso center point in continuous frames, obtained by deep learning skeleton detection algorithms (such as OpenPose), reflecting the main motion trajectory of the human body. The skeleton motion direction refers to the main direction of the skeleton center or the whole body's motion between continuous frames, which can be obtained by calculating the vector changes of continuous points. The acquisition device axis refers to the main axis direction of the viewpoint of image acquisition devices such as cameras and sensors, usually the spatial direction directly in front of the camera. The angle trend refers to the curve of the angle between the skeleton motion direction and the acquisition device main axis direction over time, reflecting the degree of alignment between the moving subject and the camera.
[0064] In S3, skeleton node pixels refer to the pixel coordinates of key human body nodes (such as wrists, knees, and tops of the head) in the image output by deep learning skeleton detection. Color attribute changes refer to the changes in color or brightness information of the corresponding pixels of skeleton nodes between consecutive frames, which are often used to help determine node state changes or occlusions. Response change amplitude refers to the significant change amplitude of node spatial position or color attributes within a short period of time, which can be used to identify important action nodes such as sudden stops and force exertion. The order of response times refers to the order in which each node is activated (reaching a specific change amplitude) during the action process, which helps to model the action in stages. The staged triggering logic of the motion refers to dividing the motion action into different stages according to the node response time and amplitude, and establishing the logical relationship of node activation in batches or time periods.
[0065] In S4, phased trigger nodes refer to the group of skeleton nodes analyzed in the previous phase that are activated sequentially in different phases of the action. Response delay nodes refer to nodes that respond (undergo critical changes) later than other nodes in the same phase, often related to action coordination and compensation. Spatial trajectory relationships refer to the movement paths and relative positional changes between response delay nodes and their spatially adjacent, activated nodes during the action process. Movement segments refer to dividing the continuous movement process into several time intervals, each corresponding to a phase of the action, used for more detailed analysis and modeling. Nodes with large spatial positional deviations refer to skeleton nodes whose actual positions deviate significantly from the movement trend or the trajectories of neighboring nodes within the current movement segment. Neighboring node trajectories refer to the temporal movement trajectories of activated skeleton nodes that are spatially close, often used to assist in estimating the positions of inactive nodes. Dynamic compensation paths refer to the reasonable movement path compensation results calculated for response delay or abnormal nodes based on spatial trajectory relationships and the movement trends of neighboring nodes.
[0066] In S5, joint angle pairs refer to the angle pairs formed by two nodes and joints (such as knee-hip-ankle) in the human skeleton, used to represent action posture features. Temporal change sequence refers to the change sequence of joint angle pairs in consecutive image frames, reflecting the angular dynamics during the action. Angle amplitude trend refers to analyzing the fluctuation or growth / decrease trend of the angle sequence on the time axis, often used to capture key action features. Hierarchical change feature segments refer to dividing the sequence into different change levels according to the angle change amplitude or fluctuation, with each level corresponding to the main or secondary change stage of the movement. Angle feature representation refers to the distribution, change form and statistical characteristics of joint angles within each level feature segment, reflecting the main dynamic information of different action stages. Feature grouping and expression form refers to the feature recombination and encoding of the angle change sequence within different levels, providing hierarchical and distinguishable input features for subsequent deep learning models.
[0067] like Figure 2 As shown, the specific steps for obtaining the dynamic change characteristics of interference are as follows:
[0068] S101: Based on the human motion image sequence, analyze the subject boundary of each frame, identify the gray-level change area outside the boundary, determine the gray-level gradient direction of each edge pixel in space, summarize the changes in the gradient direction of edge pixels in consecutive frames, aggregate the gradient direction trends between frames, and obtain the contour gradient sequence.
[0069] Subject recognition of human targets is required for each frame of the image. This process involves traversing all pixels in the image and extracting suspected human areas based on the distribution density and spatial continuity of pixel brightness values across the entire image. Large connected regions with grayscale values in the middle brightness range are identified as candidate subject regions. When performing edge detection on these regions, the brightness changes between adjacent pixels are calculated in both the horizontal and vertical directions. By comparing this change value with a preset standard deviation of brightness change point by point, it is determined whether a pixel is an edge pixel. If the brightness change trend of 1 to 3 pixels outside the edge shows significant increases or decreases, then that pixel is recorded as an edge pixel. The gradient direction of each edge pixel is summarized with angle information to form the edge gradient vector in a single frame image. Then, the gradient direction of edge pixels is compared pairwise in consecutive image frames to calculate the change in the direction angle of overlapping pixels in adjacent frames. When the change in direction is higher than a predetermined angle standard, the pixel is marked as having changed direction. The number and distribution trend of changed pixels in the whole frame are counted. This type of change information in each frame is integrated in the time dimension to generate a temporal set representing the dynamic change trend of the edge, which is used to reflect the significant gradient trend of the human body boundary during the movement process, and obtain the contour gradient sequence.
[0070] S102: Based on the contour gradient sequence, compare the corresponding pixel density state of the central region, analyze the changing characteristics of the pixel distribution in the central region of each frame, integrate the pixel density trend between consecutive frames of the central region, and obtain the regional dynamic sequence.
[0071] Further analysis is performed on the central region of each frame image. This central region is obtained by shrinking the subject boundary contour obtained in the previous step inward in both the horizontal and vertical directions at fixed proportions. During the analysis phase, the number of pixels with brightness values within the target density standard range in the central region is counted first. The effective pixel ratio in the unit region is calculated as the density index of the current frame. This process is repeated for all frames to form a set of continuous pixel density records. Then, the change amplitude of this index between any two adjacent frames is compared. If the change amplitude exceeds the preset standard, the frame is considered to have a density change. By counting the number and position of frames with increasing and decreasing density in each stage, the overall change trend is recorded in the time dimension. Then, every ten frames are used as an analysis window to observe the density change fluctuation characteristics within the window. The time position and intensity of significant increases or decreases are marked to form a dynamic change trend set based on time. This set is compared with the previously obtained contour gradient sequence on the time coordinate to confirm whether the two trends co-varied at the same time. Frames with co-variation are marked and stored uniformly. All central region density change trend information is integrated and its temporal evolution relationship is established to form a regional dynamic sequence.
[0072] S103: Based on the regional dynamic sequence, determine the direction of change between adjacent frames, identify the key frame number where the direction of change changes in the direction of change, and obtain the dynamic change characteristics of interference.
[0073] Using the acquired regional dynamic sequence as input data, the pixel density change trend between adjacent frames is compared step by step to determine whether the change direction of the current frame is consistent with that of the next frame. If the change direction changes from increasing to decreasing, or from decreasing to increasing, it is identified as a potential change inflection point. Further confirmation is needed to determine whether the change has reached a significant level. Only when the change magnitude of the current frame and the next frame both exceed the dense change judgment standard is the frame number recorded as a key frame candidate. The number set of candidate key frame numbers is formed by traversing the entire sequence. Then, the contour gradient sequence generated in the previous stage is called to synchronously check the frames marked as candidates. If the proportion of edge pixels recorded as having directional changes in the frame to the total number of edge pixels is higher than a specific judgment standard, the frame is confirmed as a key frame. Its number is bound with the central region dynamic change data and the edge direction change data to form joint feature data. The output is a disturbance dynamic change feature with time position index and multi-source change index.
[0074] like Figure 3As shown, the specific steps for obtaining attitude and viewpoint adjustment parameters are as follows:
[0075] S201: Based on the dynamic change characteristics of interference, analyze the spatial coordinates of the skeleton center of each frame in the corresponding time period, determine the motion direction of the skeleton center point of adjacent frames, integrate the continuity relationship of motion paths between frames, and obtain motion trajectory data.
[0076] Spatial coordinates of the skeleton center are extracted from each frame within a time period. These skeleton center coordinates are calculated from the convolutional results within a deep learning network after locating keypoints within each frame of the image sequence. These coordinates are stored as horizontal and vertical pixel value pairs in a two-dimensional plane, forming a coordinate sequence on the time axis. This coordinate sequence is then processed frame-by-frame to extract the difference between the horizontal and vertical coordinates of consecutive frames. Based on this difference sequence, a motion direction vector for each frame is constructed. In this process, the motion direction of each frame is generated by calculating the direction of the vector endpoint, and the direction value is normalized to 0–3 degrees in angle form. Within a 60-degree range, it is determined whether the angle difference of directional changes in three consecutive frames remains within 15 degrees. If it does, the frame segment is marked as a segment with consistent direction; otherwise, it is marked as a turning point. All motion directions are spliced together in chronological order to form a complete skeleton motion direction sequence. The continuous change pattern of motion direction in spatial angle is further detected. Combined with the node confidence output, coordinate points with confidence scores below 0.6 are removed. For missing frames, the spatial trajectory segment is reconstructed using front-to-back interpolation. The direction vector, position point, missing frame mark, and direction duration corresponding to each frame are output and summarized into continuous skeleton motion path information to obtain motion trajectory data.
[0077] S202: Based on motion trajectory data, calculate the spatial angle between the motion direction of each frame and the main axis direction of the acquisition device, compare the change process of the angle of each frame with the image sequence, identify the frame number whose angle fluctuation range is higher than the stable segment of the motion trajectory, and obtain the viewpoint offset segment.
[0078] The orientation vector of the skeleton center is extracted from each frame, and this vector is spatially compared with the main axis direction of the image acquisition device's field of view. The field of view direction of the acquisition device is calibrated as the vertical axis of the image during network initialization, i.e., the direction of the vertical center line of the image frame. The angle formed between the skeleton center vector in each frame and this axis direction is measured. A sequence of angle values from all frames is formed, and the sequence is segmented into sliding windows, with each window set to 5 frames wide. The maximum and minimum difference between the angle values is calculated as the fluctuation amplitude. If the angle fluctuation amplitude within a window exceeds 15 degrees, the frame number corresponding to that window is marked as... For unstable viewing angle segments, all unstable segments are merged and judged. If the frame interval between two consecutive segments is less than 3 frames, they are merged into one segment; otherwise, they are considered independent segments. For each segment, the difference between its average fluctuation value and the average value of its adjacent stable segments is further judged. If the difference exceeds 10 degrees, the segment is marked as a high fluctuation segment. At the same time, the viewing angle offset confidence parameter in the deep learning output is called to compensate for frames with a confidence score lower than 0.7. The average angle between two adjacent frames is used as a substitute frame value to supplement the insertion. All time periods in which viewing angle jumps or discontinuous adjustments occur are sorted out to obtain the viewing angle offset segments.
[0079] S203: Based on the viewpoint offset segment, determine the spatial correspondence between the skeleton center trajectory direction and the main axis of the acquisition device, summarize the offset change law of each frame in the interval, aggregate the orientation adjustment parameters of each interval, and obtain the attitude viewpoint adjustment parameters.
[0080] The angle between the skeleton motion direction vector and the main axis direction of the acquisition device in each frame is extracted. The variation of the angle value over time is analyzed according to the frame sequence. The rate of change of the angle in each frame is recorded as the difference between the angle values of the two adjacent frames. If the rate of change is consistent in direction and its absolute value is greater than 10 degrees in three consecutive frames, it is determined to be a continuous offset phase. Linear fitting is performed on the angle values in this segment, and its slope is extracted as the orientation offset rate parameter. This is then compared with the offset direction of the skeleton center position in the current frame. If the direction of the angle change rate matches the skeleton movement direction, the offset is confirmed. Since the main axis field of view is inconsistent, the node heatmap output by deep learning in the image frame is combined to extract the strong response area of the skeleton backbone region in the heatmap, calculate the offset value between its center and the image center, normalize the offset value and match it with the angle change, calculate the initial orientation offset angle of the starting frame, the maximum orientation deviation angle of the median frame, and the convergence angle of the ending frame for each view offset segment, statistically analyze the difference trend among the three, and record indicators such as the offset slope and the frame number of the peak change point. The orientation trend and spatial change indicators of each segment are integrated and output to obtain the attitude view adjustment parameters.
[0081] like Figure 4 As shown, the specific steps for obtaining the node timing trigger sequence are as follows:
[0082] S301: Based on the attitude and viewpoint adjustment parameters, analyze the spatial coordinates of the skeleton node pixels in consecutive frames, determine the trend of spatial coordinate changes of each node in the image sequence, and combine the changes in the color attributes of the corresponding pixels of each node to identify the skeleton nodes with spatial or color changes and obtain the node response group.
[0083] First, the two-dimensional pixel coordinate data of all nodes of the skeleton in the continuous image frames output by the deep learning module under the corrected viewpoint are read. This data forms a node position sequence with the node number as the index. Each node has a unique position coordinate in each frame. The position difference of each node is calculated frame by frame to determine the pixel movement value of each node in the horizontal and vertical directions between adjacent frames. If the inter-frame displacement in any direction exceeds 15 pixels for two consecutive frames, the node is marked as a node whose spatial position has changed. At the same time, the node color channel response image in the deep learning output is called to record the brightness of the red, green and blue channels of the pixel area where the node is located, and the brightness change is calculated in the two consecutive frames. If the brightness difference of any channel exceeds 20, the node is marked as having changed in color attribute. The spatial change mark and the color change mark are then merged. If a node meets the change criterion in either position or color, it is tentatively designated as a candidate response node. Then, based on the node confidence parameter output by deep learning, nodes with a confidence score below 0.5 are excluded to avoid false detections. The remaining nodes are used as the initial screening set of response nodes. All response nodes are aggregated by frame number. The number of response nodes in each frame is counted. If the number of response nodes exceeds 30% of the total number of skeleton nodes, the frame is considered a high response frame of the skeleton. All marked nodes in the frame are numbered, recorded, and output as a node response group.
[0084] S302: Based on the node response group, determine the response time of each skeleton node in the time series, compare the order of occurrence of the response states of each skeleton node, and organize the response timing of the skeleton nodes in consecutive frames to obtain the node timing chain.
[0085] The system retrieves the position records of response nodes in the image sequence and analyzes the image frame number where each node was first marked as a response. This number is used as the response start frame for that node. The response start frames of all nodes are sorted to construct a response timeline sequence. Next, the node response order is differentially evaluated. If the response frames of any two nodes differ by less than 2 frames and their positional distance in image space is less than 40 pixels, the two nodes are grouped into the same response cluster. All response cluster numbers are arranged chronologically to construct an aggregated time chain of node responses. Finally, the system analyzes the node number, start frame, and average response timeline within each response cluster. The response intervals are recorded to form a temporal sub-chain within the cluster. The number of frames from the start frame to the end frame of the response for each node is further counted. If the number of frames is less than 3, the response behavior is considered a short-term response and recorded as a low-stability node. If the duration exceeds 5 frames and the response state appears continuously, it is recorded as a high-stability node. Then, the nodes are sorted a second time according to their stability level and response start time. The number, response frame sequence, spatial coordinate sequence and stability of all response nodes are classified and integrated to form a node response chain with a continuous and temporal structure. The output is a node temporal chain.
[0086] S303: Based on the node timing chain, compare the correspondence with the skeleton motion stage division, determine the motion stage to which each node response time belongs, establish stage-by-stage triggering logic, and obtain the node timing trigger sequence;
[0087] The response start frame of each node is compared with the time segmentation structure of the entire action cycle, dividing the action cycle into four segments: preparation, initiation, exertion, and recovery. Assuming the entire image sequence is 40 frames long, the interval is divided into stages of 10 frames each. Node response frames are mapped to these stage intervals, and the number and frame sequence of the first responding node in each stage are recorded. If the number of responding nodes in a certain stage exceeds 25% of the total number of nodes, the stage is marked as a high-density response stage, and a node activation table is established according to the response order. This table records the stage number and the activated nodes. The nodes are grouped into a synchronous response group if there are three or more nodes with a response frame difference of no more than one frame and the spatial distance between the nodes is less than 30 pixels. The timing structure of multiple synchronous response groups and independent response nodes is extracted in each stage. The difference in the starting frame of the response group is then analyzed. The minimum value of the activation frame number of each group is used as the stage trigger point. The trigger points of each stage are combined with the node activation table to generate the node timing structure. An activation sequence index array is generated for each stage. Finally, the node timing trigger sequence is output.
[0088] like Figure 5 As shown, the specific steps for obtaining the spatial trajectory compensation result are as follows:
[0089] S401: Based on the node timing trigger sequence, analyze the skeleton nodes activated in each sub-stage, determine the spatial trajectory connection between the delayed response node and its activated neighboring nodes, compare the change characteristics of the trajectory direction of adjacent nodes, and summarize the change trend of the spatial trajectory in each action stage to obtain the stage trajectory feature group.
[0090] Using the skeleton nodes identified as active in each stage as the analysis object, the spatial coordinate sequence of each response node in each stage in consecutive image frames is first extracted. This coordinate data represents the image position of the node in each frame. Then, the response time of the nodes in each stage is compared again, and nodes with later response frame numbers are screened out and defined as delayed response nodes. Next, other nodes that are spatially adjacent to the delayed node are found. The judgment criterion is that the Euclidean distance between the nodes in the image coordinates is less than a set spatial proximity threshold of 50 pixels. The adjacent node is used as the activated control node. The coordinate sequences of these two nodes in the 5 frames before and after the response are extracted, and their trajectory direction vector in the image coordinate system is calculated. The vector direction between each frame is then used as the basis for the analysis. The change amount serves as a trajectory change characteristic indicator. By comparing the time series of trajectory direction differences between two nodes, if the change in direction difference exceeds 45 degrees within a certain time window and lasts for no less than 3 frames, it is determined that the delayed node has a trajectory direction deviation from its neighboring activated nodes. By recording such deviation relationships in each stage segment by segment, the movement trends of nodes with similar directions in each response stage are summarized. The trend direction value and the node coordinate change amount are paired for analysis to extract the node pairs that show consistent changes in the trajectory path. Then, the total number of trajectory direction change trends and the average direction change rate in each stage are counted. The node trajectory deviation, the number of consistent trajectory directions, and the trajectory direction change rate in each stage are combined and output to form a stage trajectory feature group.
[0091] S402: Based on the stage trajectory feature group, filter the skeleton nodes with spatial position deviations in each motion stage, determine their offset magnitude in the spatial trajectory, combine the trajectory changes of activated neighboring nodes, determine the spatial distribution and trajectory of abnormal nodes, and obtain the abnormal node mapping set.
[0092] Spatial position deviation analysis is performed on all skeleton nodes in each action phase. Specifically, the shortest spatial distance between each node's position in its response frame and the center trajectory curve of the node with the same trajectory direction in the corresponding phase is extracted. If this distance is greater than the deviation identification threshold, i.e., the maximum allowable error value of the reference trajectory is set to 20 pixels, then the node is determined to have a position deviation. Then, the movement path of the offset node in the frames before and after the response is checked to confirm whether there are continuous directional jumps in a short period of time. If the directional change angle in consecutive frames exceeds 30 degrees, and the node's movement amplitude is less than 10 pixels between two frames, then the node is marked as a node with inconsistent movement amplitude. The point is then compared with the trajectory of its spatial offset data and that of its neighboring nodes. The average path of the neighboring trajectory is taken as the compensation reference trajectory by the 5-frame sliding window. The directional similarity between the trajectories is used to judge. If the directional deviation rate of the node trajectory exceeds 20% and the distance error is greater than 25 pixels, it is confirmed as an abnormal node. The node number, stage information, trajectory direction difference and spatial offset value are recorded and organized into an initial set of abnormal nodes. Then, each node in the set is re-matched with its corresponding neighboring trajectory points. The node combination with a matching error of less than 10 pixels is confirmed to be valid. All valid abnormal nodes and their corresponding neighboring trajectory information are combined to form an abnormal node mapping set.
[0093] S403: Based on the abnormal node mapping set, determine the spatial relationship between the motion trajectory of each offset node and its neighboring nodes, analyze the dynamic characteristics of trajectory connection and action sequence, deduce the dynamic compensation trajectory of the offset node in the phased motion process, and obtain the spatial trajectory compensation result.
[0094] The dynamic characteristics of trajectory relationships and action sequences are analyzed using the following formula:
[0095] ;
[0096] The trajectory difference parameters are calculated, and the dynamic compensation trajectory of the offset node during the phased motion process is deduced to obtain the spatial trajectory compensation result. Representative trajectory difference parameters, This represents the total number of frames analyzed during the phased motion process. Representing the Spatial distance between frame offset nodes and neighboring nodes Representing the Spatial distance between frame offset nodes and neighboring nodes The weighting coefficients represent the trajectory changes. Representing the The amount of projected displacement change between the frame offset node and its neighboring nodes in the direction of motion. Representing the The amount of projected displacement change between the frame offset node and its neighboring nodes in the direction of motion.
[0097] The trajectory difference parameter refers to the weighted average of the changes in spatial distance between the offset node and its neighboring nodes and the changes in projected displacement in the direction of motion during the phased motion process. It is used to comprehensively reflect the spatial motion trajectory connection, change trend and local trajectory anomaly characteristics between the offset node and its neighboring nodes in consecutive frames. The formula incorporates the changes in spatial distance and the changes in projected displacement in the direction of motion into the same index. Through weighted superposition, it quantifies the trajectory deviation and dynamic coordination between nodes, which is convenient for dynamic compensation analysis in subsequent action phases.
[0098] Obtain the 3D spatial coordinate information of the offset node and its neighboring nodes in consecutive frames during the phased motion process, and calculate the frame number respectively. and spatial distance and The Euclidean distance formula is used for calculation, where the coordinates of the offset node in frame 1 are... The coordinates of the neighboring nodes are The calculation yields:
[0099] ;
[0100] In frame 2, respectively and ,have to:
[0101] 6;
[0102] In frame 3 and ,have to:
[0103] ;
[0104] In frame 4 and ,have to:
[0105] ;
[0106] Subsequently, the node was analyzed along the main motion direction. Projected displacement change Frames 1 to 4 are calculated using vector dot product. , , , According to the formula, set the weighting coefficients. Calculate the original difference terms: frames 1 to 2. , ,have to:
[0107] ;
[0108] Frames 2 to 3, , ,have to:
[0109] ;
[0110] Frames 3 to 4, , ,have to:
[0111] ;
[0112] Sum the three sets of results and divide by the number of frames. The calculated original trajectory difference parameters are as follows:
[0113] ;
[0114] right and Perform min-max normalization on each, where:
[0115] , , , ;
[0116] The normalized parameter values are:
[0117] ;
[0118] , , , ;
[0119] Based on this, the normalized difference is calculated as follows:
[0120] Frames 1 to 2, , Calculation items:
[0121] ;
[0122] Frames 2 to 3, , Calculation items:
[0123] ;
[0124] Frames 3 to 4, , Calculation items:
[0125] ;
[0126] Sum the differences in the three segments and then average them to obtain the normalized trajectory difference parameter:
[0127] ;
[0128] Trajectory difference parameters The value range is divided into three levels:
[0129] when ≤ < When the trajectory of the offset node is highly consistent with that of its neighboring nodes, and the changes in spatial distance and directional projection are stable, such nodes can be considered to have good trajectory coordination and do not enter the compensation process.
[0130] when ≤ < When this occurs, it indicates that the node trajectory has deviated to a certain extent. This deviation is a transient change or a local disturbance. It needs to be combined with other dynamic parameters for multi-dimensional comprehensive judgment to determine whether to trigger the compensation path calculation.
[0131] when If the node trajectory deviates significantly, it is determined that the condition for triggering the compensation mechanism has been met, and the next stage of spatial trajectory compensation path calculation process must be initiated immediately.
[0132] Received Falling into the second level indicates that the trajectory of the offset node has a slight deviation trend. Although it has not reached the clear compensation threshold, it is recommended to conduct further evaluation in combination with other dynamic parameters. If necessary, it can be predicted to enter the compensation process.
[0133] like Figure 6 As shown, the specific steps for obtaining the hierarchical feature representation structure are as follows:
[0134] S501: Based on the spatial trajectory compensation results, analyze the change process of the angle pairs of skeleton joints between consecutive frames in the time series, determine the amplitude trend of the angle pairs in each time period, identify the segments with fluctuation characteristics of angle change, and obtain the set of angle fluctuation segments.
[0135] The spatial coordinates of the compensated skeleton key nodes in consecutive image frames are used to construct a sequence of joint angle pairs. Each angle pair structure consisting of three key nodes is selected, such as shoulder-elbow-wrist or hip-knee-ankle. All frame data are traversed, and the included angle value of the angle pair in each frame is calculated. A time series is formed by indexing the frame number. The angle time series is divided into segments with a sliding window of every ten frames. The difference between the maximum and minimum values in each window is extracted as the angle amplitude of the segment. All amplitude data are analyzed segment by segment. If the amplitude value of a segment is greater than the angle change amplitude threshold of 25 degrees, it is marked as an angle fluctuation segment. The number of angle rising and falling frames in the segment is further recorded. By comparing whether the number of rising and falling frames is approximately equal, it is determined whether the segment has repeated fluctuation characteristics. If the difference does not exceed two frames and there are more than three direction reversals, it is defined as a fluctuation segment. Its start frame, end frame, angle extreme value pair and average change frequency are recorded. This process is repeated for all angle pairs and all frame windows. All segments that meet the fluctuation judgment conditions are summarized as the angle fluctuation segment set.
[0136] S502: Based on the set of angle fluctuation segments, select feature segments with hierarchical change characteristics, compare the differences in angle distribution and change trend of segments at each level, determine the angle feature performance of each level, and obtain grouped angle feature groups.
[0137] The angle change trend characteristics within each fluctuation segment are analyzed segment by segment. For each segment, the number and length of its angle rising and falling segments are counted, and the ratio of its rising rate to falling rate is calculated. If the ratio is greater than 1.5, it is classified as an asymmetric fluctuation segment; if the ratio is between 0.8 and 1.2, it is classified as a symmetric fluctuation segment. Then, based on the amplitude, all fluctuation segments are divided into three levels: amplitude greater than 45 degrees is the first-level feature, medium amplitude between 30 and 45 degrees is the second-level feature, and amplitude less than 30 degrees is the third-level feature. For all segments within each level, the mean angle, standard deviation, rising and falling frequency, extreme value time, and other indicators are further statistically analyzed. The indicators are averaged within the group to obtain the angle characteristic performance of that level. Then, the difference in the average values of statistical parameters between different levels is compared. For example, if the average standard deviation of the first-level feature segment is greater than 1.8 times that of the second-level feature segment, it indicates that the angle change of the first-level level is more drastic. Through this difference indicator, the angle behavior between different levels is distinguished, and the specific performance characteristics of the angle features of each level are output, forming grouped angle feature groups.
[0138] S503: Based on the grouped angle feature groups, optimize the grouping method and expression content of angle features, adjust the arrangement and encoding of angle sequences within each group, aggregate the statistical parameters and time series labels associated with hierarchical features, and obtain the hierarchical feature expression structure;
[0139] The time series angles within each level are optimized in sequence and their content is adjusted. First, within each level, the extreme value positions and amplitude trends of all angle sequences are extracted. The extreme value pairs are rearranged so that sequences within the same group have similar structural features on the time axis. For example, the starting frame is unified as the starting point of the rising segment, and the extreme values are aligned to the middle frame position. Then, the adjusted sequence is encoded into an angle trend label sequence. The label content consists of rising, falling, and stable states. Subsequently, the average frame length of the rising segment, the average frame length of the falling segment, and the distribution ratio of the stable segment in each group are statistically analyzed to construct a hierarchical statistical parameter table. This parameter table should also include key statistical indicators such as the average amplitude value of each level, the proportion of amplitude direction, and the frequency of extreme value fluctuations. Then, time series labels are added according to the indicators, such as high amplitude high frequency type, medium amplitude low frequency type, and symmetrical stable type. Labels are also added to each feature group for further classification and recognition. The output is a hierarchical feature expression structure.
[0140] like Figure 7 As shown, a hierarchical feature modeling system for motion pose recognition includes:
[0141] The feature perturbation extraction module is based on human motion image sequences. It analyzes the gray-scale change area outside the subject boundary, judges the gradient direction of edge pixels frame by frame, compares the temporal changes of the edge and the center, and detects frames with reversed direction to obtain the dynamic change features of the perturbation.
[0142] The perspective adaptation module analyzes the spatial coordinates of the skeleton center within the associated time period based on the dynamic change characteristics of interference, determines the direction of the skeleton trajectory in consecutive frames, compares the trend of the angle between the motion direction and the axis of the acquisition device, adjusts the perspective of the acquisition device, and obtains the attitude perspective adjustment parameters.
[0143] The node dynamic detection module adjusts the parameters based on the pose and viewpoint to determine the changes in the spatial coordinates and color attributes of skeleton node pixels in consecutive frames, analyzes the differences in node position and color, identifies nodes with changes in amplitude, and obtains the node temporal trigger sequence.
[0144] The trajectory compensation inference module is based on the node timing trigger sequence to determine the spatial trajectory relationship between the delayed response node and the activated neighboring node, compare the trajectory direction and trend, calculate the dynamic compensation path, and obtain the spatial trajectory compensation result.
[0145] The hierarchical feature modeling module analyzes the change sequence of joint angle pairs between consecutive frames based on the spatial trajectory compensation results, judges the angle change trend, identifies hierarchical change feature segments, and obtains the hierarchical feature expression structure.
[0146] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A hierarchical feature modeling method for motion pose recognition, characterized in that, The method includes: S1: Based on human motion image sequences, analyze the gray-scale change area outside the subject boundary, determine the gradient direction of edge pixels frame by frame, compare the temporal changes of the edge and the center, and detect frames with reversed direction to obtain the dynamic change characteristics of interference. S2: Based on the dynamic change characteristics of the interference, analyze the spatial coordinates of the skeleton center within the associated time period, determine the direction of the skeleton trajectory in consecutive frames, compare the trend of the angle between the motion direction and the axis of the acquisition device, adjust the viewing angle of the acquisition device, and obtain the attitude viewing angle adjustment parameters. The specific steps for obtaining the attitude view adjustment parameters are as follows: S201: Based on the dynamic change characteristics of the interference, analyze the spatial coordinates of the skeleton center of each frame in the corresponding time period, determine the motion direction of the skeleton center point of adjacent frames, integrate the continuity relationship of motion paths between frames, and obtain motion trajectory data. S202: Based on the motion trajectory data, calculate the spatial angle between the motion direction of each frame and the main axis direction of the acquisition device, compare the change process of the angle of each frame with the image sequence, identify the frame number whose angle fluctuation range is higher than the stable segment of the motion trajectory, and obtain the viewpoint offset segment. S203: Based on the aforementioned view offset segment, determine the spatial correspondence between the skeleton center trajectory direction and the main axis of the acquisition device, summarize the offset change pattern of each frame within the interval, aggregate the orientation adjustment parameters of each interval, and obtain the attitude view adjustment parameters. S3: Based on the posture and viewpoint adjustment parameters, determine the changes in spatial coordinates and color attributes of skeleton node pixels in consecutive frames, analyze the differences in node position and color, identify nodes with change amplitude, and obtain the node timing trigger sequence. The specific steps for obtaining the node timing trigger sequence are as follows: S301: Based on the posture and viewpoint adjustment parameters, analyze the spatial coordinates of the skeleton node pixels in consecutive frames, determine the trend of spatial coordinate changes of each node in the image sequence, and combine the changes in the color attributes of the corresponding pixels of each node to identify the skeleton nodes with spatial or color changes and obtain the node response group. S302: Based on the node response group, determine the response time of each skeleton node in the time series, compare the order in which the response states of each skeleton node occur, and organize the response timing of the skeleton nodes in consecutive frames to obtain the node timing chain. S303: Based on the node timing chain, compare the correspondence with the skeleton motion stage division, determine the motion stage to which each node response time belongs, establish stage-by-stage triggering logic, and obtain the node timing trigger sequence; S4: Based on the node timing trigger sequence, determine the spatial trajectory relationship between the delayed response node and the activated neighboring node, compare the trajectory direction and trend, calculate the dynamic compensation path, and obtain the spatial trajectory compensation result. S5: Based on the spatial trajectory compensation results, analyze the change sequence of joint point angle pairs between consecutive frames, determine the angle change trend, identify hierarchical change feature segments, and obtain the hierarchical feature expression structure. The dynamic change features of the interference include disturbance fluctuation amount, edge stability index, and inter-frame change level. The attitude and view adjustment parameters include deflection angle parameters, target field of view interval, and acquisition consistency reference. The node temporal trigger sequence includes activation node index, stage marker information, and response temporal data. The spatial trajectory compensation result includes spatial position correction amount, compensation trajectory parameters, and node dynamic mapping. The hierarchical feature expression structure includes hierarchical angle features, intra-group statistical parameters, and temporal grouping identifier.
2. The hierarchical feature modeling method for motion pose recognition according to claim 1, characterized in that, The specific steps for obtaining the dynamic change characteristics of the interference are as follows: S101: Based on the human motion image sequence, analyze the subject boundary of each frame, identify the gray-level change area outside the boundary, determine the gray-level gradient direction of each edge pixel in space, summarize the changes in the gradient direction of edge pixels in consecutive frames, aggregate the gradient direction trends between frames, and obtain the contour gradient sequence. S102: Based on the contour gradient sequence, compare the corresponding pixel density state of the central region, analyze the changing characteristics of the pixel distribution in the central region of each frame, integrate the pixel density trend between consecutive frames of the central region, and obtain the regional dynamic sequence. S103: Based on the dynamic sequence of the region, determine the direction of change between adjacent frames, identify the key frame number where the direction of change changes, and obtain the dynamic change characteristics of the interference.
3. The hierarchical feature modeling method for motion pose recognition according to claim 1, characterized in that, The specific steps for obtaining the spatial trajectory compensation result are as follows: S401: Based on the node timing trigger sequence, analyze the skeleton nodes activated in each sub-stage, determine the spatial trajectory connection between the delayed response node and its activated neighboring nodes, compare the change characteristics of the trajectory direction of adjacent nodes, and summarize the change trend of the spatial trajectory in each action stage to obtain the stage trajectory feature group. S402: Based on the stage trajectory feature group, filter out the skeleton nodes with spatial position deviations in each motion stage, determine their offset magnitude in the spatial trajectory, and combine the trajectory changes of activated neighboring nodes to determine the spatial distribution and trajectory of abnormal nodes, thereby obtaining an abnormal node mapping set. S403: Based on the abnormal node mapping set, determine the spatial relationship between the motion trajectory of each offset node and its neighboring nodes, analyze the dynamic characteristics of trajectory connection and action sequence, calculate the dynamic compensation trajectory of the offset node in the phased motion process, and obtain the spatial trajectory compensation result.
4. The hierarchical feature modeling method for motion pose recognition according to claim 1, characterized in that, The specific steps for obtaining the hierarchical feature representation structure are as follows: S501: Based on the spatial trajectory compensation results, analyze the change process of the angle pairs of skeleton joints between consecutive frames in the time series, determine the amplitude trend of the angle pairs in each time period, identify the segments with fluctuation characteristics of angle changes, and obtain the set of angle fluctuation segments. S502: Based on the set of angle fluctuation segments, filter feature segments with hierarchical change characteristics, compare the differences in angle distribution and change trend of each level of segments, determine the angle feature performance of each level, and obtain grouped angle feature groups. S503: Based on the grouped angle feature groups, optimize the grouping method and expression content of the angle features, adjust the arrangement and encoding of the angle sequences within each group, aggregate the statistical parameters and time series labels associated with the hierarchical features, and obtain the hierarchical feature expression structure.
5. The hierarchical feature modeling method for motion pose recognition according to claim 1, characterized in that, The subject boundary refers to the outer edge contour of the human body, hand, and head targets detected by deep learning or image segmentation methods in a sequence of moving images. The grayscale change region refers to the part of the image where the brightness of a pixel changes in space or time.
6. The hierarchical feature modeling method for motion pose recognition according to claim 1, characterized in that, The spatial coordinates of the skeleton center refer to the three-dimensional or two-dimensional coordinate sequence of key points of the human body center point and torso center point obtained through OpenPose in consecutive frames, and the angle trend refers to the curve of the change of the angle between the skeleton movement direction and the main axis direction of the acquisition device over time.
7. A hierarchical feature modeling system for motion posture recognition, said system being used to implement the hierarchical feature modeling method for motion posture recognition as described in any one of claims 1-6, characterized in that, The system includes: The feature perturbation extraction module is based on human motion image sequences. It analyzes the gray-scale change area outside the subject boundary, judges the gradient direction of edge pixels frame by frame, compares the temporal changes of the edge and the center, and detects frames with reversed direction to obtain the dynamic change features of the perturbation. Based on the dynamic change characteristics of the interference, the perspective adaptive module analyzes the spatial coordinates of the skeleton center within the associated time period, determines the direction of the skeleton trajectory in consecutive frames, compares the trend of the angle between the motion direction and the axis of the acquisition device, adjusts the perspective of the acquisition device, and obtains the attitude perspective adjustment parameters. The node dynamic detection module determines the changes in spatial coordinates and color attributes of skeleton node pixels in consecutive frames based on the posture and viewpoint adjustment parameters, analyzes the differences in node position and color, identifies nodes with changes in amplitude, and obtains the node temporal trigger sequence. The trajectory compensation inference module determines the spatial trajectory relationship between the delayed response node and the activated neighboring node based on the node timing trigger sequence, compares the trajectory direction and trend, calculates the dynamic compensation path, and obtains the spatial trajectory compensation result. Based on the spatial trajectory compensation results, the hierarchical feature modeling module analyzes the change sequence of joint angle pairs between consecutive frames, judges the angle change trend, identifies hierarchical change feature segments, and obtains the hierarchical feature expression structure.
Citation Information
Patent Citations
Human body skeleton motion sequence behavior identification method
CN106203363A
Human body behavior recognition method based on bone joint point data
CN111914798A