An AI-based motion gesture recognition system
By using an AI-based motion posture recognition system, which utilizes key point extraction and temporal skeleton graph structure, the system solves the problem of poor flexibility in traditional systems and achieves accurate recognition and efficient matching of complex movements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-07
AI Technical Summary
Traditional motion posture recognition systems rely on sensors or specific devices with fixed layouts, which are inflexible, difficult to adapt to diverse scenarios, and lack recognition accuracy in complex environments, making it impossible to achieve high-frequency motion capture and fine posture analysis.
An AI-based motion posture recognition system is adopted. The system detects the human skeleton through a key point extraction module, generates a multi-node temporal coordinate set, constructs a temporal skeleton graph structure, extracts node trajectory features, analyzes the action process in segments, and performs matching and recognition in combination with standard posture templates.
It achieves accurate parsing of complex movements without relying on external hardware, improving the accuracy and timeliness of posture recognition and adapting to complex scenarios and dynamic environments.
Smart Images

Figure CN121564446B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of action recognition, in particular to an AI-based motion gesture recognition system. BACKGROUND
[0002] The technical field of action recognition refers to a technical field of analyzing, recognizing and understanding the actions and postures of humans or objects through computer vision, artificial intelligence and related technical means. The technical field includes the application of core tasks such as image and video analysis, feature extraction, pattern recognition, machine learning and deep learning. Action recognition technology is widely used in intelligent monitoring, health monitoring, sports training, virtual reality and other fields. The core task is to extract key frame information from video or image, classify and recognize the action through algorithm model, and realize real-time monitoring and analysis of human or object behavior. Among them, the traditional motion gesture recognition system refers to a system for recognizing human motion state based on sensors or computer vision technology. The traditional system obtains motion data through wearing sensors, optical tracking, depth cameras and other devices, and recognizes gestures by using traditional image processing technology and feature extraction method. The system relies on fixed sensor layout or complex device installation, resulting in poor flexibility and being limited by hardware devices, which cannot perform large-scale real-time data processing. Through multi-point sensor or visual data input, the traditional motion gesture recognition system can capture and analyze human motion, but it still has certain limitations in dealing with complex scenes or dynamic environments.
[0003] The traditional motion gesture recognition system relies on fixed layout sensors or specific types of camera devices, and has problems such as strong hardware dependence and complex installation in the process of motion data acquisition. The system is limited by the flexibility of the device in actual deployment, and is difficult to adapt to diversified scenes, resulting in insufficient recognition accuracy in dynamic environments or large areas. Traditional methods extract features through static image processing or rule setting, lack of system modeling of action continuity and change trend, resulting in action recognition delay, misjudgment and inability to effectively capture complex posture conversion. The overall processing efficiency and adaptability of the system are limited, and it is difficult to meet the demand of high-frequency motion capture and fine gesture analysis. SUMMARY
[0004] The purpose of the present application is to solve the problems existing in the prior art, and to provide an AI-based motion gesture recognition system.
[0005] In order to achieve the above purpose, the application adopts the following technical scheme, an AI-based motion gesture recognition system comprises:
[0006] The key point extraction module obtains video frame data in a motion scene, detects a human skeleton in each frame of image, records pixel coordinates of multiple joints of a human body in a two-dimensional space, and generates a multi-node time sequence coordinate set;
[0007] The graph structure construction module extracts a connection relationship between joints in each frame of image based on the multi-node time sequence coordinate set, identifies adjacent joint pairs and calculates a spatial distance and an included angle value, sequentially arranges a whole set of graph structures in a time dimension according to node numbers, and generates a time sequence skeleton graph structure sequence;
[0008] The node trajectory modeling module extracts a coordinate change amplitude, a speed change interval and an edge included angle change sequence of a node in consecutive graph frames according to the time sequence skeleton graph structure sequence, compares a coordinate difference value trend and an angle change direction of the same node in a differential time period, and generates a key node trajectory feature group;
[0009] The action segmentation analysis module extracts extreme value frames, inflection point frames and still frames of a key node in a time sequence trajectory based on the key node trajectory feature group, judges a type of an action link through a frame sequence interval, divides an action process into three categories of a starting segment, a conversion segment and a termination segment, and obtains a segmented dominant node label set.
[0010] As a further scheme of the application, the multi-node time sequence coordinate set includes a node time sequence number, a spatial coordinate sequence and a node corresponding frame sequence index, the time sequence skeleton graph structure sequence includes a node connection topology, an edge attribute set and a graph sequence number, the key node trajectory feature group includes a trajectory change mode, a node spatial distribution and a connection structure feature, and the segmented dominant node label set includes an action stage identifier, a dominant node number and a key frame sequence range.
[0011] As a further scheme of the application, the key point extraction module includes:
[0012] The image frame acquisition submodule obtains video frame data in a motion scene, numbers the video frame data by using a frame sequence number and a timestamp, decodes a pixel array in the video frame data, extracts a pixel value matrix in an RGB channel, combines in a time sequence, and obtains a continuous frame image matrix set;
[0013] The skeleton coordinate recording submodule detects a target human body region bounding box position in each frame of image according to the continuous frame image matrix set, identifies a two-dimensional coordinate position of a human body key point in the bounding box, projects and converts the two-dimensional coordinates, constructs a coordinate vector of the key point in a two-dimensional space, and generates a multi-frame skeleton three-dimensional coordinate set;
[0014] The node sequence construction submodule calls the multi-frame skeleton three-dimensional coordinate set, matches key nodes with the same number in adjacent frames, identifies the spatial position change trend, calculates the displacement similarity value between nodes, constructs a continuous spatial trajectory structure on the time axis, and obtains a multi-node temporal coordinate set.
[0015] As a further aspect of the present invention, the graph structure construction module includes:
[0016] The graph node connection extraction submodule extracts the coordinate positions of joints in each frame of the image based on the multi-node temporal coordinate set, identifies the positional relationship in the image, determines whether a connection is formed according to the node number index, filters adjacent joint pairs that satisfy the spatial proximity constraint, and records the number combination and connection direction attribute to generate a set of adjacent joint connection pairs.
[0017] The spatial feature calculation submodule calls the set of adjacent joint connection pairs, obtains the three-dimensional spatial coordinate difference of each pair of connection joints, calculates the weighted connection strength value of each pair of connection edges, which represents the degree of structural connection and directional perturbation response strength between two nodes in a multi-frame image sequence, and obtains a set of weighted connection strengths.
[0018] The timing recording submodule extracts the connection strength value and direction attribute between each pair of nodes based on the weighted connection strength set, and combines it with the corresponding frame sequence information to arrange the connection relationships in the video frames in order of node number, assembling them frame by frame into a sequence of node numbers and edge attributes with timing identifiers, and obtaining a timing skeleton graph structure sequence.
[0019] As a further aspect of the present invention, the weighted connection strength value is expressed by the formula:
[0020] ;
[0021] in, Indicates joint and The weighted connection strength value, Indicates joint and Euclidean distance in a single frame image Indicates the first In-frame joints and The path perturbation length of the directional offset. Indicates the first Frame orientation difference influencing factor Indicates the first Frame dynamic inertia response factor This indicates the total number of frames.
[0022] As a further aspect of the present invention, the node trajectory modeling module includes:
[0023] The node coordinate change extraction submodule obtains two-dimensional space coordinate information of nodes in continuous frames based on the time sequence skeleton graph structure sequence, calls coordinate data of nodes with the same number in adjacent frames, calculates coordinate difference values and continuous difference trend directions in the time sequence, records coordinate difference sequences of each node between continuous frames according to frame sequence index values, and generates a node time sequence coordinate difference value sequence;
[0024] The angle change calculation submodule calls the node time sequence coordinate difference value sequence, calculates speed values of nodes between multiple frames according to a time interval constant between adjacent frames and the coordinate difference values, obtains a speed sequence composed of the speed values and defines a speed change interval according to a valley speed and a peak speed, calculates an included angle change amount between direction vectors of connected edges according to coordinates of nodes at two ends of the connected edges in adjacent frames, and generates a speed change interval and an edge included angle change sequence;
[0025] The trajectory feature recognition submodule judges speed change directions and included angle change trends of nodes in continuous frames according to the speed change interval and the edge included angle change sequence, identifies node frame sequence indexes with speed increase and decrease trend reversals or included angle change mutations, calls node coordinate positions and connected edge numbers in frames corresponding to the node frame sequence indexes, records coordinate difference values and connected edge number change situations of the nodes between the previous and subsequent frames, and generates a key node trajectory feature group.
[0026] As a further scheme of the application, the action segmentation analysis module comprises:
[0027] The key frame identification submodule detects frame indexes with zero speed change in the node trajectory sequence as static frames based on the key node trajectory feature group, determines extreme value frames according to positions of maximum values and minimum values in the coordinate difference value sequence, determines inflection point frames according to frame points with sign changes in the edge included angle sequence, and generates a node key frame index set.
[0028] The action segment division submodule calculates frame sequence interval values between adjacent key frames according to the node key frame index set, performs interval matching judgment according to the frame interval and set starting segment threshold values, transition segment threshold values and ending segment threshold values, labels action segment attribute categories in which the key frames are located, and generates an action segment category sequence.
[0029] The dominant node label extraction submodule calls key frame indexes corresponding to action segment in the action segment category sequence, extracts key node number information under the corresponding frames, counts node index values in each type of action segment, screens node index values with connected edge numbers exceeding average values, and generates a segmented dominant node label set.
[0030] As a further scheme of the application, the system further comprises a posture recognition output module:
[0031] The posture recognition output module retrieves a recorded standard posture template according to the segmented dominant node label set, finds action posture identification data corresponding to the node combination in the standard posture template, calls the action posture identification data to compare with the position of the real-time node under the image, filters the standard posture template number meeting the matching condition, and sorts according to the frequency of occurrence to generate a motion posture matching label sequence.
[0032] The motion posture matching label sequence includes a matching template number, a posture recognition label and a frequency of occurrence.
[0033] As a further scheme of the present application, the posture recognition output module comprises:
[0034] The posture template matching sub-module retrieves a node combination index recorded in the standard posture template library according to the node number combination in the segmented dominant node label set, calls action posture identification data under the associated node combination, performs corresponding frame matching comparison with the position coordinates of the corresponding node in the real-time image frame, performs filtering according to the Euclidean distance error threshold and the angle change tolerance standard between nodes, and generates a set of filtered standard posture template numbers.
[0035] The label matching sub-module calls the set of filtered standard posture template numbers, accumulatively counts the number of occurrences of the standard posture template number, sorts according to the number frequency from high to low, assigns a corresponding posture label value and records in the corresponding frame position in the time sequence, constructs a posture label time index correspondence relationship under the node sequence, and generates a motion posture matching label sequence.
[0036] Compared with the prior art, the present application has the advantages and positive effects that:
[0037] In the present application, the spatial positions of each node of the human body in the time sequence structure in the continuous image are extracted, and a graph structure sequence based on the topological relationship and motion trend between nodes is established, so that dynamic modeling and structural expression of the skeleton motion can be realized. By analyzing the trajectory features such as node position change, speed fluctuation and direction change, key paragraphs with inversion and abnormality in the action can be effectively identified, and action segmentation classification is performed in combination with the frame sequence interval and the key node activity state, so that precise analysis of complex action processes can be realized without relying on external hardware. Through the matching with the standard template, the accuracy and timeliness of posture recognition are improved. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 The system flowchart of the present application is shown in the figure.
[0039] Figure 2 The flowchart of the key point extraction module in the present application is shown in the figure.
[0040] Figure 3A flow chart of a graph structure construction module in the present application;
[0041] Figure 4 A flow chart of a node trajectory modeling module in the present application;
[0042] Figure 5 A flow chart of an action segmentation analysis module in the present application;
[0043] Figure 6 A flow chart of a pose recognition output module in the present application. DETAILED DESCRIPTION
[0044] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0045] In the description of the present application, it should be understood that the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. In addition, in the description of the present application, the meaning of "a plurality of" is two or more, unless otherwise explicitly and specifically limited.
[0046] Please refer to Figure 1 An AI-based motion pose recognition system comprises:
[0047] The key point extraction module obtains video frame data in a motion scene, detects the human skeleton in each frame of image, records the pixel coordinates of multiple joints of the human body in two-dimensional space, obtains three-dimensional coordinates, sorts the position sequence of the same number of nodes in consecutive frames on the time axis, establishes the continuous structure of the skeleton in multiple frames of image, and generates a multi-node time sequence coordinate set;
[0048] The graph structure construction module extracts the connection relationship between the key nodes in each frame of image based on the multi-node time sequence coordinate set, identifies adjacent joint pairs and calculates the spatial distance and included angle value, constructs an initial graph structure with node number, edge weight and connection direction attributes, and sequentially arranges the entire group of graph structures in the time dimension according to the node number, to generate a time sequence skeleton graph structure sequence;
[0049] The node trajectory modeling module extracts the coordinate change amplitude, velocity change range and edge angle change sequence of nodes in continuous frames based on the temporal skeleton graph structure sequence. It compares the coordinate difference trend and angle change direction of the same node in different time periods, filters out the node trajectory segments with direction reversal and abnormal acceleration fluctuation, and generates key node trajectory feature groups by combining the spatial position of the node and the number of connected edges in the previous and next frames.
[0050] The action segmentation and parsing module extracts the extreme value frames, inflection point frames and still frames of key nodes in the temporal trajectory based on the key node trajectory feature group. It determines the type of action segment by the frame sequence interval and divides the action process into three categories: start segment, transition segment and end segment, and obtains the segmentation dominant node label set.
[0051] The posture recognition output module retrieves the recorded standard posture templates based on the segmented dominant node label set, finds the action posture identification data of the corresponding node combination in the standard posture template, compares the action posture identification data with the position of the real-time node under the image, filters the standard posture template numbers that meet the matching conditions, sorts them according to the frequency of occurrence, and generates a motion posture matching label sequence.
[0052] The multi-node temporal coordinate set includes node temporal number, spatial coordinate sequence, and node corresponding frame sequence index; the temporal skeleton graph structure sequence includes node connection topology, edge attribute set, and graph sequence number; the key node trajectory feature group includes trajectory change pattern, node spatial distribution, and connection structure features; the segmented dominant node label set includes action stage identifier, dominant node number, and key frame sequence range; and the motion posture matching label sequence includes matching template number, posture recognition label, and frequency of occurrence sorting.
[0053] Please see Figure 2 The key point extraction module includes:
[0054] The image frame acquisition submodule acquires video frame data in the moving scene, uses frame sequence number and timestamp information to number the video frame data, calls the pixel array in the video frame data for decoding operation, extracts the pixel value matrix in the RGB channel, and combines them in time order to obtain a set of continuous frame image matrices.
[0055] The system reads the video file index and decodes the image data stream of each frame sequentially. It converts the compressed frame stream into a continuous RGB pixel matrix by calling the decoding interface of the FFmpeg library. A standard image with a resolution of 640×480 is selected as the output for each frame. The frame rate is set to 30fps, the acquisition period is set to 10 seconds, and a total of 300 frames are acquired. Each frame is sorted by its frame number, generating a sequence of frames, for example, a list {1, 2, ..., 300} with frame numbers from 1 to 300. The pixel matrix set corresponding to each frame forms a three-dimensional array. During image reading, each frame needs to record a timestamp. The timestamp is obtained by dividing the frame number by the frame rate; for example, the timestamp of frame 150 is 150 ÷ 30 = 5.0 seconds. The image matrix and time sequence are combined into a structure to ensure a one-to-one correspondence between image frames and time information, thus obtaining a continuous frame image matrix set.
[0056] The skeleton coordinate recording submodule detects the bounding box position of the target human body region in each frame image based on the continuous frame image matrix set, identifies the two-dimensional coordinate position of the human body key points within the bounding box, performs projection transformation on the two-dimensional coordinates, constructs the coordinate vector of the key points in two-dimensional space, and generates a multi-frame skeleton three-dimensional coordinate set.
[0057] Boundary box localization is performed on each frame of the image matrix using a sliding window scanning method based on pixel intensity differences. Boundary segments with prominent grayscale changes are identified within the image region, and their continuity and closure are assessed. A closure threshold of 0.8 is set; if the closure is greater than 0.8, it is considered a potential human body region. After identification, the region is cropped and sent to the keypoint extraction process. In keypoint extraction, the symmetry of pixel grayscale differences within the region is analyzed to identify areas such as the head, shoulders, and knees. Projection transformation is performed using coordinate difference normalization to convert the two-dimensional image coordinates into three-dimensional relative spatial coordinates. For example, in a frame, the detected left knee keypoint position is (120, 230), and the converted three-dimensional relative coordinates are (0.30, 0.575, 1.5). The conversion is based on image width, image height, and a reference depth value, uniformly set to 1.5 meters. The three-dimensional coordinates are aligned with a preset skeleton template, and the three-dimensional coordinates and numbers of each keypoint in each frame are recorded, generating a multi-frame skeleton three-dimensional coordinate set.
[0058] The node sequence construction submodule calls a multi-frame skeleton 3D coordinate set to match key nodes with the same number in adjacent frames, identify spatial position change trends, and uses the following formula:
[0059] ;
[0060] Calculate the displacement similarity values between nodes, construct a continuous spatial trajectory structure on the time axis, and obtain a multi-node temporal coordinate set;
[0061] in, Indicates the first Node and the Displacement similarity values between nodes Indicates the first Node at the The three-dimensional coordinates in the frame This indicates the time interval between matched frames for a node. Indicates the first segment within a real-time clip The mean of the frame time difference, For the first The node at the th The velocity vector at the frame, This represents the average velocity of a node across video frames. Indicates the first Node at the Angular offset between the frame and the reference skeleton direction. These are the normalized adjustment factors for time, velocity, and angle difference, respectively. The number of video frames;
[0062] Formula calculation logic: Used to measure the similarity of the 3D trajectories of two key nodes throughout the entire video sequence, calculating the Euclidean distance between the two nodes in space in each frame, i.e. After taking the square root, the original displacement value is obtained. Then, the displacement value is divided by the sum of three normalization factors, which are the normalized time interval terms. Speed difference term , Angle difference Multiply by the adjustment coefficient respectively The entire formula is in the frame to The summation is performed to obtain the cumulative evaluation value of the average change trend between nodes. The smaller the calculation result, the closer the change trajectory between nodes is. By setting the adjustment coefficient, the comparison values of time, speed and angle fall into a uniform scale, which facilitates the comparison between samples with different frame rates and motion amplitudes.
[0063] The displacement similarity value between nodes measures the consistency of the trajectory changes of two key nodes in three-dimensional space over time. The spatial distance between the two points is calculated frame by frame, and normalization is performed considering time interval, speed difference and angle change. The smaller the value, the closer the motion trajectory of the two nodes is, the more synchronized their spatial behavior is, which is convenient for constructing continuous action sequences.
[0064] For two keypoints numbered i and j, extract the set of coordinates in T consecutive frames. Let the three-dimensional coordinates of keypoint i in frame t be... The coordinates of the j-th key point are The square of the Euclidean distance between the two points is ;
[0065] Parameter meaning and calculation process:
[0066] Representing nodes respectively In the 3D coordinates in a frame;
[0067] For the first Key points and The average time interval between key points on the matching frames is set to 2 frames here, corresponding to... Second;
[0068] For the first The node at the th The velocity vector at the frame, through Calculations show that if the node's displacement between two frames is (0.01, 0.02, 0.00), then the velocity vector is (0.3, 0.6, 0.0) m / s;
[0069] The average velocity of the node over T frames is the mean of the magnitude of the velocity vector.
[0070] Indicates the first The angle between a node and its adjacent skeleton vector in frame t;
[0071] Let be the corresponding angle of the j-th node, in radians, with the angle difference controlled within the range of 0 to π; The normalization adjustment factors were set to 0.5, 0.3, and 0.2 respectively, as shown in the table below based on the experimental settings;
[0072] Table 1: Parameters for Calculating Skeleton Node Similarity
[0073] Adjustment factor Setting value Setting according to Time normalization coefficient α 0.5 Control within ±0.1 seconds according to inter-frame change amount Speed normalization coefficient β 0.3 Control within 0.5 m / s according to node speed change amount Angle normalization coefficient γ 0.2 Control within 45° according to inter-node angle difference
[0074] As shown in Table 1, the adjustment factor is set with reference to the range of node changes in the actual collected data to ensure that the value of the denominator of the similarity is within a reasonable fluctuation range and to avoid extreme cases from affecting the overall similarity.
[0075] A practical example was performed using nodes 5 (left knee) and 9 (right knee) in a 5-frame image sequence, and the coordinates were extracted as follows:
[0076] Node 5 coordinate sequence: {(0.30, 0.58, 1.50), (0.31, 0.57, 1.50), (0.33, 0.55, 1.49), (0.36, 0.52, 1.49), (0.39, 0.50, 1.48)};
[0077] Node 9 coordinate sequence: {(0.60, 0.58, 1.50), (0.61, 0.57, 1.49), (0.63, 0.55, 1.48), (0.66, 0.52, 1.47), (0.69, 0.50, 1.46)};
[0078] The velocity modulus was calculated, and the average velocity v≈0.354m / s was obtained. The angle difference was quantified into a difference in radians, ranging from 0.3 to 0.6, and calculated using the above formula:
[0079] ;
[0080] The results show that node 5 and node 9 have a moderate degree of spatial similarity in the analyzed frame sequence, which can be used as the basis for constructing node time series labels.
[0081] Please see Figure 3 The graph structure building module includes:
[0082] The graph node connection extraction submodule extracts the coordinate positions of joints in each frame of the image based on a multi-node temporal coordinate set, identifies the positional relationships in the image, determines whether a connection is formed based on the node's index number, filters adjacent joint pairs that satisfy spatial proximity constraints, and records the number combination and connection direction attributes to generate a set of adjacent joint connection pairs.
[0083] By inputting the key node coordinate set for each frame of the image, the horizontal and vertical coordinates of each node in the corresponding frame are extracted. For example, if there is a node numbered 3 in the first frame with a horizontal coordinate of 125 pixels and a vertical coordinate of 92 pixels, then a two-dimensional position vector can be constructed. Based on the node numbering information of each node in the image, a node index dictionary is constructed. Using the key-value pairs in the index dictionary, a spatial adjacency relationship is retrieved between any two nodes. For example, it is determined whether the difference between the horizontal and vertical coordinates of nodes numbered 3 and 7 is within a preset threshold in any frame. If the difference in the horizontal coordinate is less than 50 pixels and the difference in the vertical coordinate is less than 40 pixels, then an adjacency relationship is established, recorded, and a connection directionality is constructed. For example, the direction from the node with the smaller number to the node with the larger number is designated as the positive direction and recorded as... In the entire graph, traverse each pair of nodes that satisfy the adjacency relationship, add the numbered pair to the adjacency pair combination list, such as node pairs (3, 7), (4, 5), (6, 8), etc., and generate a set of adjacent joint connection pairs.
[0084] The spatial feature calculation submodule calls the set of adjacent joint connection pairs to obtain the 3D spatial coordinate difference of each pair of connected joints, using the formula:
[0085] ;
[0086] Calculate the weighted connection strength value for each pair of connected edges, which represents the degree of structural connection and directional perturbation response strength between the two nodes in a multi-frame image sequence, and obtain the weighted connection strength set;
[0087] in, Indicates joint and The weighted connection strength value, Indicates joint and Euclidean distance in a single frame image Indicates the first In-frame joints and The path perturbation length of the directional offset. Indicates the first Frame orientation difference influencing factor Indicates the first Frame dynamic inertia response factor Indicates the total number of frames;
[0088] Formula calculation logic explanation: Calculates the average three-dimensional Euclidean distance between node o and node b across all frames. It is used to measure the static proximity of two points in space, in each frame. In the middle, the perturbation length of the connection direction in this frame is calculated respectively. Then, the influence factor of directional difference will be added. and image response factor Multiplication represents the impact of frame changes on the overall connection strength. The dynamic correction value is then summed over all frames. Finally, the static distance is added to the dynamic correction term, the absolute value is taken, and then divided by the number of frames. The average connection strength value, which integrates the effects of static spatial differences and dynamic directional disturbances, is obtained. The larger the value, the weaker the connection stability; conversely, the smaller the value, the tighter the connection.
[0089] The weighted connection strength value is used to measure the strength of the overall connection between any two nodes in multiple frames of images. It integrates the spatial Euclidean distance between nodes and dynamic features such as directional perturbation and motion consistency in each frame. The lower the value, the closer the node pair is in space, the more consistent the directional changes, and the more stable the connection.
[0090] The connection strength value of each node pair under the 3D spatial coordinate difference is calculated. The specific processing flow is as follows: For any node pair (node o, node b), the 3D coordinate difference value in each frame of the image is extracted. For example, if the position of node o in the first frame is (0.23, 0.45, 0.12) and the position of node b is (0.18, 0.47, 0.10), then the 3D difference vector is... Calculate the Euclidean distance, and you will get:
[0091] ;
[0092] By iterating through the coordinate difference values of each pair of nodes in frame T=5 in this manner, the corresponding average distance value can be calculated. ;
[0093] Parameter meaning and calculation process:
[0094] : Total number of image frames, in this example, T=5;
[0095] The average Euclidean distance between node pairs o and b has been calculated to be 0.057.
[0096] The path perturbation length of the node pair in frame t needs to be quantified based on the degree of direction change. For example, if the offset angle between the path direction and the previous frame is 8°, then the perturbation length is the path length multiplied by sin(8°). Taking a path length of 0.06 as an example, the perturbation is 0.06×sin(8°)≈0.008.
[0097] The influence factor of the directional difference in frame t is based on the cosine of the included angle. It is assigned a value of 0.9 when the included angle is less than 10° and a value of 0.3 when the included angle is greater than 45°.
[0098] : Weighted response factor for frame t. The response factor is derived from frame stability evaluation and ranges from [0.6, 1.0]. If the jitter amplitude between nodes in the frame is less than 5 pixels, then γ = 1.0; otherwise, it is set to 0.7.
[0099] The experimental data are set as follows;
[0100] Table 2: Weighted Connection Strength Parameter Table
[0101]
[0102] Calculate the weighted sum of the disturbance direction terms:
[0103] ;
[0104] Substituting into the formula, we get:
[0105] ;
[0106] The result shows that the weighted connection strength between node o and node b is 0.018106, and this value will be used in subsequent connection filtering and edge construction sorting.
[0107] The advantage of the formula is that by combining the product of directional perturbation, directional difference factor and frame response factor in the calculation, the weighted connection strength value can dynamically reflect the spatial proximity and motion stability between nodes.
[0108] The timing recording submodule extracts the connection strength value and direction attribute between each pair of nodes based on the weighted connection strength set. Combined with the corresponding frame sequence information, it arranges the connection relationships in the video frames in order of node number and assembles them frame by frame into a sequence of node numbers and edge attributes with timing identifiers to obtain the timing skeleton graph structure sequence.
[0109] Each pair of node numbers is paired and encoded with its corresponding direction attribute and weighted connection value to generate the structure as follows: The list of quadruples is arranged in ascending order by node number. For example, if there is a node pair (3, 7) with a W value of 0.018106 and a direction of (0.05, -0.02, 0.02), the corresponding quadruple is (3, 7, 0.018106, (0.05, -0.02, 0.02)). The corresponding frame sequence information is read. For example, if the frame number is 1 to 5, the existence of the connection relationship in each frame is recorded as a Boolean array. For example, [1, 1, 1, 1, 1] indicates continuous existence. The number, direction, connection strength value and frame sequence Boolean array of each node pair are assembled into a complete connection structure table. The table is sorted in order by edge attributes to obtain the temporal skeleton graph structure sequence.
[0110] Please see Figure 4 The node trajectory modeling module includes:
[0111] The node coordinate change extraction submodule is based on the temporal skeleton graph structure sequence. It obtains the two-dimensional spatial coordinate information of nodes in continuous graph frames, calls the coordinate data of nodes with the same number in adjacent frames, calculates the coordinate difference and the direction of continuous difference trend in the time series, records the coordinate difference sequence of each node between continuous frames according to the frame sequence index value, and generates the node temporal coordinate difference sequence.
[0112] The skeleton detection model identifies key skeleton nodes of the human body in each frame of the image, obtains the two-dimensional spatial coordinates of each node in the frame, and archives them according to their node numbers. After skeleton recognition is completed, the coordinate data of nodes between frame t and frame t+1 are read one by one according to the frame index in the frame sequence. By calculating the coordinate change trend of nodes with the same number between frames, the displacement difference of each node between consecutive frames is obtained. To ensure the continuity and temporal sequence of displacement difference, each group of node data is uniformly processed to extract the difference sequence arranged in order in the frame. At the same time, a coordinate change record table is constructed according to the node number. A coordinate difference list is established for each node, and the direction of change in the difference list is organized to generate a corresponding trend direction vector sequence. A fixed-interval index window can be introduced to update and synchronously record the time series index of the node and the corresponding coordinate difference data in real time, thus constructing a complete two-dimensional spatial change model. For example, in a 5-frame video, the movement path of node 1 from frame 1 to frame 5 can be resolved into five sets of coordinate points. By calculating the coordinate change of the node between each two frames, the difference sequence can be obtained. Together with the direction change information, they constitute the basic set of node motion trajectory data and generate the node time series coordinate difference sequence.
[0113] The angle change calculation submodule calls the node time coordinate difference sequence, calculates the node’s velocity value between multiple frames based on the time interval constant and coordinate difference between adjacent frames, obtains the velocity sequence composed of velocity values, and defines the velocity change range based on the trough velocity and peak velocity. Based on the coordinates of the nodes at both ends of the connecting edge in adjacent frames, it calculates the angle change between the direction vectors of the connecting edge and generates the velocity change range and the angle change sequence of the connecting edge.
[0114] Given a fixed time interval, the movement speed of nodes in each frame is extracted using the calculated node displacement difference data. The speed values are organized according to the node number and frame index to form a continuous speed change sequence. This sequence is used to divide the speed change interval. The division process uses the minimum and maximum speed values as a boundary and analyzes the speed fluctuations within the interval. The connecting edges in each frame are processed. Each connecting edge consists of two nodes in the skeleton graph. By reading the coordinate information of the two end nodes in frame t and frame t+1, the change in the connection direction between the two frames is calculated. It is determined whether the angle difference reaches a predetermined threshold. The angle change values are uniformly archived as an angle change sequence. It is possible to observe whether each pair of connecting edges exhibits violent rotation or stable extension between consecutive frames. The speed change interval of the node and the angle change of the connecting edge are uniformly recorded, which can be represented as a video with highly dynamic body movements, such as waving or turning around. By matching the speed change and the angle change, the temporal changes of the node or connecting edge during the movement process can be described, generating a speed change interval and a connecting edge angle change sequence.
[0115] The trajectory feature recognition submodule determines the direction of velocity change and the trend of angle change of nodes in consecutive frames based on the velocity change interval and the sequence of angle change of the connecting edges. It identifies the frame index of nodes with reversed velocity increase / decrease trends or abrupt changes in angle change, calls the node coordinate position and number of connecting edges in the corresponding frame of the node frame index, records the coordinate difference of the node in the previous and next frames and the change in the number of connecting edges, and generates the trajectory feature group of key nodes.
[0116] The system analyzes the trajectory features of nodes frame by frame during the execution of actions. It identifies segments with drastic fluctuations in the velocity sequence, including sections where the velocity direction reverses significantly or the velocity suddenly slows down and then rises rapidly again. The system maps the segments to specific frame indices and, combined with the angle change sequence, checks whether the direction of the connecting edges shifts significantly within the same frame. If both velocity and angle changes occur simultaneously, the system identifies the node as a critical trajectory change point. It retrieves the node's two-dimensional coordinates and the number of connecting edges in the frame and compares them with previous and subsequent frames. If it finds that the node's position has shifted significantly and the number of connecting edges has increased or decreased abruptly, the system records the trajectory feature state of the node and classifies it as a critical node. In practical applications, for example, when recognizing a person performing a "turning" motion, the velocity direction of the torso center node will rapidly reverse, and the angle of the connecting edges of the arm will change drastically. The node changes from connecting 3 edges to 5 edges in this frame. Based on this, the system determines that the node is a critical trajectory point in this frame and generates a critical node trajectory feature group.
[0117] Please see Figure 5 The action segmentation and parsing module includes:
[0118] The keyframe recognition submodule is based on the key node trajectory feature group. It detects the frame index in the node trajectory sequence where the velocity change is zero as a stationary frame, calls the positions of the maximum and minimum values in the coordinate difference sequence to determine the extreme value frame, and determines the inflection point frame based on the frame point in the edge angle sequence where the sign of the rate of change changes, and generates a set of node keyframe indexes.
[0119] The system analyzes the motion state of each frame in the node trajectory, filters out the frame indices where the velocity change is zero, and these frames represent nodes that are stationary on the time axis, meaning the node's position does not move significantly between frames. Based on the extracted velocity sequence, the system compares each time point with a zero velocity value and extracts its index, constructing a list of stationary frame indices. The system then traverses the coordinate difference sequence, searching for extreme value positions within it, i.e., finding the frame points with the largest and smallest changes. By scanning the entire sequence and comparing the differences, the frame number of the extreme value frame can be determined. Finally, the system considers the changes in the angle between adjacent edges in the sequence... The conversion of the rate sign locates inflection point frames. For example, when the angle between the current frame and the next frame changes from positive to negative or vice versa, it is a rate sign conversion. The system records such points as inflection point frames of attitude change. The system merges the three types of key frame index sets: still frames, extreme value frames, and inflection point frames, and summarizes the important time points that describe the change in the motion state of the node. In a specific example, if a person raises their arm to the highest point and pauses briefly before falling, the still frame is the moment when the arm stops, the extreme value frame is the moment when the displacement is the largest during the raising or falling, and the inflection point frame is the frame where the direction is reversed. This generates a set of node key frame indexes.
[0120] The action segmentation submodule calculates the frame sequence interval between adjacent keyframes based on the node keyframe index set, performs interval matching judgment based on the frame interval and the set start segment threshold, transition segment threshold and end segment threshold, marks the action segment attribute category of the keyframe, and generates an action segment category sequence.
[0121] By sequentially comparing the frame index values between every two adjacent keyframes, the system calculates the spacing value and forms a keyframe spacing sequence. After obtaining the frame spacing data, the system classifies and matches each spacing segment according to the preset three frame spacing threshold ranges of start segment, transition segment, and end segment. The system maps each keyframe spacing sequence to the corresponding action segment category by traversing the keyframe spacing sequence, gradually marking the action stage to which each keyframe belongs. Each category is marked as "start", "transition", or "end". In specific action analysis scenarios, such as identifying the action process of "jumping - taking off - landing", the system accurately depicts the logical structure of the action and generates an action segment category sequence.
[0122] The dominant node label extraction submodule calls the keyframe index corresponding to the action segment in the action segment category sequence, extracts the key node number information under the corresponding frame, counts the node index value in each type of action segment, filters the node index value with the number of connected edges exceeding the average value, and generates a segmented dominant node label set.
[0123] Keyframe indices belonging to "start segment," "transition segment," or "end segment" are extracted one by one. Corresponding key node numbers are extracted from each category segment. By statistically analyzing the frequency of these numbers in their respective action segments, a preliminary set of highly active nodes in each segment is obtained. The number of connecting edges for each node in each frame is counted, and the average number of connecting edges for each node in the current segment is calculated. Based on this average value, node numbers with a number of connecting edges exceeding the average are selected. Dominant node labels are then selected based on node frequency. In practical applications, for example, in the "throwing" action, if hand nodes frequently appear in the start and transition segments and their number of connecting edges is consistently greater than that of the remaining nodes (e.g., both the palm and elbow have connecting edges), the system will identify them as the dominant node of that segment. The result is that the dominant node in the start segment is the right wrist, the dominant node in the transition segment is the middle of the right arm, and the dominant node in the end segment is the shoulder or chest center node. This clarifies the location of the skeletal control center for each action segment, forming a segmented dominant node label set.
[0124] Please see Figure 6 The pose recognition output module includes:
[0125] The posture template matching submodule retrieves the node combination index recorded in the standard posture template library based on the node number combination in the segmented dominant node label set, calls the action posture identification data under the associated node combination, performs corresponding frame matching and comparison with the position coordinates of the corresponding node in the real-time image frame, and filters according to the Euclidean distance error threshold and the angle change tolerance standard between nodes to generate a filtered standard posture template number set.
[0126] The system searches the standard pose template library according to the number combination format, matching the preset standard node combination index. Each standard pose template includes a set of node numbers and the relative spatial position and connection structure of the nodes in the standard pose. The system searches for template items with the same number as the currently dominant node combination, filters out the candidate template number set, and sequentially calls the node coordinate data under each matching template. The node coordinate data is compared and matched with the coordinates of nodes with the same number in the real-time image frame. During the comparison process, the system calculates the Euclidean distance between the real-time node position and the template node position node by node, and calculates the change in the angle formed by the connecting edges. The node comparison results are judged to be within the preset error range, such as the Euclidean distance being set to within 5 pixels and the included angle tolerance being set to within 10 degrees. Only template items that meet both conditions are retained in the filtering list. This process is executed independently on each image frame, and the comparison is completed in each frame. In a specific example, if the detected dominant node is a combination of three points: shoulder, elbow, and wrist, the system searches for template items that include the combination in the standard library and compares the coordinates with the positions of the three nodes in the current frame. If the distance between the shoulder and elbow in matching item A differs by 4 pixels and the angle changes by 8 degrees, the system considers A to meet the standard and generates a set of standard posture template numbers after filtering.
[0127] The label matching submodule calls the filtered set of standard posture template numbers, accumulates the number of times the standard posture template number appears, sorts them from high to low according to the frequency of the number, assigns corresponding posture label values and records them in the corresponding frame position in the time series, constructs the posture label time index comparison relationship under the node sequence, and generates a motion posture matching label sequence.
[0128] By traversing the frame number records, a statistical table is constructed that includes template numbers and their corresponding frequencies. The template numbers are sorted in descending order of statistical frequency, and the highest-ranking numbers are assigned corresponding attitude label values. For example, the number with the highest frequency is labeled "Attitude 1", the second highest is labeled "Attitude 2", and so on. The system records the attitude labels into the corresponding image frames, establishing a time-series correspondence between frame indices and attitude labels. That is, each frame corresponds to a label value, which fully describes the standard attitude numbering of the dominant node at each moment. In actual operation, if in a "jump and land" action, template number A appears 15 times, number B appears 3 times, and number C appears 2 times in 20 consecutive frames, the system will label A as "Attitude 1" and assign it to these 15 frames. The remaining frames will be assigned the values "Attitude 2" and "Attitude 3" in sequence, generating a motion attitude matching label sequence.
[0129] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. An AI-based motion posture recognition system, characterized in that, The system includes: The key point extraction module acquires video frame data in the motion scene, detects the human skeleton in each frame image, records the pixel coordinates of multiple joints of the human body in two-dimensional space, and generates a multi-node temporal coordinate set. The graph structure construction module extracts the connection relationships between joints in each frame image based on the multi-node temporal coordinate set, identifies adjacent joint pairs and calculates the spatial distance and angle values, and arranges the entire graph structure in an orderly manner in the time dimension according to the node number to generate a temporal skeleton graph structure sequence. The graph structure construction module includes: The graph node connection extraction submodule extracts the coordinate positions of joints in each frame of the image based on the multi-node temporal coordinate set, identifies the positional relationship in the image, determines whether a connection is formed according to the node number index, filters adjacent joint pairs that satisfy the spatial proximity constraint, and records the number combination and connection direction attribute to generate a set of adjacent joint connection pairs. The spatial feature calculation submodule calls the set of adjacent joint connection pairs, obtains the three-dimensional spatial coordinate difference of each pair of connection joints, calculates the weighted connection strength value of each pair of connection edges, which represents the degree of structural connection and directional perturbation response strength between two nodes in a multi-frame image sequence, and obtains a set of weighted connection strengths. The timing recording submodule extracts the connection strength value and direction attribute between each pair of nodes based on the weighted connection strength set, and combines it with the corresponding frame sequence information to arrange the connection relationships in the video frames in order of node number, assembling them frame by frame into a sequence of node numbers and edge attributes with timing identifiers, and obtaining a timing skeleton graph structure sequence. The node trajectory modeling module extracts the coordinate change amplitude, velocity change range and edge angle change sequence of nodes in continuous frames based on the temporal skeleton diagram structure sequence, compares the coordinate difference trend and angle change direction of the same node in different time periods, and generates key node trajectory feature groups. Based on the key node trajectory feature group, the action segmentation and parsing module extracts the extreme value frames, inflection point frames and still frames of the key nodes in the time sequence trajectory, determines the type of action segment by the frame sequence interval, and divides the action process into three categories: start segment, transition segment and termination segment, thus obtaining the segmentation dominant node label set.
2. The AI-based motion posture recognition system according to claim 1, characterized in that, The multi-node temporal coordinate set includes node temporal number, spatial coordinate sequence, and node corresponding frame sequence index; the temporal skeleton graph structure sequence includes node connection topology, edge attribute set, and graph sequence number; the key node trajectory feature group includes trajectory change mode, node spatial distribution, and connection structure features; and the segmented dominant node label set includes action stage identifier, dominant node number, and key frame sequence range.
3. The AI-based motion posture recognition system according to claim 1, characterized in that, The key point extraction module includes: The image frame acquisition submodule acquires video frame data in the moving scene, uses frame sequence number and timestamp information to number the video frame data, calls the pixel array in the video frame data for decoding operation, extracts the pixel value matrix in the RGB channel, and combines them in time order to obtain a set of continuous frame image matrices. The skeleton coordinate recording submodule detects the position of the bounding box of the target human body region in each frame image based on the continuous frame image matrix set, identifies the two-dimensional coordinate position of the human body key points within the bounding box, performs projection transformation on the two-dimensional coordinates, constructs the coordinate vector of the key points in the two-dimensional space, and generates a multi-frame skeleton three-dimensional coordinate set. The node sequence construction submodule calls the multi-frame skeleton three-dimensional coordinate set, matches key nodes with the same number in adjacent frames, identifies the spatial position change trend, calculates the displacement similarity value between nodes, constructs a continuous spatial trajectory structure on the time axis, and obtains a multi-node temporal coordinate set.
4. The AI-based motion posture recognition system according to claim 1, characterized in that, The weighted connection strength value is calculated using the following formula: ; in, Indicates joint and The weighted connection strength value, Indicates joint and Euclidean distance in a single frame image Indicates the first In-frame joints and The path perturbation length of the directional offset. Indicates the first Frame orientation difference influencing factor Indicates the first Frame dynamic inertia response factor This indicates the total number of frames.
5. The AI-based motion posture recognition system according to claim 1, characterized in that, The node trajectory modeling module includes: The node coordinate change extraction submodule obtains the two-dimensional spatial coordinate information of nodes in continuous frames based on the temporal skeleton graph structure sequence, calls the coordinate data of nodes with the same number in adjacent frames, calculates the coordinate difference and continuous difference trend direction in the time series, records the coordinate difference sequence of each node between continuous frames according to the frame sequence index value, and generates the node temporal coordinate difference sequence. The angle change calculation submodule calls the node time-series coordinate difference sequence, calculates the node’s velocity value between multiple frames based on the time interval constant and coordinate difference between adjacent frames, obtains the velocity sequence composed of velocity values, and defines the velocity change range based on the trough velocity and peak velocity. Based on the coordinates of the nodes at both ends of the connecting edge in adjacent frames, it calculates the angle change between the direction vectors of the connecting edge and generates the velocity change range and the angle change sequence of the connecting edge. The trajectory feature recognition submodule determines the direction of velocity change and the trend of angle change of nodes in consecutive frames based on the velocity change range and the sequence of angle change of the connecting edges. It identifies the frame index of nodes with reversed velocity increase / decrease trends or abrupt changes in angle change, calls the node coordinate position and the number of connecting edges in the corresponding frame of the node frame index, records the coordinate difference of the node in the previous and next frames and the change in the number of connecting edges, and generates a key node trajectory feature group.
6. The AI-based motion posture recognition system according to claim 5, characterized in that, The action segmentation and parsing module includes: The keyframe recognition submodule, based on the key node trajectory feature group, detects the frame index in the node trajectory sequence where the velocity change is zero as a stationary frame, calls the positions of the maximum and minimum values in the coordinate difference sequence to determine the extreme value frame, and determines the inflection point frame based on the frame point in the edge angle sequence where the sign of the rate of change changes, and generates a set of node keyframe indexes. The action segmentation submodule calculates the frame sequence interval between adjacent key frames based on the node key frame index set, performs interval matching judgment based on the frame interval and the set start segment threshold, transition segment threshold and end segment threshold, marks the action segment attribute category of the key frame, and generates an action segment category sequence. The dominant node label extraction submodule calls the keyframe index corresponding to the action segment in the action segment category sequence, extracts the key node number information under the corresponding frame, counts the node index value in each type of action segment, filters the node index value with the number of connected edges exceeding the average value, and generates a segmented dominant node label set.
7. The AI-based motion posture recognition system according to claim 1, characterized in that, The system also includes an attitude recognition output module: The posture recognition output module retrieves the recorded standard posture templates based on the segmented dominant node label set, finds the action posture identification data of the corresponding node combination in the standard posture template, compares the action posture identification data with the position of the real-time node in the image, filters the standard posture template numbers that meet the matching conditions, sorts them according to the frequency of occurrence, and generates a motion posture matching label sequence. The motion posture matching label sequence includes the matching template number, posture recognition label, and frequency of occurrence sorting.
8. The AI-based motion posture recognition system according to claim 7, characterized in that, The posture recognition output module includes: The posture template matching submodule retrieves the node combination index recorded in the standard posture template library based on the node number combination in the segmented dominant node label set, calls the action posture identification data under the associated node combination, performs corresponding frame matching and comparison with the position coordinates of the corresponding node in the real-time image frame, and filters according to the Euclidean distance error threshold and the angle change tolerance standard between nodes to generate a filtered standard posture template number set. The label matching submodule calls the filtered set of standard posture template numbers, accumulates the number of times the standard posture template number appears, sorts them from high to low according to the frequency of the number, assigns corresponding posture label values and records them in the corresponding frame position in the time series, constructs the posture label time index comparison relationship under the node sequence, and generates a motion posture matching label sequence.
Citation Information
Patent Citations
Pattern recognition-based unhooking and rehooking AI accurate recognition grabbing system and method
CN120807957A
Motion posture recognition method and system based on deep learning
CN121305683A