Motion trajectory tracking and action recognition method and model construction method and device

By combining a pre-trained network model with deformable convolution and attention mechanisms with motion continuity constraints, the problems of accuracy and real-time performance in action recognition under complex motion scenarios are solved, achieving efficient action recognition results.

CN120953315APending Publication Date: 2025-11-14CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510880864.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing computer vision processing solutions struggle to accurately identify motion features in complex motion scenarios, especially under varying lighting conditions and complex backgrounds. This results in low accuracy and robustness of action recognition, complex processing logic, high computational load, and poor real-time performance.

Method used

Feature extraction is performed using a pre-trained network model with deformable convolution and attention mechanisms, key point location prediction is performed using a pre-trained network model with motion continuity constraints, and action recognition is performed through cross-modal feature fusion and spatiotemporal dependency analysis.

Benefits of technology

It improves the accuracy and robustness of motion recognition in complex motion scenarios, reduces redundant calculations, and achieves compatibility between real-time performance and accuracy, making it suitable for sports training and competition analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953315A_ABST
    Figure CN120953315A_ABST
Patent Text Reader

Abstract

The invention relates to a motion trail tracking and action recognition method and device and a model construction method and device, and the method comprises the steps: carrying out the target recognition of a time sequence video frame of a to-be-analyzed video, and obtaining a motion subject recognition region; based on a first pre-training network model of a deformable convolution and attention mechanism, performing feature extraction on the motion subject recognition region to obtain a visual feature vector representing the pose of the motion subject; based on a second pre-training network model of motion continuity constraint, performing position prediction on the key point of the motion subject to obtain an attitude feature vector used for representing the position of the key point of the motion subject; performing cross-modal feature fusion on the visual feature vector and the attitude feature vector to obtain a fused feature vector; and based on a third pre-training network model, carrying out space-time dependency analysis and action classification processing on the fusion feature vectors which are arranged in the time sequence to obtain an action recognition result of the motion subject. Accuracy and timeliness of action recognition can be improved, and robustness is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the fields of video processing and artificial intelligence technology, and in particular to a method, model building method and apparatus for motion trajectory tracking and action recognition. Background Technology

[0002] For various sports, motion tracking based on high-precision sensors or on-site filming devices is beneficial for subsequent analysis of the subject's motion performance and for targeted training and improvement. Currently, motion analysis of captured video data mostly employs computer vision processing to distinguish between foreground and background and locate the motion area, followed by image analysis techniques to recognize the motion. For example, background modeling is used to establish a stable background model to differentiate between foreground and background.

[0003] In realizing the concept disclosed herein, the inventors discovered at least the following technical problems in the related technologies: For some sports with complex limb movements (such as freestyle skiing in a halfpipe, skating, curling, etc.), when athletes are in high-speed motion scenarios, existing computer vision processing solutions either struggle to capture and recognize rich motion features, resulting in numerous recognition errors or missed detections under different lighting conditions and complex background changes, leading to low accuracy and robustness of motion recognition; or the processing logic is complex, or the image area being located is wide, resulting in a large computational load, poor real-time data processing, and difficulty in achieving compatibility between real-time computation and accuracy. Summary of the Invention

[0004] To solve the above-mentioned technical problems, or at least partially solve them, embodiments of this disclosure provide a method, model building method, and apparatus for motion trajectory tracking and action recognition.

[0005] In a first aspect, embodiments of this disclosure provide a method for motion trajectory tracking and action recognition. The method includes: performing target recognition on temporal video frames of a video to be analyzed to obtain a moving subject recognition region; extracting features from the moving subject recognition region using a first pre-trained network model based on deformable convolution and attention mechanisms to obtain a visual feature vector representing the pose of the moving subject; predicting the position of key points of the moving subject using a second pre-trained network model based on motion continuity constraints to obtain a pose feature vector representing the position of the key points of the moving subject; fusing the visual feature vector and the pose feature vector across modalities to obtain a fused feature vector; and performing spatiotemporal dependency analysis and action classification processing on the temporally arranged fused feature vector using a third pre-trained network model to obtain the action recognition result of the moving subject.

[0006] In some embodiments, the first pre-trained network model includes: a position offset prediction submodule, a deformable overlay submodule, a feature extraction submodule, a batch normalization processing submodule, a nonlinear processing submodule, and a channel-spatial dual attention submodule. The position offset prediction submodule is used to predict the position offset of sampled pixels in the moving subject recognition region of each video frame based on the first pre-trained network layer. The deformable overlay submodule is used to update the sampled pixel positions according to the position offset, obtaining updated sampled pixel positions. The feature extraction submodule is used to extract features from the updated sampled pixel positions based on the second pre-trained network layer, obtaining a preliminary output feature map. The batch normalization processing submodule is used to batch normalize the preliminary output feature maps of each video frame, obtaining a normalized output feature map corresponding to each video frame. The nonlinear processing submodule is used to perform nonlinear mapping processing on the normalized output feature maps based on an activation function, obtaining an intermediate output feature map. The aforementioned channel-space dual attention submodule is used to combine channel attention mechanism and spatial attention mechanism to perform attention weighting processing on the intermediate output feature map to obtain a visual feature vector representing the pose of the moving subject.

[0007] In some embodiments, the second pre-trained network model includes at least one of a keypoint location prediction submodule, a temporal consistency constraint submodule, or a limb constraint submodule, and a keypoint fusion submodule. The keypoint location prediction submodule is used to predict the location probability distribution of each keypoint in each video frame based on the third pre-trained network layer, obtaining multiple single-point heatmaps for each video frame. The temporal consistency constraint submodule is used to analyze whether the difference between the single-point heatmaps of corresponding keypoints in adjacent frames is less than a preset threshold based on the single-point heatmaps of each video frame, and to perform position correction on target keypoints exceeding the preset threshold based on the Kalman filter algorithm, obtaining a corresponding corrected single-point heatmap and outputting it to the keypoint fusion submodule. The limb constraint submodule is used to analyze the single-point heatmaps of multiple keypoints corresponding to each video frame based on limb connection relationships, and to perform position correction on target keypoints that do not conform to limb connection relationships, obtaining a corrected single-point heatmap and outputting it to the keypoint fusion submodule. The aforementioned key point fusion submodule is used to flatten the features of multiple single-point heatmaps in each video frame to obtain key point position features. Based on the fourth pre-trained network layer, the key point position features are fused according to the motion continuity of adjacent frames to obtain the pose feature vector in each video frame used to represent the position of key points of the moving subject.

[0008] In some embodiments, the loss function of the second pre-trained network model during the training phase is a weighted sum of the heatmap loss function and the motion continuity loss function. The heatmap loss function is used to represent the deviation between the probability distribution of key point positions predicted by the third pre-trained network layer and the actual probability distribution of positions. The continuity loss function is used to represent the degree to which the changes in the positions of corresponding key points in adjacent frames conform to motion continuity when the fourth pre-trained network layer performs feature fusion.

[0009] In some embodiments, the temporal video frames of the video to be analyzed are obtained through at least one of the following preprocessing steps: frame sampling based on a preset sampling rate and resolution unification and pixel normalization processing; or, data augmentation processing, wherein the data augmentation includes at least one of the following processing methods: random flipping, brightness variation, and contrast adjustment; or, splicing or merging of data source segments at least one of these processing steps to obtain an integrated video to be analyzed and then performing frame sampling. Target recognition is performed on the temporal video frames of the video to be analyzed to obtain a moving subject recognition region, including: based on a target detection algorithm, moving subject recognition is performed on each temporal video frame of the video to be analyzed to obtain one or more subject boundary recognition boxes corresponding to moving subjects; based on the distribution information within the subject boundary recognition boxes, the subject boundary recognition boxes are finely adjusted to obtain the moving subject recognition region. The above visual feature vector and the above posture feature vector are fused across modalities to obtain a fused feature vector, including: comparing whether the dimensions of the visual feature vector and the posture feature vector are consistent; if the dimensions of the visual feature vector and the posture feature vector are inconsistent, performing dimension alignment mapping processing on the posture feature vector to the visual feature vector; and performing feature fusion processing on the dimension-aligned visual feature vector and the posture feature vector to obtain a fused feature vector.

[0010] In some embodiments, the third pre-trained network model includes a temporal dependency determination submodule and an action classification processing submodule. The temporal dependency determination submodule performs spatiotemporal dependency analysis on the fused feature vectors arranged temporally based on the fifth pre-trained network layer to obtain the temporal predicted state and the latent vectors corresponding to each time step. The action classification processing submodule performs action classification based on the sixth pre-trained network layer, using the temporal predicted state and the latent vectors corresponding to each time step, to obtain the action recognition result of the moving subject.

[0011] Secondly, embodiments of this disclosure provide a method for constructing a motion trajectory tracking and action recognition model. The above construction method includes: performing target recognition on temporal video frames of multiple videos in the training dataset to obtain a training motion subject recognition region; extracting features from the training motion subject recognition region using a first network model based on deformable convolution and attention mechanisms to obtain a training visual feature vector representing the pose of the motion subject; predicting the position of key points of the motion subject using a second network model based on motion continuity constraints to obtain a training posture feature vector representing the position of key points of the motion subject; fusing the above training visual feature vector and the above training posture feature vector across modalities to obtain a training fused feature vector; performing spatiotemporal dependency analysis and action classification processing on the temporally arranged training fused feature vector based on a third network model to obtain the training action recognition result of the motion subject; using the action labels of multiple videos in the training dataset as training labels to train the first network model, the second network model, and the third network model, resulting in a first pre-trained network model, a second pre-trained network model, and a third pre-trained network model after training; and constructing a motion trajectory tracking and action recognition model based on the first pre-trained network model, the second pre-trained network model, and the third pre-trained network model.

[0012] Thirdly, embodiments of this disclosure provide an apparatus for motion trajectory tracking and action recognition. The apparatus includes: a first target detection module, a first adaptive feature extraction module, a first pose prediction module, a first multimodal fusion module, and a first temporal action analysis module. The first target detection module is used to perform target recognition on temporal video frames of a video to be analyzed, obtaining a moving subject recognition region. The first adaptive feature extraction module is used to extract features from the moving subject recognition region based on a first pre-trained network model using deformable convolution and attention mechanisms, obtaining a visual feature vector representing the pose of the moving subject. The first pose prediction module is used to predict the position of key points of the moving subject based on a second pre-trained network model using motion continuity constraints, obtaining a pose feature vector representing the position of the key points of the moving subject. The first multimodal fusion module is used to perform cross-modal feature fusion of the visual feature vector and the pose feature vector, obtaining a fused feature vector. The first temporal action analysis module is used to perform spatiotemporal dependency analysis and action classification processing on the temporally arranged fused feature vector based on a third pre-trained network model, obtaining the action recognition result of the moving subject.

[0013] Fourthly, embodiments of this disclosure provide an apparatus for constructing a motion trajectory tracking and action recognition model. The apparatus includes: a second target detection module, a second adaptive feature extraction module, a second pose prediction module, a second multimodal fusion module, a second temporal action analysis module, a training module, and a model construction module. The second target detection module is used to perform target recognition on temporal video frames from multiple videos in a training dataset to obtain a training motion subject recognition region. The second adaptive feature extraction module is used to extract features from the training motion subject recognition region based on a first network model using deformable convolution and attention mechanisms to obtain a training visual feature vector representing the pose of the motion subject. The second pose prediction module is used to predict the position of key points of the motion subject based on a second network model using motion continuity constraints to obtain a training pose feature vector representing the position of the key points of the motion subject. The second multimodal fusion module is used to perform cross-modal feature fusion of the training visual feature vector and the training pose feature vector to obtain a training fused feature vector. The second temporal action analysis module, based on the third network model, performs spatiotemporal dependency analysis and action classification on the temporally arranged fusion feature vectors used for training, obtaining the action recognition results for the moving subject. The training module uses action labels from multiple videos in the training dataset as training labels to train the first, second, and third network models, resulting in the first, second, and third pre-trained network models. The model construction module constructs motion trajectory tracking and action recognition models based on the first, second, and third pre-trained network models.

[0014] Fifthly, embodiments of this disclosure provide an electronic device. The electronic device includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, communication interface, and memory communicate with each other via the communication bus; the memory stores computer programs; and the processor, when executing the program stored in the memory, implements the motion trajectory tracking and action recognition method provided in the first aspect embodiment or the motion trajectory tracking and action recognition model construction method provided in the second aspect embodiment.

[0015] Sixthly, embodiments of this disclosure provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the motion trajectory tracking and action recognition method provided in the first aspect embodiment or the motion trajectory tracking and action recognition model construction method provided in the second aspect embodiment.

[0016] The technical solutions provided in the embodiments of this disclosure have at least some or all of the following advantages:

[0017] By performing target recognition on temporal video frames of the video to be analyzed, the region for identifying the moving subject is obtained. Since the first pre-trained network model is based on deformable convolution and attention mechanisms, it can adaptively adjust the receptive field during the feature extraction stage and focus on motion regions (such as detailed regions like jumping and flipping) and core features. The resulting visual feature vector is concentrated on the moving subject and reflects the detailed changes in the subject's movements, improving robustness to blurred and occluded scenes, reducing redundant computation, and increasing the inference efficiency for subsequent moving subject tracking and action classification. Since the second pre-trained network model is based on motion continuity constraints, predicting the position of key points based on the second pre-trained network model improves the accuracy of position prediction and conforms to the laws of motion continuity, effectively enhancing the posture description capability, significantly reducing key point jitter in high-speed motion scenes, and improving the robustness of posture estimation. The system ensures reliability. Then, by fusing visual and posture feature vectors across modalities, a fused feature vector is obtained. This fully utilizes visual features and structured posture information, with the two mutually reinforcing and assisting each other. The generated fused feature vector has the advantage of comprehensive and multi-dimensional description, adapting to different lighting conditions, complex background angles, and detailed changes in movement. Subsequently, spatiotemporal dependency analysis and action classification are performed on the temporally arranged fused feature vector to effectively track the moving subject and construct motion dependencies, thereby distinguishing similar but different action types and obtaining relatively accurate action recognition results. Furthermore, since the above processing utilizes pre-trained network models, the process is efficient, combining accuracy and real-time performance, and can be applied to training analysis or competition analysis of various sports (such as freestyle skiing, skating, curling, etc.). Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0019] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating a method for motion trajectory tracking and action recognition according to an embodiment of the present disclosure is shown schematically.

[0021] Figure 2 A block diagram of the structure and a schematic diagram of the processing of a first pre-trained network model according to an embodiment of the present disclosure are shown schematically.

[0022] Figure 3 The diagram schematically illustrates the network structure of a first pre-trained network model and the detailed structure of the attention processing network according to an embodiment of the present disclosure.

[0023] Figure 4 A block diagram of the structure and a schematic diagram of the processing of a second pre-trained network model according to an embodiment of the present disclosure are shown schematically.

[0024] Figure 5 The diagram illustrates the process of multiple training iterations corresponding to the training phase of a second pre-trained network model according to an embodiment of the present disclosure.

[0025] Figure 6 A block diagram of the structure and a schematic diagram of the processing of a third pre-trained network model according to an embodiment of the present disclosure are shown schematically.

[0026] Figure 7 A flowchart illustrating a method for constructing a motion trajectory tracking and action recognition model according to an embodiment of the present disclosure is shown schematically.

[0027] Figure 8 The diagram illustrates the process of constructing a motion trajectory tracking and motion recognition model according to an embodiment of the present disclosure.

[0028] Figure 9 A schematic block diagram of an electronic device provided in an embodiment of the present disclosure is shown. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0030] The first exemplary embodiment of this disclosure provides a method for motion trajectory tracking and action recognition.

[0031] Figure 1 A flowchart illustrating a method for motion trajectory tracking and action recognition according to an embodiment of the present disclosure is shown schematically.

[0032] Reference Figure 1As shown, the motion trajectory tracking and action recognition method provided in this embodiment includes the following steps: S110, S120, S130, S140 and S150.

[0033] In step S110, target recognition is performed on the temporal video frames of the video to be analyzed to obtain the moving subject recognition region.

[0034] In the embodiments of this disclosure, the data source of the video to be analyzed can be video data captured within a certain time period for the same sports scene. This video data can be a single video, multiple videos captured from different perspectives, multiple video clips captured from different camera positions at different time points, or multi-source video data captured by multiple terminals (such as professional sports cameras, smartphones, etc.) or multiple terminals (which can be of the same or different categories, emphasizing the number of devices, such as multiple cameras or multiple smartphones, etc.). The aforementioned sports scenes include, but are not limited to: single-person skiing training scenes in a halfpipe, multi-person skiing training scenes in a halfpipe, multi-person skiing competition scenes in a halfpipe, skating training scenes (such as short track speed skating, figure skating, etc.), skating competition scenes, curling training scenes, curling competition scenes, etc. In other application scenarios, it can also be used for tracking and motion recognition of wild animals to promote the understanding of wild animals or the advancement of bionic technology.

[0035] Since the data sources of the videos to be analyzed may differ, preprocessing is performed on the videos to obtain the corresponding sampled temporal video frames. These are image frames based on temporal arrangement, obtained by sampling at a preset sampling rate. For example, the temporal video frames can be represented as {I1, I2, I3, ..., I...} m}, where m represents the total number of sampling frames, I1~I m These represent the image frames corresponding to the first sampling timestamp to the mth sampling timestamp, respectively.

[0036] In some embodiments, the temporal video frames of the video to be analyzed are obtained by the following preprocessing: frame sampling based on a preset sampling rate and resolution unification (e.g., resolution size of 640×480, the value is only for example) and pixel normalization (corresponding to color space normalization, such as dividing by 255 or normalizing according to ImageNet mean and variance).

[0037] To control the number of input frames and adapt to the model's processing capabilities, the video to be analyzed is sampled frame by frame, that is, some frames are retained while the rest are skipped.

[0038] Assuming the original frame rate is f in The target frame rate is f out Then the sampling interval s is: The output frame sequence is: {F is}, Where N f This indicates the total number of video frames.

[0039] In the preprocessing operation to unify the resolution, all image frames in the output frame sequence are scaled to a uniform width × height (W′, H′).

[0040] Assuming the original resolution is (W, H), then the scaling factor is: The pixels (x, y) in each image frame are resampled using bilinear interpolation:

[0041] In the preprocessing operation of pixel normalization, the distribution of the input image is made close to the standard normal distribution in order to speed up the processing.

[0042] For example, standardizing the ImageNet image frame dataset:

[0043] Where, μ c and σ c These are the mean and standard deviation of the ImageNet dataset across each channel c.

[0044] As an example, the specific values ​​are as follows:

[0045] μ=[0.485,0.456,0.406], σ=[0.229,0.224,0.225].

[0046] In some embodiments, to improve the image quality of the data source, a data augmentation preprocessing operation can be performed. This data augmentation includes at least one of the following processing methods: random flipping, brightness variation, and contrast adjustment. The data-augmented video to be analyzed is then sampled to obtain time-series video frames.

[0047] For example, in image denoising and enhancement, removing salt-and-pepper noise is defined as:

[0048] in This represents a neighborhood window centered at (x, y).

[0049] Histogram equalization is applied to local areas to limit contrast and avoid over-enhancement.

[0050] I clahe (x,y)=CLAHE(I(x,y),clipLimit=C,gridSize=G).

[0051] In some embodiments, for data source segments or multiple videos, the video to be analyzed is integrated based on splicing or fusion processing. For example, the time-series video frames of the video to be analyzed are obtained by the following preprocessing: performing at least one process of splicing or fusion of data source segments to obtain the integrated video to be analyzed and performing frame sampling.

[0052] The preprocessing operations in the above embodiments can be used individually or in combination. They can be consistent with the preprocessing operations in the training phase, or optional preprocessing operations performed in some pre-training phases can be omitted in the analysis phase. By preprocessing the video to be analyzed and then sampling the frames to obtain time-series video frames, at least one of the following effects can be achieved: improving image quality, reducing noise interference, reducing the computational load of subsequent processing, or improving the quality of algorithm output results, etc.

[0053] In some embodiments, the target recognition method in step S110 above can adopt various image region recognition methods, such as, but not limited to, image semantic recognition methods, target detection algorithms, etc. The following description uses a target detection algorithm as an example.

[0054] In some embodiments, step S110 above, which involves performing target recognition on temporal video frames of the video to be analyzed to obtain a moving subject recognition region, includes:

[0055] Based on the object detection algorithm, moving subject recognition is performed on each temporal video frame of the video to be analyzed to obtain one or more subject boundary recognition boxes corresponding to the moving subjects.

[0056] Based on the distribution information within the subject boundary recognition box, the aforementioned subject boundary recognition box is finely adjusted to obtain the moving subject recognition area.

[0057] In some embodiments, the above-mentioned target detection algorithm uses the SSD (Single Shot Multi-Box Detector) detection algorithm. The SSD detection algorithm is a single-stage target detection algorithm based on deep learning. It combines convolutional neural networks with multi-scale anchor point mechanisms, and can predict targets of different sizes based on multi-scale feature maps. By introducing a prior box mechanism, it has both high real-time performance and detection accuracy.

[0058] The specific execution process of the SSD detection algorithm is shown in the following example:

[0059] A default bounding box is set on the feature map of the k-th layer, and the size of the default bounding box is s. k The definition is as follows:

[0060] Where K is the number of feature layers, s minFor the minimum bounding box size (relative to the input image), s max This is the maximum frame size (relative to the input image).

[0061] For each frame, set multiple aspect ratios. The corresponding width and height are: For each default box d, predict its offset relative to the ground truth box g:

[0062] The loss function L of the SSD detection algorithm consists of two parts: classification loss L conf and positioning loss L loc : Where N represents the number of default boxes that match the actual target; Used to indicate whether the j-th default box matches the i-th real box; p i c represents the predicted class probability; j Indicates the true category label; t i Indicates the parameters of the prediction box; This represents the parameters of the true bounding box; α represents the tradeoff coefficient (usually taken as 1).

[0063] The classification loss uses the softmax loss:

[0064] The localization loss uses the Smooth L1 Loss function:

[0065] Finally, Non-Maximum Suppression (NMS) is performed to eliminate redundant prediction boxes before the final output. If the IoU parameter of two boxes exceeds a threshold (e.g., 0.5), the one with the higher score is retained.

[0066] After identifying moving subjects based on object detection algorithms, each detected moving subject corresponds to a subject boundary recognition box b = (x min y min x max y max ).

[0067] In some embodiments, by analyzing the distribution information of the subject boundary recognition box identified by the target detection algorithm, such as analyzing the relative distribution ratio and relative positional relationship between the blank area / background area and the distribution area of ​​the moving subject in the edge region of the subject boundary recognition box, the boundary recognition box can be finely adjusted, such as by image cropping, to obtain a more accurate moving subject recognition area. This can reduce the image size and the amount of subsequent feature extraction processing.

[0068] For example, assuming the original image size is (W, H), then the pixel range corresponding to the bounding box is: x∈[x min x max ], y∈[y min y max The cropped sub-image is represented as: I roi =I(x, y)x∈[x min x max ], y∈[y min y max ].

[0069] In some embodiments, the moving subject recognition region obtained in step S110 may undergo post-processing to meet the input format requirements of the subsequent step S120. It is understood that this post-processing step is optional; it is not necessary to perform this step if the same object detection and fine-tuning cropping processes have been performed and the processing results meet a uniform format.

[0070] In some examples, post-processing involves scaling the moving subject recognition region to another uniform size, such as 224×224. Pixels are then resampled using bilinear interpolation. This process ensures that all inputs have the same dimensions for step S120 and remain consistent with the training phase, facilitating efficient processing.

[0071] In step S120, based on the first pre-trained network model with deformable convolution and attention mechanism, features are extracted from the moving subject recognition region to obtain a visual feature vector representing the pose of the moving subject.

[0072] Figure 2 A schematic diagram of the structure and processing of a first pre-trained network model according to an embodiment of the present disclosure is shown.

[0073] In some embodiments, refer to Figure 2 As shown, the first pre-trained network model includes: a position offset prediction submodule 210, a deformable superposition submodule 220, a feature extraction submodule 230, a batch normalization processing submodule 240, a nonlinear processing submodule 250, and a channel-space dual attention submodule 260.

[0074] The aforementioned position offset prediction submodule 210 is used to predict the position offset of the sampled pixel points in the moving subject recognition region of each video frame based on the first pre-trained network layer.

[0075] The aforementioned deformable overlay submodule 220 is used to update the position of the sampled pixel point according to the aforementioned position offset, so as to obtain the updated position of the sampled pixel point.

[0076] The aforementioned feature extraction submodule 230 is used to extract features from the updated sampled pixel positions based on the second pre-trained network layer to obtain a preliminary output feature map.

[0077] The batch normalization processing submodule 240 described above is used to perform batch normalization processing on the preliminary output feature maps of each video frame to obtain the normalized output feature maps corresponding to each video frame.

[0078] The aforementioned nonlinear processing submodule 250 is used to perform nonlinear mapping processing on the normalized output feature map based on the activation function to obtain an intermediate output feature map.

[0079] The aforementioned channel-space dual attention submodule 260 is used to combine the channel attention mechanism and the spatial attention mechanism to perform attention weighting processing on the intermediate output feature map to obtain a visual feature vector representing the pose of the moving subject.

[0080] In some embodiments, the first pre-trained network model is a multi-resolution attention receptive field convolutional neural network, used to extract high-dimensional spatial feature vectors from the moving subject recognition region of each image frame to describe human posture and motion state (spatial position, velocity, acceleration, rotational angular velocity, rotational angular acceleration, etc.), which are visual feature vectors.

[0081] Figure 3 The diagram schematically illustrates the network structure of a first pre-trained network model and the detailed structure of the attention processing network according to an embodiment of the present disclosure.

[0082] Reference Figure 3 As shown, an example of a refined network structure corresponding to the first pre-trained network model (multi-resolution attention receptive field convolutional neural network) is provided, which includes: pre-trained original convolutional layers, multiple sequentially connected pre-trained multi-resolution attention convolutional residual blocks, average pooling layers, flattening layers and fully connected layers.

[0083] The pre-trained original convolutional layers described above are used to perform the following operations: convolution operation, batch regularization (or batch normalization) operation, nonlinear transformation based on activation function to increase the model's nonlinear processing capability, and max pooling operation.

[0084] The pre-trained multi-resolution attention convolutional residual block includes: convolutional residual blocks and scaled residual blocks, an offset prediction network, and an attention processing network.

[0085] The average pooling layer described above is used to reduce the dimensionality of feature maps, decrease computational cost and the number of parameters, and prevent overfitting. Its main functions include: extracting key features while preserving crucial image information; reducing feature map size and computational complexity; enabling parameter sharing and enhancing the model's generalization ability; and reducing spatial sensitivity and improving the model's robustness.

[0086] The aforementioned flattening layer performs a flattening (or feature map flattening) operation, expanding a multi-dimensional feature map into a one-dimensional vector, which is then passed as input to the fully connected layer. Its main functions include: compressing the information extracted from the feature map into a vector, which can then be fed into the fully connected layer for tasks such as classification or regression; and improving the model's processing efficiency by reducing computation and the number of parameters due to the flattening operation.

[0087] based on Figure 3 The network structure of the first pre-trained network model in the example, assuming the input feature map is... Where H and W represent the height and width of the feature map, respectively; C in Indicates the number of input channels.

[0088] First, an offset is predicted for each sampling point. This is typically done through a small sub-network (the first pre-trained network layer, i.e., the offset prediction network). This sub-network is an additional convolutional layer whose input is the input feature map X of the current layer, and whose output is the predicted offset ΔP for each sampling point. Let the output of this sub-network be... Here, N is the number of sampling points in the convolution kernel. Each sampling point has two offset values ​​(corresponding to offsets along the x-axis and y-axis, respectively), so the total output dimension is H×W×2N. Calculation formula: ΔP=Conv offset (X), here, Conv offset This indicates a convolution operation used to predict offsets.

[0089] Based on the predicted offset ΔP, the position of each sampling point can be updated (implemented based on the deformable overlay submodule 220, e.g., based on the residual block with added scale). The original set of sampling point positions can be represented as P = {p n}, where p n This is a fixed offset relative to the center point (x, y). The updated sampling point position becomes: P′=P+ΔP, specifically, for each sampling point p n Its new sampling position is p′ n =p n +Δp n , where Δp n It is the corresponding offset obtained from ΔP.

[0090] Next, convolution operations are performed using the updated sampling point positions (based on the pre-trained original convolutional layer). Let the convolution kernel weights be... Then output feature map The calculation method is as follows: For each output position (x′, y′), its value is: Where: (dx′ n ,dy′ n ) is the new sampling position relative to the output position (x′, y′), obtained by adding an offset to the original sampling position. n These are the corresponding convolutional kernel weights. Since the actual sampling locations may not be integer coordinates, interpolation methods are needed to calculate the feature values ​​for these non-integer coordinates.

[0091] After obtaining the initial output feature map, it is usually batch normalized (based on the pre-trained original convolutional layers) to accelerate training and stabilize the model. The output of the batch normalized model is denoted as Z1: Z1 = BN(Y). Then, the ReLU activation function (based on the pre-trained original convolutional layers) is applied to perform a nonlinear transformation on the batch normalized feature map: z1 = ReLU(Z1).

[0092] Further processing of the feature map enhances the representation of key information. (Refer to...) Figure 3 The diagram illustrates the channel attention and spatial attention mechanisms of the attention processing network. The new feature map obtained based on this dual attention mechanism is: z2 = M s (M c (z1))⊙z1, the final output is: y=z2+x. This mechanism enables the model to dynamically adjust its receptive field, better adapt to complex spatial transformations, thereby improving the effectiveness and robustness of feature extraction.

[0093] Since the first pre-trained network model is based on deformable convolution and attention mechanism, it can adaptively adjust the receptive field and focus on the motion region (such as detailed regions such as jumping and flipping) and pay attention to core features during the feature extraction stage. The resulting visual feature vector can be concentrated on the moving subject and reflect the detailed changes in the moving subject's actions. This can improve the robustness to blurred and occluded scenes, reduce redundant calculations, and improve the inference efficiency of subsequent moving subject tracking and action classification.

[0094] In step S130, the second pre-trained network model based on motion continuity constraints predicts the positions of key points of the moving subject, and obtains the posture feature vector used to characterize the positions of key points of the moving subject.

[0095] Figure 4 A schematic diagram illustrating the structure and processing of a second pre-trained network model according to an embodiment of this disclosure is provided. (In conjunction with...) Figure 1 and Figure 4 As shown, the input to the second pre-trained network model can be the moving subject recognition region of a temporal video frame, or it can be the visual feature vector extracted in step S120, as shown in the figure. Figure 4 The dashed arrows running in parallel are shown in the middle.

[0096] In some embodiments, refer to Figure 4 As shown, the second pre-trained network model includes at least one of the following: key point location prediction submodule 410, time consistency constraint submodule 421 or limb constraint submodule 422, and key point fusion submodule 430.

[0097] The aforementioned key point location prediction submodule 410 is used to predict the location distribution probability of each key point in each video frame based on the third pre-trained network layer, and to obtain multiple single-point heatmaps for each video frame.

[0098] Reference Figure 4 As shown in the information flow indicated by the single-dot dashed box, the aforementioned time consistency constraint submodule 421 is used to analyze whether the difference between the single-dot heatmaps of corresponding key points in adjacent frames is less than a preset threshold based on the single-dot heatmaps of each video frame, and to perform position correction for target key points that exceed the preset threshold based on the Kalman filter algorithm, thereby obtaining the corresponding corrected single-dot heatmap and outputting it to the key point fusion submodule.

[0099] Reference Figure 4 As indicated by the double-dotted line box in the information flow, the aforementioned limb constraint submodule 422 is used to analyze the single-point heatmap of multiple key points corresponding to each video frame based on the limb connection relationship and correct the position of target key points that do not conform to the limb connection relationship, thereby obtaining the corrected single-point heatmap and outputting it to the keypoint fusion submodule.

[0100] The aforementioned key point fusion submodule 430 is used to flatten the features of multiple single-point heatmaps in each video frame to obtain key point position features. Based on the fourth pre-trained network layer, the key point position features are fused according to the motion continuity of adjacent frames to obtain the pose feature vector in each video frame used to represent the position of key points of the moving subject.

[0101] In some embodiments, the second pre-trained network model described above is a spatiotemporally consistent pose prediction model that predicts the spatial location of key points of a moving subject (e.g., the joints of the limbs, the head, etc.) to achieve pose estimation, providing a more refined motion description than bounding boxes and assisting in the refined recognition of actions.

[0102] In some embodiments, a keypoint distribution map and a partial confidence map are first generated, resulting in a single-point heatmap used to predict the location distribution probability of each keypoint.

[0103] Assume the set of key points is Where K is the total number of keypoints. For each frame image I t The model outputs a heatmap. Each channel corresponds to a single-point heatmap. This represents the probability distribution of the location of the k-th keypoint in the image. The higher the probability of a location appearing, the higher its heat map intensity. For a single image frame, the heat maps of multiple keypoints can be integrated (e.g., by channel, with one channel corresponding to one keypoint; or by overlaying the heat maps according to the relative positions of different keypoints) to obtain a pose heat map.

[0104] In an embodiment including the limb constraint submodule 422, based on obtaining multiple single-point heatmaps corresponding to each video frame, limb connection relationships are established through the limb connection field (describing the direction and length information between two key points), and the position of target key points that do not conform to the limb connection relationship is corrected. Alternatively, in an embodiment including the temporal consistency constraint submodule 421, position correction can also be performed based on motion consistency constraints, whereby the difference between the key point heatmaps of adjacent frames should be kept within a reasonable range: ||H t -H t-1 If the condition ||2<∈ is not met, then keypoint drift or false detection is considered to exist. Keypoint positions are corrected using a technique similar to Kalman filtering to ensure a smooth transition between keypoint positions in adjacent frames and maintain the stability of the keypoint position sequence.

[0105] Figure 5 The diagram illustrates the process of multiple training iterations corresponding to the training phase of a second pre-trained network model according to an embodiment of the present disclosure.

[0106] The second pre-trained network model is obtained by training the second network model, referring to... Figure 5 As shown, during the training phase, the input is processed. The moving subject recognition region or corresponding visual feature vector of each video frame is processed by convolution X (corresponding to the third pre-trained network layer after training) to obtain a partial confidence map, that is, a single-point heatmap corresponding to the positional probability distribution of a certain key point in each video frame. The corresponding loss function is the heatmap loss function (in... Figure 5 The loss function described in the text is a continuous dynamic partial confidence map loss function; adjacent frames are processed by convolution Y (corresponding to the fourth pre-trained network layer after training) to obtain the limb connection field, which changes dynamically with time. In this stage, convolution Y performs feature fusion processing on the key point position features according to the motion continuity of adjacent frames, and the corresponding loss function is the continuous loss function (in the text). Figure 5The description is as a continuous dynamic limb connection field loss function.

[0107] During the training phase, the corresponding loss function is a weighted sum of the heatmap loss function and the motion continuity loss function. The heatmap loss function is used to represent the deviation between the probability distribution of keypoint locations predicted by the third pre-trained network layer and the actual probability distribution; the continuity loss function is used to represent the degree to which the changes in the corresponding keypoint locations in adjacent frames conform to motion continuity when the fourth pre-trained network layer performs feature fusion.

[0108] In this embodiment, in addition to the heatmap loss function, a motion continuity loss function is defined to further enhance the temporal coherence of the key point sequence. This motion continuity loss function measures the degree to which the positional changes of corresponding key points in adjacent frames conform to motion continuity. To simplify the representation, the displacement changes of key points between consecutive frames can be directly used for simplified calculation.

[0109] Assuming the keypoint coordinates in frame t are pt, the motion continuity loss function can be expressed as: in: That is, the difference in keypoint coordinates between two consecutive frames, where ||||2 represents the L2 norm, pointing to the square root of the sum of the squares of each element in the variable; L motion It involves squaring the L2 norm of each pair of adjacent image frames and then summing the results over all frames.

[0110] Therefore, during the training phase, the loss function of the second network model is the heatmap loss function L. heatmap and motion continuity loss function L motion The weighted sum (which, when it includes at least one of the limb constraint submodules or the time consistency constraint submodule, is described as a continuous dynamic limb connection field loss function) can be expressed as:

[0111] L total =L heatmap +λ·L motion , where λ is the balancing coefficient or weighting coefficient, used to adjust the relative importance between the two loss terms.

[0112] Since the second pre-trained network model is based on motion continuity constraints, predicting the position of key points based on the second pre-trained network model can improve the accuracy of position prediction and conform to the law of motion continuity, effectively enhance the posture description capability, significantly reduce key point jitter in high-speed motion scenarios, and improve the robustness and reliability of posture estimation.

[0113] The second pre-trained network model in this embodiment can not only generate keypoint heatmaps for each frame, but also ensure the temporal coherence and stability of these keypoints. In some embodiments, in addition to outputting pose feature vectors, an optimized keypoint position sequence can also be output, in which the positions are more accurate and conform to the physical laws of motion. Moreover, these keypoint position sequences are connected according to the limb relationships of the moving subject to form a motion skeleton structure. For example, assuming the set of limbs is ε, for each edge e = (j i j j Given ε, output a two-dimensional vector field. This indicates the direction the region points towards the limb. Using a keypoint connection algorithm, the most probable location points in the heatmap are matched with the directional information in the limb connection field, and multiple most probable location points are connected to obtain a complete motion skeleton structure. This motion skeleton structure can be used as another output, for example, it can be output together with the subsequently recognized action results to improve the intuitiveness of overall trajectory tracking and action recognition.

[0114] In step S140, the above visual feature vector and the above pose feature vector are fused across modal features to obtain a fused feature vector.

[0115] In some embodiments, step S140 above involves cross-modal feature fusion of the visual feature vector and the pose feature vector to obtain a fused feature vector, including:

[0116] Compare whether the dimensions of the visual feature vector and the pose feature vector are consistent;

[0117] When the dimensions of the visual feature vector and the pose feature vector are inconsistent, the pose feature vector is dimensionally aligned and mapped to the visual feature vector.

[0118] The dimension-aligned visual feature vector and pose feature vector are subjected to gated feature fusion processing to obtain the fused feature vector.

[0119] For example, visual feature vector representation: d f Given the dimension of the visual feature vector, the pose feature vector is represented as: d h Let be the dimension of the pose feature vector. After dimension alignment mapping, cross-modal feature fusion is performed. In some embodiments, the fusion ratio is controlled by gating.

[0120] γ t =σ(W g [f t h t ]+b g ), vt =γ t ·f t +(1-γ t )·h t ,

[0121] Final output: fused feature vector

[0122] In step S150, based on the third pre-trained network model, spatiotemporal dependency analysis and action classification are performed on the temporally arranged fusion feature vectors to obtain the action recognition results of the moving subject.

[0123] Figure 6 A block diagram of the structure and a schematic diagram of the processing of a third pre-trained network model according to an embodiment of the present disclosure are shown schematically.

[0124] In some embodiments, refer to Figure 6 As shown, the third pre-trained network model includes: a temporal dependency determination submodule 610 and an action classification processing submodule 620.

[0125] The aforementioned temporal dependency determination submodule 610 is used to perform spatiotemporal dependency analysis on the fusion feature vector of temporal arrangement based on the fifth pre-trained network layer, so as to obtain the temporal prediction state and the hidden vector corresponding to each time step.

[0126] The aforementioned action classification processing submodule 620 is used to perform action classification based on the sixth pre-trained network layer, the aforementioned temporal prediction state, and the latent vectors corresponding to each time step, to obtain the action recognition result of the moving subject.

[0127] In some embodiments, the fifth pre-trained network layer described above is trained based on a Long Short-Term Memory (LSTM) network. The forward and backward processing of the LSTM network is a crucial part of ensuring the network can effectively capture long-term dependencies. The LSTM network is trained using the Backpropagation Through Time (BPTT) method (which optimizes network performance by adjusting weights and biases in the neural network through backpropagation in the time dimension). This method allows the partial derivatives of the model parameters to be calculated and updated through two forward and backward processing steps.

[0128] Specifically, an LSTM network processes a sequence of length S, with the forward operation starting from s=1 and the backward operation starting from s=S.

[0129] During the forward processing, the LSTM network processes each time step of the input sequence sequentially. For each time step s, the LSTM module first computes the values ​​of the forget gate, input gate, internal state, and output gate. These computation steps are as follows:

[0130] Forgotten Gate:

[0131]

[0132] Input Gate:

[0133]

[0134] Internal state:

[0135]

[0136] Output gate:

[0137]

[0138] Status output:

[0139]

[0140] During the backward processing, the LSTM network propagates the error forward step by step, starting from the last time step S. This process involves calculating the gradients of each gate and state, and updating the model parameters based on these gradients. The specific backward processing steps are as follows:

[0141] Gradient of the state output:

[0142]

[0143] Gradient of the output gate:

[0144]

[0145] Gradient of internal state:

[0146]

[0147] Gradient of the forget gate:

[0148]

[0149] Gradient of the input gate:

[0150]

[0151] Through the aforementioned forward and backward processing steps, the LSTM network can effectively capture long-term dependencies in the input sequence and continuously optimize the model parameters using the BPTT method, thereby improving the model's performance. This mechanism gives LSTM a powerful memory capacity within the network structure, enabling it to maintain high accuracy and stability when processing long sequence data.

[0152] Based on steps S110-S150 above, target recognition is performed on the temporal video frames of the video to be analyzed to obtain the moving subject recognition region. Since the first pre-trained network model is based on deformable convolution and attention mechanisms, it can adaptively adjust the receptive field and focus on the motion region (such as detailed regions like jumping and flipping) and core features during the feature extraction stage. The resulting visual feature vector can concentrate on the moving subject and reflect the detailed changes in the moving subject's movements, improving robustness to blurred and occluded scenes, reducing redundant computation, and improving the inference efficiency for subsequent moving subject tracking and action classification. Since the second pre-trained network model is based on motion continuity constraints, keypoint position prediction based on the second pre-trained network model can improve the accuracy of position prediction and conform to the motion continuity law, effectively enhancing posture description capabilities, significantly reducing keypoint jitter in high-speed motion scenes, and improving... The robustness and reliability of pose estimation are assessed. Subsequently, cross-modal feature fusion of visual and pose feature vectors yields a fused feature vector. This fully utilizes visual features and structured pose information, with the two mutually reinforcing and assisting each other. The generated fused feature vector offers comprehensive and multi-dimensional description, adapting to different lighting conditions, complex background angles, and detailed changes in motion. Further spatiotemporal dependency analysis and action classification are performed on the temporally arranged fused feature vector to effectively track the moving subject and construct motion dependencies, thereby distinguishing similar but temporally different action types and obtaining relatively accurate action recognition results. Furthermore, since the above processing utilizes pre-trained network models, the process is highly efficient, combining accuracy and real-time performance. It can be applied to training and competition analysis in various sports (such as freestyle skiing, skating, and curling). Therefore, the solution provided in this embodiment can achieve efficient and stable tracking and action recognition of multiple targets, significantly improving the accuracy, timeliness, and robustness of target detection, tracking, and action recognition.

[0153] A second exemplary embodiment of this disclosure provides a method for constructing a motion trajectory tracking and action recognition model.

[0154] Figure 7 A flowchart illustrating a method for constructing a motion trajectory tracking and action recognition model according to an embodiment of the present disclosure is shown schematically. Figure 8 The diagram illustrates the process of constructing a motion trajectory tracking and motion recognition model according to an embodiment of the present disclosure.

[0155] Reference Figure 7 and Figure 8 As shown, the method for constructing a motion trajectory tracking and action recognition model provided in this embodiment includes steps S710 to S770.

[0156] In step S710, target recognition is performed on time-series video frames of multiple videos in the training dataset to obtain the moving subject recognition region for training.

[0157] In step S720, based on the first network model with deformable convolution and attention mechanism, feature extraction is performed on the recognition region of the moving subject for training to obtain the visual feature vector representing the pose of the moving subject for training.

[0158] In step S730, the second network model based on motion continuity constraints predicts the positions of the key points of the moving subject, and obtains the training posture feature vector used to characterize the positions of the key points of the moving subject.

[0159] In step S740, the above-mentioned training visual feature vector and the above-mentioned training pose feature vector are fused across modalities to obtain the training fused feature vector.

[0160] In step S750, based on the third network model, the spatiotemporal dependency analysis and action classification processing of the time-series arranged training fusion feature vectors are performed to obtain the training action recognition results of the moving subject.

[0161] In step S760, the action labels of multiple videos in the training dataset are used as training labels to train the first network model, the second network model, and the third network model. After training, the first pre-trained network model, the second pre-trained network model, and the third pre-trained network model are obtained.

[0162] In step S770, a motion trajectory tracking and action recognition model is constructed based on the first pre-trained network model, the second pre-trained network model, and the third pre-trained network model.

[0163] The constructed motion trajectory tracking and action recognition models can be compared Figure 8 and Figure 1 As shown, it mainly includes: a first pre-trained network model, a second pre-trained network model, and a third pre-trained network model; in some embodiments, it may further include an SSD object detection model. The specific connection relationships between the various models can be understood by referring to the data processing flow above, and will not be repeated here.

[0164] In this embodiment, the first network model is a description of the first pre-trained network model during the training phase, the second network model is a description of the second pre-trained network model during the training phase, and the third network model is a description of the third pre-trained network model during the training phase. The structure of the first network model can refer to the structural block diagram and specific network structure example of the first pre-trained network model in the first embodiment, the structure of the second network model can refer to the structural block diagram and specific network structure example of the first pre-trained network model in the first embodiment, and the structure of the third network model can refer to the structural block diagram and specific network structure example of the third pre-trained network model in the first embodiment. Further details will not be provided here.

[0165] A third exemplary embodiment of this disclosure provides an apparatus for motion trajectory tracking and motion recognition.

[0166] The aforementioned motion trajectory tracking and action recognition device includes: a first target detection module, a first adaptive feature extraction module, a first pose prediction module, a first multimodal fusion module, and a first temporal action analysis module.

[0167] The aforementioned first target detection module is used to perform target recognition on the temporal video frames of the video to be analyzed, and obtain the moving subject recognition region.

[0168] The aforementioned first adaptive feature extraction module is used to extract features from the moving subject recognition region based on the first pre-trained network model with deformable convolution and attention mechanisms, and obtain a visual feature vector representing the pose of the moving subject.

[0169] The aforementioned first pose prediction module is used to predict the positions of key points of the moving subject based on the second pre-trained network model with motion continuity constraints, and to obtain pose feature vectors that characterize the positions of key points of the moving subject.

[0170] The aforementioned first multimodal fusion module is used to perform cross-modal feature fusion of the aforementioned visual feature vector and the aforementioned pose feature vector to obtain a fused feature vector.

[0171] The aforementioned first temporal action analysis module is used to perform spatiotemporal dependency analysis and action classification processing on the fusion feature vectors arranged in a temporal sequence based on the third pre-trained network model, so as to obtain the action recognition results of the moving subject.

[0172] In some embodiments, the first multimodal fusion module includes:

[0173] The dimension comparison submodule is used to compare whether the dimensions of the visual feature vector and the pose feature vector are consistent.

[0174] The alignment mapping submodule is used to perform dimension alignment mapping from the pose feature vector to the visual feature vector when the dimensions of the visual feature vector and the pose feature vector are inconsistent.

[0175] The feature fusion submodule is used to perform feature fusion processing on dimension-aligned visual feature vectors and pose feature vectors to obtain fused feature vectors.

[0176] More details and beneficial effects of this embodiment can be found in the relevant description of the first embodiment, which will not be repeated here.

[0177] The fourth exemplary embodiment of this disclosure provides an apparatus for constructing a motion trajectory tracking and action recognition model.

[0178] The aforementioned device for constructing motion trajectory tracking and action recognition models includes: a second target detection module, a second adaptive feature extraction module, a second pose prediction module, a second multimodal fusion module, a second temporal action analysis module, a training module, and a model construction module.

[0179] The second target detection module described above is used to perform target recognition on time-series video frames of multiple videos in the training dataset to obtain the recognition region of the moving subject for training.

[0180] The aforementioned second adaptive feature extraction module is used to extract features from the recognition region of the moving subject during training based on the first network model using deformable convolution and attention mechanisms, thereby obtaining a training visual feature vector representing the pose of the moving subject.

[0181] The aforementioned second pose prediction module is used to predict the positions of key points of the moving subject based on the second network model with motion continuity constraints, and to obtain a training pose feature vector that represents the position of the key points of the moving subject.

[0182] The second multimodal fusion module is used to perform cross-modal feature fusion of the training visual feature vector and the training pose feature vector to obtain the training fused feature vector.

[0183] The aforementioned second temporal action analysis module is used to perform spatiotemporal dependency analysis and action classification processing on the temporally arranged fusion feature vectors for training based on the third network model, so as to obtain the training action recognition results of the moving subject.

[0184] The training module described above is used to train the first network model, the second network model, and the third network model by using the action labels of multiple videos in the training dataset as training labels. After training, the first pre-trained network model, the second pre-trained network model, and the third pre-trained network model are obtained.

[0185] The aforementioned model building module is used to construct motion trajectory tracking and action recognition models based on the first pre-trained network model, the second pre-trained network model, and the third pre-trained network model.

[0186] More details and beneficial effects of this embodiment can be found in the descriptions of the first and second embodiments, which will not be repeated here.

[0187] Any plurality of the functional modules included in the apparatus of the third embodiment or the apparatus of the fourth embodiment described above can be combined into one module, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. At least one of the functional modules included in the apparatus of the third embodiment or the apparatus of the fourth embodiment described above can be at least partially implemented as hardware circuitry, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or implemented by any other reasonable means of integrating or packaging circuitry, or implemented in any one of software, hardware, and firmware methods, or in a suitable combination of any of these. Alternatively, at least one of the functional modules included in the apparatus of the third embodiment or the apparatus of the fourth embodiment described above can be at least partially implemented as a computer program module, which, when run, can perform corresponding functions.

[0188] The fifth exemplary embodiment of this disclosure provides an electronic device.

[0189] Figure 9 A schematic block diagram of an electronic device provided in an embodiment of the present disclosure is shown.

[0190] Reference Figure 9 As shown, the electronic device 900 provided in this embodiment includes a processor 901, a communication interface 902, a memory 903, and a communication bus 904. The processor 901, the communication interface 902, and the memory 903 communicate with each other through the communication bus 904. The memory 903 is used to store computer programs. When the processor 901 executes the program stored in the memory, it implements the motion trajectory tracking and action recognition method provided in the first embodiment or the motion trajectory tracking and action recognition model construction method provided in the second embodiment.

[0191] A sixth exemplary embodiment of this disclosure also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program that, when executed by a processor, implements the motion trajectory tracking and action recognition method provided in the first embodiment or the method for constructing a motion trajectory tracking and action recognition model provided in the second embodiment.

[0192] The computer-readable storage medium may be included in the device or apparatus described in the above embodiments; or it may exist independently and not assembled into the device or apparatus. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.

[0193] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0194] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solutions provided in this disclosure comply with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.

[0195] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0196] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for motion trajectory tracking and action recognition, characterized in that, include: Target recognition is performed on the temporal video frames of the video to be analyzed to obtain the moving subject recognition region; The first pre-trained network model based on deformable convolution and attention mechanism extracts features from the moving subject recognition region to obtain a visual feature vector representing the pose of the moving subject. The second pre-trained network model based on motion continuity constraints predicts the positions of key points of the moving subject and obtains a posture feature vector to represent the position of the key points of the moving subject. The visual feature vector and the pose feature vector are fused across modalities to obtain a fused feature vector; Based on the third pre-trained network model, spatiotemporal dependency analysis and action classification are performed on the fusion feature vectors arranged in time sequence to obtain the action recognition results of the moving subject.

2. The method according to claim 1, characterized in that, The first pre-trained network model includes: The position offset prediction submodule is used to predict the position offset of sampled pixels in the moving subject recognition region of each video frame based on the first pre-trained network layer. The deformable overlay submodule is used to update the position of the sampled pixel point according to the position offset to obtain the updated position of the sampled pixel point. The feature extraction submodule is used to extract features from the updated sampled pixel locations based on the second pre-trained network layer to obtain a preliminary output feature map. The batch normalization processing submodule is used to perform batch normalization processing on the preliminary output feature maps of each video frame to obtain the normalized output feature maps corresponding to each video frame. The nonlinear processing submodule is used to perform nonlinear mapping processing on the normalized output feature map based on the activation function to obtain an intermediate output feature map; The channel-space dual attention submodule is used to combine channel attention mechanism and spatial attention mechanism to perform attention weighting processing on the intermediate output feature map to obtain a visual feature vector representing the pose of the moving subject.

3. The method according to claim 1 or 2, characterized in that, The second pre-trained network model includes: The key point location prediction submodule is used to predict the location probability distribution of each key point in each video frame based on the third pre-trained network layer, and to obtain multiple single-point heatmaps for each video frame. The system includes at least one of a time consistency constraint submodule or a limb constraint submodule. The time consistency constraint submodule is used to analyze whether the difference between the single-point heatmaps of corresponding key points in adjacent frames is less than a preset threshold based on the single-point heatmaps of each video frame, and to perform position correction on target key points that exceed the preset threshold based on the Kalman filter algorithm, thereby obtaining the corresponding corrected single-point heatmap and outputting it to the keypoint fusion submodule. The limb constraint submodule is used to analyze the single-point heatmaps of multiple key points corresponding to each video frame based on limb connection relationships, and to perform position correction on target key points that do not conform to the limb connection relationships, thereby obtaining the corrected single-point heatmap and outputting it to the keypoint fusion submodule. The key point fusion submodule is used to flatten the features of multiple single-point heatmaps in each video frame to obtain key point position features. Based on the fourth pre-trained network layer, the key point position features are fused according to the motion continuity of adjacent frames to obtain the pose feature vector in each video frame used to represent the position of key points of the moving subject.

4. The method according to claim 3, characterized in that, The loss function of the second pre-trained network model during the training phase is a weighted sum of the heatmap loss function and the motion continuity loss function. The heatmap loss function is used to represent the deviation between the probability distribution of key point positions predicted by the third pre-trained network layer and the actual probability distribution of positions. The continuity loss function is used to represent the degree to which the changes in the positions of corresponding key points in adjacent frames conform to motion continuity when the fourth pre-trained network layer performs feature fusion.

5. The method according to claim 1, characterized in that, The temporal video frames of the video to be analyzed are obtained through at least one of the following preprocessing: frame sampling based on a preset sampling rate and resolution unification and pixel normalization; or, data enhancement processing is performed, the data enhancement including at least one of the following processing methods: random flipping, brightness variation and contrast adjustment; Alternatively, at least one process, such as splicing or merging data source segments, can be performed to obtain an integrated video to be analyzed, and then frame sampling can be performed. Target recognition is performed on the temporal video frames of the video to be analyzed to obtain the moving subject recognition region, including: Based on the object detection algorithm, moving subject recognition is performed on each temporal video frame of the video to be analyzed to obtain one or more subject boundary recognition boxes corresponding to the moving subjects; according to the distribution information within the subject boundary recognition boxes, the subject boundary recognition boxes are finely adjusted to obtain the moving subject recognition region; The visual feature vector and the pose feature vector are fused across modalities to obtain a fused feature vector, including: Compare whether the dimensions of the visual feature vector and the pose feature vector are consistent; if the dimensions of the visual feature vector and the pose feature vector are inconsistent, perform dimension alignment mapping on the pose feature vector to the visual feature vector; perform feature fusion processing on the dimension-aligned visual feature vector and the pose feature vector to obtain the fused feature vector.

6. The method according to claim 1, characterized in that, The third pre-trained network model includes: The temporal dependency determination submodule is used to perform spatiotemporal dependency analysis on the fusion feature vector of temporal arrangement based on the fifth pre-trained network layer, and obtain the temporal prediction state and the hidden vector corresponding to each time step. The action classification processing submodule is used to perform action classification based on the sixth pre-trained network layer, the temporal prediction state, and the latent vectors corresponding to each time step, to obtain the action recognition result of the moving subject.

7. A method for constructing a motion trajectory tracking and action recognition model, characterized in that, include: Target recognition is performed on time-series video frames from multiple videos in the training dataset to obtain the recognition region for moving subjects during training. The first network model based on deformable convolution and attention mechanism extracts features from the recognition region of the moving subject for training, and obtains the visual feature vector representing the pose of the moving subject for training. The second network model based on motion continuity constraints predicts the positions of key points of the moving subject and obtains a training posture feature vector to represent the positions of key points of the moving subject. The training visual feature vector and the training pose feature vector are fused across modal features to obtain the training fused feature vector; Based on the third network model, spatiotemporal dependency analysis and action classification are performed on the time-series arranged training fusion feature vectors to obtain the training action recognition results of the moving subject. The action labels of multiple videos in the training dataset are used as training labels to train the first network model, the second network model, and the third network model. After training, the first pre-trained network model, the second pre-trained network model, and the third pre-trained network model are obtained. Based on the first pre-trained network model, the second pre-trained network model, and the third pre-trained network model, a motion trajectory tracking and action recognition model is constructed.

8. A device for motion trajectory tracking and action recognition, characterized in that, include: The first target detection module is used to identify targets in the time-series video frames of the video to be analyzed, and to obtain the moving subject recognition region; The first adaptive feature extraction module is used to extract features from the moving subject recognition region of the first pre-trained network model based on deformable convolution and attention mechanism, and obtain a visual feature vector representing the pose of the moving subject. The first posture prediction module is used to predict the position of key points of the moving subject based on the second pre-trained network model with motion continuity constraints, and obtain the posture feature vector used to represent the position of key points of the moving subject. The first multimodal fusion module is used to perform cross-modal feature fusion of the visual feature vector and the pose feature vector to obtain a fused feature vector; The first temporal action analysis module is used to perform spatiotemporal dependency analysis and action classification on the fusion feature vectors arranged in a temporal sequence based on the third pre-trained network model, so as to obtain the action recognition results of the moving subject.

9. A device for constructing a motion trajectory tracking and action recognition model, characterized in that, include: The second object detection module is used to perform object recognition on time-series video frames of multiple videos in the training dataset to obtain the recognition region of moving subjects for training. The second adaptive feature extraction module is used to extract features from the recognition region of the moving subject in training based on the first network model based on deformable convolution and attention mechanism, so as to obtain the training visual feature vector representing the pose of the moving subject. The second pose prediction module is used to predict the position of key points of the moving subject in the second network model based on motion continuity constraints, and obtain the pose feature vector used to represent the position of key points of the moving subject for training. The second multimodal fusion module is used to perform cross-modal feature fusion of the training visual feature vector and the training pose feature vector to obtain the training fused feature vector; The second temporal action analysis module is used to perform spatiotemporal dependency analysis and action classification on the training fusion feature vectors arranged in a temporal sequence based on the third network model, so as to obtain the training action recognition results of the moving subject. The training module is used to train the first network model, the second network model, and the third network model by using the action labels of multiple videos in the training dataset as training labels. After training, the first pre-trained network model, the second pre-trained network model, and the third pre-trained network model are obtained. The model building module is used to build motion trajectory tracking and action recognition models based on the first pre-trained network model, the second pre-trained network model, and the third pre-trained network model.

10. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method of any one of claims 1-7.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1-7.