Spatial interaction precise recognition method and system based on multimodal fusion

By generating spatiotemporal attention maps to enhance visual features and combining feature complementary correction and image completion techniques, the problems of low action recognition accuracy and data missing in multimodal fusion methods are solved, achieving higher recognition accuracy and stability.

CN120234654BActive Publication Date: 2025-09-19ZHONGTIAN ZHILING (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510705395.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-19
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

Existing multimodal fusion methods lack effective spatiotemporal attention mechanisms and feature complementary correction mechanisms, resulting in low accuracy in complex human motion recognition and lack of data completion capabilities in actual application environments, affecting the confidence of recognition results and system stability.

Method used

By collecting human skeleton information and visual image information, a spatiotemporal attention map is generated and visual features are enhanced. Feature fusion is performed by combining feature complementary correction and cross-modal attention mechanism. When the recognition confidence is low, the spatiotemporal convolutional generative adversarial network guided by the action causal relationship map is used to complete the image.

Benefits of technology

It improves the accuracy and robustness of spatial interactive action recognition, enhances recognition performance in complex environments, solves image missing and occlusion problems, and improves the system's recognition accuracy and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234654B_ABST
    Figure CN120234654B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for accurate spatial interaction recognition based on multimodal fusion, which relates to the field of artificial intelligence technology. The method includes collecting human skeleton, visual images and voice information, using a feature extraction network to generate a spatiotemporal attention map to enhance visual features, performing feature complementary correction, and fusing the three modal information to obtain a unified feature representation; when the recognition confidence is low, a spatiotemporal convolution-based adversarial network guided by an action causal relationship map is used to complete the missing image frames, thereby improving the accuracy and robustness of spatial interaction recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to artificial intelligence technology, and in particular to a method and system for accurately identifying spatial interactions based on multimodal fusion. Background Art

[0002] With the development of human-computer interaction technology, spatial interaction technology, as an important human-computer interaction method, has been widely used in smart homes, intelligent driving, virtual reality and other fields. Traditional spatial interaction recognition methods mainly rely on single-modal data, such as visual images or skeletal information, for action recognition. However, single-modal information often has difficulty accurately identifying complex human actions, especially in real-world environments that are easily affected by factors such as lighting and occlusion.

[0003] Current multimodal fusion methods mainly include feature-level fusion and decision-level fusion. Feature-level fusion is to fuse multimodal features during the feature extraction stage, while decision-level fusion is to make a comprehensive decision based on the results of each modality's individual recognition. Existing spatial interaction recognition technologies based on multimodal fusion have the following defects and shortcomings:

[0004] The multimodal feature fusion methods in the existing technology lack an effective spatiotemporal attention mechanism and cannot fully utilize the spatiotemporal features of joint points in human skeleton information to enhance visual image information, resulting in limited expressive ability of fused features and difficulty in accurately identifying complex spatial interactive actions.

[0005] Existing multimodal fusion methods often ignore the complementarity and consistency between modalities when processing data from different modalities, and lack an effective feature complementarity correction mechanism, resulting in large information loss during the fusion process and affecting the final recognition accuracy.

[0006] Existing technologies lack effective solutions for dealing with data missing problems in actual application environments. In particular, when missing frames occur in visual image information, effective data completion cannot be performed, resulting in low confidence in the recognition results and affecting the reliability and stability of the system. Summary of the Invention

[0007] The embodiments of the present invention provide a method and system for accurately identifying spatial interactions based on multimodal fusion, which can solve the problems in the prior art.

[0008] A first aspect of an embodiment of the present invention provides a method for accurately identifying spatial interactions based on multimodal fusion, comprising:

[0009] Collect human skeleton information, visual image information and voice information;

[0010] The human skeleton information and visual image information are input into a feature extraction network. The feature extraction network generates a spatiotemporal attention map based on the spatiotemporal features of the joints in the human skeleton information. The spatiotemporal attention map is applied to the visual image information to obtain enhanced visual features. Feature complementary correction is performed on the human skeleton information and the enhanced visual features. Based on temporal modeling, feature alignment and cross-modal attention mechanism, the corrected human skeleton features and visual features are fused with speech information features to obtain a unified feature representation.

[0011] Action recognition is performed on the unified feature representation. When the confidence of the recognition result is lower than a preset confidence threshold, the missing image frames in the visual image information are supplemented based on a spatiotemporal convolutional generative adversarial network guided by the action causal relationship graph to obtain supplemented visual image information; the supplemented visual image information is re-featured with the human skeleton information and voice information to obtain the final action recognition result.

[0012] In an optional embodiment,

[0013] The feature extraction network generates a spatiotemporal attention map based on the spatiotemporal features of the joints in the human skeleton information, and the spatiotemporal attention map is applied to the visual image information to obtain enhanced visual features, including:

[0014] Obtain a skeleton sequence comprising T frames, wherein each frame of the skeleton sequence comprises N joint points, each joint point comprising two-dimensional coordinates and confidence information, wherein the two-dimensional coordinates are used to represent the position of the joint point in the image; and calculate a displacement vector of the joint point based on the two-dimensional coordinates of adjacent frames;

[0015] The skeleton sequence is constructed into a graph structure, wherein the graph structure includes a vertex set and an edge set, wherein the vertex set corresponds to the joint points in the skeleton sequence, and the edge set corresponds to the bone connection in the skeleton sequence; the graph structure is input into a graph convolutional network, and the features of the joint points are enhanced based on the neighbor node set of the joint points to obtain enhanced joint point features; multi-scale spatiotemporal feature extraction is performed on the enhanced joint point features to generate spatiotemporal position encoding and fuse it with the spatiotemporal features to obtain encoding features, and spatiotemporal self-attention features are calculated; weight fusion and constraint optimization are performed on the spatiotemporal self-attention features to obtain a spatiotemporal attention graph;

[0016] Based on the displacement vector, the motion amplitude value in the spatiotemporal attention map is calculated, and the motion amplitude value is determined by the Euclidean norm of the displacement vector at the corresponding position; the motion amplitude value is normalized to obtain an enhancement coefficient, and the product of the enhancement coefficient and the spatiotemporal attention map is applied to the visual image information to enhance the feature expression of the target action area to obtain enhanced visual features.

[0017] In an optional embodiment,

[0018] Generating a spatiotemporal attention map involves:

[0019] Performing multi-scale downsampling on the enhanced joint point features to obtain a multi-scale spatial feature map, and performing a temporal convolution operation on the multi-scale spatial feature map within a preset time window to obtain a temporal dynamic feature;

[0020] Generate spatiotemporal position coding, fuse the spatiotemporal position coding with the temporal dynamic features to obtain coding features; calculate spatiotemporal self-attention features based on the coding features, and use the spatiotemporal self-attention features to characterize the spatiotemporal dependency relationship between joint points;

[0021] Performing global average pooling on the spatiotemporal self-attention features to obtain feature importance weights, and applying the feature importance weights to the multi-scale spatial feature map and the temporal dynamic features to obtain an initial attention map;

[0022] The initial attention map is optimized based on temporal smoothness constraints, spatial structure constraints and sparse constraints, and boundary enhancement is performed to obtain a spatiotemporal attention map.

[0023] In an optional embodiment,

[0024] Performing feature complementary correction on the human skeleton information and the enhanced visual features includes:

[0025] Based on human skeleton information and enhanced visual features, the skeleton feature reliability and visual feature reliability are calculated; the skeleton feature reliability is obtained by calculating the geometric consistency, motion smoothness and structural integrity of the joint points, and the visual feature reliability is obtained by calculating the temporal consistency and scene context;

[0026] When it is detected that the reliability of the bone feature is lower than the first threshold, correction is performed based on the features of the corresponding position in the visual image information; when it is detected that the reliability of the visual feature is lower than the second threshold, correction is performed based on the joint point position in the human bone information; when the reliability of the bone feature is not lower than the first preset threshold and the reliability of the visual feature is not lower than the second preset threshold, the human bone information and enhanced visual features are used as the corrected bone features and visual features.

[0027] In an optional embodiment,

[0028] Performing feature complementarity correction involves:

[0029] A complementary correction weight matrix is ​​generated based on the reliability of the skeletal features and the reliability of the visual features; the skeletal features and the enhanced visual features are bidirectionally corrected using the complementary correction weight matrix, including: weighting the appearance information of the enhanced visual features with a first correction weight to obtain a spatial correction map, and fusing the spatial correction map with the skeletal features to obtain a corrected skeletal feature; weighting the motion trajectory information of the skeletal features with a second correction weight to obtain a motion correction map, and fusing the motion correction map with the enhanced visual features to obtain a corrected visual feature;

[0030] Temporal consistency and cross-modal consistency verification are performed on the corrected skeletal features and visual features to generate residual correction terms, which are then fused with the corresponding features to obtain the final corrected skeletal features and visual features.

[0031] In an optional embodiment,

[0032] Action recognition is performed on the unified feature representation. When the confidence level of the recognition result is lower than a preset confidence threshold, missing image frames in the visual image information are supplemented using a spatiotemporal convolutional generative adversarial network guided by an action causal relationship graph. The supplemented visual image information includes:

[0033] Action recognition is performed on the unified feature representation. Spatiotemporal features are extracted through a 3D convolutional network and combined with a temporal attention mechanism for feature fusion to obtain the action recognition result and its confidence level. When the confidence level of the action recognition result falls below a preset confidence threshold, an action causal relationship graph is constructed based on hierarchical action decomposition and temporal dependency fusion.

[0034] A causal attention mechanism is constructed based on the action causal relationship graph, and the causal attention mechanism is integrated into a generative adversarial network. The generative adversarial network includes a generator and a discriminator. The generator adopts an encoder-decoder structure and introduces the causal attention mechanism. The discriminator discriminates the spatiotemporal features of the generated results.

[0035] For missing image frames in visual image information, the generative adversarial network is used to complete them. During the completion process, reconstruction loss, adversarial loss and causal consistency constraints are combined for optimization. The causal consistency constraints are obtained based on the deviation between the generated results calculated by the action causal relationship graph and the predicted action sequence; the temporal smoothness constraints are applied to the completed visual image information to generate the final completion result.

[0036] In an optional embodiment,

[0037] The action causal relationship graph is constructed based on hierarchical action decomposition and temporal dependency fusion, including:

[0038] An action sequence is obtained based on the unified feature representation, the action sequence is decomposed into an atomic action set, and an action combination pattern in the atomic action set is identified through an action combination rule to generate a combined action set; action transition probabilities in the atomic action set and the combined action set are calculated to obtain an atomic action layer transition probability and a combined action layer transition probability, and a mapping probability from an atomic action to a combined action and a decomposition probability from a combined action to an atomic action are calculated;

[0039] Construct a dynamic time window, the dynamic time window containing the preceding and following adjacent actions of the target action, and calculate a local temporal dependency score based on the dynamic time window; count the number of co-occurrences of actions in the action sequence to construct an action co-occurrence matrix, and calculate a global temporal dependency score; perform a weighted fusion of the local temporal dependency score and the global temporal dependency score to obtain a fused dependency score, and update the atomic action layer transition probability, the combined action layer transition probability, the mapping probability, and the decomposition probability based on the fused dependency score;

[0040] The actions in the atomic action set and combined action set are taken as nodes. Based on the action recognition results, the updated atomic action layer transfer probability, combined action layer transfer probability, mapping probability and decomposition probability are adjusted online and normalized post-processing is performed as the edge weights to construct an action causal relationship graph with a hierarchical structure.

[0041] A second aspect of an embodiment of the present invention provides a spatial interaction accurate recognition system based on multimodal fusion, including:

[0042] The first unit is used to collect human skeleton information, visual image information and voice information;

[0043] The second unit is configured to input the human skeleton information and visual image information into a feature extraction network, wherein the feature extraction network generates a spatiotemporal attention map based on the spatiotemporal features of the joints in the human skeleton information, and applies the spatiotemporal attention map to the visual image information to obtain enhanced visual features; perform feature complementary correction on the human skeleton information and the enhanced visual features; and fuse the corrected human skeleton features and visual features with speech information features based on temporal modeling, feature alignment, and cross-modal attention mechanisms to obtain a unified feature representation;

[0044] The third unit is used to perform action recognition on the unified feature representation. When the confidence of the recognition result is lower than a preset confidence threshold, the missing image frames in the visual image information are supplemented based on the spatiotemporal convolutional generative adversarial network guided by the action causal relationship graph to obtain the supplemented visual image information; the supplemented visual image information is re-featured with the human skeleton information and voice information to obtain the final action recognition result.

[0045] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:

[0046] processor;

[0047] a memory for storing processor-executable instructions;

[0048] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0049] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0050] The present invention collects three modal data types, namely human skeleton information, visual image information and voice information, and enhances visual features based on spatiotemporal attention maps, thereby achieving effective fusion of multimodal information and improving the accuracy of spatial interactive action recognition.

[0051] The present invention adopts a feature complementary correction mechanism to complementarily correct human skeletal features with visual features. It combines temporal modeling, feature alignment and cross-modal attention mechanism to solve the heterogeneity problem between different modal information, so that each modal information can complement and enhance each other, thereby improving the robustness of action recognition.

[0052] The present invention introduces a spatiotemporal convolutional generative adversarial network guided by an action causal relationship graph, and performs visual image frame completion processing for situations with low recognition confidence, effectively solving problems such as image missing and occlusion that may occur in practical applications, and further improving the system's recognition performance and adaptability in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 Schematic diagram of the process of a method for accurate spatial interaction recognition based on multimodal fusion according to an embodiment of the present invention;

[0054] Figure 2 Comparison of attention accuracy between multi-scale spatiotemporal feature fusion and triple constraint optimization;

[0055] Figure 3 Performance comparison chart on action prediction accuracy. DETAILED DESCRIPTION

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0057] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0058] Figure 1 FIG. 1 is a flow chart of a method for accurately identifying spatial interactions based on multimodal fusion according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0059] Collect human skeleton information, visual image information and voice information;

[0060] The human skeleton information and visual image information are input into a feature extraction network. The feature extraction network generates a spatiotemporal attention map based on the spatiotemporal features of the joints in the human skeleton information. The spatiotemporal attention map is applied to the visual image information to obtain enhanced visual features. Feature complementary correction is performed on the human skeleton information and the enhanced visual features. Based on temporal modeling, feature alignment and cross-modal attention mechanism, the corrected human skeleton features and visual features are fused with speech information features to obtain a unified feature representation.

[0061] Action recognition is performed on the unified feature representation. When the confidence of the recognition result is lower than a preset confidence threshold, the missing image frames in the visual image information are supplemented based on a spatiotemporal convolutional generative adversarial network guided by the action causal relationship graph to obtain supplemented visual image information; the supplemented visual image information is re-featured with the human skeleton information and voice information to obtain the final action recognition result.

[0062] In an optional embodiment, the feature extraction network generates a spatiotemporal attention map based on the spatiotemporal features of the joints in the human skeleton information, and applies the spatiotemporal attention map to the visual image information to obtain enhanced visual features, including:

[0063] Obtain a skeleton sequence comprising T frames, wherein each frame of the skeleton sequence comprises N joint points, each joint point comprising two-dimensional coordinates and confidence information, wherein the two-dimensional coordinates are used to represent the position of the joint point in the image; and calculate a displacement vector of the joint point based on the two-dimensional coordinates of adjacent frames;

[0064] The skeleton sequence is constructed into a graph structure, wherein the graph structure includes a vertex set and an edge set, wherein the vertex set corresponds to the joint points in the skeleton sequence, and the edge set corresponds to the bone connection in the skeleton sequence; the graph structure is input into a graph convolutional network, and the features of the joint points are enhanced based on the neighbor node set of the joint points to obtain enhanced joint point features; multi-scale spatiotemporal feature extraction is performed on the enhanced joint point features to generate spatiotemporal position encoding and fuse it with the spatiotemporal features to obtain encoding features, and spatiotemporal self-attention features are calculated; weight fusion and constraint optimization are performed on the spatiotemporal self-attention features to obtain a spatiotemporal attention graph;

[0065] Based on the displacement vector, the motion amplitude value in the spatiotemporal attention map is calculated, and the motion amplitude value is determined by the Euclidean norm of the displacement vector at the corresponding position; the motion amplitude value is normalized to obtain an enhancement coefficient, and the product of the enhancement coefficient and the spatiotemporal attention map is applied to the visual image information to enhance the feature expression of the target action area to obtain enhanced visual features.

[0066] Exemplarily, a human skeleton sequence data containing T frames is obtained, where T can be set to 64 frames, and each frame contains N joint points, where N can be 18 standard human joint points, such as the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. Each joint point contains a two-dimensional coordinate (x, y) and a confidence value c ranging from 0 to 1, which is used to indicate the reliability of the joint point detection. For the i-th joint point between the t-th frame and the t+1-th frame, its displacement vector is calculated as the current frame coordinate minus the previous frame coordinate. For example, for a waving action, the wrist joint point will have a large horizontal displacement between consecutive frames.

[0067] The skeleton sequence is constructed as a graph structure G=(V,E), where V is a vertex set and E is an edge set. The vertex set V corresponds to N joint points, and the edge set E corresponds to the bone connections. For example, there is an edge between the shoulder and the elbow, and there is an edge between the elbow and the wrist. In actual implementation, these connections can be predefined, such as the shoulder (number 1) is connected to the elbow (number 2), the elbow is connected to the wrist (number 3), and so on. The constructed graph structure is input into the graph convolutional network (GCN). GCN enhances the joint point features through the following steps: For each joint point i, determine its neighbor node set N(i), including the joint points directly connected to it. For example, for the elbow node, its neighbor node set includes the shoulder and wrist. Then, the spatial weight matrix W between the neighbor nodes is calculated, and the weight value w ij Calculated based on the Euclidean distance between joint points i and j: w ij = exp(-||p i -p j ||2 / σ2), where p i and p jare the spatial positions of nodes i and j, respectively, and σ is an adjustable scaling parameter. This embodiment adopts a two-layer GCN structure with an input feature dimension of 3 (x, y, c), an intermediate layer feature dimension of 64, and an output feature dimension of 128. Taking the first layer of GCN as an example, for joint point i, its updated feature is the weighted sum of its own feature and the features of all connected joints. In this way, the feature representation of each joint not only contains its own information, but also integrates the structural and motion information of neighboring nodes, thereby obtaining an enhanced joint feature.

[0068] Multi-scale spatiotemporal feature extraction is performed on the enhanced joint point features to generate spatiotemporal position encoding and fused with the spatiotemporal features to finally obtain the spatiotemporal attention map.

[0069] For the i-th joint in frame t, its motion amplitude is determined by calculating the Euclidean norm of the joint's displacement vector between two adjacent frames. Specifically, the square root of the difference between the squared abscissa and the squared ordinate of the joint in the current frame t and the next frame t+1 is added together. All motion amplitudes are normalized to obtain an enhancement coefficient between 0 and 1. This normalization is performed by subtracting the minimum motion amplitude from the current motion amplitude, and then dividing the result by the difference between the maximum and minimum motion amplitudes. To prevent the denominator from being zero, a small constant, such as 0.00001, is added to the denominator. The enhancement coefficient is element-wise multiplied by the spatiotemporal attention map and then applied to the visual features extracted from the input video. The dimensions of the visual features include the temporal dimension T, the feature map height H, the width W, and the number of channels C. The region surrounding each joint is located in the feature map according to its two-dimensional coordinates, and the features of that region are weightedly enhanced by multiplying the corresponding enhancement coefficient by the attention value. For example, if the enhancement coefficient of a joint point is calculated to be 0.8 and the attention value of the corresponding position is 0.7, the two values ​​are multiplied together to obtain 0.56 as the final weight coefficient to enhance the features of this area. In this way, the final enhanced visual features highlight the characteristics of the target action area.

[0070] The present invention achieves effective feature enhancement of the target action area by generating a spatiotemporal attention map based on the spatiotemporal features of skeletal joints and applying it to visual image information. The skeletal sequence is constructed as a graph structure and features are extracted through a graph convolutional network. Multi-scale spatiotemporal feature extraction and self-attention mechanism are combined to accurately capture the spatial relationship and temporal changes between joints. The motion amplitude is calculated by displacement vector and combined with the spatiotemporal attention map to achieve adaptive enhancement of visual features, significantly improve the perception ability of dynamic action areas, enhance the discriminability and robustness of feature expression, and provide more accurate feature representation for subsequent action recognition tasks.

[0071] In an optional embodiment, generating a spatiotemporal attention map includes:

[0072] Performing multi-scale downsampling on the enhanced joint point features to obtain a multi-scale spatial feature map, and performing a temporal convolution operation on the multi-scale spatial feature map within a preset time window to obtain a temporal dynamic feature;

[0073] Generate spatiotemporal position coding, fuse the spatiotemporal position coding with the temporal dynamic features to obtain coding features; calculate spatiotemporal self-attention features based on the coding features, and use the spatiotemporal self-attention features to characterize the spatiotemporal dependency relationship between joint points;

[0074] Performing global average pooling on the spatiotemporal self-attention features to obtain feature importance weights, and applying the feature importance weights to the multi-scale spatial feature map and the temporal dynamic features to obtain an initial attention map;

[0075] The initial attention map is optimized based on temporal smoothness constraints, spatial structure constraints and sparse constraints, and boundary enhancement is performed to obtain a spatiotemporal attention map.

[0076] For example, convolution kernels of different sizes are used to downsample joint features. Three different convolution kernel sizes, such as 3×3, 5×5, and 7×7, are used to convolve the joint features. The stride size for each convolution kernel is set to 2, and the padding parameter is set to 1, resulting in feature maps of three different scales. For example, for joint features with an input size of 64×25 (where 64 represents the feature dimension and 25 represents the number of joints), three spatial feature maps of sizes 32×13, 16×7, and 8×4 are obtained after downsampling. These three feature maps of different scales are resized to the same spatial size of 32×13 using bilinear interpolation. They are then weighted and fused based on feature importance, with larger-scale feature maps given higher weights. The weights for the three scales can be set to 0.4, 0.35, and 0.25, respectively.

[0077] Temporal convolution is performed on the multi-scale spatial feature maps within a preset time window to obtain temporal dynamic features. The time window length is set to T (for example, T = 10 frames). One-dimensional temporal convolution is performed on the multi-scale spatial feature maps within each time window. One-dimensional convolution is used with a kernel size of 3, a stride of 1, and a padding of 1 to maintain the temporal dimension. For each scale spatial feature map, the corresponding temporal dynamic features are obtained after temporal convolution.

[0078] Spatiotemporal position coding consists of two parts: temporal position coding and spatial position coding. Temporal position coding is used to represent the relative position relationship between different time frames. It is generated through a combination of sine and cosine functions. For each time step, a position coding vector with the same feature dimension is generated, so that the features of different time steps can distinguish their temporal relationship. Spatial position coding is used to represent the relative position relationship between different joints in the human skeleton. First, an adjacency matrix representing the connection relationship of the joints is constructed, and the shortest path distance between any two joints is calculated. Then, based on this distance information, a coding vector is generated for each joint that reflects its relative position in the skeleton structure. Finally, the temporal position coding and spatial position coding are spliced ​​in the feature dimension to obtain a complete spatiotemporal position coding.

[0079] The spatiotemporal position encoding is fused with the temporal dynamic features to obtain the encoded features. This fusion operation is achieved through element-wise addition. That is, for each temporal dynamic feature, it is added to the spatiotemporal position encoding of the corresponding dimension to obtain the encoded features. Based on the encoded features, spatiotemporal self-attention features are calculated to characterize the spatiotemporal dependencies between joints. Specifically, the encoded features at each scale are passed through three independent linear transformation layers to obtain the query vector, key vector, and value vector, respectively. The output dimension of the linear transformation layer can be set to one-eighth of the input dimension. Then, the query vector and the key vector are matrix multiplied and divided by the scaling factor (for example, 8). Softmax normalization is performed to obtain the attention weight matrix. Finally, the attention weight matrix is ​​multiplied by the value vector to obtain the self-attention features. Multiple attention heads can be set, each head independently calculating the attention features. Finally, the features of all heads are concatenated and linearly transformed to obtain the final spatiotemporal self-attention features.

[0080] The spatiotemporal self-attention features are average pooled across both the temporal and spatial dimensions to obtain channel-level feature vectors. These are then mapped to a value between 0 and 1 using a sigmoid function to form feature importance weights. For example, for a spatiotemporal self-attention feature of size 64 × T × 25, global average pooling yields a 64-dimensional weight vector. These feature importance weights are then multiplied channel-wise with the multi-scale spatial feature map and the temporal dynamics feature. The results are then weighted and summed to form the initial attention map.

[0081] The temporal smoothness constraint is implemented by calculating the differences between the attention maps of adjacent time frames. This difference is calculated using the Euclidean distance metric. The difference values ​​of the attention maps of each pair of adjacent frames are calculated and summed to obtain the temporal smoothness loss. The spatial structure constraint considers the natural connectivity of the human skeleton. For each pair of adjacent joints, an attenuation coefficient is calculated based on their distance in the skeleton graph. The closer the joints are, the greater the penalty for the difference in their attention values, thus obtaining the spatial structure loss. The sparse constraint is implemented by calculating the L1 norm of the attention map, which forces the attention values ​​to be concentrated on a few key joints. The loss terms of these three constraints are weighted and combined, with weights set to 0.3, 0.5, and 0.2, respectively, to obtain the total constraint loss, which is used to optimize the initial attention map.

[0082] Boundary enhancement is achieved through the following steps: applying the Sobel operator in the horizontal and vertical directions to the optimized attention map for edge detection, resulting in horizontal and vertical edge response maps; taking the square root of the sum of the squares of the edge response maps in both directions to obtain an edge strength map; normalizing the edge strength map to a value between 0 and 1; and weighted fusion of the normalized edge strength map with the original attention map, with fusion weights set to 0.85 and 0.15. In this way, the boundaries between the joint point regions in the attention map are enhanced, making the attention distribution of different joint points clearer, and ultimately obtaining a boundary-enhanced spatiotemporal attention map.

[0083] Figure 2 This is a comparison chart of the attention accuracy of multi-scale spatiotemporal feature fusion and triple constraint optimization. The horizontal axis represents the complexity of the action, ranging from low to extremely high; the vertical axis represents the percentage of attention accuracy, ranging from 50% to 100%. The three compared methods are: the present invention (multi-scale + triple constraint): using multi-scale feature extraction and a triple optimization strategy of temporal smoothing constraint, spatial structure constraint, and sparse constraint; single scale + single constraint: using only single-scale feature extraction and a single constraint; multi-scale + single constraint: using multi-scale feature extraction but only a single constraint. It can be clearly seen from the figure that the present invention achieves the highest attention accuracy at all complexities, especially in high-complexity and extremely high-complexity action scenes. In low-complexity scenes, the present invention achieves an accuracy of 94.7%, while the single-scale + single constraint method only achieves 82.2%. In extremely high-complexity scenes, the present invention still maintains a high accuracy of 83.2%, while the single-scale + single constraint method drops to 56.7%. This result fully demonstrates the superiority and stability of multi-scale spatiotemporal feature fusion and triple-constraint optimization in complex action scenarios. In particular, as the complexity of the action increases, the performance gap between the method of the present invention and the traditional method becomes more and more significant.

[0084] The spatiotemporal attention map generation method proposed in the present invention extracts the multi-scale spatiotemporal features of joint points through multi-scale downsampling and temporal convolution operations, and combines spatiotemporal position encoding to enhance the spatiotemporal expression ability of features; introduces a spatiotemporal self-attention mechanism to capture the complex dependencies between joint points, and generates feature importance weights through global average pooling to achieve adaptive enhancement of features; designs a triple constraint optimization strategy and boundary enhancement technology to ensure the temporal smoothness, spatial structure rationality and regional focus of attention distribution, effectively improves the accuracy and discriminability of the attention map, and provides high-quality guidance information for subsequent visual feature enhancement.

[0085] In an optional embodiment, performing feature complementary correction on the human skeleton information and the enhanced visual features includes:

[0086] Based on human skeleton information and enhanced visual features, the skeleton feature reliability and visual feature reliability are calculated; the skeleton feature reliability is obtained by calculating the geometric consistency, motion smoothness and structural integrity of the joint points, and the visual feature reliability is obtained by calculating the temporal consistency and scene context;

[0087] When it is detected that the reliability of the bone feature is lower than the first threshold, correction is performed based on the features of the corresponding position in the visual image information; when it is detected that the reliability of the visual feature is lower than the second threshold, correction is performed based on the joint point position in the human bone information; when the reliability of the bone feature is not lower than the first preset threshold and the reliability of the visual feature is not lower than the second preset threshold, the human bone information and enhanced visual features are used as the corrected bone features and visual features.

[0088] For example, when calculating skeletal feature reliability, the assessment of geometric consistency primarily examines whether the distances between adjacent joints conform to the anatomical characteristics of the human skeleton. The system pre-stores standard length ratio data for various skeletal segments, such as an approximately 1.1:1 ratio for the upper arm to the lower arm, and a 1:1.1 ratio for the thigh to the lower leg. During detection, the distances between the detected adjacent joints are calculated and compared with the expected bone lengths. If the actual length of a skeletal segment deviates from the expected length by more than 30%, the geometric consistency score for that segment will be lowered. For example, if the detected upper arm length is 40 cm and the lower arm length is 60 cm, this clearly violates the normal skeletal proportions, and the system will lower the geometric consistency score for that segment. To assess motion smoothness, the motion trajectories of the joints are analyzed across multiple consecutive frames. In normal human motion, the position changes of joints should be smooth and continuous, without sudden changes. The system calculates the rate of change of each joint's position across five consecutive frames. If the position change of a joint between consecutive frames exceeds a preset threshold (e.g., a distance change of more than 30 cm or an angle change of more than 45 degrees), the joint is considered to have low motion smoothness. For example, if it is detected that the elbow joint has suddenly moved 50 cm between two consecutive frames, while other parts of the human body have not changed much, the system will lower the motion smoothness score of this joint. The structural integrity assessment mainly examines whether the detected bone structure is complete, checks whether all key joints (such as head, shoulders, hips, knees, ankles, etc.) are detected, and whether the connection relationship between these joints is correct. If the system detects that some key joints are missing, or the connection relationship between the joints is abnormal (such as the left ankle is connected to the right knee), it is considered that the integrity of the bone structure is damaged. For example, in a certain detection, the system identified all the joints of the upper body, but only detected the joints of the left leg in the lower body, and the right leg was completely missing. At this time, the system will significantly reduce the bone structure integrity score.

[0089] The final skeletal feature reliability score was calculated by combining the scores of the above three aspects, ranging from 0 to 1. The higher the score, the higher the reliability.

[0090] When calculating visual feature reliability, temporal consistency assessment focuses on the changes in visual features across multiple consecutive frames. This involves analyzing the degree of change in the feature vector at each pixel position across 10 consecutive frames. For normal human motion, the visual features of the corresponding region should exhibit a certain degree of consistency. The standard deviation of the feature vector in each region along the temporal dimension is calculated. If the standard deviation is too large (e.g., exceeding 200% of the regional average), the region is considered to have low temporal consistency. For example, when a person moves quickly from a bright area to a shadowed area, this may cause a sudden change in clothing color, resulting in a lower temporal consistency score for this region. Scene context assessment examines the degree of match between visual features and the scene environment. Visual feature reliability is assessed by analyzing environmental factors in the image, such as lighting conditions and background complexity. Visual feature reliability is generally lower in environments with complex backgrounds or uneven lighting. For example, in strong backlighting, a person's outline may become blurred, reducing the scene context score for the corresponding region. Similarly, when objects in the background have a similar color to the person's clothing, their boundaries may become unclear, similarly reducing the scene context score for these regions.

[0091] By comprehensively evaluating the temporal consistency and scene context, the system calculates the final visual feature reliability score, which ranges from 0 to 1.

[0092] When the reliability of the detected skeletal features falls below a first threshold (e.g., 0.6), corrections are performed using features from the corresponding position in the visual image information. The corresponding joint point is located on the visual feature map, and visual features are extracted from that location and its surrounding area (e.g., a 20×20 pixel area). The joint point's position is then re-estimated based on these visual features. For example, if the reliability of the skeletal features detected for the right knee joint is only 0.45 (below the threshold of 0.6), visual features from that area (e.g., a clear leg edge line) are used to accurately re-locate the right knee joint.

[0093] When the reliability of a detected visual feature is lower than a second threshold (e.g., 0.7), the feature will be corrected based on the joint locations in the human skeleton information. Based on the reliable joint locations, the feature weights of the corresponding area on the visual feature map will be adjusted, increasing the weight of features consistent with the skeletal structure and reducing the weight of inconsistent features. For example, if the system detects that the reliability of the visual features in a person's arm area is 0.65 (lower than the threshold of 0.7), while the skeletal features in that area (e.g., shoulder, elbow, and wrist joint locations) are relatively reliable, the system will adjust the visual features of the arm area based on these joint locations to make them more consistent with the constraints of the skeletal structure.

[0094] The feature complementary correction method proposed in the present invention realizes the complementary advantages of multimodal information by accurately evaluating the reliability of skeletal features and visual features; the reliability of skeletal features is evaluated from three dimensions: geometric consistency, motion smoothness, and structural integrity, and the reliability of visual features is evaluated from two aspects: temporal consistency and scene context, thus establishing a comprehensive feature quality evaluation system; the adaptive correction mechanism based on reliability scoring can intelligently introduce complementary modal information for correction when single modal information is unreliable, effectively solving the problem of feature degradation in complex scenes, significantly improving the robustness and accuracy of feature representation, and providing high-quality feature input for subsequent action recognition tasks.

[0095] In an optional embodiment, performing feature complementation correction includes:

[0096] A complementary correction weight matrix is ​​generated based on the reliability of the skeletal features and the reliability of the visual features; the skeletal features and the enhanced visual features are bidirectionally corrected using the complementary correction weight matrix, including: weighting the appearance information of the enhanced visual features with a first correction weight to obtain a spatial correction map, and fusing the spatial correction map with the skeletal features to obtain a corrected skeletal feature; weighting the motion trajectory information of the skeletal features with a second correction weight to obtain a motion correction map, and fusing the motion correction map with the enhanced visual features to obtain a corrected visual feature;

[0097] Temporal consistency and cross-modal consistency verification are performed on the corrected skeletal features and visual features to generate residual correction terms, which are then fused with the corresponding features to obtain the final corrected skeletal features and visual features.

[0098] For example, when generating a complementary correction weight matrix based on skeletal feature reliability and visual feature reliability, the skeletal feature reliability score and the visual feature reliability score are first normalized to the range of 0 to 1. For each joint position, if its skeletal feature reliability is 0.8 and its visual feature reliability is 0.6, the first correction weight is set to 0.6 and the second correction weight is set to 0.8, reflecting the complementary effect of the high-reliability modality on the low-reliability modality.

[0099] During the bidirectional correction process, the appearance information of the enhanced visual features needs to be processed. For example, if the visual features of the arm region contain 64 channels, each channel represents a different appearance feature, such as edges and textures. These features are multiplied by the first correction weight. Regions with higher weights (such as the elbow region with a reliability of 0.9) retain their original features, while regions with lower weights (such as the wrist region with a reliability of 0.3) are suppressed. The resulting spatial correction image is then fused with the skeletal features using a weighted summation method, with weights set to 0.6 and 0.4.

[0100] To process the motion trajectory information of skeletal features, the position change sequence of each joint point within 10 consecutive frames is extracted. For example, for the wrist joint, its horizontal and vertical displacements are recorded over 10 consecutive frames. This trajectory information is multiplied by a second correction weight. For joints with reliable motion trajectories (such as the shoulder with a reliability of 0.85), their complete motion information is retained; for joints with unreliable motion trajectories (such as the ankle with a reliability of 0.4), their influence is reduced. The resulting motion correction map is then fused with the enhanced visual features, with fusion weights set to 0.55 and 0.45.

[0101] During the temporal consistency verification phase, the corrected features are analyzed in a time window with a window size of 15 frames. The feature change trends of each joint across these frames are analyzed, including changes in position and appearance. If a joint feature undergoes drastic changes between consecutive frames, such as a sudden change in position exceeding 25 centimeters or a change in feature value exceeding 70% of its original value, a temporal inconsistency is considered. For example, if the left hand is detected to have a sudden displacement of 40 centimeters in frame 8, while the displacement in all other frames is within 5 centimeters, a correction term is generated for the feature in frame 8. Cross-modal consistency verification primarily examines the degree of match between skeletal features and visual features. At each joint position, the skeletal localization result is compared with the position indicated by the visual features. If the difference between the two exceeds a preset threshold (e.g., 15 pixels), a corresponding residual correction term is generated. For example, if the skeletal features indicate that the right elbow is located at image coordinates (200, 300), while the visual features indicate (180, 320), the positional deviation is calculated and a correction is generated. The generation of the residual correction term takes into account the verification results of temporal consistency and cross-modal consistency. For the right elbow in the above example, if the temporal consistency verification shows a positional deviation of 10 pixels and the cross-modal consistency verification shows a positional deviation of 20 pixels, a residual correction term of 15 pixels is generated. The residual correction term is fused with the corresponding feature, with a fusion ratio of 0.85 for the original feature and 0.15 for the residual correction term.

[0102] Through a complementary correction mechanism, the present invention effectively utilizes the advantages of the two modalities to complement each other's deficiencies, thereby improving the reliability and accuracy of features; through a two-way correction and multi-level verification mechanism, the quality and reliability of features are significantly improved, providing a better feature foundation for subsequent action recognition and behavior analysis tasks.

[0103] In an optional embodiment, action recognition is performed on the unified feature representation. When the confidence level of the recognition result is lower than a preset confidence threshold, missing image frames in the visual image information are supplemented using a spatiotemporal convolutional generative adversarial network guided by an action causal relationship graph. The supplemented visual image information includes:

[0104] Action recognition is performed on the unified feature representation. Spatiotemporal features are extracted through a 3D convolutional network and combined with a temporal attention mechanism for feature fusion to obtain the action recognition result and its confidence level. When the confidence level of the action recognition result falls below a preset confidence threshold, an action causal relationship graph is constructed based on hierarchical action decomposition and temporal dependency fusion.

[0105] A causal attention mechanism is constructed based on the action causal relationship graph, and the causal attention mechanism is integrated into a generative adversarial network. The generative adversarial network includes a generator and a discriminator. The generator adopts an encoder-decoder structure and introduces the causal attention mechanism. The discriminator discriminates the spatiotemporal features of the generated results.

[0106] For missing image frames in visual image information, the generative adversarial network is used to complete them. During the completion process, reconstruction loss, adversarial loss and causal consistency constraints are combined for optimization. The causal consistency constraints are obtained based on the deviation between the generated results calculated by the action causal relationship graph and the predicted action sequence; the temporal smoothness constraints are applied to the completed visual image information to generate the final completion result.

[0107] For example, time series modeling is performed on the corrected skeletal, visual, and speech features. For the skeletal feature sequence, a bidirectional long short-term memory network is used to process 32 consecutive frames of data with 256 hidden units to extract motion feature information at each time step. For the visual feature sequence, a temporal convolutional network is used to process 32 consecutive image frames with a convolution kernel size of 3 and 128 channels to extract the temporal dependencies of the visual features. For the speech features, the audio signal is divided into time segments aligned with the video frames. 32-dimensional Mel-frequency cepstral coefficient features are extracted from each segment and then extracted using a one-dimensional convolutional network. During feature alignment, the features of the three modalities are mapped to the same feature space with a uniform feature dimension of 256. A time window mechanism is used for feature alignment with a window size of 8 frames. Within each time window, a similarity matrix is ​​calculated between features from different modalities to guide feature alignment. For example, if a high correlation is detected between the motion of hand joints and the energy variation of the speech signal in a certain action segment, the alignment weight between the two modal features is strengthened. The cross-modal attention mechanism achieves feature enhancement by calculating the mutual relationships between features from different modalities. For each modality, two feature transformation branches are constructed to generate query features and key-value features, respectively. For example, the query features generated by skeletal features can be weighted with the key-value features of visual features to achieve skeleton-guided visual feature enhancement. To balance the importance of different modalities, learnable modality importance weights are set, initially set to one-third. In this way, features from the three modalities are fused into a unified feature representation, maintaining the feature dimensionality at 256.

[0108] When using a unified feature representation for action recognition, a 3D convolutional network is used to extract spatiotemporal features. The network consists of five convolutional modules, each consisting of two 3D convolutional layers. The first convolutional module takes as input a 32-frame × 256-dimensional feature sequence. Features are extracted using a 3×3×3 convolution kernel with a stride of 1, resulting in 64 output feature channels. Subsequent convolutional modules increase the number of feature channels to 128, 256, 512, and 512, respectively. A max-pooling layer with a stride of 2 is then used to reduce the spatiotemporal resolution of the features. To capture long-term dependencies in action sequences, a temporal attention mechanism is introduced based on the convolutional features. The feature sequence is divided into four time segments, each containing eight frames. For each time segment, a correlation score is calculated with respect to the other segments. This correlation score takes into account the semantic similarity and temporal position of the features. Segments with higher similarity and closer temporal distance receive higher attention weights. For example, in a waving action, the key segments of arm raising and lowering receive higher cross-correlation attention weights. The features with added temporal attention weights are mapped to the action category space through two fully connected layers. The output dimension of the first fully connected layer is 1024, and the output dimension of the second layer is the number of action categories. A softmax function is used to obtain the probability distribution of each action category. The category with the highest probability is the recognition result, and its probability value is used as the confidence level.

[0109] When the confidence level of the action recognition result falls below a preset confidence threshold (e.g., 0.75), the visual image completion module is activated. An action causal relationship graph is constructed based on hierarchical action decomposition and temporal dependency fusion.

[0110] The causal attention mechanism is constructed as follows: From the action causal graph, the set of preceding and subsequent possible action nodes at the current moment is extracted, each containing the five most relevant action nodes. Feature embedding is performed on these nodes, mapping each action node into a 64-dimensional vector, while the edge relationships (transition probabilities) between nodes are mapped into 32-dimensional vectors. These node and edge features are processed using a graph attention network, which consists of three graph attention layers, each with eight attention heads, and outputs a 128-dimensional graph feature vector. To enable interaction with visual features, a cross-attention module is designed. This module interacts the visual feature map F (512×8×8) output by the encoder with the graph feature vector G. First, G is expanded to a 512-dimensional vector through linear projection. The dot product similarity between G and each spatial position feature in F is calculated to form an attention map A (8×8). Finally, A is applied to F to obtain a causally weighted feature map F' (512×8×8). Furthermore, a dynamic gating mechanism is introduced to adaptively adjust the influence of causal attention based on the current confidence level. Specifically, a sigmoid function is used to map the confidence level to a weight α between 0 and 1. The final feature is F'' = α·F' + (1-α)·F, achieving an adaptive fusion of causal and visual information. This causal attention mechanism effectively introduces action semantics into the visual generation process, ensuring that the generated results conform to the action logic.

[0111] The generative adversarial network consists of two parts: a generator and a discriminator. The generator uses an encoder-decoder structure. The encoder processes the input image sequence and maps it to a latent feature space; the decoder reconstructs the latent features into a target image. The encoder contains six downsampling convolutional blocks, each consisting of a convolutional layer, an instance normalization layer, and a LeakyReLU activation function. The convolutional layers use 4×4 kernels with a stride of 2 and padding of 1. The number of channels is 64, 128, 256, 512, 512, and 512, respectively. The decoder contains six upsampling transposed convolutional blocks, each consisting of a transposed convolutional layer, an instance normalization layer, and a ReLU activation function. The transposed convolutional layers use a 4×4 kernel with a stride of 2 and padding of 1. The number of channels is 512, 512, 256, 128, 64, and 3 (RGB channels), respectively.

[0112] The causal attention mechanism is integrated into the bottleneck layer between the encoder and decoder, fusing the feature map (512×8×8) output by the encoder with the action causal relationship encoding result. The nodes and edges related to the current action in the action causal relationship map are extracted and encoded into a 128-dimensional vector using a graph convolutional network. This vector is mapped to the same dimension as the number of channels in the feature map (512) using a fully connected layer. The similarity between this vector and each spatial position in the feature map is calculated to generate an attention map. The attention map is applied to the original feature map to obtain a weighted feature map as the input to the decoder.

[0113] The discriminator uses the PatchGAN structure to perform spatiotemporal feature discrimination on the generated results. This discriminator not only judges the authenticity of a single image frame, but also evaluates temporal coherence. The discriminator input is a continuous sequence of 3 frames, and extracts spatiotemporal features through 3D convolution. The specific structure contains 5 3D convolution blocks, each of which consists of a 3D convolution layer, an instance normalization layer, and a LeakyReLU activation function. The convolution layer uses a 4×4×4 convolution kernel (time×height×width), a stride of (1,2,2), and the number of channels is 64, 128, 256, 512, and 1, respectively. The final output is a discriminant score map, which represents the authenticity score of each part of the input sequence.

[0114] The above-mentioned generative adversarial network is used to complete missing image frames in visual image information. Missing frames in the video sequence are detected, and then 2-3 frames before and after the missing frame are extracted as context information input to the generative network. For example, for a 30-frame-per-second video, if the 15th frame is missing, the system will extract frames 12-14 and 16-18 as conditional inputs to generate the completion result of frame 15. During the completion process, three loss functions are combined for optimization: reconstruction loss, adversarial loss, and causal consistency constraint. The reconstruction loss measures the pixel-level difference between the generated image and the real image. The L1 loss is used to calculate the sum of the absolute differences between the corresponding pixel values ​​of the generated frame and the real frame, and then the average is taken. The adversarial loss comes from the feedback of the discriminator and adopts the form of WGAN-GP (Wasserstein GAN with gradient penalty) to improve the authenticity and temporal coherence of the generated image. The causal consistency constraint is based on the deviation between the generated result and the predicted action sequence calculated based on the action causal relationship graph. The generated image frame is input into the pre-trained action recognition model to obtain the action prediction result of the frame; at the same time, based on the action causal relationship graph and the known previous action, the most likely action of the current frame is predicted; then the difference between the two results is calculated as the causal consistency constraint. In practice, cross entropy is used to calculate the difference between the probability distributions of two actions. The three loss functions are combined with different weights. In this embodiment, the reconstruction loss weight is 0.6, the adversarial loss weight is 0.2, and the causal consistency constraint weight is 0.2. These weights can be adjusted according to specific application requirements.

[0115] The completed visual image information is constrained for temporal smoothness to ensure a smooth visual transition between the generated image frame and the preceding and following frames, avoiding flickering or discontinuities. Optical flow is calculated for three consecutive frames (previous frame, completed frame, and next frame), and interpolation and smoothing are then performed based on this optical flow information. A bidirectional optical flow calculation method is used: forward optical flow from the previous frame to the completed frame, and backward optical flow from the next frame to the completed frame. The two are then weighted averaged based on temporal position to obtain the final flow field. Based on this flow field, local adjustments are made to the completed frame to ensure the naturalness of the motion transition.

[0116] The completed visual image information is re-integrated with the human skeleton information and voice information to obtain the final action recognition result. If the confidence level of the recognition result is not lower than the preset confidence threshold, the action recognition result is regarded as the final action recognition result.

[0117] The present invention introduces an action causal relationship graph to guide the visual image completion process, effectively solving the limitations of traditional methods in dealing with the problem of missing information in action recognition. It not only considers the visual quality of the image, but also emphasizes the logical coherence of the action, significantly improving the accuracy and robustness of action recognition. Especially in complex environments and scenarios with rapidly changing actions, it can effectively restore missing information, providing reliable technical support for intelligent monitoring, human-computer interaction, virtual reality and other fields, and has broad application prospects.

[0118] In an optional embodiment, constructing an action causal relationship graph based on hierarchical action decomposition and temporal dependency fusion includes:

[0119] An action sequence is obtained based on the unified feature representation, the action sequence is decomposed into an atomic action set, and an action combination pattern in the atomic action set is identified through an action combination rule to generate a combined action set; action transition probabilities in the atomic action set and the combined action set are calculated to obtain an atomic action layer transition probability and a combined action layer transition probability, and a mapping probability from an atomic action to a combined action and a decomposition probability from a combined action to an atomic action are calculated;

[0120] Construct a dynamic time window, the dynamic time window containing the preceding and following adjacent actions of the target action, and calculate a local temporal dependency score based on the dynamic time window; count the number of co-occurrences of actions in the action sequence to construct an action co-occurrence matrix, and calculate a global temporal dependency score; perform a weighted fusion of the local temporal dependency score and the global temporal dependency score to obtain a fused dependency score, and update the atomic action layer transition probability, the combined action layer transition probability, the mapping probability, and the decomposition probability based on the fused dependency score;

[0121] The actions in the atomic action set and combined action set are taken as nodes. Based on the action recognition results, the updated atomic action layer transfer probability, combined action layer transfer probability, mapping probability and decomposition probability are adjusted online and normalized post-processing is performed as the edge weights to construct an action causal relationship graph with a hierarchical structure.

[0122] For example, based on the unified feature representation, action sequences are generated. Continuous video is segmented using a sliding window approach with a window size of 2 seconds and a step size of 0.5 seconds. The video segments within each window are then identified using a pre-trained action recognition model to determine the action category. For example, for a 30-second video, the system extracts 57 windows, generating an action sequence such as "stand - bend - pick up an object - stand upright - walk - put down an object."

[0123] Action sequences are decomposed into a set of atomic actions. Atomic actions are indecomposable basic units of movement and form the basis of complex actions. The system pre-defines 50 atomic actions, including "bend the elbow," "raise the arm," "lean the torso forward," and "rotate the head." The decomposition process uses a pre-trained pose estimation model to extract key points. The atomic actions are then determined through rule matching and template comparison. For example, the action "pick up an object" might be decomposed into atomic actions such as "extend the arm," "clench the fingers," "bend the elbow," and "retract the arm." Action combination rules are used to identify action combination patterns within the set of atomic actions and generate a set of combined actions. These action combination rules are designed based on a temporal pattern mining algorithm to identify frequently occurring atomic action sequences. With a minimum support of 0.05 and a minimum confidence of 0.6, an improved PrefixSpan algorithm is used to mine frequent sequential patterns. The mined patterns undergo semantic filtering to form combined actions. For example, the system might find that the combination of "extend the arm," "clench the fingers," and "bend the elbow" occurs frequently and define it as the combined action "grasp." In this embodiment, the system recognizes a total of 120 common combination actions.

[0124] Calculate the transition probabilities for the atomic and combined action sets. For each atomic action set, count the number of transitions between adjacent atomic actions and divide this by the total number of occurrences of the preceding action to obtain the atomic-action-level transition probability. For example, if "stretch out your arms" is followed by "clench your fingers" 200 times, and "stretch out your arms" appears 250 times in total, the transition probability is 0.8. Similarly, calculate the combined action-level transition probability. For the mapping probability from atomic actions to combined actions, count the number of times an atomic action belongs to a specific combined action and divide it by the total number of occurrences of that atomic action. For example, if "clench your fingers" appears 350 times in "grasp," and 400 times in total, the mapping probability is 0.875. The decomposition probability from combined actions to atomic actions is the probability that a combined action contains a specific atomic action. This is calculated by dividing the number of times a combined action contains that atomic action by the total number of combined actions.

[0125] The size of the dynamic time window is adaptively adjusted based on the duration of the action, typically encompassing the two actions preceding and following the target action. Within the window, the system considers the time interval and order relationship between actions to calculate a local temporal dependency score. Time intervals are processed using an exponential decay function, with shorter intervals receiving higher weights. Order relationships are weighted differently based on relative position: directly adjacent actions receive a weight of 1, actions separated by one interval receive a weight of 0.7, and actions separated by two intervals receive a weight of 0.4. The local temporal dependency score is the product of the time interval weight and the order relationship weight. For example, if actions A and B are closely connected (with a time interval of 0.2 seconds), the time interval weight is 0.98; if A is directly followed by B, the order relationship weight is 1, resulting in a local dependency score of 0.98. Simultaneously, the number of co-occurrences of actions in the action sequence is counted to construct an action co-occurrence matrix, which is then used to calculate the global temporal dependency score. Action co-occurrence refers to the occurrence of two actions in the same sequence, regardless of their specific position. The number of co-occurrences of each action pair in the training data is counted to construct an N×N co-occurrence matrix (N is the total number of actions). The global temporal dependency score is calculated using point-wise mutual information, measuring the degree of non-randomness of the co-occurrence of two actions. Specifically, the logarithm of the ratio of the number of times two actions co-occur and the product of their independent probabilities of occurrence is calculated. A higher value indicates stronger dependency. For example, if the actions "pick up an object" and "put down an object" co-occur 850 times in 1000 sequences, with respective probabilities of 0.3 and 0.25, their global dependency score is high, approximately 1.55.

[0126] The local temporal dependency score is weighted and fused with the global temporal dependency score to produce a fused dependency score. This fusion is linearly weighted, with weights set based on the application scenario. In this example, the local dependency score is weighted 0.7, and the global dependency score is weighted 0.3. For example, for the action pair "pick up object - put down object," if the local dependency score is 0.85 and the global dependency score is 1.55, the fused dependency score is 0.85 × 0.7 + 1.55 × 0.3 = 1.06.

[0127] Based on the fused dependency scores, the transition probabilities at the atomic action level, the combined action level, the mapping probabilities, and the decomposition probabilities are updated. The update process uses a weighted average method, with the original probability weighted at 0.6 and the dependency score influence weighted at 0.4. The dependency scores are mapped to the 0-1 range using a sigmoid function and then weighted averaged with the original probabilities. For example, if the original transition probability from "extending the arm" to "clenching the fingers" is 0.8 and the corresponding fused dependency score is 1.06, which is 0.74 after mapping, the updated transition probability is 0.8 × 0.6 + 0.74 × 0.4 = 0.776.

[0128] A hierarchical action causal relationship graph is constructed, using actions from atomic and composite action sets as nodes. Edges in the graph are determined based on updated probabilities. These include edges within the atomic action layer (transitions between atomic actions), edges within the composite action layer (transitions between composite actions), and edges between layers (mapping edges from atomic actions to composite actions and decomposing edges from composite actions to atomic actions). These edge weights are adjusted online based on the action recognition results, with the adjustment amount proportional to the recognition confidence. For example, if the system recognizes the action "pick up an object" with a confidence level of 0.9, the edge weight associated with that action is increased by a larger amount (e.g., 0.1); when the confidence level is lower (e.g., 0.6), the adjustment amount is smaller (e.g., 0.03). The adjusted weights are normalized to ensure that the sum of all edge weights originating from the same node is 1.

[0129] In practical applications, the graph is optimized for storage and indexing, using an adjacency list structure. Action nodes are organized hierarchically, with each node recording its outgoing target node and weight. To improve retrieval efficiency, the system establishes a hash map from action names to node IDs. Furthermore, the graph is updated after processing every 1,000 action sequences to ensure it adapts to new action patterns.

[0130] Figure 3This is a performance comparison chart on action prediction accuracy, showing the prediction accuracy comparison results of three different methods in action scenarios of various complexities. The horizontal axis represents the complexity of the action scenario, from left to right, it is "single-person simple action", "single-person complex action", "two-person interactive action", "multi-person complex interaction" and "cross-scene multi-person interaction", with the complexity gradually increasing; the vertical axis represents the percentage of action prediction accuracy, ranging from 50% to 100%. The figure compares three different action prediction methods: the present invention (hierarchical action causal graph) adopts a two-layer structure of atomic action layer and combined action layer, integrating local temporal dependency and global temporal dependency; the single-layer action transition model only uses a single layer to model the action transition relationship; the traditional Markov model is based on the state transition probability modeling of fixed time steps. It can be clearly seen from the chart that as the complexity of the action scenario increases, the prediction accuracy of the three methods shows a downward trend, but the decline of the method of the present invention is the smallest. In simple, single-person action scenarios, the differences between the three methods are relatively small (95.2% for the present invention, 92.7% for the single-layer model, and 90.8% for the Markov model). However, in highly complex scenarios involving multi-person interactions across multiple scenes, the accuracy gap between the present method (79.4%) and the other two methods (63.1% for the single-layer model and 60.7% for the Markov model) widens significantly. This performance difference is primarily attributed to the present method's hierarchical structure, which more effectively captures long-range dependencies in complex action sequences, and its strategy of fusing local and global temporal dependencies, which provides more comprehensive action context.

[0131] The present invention constructs an action causal relationship graph through hierarchical action decomposition and temporal dependency fusion, effectively solving the problem that traditional methods are difficult to simultaneously capture the fine-grained features and high-level semantics of actions; it not only considers the direct transfer relationship between actions, but also integrates local and global temporal dependency information, significantly improving the accuracy of action prediction and understanding; the hierarchical structure design enables the graph to express the details of basic action units and reflect the overall semantics of combined actions, providing strong knowledge support for tasks such as video completion and action prediction, and has broad application prospects.

[0132] A second aspect of an embodiment of the present invention provides a spatial interaction accurate recognition system based on multimodal fusion, including:

[0133] The first unit is used to collect human skeleton information, visual image information and voice information;

[0134] The second unit is configured to input the human skeleton information and visual image information into a feature extraction network, wherein the feature extraction network generates a spatiotemporal attention map based on the spatiotemporal features of the joints in the human skeleton information, and applies the spatiotemporal attention map to the visual image information to obtain enhanced visual features; perform feature complementary correction on the human skeleton information and the enhanced visual features; and fuse the corrected human skeleton features and visual features with speech information features based on temporal modeling, feature alignment, and cross-modal attention mechanisms to obtain a unified feature representation;

[0135] The third unit is used to perform action recognition on the unified feature representation. When the confidence of the recognition result is lower than a preset confidence threshold, the missing image frames in the visual image information are supplemented based on the spatiotemporal convolutional generative adversarial network guided by the action causal relationship graph to obtain the supplemented visual image information; the supplemented visual image information is re-featured with the human skeleton information and voice information to obtain the final action recognition result.

[0136] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:

[0137] processor;

[0138] a memory for storing processor-executable instructions;

[0139] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0140] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0141] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A spatial interaction accurate recognition method based on multimodal fusion, characterized by: include: Collect human skeleton information, visual image information and voice information; Inputting the human skeleton information and visual image information into a feature extraction network, the feature extraction network generating a spatiotemporal attention map based on the spatiotemporal features of the joints in the human skeleton information, and applying the spatiotemporal attention map to the visual image information to obtain enhanced visual features; Performing feature complementary correction on the human skeleton information and enhanced visual features, including: generating a complementary correction weight matrix based on the skeleton feature reliability and the visual feature reliability; using the complementary correction weight matrix to perform bidirectional correction on the skeleton features and the enhanced visual features, including: weighting the appearance information of the enhanced visual feature with a first correction weight to obtain a spatial correction map, and fusing the spatial correction map with the skeleton features to obtain corrected skeleton features; weighting the motion trajectory information of the skeleton features with a second correction weight to obtain a motion correction map, and fusing the motion correction map with the enhanced visual features to obtain corrected visual features; based on temporal modeling, feature alignment and cross-modal attention mechanism, fusing the corrected human skeleton features and visual features with speech information features to obtain a unified feature representation; Action recognition is performed on the unified feature representation. When the confidence of the recognition result is lower than a preset confidence threshold, the missing image frames in the visual image information are supplemented based on the spatiotemporal convolutional generative adversarial network guided by the action causal relationship graph, including: constructing an action causal relationship graph based on hierarchical action decomposition and temporal dependency fusion; constructing a causal attention mechanism based on the action causal relationship graph, and integrating the causal attention mechanism into the generative adversarial network, the generative adversarial network including a generator and a discriminator, the generator adopts an encoder-decoder structure and introduces the causal attention mechanism, and the discriminator performs spatiotemporal feature discrimination on the generated results; the missing image frames in the visual image information are supplemented using the generative adversarial network to obtain the supplemented visual image information; the supplemented visual image information is re-featured with the human skeleton information and voice information to obtain the final action recognition result.

2. The method according to claim 1, characterized in that The feature extraction network generates a spatiotemporal attention map based on the spatiotemporal features of the joints in the human skeleton information, and the spatiotemporal attention map is applied to the visual image information to obtain enhanced visual features, including: Obtain a skeleton sequence comprising T frames, wherein each frame of the skeleton sequence comprises N joint points, each joint point comprising two-dimensional coordinates and confidence information, wherein the two-dimensional coordinates are used to represent the position of the joint point in the image; and calculate a displacement vector of the joint point based on the two-dimensional coordinates of adjacent frames; The skeleton sequence is constructed into a graph structure, wherein the graph structure includes a vertex set and an edge set, wherein the vertex set corresponds to the joint points in the skeleton sequence, and the edge set corresponds to the bone connection in the skeleton sequence; the graph structure is input into a graph convolutional network, and the features of the joint points are enhanced based on the neighbor node set of the joint points to obtain enhanced joint point features; multi-scale spatiotemporal feature extraction is performed on the enhanced joint point features to generate spatiotemporal position encoding and fuse it with the spatiotemporal features to obtain encoding features, and spatiotemporal self-attention features are calculated; weight fusion and constraint optimization are performed on the spatiotemporal self-attention features to obtain a spatiotemporal attention graph; Based on the displacement vector, the motion amplitude value in the spatiotemporal attention map is calculated, and the motion amplitude value is determined by the Euclidean norm of the displacement vector at the corresponding position; the motion amplitude value is normalized to obtain an enhancement coefficient, and the product of the enhancement coefficient and the spatiotemporal attention map is applied to the visual image information to enhance the feature expression of the target action area to obtain enhanced visual features.

3. The method according to claim 2, characterized in that Generating a spatiotemporal attention map involves: Performing multi-scale downsampling on the enhanced joint point features to obtain a multi-scale spatial feature map, and performing a temporal convolution operation on the multi-scale spatial feature map within a preset time window to obtain a temporal dynamic feature; Generate spatiotemporal position coding, fuse the spatiotemporal position coding with the temporal dynamic features to obtain coding features; calculate spatiotemporal self-attention features based on the coding features, and use the spatiotemporal self-attention features to characterize the spatiotemporal dependency relationship between joint points; Performing global average pooling on the spatiotemporal self-attention features to obtain feature importance weights, and applying the feature importance weights to the multi-scale spatial feature map and the temporal dynamic features to obtain an initial attention map; The initial attention map is optimized based on temporal smoothness constraints, spatial structure constraints and sparse constraints, and boundary enhancement is performed to obtain a spatiotemporal attention map.

4. The method according to claim 1, wherein Performing feature complementary correction on the human skeleton information and the enhanced visual features includes: Based on human skeleton information and enhanced visual features, the skeleton feature reliability and visual feature reliability are calculated; the skeleton feature reliability is obtained by calculating the geometric consistency, motion smoothness and structural integrity of the joint points, and the visual feature reliability is obtained by calculating the temporal consistency and scene context; When it is detected that the reliability of the bone feature is lower than the first threshold, correction is performed based on the features of the corresponding position in the visual image information; when it is detected that the reliability of the visual feature is lower than the second threshold, correction is performed based on the joint point position in the human bone information; when the reliability of the bone feature is not lower than the first preset threshold and the reliability of the visual feature is not lower than the second preset threshold, the human bone information and enhanced visual features are used as the corrected bone features and visual features.

5. The method according to claim 4, characterized in that Performing feature complementarity correction involves: Temporal consistency and cross-modal consistency verification are performed on the corrected skeletal features and visual features to generate residual correction terms, which are then fused with the corresponding features to obtain the final corrected skeletal features and visual features.

6. The method according to claim 1, characterized in that Action recognition is performed on the unified feature representation. When the confidence level of the recognition result is lower than a preset confidence threshold, missing image frames in the visual image information are supplemented using a spatiotemporal convolutional generative adversarial network guided by an action causal relationship graph. The supplemented visual image information includes: Action recognition is performed on the unified feature representation. Spatiotemporal features are extracted through a 3D convolutional network and combined with a temporal attention mechanism for feature fusion to obtain the action recognition result and its confidence level. When the confidence level of the action recognition result falls below a preset confidence threshold, an action causal relationship graph is constructed based on hierarchical action decomposition and temporal dependency fusion. During the completion process, reconstruction loss, adversarial loss and causal consistency constraints are combined for optimization. The causal consistency constraints are obtained based on the deviation between the generated results calculated by the action causal relationship graph and the predicted action sequence; the temporal smoothness constraints are applied to the completed visual image information to generate the final completion result.

7. The method according to claim 6, characterized in that The action causal relationship graph is constructed based on hierarchical action decomposition and temporal dependency fusion, including: An action sequence is obtained based on the unified feature representation, the action sequence is decomposed into an atomic action set, and an action combination pattern in the atomic action set is identified through an action combination rule to generate a combined action set; action transition probabilities in the atomic action set and the combined action set are calculated to obtain an atomic action layer transition probability and a combined action layer transition probability, and a mapping probability from an atomic action to a combined action and a decomposition probability from a combined action to an atomic action are calculated; Construct a dynamic time window, the dynamic time window containing the preceding and following adjacent actions of the target action, and calculate a local temporal dependency score based on the dynamic time window; count the number of co-occurrences of actions in the action sequence to construct an action co-occurrence matrix, and calculate a global temporal dependency score; perform a weighted fusion of the local temporal dependency score and the global temporal dependency score to obtain a fused dependency score, and update the atomic action layer transition probability, the combined action layer transition probability, the mapping probability, and the decomposition probability based on the fused dependency score; The actions in the atomic action set and combined action set are taken as nodes. Based on the action recognition results, the updated atomic action layer transfer probability, combined action layer transfer probability, mapping probability and decomposition probability are adjusted online and normalized post-processing is performed as the edge weights to construct an action causal relationship graph with a hierarchical structure.

8. A spatial interaction precision recognition system based on multimodal fusion, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to collect human skeleton information, visual image information and voice information; A second unit is configured to input the human skeleton information and visual image information into a feature extraction network, wherein the feature extraction network generates a spatiotemporal attention map based on the spatiotemporal features of the joints in the human skeleton information, and applies the spatiotemporal attention map to the visual image information to obtain enhanced visual features; Performing feature complementary correction on the human skeleton information and enhanced visual features; fusing the corrected human skeleton features and visual features with speech information features based on temporal modeling, feature alignment, and cross-modal attention mechanisms to obtain a unified feature representation; A third unit is configured to perform action recognition on the unified feature representation, and when the confidence level of the recognition result is lower than a preset confidence threshold, the missing image frames in the visual image information are supplemented by a spatiotemporal convolutional generative adversarial network guided by an action causal relationship graph to obtain supplemented visual image information; The completed visual image information is re-featured with the human skeleton information and voice information to obtain the final action recognition result.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Sign language recognition method, device and system based on vision and skeleton information fusion

    CN116152926A

  • Multi-mode video anomaly detection method combining RGB appearance, skeleton posture and audio information and related equipment

    CN119007288A