A refined behavior recognition method and system for virtual teaching scenarios

By using a multi-layered perception adaptive behavior measurement model, the problem of poor behavior recognition accuracy in virtual teaching scenarios is solved. It achieves simultaneous capture of local details and global semantics, thereby improving the accuracy and stability of behavior recognition in virtual teaching scenarios.

CN121861730BActive Publication Date: 2026-05-26GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG UNIV OF TECH
Filing Date
2026-03-19
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing methods for refined behavior recognition in virtual teaching scenarios suffer from performance degradation when faced with changing factors such as lighting fluctuations, viewpoint switching, and individual posture differences, resulting in poor accuracy of behavior recognition results.

Method used

A multi-layer perceptual adaptive behavior measurement model is adopted, including a feature extraction network, a feature enhancement network, and a pose decoding network. Through multi-scale feature extraction, feature enhancement, and pose decoding, behavior recognition results are generated. This model integrates features through a pyramid attention module, a local-global context fusion module, and a multi-head attention feature fusion module, and combines adaptive activation adjustment, deformable convolution enhancement, and dynamic feature calibration to improve environmental adaptability.

Benefits of technology

It significantly enhances the ability to recognize refined behaviors in virtual teaching scenarios, improves the accuracy and stability of behavior recognition results, and can capture both local details and global semantic information simultaneously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121861730B_ABST
    Figure CN121861730B_ABST
Patent Text Reader

Abstract

This invention discloses a refined behavior recognition method and system for virtual teaching scenarios, solving the technical problem that existing refined behavior recognition methods for virtual teaching scenarios result in poor accuracy of the final behavior recognition results. The method includes acquiring behavior images of a virtual teaching scenario and inputting these images into a multi-layered perception adaptive behavior measurement model, which includes a feature extraction network, a feature enhancement network, and a pose decoding network. The feature extraction network extracts multi-scale features from the behavior images of the virtual teaching scenario, outputting multi-scale fused features. The feature enhancement network enhances the multi-scale fused features, outputting target enhanced features. The target enhanced features are then input into the pose decoding network for pose decoding to generate the behavior recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for refined behavior recognition in virtual teaching scenarios. Background Technology

[0002] As an educational model based on new educational information technology, virtual teaching can break through the time and space limitations of traditional teaching and provide learners with a safe and repeatable practice environment, especially showing great potential in the field of high-risk and high-cost skills training.

[0003] However, the in-depth development of virtual teaching faces severe technical bottlenecks: it is difficult to capture and evaluate learners' refined operational behaviors in the virtual environment in real time and accurately, including: inaccurate capture of subtle changes in movements, insufficient stability in recognition under complex environments, and the difficulty in balancing real-time performance and accuracy. These issues affect the human-computer interaction and assessment effectiveness of virtual teaching systems. A large-scale survey by Mortazavi et al. showed that approximately 40% of learners experienced a decline in their learning experience due to the inability to obtain accurate behavioral assessments. Therefore, overcoming the technical bottlenecks in refined behavior recognition has become crucial for promoting the development of virtual teaching.

[0004] Existing methods for refined behavior recognition in virtual teaching scenarios introduce multi-attention mechanisms to dynamically adjust feature weights, improving the model's ability to focus on key areas. However, these methods lack the ability to dynamically adjust online. When faced with common factors in virtual teaching, such as lighting fluctuations, viewpoint switching, and individual posture differences, the model's performance significantly degrades, resulting in poor accuracy of the final behavior recognition results. Summary of the Invention

[0005] This invention provides a refined behavior recognition method and system for virtual teaching scenarios, which solves the technical problem that existing refined behavior recognition methods for virtual teaching scenarios result in poor accuracy of the final behavior recognition results.

[0006] The first aspect of this invention provides a refined behavior recognition method for virtual teaching scenarios, comprising:

[0007] Acquire behavioral images of a virtual teaching scenario and input the behavioral images of the virtual teaching scenario into a behavior measurement model based on multi-layer perception and adaptive behavior, wherein the behavior measurement model based on multi-layer perception and adaptive behavior includes a feature extraction network, a feature enhancement network, and a pose decoding network;

[0008] The feature extraction network is used to extract multi-scale features from the behavioral images of the virtual teaching scene and output multi-scale fused features.

[0009] The feature enhancement network described above is used to enhance the multi-scale fused features, and the target enhanced features are output.

[0010] The target enhancement features are input into the pose decoding network for pose decoding to generate behavior recognition results.

[0011] Optionally, the step of extracting multi-scale features from the virtual teaching scene behavior images through the feature extraction network and outputting multi-scale fused features includes:

[0012] The behavioral images of the virtual teaching scene are feature-encoded to output basic behavioral features;

[0013] High-resolution feature extraction and low-resolution feature extraction are performed on the basic behavioral features respectively, and high-resolution behavioral features and low-resolution behavioral features are output.

[0014] Multi-scale feature fusion is performed on the high-resolution features and low-resolution features of the behavior to output multi-scale fused features.

[0015] Optionally, the feature enhancement network includes a pyramid attention module, a local-global context fusion module, and a multi-head attention feature fusion module; the step of using the feature enhancement network to enhance the multi-scale fused features and outputting the target enhanced features includes:

[0016] The multi-scale fusion features are integrated using the pyramid attention module to output multi-scale integrated features.

[0017] A local-global context fusion module is used to perform local-global context fusion on the multi-scale integrated features to generate multi-scale local-global fused features.

[0018] The multi-scale local-global fusion features are input into the multi-head attention feature fusion module for deep fusion, and the enhanced fusion features are output.

[0019] An adaptive activation adjustment mechanism is used to adaptively adjust the enhanced fusion features, and the adjusted enhanced fusion features are output.

[0020] The adjusted enhanced fusion features are deformed by performing deformed convolution on the adaptive spatial sampling mechanism to generate deformed convolution enhanced features.

[0021] An adaptive feature calibration mechanism is used to adaptively calibrate the deformable convolution enhancement features to generate target enhancement features.

[0022] Optionally, the step of integrating the multi-scale fusion features through the pyramid attention module and outputting multi-scale integrated features includes:

[0023] Perform multi-scale pooling on the multi-scale fusion feature and output the multi-scale pooled feature corresponding to the multi-scale fusion feature.

[0024] Spatial attention and channel attention are calculated for the multi-scale pooling features respectively, and the spatial attention and channel attention corresponding to the multi-scale pooling features are output.

[0025] Adaptively weighted and fused the spatial attention and channel attention corresponding to the multi-scale pooling features to output the fused attention corresponding to the multi-scale pooling features;

[0026] The multi-scale pooling features and the corresponding fusion attention are integrated to output multi-scale integrated features.

[0027] Optionally, the step of using a local-global context fusion module to perform local-global context fusion on the multi-scale integrated features to generate multi-scale local-global fused features includes:

[0028] Local fine-grained features are extracted from the multi-scale integrated features to output multi-scale local fine-grained features;

[0029] The multi-scale local fine features are dynamically weighted to generate multi-scale local weighted features;

[0030] Global context features are extracted from the multi-scale integrated features to generate multi-scale global context features;

[0031] The multi-scale global context features and the multi-scale local weighted features are fused based on a gating mechanism to generate multi-scale local-global fused features.

[0032] Optionally, the step of inputting the multi-scale local-global fusion features into a multi-head attention feature fusion module for deep fusion and outputting enhanced fusion features includes:

[0033] The multi-scale local-global fusion features are compressed to output multi-scale compressed features;

[0034] Perform multi-head self-attention calculation on the multi-scale compressed features and output the multi-head attention weights corresponding to the multi-scale compressed features;

[0035] The multi-head attention weights corresponding to the multi-scale compressed features are concatenated to obtain multi-scale attention features.

[0036] Calculate the target weights corresponding to the multi-scale attention features;

[0037] The multi-scale attention features and their corresponding target weights are weighted and summed to output the enhanced fusion features.

[0038] Optionally, the step of inputting the target enhancement features into the pose decoding network for pose decoding to generate behavior recognition results includes:

[0039] Pose attention calculation is performed on the target enhancement features to obtain attention enhancement features;

[0040] Based on the attention enhancement features, perform heatmap prediction and output a keypoint heatmap;

[0041] Based on the attention enhancement features, displacement field prediction is performed, and the key point displacement field is output.

[0042] Behavior pose estimation is performed based on the key point heatmap and the key point displacement field to generate behavior recognition results.

[0043] The second aspect of this invention provides a refined behavior recognition system for virtual teaching scenarios, comprising:

[0044] The acquisition module is used to acquire behavioral images of a virtual teaching scene and input the behavioral images of the virtual teaching scene into a behavior measurement model based on multi-layer perception and adaptive behavior, wherein the behavior measurement model based on multi-layer perception and adaptive behavior includes a feature extraction network, a feature enhancement network and a pose decoding network.

[0045] The extraction module is used to extract multi-scale features from the virtual teaching scene behavior images through the feature extraction network and output multi-scale fused features;

[0046] The enhancement module is used to enhance the multi-scale fused features using the feature enhancement network and output the target enhanced features;

[0047] The decoding module is used to input the target enhancement features into the pose decoding network for pose decoding and generate behavior recognition results.

[0048] A third aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the refined behavior recognition method for virtual teaching scenarios as described above.

[0049] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the refined behavior recognition method for virtual teaching scenarios as described above.

[0050] As can be seen from the above technical solutions, the present invention has the following advantages:

[0051] The present invention provides a method for refined behavior recognition in virtual teaching scenarios. It acquires behavior images of virtual teaching scenarios and inputs these images into a multi-layered perception adaptive behavior measurement model. This model includes a feature extraction network, a feature enhancement network, and a pose decoding network. The feature extraction network extracts multi-scale features from the virtual teaching scenario behavior images, outputting multi-scale fused features. The feature enhancement network then enhances these multi-scale fused features, outputting target-enhanced features. These target-enhanced features are input into the pose decoding network for pose decoding, generating behavior recognition results. Based on this method, the present invention, through multi-scale feature extraction by the feature extraction network, preserves both local details of learner operations in the virtual teaching scenario and captures global action semantics. Simultaneously, the feature enhancement network further enhances the multi-scale fused features, achieving deep integration of multi-level features and significantly improving the ability to capture refined teaching actions, thereby enhancing the accuracy of refined behavior recognition results in virtual teaching scenarios. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 This is a flowchart illustrating the steps of a refined behavior recognition method for virtual teaching scenarios provided in Embodiment 1 of the present invention.

[0054] Figure 2 This is a schematic diagram of the structure of the behavior measurement model based on multilayer perception and adaptation provided in Embodiment 1 of the present invention;

[0055] Figure 3 This is a schematic diagram of the feature enhancement network provided in Embodiment 1 of the present invention;

[0056] Figure 4 This is a schematic diagram of the framework of the adaptive attitude estimation strategy provided in Embodiment 1 of the present invention;

[0057] Figure 5 This is a schematic diagram of the hardware and software architecture provided in Embodiment 1 of the present invention;

[0058] Figure 6 This is a structural block diagram of a refined behavior recognition system for virtual teaching scenarios provided in Embodiment 2 of the present invention. Detailed Implementation

[0059] This invention provides a method and system for refined behavior recognition in virtual teaching scenarios, which solves the technical problem that existing refined behavior recognition methods for virtual teaching scenarios result in poor accuracy of the final behavior recognition results.

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It should be noted that in the optional embodiments of the present invention, the object information and other related data involved require the permission or consent of the object when the embodiments of the present invention are applied to specific products or technologies, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments of the present invention involve data related to the object, it needs to be obtained with the authorization and consent of the object, the authorization and consent of the relevant departments, and in compliance with the relevant laws, regulations, and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required, and the embodiments also need to be implemented with the authorization and consent of the object.

[0061] Please see Figure 1 , Figure 1 The flowchart illustrates the steps of a refined behavior recognition method for virtual teaching scenarios provided in Embodiment 1 of the present invention.

[0062] This invention provides a refined behavior recognition method for virtual teaching scenarios, comprising:

[0063] Step 101: Obtain behavioral images of the virtual teaching scene and input the behavioral images of the virtual teaching scene into the behavior measurement model based on multilayer perception and adaptation. The behavior measurement model based on multilayer perception and adaptation includes a feature extraction network, a feature enhancement network, and a pose decoding network.

[0064] Behavioral images in virtual teaching scenarios refer to visual data that records learners' operational behaviors in a virtual teaching environment, including learners' fine operational actions, interaction processes with the virtual teaching scenario, and related scenario information.

[0065] It should be noted that the process involves acquiring behavioral images of a virtual teaching scenario and inputting these images into a multi-layered perception adaptive behavior measurement model. This model includes a feature extraction network, a feature enhancement network, and a pose decoding network. The feature extraction network extracts multi-scale features from the virtual teaching scenario behavioral images and outputs multi-scale fused features. The feature enhancement network then strengthens these multi-scale fused features to obtain target enhanced features. Finally, the pose decoding network performs pose analysis on the target enhanced features, ultimately generating refined behavior recognition results in the virtual teaching scenario.

[0066] Among them, such as Figure 2 As shown, this invention, building upon existing research on feature representation and dynamic adaptation, proposes a Multi-level Perceptual Adaptive Behavior Measurement (MPABM) model. MPABM achieves multi-level feature fusion from local details to global semantics through cascaded processing of a Pyramid Attention Module (PAM), a Local Global Context Module (LGCM), and an Adaptive Feature Fusion Module (AFFM). Then, it combines an adaptive strategy of adaptive activation adjustment, deformable convolution enhancement, and dynamic feature calibration to improve robustness to complex virtual teaching environments. Finally, to verify its effectiveness, this invention constructs a Virtual Teaching Behavior (VTB) dataset. Experiments show that MPABM significantly outperforms existing methods in both accuracy and efficiency, promoting the refined development of educational information technology in virtual teaching.

[0067] The multilayer perceptual adaptive behavior measurement model (MPABM) consists of a feature extraction network, a feature enhancement network, and a pose decoding network, forming a three-layer MPABM architecture. This architecture employs a progressive processing strategy, constructing multi-scale feature representations through the feature extraction network, performing deep semantic fusion via the feature enhancement network, and finally achieving adaptive output through the pose decoding network. This systematically addresses the dual challenges of insufficient feature representation depth and limited environmental adaptability in refined behavior recognition within virtual teaching scenarios. Figure 2 As shown, Figure 2Part (A) in the diagram represents the feature extraction network, which constructs a multi-scale feature pyramid based on the HRNet backbone (High-Resolution Network). This network comprises four components: initial feature extraction, a high-resolution branch, a low-resolution branch, and multi-scale feature fusion, transforming the input image into a multi-scale feature representation. Figure 2 Part (B) is the feature enhancement network, which consists of three cascaded modules: PAM, LGCM, and AFFM. It receives multi-scale features and generates highly semantically fused features through layer-by-layer processing. Figure 2 Part (C) in the diagram is the pose decoding network, which consists of a pose attention (PA) module, a heatmap branch, and a displacement field branch.

[0068] Step 102: Extract multi-scale features from the behavioral images of the virtual teaching scene using a feature extraction network, and output multi-scale fused features.

[0069] It should be noted that the feature extraction network is used to extract multi-scale features from the behavior images of the virtual teaching scene and output multi-scale fusion features. Specifically, the behavior images of the virtual teaching scene are first encoded to obtain basic behavioral features. Then, high-resolution behavioral features and low-resolution behavioral features are extracted through high-resolution and low-resolution branches respectively. Finally, the two types of features are fused to output multi-scale fusion features.

[0070] Furthermore, step 102 may include the following sub-steps:

[0071] S21. Encode the behavioral images in the virtual teaching scenario and output the basic behavioral features;

[0072] S22. Perform high-resolution feature extraction and low-resolution feature extraction on the basic behavioral features respectively, and output high-resolution behavioral features and low-resolution behavioral features.

[0073] S23. Perform multi-scale feature fusion on high-resolution behavioral features and low-resolution behavioral features, and output multi-scale fused features.

[0074] Multi-scale fusion features are feature representations obtained by combining feature concatenation and weighted summation to integrate the advantages of features at different scales. They cover the four scale features shown in the figure: 32-channel input features with a spatial resolution of H / 4×W / 4, 64-channel F2 features with a spatial resolution of H / 8×W / 8, 128-channel F3 features with a spatial resolution of H / 16×W / 16, and 256-channel F4 features with a spatial resolution of H / 32×W / 32. In the integration process, this feature retains the local details of learner operations in the virtual teaching scenario carried by the first two relatively high-resolution features, and also incorporates the overall behavioral semantic information enhanced by the latter two low-resolution features, ultimately forming a feature result that takes into account both detailed expression and semantic depth.

[0075] It should be noted that feature encoding is performed on behavioral images in virtual teaching scenarios. Basic convolution and normalization operations transform the raw visual data into structured behavioral features, providing a unified and effective input for subsequent branch extraction. High-resolution and low-resolution feature extraction are then performed on these behavioral features. The high-resolution branch maintains the feature space resolution through lightweight convolution, outputting high-resolution behavioral features containing local details such as hand operations and tool interactions. The low-resolution branch enhances semantic abstraction through stride convolution, outputting low-resolution behavioral features reflecting the overall operational posture and action intent. Multi-scale feature fusion is then performed on the high-resolution and low-resolution behavioral features, combining feature concatenation and weighted summation to integrate the advantages of both types of features, outputting multi-scale fused features that simultaneously carry local details and global semantics. This step enhances the expressive power of features for complex behaviors in virtual teaching scenarios, providing high-quality feature support for subsequent feature enhancement and posture decoding.

[0076] Step 103: Use a feature enhancement network to enhance the multi-scale fused features and output the target enhanced features.

[0077] The feature enhancement network includes a pyramid attention module, a local-global context fusion module, and a multi-head attention feature fusion module.

[0078] It should be noted that refined behavior recognition in virtual teaching scenarios requires capturing both local details and global semantic information of actions. To address the shortcomings of existing single feature extraction strategies in handling complex scenarios, this invention proposes a multi-level feature perception (MFP) mechanism, i.e., a feature enhancement network, which constructs a complete feature representation from local details to global semantics through cascaded processing of PAM, LGCM, and AFFM. Figure 3 As shown, the multi-scale features generated by the feature extraction network first enter... Figure 3The (A) part is used as input and processed layer by layer by the three modules PAM, LGCM and AFFM to realize the transformation from multi-scale basic features to high semantic fusion features.

[0079] Specifically, step 103 may include the following sub-steps:

[0080] S31. Multi-scale feature integration is performed on the multi-scale fusion features through the pyramid attention module, and the multi-scale integrated features are output.

[0081] It should be noted that the pyramid attention module integrates multi-scale features of the multi-scale fusion feature. First, pooling features of different scales are constructed based on the feature. Then, the corresponding spatial attention map and channel attention map are calculated. After weighting and adjusting the pooling features using these attention maps, the adjusted features are combined with the four scale features contained in the original multi-scale fusion feature to output the multi-scale integrated feature.

[0082] Further, step S31 may include the following sub-steps:

[0083] S311. Perform multi-scale pooling on the multi-scale fusion features and output the multi-scale pooling features corresponding to the multi-scale fusion features.

[0084] S312. Calculate spatial attention and channel attention for the multi-scale pooling features respectively, and output the spatial attention and channel attention corresponding to the multi-scale pooling features.

[0085] S313. Adaptively weighted fuse the spatial attention and channel attention corresponding to the multi-scale pooling features to output the fused attention corresponding to the multi-scale pooling features;

[0086] S314. Integrate the multi-scale pooling features and the corresponding fusion attention to output multi-scale integrated features.

[0087] It should be noted that, as Figure 3 As shown in section (B), the PAM module (Pyramid Attention Module) receives multi-scale fused features. As input, by constructing a feature pyramid and combining spatial and channel dual attention mechanisms, adaptive perception and weighted fusion of spatial information at different granularities are achieved. Given input features For simplicity, it will be referred to below as The module first constructs a multi-scale feature pyramid through pooling operations at different scales:

[0088] ;

[0089] Where C, H, and W represent the number of channels, height, and width of the feature map, respectively; This represents pooling operations at different scales, generating feature maps with different spatial resolutions. For each pooling scale *l*, the feature map (i.e., the multi-scale pooling feature map) is... Spatial attention and channel attention are calculated separately. Spatial attention is implemented in the following way:

[0090] ;

[0091] in, Multi-scale pooling features Corresponding spatial attention; Conv k×k This represents a convolution operation with a kernel size of k×k, BN represents batch normalization, and GELU is the activation function (Gaussian Error Linear Units). The function is sigmoid. Channel attention employs a squeeze-and-excitation mechanism:

[0092] ;

[0093] in, Multi-scale pooling features The corresponding channel attention; GAP represents global average pooling, and MLP is a two-layer perceptron. The outputs of the two attention methods are fused through adaptive weighting.

[0094] ;

[0095] in, Multi-scale pooling features Corresponding fusion attention; These are learnable weight parameters. The final feature enhancement is achieved through attention weighting and upsampling, ensuring the effective integration of multi-scale information.

[0096] ;

[0097] in, It is a multi-scale integration feature; This represents element-wise multiplication. Features at different scales are upsampled to the original resolution. The pyramid attention module effectively captures spatial information from fine to coarse granular through a multi-scale pyramid structure, while the dual attention mechanism enables joint modeling of spatial location and channel importance. This completes the attention-weighted fusion of multi-scale features in the first layer, laying the foundation for the recognition of refined operational behaviors.

[0098] S32. Use the local-global context fusion module to perform local-global context fusion on the multi-scale integrated features to generate multi-scale local-global fused features;

[0099] It should be noted that the local-global context fusion module is used to perform local-global context fusion on the multi-scale integrated features. First, multi-scale local detail features are extracted from the feature through convolution with different kernel sizes. Then, global semantic context features are modeled on it. Subsequently, the local detail features and global semantic features are weighted and fused through a gating mechanism to generate multi-scale local-global fused features.

[0100] Furthermore, step S32 may include the following sub-steps:

[0101] S321. Extract local fine features from the multi-scale integrated features and output the multi-scale local fine features.

[0102] S322. Dynamically weight the multi-scale local fine features to generate multi-scale local weighted features;

[0103] S323. Extract global context features from the multi-scale integrated features to generate multi-scale global context features;

[0104] S324. Based on the gating mechanism, multi-scale global context features and multi-scale local weighted features are fused to generate multi-scale local-global fused features.

[0105] It should be noted that, as Figure 3 As shown in section (C), the local-global context fusion module achieves deep fusion of local detailed features and global semantic information through a dual-branch parallel architecture. This module receives the enhanced features output by PAM. That is, multi-scale integration features As input (denoted as x for simplicity), the local context branch uses multi-scale depthwise separable convolution to extract local fine features in parallel, resulting in multi-scale local fine features. :

[0106] ;

[0107] in, Indicates the kernel size. For depthwise separable convolutions, PWConv is a pointwise convolution. Local feature fusion employs attention-guided dynamic weighting.

[0108] ;

[0109] in, Multi-scale local weighted features; dynamic weighting coefficients We obtain the following by normalization using the softmax function:

[0110] ;

[0111] Wherein, GAP stands for Global Average Pooling; FC stands for Fully Connected Layer. This is the index of the convolution kernel size, and its value is the same as k.

[0112] Meanwhile, this invention designs a global context modeling branch:

[0113] ;

[0114] in, Multi-scale global context features; and These represent the LayerNorm and GELU activation functions, respectively. and These are the parameters for the fully connected layer. The fusion of local features and global context is achieved through a gating mechanism:

[0115] ;

[0116] in, It is a multi-scale local-global fusion feature; Indicates feature splicing, This module employs a dual-branch parallel processing architecture. On one hand, it utilizes multi-scale deep convolution to extract local fine features, and on the other hand, it establishes contextual semantics through global pooling. The two sets of features complement each other through a gating fusion mechanism, completing the deep fusion of local and global information in the second layer. This ensures that the model can capture local details such as hand operations while also understanding the global semantics of the overall action.

[0117] S33. Input the multi-scale local-global fusion features into the multi-head attention feature fusion module for deep fusion and output the enhanced fusion features;

[0118] It should be noted that the multi-scale local-global fusion features are input into the multi-head attention feature fusion module for deep fusion. First, the channels of each scale contained in the feature are compressed to a unified dimension through 1×1 convolution. Then, cross-scale feature associations are established through the multi-head self-attention mechanism to achieve interaction. Finally, the features after interaction at each scale are weighted and summed through learnable weights to output the enhanced fusion features.

[0119] Furthermore, step S33 may include the following sub-steps:

[0120] S331. Compress the multi-scale local-global fusion features and output the multi-scale compressed features;

[0121] S332. Perform multi-head self-attention calculation on the multi-scale compressed features and output the multi-head attention weights corresponding to the multi-scale compressed features.

[0122] S333. Concatenate the multi-head attention weights corresponding to the multi-scale compressed features to obtain multi-scale attention features;

[0123] S334. Calculate the target weights corresponding to the multi-scale attention features;

[0124] S335. Perform a weighted summation of the multi-scale attention features and their corresponding target weights to output the enhanced fusion features.

[0125] It should be noted that, as Figure 3 As shown in section (D), the adaptive feature fusion module is responsible for integrating the multi-level features processed by the first two layers into a unified representation. This module receives deep fusion features from four scales as input and achieves effective utilization of cross-scale information through three steps: channel compression, multi-head attention, and weighted fusion. Since features at different scales may have different numbers of channels after processing by PAM and LGCM, the module first compresses all scale features to a unified channel dimension using a 1x1 convolution, i.e., compressing the multi-scale local-global fusion features, and outputting multi-scale compressed features:

[0126] ;

[0127] Where i is the scale index (values ​​1, 2, 3, and 4, corresponding to four scale features); For the multi-scale compression feature of the i-th scale; The multi-scale local-global fusion features at the i-th scale include 32-channel local-global fusion features with spatial resolution H / 4×W / 4, 64-channel local-global fusion features with spatial resolution H / 8×W / 8, 128-channel local-global fusion features with spatial resolution H / 16×W / 16, and 256-channel local-global fusion features with spatial resolution H / 32×W / 32.

[0128] The compressed features interact across scales via a multi-head self-attention mechanism. For each head h, the query, key, and value are computed:

[0129] ;

[0130] in, Let i be the query feature (Query) corresponding to the i-th scale and the h-th head. This is the query weight matrix corresponding to the h-th head; The key feature (Key) is the key corresponding to the i-th scale and the h-th head. This is the key weight matrix corresponding to the h-th head; Value is the value feature corresponding to the i-th scale and the h-th head. Let h be the value weight matrix corresponding to the h-th head; h is the head index in multi-head attention.

[0131] Attention weights and output are calculated as follows:

[0132] ;

[0133] in, Let be the multi-head attention weights corresponding to the i-th scale and the h-th head; for Activation function; This refers to the dimension parameter for attention. Attention results from different heads are merged through a concatenation operation.

[0134] ;

[0135] in, Here, H represents the multi-scale attention feature corresponding to the i-th scale, and H is the number of heads. This is the output projection matrix. Finally, the enhanced fusion features are obtained through weighted summation. :

[0136] ;

[0137] Among them, the target weight Generate based on feature importance using the SE module (Squeeze-and-Excitation Networks):

[0138] ;

[0139] in, for Activation function; i and j are scale indices, with values ​​of 1, 2, 3, and 4, corresponding to features at four different scales, used to traverse all scales to calculate weights; To sum the exponential results across four scales, this module establishes cross-scale feature relationships through a multi-head attention mechanism. Different heads focus on different feature subspaces, allowing the fusion process to fully utilize the complementary information of features at each scale. Learnable weight parameters are used to integrate features, ensuring that features at different scales are allocated their contributions reasonably according to their importance. This three-layer progressive feature fusion mechanism enables the model to effectively capture refined action features and global contextual information in teaching scenarios, providing a reliable feature foundation for subsequent pose estimation.

[0140] S34. Adaptive activation adjustment mechanism is used to adaptively activate and adjust the enhanced fusion features, and the adjusted enhanced fusion features are output.

[0141] It should be noted that the complexity and dynamism of virtual teaching environments pose a significant challenge to the stability of behavior recognition systems. Due to differences in learners' operating habits, changes in perspective, and environmental interference, existing fixed-parameter recognition methods often struggle to maintain stable performance. Feature enhancement networks output fused features... Adaptive processing is needed to cope with these dynamic changes. Therefore, this invention proposes an adaptive pose estimation strategy based on dynamic feature enhancement. By introducing adaptive adjustment mechanisms at three levels—activation function, spatial sampling, and feature distribution—the model achieves rapid response to environmental changes. This strategy comprises the cascaded processing of three core components: adaptive activation adjustment (A), adaptive spatial sampling (B), and adaptive feature calibration (C). These three components work together to form a complete adaptive technical route from activation function to spatial transformation and feature normalization, as shown below. Figure 4 As shown.

[0142] Furthermore, such as Figure 4 As shown in section (A), the adaptive activation adjustment mechanism receives fused features. As input, the activation function shape is dynamically adjusted according to its distribution characteristics. This mechanism generates calibration features through the collaborative processing of parameter generation and adaptive function branches, which are then applied through activation, channel modulation, and residual connections in the output integration stage, ultimately outputting enhanced features. Given input features... (For simplicity, it will be denoted as x in the following text), the adaptive activation function is defined as:

[0143] ;

[0144] in, For the adjusted enhanced fusion features; dynamic parameters , and The parameter generation branch adaptively calculates based on global statistics of the input features:

[0145] ;

[0146] GAP stands for Global Average Pooling. For GELU activation function, For the sigmoid function, ensure that the parameter values ​​are within a reasonable range. and Learnable parameters are generated for the parameter generation network. This adaptive activation mechanism can dynamically adjust the shape of the input features according to their distribution characteristics, providing more flexible nonlinear mapping capabilities, and is particularly suitable for handling differences in the behavioral patterns of different learners in teaching scenarios. Compared to fixed ReLU or Leaky ReLU, this adaptive mechanism can learn optimal scaling coefficients for positive, negative, and linear parts respectively, thereby better preserving and enhancing key feature information. The activated and adjusted features (i.e., the adjusted and enhanced fusion features) are denoted as... = This serves as the input for the next stage of deformable convolutional enhancement. This mechanism achieves adaptation at the activation function level, enabling the nonlinear transformation to automatically adjust according to the input content, providing optimized feature representation for subsequent spatial adaptive processing.

[0147] S35. The adjusted enhanced fusion features are deformed by convolution through an adaptive spatial sampling mechanism to generate deformed convolution enhanced features.

[0148] It should be noted that, as Figure 4 As shown in section (B), the deformable convolution enhancement mechanism receives activated-adjusted features. As input, the convolutional kernel achieves geometrically adaptive modeling of pose changes by learning spatially adaptive sampling positions. This mechanism generates spatial transformation information through transformation matrix and offset field branches, and then generates deformable features through deformable convolution. Given input features... (For simplicity, hereinafter referred to as x), the mechanism first generates a spatially adaptive position offset field:

[0149] ;

[0150] Among them, the offset field N is the number of sampling points (N=9 corresponds to a 3×3 convolution kernel). The offset at each position... The sampling position is dynamically adjusted based on the input content, thanks to adaptive learning via the network. To achieve more accurate geometric adaptive transformation, a two-dimensional affine transformation matrix generation mechanism was designed:

[0151] ;

[0152] The output of adaptive spatial sampling is calculated through deformable convolution:

[0153] ;

[0154] in, This is the sampling offset for standard convolution. This indicates the sampling position after adaptive adjustment and deformation. These are the convolution weights. Because... The coordinates are usually non-integer coordinates, and bilinear interpolation is used to calculate the eigenvalues.

[0155] ;

[0156] in, Represents the spatial location of all integers. Let be the bilinear interpolation kernel function. The feature enhanced by deformable convolution (i.e., the deformable convolution enhanced feature) is denoted as . = (The enhanced features are obtained by performing deformable convolution on the activated and adjusted features, serving as input for the next stage of dynamic feature calibration.) This adaptive spatial sampling mechanism allows the convolution operation to dynamically adjust the shape and position of the receptive field according to the input content, thereby better capturing key feature regions in the taught actions. Compared to fixed regular grid sampling, adaptive spatial sampling can automatically focus on important regions around key points, effectively handling geometric deformation and occlusion issues in pose. This mechanism achieves adaptation at the spatial sampling level, enabling the convolution kernel to flexibly adjust the sampling position according to the target pose features, significantly enhancing the model's geometric adaptability to complex pose changes.

[0157] S36. Adaptive feature calibration mechanism is used to perform adaptive feature calibration on deformable convolution enhancement features to generate target enhancement features.

[0158] It should be noted that, as Figure 4 As shown in section (C), the dynamic feature calibration mechanism receives features enhanced by deformable convolution (i.e., deformable convolution enhanced features). As input, the feature distribution is dynamically adjusted according to task conditions. This mechanism generates the final calibrated features through parallel processing of feature normalization and feature refinement branches, via feature integration. Given input features... (For simplicity, it will be denoted as x in the following text), first calculate the characteristic statistics:

[0159] ;

[0160] Then, adaptive feature calibration is performed using conditional normalization:

[0161] ;

[0162] in and The scaling parameter is determined by task condition c. and offset parameters It is generated dynamically through learning networks.

[0163] ;

[0164] in, The input feature x is the feature value at spatial location (h, w), where h is the height index and w is the width index; The global mean of the input feature x; The global variance of the input feature x; is an extremely small constant; c represents conditional information, serving as the basis for dynamic calibration. This information typically originates from task-related guiding signals or feature statistics, used to generate scaling and offset parameters adapted to the current scene; MLP stands for Multilayer Perceptron; the task condition c encodes environmental factors such as scene type, lighting conditions, and viewpoint information, enabling adaptive adjustment of feature calibration for different teaching scenarios. The calibrated features (i.e., target enhancement features) are denoted as . = This serves as the final input to the pose decoding network. This adaptive calibration mechanism can dynamically adjust the feature distribution according to different teaching task scenarios, significantly improving the model's generalization ability and recognition stability in complex teaching environments. Compared with traditional batch normalization, adaptive feature calibration can flexibly adjust the normalization parameters according to specific task features, avoiding information loss caused by uniform processing and better preserving the feature differences in different scenarios.

[0165] Step 104: Input the target enhancement features into the pose decoding network for pose decoding and generate behavior recognition results.

[0166] It should be noted that the target augmentation features are input into the pose decoding network for pose decoding. First, the high-dimensional target augmentation features are mapped to low-dimensional pose feature vectors through a fully connected layer. Then, the temporal modeling module captures the dynamic correlation information between feature vectors. Finally, the Softmax classifier outputs the probability distribution of various behaviors to generate behavior recognition results.

[0167] Furthermore, step 104 may include the following sub-steps:

[0168] S41. Perform pose attention calculation on the target enhancement features to obtain attention enhancement features;

[0169] S42. Based on the attention enhancement features, predict the heatmap and output the key point heatmap;

[0170] S43. Based on the attention enhancement features, predict the displacement field and output the key point displacement field.

[0171] S44. Based on the key point heatmap and key point displacement field, perform behavior pose estimation and generate behavior recognition results.

[0172] Pose attention computation refers to the process of weighted optimization of key regions and channels of target enhancement features through a mechanism that combines spatial attention and channel attention, in order to enhance effective information related to pose recognition.

[0173] Heatmap prediction refers to the processing operation based on attention-enhanced features, which models the spatial distribution of key points through convolutional layers and outputs a two-dimensional probability heatmap for each key point. The peak of the heatmap corresponds to the candidate position of the key point.

[0174] A keypoint heatmap refers to a two-dimensional probability heatmap output by heatmap prediction. Each heatmap corresponds to a preset human keypoint and is used to characterize the distribution of candidate positions of keypoints in the feature map.

[0175] Displacement field prediction refers to the processing operation that uses attention-enhanced features to restore resolution through transposed convolution and predict the offset of candidate keypoint positions, outputting the displacement field of keypoints.

[0176] The key point displacement field refers to the offset information in the displacement field prediction output, which is used to correct the deviation of the peak coordinates of the key point heatmap and improve the accuracy of key point positioning.

[0177] Behavioral pose estimation refers to the entire process of determining precise key point coordinates by combining key point heatmaps and displacement fields, constructing a human pose skeleton, and making behavioral judgments.

[0178] It should be noted that the target enhancement features are subjected to pose attention calculation. Spatial attention mechanism is used to capture the spatial correlation of key pose regions in the feature map, and channel attention mechanism is combined to strengthen the feature channel weights related to pose recognition. The target enhancement features are then weighted and optimized to obtain attention-enhanced features. Heatmap prediction is performed based on the attention-enhanced features. Stacked convolutional layers are used to model the spatial distribution of human key points and output a two-dimensional probability heatmap corresponding to each preset key point. The peak position in the heatmap corresponds to the candidate coordinates of the key point. Displacement field prediction is performed based on the attention-enhanced features. Transposed convolutional layers are used to restore the spatial resolution of the feature map and predict the offset information of each candidate position of the key point, outputting the key point displacement field. Behavioral pose estimation is performed based on the key point heatmap and the key point displacement field. The peak coordinates of the key point heatmap are extracted as the initial position, and the offset of the key point displacement field is superimposed to obtain the accurate human key point coordinates. A human pose skeleton is constructed based on the key point coordinates, and then a classifier is used to perform feature matching and behavior determination on the pose skeleton to generate behavior recognition results. This invention focuses on key posture information through a posture attention mechanism and improves the accuracy of key point localization by combining heatmap and displacement field fusion prediction. It effectively solves the problem that posture features are easily interfered with in complex teaching scenarios and enhances the robustness and reliability of behavior recognition results.

[0179] It is worth mentioning that the experimental evaluation of this invention was conducted on a self-built VTB dataset (Virtual Teaching Behavior Dataset). The test environment was a workstation equipped with an RTX 3090 GPU (NVIDIA RTX 3090 Graphics Processing Unit), implemented using the PyTorch framework. Figure 5 shows the complete hardware and software architecture. The evaluation metrics used were standard metrics in the field of human pose estimation, namely Average Precision (AP) and Average Recall (AR). These two metrics characterize the model's detection capability and keypoint capture capability from different perspectives. The comparative experiments selected representative methods such as RTMO-S (Real-Time Multi-Object-Small), YOLOXPose-S (You Only Look Once eXtreme Pose-Small) / L (YOLOXPose-Large), DEKR (Deep Keypoint Regression), and CID (Contextual Information Distillation). These methods represent lightweight solutions that prioritize real-time performance and accuracy-first solutions based on high-resolution networks, respectively. All methods were tested in the same environment to ensure the fairness of the comparison.

[0180] Furthermore, understanding the underlying mechanisms behind MPABM's performance improvements requires analyzing the contributions of each core component. This invention constructs a simplified model without MFP and APE as a starting point. This configuration achieves 92.1% AP and 93.9% AR on the VTB dataset, with an inference time of 336.18 milliseconds. After adding the MFP mechanism, the model exhibits an interesting dual improvement: AP and AR jump to 94.1% and 95.5% respectively, and more surprisingly, the inference time drops dramatically to 158.04 milliseconds, representing an efficiency improvement of more than half. This seemingly contradictory phenomenon actually reveals the ingenuity of the MFP design. The pyramid attention module captures rich spatial information through multi-scale perception, the local-global context module bridges the gap between details and semantics, and the adaptive feature fusion module avoids redundant computation through intelligent compression. When testing the APE strategy separately, another pattern is observed: AP and AR improve to 94.9% and 95.7% respectively, while the inference time increases to 319.91 milliseconds. Compared to MFP, which focuses on improving feature quality, APE pays more attention to enhancing adaptability. By introducing dynamic adjustments at three levels—activation function, spatial sampling, and feature calibration—the model can cope with complex factors such as differences in learners' operating habits, changes in perspective, and environmental interference in virtual teaching scenarios.

[0181] The complementarity between the two modules is far more valuable than when used alone. The multi-level feature representation constructed by MFP provides a rich foundation for APE's adjustments. Imagine if the input features themselves lack discriminative power; even the most flexible adaptive mechanism would struggle to function. Conversely, APE's dynamic adjustment capabilities compensate for MFP's shortcomings in the face of environmental changes. Even with highly refined feature extraction, a fixed processing flow is bound to experience performance fluctuations when illumination changes or pose deformations occur. This synergy is particularly evident in real-world scenarios: when handling fine hand movements, MFP accurately captures subtle finger movements, while APE dynamically adjusts its receptive field based on the specific pose; when dealing with perspective shifts, MFP's multi-scale representation provides sufficient information redundancy, while APE locks in key regions through spatial adaptive sampling.

[0182] The complete model, integrating MFP and APE, achieved 95.1% AP and 96.1% AR on the VTB dataset, with an inference time of 307.65 milliseconds. Compared to the baseline configuration, this represents an improvement of 3.0% AP and 2.2% AR, and an efficiency improvement of 8.5%. This result confirms the design philosophy of this invention: refined behavior recognition in virtual teaching scenarios requires both deep feature representation capabilities to capture operational details and flexible environmental adaptability to cope with complex changes; neither is dispensable. Notably, the performance gain of the complete model (3.0%) exceeds the simple sum of the contributions from MFP (2.0%) and APE (2.8%), and this additional gain comes from the enhanced interaction between the two. High-quality features make dynamic adjustments more accurate, while adaptive processing further unleashes the potential of feature representation.

[0183] Furthermore, for comparative experimental analysis, examining MPABM's performance within a broader technological context provides a clearer understanding of its technological positioning. Lightweight models like RTMO-S (Real-Time Multi-Object Pose Estimation) and YOLOXPose (You Only Look Once eXtreme Pose Estimation) are favored in industrial applications. They sacrifice inference speed for simplified network structure, a reasonable trade-off on resource-constrained edge devices. RTMO-S achieved 73.4% AP on the VTB dataset, while YOLOXPose-S and YOLOXPose-L reached 82.7% and 81.5% respectively, but still lagged behind MPABM's 95.1% by 12.4% to 21.7%. This gap has a substantial impact in virtual teaching scenarios. When learners perform operations requiring millimeter-level precision, such as instrument calibration, lightweight models often fail to accurately locate finger joints, leading to biased operation evaluations. The limitations of network capacity make it difficult for these models to simultaneously model global action intentions and local operational details, which is precisely the core requirement of virtual teaching assessment.

[0184] A more compelling comparison comes from methods based on the same HRNet-W32 (High-Resolution Network with Width 32) backbone. DEKR employs a keypoint grouping and decoupling strategy, demonstrating robust performance on multi-person pose estimation tasks, achieving 92.1% AP and 93.9% AR on the VTB dataset, which can be considered the baseline level for this architecture. CID builds upon this by introducing context instance decoupling, improving AR to 94.3%, but without improving AP. MPABM achieves 95.1% AP and 96.1% AR under the same computational resources, representing improvements of 3.0% and 2.2% compared to DEKR, and 3.0% and 1.8% compared to CID. This performance leap under the same hardware conditions directly reflects the value of algorithmic innovation, achieving a qualitative breakthrough not through larger models or more computation, but through a more reasonable feature processing flow and a more flexible adaptive mechanism. In actual testing, this invention observed that MPABM has a particularly prominent advantage in handling partially occluded scenes. When the learner's hand is occluded by the object being manipulated, traditional methods often result in keypoint drift, while MPABM, with the help of PAM's multi-scale perception and spatial adaptive sampling of deformable convolution, can infer the complete pose structure from a limited visible area.

[0185] Reviewing these comparative experiments, MPABM's technological advantages stem from its systematic innovative design rather than a single breakthrough. The multi-level feature fusion structure comprised of PAM, LGCM, and AFFM is not simply about deepening the network; rather, it establishes a progressive extraction path from multi-scale basic features to highly semantically fused features, with each layer undertaking a clearly defined functional role. Adaptive activation adjustment enables nonlinear transformations to automatically adjust according to the input content, deformable convolution enhancement allows the receptive field of the convolution kernel to flexibly deform with the target pose, and dynamic feature calibration ensures the rationality of feature distribution under different scenarios. This design philosophy, which deeply couples deep feature engineering with flexible adaptive mechanisms, allows MPABM to not only surpass existing methods in numerical metrics but also demonstrate a profound understanding of the complexity of virtual teaching scenarios in practical applications, providing a practical and feasible technical path for the refined development of educational information technology in the field of virtual teaching.

[0186] As a comparison of technical effects, existing technologies can be referenced. Regarding feature representation, Li et al.'s multi-level feature learning framework generates multiple feasible pose hypotheses through three-stage spatiotemporal representation learning, effectively handling deep ambiguity and self-occlusion problems. However, it suffers from high computational complexity, and the modeling of dependencies between multi-level features still needs improvement. Fan et al. extended this approach to the spatiotemporal domain, designing a multi-scale, multi-level spatiotemporal feature framework that achieves comprehensive modeling of dynamic human motion. However, it suffers from insufficient local detail feature extraction capabilities, especially in teaching scenarios requiring simultaneous understanding of global action intent and local operational details; this shallow fusion strategy struggles to meet the demands of refined recognition. In terms of environmental adaptation, Raychaudhuri et al. proposed the passive domain adaptive framework POST (Prior-guided Self-training), achieving privacy-preserving cross-domain transfer through dual-space adaptation and pose prior regularization. However, it faces challenges in real-time performance and lacks dynamic parameter adjustment capabilities. Zhang et al.'s PeVL framework (Pose-Enhanced Vision-Language Model) introduces multimodal fusion, enabling fine-grained recognition at both the action and sub-action semantic levels. However, the generalization ability and real-time performance of this method in specific virtual teaching scenarios still need further verification.

[0187] This invention addresses the technical challenges of refined behavior measurement in virtual teaching scenarios by proposing a refined behavior recognition method for such scenarios. Through the synergistic design of a multi-level feature fusion mechanism and an adaptive pose estimation strategy, it systematically solves the dual bottlenecks of insufficient feature representation depth and limited environmental adaptability faced by existing methods in virtual teaching environments. Experimental results show that MPABM achieves breakthrough performance on the self-built VTB dataset, with an average precision (AP@0.5) of 95.1% and an average recall (AR) of 96.1%, significantly surpassing existing mainstream methods. Ablation experiments confirm that the multi-level feature perception mechanism (MFP) and the adaptive pose estimation strategy (APE) are key to the performance improvement, contributing 2.0% and 2.8% to the AP improvement, respectively. Their synergistic effect brings a final gain of 3.0%, while inference efficiency is improved by 8.5%. This fully verifies the effectiveness and advancement of the technical approach combining "high-quality feature representation" and "flexible environmental adaptation" in solving the problem of refined behavior recognition in virtual teaching.

[0188] Specifically, at the feature representation level, a complete link from multi-scale feature depth perception to cross-scale semantic intelligence integration is constructed through the cascaded design of the Pyramid Attention Module (PAM), Local-Global Context Module (LGCM), and Adaptive Feature Fusion Module (AFFM), significantly enhancing the ability to capture refined teaching actions. At the environmental adaptation level, by constructing an adaptive strategy integrating adaptive activation adjustment, deformable convolution enhancement, and dynamic feature calibration, comprehensive dynamic adjustments from neuron activation and spatial sampling to feature distribution are achieved, giving the model inherent resilience to cope with the dynamics of the virtual teaching environment and individual differences among learners. These two technological innovations complement each other, forming a complete technical framework from feature extraction to pose decoding, providing a new paradigm for high-precision, high-efficiency, and highly robust behavior measurement in virtual teaching scenarios.

[0189] In this embodiment of the invention, a method for refined behavior recognition in virtual teaching scenarios is provided. The method acquires behavior images of the virtual teaching scenario and inputs these images into a multi-layered perception adaptive behavior measurement model. This model includes a feature extraction network, a feature enhancement network, and a pose decoding network. The feature extraction network extracts multi-scale features from the virtual teaching scenario behavior images, outputting multi-scale fused features. The feature enhancement network then enhances these multi-scale fused features, outputting target-enhanced features. These target-enhanced features are input into the pose decoding network for pose decoding, generating behavior recognition results. Based on this approach, the invention, through multi-scale feature extraction by the feature extraction network, preserves both local details of learner operations in the virtual teaching scenario and captures global action semantics. Furthermore, the feature enhancement network further enhances the multi-scale fused features, achieving deep integration of multi-level features and significantly improving the ability to capture refined teaching actions, thereby enhancing the accuracy of refined behavior recognition results in virtual teaching scenarios.

[0190] Please see Figure 6 , Figure 6 This is a structural block diagram of a refined behavior recognition system for virtual teaching scenarios provided in Embodiment 2 of the present invention.

[0191] This invention provides a refined behavior recognition system for virtual teaching scenarios, comprising:

[0192] The acquisition module 601 is used to acquire behavior images of virtual teaching scenarios and input the behavior images of virtual teaching scenarios into a behavior measurement model based on multi-layer perception and adaptive behavior. The behavior measurement model based on multi-layer perception and adaptive behavior includes a feature extraction network, a feature enhancement network, and a pose decoding network.

[0193] Extraction module 602 is used to extract multi-scale features from virtual teaching scene behavior images through a feature extraction network and output multi-scale fused features;

[0194] Enhancement module 603 is used to enhance the multi-scale fused features using a feature enhancement network and output the target enhanced features;

[0195] The decoding module 604 is used to input the target enhancement features into the pose decoding network for pose decoding and generate behavior recognition results.

[0196] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system and modules described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0197] This invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor performs the steps of the refined behavior recognition method for virtual teaching scenarios as described in the above embodiments.

[0198] This invention also provides a computer-readable storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements the steps of the refined behavior recognition method for virtual teaching scenarios as described in the above embodiments.

[0199] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0200] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0201] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0202] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0203] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for fine-grained behavior recognition for virtual teaching scenarios, characterized in that, include: Acquire behavioral images of a virtual teaching scenario and input the behavioral images of the virtual teaching scenario into a behavior measurement model based on multi-layer perception and adaptive behavior, wherein the behavior measurement model based on multi-layer perception and adaptive behavior includes a feature extraction network, a feature enhancement network, and a pose decoding network; The feature extraction network is used to extract multi-scale features from the behavioral images of the virtual teaching scene and output multi-scale fused features. The feature enhancement network is used to enhance the multi-scale fused features and output the target enhanced features. The target enhancement features are input into the pose decoding network for pose decoding to generate behavior recognition results; The feature enhancement network includes a pyramid attention module, a local-global context fusion module, and a multi-head attention feature fusion module; the step of using the feature enhancement network to enhance the multi-scale fused features and outputting the target enhanced features includes: The multi-scale fusion features are integrated using the pyramid attention module to output multi-scale integrated features. A local-global context fusion module is used to perform local-global context fusion on the multi-scale integrated features to generate multi-scale local-global fused features. The multi-scale local-global fusion features are input into the multi-head attention feature fusion module for deep fusion, and the enhanced fusion features are output. An adaptive activation adjustment mechanism is used to adaptively adjust the enhanced fusion features, and the adjusted enhanced fusion features are output. The adjusted enhanced fusion features are deformed by performing deformed convolution on the adaptive spatial sampling mechanism to generate deformed convolution enhanced features. An adaptive feature calibration mechanism is used to adaptively calibrate the deformable convolutional enhancement features to generate the target enhancement features; The step of using a local-global context fusion module to perform local-global context fusion on the multi-scale integrated features to generate multi-scale local-global fused features includes: Local fine-grained features are extracted from the multi-scale integrated features to output multi-scale local fine-grained features; The multi-scale local fine features are dynamically weighted to generate multi-scale local weighted features; Global context features are extracted from the multi-scale integrated features to generate multi-scale global context features; The multi-scale global context features and the multi-scale local weighted features are fused based on a gating mechanism to generate multi-scale local-global fused features; The step of inputting the multi-scale local-global fusion features into a multi-head attention feature fusion module for deep fusion and outputting enhanced fusion features includes: The multi-scale local-global fusion features are compressed to output multi-scale compressed features; Perform multi-head self-attention calculation on the multi-scale compressed features and output the multi-head attention weights corresponding to the multi-scale compressed features; The multi-head attention weights corresponding to the multi-scale compressed features are concatenated to obtain multi-scale attention features. Calculate the target weights corresponding to the multi-scale attention features; The multi-scale attention features and their corresponding target weights are weighted and summed to output the enhanced fusion features.

2. The method of claim 1, wherein the method is a method of fine-grained behavior recognition for a virtual teaching scenario. The step of extracting multi-scale features from the virtual teaching scene behavior images through the feature extraction network and outputting multi-scale fused features includes: The behavioral images of the virtual teaching scene are feature-encoded to output basic behavioral features; High-resolution feature extraction and low-resolution feature extraction are performed on the basic behavioral features respectively, and high-resolution behavioral features and low-resolution behavioral features are output. Multi-scale feature fusion is performed on the high-resolution features and low-resolution features of the behavior to output multi-scale fused features.

3. The method of claim 1, wherein the method is a method of fine-grained behavior recognition for a virtual teaching scenario. The process of integrating the multi-scale fused features using a pyramid attention module to output multi-scale integrated features includes: Perform multi-scale pooling on the multi-scale fusion feature and output the multi-scale pooled feature corresponding to the multi-scale fusion feature. Spatial attention and channel attention are calculated for the multi-scale pooling features respectively, and the spatial attention and channel attention corresponding to the multi-scale pooling features are output. Adaptively weighted and fused the spatial attention and channel attention corresponding to the multi-scale pooling features to output the fused attention corresponding to the multi-scale pooling features; The multi-scale pooling features and the corresponding fusion attention are integrated to output multi-scale integrated features.

4. The method of claim 1, wherein the method is a method of fine-grained behavior recognition for a virtual teaching scenario. The step of inputting the target enhancement features into the pose decoding network for pose decoding and generating behavior recognition results includes: Pose attention calculation is performed on the target enhancement features to obtain attention enhancement features; Based on the attention enhancement features, perform heatmap prediction and output a keypoint heatmap; Based on the attention enhancement features, displacement field prediction is performed, and the key point displacement field is output. Behavior pose estimation is performed based on the key point heatmap and the key point displacement field to generate behavior recognition results.

5. A system for fine-grained behavior recognition in a virtual teaching scenario, applied to the method for fine-grained behavior recognition in a virtual teaching scenario according to claim 1, characterized in that, include: The acquisition module is used to acquire behavioral images of a virtual teaching scene and input the behavioral images of the virtual teaching scene into a behavior measurement model based on multi-layer perception and adaptive behavior, wherein the behavior measurement model based on multi-layer perception and adaptive behavior includes a feature extraction network, a feature enhancement network and a pose decoding network. The extraction module is used to extract multi-scale features from the virtual teaching scene behavior images through the feature extraction network and output multi-scale fused features; The enhancement module is used to enhance the multi-scale fused features using the feature enhancement network and output the target enhanced features; The decoding module is used to input the target enhancement features into the pose decoding network for pose decoding and generate behavior recognition results.

6. An electronic device, comprising: The method includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the refined behavior recognition method for virtual teaching scenarios as described in any one of claims 1-4.

7. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed, it implements the refined behavior recognition method for virtual teaching scenarios as described in any one of claims 1-4.