Interactive teaching evaluation auxiliary system and device based on multi-modal fusion
By constructing a user training motion profile database and a real-time evaluation and feedback mechanism, the problems of one-sided motion evaluation and lack of precision in guidance in traditional physical education teaching have been solved, enabling accurate recommendation of personalized programs and continuous optimization of movements.
Patent Information
- Application Number
- CN202511249695.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Traditional physical education classroom teaching models struggle to quantify, record, and systematically analyze students' movement standardization and training intensity. Personalized teaching goals are difficult to implement, and the lack of a multimodal data fusion assessment system and personalized program matching mechanism leads to a one-sided portrayal of athletic ability and a lack of precise guidance.
The interactive teaching assessment assistance system based on multimodal fusion constructs a user training action profile library through the profile module, generates the optimal guidance plan in combination with the strategy recommendation module, evaluates the actions in real time and provides feedback on the spatiotemporal change chain of non-standard actions through the auxiliary adjustment module, and makes real-time adjustments through the source tracing feedback module, forming a closed-loop optimization.
It enables accurate assessment of multimodal data and personalized solution matching, improves the scientific nature and interactivity of assessment assistance, shortens the movement correction cycle, and enhances the scientific nature and effectiveness of training.
Smart Images

Figure CN120807242B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of assessment assistance, and particularly relates to an interactive teaching assessment assistance system and device based on multimodal fusion. Background Technology
[0002] Against the backdrop of the deepening advancement of educational informatization, traditional physical education classroom teaching models have significant limitations: teachers rely on on-site observation and experience-based judgment, making it difficult to quantify, record, and systematically analyze students' movement standardization and training intensity. Furthermore, individual differences make it difficult to implement personalized teaching goals, such as the difficulty in capturing subtle deviations in joint angles during track and field throwing and the difficulty in adjusting the load based on real-time physiological movements during physical training. At the same time, existing models also suffer from problems such as the lack of a multimodal data fusion assessment system, the absence of a personalized program matching mechanism, and insufficient movement assessment and traceability capabilities. This results in a one-sided portrayal of students' athletic abilities, difficulty in generating suitable guidance programs, and a lack of precise basis for teachers' corrections, thus hindering the improvement of physical education teaching quality. To address this, this invention provides an interactive teaching assessment auxiliary system and device based on multimodal fusion. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention proposes an interactive teaching assessment assistance system and device based on multimodal fusion. This system collects multi-dimensional information through a profiling module, and constructs a user training action profiling library by combining multimodal feature extraction and knowledge graph algorithms. A strategy recommendation module generates the optimal guidance scheme based on the profiling library, real-time information, and a teaching guidance strategy library, using a correction scheme matching model. An auxiliary adjustment module collects training action images in real time, and uses a spatiotemporal feature extraction model and assessment algorithm to obtain a continuous corrective action sequence and standard score after guidance. A source feedback module, combining an action standard knowledge base, obtains the horizontal and vertical spatiotemporal change chain of non-standard actions through a forward reasoning model, providing feedback to the instructor for real-time adjustment until the action meets the standard. This application achieves accurate assessment and personalized scheme matching through multimodal data fusion, effectively improving the scientific rigor and interactivity of the assessment assistance.
[0004] To achieve the above objectives, the present invention provides the following technical solution:
[0005] The interactive teaching assessment support system based on multimodal fusion includes: a profiling module, a strategy recommendation module, an auxiliary adjustment module, and a source feedback module.
[0006] The profiling module is used to collect basic user information, historical training achievement information, and corresponding guidance plan information and standard action information knowledge base of the configured student terminal, and combine them with knowledge graph algorithm to obtain user training action profiling library;
[0007] The strategy recommendation module is used to obtain the optimal matching guidance plan for the corresponding user by combining the user training action profile library with the real-time collected user basic information and the configured teaching guidance strategy library and through the correction plan matching model.
[0008] The auxiliary adjustment module is used to provide auxiliary guidance according to the matching guidance plan, and collect the training action information of the student in real time. Combined with the preset spatiotemporal feature extraction model and comprehensive evaluation algorithm, it obtains the continuous corrective action sequence after guidance and the corresponding correction standard evaluation score.
[0009] The source tracing and feedback module is used to obtain the spatiotemporal change chain of non-standard movements in the corresponding teaching scenario based on the continuous corrective movement sequence after guidance and the corresponding correction standard evaluation score, combined with the standard movement information knowledge base, through a forward reasoning model. The spatiotemporal change chain of non-standard movements is then fed back to the alarm device configured on the teacher's end for real-time early warning. The module also collects feedback from the teacher's end to adjust the matching guidance plan in real time until the corresponding user's training movements meet the conditions. At the same time, the guidance plan for the user's training movements that meet the conditions is stored in the standard movement information knowledge base for management.
[0010] Specifically, the spatiotemporal variation chain of non-standard movements includes the spatiotemporal variation chain of lateral non-standard movements constructed by different training movements under the same guidance scheme in the current training cycle, and the spatiotemporal variation chain of longitudinal non-standard movements constructed by the same training movement under the same guidance scheme in different training cycles.
[0011] The profiling module includes a data acquisition unit, a first data analysis unit, and a second data analysis unit;
[0012] The data acquisition unit is used to collect basic user information, historical motion assessment text and image information, corresponding guidance scheme information and standard motion information knowledge base for configuring students, and to classify, arrange and standardize the collected information according to user and collection timestamp.
[0013] The first data analysis unit is used to obtain user action evaluation information triplet and corresponding action spatiotemporal feature set by combining user historical motion evaluation action text information and image information with entity relationship extraction algorithm and image recognition algorithm.
[0014] The second data analysis unit is used to obtain the motion correlation at continuous time points and the probability that the next motion will be non-standard under the condition that the previous motion is non-standard, based on the standard motion information knowledge base and the correlation analysis algorithm with built-in expert experience.
[0015] Specifically, the portrait module also includes portrait units;
[0016] The profiling unit is used to construct a user training action profiling library based on user basic information, user action evaluation information triplet, corresponding action spatiotemporal feature set, standard action information knowledge base, action correlation at continuous time points, and the probability of subsequent actions being non-standard under the condition that the previous action is non-standard, through graph algorithm. The user training action profiling library includes a baseline layer, a training layer, and an inference layer.
[0017] The baseline layer stores basic user information; the training layer stores the spatiotemporal features of training actions in each historical guidance scheme, along with corresponding training evaluation scores and action guidance correction information; the inference layer stores a standard action information knowledge base and, by combining the action correlations at consecutive time points and the probability of subsequent actions being non-standard under the condition that the previous action was non-standard, constructs causal inference connections between consecutive actions to obtain the causal inference chain of training actions. This chain is then mapped to the training layer under the corresponding user to evaluate and infer correct the user's action information.
[0018] Specifically, the auxiliary adjustment module includes a spatiotemporal feature extraction unit and an evaluation unit;
[0019] The spatiotemporal feature extraction unit is used to obtain the spatiotemporal correlation feature space of the user's optimal matching guidance scheme by combining the training action information corresponding to the optimal matching guidance scheme with a preset spatiotemporal graph convolution model.
[0020] The evaluation unit is used to obtain the training evaluation score of each action in the optimal matching guidance scheme by combining the spatiotemporal correlation feature space of the user's optimal matching guidance scheme with the causal inference chain of training actions through a comprehensive evaluation algorithm.
[0021] Specifically, the source tracing feedback module includes a reverse reasoning unit and a feedback adjustment unit;
[0022] The reverse reasoning unit is used to perform forward causal reasoning of non-standard actions based on the training evaluation score of each action in the optimal matching guidance scheme, the action spatiotemporal change chain of historical longitudinal non-standard actions, and the causal reasoning connection in the causal reasoning chain of training actions. Based on the training evaluation score and the color mapping space of HSV, it obtains the action spatiotemporal change chain of real-time lateral non-standard actions and the action spatiotemporal change chain of real-time longitudinal non-standard actions.
[0023] The feedback adjustment unit is used to feed back the spatiotemporal change chains of real-time lateral and longitudinal non-standard movements, along with their corresponding training evaluation scores, to the profiling unit for storage. Simultaneously, it judges the optimal matching guidance scheme's comprehensive training evaluation score against a preset evaluation threshold. If the comprehensive training evaluation score is greater than the preset threshold but any movement does not meet the corresponding standard movement information, the corresponding movement information and the non-standard movement tracing information obtained through the spatiotemporal change chain of the lateral non-standard movements are fed back to the corresponding instructor for guidance and correction. If the comprehensive training evaluation score is less than or equal to the preset evaluation threshold, the corresponding guidance scheme information, training evaluation score, and tracing results are fed back to the corresponding instructor for real-time adjustment of the optimal matching guidance scheme.
[0024] Specifically, the process of constructing the correction scheme matching model includes:
[0025] Based on the spatiotemporal feature information of training actions and the corresponding training evaluation scores in each historical guidance scheme in the user training action profile library, the spatiotemporal features of the same training action and the trend of training evaluation score changes in different training cycles are extracted through the non-standard action layer in the correction scheme matching model to obtain the user's longitudinal non-standard action feature change space and training evaluation score fluctuation coefficient.
[0026] Based on the user's longitudinal non-standard action feature change space and training evaluation score fluctuation coefficient, combined with the action spatiotemporal change chain of the lateral non-standard action in the user training action profile library, the degree of correlation coverage adjustment of the current non-standard action to the related actions in the same training cycle is obtained.
[0027] Simultaneously, based on the fluctuation coefficient of the training evaluation score of the associated actions under the corresponding degree of associated coverage adjustment in different training cycles after adjusting the current non-standard actions, the longitudinal effect transfer rate of the same associated action in different training cycles is obtained. Specifically, the construction process of the correction scheme matching model also includes:
[0028] Based on the training actions within each training cycle, a horizontal action node sequence is constructed, and horizontal association connections are constructed by adjusting the degree of association coverage of related actions in the same training cycle after the current non-standard action is adjusted.
[0029] Based on the fluctuation coefficient of the training evaluation score of the same training action of a user in different training cycles and the longitudinal effect transfer rate of the same related action in different training cycles, a vertical correlation transfer connection is constructed.
[0030] Based on the constructed horizontal action nodes of different training cycles, combined with horizontal and vertical correlation transfer connections and topological space layers, the user's historical training action effect correlation transfer space is obtained.
[0031] Based on the user's historical training action effect association transfer space and the teaching guidance strategy library, a matching search is performed by configuring a matching layer with a maximum adjustment evaluation effect function to generate a matching guidance scheme sequence and a corresponding matching loss function. The maximum adjustment evaluation effect function is a function constructed by fitting the corresponding non-standard action and related action correction standard evaluation score after the matching scheme is trained, the probability of the subsequent action being non-standard under the condition that the previous action is non-standard, and the effect transfer rate with a linear fitting function and taking the maximum value.
[0032] The model is trained based on the matching guidance scheme sequence and the corresponding matching loss function, combined with the preset loss function threshold and the set model training period. The trained correction scheme matching model is obtained, and the optimal matching guidance scheme corresponding to the maximum adjustment evaluation effect is output.
[0033] Specifically, the construction process of the spatiotemporal feature extraction model includes:
[0034] The system acquires training action information and corresponding standard action image information for consecutive time periods within different training cycles of the user and performs enhancement and denoising processing. Based on the processed images and the expert labeling algorithm pre-trained with skeletal node information, it performs key skeletal node labeling and center of gravity node labeling under different actions. It obtains the first continuous action labeling image sequence of the same training cycle and the second labeling image sequence of the same training action corresponding to different training cycles.
[0035] The first continuous action labeled image sequence of the same training period and the second labeled image sequence of the same training action corresponding to different training periods are simultaneously input into the OpenPose action estimation layer to obtain the continuous action node feature space of the same training period and the node bias feature space of the same training action corresponding to different training periods constructed with the standard continuous action node feature space.
[0036] Based on the feature space of continuous action nodes in the same training cycle, combined with the body structure, the connection edges between each node are constructed to obtain the undirected node graph corresponding to each image.
[0037] Based on the priority of the limb action implementation of the corresponding node in each image at consecutive time points in the undirected node graph corresponding to each image, a continuous action connection is constructed.
[0038] Based on the undirected node graph and continuous action connections corresponding to each image, a continuous action topology space is constructed through a topology space layer.
[0039] Specifically, the construction process of the spatiotemporal feature extraction model also includes:
[0040] Based on the continuous action topology space combined with the node deviation feature space corresponding to the same training action in different training cycles, a continuous action topology space containing the trend of action deviation change of the same action in different training cycles is obtained through a time trend extraction layer.
[0041] Simultaneously, the continuous action topology space is input into the limb curvature recognition layer to divide the action into stages, and the stage feature sequence of each action in the continuous action is obtained; the stage feature sequence is constructed from the node spatial coordinates corresponding to the start time, the node spatial coordinates corresponding to the turning time, and the node spatial coordinates corresponding to the end time of the corresponding action.
[0042] Based on the root node and centroid node of each action in the continuous action topology space, combined with a preset node partitioning strategy, the region node feature set of the undirected node graph corresponding to each image is obtained.
[0043] The continuous action topology space containing the trend of action deviation changes of the same action in different training cycles, the stage feature sequence of each action in the continuous action, and the region node feature set are input into the feature fusion layer constructed by the convolutional attention network to obtain the continuous action fusion topology space.
[0044] Based on the continuous action fusion topology space combined with the standard continuous action node feature space, the output layer obtains the continuous action node coordinate sequence, node coordinate deviation loss, and the corresponding action deviation change trend and trend confidence. At the same time, the continuous action fusion topology space combined with the standard continuous action node feature space is input into the association causal layer to obtain the forward cumulative contribution path and path confidence corresponding to non-standard actions.
[0045] The loss function, which is constructed based on node coordinate deviation loss, trend confidence and path confidence, is used to train the spatiotemporal feature extraction model by combining the preset loss function threshold and training period.
[0046] The interactive teaching assessment aid device based on multimodal fusion includes: a data acquisition component and a processor configured within the device; the data acquisition component integrates a profiling module; the processor integrates a strategy recommendation module, an auxiliary adjustment module, and a source feedback module.
[0047] The data acquisition component is used to collect basic user information, historical training achievement information, corresponding guidance plan information, and standard action information knowledge base of the configuration student terminal, and combine them with the built-in multimodal feature extraction algorithm and knowledge graph algorithm to obtain a user training action profile library.
[0048] The processor, based on a user training action profile library, real-time collected user basic information, and a configured teaching guidance strategy library, generates an optimal matching guidance scheme through a correction scheme matching model. While the user trains according to this scheme, it collects training action images in real time, and combines a spatiotemporal feature extraction model and an action evaluation algorithm to obtain a continuous corrective action sequence and corresponding correction standard evaluation scores after guidance. Based on the continuous corrective action sequence and corresponding correction standard evaluation scores, and a standard action information knowledge base, it obtains the horizontal and vertical spatiotemporal change chains of non-standard actions through a forward inference model, and feeds these chains back to the configured teacher. Simultaneously, it collects feedback from the teacher and adjusts the matching guidance scheme in real time until the corresponding user training action meets the conditions. The guidance scheme for user training actions that meet the conditions is then stored in the standard action information knowledge base for management.
[0049] Compared with the prior art, the beneficial effects of the present invention are:
[0050] This invention addresses the shortcomings of existing technologies by constructing a user training action profile library through multimodal feature extraction and graph algorithms in the profile module, achieving deep integration of user information and action features. It enhances the personalization and targeting of guidance plans by leveraging a correction scheme matching model combined with cross-sectional correlation analysis. Furthermore, it utilizes a spatiotemporal feature extraction model to accurately extract and evaluate action features, improving the accuracy of action evaluation. Finally, it obtains and dynamically adjusts the spatiotemporal change chain through forward inference in the source feedback module, ensuring continuous achievement of action standards. Ultimately, this invention achieves a closed loop from personalized scheme recommendation to accurate evaluation and dynamic optimization, effectively improving the scientific rigor and effectiveness of training. Attached Figure Description
[0051] Figure 1 This is a block diagram of the interactive teaching assessment support system based on multimodal fusion according to Embodiment 1 of the present invention;
[0052] Figure 2 This is a diagram of the spatiotemporal feature extraction model architecture of Embodiment 1 of the present invention. Detailed Implementation
[0053] Example 1
[0054] Please see Figure 1 The present invention provides an embodiment of an interactive teaching assessment assistance system based on multimodal fusion, comprising: a profiling module, a strategy recommendation module, an assistance adjustment module, and a source feedback module;
[0055] The profiling module is used to collect basic user information, historical training achievement information, and corresponding guidance plan information and standard action information knowledge base of the configured student terminal, and combine them with knowledge graph algorithm to obtain user training action profiling library;
[0056] The strategy recommendation module is used to obtain the optimal matching guidance plan for the corresponding user by combining the user training action profile library with the real-time collected user basic information and the configured teaching guidance strategy library and through the correction plan matching model.
[0057] The auxiliary adjustment module is used to provide auxiliary guidance according to the matching guidance plan, and collect the training action information of the student in real time. Combined with the preset spatiotemporal feature extraction model and comprehensive evaluation algorithm, it obtains the continuous corrective action sequence after guidance and the corresponding correction standard evaluation score.
[0058] The source tracing and feedback module is used to obtain the spatiotemporal change chain of non-standard movements in the corresponding teaching scenario based on the continuous corrective action sequence after guidance, the corresponding correction standard evaluation score, and the standard movement information knowledge base, through a forward inference model. This spatiotemporal change chain of non-standard movements is then fed back to the alarm device configured on the teacher's end for real-time early warning. The module also collects feedback from the teacher's end to adjust the matching guidance plan in real time until the corresponding user's training movement meets the conditions. Simultaneously, the guidance plan for the user's training movement that meets the conditions is stored in the standard movement information knowledge base for management. It should be further noted that the spatiotemporal change chain of non-standard movements in this embodiment includes the spatiotemporal change chain of horizontal non-standard movements constructed from different training movements under the same guidance plan in the current training cycle, and the spatiotemporal change chain of vertical non-standard movements constructed from the same training movement under the same guidance plan in different training cycles. The forward inference model in this embodiment is preferably a deep inference algorithm, whereby the user derives the continuous non-standard movement path that leads to the current non-standard movement. It should also be noted that the student's end and the teacher's end in this embodiment are implemented by monitoring equipment configured by those skilled in the art.
[0059] It should be further noted that the portrait module in this embodiment includes a data acquisition unit, a first data analysis unit, a second data analysis unit, and a portrait unit;
[0060] The data acquisition unit is used to collect basic user information, historical motion assessment text and image information, corresponding guidance scheme information and standard motion information knowledge base for configuring students, and to classify, arrange and standardize the collected information according to user and collection timestamp.
[0061] It should be further explained that the classification, arrangement, and standardization preprocessing process in this embodiment is as follows:
[0062] Based on the collected user basic information, historical motion evaluation text information, guidance plan information, and standard motion information knowledge base, the data is grouped and clustered by unique user identifiers. Within each user group, the data is then sorted in ascending order by collection timestamp to achieve classification. Simultaneously, data cleaning filters out samples with more than 30% missing fields, removes outliers from numerical data that deviate from the mean by more than three standard deviations, standardizes the format to a unified date format of "YYYY-MM-DD HH:MM:SS", converts text information to UTF-8 encoding, and maps numerical data such as training evaluation scores to a standardized range of 0-100. Text preprocessing uses Jieba segmentation to segment the evaluation text, removes meaningless words from a stop word list, and maps technical terms to unified terms in the motion standard knowledge base. This results in a standardized information dataset that is user-grouped, chronologically arranged, uniformly formatted, and clean.
[0063] The first data analysis unit is used to obtain user action evaluation information triplet and corresponding action spatiotemporal feature set by combining user historical motion evaluation action text information and image information with entity relationship extraction algorithm and image recognition algorithm.
[0064] It should be further explained that, in this embodiment, the entity relationship extraction algorithm and image recognition algorithm are implemented in one specific way as follows: Based on the user's historical motion evaluation action text information, word segmentation, part-of-speech tagging, and named entity recognition are performed first. Then, a pre-trained language model is used to extract deep semantic features of the text. An entity relationship extraction framework is constructed by combining bidirectional LSTM and CRF models to identify the relationships between entities and obtain user action evaluation information triples. Based on the user's training action information, image denoising and scale normalization preprocessing are performed first. Then, a target detection algorithm is used to locate the human body region. The coordinates of 18 key skeletal nodes are extracted by combining the action estimation model. The motion trajectory of the skeletal nodes in continuous frame images is spatiotemporally feature-encoded using a 3D convolutional neural network to obtain the corresponding action spatiotemporal feature set. It should be noted that the named entity recognition in this embodiment includes, but is not limited to, recognizing entities such as action names, evaluation indicators, and score values.
[0065] The second data analysis unit is used to obtain the motion correlation at continuous time points and the probability that the next motion will be non-standard under the condition that the previous motion is non-standard, based on the standard motion information knowledge base and the correlation analysis algorithm with built-in expert experience.
[0066] It should be further explained that, in this embodiment, the association analysis algorithm incorporating expert experience is implemented in one specific way as follows:
[0067] Based on continuous movement sequence data in the standard movement information knowledge base and the built-in expert experience rule base, the expert experience is first transformed into algorithmic constraints. For example, the influence weight range of lower limb movements on core movements is specified. Then, an improved association rule algorithm is used to mine frequent patterns in historical movement evaluation data, extract the associated itemsets between continuous movements, and combine the Bayesian probability formula to calculate the probability that the subsequent movement will be non-standard when the previous movement is non-standard, thus obtaining the movement association relationship and corresponding conditional probability at continuous time points. Its function is to provide causal reasoning basis for movement evaluation, support the source analysis of subsequent non-standard movements, and the targeted adjustment of guidance schemes. The expert experience rule base includes, but is not limited to, rules such as movement dependency, joint linkage priority, and connection timing threshold.
[0068] It should be further explained that the implementation process of the improved association rule algorithm in this embodiment includes:
[0069] Based on expert experience, different weights for action association pairs are preset. For example, the basic weights for the association between lower limb actions and core actions, and the weighting coefficients for joint-linked actions are specified. The corrected support is obtained by multiplying the original itemset support (number of itemset occurrences / total number of transactions) by the corresponding expert weight. The corrected confidence is obtained by multiplying the original confidence (probability of subsequent actions occurring when the preceding action occurs) by the product of the expert weight and the importance coefficient of the itemset in the training process (e.g., the coefficient of basic actions is higher than that of auxiliary actions). Minimum corrected support and confidence thresholds are set, and historical continuous action evaluation data is scanned transaction by transaction. Itemets that simultaneously meet both thresholds are selected as frequent itemsets, and the corrected strong association rules are obtained, thus realizing the accurate quantitative mining of the association relationship between continuous actions.
[0070] The profiling unit is used to construct a user training action profiling library based on user basic information, user action evaluation information triplet, corresponding action spatiotemporal feature set, standard action information knowledge base, action correlation at continuous time points, and the probability of subsequent actions being non-standard under the condition that the previous action is non-standard, through graph algorithm. The user training action profiling library includes a baseline layer, a training layer, and an inference layer.
[0071] The baseline layer stores basic user information; the training layer stores the spatiotemporal features of training actions in each historical guidance scheme, along with corresponding training evaluation scores and action guidance correction information; the inference layer stores a standard action information knowledge base and, by combining the action correlations at consecutive time points and the probability of subsequent actions being non-standard under the condition that the previous action was non-standard, constructs causal inference connections between consecutive actions to obtain the causal inference chain of training actions. This chain is then mapped to the training layer under the corresponding user to evaluate and infer correct the user's action information.
[0072] It should be further explained that the detailed process of constructing the user training action profile library in this embodiment includes:
[0073] Based on user basic information converted into a 12-dimensional feature vector, historical training evaluation information triplet encoding into entity relation vectors, action spatiotemporal feature space extraction into a 300-dimensional skeletal trajectory sequence, action standard knowledge base rules converted into 200-dimensional constraint features, and continuous action causal association and conditional probability constructed into a 100-dimensional association matrix, three types of nodes are defined: user nodes integrate basic information and historical evaluation vectors, action nodes fuse spatiotemporal features and training evaluation scores, and standard nodes encapsulate knowledge base rules and association probabilities. Edge connections are constructed by calculating the training frequency weights (training times / total times) of users and actions, the causal probability weights between actions, and the deviation weights between actions and standards (Euclidean distance between actual features and standard features). A graph convolutional network is used to aggregate neighbor features layer by layer. The first layer aggregates the interaction features between users and actions, the second layer incorporates the causal relationships between actions, and the third layer combines standard rule constraints. Attention mechanisms are used to assign higher weights to high-frequency training actions and strong causal relationships. After five rounds of iteration to update node embeddings, user node features are finally mapped to the baseline layer. Action node sequences and evaluation data are integrated into the training layer, and standard node rules and inference logic are constructed into the inference layer, resulting in a user training action profile library with a three-layer structure. The 12-dimensional feature vector includes, but is not limited to, structured data such as user age, gender, and body movements. The causal probability weights between actions are the probability values that a non-standard action in the previous action will cause a non-standard action in the next action.
[0074] It should be further noted that the auxiliary adjustment module in this embodiment includes a spatiotemporal feature extraction unit and an evaluation unit;
[0075] The spatiotemporal feature extraction unit is used to obtain the spatiotemporal correlation feature space of the user's optimal matching guidance scheme by combining the training action information corresponding to the optimal matching guidance scheme with a preset spatiotemporal graph convolution model.
[0076] The evaluation unit is used to obtain the training evaluation score of each action in the optimal matching guidance scheme by combining the spatiotemporal correlation feature space of the user's optimal matching guidance scheme with the causal inference chain of training actions through a comprehensive evaluation algorithm.
[0077] It should be further noted that the source tracing feedback module in this embodiment includes a reverse reasoning unit and a feedback adjustment unit;
[0078] The reverse reasoning unit is used to perform forward causal reasoning of non-standard actions based on the training evaluation score of each action in the optimal matching guidance scheme, the action spatiotemporal change chain of historical longitudinal non-standard actions, and the causal reasoning connection in the causal reasoning chain of training actions. Based on the training evaluation score and the color mapping space of HSV, it obtains the action spatiotemporal change chain of real-time lateral non-standard actions and the action spatiotemporal change chain of real-time longitudinal non-standard actions.
[0079] For example, the HSV color mapping space in this embodiment is as follows: In this embodiment, 0 points correspond to pure red (H=0), 100 points correspond to pure green (H=120), intermediate scores are gradient-transitioned in 5-point intervals, S and V are fixed at 1, maintaining saturation and brightness. Specific implementations can be customized by those skilled in the art based on usage habits. This involves first processing the lateral chains: sorting non-standard actions according to the execution order within the same training cycle, adding an ID and non-standard feature label to each action node, filling the node with the corresponding HSV color, and adjusting the color transparency of the connecting lines according to the action association probability intensity; for example, higher intensity results in higher transparency. The process begins by first establishing a horizontal temporal chain; then, a vertical chain is processed: the same non-standard movement is sorted by timestamps of different training cycles, a cycle number and score change value are added to each cycle node, and the node is filled with the corresponding HSV color of the score. Adjacent nodes are connected by hue gradient lines according to time progression, forming a vertical trend chain. This results in a real-time spatiotemporal change chain of non-standard movements with color markings and clearly defined node attributes. Its function is to transform abstract score data into an intuitive colorized chain, clearly presenting the correlation of movements within the same cycle and the evolution of movements across different cycles. This helps instructors quickly locate key non-standard movements and their changing trends, improving correction efficiency.
[0080] The feedback adjustment unit is used to feed back the spatiotemporal change chains of real-time lateral and longitudinal non-standard movements, along with their corresponding training evaluation scores, to the profiling unit for storage. Simultaneously, it judges the optimal matching guidance scheme's comprehensive training evaluation score against a preset evaluation threshold. If the comprehensive training evaluation score is greater than the preset threshold but any movement does not meet the corresponding standard movement information, the corresponding movement information and the non-standard movement tracing information obtained through the spatiotemporal change chain of the lateral non-standard movements are fed back to the corresponding instructor for guidance and correction. If the comprehensive training evaluation score is less than or equal to the preset evaluation threshold, the corresponding guidance scheme information, training evaluation score, and tracing results are fed back to the corresponding instructor for real-time adjustment of the optimal matching guidance scheme.
[0081] It should be further explained that the process of obtaining the training evaluation score in this embodiment is as follows: the spatiotemporal feature extraction unit of the auxiliary adjustment module extracts the spatiotemporal correlation feature space of the training action information of the optimal matching guidance scheme using a preset spatiotemporal graph convolution model. The evaluation unit compares the spatiotemporal correlation feature space with the causal inference chain of the training action and calculates the training evaluation score of each action through a comprehensive evaluation algorithm. The comprehensive training evaluation score of the optimal matching guidance scheme is obtained by weighting and summing the training evaluation scores of all actions in the scheme according to the importance weight of the actions (the core actions have a higher weight than the auxiliary actions). The feedback adjustment unit compares the comprehensive training evaluation score with the preset evaluation threshold to realize the adjustment judgment of the guidance scheme or action.
[0082] This process constructs a complete training and evaluation closed loop from information collection to solution optimization through multi-module collaboration, bringing multi-dimensional practical value. The profiling module integrates basic user information, historical evaluations, and a standard knowledge base. Utilizing multimodal feature extraction and knowledge graph algorithms to build a user training action profiling library, it structurally links scattered information, providing comprehensive feature support for subsequent solution matching. This ensures that each recommendation is rooted in the user's unique movement history and physical characteristics. The strategy recommendation module, based on the profiling library and real-time information, filters solutions through a corrective solution matching model, avoiding the blindness of generic solutions and ensuring that recommended guidance accurately matches the user's current actions and improvement needs. The auxiliary adjustment module collects action images in real time and converts them into spatiotemporal features. Combined with a standard inference chain, it generates training evaluation scores, achieving dynamic monitoring of the training process and allowing non-standard actions to be captured promptly. The source feedback module generates an intuitive spatiotemporal change chain of non-standard actions through reverse reasoning. This helps instructors quickly locate the root cause of problems and updates feedback information to the profiling library in real time, driving continuous solution optimization. This end-to-end design, encompassing information integration, precise recommendations, real-time evaluation, and dynamic adjustments, not only enhances the accuracy of motion assessment and the relevance of solution matching, but also continuously aligns with user training goals through closed-loop iteration, effectively shortening the motion correction cycle and strengthening the stability and sustainability of training results.
[0083] It should be further explained that the construction process of the correction scheme matching model in this embodiment includes:
[0084] Based on the spatiotemporal features of training actions and corresponding training evaluation scores in each historical guidance scheme in the user training action profile database, the spatiotemporal features of the same training action and the trend of training evaluation score changes in different training cycles are extracted through the non-standard action layer in the correction scheme matching model, so as to obtain the user's longitudinal non-standard action feature change space and training evaluation score fluctuation coefficient.
[0085] It should be further explained that the implementation process of the training evaluation score fluctuation coefficient in this embodiment is as follows: Based on the training evaluation scores of the same training action in different training cycles, a score sequence is first formed by sorting by training cycle, outliers in the sequence that deviate from the mean of adjacent scores by more than twice are removed, and then the arithmetic mean and sample standard deviation of the remaining scores are calculated. The ratio of the sample standard deviation to the arithmetic mean is used as the training evaluation score fluctuation coefficient to obtain a quantitative value of the degree of score fluctuation of the action in different cycles. Its function is to reflect the stability of the same action in long-term training. The larger the fluctuation coefficient, the more unstable the score change is, providing a quantitative basis for identifying persistent non-standard action characteristics and subsequent matching correction schemes.
[0086] Based on the user's longitudinal non-standard action feature change space and training evaluation score fluctuation coefficient, combined with the action spatiotemporal change chain of the lateral non-standard action in the user training action profile library, the degree of correlation coverage adjustment of the current non-standard action to the related actions in the same training cycle is obtained.
[0087] It should be further explained that the specific implementation process of the associated coverage adjustment degree in this embodiment includes:
[0088] The adjustment scheme for the current non-standard action based on the user's longitudinal non-standard action feature change space includes, but is not limited to, action amplitude correction values, force timing adjustment parameters, and training evaluation score fluctuation coefficients. It combines a list of related actions causally correlated with the current action within the same training cycle in the spatiotemporal change chain of the lateral non-standard action. Optionally, in this embodiment, related actions with a probability ≥ 0.4 are selected through the causal probability matrix of the inference layer. First, the non-standard score vectors of each related action before adjustment are extracted. Then, the adjustment parameters of the current action are input into the physical motion simulation model to calculate the theoretical score vectors of the related actions after adjustment, thus obtaining the score improvement vector for each related action. The improvement vectors of each related action are weighted using the training evaluation score fluctuation coefficient; specifically, related actions with larger fluctuation coefficients are given higher weights in the spatial dimension. The ratio of the sum of the magnitudes of the weighted improvement vectors to the total number of related actions is used as the degree of correlation coverage adjustment after the current non-standard action is adjusted within the same training cycle. Its function is to accurately quantify the multi-dimensional improvement effect of the current action adjustment on related actions in the same cycle, providing a quantitative basis for formulating action combination adjustment strategies.
[0089] Simultaneously, based on the fluctuation coefficient of the training evaluation score of the associated movements under the degree of related coverage adjustment in different training cycles after the current non-standard movements are adjusted, the longitudinal effect transfer rate of the same associated movement in different training cycles can be obtained.
[0090] It should be further explained that the implementation process of the vertical effect mobility in this embodiment includes:
[0091] Based on the evaluation results of the effect of the current non-standard movement adjustment on the related movements under different training cycles (i.e., the degree of improvement of the current movement on subsequent causally related movements, which is measured by the fluctuation coefficient of the training evaluation scores of the causally related movements before and after improvement), the evaluation data of the same related movement in three consecutive training cycles are selected. In this embodiment, the ratio of the spatial deviation improvement score of the movement in the (n+1)th cycle to the spatial deviation improvement score in the nth cycle, the ratio of the temporal deviation improvement score, and the ratio of the strength deviation improvement score are extracted. Then, the three ratios are weighted according to the cycle interval by the time decay coefficient. The sum of the weighted ratios is divided by 3 to obtain the longitudinal effect transfer rate, which is the quantitative value of the effect continuity of the same related movement in different cycles. Its function is to quantify the ability of the adjustment effect to be transferred in cross-cycle training. The higher the transfer rate, the more stable and lasting the effect is, providing a basis for designing a coherent cross-cycle correction scheme.
[0092] Based on the training actions within each training cycle, a horizontal action node sequence is constructed, and horizontal association connections are constructed by adjusting the degree of association coverage of related actions in the same training cycle after the current non-standard action is adjusted.
[0093] Based on the fluctuation coefficient of the training evaluation score of the same training action of a user in different training cycles and the longitudinal effect transfer rate of the same related action in different training cycles, a vertical correlation transfer connection is constructed.
[0094] Based on the constructed horizontal action nodes of different training cycles, combined with horizontal and vertical correlation transfer connections and topological space layers, the user's historical training action effect correlation transfer space is obtained.
[0095] It should be further explained that the process of constructing the user's historical training action effect association transfer space in this embodiment includes:
[0096] Based on lateral action nodes (including but not limited to those containing action IDs, training evaluation scores, and spatiotemporal feature vectors), lateral associative connections (directed edges with weights for associative coverage adjustment), and vertical associative transfer connections (cross-cycle edges with weights for performance transfer rates) across different training cycles, each training cycle is first treated as an independent two-dimensional plane. Lateral action nodes are arranged in an ordered manner according to the training execution sequence within the plane. Lateral associative connections connect nodes within the same plane using arrows with weighted values; specifically, the weights are directly labeled on the lines. Then, between adjacent cycle planes, vertical associative transfer connections connect the same action nodes perpendicularly, with performance transfer rates labeled on the lines, forming a multi-layered three-dimensional topological skeleton. Next, the local topological strength of each node is calculated, and the local topological... The strength is specifically defined as the ratio of the total weight of the edges associated with a node to the number of connected nodes. Nodes with the top 20% local topological strength are selected as key nodes. A manifold learning algorithm is used to map the multi-layer topological skeleton into three-dimensional space, maintaining the distance ratio between key nodes while scaling down non-key nodes proportionally. Finally, Delaunay triangulation is used to mesh the nodes in the three-dimensional space, with the mesh cell density positively correlated with the connection weight, thus obtaining the user's historical training action effect association transfer space. Its function is to transform abstract action association relationships into a quantifiable three-dimensional topological structure, intuitively presenting the horizontal action association strength and vertical effect transfer rules, providing spatial structural association basis for subsequent matching guidance schemes, and improving the accuracy of scheme matching.
[0097] Based on the user's historical training action effect association transfer space and the teaching guidance strategy library, a matching search is performed through a matching layer configured with the maximum adjustment evaluation effect function to generate a matching guidance scheme sequence and the corresponding matching loss function. It should be further explained that the maximum adjustment evaluation effect function in this embodiment is a function constructed by fitting the correction standard evaluation score of the corresponding non-standard action and related action after the matching scheme is trained, the probability of the subsequent action being non-standard under the condition that the previous action is non-standard, and the effect transfer rate with a linear fitting function and taking the maximum value.
[0098] It should be further explained that the specific implementation process of the matching layer in this embodiment for matching search includes:
[0099] Based on the evaluation scores of the correction criteria for non-standard actions and related actions, the probability of a subsequent action becoming non-standard given a previous non-standard action, and the effect transfer rate, the importance levels of these three factors are preset through expert experience. A linear combination function containing these three variables is constructed. This function is fitted using historical matching data (by adjusting the variable coefficients to minimize the error between the function value and the actual improvement). The form corresponding to the maximum value of the fitted function is taken as the maximum adjustment evaluation effect function. Then, based on the topological nodes of the user's historical training action effect correlation transfer space and the scheme features in the teaching guidance strategy library, the matching layer first selects schemes from the strategy library that match the action type of the key nodes in the topological space as the initial candidate set, calculates the function value of each candidate scheme, and retains the schemes with higher function values. The retained guidance schemes are then optimized by action sequence, the function value is recalculated, and the optimization is repeated multiple times to generate a sequence of matching guidance schemes. The deviation between the actual function value and the fitted function value of the optimized scheme is used as the matching loss function. Its purpose is to quantify the improvement potential of the scheme through the constructed evaluation function, and to select the scheme sequence that maximizes the improvement effect of non-standard actions by combining multiple rounds of optimization search, taking into account the influence of related actions and cross-cycle effect transfer, providing high-quality scheme samples for model training, and improving the effectiveness of the final recommended scheme. The features of the teaching guidance strategy library include, but are not limited to, action type and training parameters;
[0100] The model is trained based on the matching guidance scheme sequence and the corresponding matching loss function, combined with the preset loss function threshold and the set model training period. The trained correction scheme matching model is obtained, and the optimal matching guidance scheme corresponding to the maximum adjustment evaluation effect is output.
[0101] It should be further explained that the correction scheme matching model in this embodiment is constructed through a multi-stage progressive approach, forming a scheme generation mechanism that accurately adapts to users' training needs, bringing significant practical value. The non-standard action layer extracts longitudinal non-standard action features and training evaluation score fluctuation coefficients to accurately capture the stability differences of the same action in different cycles, providing a quantitative basis for identifying persistent action defects and avoiding misjudgment of random fluctuations. The calculation of the degree of correlation coverage adjustment quantifies the improvement range of the current action adjustment on related actions in the same cycle, ensuring that single action correction can drive the overall training effect improvement and avoiding the disconnect between local adjustments and related actions. The longitudinal effect transfer rate focuses on the continuity of cross-cycle effects. By measuring the ability of the same related action to transfer improvement in different cycles, it provides support for designing coherent long-term guidance schemes and prevents short-term effects from being unsustainable. The user's historical training action effect correlation transfer space transforms the abstract action correlation relationship into a perceptible three-dimensional topological structure, intuitively presenting the horizontal correlation strength and vertical transfer law, giving scheme matching a concrete basis at the spatial structure level. The matching layer constructs a maximum adjustment evaluation function and combines it with multi-round optimization search to select the sequence of solutions that maximizes improvement from the teaching guidance strategy library. It considers the impact of related actions and cross-cycle transfer, ensuring that the solutions are both relevant to current needs and have long-term effectiveness. The entire construction process forms a complete closed loop from problem identification and impact quantification to solution optimization. Through the synergistic effect of each stage, the final optimal matching guidance solution accurately targets non-standard movements, balancing short-term improvement and long-term stability, effectively enhancing the efficiency of movement correction and the sustainability of training results.
[0102] Further explanation is needed; please refer to [link / reference]. Figure 2 In this embodiment, the construction process of the spatiotemporal feature extraction model includes:
[0103] The system acquires training action information and corresponding standard action image information for consecutive time periods within different training cycles of the user and performs enhancement and denoising processing. Based on the processed images and the expert labeling algorithm pre-trained with skeletal node information, it performs key skeletal node labeling and center of gravity node labeling under different actions. It obtains the first continuous action labeling image sequence of the same training cycle and the second labeling image sequence of the same training action corresponding to different training cycles.
[0104] The first continuous action labeled image sequence of the same training period and the second labeled image sequence of the same training action corresponding to different training periods are simultaneously input into the OpenPose action estimation layer to obtain the continuous action node feature space of the same training period and the node bias feature space of the same training action corresponding to different training periods constructed with the standard continuous action node feature space.
[0105] It should be further explained that the implementation process of the OpenPose motion estimation layer in this embodiment includes:
[0106] Based on the processed first continuous action-labeled image sequence of the same training period and the second labeled image sequence of the same training action corresponding to different training periods, the OpenPose action estimation algorithm first uses a feature extraction network to perform layer-by-layer feature encoding on the images to obtain multi-scale feature maps containing human contours and joint candidate regions. Then, a keypoint detection algorithm is used to predict the heatmap (reflecting the probability of node existence) and limb association vector field (reflecting the connection strength between nodes) of each skeletal node on the feature map. Based on the heatmap, candidate nodes with high confidence are selected, and the candidate nodes are grouped into limbs by combining the association vector field to form a complete skeletal node set for a single frame image, including the coordinates and connection relationships of each node. The skeletal node sets of consecutive frames are associated in chronological order through inter-frame nodes. The matching algorithm combines node displacement continuity and limb length constraints to construct a sequence of continuous action nodes for the same training cycle. It extracts the spatiotemporal variation features of nodes in the sequence, including but not limited to displacement trajectory and velocity changes, to obtain a feature space of continuous action nodes for the same training cycle. Simultaneously, it compares the node features of the same training action corresponding to different training cycles with the feature space of standard continuous action nodes frame by frame, calculates the coordinate deviation and connection relationship deviation of each node, and integrates them to form a node deviation feature space of the same training action corresponding to different training cycles. Its function is to accurately extract the spatiotemporal features of the skeletal nodes of the training action and the deviation from the standard action, providing fine-grained node-level feature support for subsequent construction of continuous action topology space and analysis of action deviation trends, thereby improving the accuracy of action recognition and evaluation.
[0107] Based on the feature space of continuous action nodes in the same training cycle, combined with the body structure, the connection edges between each node are constructed to obtain the undirected node graph corresponding to each image.
[0108] Based on the priority of the limb action implementation of the corresponding node in each image at consecutive time points in the undirected node graph corresponding to each image, a continuous action connection is constructed. For example, if action A is a hand raising action and action B is an arm bending action, action A is implemented before action B. Based on the implementation time priority of action A and action B with respect to the elbow joint node, a continuous action connection between action A and action B with respect to the same elbow joint node is constructed.
[0109] Based on the undirected node graph and continuous action connections corresponding to each image, a continuous action topology space is constructed through a topology space layer.
[0110] It should be further explained that the specific implementation process of the topology space layer in this embodiment includes:
[0111] Based on an undirected node graph (including coordinates of each skeleton node and connecting edges constructed according to body structure) and continuous action connections (including limb affiliation and action implementation priority) within the same training cycle, a topology algorithm in the topology space layer first converts the node coordinates of the single-frame undirected node graph into 3D vectors, and the connecting edges into adjacency matrices representing node associations. The adjacency matrices at consecutive time points are arranged chronologically, and the displacement deviation of nodes between adjacent frames and the intensity of change in connecting edges are calculated. Connecting edges with intensity of change below a set standard are selected as stable topological skeletons. Nodes are weighted according to limb priority (e.g., core limbs are prioritized over limbs), with the weight of core nodes dynamically adjusted according to the amplitude of the action. The weighted node vectors and the stable topological skeleton are then mapped to 3D space. This method employs a node distance-based clustering approach to group nodes with similar action features into a single class, retaining the most closely related nodes in each class as topological vertices. It calculates the path lengths between topological vertices (accumulated based on edge weights) and uses the shortest path length as the topological edge to construct a topological network for a single frame's action. The topological networks of each frame are then concatenated along the time axis, maintaining temporal continuity through inter-frame topological vertex matching (based on coordinate similarity), forming a dynamically changing continuous action topological space. Its function is to transform discrete action nodes and connections into a structured spatiotemporal topological structure, preserving key action features and their temporal correlation, providing a unified spatial framework for subsequent analysis of action stage division and deviation trends, and improving the spatiotemporal coherence of action evaluation and the accuracy of feature extraction.
[0112] Based on the continuous action topology space combined with the node deviation feature space corresponding to the same training action in different training cycles, a continuous action topology space containing the trend of action deviation change of the same action in different training cycles is obtained through a time trend extraction layer.
[0113] It should be further explained that the implementation method of the time trend extraction layer in this embodiment is as follows:
[0114] Based on the node coordinates and connection strength of each frame in the continuous action topology space, and the node deviation features corresponding to the same action in different training cycles, including but not limited to node coordinate deviation, connection angle deviation, and temporal deviation, a temporal convolutional network algorithm configured with a time trend extraction layer is used. First, the topology space features and deviation features are combined into a three-dimensional input sequence according to the training cycle time order. Specifically, each time step contains 18 skeletal node coordinates, 12 limb connection strength values, and 6 deviation indices. The sequence is then standardized. A temporal convolutional network with 5 layers of causal convolution is constructed. The first two layers use small dilation rates (1 and 2, respectively) to extract short-term deviation fluctuation features of adjacent training frames (such as instantaneous changes in node deviation within a single cycle). The middle two layers use medium dilation rates (4 and 8) to capture medium-term trend features within the same training cycle (such as cumulative deviation changes over 5 consecutive frames). The last layer uses a large dilation rate (1...). 6) By associating features from different training cycles and integrating the outputs of each layer through residual connections, the network captures long-term deviation trends across cycles (such as deviation improvement or deterioration trajectories over three consecutive cycles). Using the actual historical deviation change curves as supervision labels, the network calculates the mean absolute error between the predicted trend value and the actual label. Convolutional kernel parameters are updated through backpropagation. After training, the continuous action topology space sequence is input into the network, and the deviation trend features of each node in different cycles are output, including but not limited to trend slope, fluctuation amplitude, and stable interval. These trend features are embedded into the node attributes of the original topology space to obtain a continuous action topology space containing the trend of action deviation changes in the same action across different training cycles. Its function is to accurately capture the short-term fluctuations and long-term evolution patterns of action deviation in the time dimension, providing time-series-level feature support for identifying persistent action defects and evaluating the stability of training effects, thereby improving the accuracy of action deviation trend prediction.
[0115] Simultaneously, the continuous action topology space is input into the limb curvature recognition layer to divide the action into stages, and the stage feature sequence of each action in the continuous action is obtained; the stage feature sequence is constructed from the node spatial coordinates corresponding to the start time, the node spatial coordinates corresponding to the turning time, and the node spatial coordinates corresponding to the end time of the corresponding action.
[0116] It should be further explained that the specific implementation process of the limb curvature recognition layer in this embodiment includes:
[0117] Based on the coordinate sequence of skeletal nodes of each limb in the continuous motion topology space, for example, the shoulder-elbow-wrist nodes of the upper limb and the hip-knee-ankle nodes of the lower limb, the node sequences of each limb are first smoothed. Optionally, a sliding window averaging method is used to filter high-frequency noise, with the window size set to 5 frames. Then, the vectors between adjacent nodes are calculated. For example, the shoulder-elbow vector is obtained by subtracting the shoulder node from the elbow node, and the elbow-wrist vector is obtained by subtracting the elbow node from the wrist node. The cosine value of the angle between adjacent vectors is calculated using the vector dot product formula and converted into an angle value as the limb joint angle. Cubic spline curve fitting is performed on the joint angle sequence of continuous frames, and the curvature value of the fitted curve is calculated. Specifically, curvature = angle change rate / limb length, where the limb length is the sum of the Euclidean distances of the limb node sequence, obtaining the real-time curvature change curve of each limb in continuous motion. A curvature threshold is set, specifically, the starting point threshold is when the curvature starts from 0 and continuously increases to 5%. The turning point threshold is the moment when the absolute value of curvature reaches its peak, and the ending point threshold is the moment when the curvature falls back to within 10% of its initial value. By scanning the curvature change curve, the frame that first exceeds the starting point threshold is marked as the starting moment, the frame corresponding to the peak curvature is marked as the turning point moment, and the frame that first falls below the ending point threshold is marked as the ending moment. The spatial coordinates of the nodes corresponding to these three moments are extracted and combined in chronological order to form a stage feature sequence. At the same time, the action stages are divided according to the time interval of the key frames. For example, the first stage is from the starting moment to the turning point, and the second stage is from the turning point to the ending moment. If there are multiple turning points, the stages are subdivided according to the number of peaks. Its function is to accurately locate the key stage nodes of the action by quantifying the curvature change of limb deformation, providing fine-grained temporal feature support for subsequent analysis of the standard of action in each stage and identification of stage defects, thereby improving the objectivity and accuracy of action stage division.
[0118] Based on the root node and centroid node of each action in the continuous action topology space, combined with a preset node partitioning strategy, the region node feature set of the undirected node graph corresponding to each image is obtained.
[0119] It should be further explained that the implementation process of the preset node partitioning strategy in this embodiment includes:
[0120] Based on the root node coordinates of each image in the continuous motion topology space, exemplarily, the hip joint node and center of gravity node coordinates are preset as the spatiotemporal reference points of all joint nodes and other joint node coordinates. The average coordinates of all joint nodes in five consecutive frames are first calculated, and this average is used as the stable coordinates of the skeleton's center of gravity. The root node is determined (in this embodiment, the joint node with the most connected nodes in the topology space, such as the hip joint, is selected as the root node). The directly adjacent nodes of the root node are extracted, exemplarily, the waist, left thigh, and right thigh nodes. The straight-line distance from the root node to the skeleton's center of gravity is calculated, and then the straight-line distance from each adjacent node to the skeleton's center of gravity is calculated separately. The distances of adjacent nodes to the root node and the center of gravity are compared: adjacent nodes with a distance less than the straight-line distance from the root node to the skeleton's center of gravity are grouped into the centripetal group, and adjacent nodes with a distance greater than the straight-line distance from the root node to the skeleton's center of gravity are grouped into the centrifugal group. The root node itself... Each node is treated as a subset 0. A 3D partition mask is generated for each node, where the three dimensions correspond to the three subsets respectively. The dimension value of the subset to which the node belongs is set to 1, and the other two dimension values are set to 0. The spatial coordinates, connection strength with adjacent nodes, and deviation features in the topological space of each node are extracted. These features are associated and combined with the corresponding 3D partition mask according to the node ID to form the regional feature items of a single node. The regional feature items of all nodes are integrated to obtain the regional node feature set of the undirected node graph corresponding to each image. Its function is to partition nodes by spatial position and centroid relationship, clearly distinguish the functional attributes of nodes in different regions. In this embodiment, centripetal group nodes mainly affect the stability of the action, while centrifugal group nodes mainly affect the amplitude of the action. This provides structured partition feature support for subsequent analysis of the action standard of each regional node and identification of regional action defects, thereby improving the regional targeting of action evaluation.
[0121] The continuous action topology space containing the trend of action deviation changes of the same action in different training cycles, the stage feature sequence of each action in the continuous action, and the region node feature set are input into the feature fusion layer constructed by the convolutional attention network to obtain the continuous action fusion topology space.
[0122] It should be further explained that the implementation process of the feature fusion layer in this embodiment includes:
[0123] Based on a continuous action topology space containing the trend of action deviation changes of the same action in different training cycles, including but not limited to node coordinates, deviation trend slope, and fluctuation amplitude; a phase feature sequence of each action in the continuous action, including but not limited to node coordinates and timestamps at the start, transition, and end times; and a regional node feature set, including but not limited to node partition masks, spatial coordinates, and connection strength, a feature fusion layer first converts the three types of features into a tensor of a unified dimension: the topology space features are unfolded into a three-dimensional tensor according to time frames; the phase feature sequence is embedded into the feature channels of the corresponding frame according to the timestamp; and the regional node feature set is concatenated with the node features through partition masks to form a spatial partition tensor. A convolutional attention network is constructed. The first layer uses a 1×1 convolution kernel to reduce the dimensionality of the topology space tensor to retain key deviation trend features, while performing a 3×3 convolution on the regional partition tensor to extract local spatial correlation features. The second layer introduces a channel attention mechanism to calculate each In this embodiment, the importance weights of feature channels are assigned as follows: deviation trend features are assigned a weight of 0.4, stage features are assigned a weight of 0.3, and region features are assigned a weight of 0.3. These weights are dynamically adjusted through global average pooling and sigmoid activation to perform weighted fusion of the convolutional features. The third layer uses a self-attention mechanism to capture the temporal correlation across frames. For example, the connection strength of stage features in adjacent frames is calculated by generating an attention map through the inter-frame feature similarity matrix to highlight key frame features. Finally, the weighted fused features are combined with the original topological space structure, and each feature item is associated by node ID to reconstruct a three-dimensional topological structure containing multi-dimensional feature associations, thus obtaining a continuous action fusion topological space. Its function is to integrate spatial topology, temporal stage, and regional partition features through a convolutional attention network, highlight the association weights of key features, provide more comprehensive fusion feature support for subsequent action evaluation, and improve the accuracy and robustness of action standard judgment.
[0124] Based on the continuous action fusion topology space combined with the standard continuous action node feature space, the output layer obtains the continuous action node coordinate sequence, node coordinate deviation loss, and the corresponding action deviation change trend and trend confidence. At the same time, the continuous action fusion topology space combined with the standard continuous action node feature space is input into the association causal layer to obtain the forward cumulative contribution path and path confidence corresponding to the non-standard action. In this embodiment, the output layer is preferably implemented by the support vector machine algorithm.
[0125] The loss function, which is constructed based on node coordinate deviation loss, trend confidence and path confidence, is used to train the spatiotemporal feature extraction model by combining the preset loss function threshold and training period.
[0126] It should be further explained that the specific implementation of the output layer in this embodiment is as follows:
[0127] Based on the continuous action fusion topology space (including but not limited to node coordinates, deviation trends, stage features, and regional attributes) and the standard continuous action node feature space (including but not limited to standard coordinates, temporal templates, and spatial constraints), a support vector machine algorithm configured in the output layer is used to first flatten the node features of the fusion topology space into one-dimensional vectors, and simultaneously process the standard continuous action node feature space into reference vectors. A multi-output support vector regression machine is constructed, and an independent regression model is trained for each node. The kernel function adopts the radial basis function, and the regularization parameter is optimized through cross-validation. With the standard coordinates as the target value, the mapping relationship between the node features of the current fusion topology space and the standard coordinates is fitted, and the continuous action node coordinate sequence is output. The Euclidean distance between the predicted coordinates and the standard coordinates is calculated as the node coordinate deviation loss. At the same time, based on the deviation distribution of the training samples, the 95% confidence interval of the loss is estimated through Bootstrap resampling. The rate of change of the deviation loss in the time dimension is used as the deviation change trend of the action, and the stability of the trend is used as the trend confidence. Its function is to accurately predict the action node coordinates and quantify the deviation through the nonlinear mapping capability of the support vector machine, providing coordinate-level guidance for action correction.
[0128] It should be further explained that the specific implementation of the causal layer in this embodiment is as follows:
[0129] Based on the differences between the continuous action fusion topology space and the standard continuous action node feature space, a causal inference algorithm configured with an associated causal layer is first constructed to build a causal graph of action nodes. Nodes represent action components, edges represent causal relationships, and weights represent conditional probabilities. Conditional dependencies between action components are statistically analyzed from historical training data; for example, the probability of hip joint compensation due to knee joint angle abnormalities. Current difference features are used as intervention variables, and counterfactual inference is used to calculate the causal effect of each difference on subsequent actions; for example, the influence of shoulder joint position deviation in the current frame on the elbow joint trajectory in the next frame. The causal effect values of each node are accumulated to form a forward cumulative contribution path. A Bayesian network is used to combine prior causal knowledge with the likelihood of current data to evaluate path confidence. Posterior probabilities are calculated for each causal relationship on the path, and the product of the posterior probabilities of all relationships on the path is used as the path confidence. Its function is to locate the root cause and propagation path of non-standard actions through causal inference, providing a theoretical basis for developing targeted correction strategies.
[0130] Example 2
[0131] Another embodiment of the present invention provides: an interactive teaching assessment auxiliary device based on multimodal fusion, comprising: a data acquisition component and a processor configured within the device, wherein the data acquisition component integrates a profiling module; and the processor integrates a strategy recommendation module, an auxiliary adjustment module, and a source feedback module.
[0132] The data acquisition component is used to collect basic user information, historical training achievement information, corresponding guidance plan information, and standard action information knowledge base of the configuration student terminal, and combine them with the built-in multimodal feature extraction algorithm and knowledge graph algorithm to obtain a user training action profile library.
[0133] The processor, based on a user training action profile library, real-time collected user basic information, and a configured teaching guidance strategy library, generates an optimal matching guidance scheme through a correction scheme matching model. While the user trains according to this scheme, it collects training action images in real time, and combines a spatiotemporal feature extraction model and an action evaluation algorithm to obtain a continuous corrective action sequence and corresponding correction standard evaluation scores after guidance. Based on the continuous corrective action sequence and corresponding correction standard evaluation scores, and a standard action information knowledge base, it obtains the horizontal and vertical spatiotemporal change chains of non-standard actions through a forward inference model, and feeds these chains back to the configured teacher. Simultaneously, it collects feedback from the teacher and adjusts the matching guidance scheme in real time until the corresponding user training action meets the conditions. The guidance scheme for user training actions that meet the conditions is then stored in the standard action information knowledge base for management.
[0134] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art, under the guidance of the present invention, can make changes, modifications, substitutions and variations to the above embodiments without departing from the spirit and scope of the claims. All of these variations are within the protection scope of the present invention.
[0135] If the technical solution disclosed herein involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution disclosed herein involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
Claims
1. An interactive teaching assessment support system based on multimodal fusion, characterized in that, include: The module includes a profile module, a strategy recommendation module, an auxiliary adjustment module, and a source tracing and feedback module. The profiling module is used to collect basic user information, historical training achievement information, and corresponding guidance plan information and standard action information knowledge base of the student terminal, and combine them with knowledge graph algorithm to obtain user training action profiling library; the profiling module includes a data acquisition unit, a first data analysis unit, a second data analysis unit, and a profiling unit. The data acquisition unit is used to collect basic user information, historical motion assessment text and image information, corresponding guidance scheme information and standard motion information knowledge base of the user's terminal, and classify, arrange and standardize the collected information according to the user and the collection timestamp. The first data analysis unit is used to obtain user action evaluation information triples and corresponding action spatiotemporal feature sets by combining user historical motion evaluation action text information and image information with entity relationship extraction algorithm and image recognition algorithm; The second data analysis unit is used to obtain the motion correlation relationship at continuous time points and the probability that the next motion will be non-standard under the condition that the previous motion is non-standard, based on the standard motion information knowledge base and the correlation analysis algorithm with built-in expert experience. The profiling unit is used to construct a user training action profiling library based on user basic information, user action evaluation information triplet, corresponding action spatiotemporal feature set, standard action information knowledge base, action correlation at continuous time points, and the probability of subsequent actions being non-standard under the condition that the previous action is non-standard, through graph algorithm. The strategy recommendation module is used to obtain the optimal matching guidance scheme for the corresponding user by combining the user training action profile library with the real-time collected user basic information and the configured teaching guidance strategy library, and through the correction scheme matching model. The auxiliary adjustment module is used to provide auxiliary guidance according to the matching guidance scheme, and to collect the training action information of the student in real time. Combined with the preset spatiotemporal feature extraction model and comprehensive evaluation algorithm, it obtains the continuous action correction space after guidance and the corresponding correction standard evaluation score. The source tracing feedback module is used to obtain the spatiotemporal change chain of non-standard movements in the corresponding teaching scenario based on the continuous corrective movement sequence after guidance and the corresponding correction standard evaluation score, combined with the standard movement information knowledge base, through a forward reasoning model; and to feed back the spatiotemporal change chain of non-standard movements to the alarm device configured on the teacher's end for real-time early warning, and to collect feedback from the teacher's end to adjust the matching guidance plan in real time until the corresponding user's training movements meet the conditions, and at the same time, to store the guidance plan of the user's training movements that meet the conditions in the standard movement information knowledge base for management.
2. The interactive teaching assessment support system based on multimodal fusion as described in claim 1, characterized in that, The spatiotemporal variation chain of non-standard movements includes the spatiotemporal variation chain of lateral non-standard movements constructed by different training movements under the same guidance scheme in the current training cycle, and the spatiotemporal variation chain of longitudinal non-standard movements constructed by the same training movement under the same guidance scheme in different training cycles.
3. The interactive teaching assessment support system based on multimodal fusion as described in claim 2, characterized in that, The user training action profile database includes a baseline layer, a training layer, and an inference layer; The baseline layer stores basic user information; the training layer stores the spatiotemporal features of training actions in each historical guidance scheme, along with corresponding training evaluation scores and action guidance correction information; the inference layer stores a standard action information knowledge base and, by combining the action correlation at consecutive time points and the probability of subsequent actions being non-standard under the condition that the previous action was non-standard, constructs causal inference connections between consecutive actions to obtain a causal inference chain for training actions. This chain is then mapped to the training layer under the corresponding user to evaluate and infer correct the user's action information.
4. The interactive teaching assessment support system based on multimodal fusion as described in claim 3, characterized in that, The auxiliary adjustment module includes a spatiotemporal feature extraction unit and an evaluation unit; The spatiotemporal feature extraction unit is used to obtain the associated spatiotemporal feature space of the user's optimal matching guidance scheme by combining the training action information corresponding to the optimal matching guidance scheme with a preset spatiotemporal graph convolution model. The evaluation unit is used to obtain the training evaluation score of each action in the optimal matching guidance scheme by combining the spatiotemporal correlation feature space of the user's optimal matching guidance scheme with the causal inference chain of the training actions through a comprehensive evaluation algorithm.
5. The interactive teaching assessment support system based on multimodal fusion as described in claim 4, characterized in that, The source tracing feedback module includes a reverse reasoning unit and a feedback adjustment unit; The reverse reasoning unit is used to perform forward causal reasoning of non-standard actions based on the training evaluation score of each action in the optimal matching guidance scheme, the action spatiotemporal change chain of historical longitudinal non-standard actions, and the causal reasoning connection in the causal reasoning chain of the training action, and to obtain the action spatiotemporal change chain of real-time lateral non-standard actions and the action spatiotemporal change chain of real-time longitudinal non-standard actions based on the training evaluation score and the color mapping space of HSV. The feedback adjustment unit is used to feed back the spatiotemporal change chains of real-time lateral non-standard movements and real-time longitudinal non-standard movements, along with their corresponding training evaluation scores, to the profiling unit for storage. Simultaneously, it judges the optimal matching guidance scheme's comprehensive training evaluation score against a preset evaluation threshold. If the comprehensive training evaluation score is greater than the preset threshold but any movement does not meet the corresponding standard movement information, the corresponding movement information and the non-standard movement tracing information obtained through the spatiotemporal change chain of the lateral non-standard movements are fed back to the corresponding instructor for guidance and correction. If the comprehensive training evaluation score is less than or equal to the preset evaluation threshold, the corresponding guidance scheme information, training evaluation score, and tracing results are fed back to the corresponding instructor for real-time adjustment of the optimal matching guidance scheme.
6. The interactive teaching assessment support system based on multimodal fusion as described in claim 5, characterized in that, The process of constructing the correction scheme matching model includes: Based on the spatiotemporal feature information of training actions and the corresponding training evaluation scores in each historical guidance scheme in the user training action profile library, the spatiotemporal features of the same training action and the trend of training evaluation score changes in different training cycles are extracted through the non-standard action layer in the correction scheme matching model to obtain the user's longitudinal non-standard action feature change space and training evaluation score fluctuation coefficient. Based on the user's longitudinal non-standard action feature change space and training evaluation score fluctuation coefficient, combined with the action spatiotemporal change chain of the lateral non-standard action in the user training action profile library, the degree of correlation coverage adjustment of the current non-standard action to the related actions in the same training cycle is obtained. Simultaneously, based on the fluctuation coefficient of the training evaluation score of the associated movements under the degree of related coverage adjustment in different training cycles after the current non-standard movements are adjusted, the longitudinal effect transfer rate of the same associated movement in different training cycles can be obtained.
7. The interactive teaching assessment support system based on multimodal fusion as described in claim 6, characterized in that, The process of constructing the correction scheme matching model also includes: Based on the training actions within each training cycle, a horizontal action node sequence is constructed, and horizontal association connections are constructed by adjusting the degree of association coverage of related actions in the same training cycle after the current non-standard action is adjusted. Based on the fluctuation coefficient of the training evaluation score of the same training action of a user in different training cycles and the longitudinal effect transfer rate of the same related action in different training cycles, a vertical correlation transfer connection is constructed. Based on the constructed horizontal action nodes of different training cycles, combined with horizontal and vertical correlation transfer connections and topological space layers, the user's historical training action effect correlation transfer space is obtained. Based on the user's historical training action effect association transfer space and the teaching guidance strategy library, a matching search is performed through a matching layer configured with the maximum adjustment evaluation effect function to generate a matching guidance scheme sequence and a corresponding matching loss function; the maximum adjustment evaluation effect function is a function constructed by fitting the corresponding non-standard action and related action correction standard evaluation score after the matching scheme is trained, the probability of the subsequent action being non-standard under the condition that the previous action is non-standard, and the effect transfer rate with a linear fitting function and taking the maximum value. The model is trained based on the matching guidance scheme sequence and the corresponding matching loss function, combined with the preset loss function threshold and the set model training period. The trained correction scheme matching model is obtained, and the optimal matching guidance scheme corresponding to the maximum adjustment evaluation effect is output.
8. The interactive teaching assessment support system based on multimodal fusion as described in claim 7, characterized in that, The construction process of the spatiotemporal feature extraction model includes: The system acquires training action information and corresponding standard action image information for consecutive time periods within different training cycles of the user and performs enhancement and denoising processing. Based on the processed images and the expert labeling algorithm pre-trained with skeletal node information, it performs key skeletal node labeling and center of gravity node labeling under different actions. It obtains the first continuous action labeling image sequence of the same training cycle and the second labeling image sequence of the same training action corresponding to different training cycles. The first continuous action labeled image sequence of the same training period and the second labeled image sequence of the same training action corresponding to different training periods are simultaneously input into the OpenPose action estimation layer to obtain the continuous action node feature space of the same training period and the node bias feature space of the same training action corresponding to different training periods constructed with the standard continuous action node feature space. Based on the feature space of continuous action nodes in the same training cycle, combined with the body structure, the connection edges between each node are constructed to obtain the undirected node graph corresponding to each image. Based on the priority of the limb action implementation of the corresponding node in each image at consecutive time points in the undirected node graph corresponding to each image, a continuous action connection is constructed. Based on the undirected node graph and continuous action connections corresponding to each image, a continuous action topology space is constructed through a topology space layer.
9. The interactive teaching assessment support system based on multimodal fusion as described in claim 8, characterized in that, The construction process of the spatiotemporal feature extraction model also includes: Based on the continuous action topology space combined with the node deviation feature space corresponding to the same training action in different training cycles, a continuous action topology space containing the trend of action deviation change of the same action in different training cycles is obtained through a time trend extraction layer. Simultaneously, the continuous action topology space is input into the limb curvature recognition layer to divide the action into stages, thereby obtaining the stage feature sequence of each action in the continuous action; the stage feature sequence is constructed from the node spatial coordinates corresponding to the start time, the node spatial coordinates corresponding to the turning time, and the node spatial coordinates corresponding to the end time of the corresponding action. Based on the root node and centroid node of each action in the continuous action topology space, combined with a preset node partitioning strategy, the region node feature set of the undirected node graph corresponding to each image is obtained. The continuous action topology space containing the trend of action deviation changes of the same action in different training cycles, the stage feature sequence of each action in the continuous action, and the region node feature set are input into the feature fusion layer constructed by the convolutional attention network to obtain the continuous action fusion topology space. Based on the continuous action fusion topology space combined with the standard continuous action node feature space, the output layer obtains the continuous action node coordinate sequence, node coordinate deviation loss, and the corresponding action deviation change trend and trend confidence. At the same time, the continuous action fusion topology space combined with the standard continuous action node feature space is input into the association causal layer to obtain the forward cumulative contribution path and path confidence corresponding to non-standard actions. The loss function, which is constructed based on node coordinate deviation loss, trend confidence and path confidence, is used to train the spatiotemporal feature extraction model by combining the preset loss function threshold and training period.
10. An interactive teaching assessment aid device based on multimodal fusion, implemented based on any one of claims 1-9, characterized in that, include: A data acquisition component and processor are configured within the device, wherein the data acquisition component integrates a portrait module; The processor integrates a strategy recommendation module, an auxiliary adjustment module, and a source tracing feedback module. The data acquisition component is used to collect basic user information, historical training achievement information, and corresponding guidance scheme information and standard action information knowledge base of the student terminal, and combine them with the built-in multimodal feature extraction algorithm and knowledge graph algorithm to obtain a user training action profile library. The processor, based on a user training action profile library, real-time collected user basic information, and a configured teaching guidance strategy library, generates an optimal matching guidance scheme through a correction scheme matching model. While the user trains according to this scheme, it collects training action images in real time, and combines a spatiotemporal feature extraction model and an action evaluation algorithm to obtain a continuous corrective action sequence and corresponding correction standard evaluation scores after guidance. Based on the continuous corrective action sequence and corresponding correction standard evaluation scores, and a standard action information knowledge base, it obtains the horizontal and vertical spatiotemporal change chains of non-standard actions through a forward reasoning model, and feeds these chains back to the configured teacher. Simultaneously, it collects feedback from the teacher and adjusts the matching guidance scheme in real time until the corresponding user training action meets the conditions. The guidance scheme for user training actions that meet the conditions is then stored in the standard action information knowledge base for management.
Citation Information
Patent Citations
Interactive physical education method based on large model
CN120277511A
Generative and multi-modal sensing integrated agent learning system
CN120277628A
Cited By
Method and system for evaluating and guiding personalized growth of teenagers based on multi-dimensional ability model
CN122288244A