Limb motion learning clue labeling method and system based on cross-modal mapping
By using cross-modal mapping methods and leveraging speech-to-text conversion and motion feature extraction techniques to generate visual cues, the problem of low efficiency in motion instruction in online fitness courses is solved, achieving intuitive presentation of key motion points and improving learning outcomes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2025-07-10
- Publication Date
- 2026-05-15
AI Technical Summary
In online fitness courses, movement instruction mainly relies on verbal and textual prompts, making it difficult for learners to efficiently understand the key points of the movements. Existing visualization production processes are cumbersome, inefficient, and fail to meet learning needs.
By employing a cross-modal mapping method, speech-to-text conversion and BERT structure parsing of text entities are used. Combined with BlazePose and Bodypix models to extract action features, visual cues are generated and placed in corresponding spatiotemporal locations to form intuitive body movement learning cues.
It presents the key points of the movements intuitively and vividly, improves the learning effect, simplifies the production process, is highly scalable, is suitable for learning multiple types of movements, supports personalized customization, and is applicable to fitness movements and other movement learning fields.
Smart Images

Figure CN120977153B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent motion-assisted learning technology, specifically to a method and system for annotating limb movement learning cues based on cross-modal mapping. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] With the rapid development of information technology and the increasing health awareness of the public, online fitness has become an important way to participate in exercise. Through online fitness platforms and applications, users can participate in various fitness courses anytime, anywhere, learning and practicing exercises through digital media such as live video streaming and pre-recorded videos, thus meeting their fitness needs. Among these, some health-oriented physical exercises, such as Tai Chi and yoga, have become widely popular. However, because these exercises are relatively complex and require precise postures, learners often find it difficult to master the movements efficiently and quickly, resulting in learning outcomes that fall short of expectations.
[0004] To improve the efficiency and effectiveness of motor learning, online motor instruction courses primarily rely on verbal explanations, supplemented by textual prompts (such as subtitles) for guidance. However, due to the abstract nature of verbal instruction, learners sometimes struggle to directly understand the instructor's intentions and expressions, creating a barrier to content delivery. Leveraging multimedia technology, some courses incorporate visual cues such as graphics and animations into videos, using intuitive visual information to help learners understand key aspects of the movements and compensate for the limitations of a single verbal modality. However, most of these courses rely on manually creating visualizations using video editing software, a cumbersome and inefficient process that fails to meet application requirements. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method and system for annotating limb movement learning cues based on cross-modal mapping. By constructing optimized graphical cues, the key points of movements can be presented more intuitively and vividly, helping learners to clearly understand the key points of movements through intuitive visual prompts and improving learning effectiveness.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, the present invention provides a method for annotating limb movement learning cues based on cross-modal mapping.
[0008] A method for annotating limb movement learning cues based on cross-modal mapping includes the following process:
[0009] Perform language parsing on the audio narration in the action instruction video, or directly extract the audio narration subtitles from the action instruction video to obtain sentence-by-sentence text and corresponding time markers;
[0010] The sentence-by-sentence text is parsed to extract the entities in the sentence-by-sentence text and the roles corresponding to each entity;
[0011] Extracting the movement characteristics of the demonstrator from instructional videos;
[0012] The identified entities are mapped onto graphics to generate visual cues;
[0013] Based on visual cues, language parsing results, and motion features, the visual cues are placed in their corresponding time and space locations to obtain the final annotation results for limb movement learning cues.
[0014] In one implementation of the first aspect of the present invention, language parsing is performed on the audio narration in the action instruction video, including:
[0015] A speech-to-text neural network is used to extract sentence-by-sentence text from the audio narration of action instruction videos, and a timetable corresponding to the start and end times of each sentence-by-sentence text is obtained. The video actions corresponding to each start and end time period are segmented to generate text subtitles.
[0016] In one implementation of the first aspect of the present invention, parsing the sentence-by-sentence text and extracting entities and their corresponding roles from the sentence-by-sentence text includes:
[0017] The BERT architecture is used to generate text representations of input words in sentence-by-sentence text. Sub-word, sentence, and position embeddings are generated through word segmentation and high-dimensional mapping. A two-stage bidirectional processing method is used with a Transformer architecture. In the first stage, large-scale pre-training is performed on a pre-set encyclopedia corpus. Feature extraction is carried out by occlusion language and sentence prediction tasks while taking into account context. Supervised fine-tuning is used in the second stage to obtain multi-dimensional vector representations of sentence-by-sentence text. Based on the multi-dimensional vector representations of sentence-by-sentence text, a bidirectional long short-term memory network is used to extract key contextual features of entities. A conditional random field structure is used to post-process the output of the bidirectional long short-term memory network. The sequence is labeled using a global optimal solution.
[0018] In one implementation of the first aspect of the present invention, the entity includes: object, action, direction, data, and metaphor;
[0019] The object is the agent or the passive argument, the action is the action predicate, the direction is the direction argument, the data is the data argument, and the metaphor is the metaphor argument.
[0020] As a further limitation of the first aspect of the invention, the motion characteristics of the subject demonstrating the action are extracted from the action instruction video, including:
[0021] The BlazePose network is used for human pose estimation, to identify the key skeletal coordinates of the demonstration subject in the video, and to repair misaligned and missing points and smooth and denoise the data to obtain the skeletal joint coordinates.
[0022] The Bodypix human semantic segmentation model is used to identify and segment various body parts, and the recognition edges are optimized and smoothed.
[0023] A fusion model based on CNN and Transformer is used to identify and locate the demonstrated movements in the action instruction video, and to identify the subcategories and time positions of the movements in the action instruction video.
[0024] As a further limitation of the first aspect of the present invention, for the object entity, the position obtained by posture estimation and part segmentation is visualized by color block or box selection effect to highlight the position of the drawn muscles and joints.
[0025] For motion entities, motion trajectories are drawn using joint positions obtained from pose estimation and combined with other entities for joint description.
[0026] For directional entities, the joint positions obtained by pose estimation are visualized using arrowhead shapes of different sizes according to the range of motion.
[0027] For data entities, visualize predefined animation formats;
[0028] For metaphorical entities, predefined animation forms are visualized with reference to pose estimation and part segmentation results.
[0029] As a further limitation of the first aspect of the present invention, based on visual cues, language parsing results, and motion features, the visual cues are placed into corresponding temporal and spatial locations to obtain the final limb movement learning cue annotation results, including:
[0030] Based on motion characteristics, the time boundaries of each motion stage in the motion instruction video are obtained. The same stage of motion is defined as a motion-semantic time unit, including multiple motion steps corresponding to each sentence. The start time of each sentence is used as the presentation time of the corresponding visual cue visualization, and the time position is arranged accordingly.
[0031] Based on the topological structure of the human body, a four-level pyramid is defined, including joints, limbs, half-torso, and full torso. The description of joints is directly located to the key point position, which serves as the reference position for the layout.
[0032] For limb regions consisting of two or more joints, including the upper and lower limbs, corresponding mapping is performed using region segmentation coverage and skeletal connections.
[0033] For the half-torso and full-torso, the main skeletal lines are used to represent them;
[0034] For the direction argument, calculate the movement direction based on the keypoint trajectory of the specified joint, and place the arrow at the current time. and a given time threshold range Defined pointer location place, and towards The trailing edge is gradually enhanced with color and transparency gradients to improve its indicativeness;
[0035] For joint angle content in data arguments, the angle graphic data is embedded into the corresponding joint area by combining the joint connection line, and the angle value of the current action is superimposed and displayed; for distance values, they are converted into dashed lines between joints and the corresponding values are displayed.
[0036] For metaphorical arguments, a set of graphical metaphors is used to pre-define some high-frequency metaphor types and match them with corresponding body movements.
[0037] Secondly, the present invention provides a limb movement learning cue annotation system based on cross-modal mapping.
[0038] A limb movement learning cue annotation system based on cross-modal mapping, comprising:
[0039] The text extraction unit is configured to: perform language parsing on the audio narration in the action instruction video, or directly extract the audio narration subtitles in the action instruction video to obtain sentence-by-sentence text and corresponding time markers;
[0040] The semantic parsing unit is configured to: parse the sentence-by-sentence text and extract entities in the sentence-by-sentence text and the roles corresponding to each entity;
[0041] The motion feature extraction unit is configured to extract motion features of the subject demonstrating the motion from the motion instruction video.
[0042] The cross-modal action cue generation unit is configured to map the identified entities to graphics and generate visual cues.
[0043] The spatiotemporal motion cue layout unit is configured to: based on visual cues, language parsing results, and motion features, lay out the visual cues to the corresponding time and space locations to obtain the final limb motion learning cue annotation results.
[0044] Thirdly, the present invention provides a computer device, comprising: a processor and a computer-readable storage medium;
[0045] A processor, adapted to execute computer programs;
[0046] A computer-readable storage medium storing a computer program that, when executed by the processor, implements the limb movement learning cue annotation method based on cross-modal mapping as described in the first aspect of the present invention.
[0047] Fourthly, the present invention provides a computer-readable storage medium storing a computer program adapted to be loaded by a processor and executed as described in the first aspect of the present invention for labeling limb movement learning cues based on cross-modal mapping.
[0048] Compared with the prior art, the beneficial effects of the present invention are:
[0049] 1. This invention innovatively proposes a method for labeling limb movement learning cues based on cross-modal mapping. It is simple, efficient, and low-cost. By parsing the language of spoken explanations to obtain sentence-by-sentence text and time stamps, and then parsing the text to extract entities and roles, the process is clear and easy to implement. It can automatically generate graphical visual cues and embed them into videos, achieving enhanced content through virtual-real fusion. It maps identified entities to graphically generated visual cues and accurately places them in corresponding spatiotemporal locations based on language parsing and motion characteristics, presenting key movement points in a more intuitive and vivid way. Learners can clearly understand the key points of the movements with the help of visual prompts, greatly improving learning effectiveness. It supports coaches in efficiently and quickly automatically generating enhanced movement instruction videos, simplifying the manual production process, reducing the difficulty of creation, and offering strong scalability. It allows online learners to upload approved video tutorials for production, providing flexibility and personalized customization capabilities. It has a wide range of applications, not only suitable for learning various fitness movements such as dance, Tai Chi, and yoga, but also easily extended to other movement learning fields, helping to enhance interest in movement learning and providing strong support for promoting online movement learning and national fitness.
[0050] 2. This invention innovatively proposes a method for annotating body movement learning cues based on cross-modal mapping. Employing a speech-to-text conversion neural network, it can accurately and efficiently extract sentence-by-sentence text from audio explanations, while simultaneously generating a timetable corresponding to the start and end times of each sentence, providing a precise time reference for subsequent operations. The video actions corresponding to each start and end time period are segmented, closely linking video content with text information, facilitating in-depth analysis of specific actions. The generated text subtitles not only provide learners with clear textual assistance, facilitating comprehension of the explanations, but are also particularly suitable for learners with hearing impairments or in noisy environments. This lays a solid foundation for the entire method of annotating body movement learning cues based on cross-modal mapping, enabling subsequent operations such as entity extraction, mapping graphics, and layout visualization of cues to be carried out systematically based on accurate time and textual information, thereby improving the accuracy and practicality of the final annotation results.
[0051] 3. This invention innovatively proposes a method for annotating limb movement learning cues based on cross-modal mapping. It uses a BERT structure to process sentence-by-sentence text, generating and fusing multiple embedding representations through word segmentation and high-dimensional mapping. Then, it undergoes a two-stage bidirectional Transformer process: large-scale pre-training with contextual feature extraction followed by supervised fine-tuning, accurately obtaining multi-dimensional vector representations of sentence-by-sentence text and fully mining semantic information, laying a solid foundation for subsequent processing. Based on these multi-dimensional vector representations, a bidirectional long short-term memory network is used to extract key features of entity context, effectively capturing the semantic relationships of entities in the text. A conditional random field structure is then used for post-processing, labeling sequences using a globally optimal solution, accurately identifying entities and their corresponding roles in sentence-by-sentence text. This series of operations is closely integrated, from semantic understanding to feature extraction to accurate annotation, providing an accurate basis for subsequently mapping entities to graphics to generate visual cues. This ensures that the final limb movement learning cue annotation results accurately reflect the key points of the movements, improving learners' understanding and learning outcomes.
[0052] 4. This invention innovatively proposes a method for annotating limb movement learning cues based on cross-modal mapping. It clearly defines entities encompassing objects, actions, directions, data, and metaphors, and clearly defines the corresponding roles of each entity, providing a clear framework for subsequent precise processing. Objects, as the agent or receiver arguments, accurately define the executor and receiver of the action, allowing learners to clearly understand the initiator and target of the action. Actions, as predicates, highlight the core action content, which is key to understanding the essentials of the action. Direction arguments clarify the direction of the action, preventing learners from making deviations due to unclear direction. Data arguments quantify action-related parameters, such as speed and amplitude, making action standards more specific and measurable. Metaphor arguments provide a visual aid to understand complex actions. This meticulous classification and role definition ensures that the entity information extracted from the text is complete and accurate, laying a solid foundation for generating visual cues. Ultimately, this allows the final annotation results to comprehensively and accurately present the key points of the action, improving learning effectiveness.
[0053] 5. This invention innovatively proposes a method for annotating limb movement learning cues based on cross-modal mapping. It employs the BlazePose network for human pose estimation, accurately identifying the key skeletal coordinates of the demonstrator and correcting errors and omissions through smoothing and denoising, ensuring the accuracy of skeletal joint coordinates and providing a reliable foundation for subsequent movement analysis. The Bodypix human semantic segmentation model identifies and segments various body parts and optimizes edges, making the definition of body parts clearer and facilitating a deeper understanding of the involvement of each part in the movement. A fusion model based on CNN and Transformer identifies and locates the demonstrator's movements, accurately determining the category and temporal location of the movement subclasses, comprehensively grasping the details and progress of the movement. By combining these three technologies, from skeletons and body parts to the overall movement, motion features are extracted comprehensively, providing rich and accurate information for generating visual cues. This allows the final annotation results to present the key points of the movement more meticulously and accurately, improving learners' understanding and mastery of the movement.
[0054] 6. This invention innovatively proposes a method for annotating limb movement learning cues based on cross-modal mapping. For object entities, muscle and joint positions are highlighted by color blocks or box selections, allowing learners to quickly focus on key parts and clarify the force points and areas of action of the movement. Movement entities are drawn with motion trajectories and described in combination with other entities, making the dynamic process of the movement clear and facilitating learners' understanding of the continuity and logic of the movement. Directional entities are visualized with arrows of different sizes, intuitively showing changes in the direction of the movement and avoiding errors caused by ambiguity in direction. Data entities are visualized using predefined animations, transforming abstract data into intuitive images, allowing learners to easily grasp movement-related parameters. Metaphorical entities are visualized with reference to posture estimation and part segmentation results, using a visual approach to assist in understanding complex movements, stimulating learners' interest, and improving the overall efficiency of learners' understanding and mastery of movement essentials.
[0055] 7. This invention innovatively proposes a method for annotating limb movement learning cues based on cross-modal mapping. It determines the temporal boundaries of movement stages through motion features, defines motion-semantic time units, and lays out visual cues according to the explanation time, ensuring that the cues are presented synchronously with the explanation, facilitating learners' understanding of the correspondence between movement and explanation. A four-level pyramid is defined based on the human body's topology, providing precise layout references for different parts and ensuring accurate positioning of visual cues. Different limb regions are mapped using methods such as region segmentation and skeletal connections; the half-body and full-body torso are presented using connections of major bones, clearly demonstrating the movement structure. Directional arguments are calculated using keypoint trajectories and arrows are laid out, with gradient effects enhancing indicativeness. Data arguments visualize angle and distance values embedded in corresponding regions, intuitively presenting movement parameters. Metaphorical arguments pre-set graphical metaphor sets and match movements, using visual aids to assist in understanding complex movements, reducing learning difficulty, and comprehensively improving learners' mastery of movements.
[0056] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0057] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0058] Figure 1 A flowchart illustrating a cross-modal mapping method for annotating limb movement learning cues, provided as an exemplary embodiment of the present invention;
[0059] Figure 2 A schematic diagram of a cross-modal mapping method for annotating limb movement learning cues, provided as an exemplary embodiment of the present invention;
[0060] Figure 3 A deep hybrid named entity recognition model architecture based on BERT-BiLSTM-CRF is provided as an exemplary embodiment of the present invention;
[0061] Figure 4 A flowchart illustrating a semantic role labeling method based on a prompting learning approach, provided as an exemplary embodiment of the present invention;
[0062] Figure 5 A schematic diagram illustrating the construction of a semantic role annotation template based on cue learning, as provided in an exemplary embodiment of the present invention;
[0063] Figure 6 A schematic diagram of human motion feature extraction provided as an exemplary embodiment of the present invention;
[0064] Figure 7 A schematic diagram of the design space for enhanced visual cues in VR motion learning and training, provided as an exemplary embodiment of the present invention;
[0065] Figure 8 A schematic diagram of a content selection interface provided for an exemplary embodiment of the present invention;
[0066] Figure 9 A schematic diagram of a content creation interface provided for an exemplary embodiment of the present invention;
[0067] Figure 10 A schematic diagram of generated partially enhanced video content provided as an exemplary embodiment of the present invention;
[0068] Figure 11 A flowchart illustrating a cross-modal mapping limb movement learning cue annotation system provided as an exemplary embodiment of the present invention;
[0069] Figure 12 A schematic diagram of a computer device provided for an exemplary embodiment of the present invention. Detailed Implementation
[0070] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0071] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0072] This implementation proposes a cross-modal mapping method for annotating limb movement learning cues, such as... Figure 1 As shown, the process includes the following:
[0073] S1: Users upload action instruction videos containing verbal explanations. The server first uses a speech-to-text conversion network to convert the verbal explanations in the video into sentence-by-sentence text and extracts the corresponding time stamps.
[0074] S2: Using the semantic parsing module, the extracted sentence-by-sentence text is parsed to extract the entity categories and the corresponding role relationships of each entity, and to extract action elements and their key descriptions.
[0075] S3: Using the motion feature extraction module, extract the motion features of the subject demonstrating the motion from the video, including multi-level motion features such as posture, limb outline, composition of action subclasses, and time and position.
[0076] S4: Using the spatiotemporal action cue generation module, based on the language parsing results, the identified entities are mapped to graphics to generate visual cues;
[0077] S5: Based on the generated cues, language analysis, and motion understanding results, the spatiotemporal motion cue layout module is used to place the visual cues in appropriate time and space positions to ensure accurate embedding of the visualization effect, thereby generating an enhanced motion instruction video.
[0078] In step S1 of this implementation, a speech-to-text conversion neural network is used to obtain a timetable corresponding to the start and end times of each sentence of text from the video audio narration, and the video action corresponding to each subtitle segment is segmented and an SRT subtitle file is generated.
[0079] In step S2 of this implementation, specifically, it includes:
[0080] S201: Naming entities and semantic roles are assigned to the corpus components in the language used to explain motion.
[0081] S202: Employs a named entity recognition module based on the deep hybrid model BERT-Bi-LSTM-CRF to identify visual entities with specific meanings from the explanatory text and mark their locations;
[0082] S203: Using the GPT language model combined with prompting learning method, we analyze the semantic relevance and role of text entities, and determine the semantic relationship between each component in the sentence and the predicate.
[0083] In step S201 of this implementation method, named entities and semantic roles are labeled for the motion explanation language, specifically including:
[0084] Define the following types of entities that are key points in the explanation of motion:
[0085] E1-Object (OBJ): refers to the subject performing the action, such as joints, muscles, limbs, knees, upper arms and other body parts, which act as the [actor] or [receiver] of the action;
[0086] E2 - Action (MOT): refers to an action performed by a body part or a static posture formed, such as "bending (knee)" or "stepping (with left foot)," as the [action predicate] role;
[0087] E3 - Direction (DIR): Describes the direction of an action, such as "(stepping to the right)", and acts as a [direction argument].
[0088] E4-Data (DAT): Describes the quantitative key points of the movement, defines specific numerical standards such as the angle and distance of the movement, such as "lower leg perpendicular to the ground (90°)" and "oblique (45°) direction", as the role of [data argument];
[0089] E5 - Metaphor (MET): refers to descriptive statements that use analogy, categorization, symbolism, etc. to express the meaning and experience of an action, such as "holding hands as if holding (a ball)" or "standing tall and firm like (a tree)," as the role of [metaphorical argument].
[0090] In this implementation, a corpus of motion explanations is constructed. The text data is cleaned using regular expressions and rules to remove garbled characters and unnecessary spaces and symbols. Entities and roles are labeled in accordance with language annotation specifications. Taking the action predicate as the center, the core arguments related to the central word are labeled, including the corresponding agent, receiver, data arguments, metaphor arguments, and direction arguments, which constitute semantic units with action implementation correlation.
[0091] In step S202 of this implementation, the specific method for performing named entity recognition is as follows:
[0092] Input of explanatory texts for sports First, the BERT structure is used to process the input words in the sentence. The text representation is generated by word segmentation and high-dimensional mapping to produce word, sentence, and positional embeddings for sentences. These embeddings are then fused together, and a two-stage bidirectional training process is performed using a Transformer architecture. The first stage involves large-scale pre-training on the Chinese Wikipedia corpus, extracting features while considering context through Masked Language Model (MLM) and Next Sentence Prediction (NSP) tasks. Simultaneously, a second stage of fine-tuning is conducted using supervised fine-tuning to obtain the text representation. 3D vector representation .
[0093] Based on the BERT encoder output, a bidirectional long short-term memory (BiLSTM) network is used to further extract key contextual features for named entity recognition in physical education teaching. Gating units are used in conjunction with a forward-backward structure to extract temporal information.
[0094] (1);
[0095] in, This represents the feature representation obtained by forward LSTM computation from front to back. This is the feature representation obtained by back-to-forward computation using a reverse LSTM. This is the result of concatenating bidirectional feature vectors.
[0096] To further model the dependencies between adjacent labels, a conditional random field structure is used to post-process the output of the BiLSTM model. The sequence is labeled using a global optimum approach. The hidden layer input results generated by the BiLSTM are linearly transformed to obtain the output score matrix corresponding to each entity label. Its size is ,in The number of words. For the number of tags, For the first sentence The word was marked as the first The probability of each label for the obtained predicted label sequence Its scoring function is:
[0097] (2);
[0098] in, For the transition fraction matrix, For from the label Move to label The score.
[0099] The conditional probability of the CRF generated by the predicted sequence Y is:
[0100] (3);
[0101] The Viterbi algorithm is used to maximize the conditional probability of the CRF to obtain the globally optimal labeled sequence.
[0102] (4);
[0103] in, For real labeled sequences, For all possible labeled sequences.
[0104] In step S203 of this implementation, semantic role labeling is performed, specifically including:
[0105] Construct prompts for the input sentence This study combines cue learning, few-shot demonstration, and self-validation methods to characterize entities within sentences, analyzing the semantic relevance and roles of text entities. The cue learning method constructs semantic role-labeling cue templates including a first-sentence task, a second-sentence example (few-shot demonstration), and a requirement input, guiding GPT to accurately identify semantic roles in complex multi-relational motion data. Simultaneously, the k-Nearest-Neighbor method is used to retrieve input sequences from the training set. The algorithm calculates the similarity between the input and the representation in the training samples, finds the nearest neighbors for each label, and constructs a few-sample demonstration set using the sentences associated with it. Finally, a self-verification strategy is used to verify the accuracy of the semantic role relationship labeling, that is, the model is asked to improve the accuracy of self-verification by prompting words.
[0106] In step S3 of this implementation, the specific methods include:
[0107] S301: The BlazePose network is used for human pose estimation, to identify the key skeletal coordinates of the subject in the video, and to repair misaligned and missing points and smooth and denoise the video, so as to obtain more accurate skeletal joint coordinates.
[0108] S302: The Bodypix human semantic segmentation model is used to identify and segment various body parts, and the identification edges are optimized and smoothed to improve the continuity and stability of the segmentation results.
[0109] S303: Employs a fine-grained motion understanding method based on CNN and Transformer to identify and locate demonstration actions in videos, and to identify the subcategories and temporal positions of actions in the video.
[0110] In step S303, motion recognition and localization are performed, specifically including:
[0111] Spatiotemporal data augmentation and dual-view sampling are performed on the action instruction video, and the data is input into the feature encoder. Spatiotemporal information is extracted and aggregated across layers through a CNN-based hybrid appearance feature encoding structure and a Transformer-based multi-scale temporal feature encoding structure, fully integrating spatiotemporal context information. Sequence contrast loss is applied to minimize the embedding similarity between views, and frame-level dense modeling is performed to obtain the start and end time boundaries of each fine-grained action stage in the video.
[0112] In step S4 of this implementation, specifically, it includes:
[0113] S401: Construct a design space and mapping scheme from linguistic components to visual cues, such as using color blocks and hollow boxes to describe body parts (muscles, joints); using arrows and lines to highlight the direction and position of the action; displaying the key data points of the action in a text-embedded graphical form; and describing the metaphorical entities of the action with predefined common imagery animations.
[0114] S402: Based on the semantic parsing results, map the entities involved in the explanation to their corresponding visualization forms.
[0115] In step S401, the sources for generating the design space and mapping scheme include visual cues contained in high-usage videos on various platforms, as well as previous research on text-driven visual content generation, sports visualization, data-driven video and animation graphics, and enhanced sports videos. Specific mapping scheme content includes:
[0116] E1 objects are visualized using color blocks or bounding boxes to highlight the locations of muscles and joints obtained through pose estimation and part segmentation.
[0117] E2 actions use joint positions obtained from pose estimation to draw motion trajectories and combine them with other entities to describe them together (e.g., combine them with arrows to describe directional actions).
[0118] In the E3 direction, the joint positions obtained through pose estimation are visualized using arrowhead shapes of different sizes according to the range of motion;
[0119] E4 data allows for visualization of predefined animation formats; for example, various angle graphics can be preset for the angles formed by the body, and scale animations can be preset for the distances formed by movements, with embedded numbers.
[0120] E5 metaphors visualize predefined animation forms by referencing pose estimation and part segmentation results; for example, "holding a ball" is mapped to a spherical animated object; "body like a tree" is mapped to an animated "tree" image.
[0121] In step S5 of this implementation, specifically, it includes:
[0122] S501: Based on the fine-grained recognition results of the actions, the temporal boundaries of each action stage in the video are obtained. Actions in the same stage are defined as a motion-semantic temporal unit, which includes multiple action steps corresponding to each sentence's explanation. The start time of each sentence's explanation is used as the presentation time of the corresponding visual cue visualization, and the temporal positions are arranged accordingly.
[0123] S502: Based on the extracted human body posture, position, and body parts, different visualization methods are laid out in the video, specifically including:
[0124] According to the definition of human topology, there is a four-level pyramid including joints, limbs, half-torso, and full torso. The joint description is directly located to the key point position, and other arguments are arranged accordingly based on the reference position.
[0125] For limb regions consisting of two or more joints, including the upper limbs and lower limbs, the corresponding mapping is performed by region segmentation coverage and skeletal connection.
[0126] For the half-torso and full-torso, considering their large coverage area, in order to avoid excessive visualization effects affecting the movement effect, the main skeletal lines are used for presentation; for example, when describing "upper body upright", the visualization is done by connecting the midpoints of the neck to the two hips.
[0127] For the direction argument, calculate the movement direction based on the keypoint trajectory of the specified joint, and place the arrow at the current time. and a given time threshold range Defined pointer location place, and towards The trailing edge is gradually enhanced with color and transparency gradients to improve its indicativeness;
[0128] For joint angle content in data arguments, the angle graphic data is embedded into the corresponding joint area by combining the joint connection line, and the angle value of the current action is superimposed and displayed; for distance values, they are converted into dashed lines between joints and the corresponding values are displayed.
[0129] For metaphorical arguments, a set of graphical metaphors is used to pre-define some high-frequency metaphor types and match them with corresponding body movements; for example, for the semantic matching description of "holding a ball with both hands", the ball-shaped object is placed between the hands; for "the body stands tall like a tree", the overall direction of the torso is identified and the image of "trees" is placed around the torso.
[0130] like Figure 2 As shown, an exemplary method flow is provided:
[0131] (1) Users upload action instruction videos containing text explanations (oral explanations or text explanations). The video content can be from the Internet or recorded by themselves.
[0132] (2) Determine whether the video contains subtitle files. If not, proceed to step (3). If yes, proceed to step (4).
[0133] (3) The system converts the audio narration in the video into a subtitle file containing text content and timestamps, which users can proofread and modify;
[0134] (4) Named entity recognition and semantic role annotation are performed on the explanatory text. Users can view and interactively modify the language parsing results.
[0135] (5) Perform human pose estimation, part segmentation and recognition and localization on video content. Users can view and modify the extracted motion features;
[0136] (6) Generate the mapped clues and place them in the appropriate spatiotemporal locations of the video;
[0137] (7) Determine whether the special effects need to be modified. If yes, proceed to step (8). If no modification is needed, proceed to step (9).
[0138] (8) Users can edit and adjust the color, size, position, and other effects of the clues, and make interactive modifications and optimizations;
[0139] (9) Determine whether the production process is complete. If yes, proceed to step (10); otherwise, proceed to step (6).
[0140] (10) Export video.
[0141] like Figure 3 The diagram shows the structure of the named entity recognition model in the semantic parsing module. It uses a deep hybrid model, BERT-Bi-LSTM-CRF, to perform named entity recognition on the text, extracting action-related entities (such as joints, muscles, limb parts, etc.). Figure 4 The diagram shows the flowchart of the semantic role annotation method in the semantic parsing module. It further parses the role relationships of actions in the text, such as agent and receiver, through template hints, few-shot demonstrations, and self-verification, to clarify the specific execution method of the action; for example... Figure 5 The image shows a semantic role annotation template in the semantic role annotation method, including task settings, few-shot demonstration, and requirement input; as shown... Figure 6 The diagram shows the action recognition flowchart. It utilizes the BlazePose network for human pose estimation, extracting key skeletal points from the video. Then, the Bodypix model segments body parts for further precise action feature localization. Finally, CNN and Transformer networks perform fine-grained action recognition, outputting the temporal boundaries of each action stage to support the generation of accurate learning cues. Figure 7 As shown, this is the construction space for the enhanced visualization clues, including entity categories, feature categories, and the clue presentation methods corresponding to each entity; for example... Figure 8 As shown, this is the content selection interface. Users can choose from the built-in tutorials, including enhancement videos of different types and difficulty levels for training; they can also upload their own videos, which will be automatically generated and interactively modified to create enhancement exercise instruction videos for training; for example... Figure 9 As shown, to enhance the video content creation interface, visual elements such as graphics and animations are combined to display key information about motion learning. Based on the motion descriptions extracted from the text, graphics such as color blocks, arrows, and trajectories are generated to help learners intuitively understand the execution methods and key points of the motions; for example... Figure 10The image shown is an example of a partially generated instructional video. Through animated prompts, it helps users master the details of the movements, especially in terms of joint angles, direction of movement, and force distribution, thereby deepening learners' understanding of the key points of the movements. Users can see customizable visual cues and instructional prompts on this interface and can interact with the provided learning content.
[0142] Figure 11 This paper presents a limb movement learning cue annotation system based on cross-modal mapping, including:
[0143] The text extraction unit is configured to perform language parsing on the audio explanations in the action instruction video to obtain sentence-by-sentence text and corresponding time markers;
[0144] The semantic parsing unit is configured to: parse the sentence-by-sentence text and extract entities in the sentence-by-sentence text and the roles corresponding to each entity;
[0145] The motion feature extraction unit is configured to extract motion features of the subject demonstrating the motion from the motion instruction video.
[0146] The cross-modal action cue generation unit is configured to map the identified entities to graphics and generate visual cues.
[0147] The spatiotemporal motion cue layout unit is configured to: based on visual cues, language parsing results, and motion features, lay out the visual cues to the corresponding time and space locations to obtain the final limb motion learning cue annotation results.
[0148] It is understood that the aforementioned units can be individually or entirely merged into one or more other units, or some of the units can be further divided into multiple functionally smaller units. This achieves the same operation without affecting the technical effects of the embodiments of this application. The aforementioned units are based on logical functional division. In practical applications, the function of one unit can be implemented by multiple units, or the function of multiple units can be implemented by one unit. In other embodiments of this application, the system may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented collaboratively by multiple units.
[0149] According to another embodiment of this application, the system described in this embodiment can be constructed by running a computer program (including program code) capable of performing the steps involved in the corresponding method of the present invention on a general-purpose computing device, such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the aforementioned computing device through the computer-readable recording medium, and run therein.
[0150] Figure 12 A computer device is shown, comprising a processor, a communication interface, and a computer-readable storage medium. The processor, communication interface, and computer-readable storage medium are connected via a bus or other means.
[0151] The communication interface is used to receive and send data. The computer-readable storage medium can be stored in the memory of the electronic device. The computer-readable storage medium is used to store computer programs, which include program instructions. The processor is used to execute the program instructions stored in the computer-readable storage medium.
[0152] A processor is the computing and control core of an electronic device. It is suitable for implementing one or more instructions, specifically for loading and executing one or more instructions to achieve the corresponding method flow or function.
[0153] The processor is configured to perform the following process:
[0154] The audio narration in the action instruction video is parsed to obtain sentence-by-sentence text and corresponding time markers;
[0155] The sentence-by-sentence text is parsed to extract the entities in the sentence-by-sentence text and the roles corresponding to each entity;
[0156] Extracting the movement characteristics of the demonstrator from instructional videos;
[0157] The identified entities are mapped onto graphics to generate visual cues;
[0158] Based on visual cues, language parsing results, and motion features, the visual cues are placed in their corresponding time and space locations to obtain the final annotation results for limb movement learning cues.
[0159] This invention also provides a computer-readable storage medium, which is a memory device in an electronic device for storing programs and data. It is understood that the computer-readable storage medium here may include both built-in storage media in the electronic device and extended storage media supported by the electronic device. The computer-readable storage medium provides storage space for storing the processing system of the electronic device.
[0160] Furthermore, this storage space also contains one or more instructions suitable for loading and execution by the processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM memory or unstable memory, such as at least one disk storage device; optionally, it can also be at least one computer-readable storage medium located remotely from the aforementioned processor.
[0161] In one embodiment, the computer-readable storage medium stores one or more instructions; the processor loads and executes the one or more instructions stored in the computer-readable storage medium to perform the following process:
[0162] The audio narration in the action instruction video is parsed to obtain sentence-by-sentence text and corresponding time markers;
[0163] The sentence-by-sentence text is parsed to extract the entities in the sentence-by-sentence text and the roles corresponding to each entity;
[0164] Extracting the movement characteristics of the demonstrator from instructional videos;
[0165] The identified entities are mapped onto graphics to generate visual cues;
[0166] Based on visual cues, language parsing results, and motion features, the visual cues are placed in their corresponding time and space locations to obtain the final annotation results for limb movement learning cues.
[0167] The present invention also provides a computer program product or computer program comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the following process:
[0168] The audio narration in the action instruction video is parsed to obtain sentence-by-sentence text and corresponding time markers;
[0169] The sentence-by-sentence text is parsed to extract the entities in the sentence-by-sentence text and the roles corresponding to each entity;
[0170] Extracting the movement characteristics of the demonstrator from instructional videos;
[0171] The identified entities are mapped onto graphics to generate visual cues;
[0172] Based on visual cues, language parsing results, and motion features, the visual cues are placed in their corresponding time and space locations to obtain the final annotation results for limb movement learning cues.
[0173] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0174] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic cable, digital cable) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0175] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for annotating limb movement learning cues based on cross-modal mapping, characterized in that, Includes the following processes: Perform language parsing on the audio narration in the action instruction video, or directly extract the audio narration subtitles from the action instruction video to obtain sentence-by-sentence text and corresponding time markers; The sentence-by-sentence text is parsed to extract the entities in the sentence-by-sentence text and the roles corresponding to each entity; The entities include: objects, actions, directions, data, and metaphors; The object is the agent or the passive argument, the action is the action predicate, the direction is the direction argument, the data is the data argument, and the metaphor is the metaphor argument. Extracting the movement characteristics of the demonstrator from instructional videos; The identified entities are mapped onto graphics to generate visual cues; Based on visual cues, language parsing results, and motion features, the visual cues are placed into corresponding temporal and spatial locations to obtain the final annotation results for limb movement learning cues, including: Based on motion characteristics, the time boundaries of each motion stage in the motion instruction video are obtained. The same stage of motion is defined as a motion-semantic time unit, including multiple motion steps corresponding to each sentence. The start time of each sentence is used as the presentation time of the corresponding visual cue visualization, and the time position is arranged accordingly. Based on the topological structure of the human body, a four-level pyramid is defined, including joints, limbs, half-torso, and full torso. The description of joints is directly located to the key point position, which serves as the reference position for the layout. For limb regions consisting of two or more joints, including the upper and lower limbs, corresponding mapping is performed using region segmentation coverage and skeletal connections. For the half-torso and full-torso, the main skeletal lines are used to represent them; For the direction argument, calculate the movement direction based on the keypoint trajectory of the specified joint, and place the arrow at the current time. and a given time threshold range Defined pointer location place, and towards The trailing edge is gradually enhanced with color and transparency gradients to improve its indicativeness; For joint angle content in data arguments, the angle graphic data is embedded into the corresponding joint area by combining the joint connection line, and the angle value of the current action is superimposed and displayed; for distance values, they are converted into dashed lines between joints and the corresponding values are displayed. For metaphorical arguments, a set of graphical metaphors is used to pre-define some high-frequency metaphor types and match them with corresponding body movements.
2. The method for annotating limb movement learning cues based on cross-modal mapping as described in claim 1, characterized in that, Perform language analysis on the audio narration in the instructional videos, including: A speech-to-text neural network is used to extract sentence-by-sentence text from the audio narration of action instruction videos, and a timetable corresponding to the start and end times of each sentence-by-sentence text is obtained. The video actions corresponding to each start and end time period are segmented to generate text subtitles.
3. The method for annotating limb movement learning cues based on cross-modal mapping as described in claim 1, characterized in that, The sentence-by-sentence text is parsed to extract entities and their corresponding roles, including: The BERT structure is used to generate text representations of input words in sentence-by-sentence text. Sub-words, sentences, and position embeddings of sentences are generated through word segmentation and high-dimensional mapping. The Transformer structure is used for bidirectional two-stage processing. In the first stage, large-scale pre-training is performed on a pre-set encyclopedia corpus. Feature extraction is performed by occluding language and sentence prediction tasks while taking into account the context. The second stage of fine-tuning is performed using supervised fine-tuning to obtain multi-dimensional vector representations of sentence-by-sentence text. Based on the multidimensional vector representation of sentence-by-sentence text, the key contextual features of entities are extracted using a bidirectional long short-term memory network. The output of the bidirectional long short-term memory network is post-processed using a conditional random field structure, and the sequence is labeled using a global optimal solution.
4. The method for annotating limb movement learning cues based on cross-modal mapping as described in claim 1, characterized in that, Extract the movement characteristics of the subject demonstrating the movement from the instructional video, including: The BlazePose network is used for human pose estimation, to identify the key skeletal coordinates of the demonstration subject in the video, and to repair misaligned and missing points and smooth and denoise the data to obtain the skeletal joint coordinates. The Bodypix human semantic segmentation model is used to identify and segment various body parts, and the recognition edges are optimized and smoothed. A fusion model based on CNN and Transformer is used to identify and locate the demonstrated movements in the action instruction video, and to identify the subcategories and time positions of the movements in the action instruction video.
5. The method for annotating limb movement learning cues based on cross-modal mapping as described in claim 4, characterized in that, For object entities, the positions obtained by pose estimation and part segmentation are visualized using color blocks or bounding boxes to highlight the locations of muscles and joints. For motion entities, motion trajectories are drawn using joint positions obtained from pose estimation and combined with other entities for joint description. For directional entities, the joint positions obtained by pose estimation are visualized using arrowhead shapes of different sizes according to the range of motion. For data entities, visualize predefined animation formats; For metaphorical entities, predefined animation forms are visualized with reference to pose estimation and part segmentation results.
6. A limb movement learning cue annotation system based on cross-modal mapping, characterized in that, include: The text extraction unit is configured to: perform language parsing on the audio narration in the action instruction video, or directly extract the audio narration subtitles in the action instruction video to obtain sentence-by-sentence text and corresponding time markers; The semantic parsing unit is configured to: parse the sentence-by-sentence text and extract entities in the sentence-by-sentence text and the roles corresponding to each entity; The entities include: objects, actions, directions, data, and metaphors; The object is the agent or the passive argument, the action is the action predicate, the direction is the direction argument, the data is the data argument, and the metaphor is the metaphor argument. The motion feature extraction unit is configured to extract motion features of the subject demonstrating the motion from the motion instruction video. The cross-modal action cue generation unit is configured to map the identified entities to graphics and generate visual cues. The spatiotemporal motion cue layout unit is configured to: based on visual cues, language parsing results, and motion features, place visual cues into corresponding temporal and spatial locations to obtain the final limb motion learning cue annotation results, including: Based on motion characteristics, the time boundaries of each motion stage in the motion instruction video are obtained. The same stage of motion is defined as a motion-semantic time unit, including multiple motion steps corresponding to each sentence. The start time of each sentence is used as the presentation time of the corresponding visual cue visualization, and the time position is arranged accordingly. Based on the topological structure of the human body, a four-level pyramid is defined, including joints, limbs, half-torso, and full torso. The description of joints is directly located to the key point position, which serves as the reference position for the layout. For limb regions consisting of two or more joints, including the upper and lower limbs, corresponding mapping is performed using region segmentation coverage and skeletal connections. For the half-torso and full-torso, the main skeletal lines are used to represent them; For the direction argument, calculate the movement direction based on the keypoint trajectory of the specified joint, and place the arrow at the current time. and a given time threshold range Defined pointer location place, and towards The trailing edge is gradually enhanced with color and transparency gradients to improve its indicativeness; For joint angle content in data arguments, the angle graphic data is embedded into the corresponding joint area by combining the joint connection line, and the angle value of the current action is superimposed and displayed; for distance values, they are converted into dashed lines between joints and the corresponding values are displayed. For metaphorical arguments, a set of graphical metaphors is used to pre-define some high-frequency metaphor types and match them with corresponding body movements.
7. A computer device, characterized in that, include: Processor and computer-readable storage media; A processor, adapted to execute computer programs; A computer-readable storage medium storing a computer program, which, when executed by the processor, implements the limb movement learning cue annotation method based on cross-modal mapping as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded by a processor and executed as described in any one of claims 1 to 5, which is a method for annotating limb movement learning cues based on cross-modal mapping.