Teaching interaction-oriented digital human three-dimensional reconstruction system

By generating fused feature vectors through multimodal sensor acquisition and dynamic weight distribution algorithms, combined with pre-trained model matching role models and interaction rules, the problem of insufficient information capture by digital humans in teaching interactions is solved, personalized and efficient teaching interaction responses are achieved, and the practicality and immersion of the teaching experience are improved.

CN120599155AActive Publication Date: 2025-09-05XIAMEN YIXUE SOFTWARE CO LTD

Patent Information

Application Number
CN202511094018.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-05
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Existing digital humans find it difficult to fully capture multimodal information during teaching interactions, resulting in poor interaction effects, insufficient flexibility in response strategies, and difficulty in providing personalized and precise teaching experiences.

Method used

Voice, gesture, eye and facial data are collected through a multimodal sensor array, and the teaching scene parameters are obtained by combining with the preset database. The multimodal fusion feature vector is generated using a dynamic weight distribution algorithm, and the role model and interaction rules are dynamically matched to achieve directional interactive response. The system parameters are adjusted through reinforcement learning to form an adaptive optimization loop.

Benefits of technology

It improves the adaptability and effectiveness of teaching interaction, can provide personalized responses according to different teaching scenarios, enhance the practicality and immersion of the teaching experience, and continuously optimize interactive performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599155A_ABST
    Figure CN120599155A_ABST
Patent Text Reader

Abstract

The invention provides a teaching interaction-oriented digital human three-dimensional reconstruction system, and relates to the technical field of human-computer interaction, and the system comprises a data collection module which is used for collecting a voice signal, gesture action data, eye tracking data and facial expression data of a user in real time through a multi-modal sensor array, and obtaining a teaching type identifier, an interaction stage state parameter and a participant role identifier of the current teaching scene through a preset teaching scene database. According to the invention, accurate adaptation of teaching scenes and intelligent generation and dynamic optimization of interaction strategies are realized, the accuracy and suitability of teaching interaction are improved, and personalized and high-quality teaching interaction experience is brought to users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer interaction, and in particular to a digital human three-dimensional reconstruction system for teaching interaction. Background Art

[0002] In traditional teaching models, teacher-student interaction is often limited by time and space, and comprehensive feedback is difficult to provide to every student, hindering teaching effectiveness. While the application of digital humans in teaching scenarios has achieved some success, there is still room for improvement in the depth and breadth of teaching interactions. Some digital human systems have limited information processing capabilities, primarily focusing on processing voice or visual information. Their ability to comprehensively process and analyze multimodal information, including the user's voice signals, gestures, eye tracking data, and facial expressions, is relatively weak.

[0003] For example, when explaining complex mathematical formulas, students may convey their confusion through puzzled eyes, slightly frowned brows, and hesitant questions. However, due to deficiencies in multimodal information fusion and analysis, these digital humans are prone to being unable to accurately capture students' true intentions and emotional states, resulting in poor interaction effects.

[0004] In addition, the flexibility of the response strategies of many digital humans needs to be strengthened, and it is difficult to dynamically adjust the interaction methods according to the characteristics of different teaching scenarios and interaction stages. This makes it difficult for users to obtain personalized and precise experiences during the interaction process, limiting the application potential of digital humans in teaching interaction scenarios. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a digital human three-dimensional reconstruction system for teaching interaction, thereby improving the quality of teaching interaction.

[0006] In order to solve the above technical problems, the technical solutions of the present invention are as follows: In the first aspect, a digital human 3D reconstruction system for teaching interaction includes: The data acquisition module is used to collect the user's voice signals, gesture data, eye tracking data and facial expression data in real time through a multimodal sensor array, and obtain the teaching type identifier, interaction stage state parameters and participant role identifier of the current teaching scene through a preset teaching scene database; The multimodal fusion module is used to convert speech signals into text, extract three-dimensional spatial trajectory features from gesture data, calculate the spatial coordinates of the gaze focus and encode expression features from eye tracking data and facial expression data, and generate a multimodal fusion feature vector through a dynamic weight allocation algorithm. The strategy matching module is used to input the multimodal fusion feature vector into the pre-trained teaching scenario classification model, dynamically match the corresponding role model and the associated interaction rule set from the three-dimensional role template library according to the output teaching type identifier, and obtain the response strategy matched at the current stage; An interactive response module, used to implement targeted interactive responses based on the corresponding role model and associated interaction rule set; The optimization loop module is used to capture the user's operation delay data and voice feedback accuracy in real time through feedback sensors, dynamically adjust the dynamic weight distribution algorithm parameters through reinforcement learning agents, and update the interaction rule set and device control protocol of the three-dimensional character template to form an adaptive optimization loop.

[0007] Furthermore, speech signals are converted into text, gesture data are subjected to three-dimensional trajectory feature extraction, eye tracking data and facial expression data are subjected to gaze focus spatial coordinate calculation and expression feature encoding, and a multimodal fusion feature vector is generated through a dynamic weight allocation algorithm, including: The speech signal is subjected to noise suppression and speech segmentation processing, the time-frequency features of the speech signal are extracted, and a semantic text sequence is directly generated through an end-to-end speech recognition network. At the same time, the emotional tendency intensity value is calculated based on the speech spectrum energy distribution; the gesture action data is calibrated for the joint point spatial coordinates, and a smooth three-dimensional spatial trajectory is generated based on the calibration results using the cubic spline interpolation algorithm, and the direction vector, curvature characteristics and instantaneous velocity parameters of the trajectory are extracted; the pupil's gaze focus spatial coordinates in the three-dimensional teaching scene are calculated based on the eye tracking data, and the core area and diffusion range of the user's attention distribution are determined through the density clustering algorithm; the facial expression data is subjected to facial muscle group movement amplitude detection, and combined with the predefined expression coding rule library, expression category labels and emotional intensity level parameters are generated; According to the teaching type identifier and the interaction stage state parameters, the semantic text sequence and emotional tendency intensity value, the extracted trajectory curvature features, the attention core area and the emotional intensity level parameters are input into the dynamic weight allocator, and a multimodal fusion feature vector is generated through weighted splicing and feature dimensionality reduction.

[0008] Furthermore, the speech signal is subjected to noise suppression and speech segmentation processing, the time-frequency features of the speech signal are extracted, and a semantic text sequence is directly generated through the end-to-end speech recognition network. At the same time, the emotional tendency intensity value is calculated based on the speech spectrum energy distribution, including: The adaptive filtering algorithm is used to suppress background noise in speech signals, and the speech activity detection algorithm is used to segment effective speech segments and extract the time-frequency spectrum characteristics and energy distribution parameters of each speech segment. Input the time-frequency spectrum features into the end-to-end speech recognition network, output the semantic text sequence corresponding to the speech segment, and extract the recognition probability value of each word based on the word-level probability distribution output by the network; Based on the word-level probability values, the average probability value of all words in the text sequence is calculated as the basic confidence score, and the score is dynamically calibrated in combination with the energy distribution parameters of the speech segment to generate a global confidence score; Based on the predefined emotional keyword library, keyword matching is performed on the semantic text sequence, and the frequency and contextual association weight of emotional keywords are counted; Based on the speech energy distribution parameters, the high-frequency energy proportion and intonation fluctuation amplitude of the speech segment are analyzed, and the emotional tendency intensity value is generated by combining the frequency of emotional keywords and context-related weights.

[0009] Furthermore, the multimodal fusion feature vector is input into the pre-trained teaching scenario classification model. According to the output teaching type identifier, the corresponding role model and the associated interaction rule set are dynamically matched from the 3D role template library to obtain the response strategy matched at the current stage, including: A teaching scenario classification model is trained based on a multimodal training dataset containing multimodal feature vectors of speech text sequences, gesture trajectory features, attention core areas, and emotion intensity parameters, as well as corresponding teaching type labels. A probability distribution of teaching type identifiers is generated through a cross-modal feature fusion layer and a classifier. The multimodal fusion feature vector is input into the teaching scene classification model, the joint weight of each modal feature is calculated through the cross-modal feature fusion layer, the probability distribution of the current teaching type identifier is output, and the scene classification result is obtained; According to the teaching type identifier, the corresponding role model and the associated interaction rule set are matched from the three-dimensional role template library; Based on the interaction stage state parameters, the response strategy matching the current stage is dynamically extracted from the associated interaction rule set, including voice response delay compensation parameters, gesture action synchronization threshold and gaze focus rendering priority.

[0010] Furthermore, according to the teaching type identifier, a corresponding role model and an associated interaction rule set are matched from the three-dimensional role template library, including: When the teaching type identifier is classroom teaching module, the virtual teacher behavior model and the associated semantic guidance rule set, knowledge point association mapping table and visual dynamic demonstration strategy are extracted; When the teaching type identifier is the off-class practice team module, the student collaboration logic model and task priority allocation rules, abnormal operation detection protocol and collaboration status feedback mechanism are extracted; When the teaching type identifier is the enterprise practice module, the enterprise tutor scenario model and equipment operation instruction mapping table, process optimization constraints and production environment real-time data interface are extracted.

[0011] Furthermore, based on the corresponding role model and associated interaction rule set, a targeted interactive response is achieved, including: Generate behavioral logic instructions corresponding to the teaching type identifier based on the role model and the associated interaction rule set; The behavioral logic instructions are input into the three-dimensional virtual image parameterized driving engine, which drives the interactive response of the digital human by converting the guided question-and-answer instructions into playable voice waveform data of the speech synthesis model, converting the visual demonstration action sequence into skeletal driving parameters and corresponding limb motion trajectories, and mapping the device control signals into the device operation instruction flow of the virtual production environment.

[0012] Furthermore, based on the role model and the associated interaction rule set, behavioral logic instructions that strictly correspond to the teaching type identifier are generated, including: When the role model is a virtual teacher behavior model, guided question-answering instructions and visual demonstration action sequences associated with knowledge points are generated based on the semantic guidance rule set; When the role model is a student collaboration logic model, collaborative task assignment instructions and abnormal operation detection signals are generated based on the task priority assignment protocol; When the role model is an enterprise mentor scenario model, device control signals and process optimization parameter adjustment instructions are generated based on the device operation instruction mapping relationship.

[0013] Furthermore, feedback sensors are used to capture the user's interactive response latency data and voice feedback accuracy in real time. A reinforcement learning agent is used to dynamically adjust the dynamic weight distribution algorithm parameters and update the interaction rule set and device control protocol of the 3D character template, forming an adaptive optimization loop, including: The user's operational data on interactive responses is captured in real time through feedback sensors, including the response delay time and semantic matching accuracy of user voice input in voice interaction scenarios, the synchronization delay difference between user actions and digital human skeleton drive parameters in gesture interaction scenarios, and the offset between the user's gaze area and the rendering focus in visual focus scenarios. Quantify the difference between the interactive response operation data and the expected response target to generate optimization indicators, including voice response delay compensation deviation value, gesture synchronization error coefficient and visual focus deviation; Based on the optimization indicators, the weight distribution ratio of the multimodal fusion feature vector, the voice response delay compensation parameter, the gesture action synchronization threshold, and the associated interaction rule set are adjusted through the reinforcement learning agent to obtain the adjusted parameters and rule set; The adjusted parameters and rule sets are synchronized to the 3D character template library and multimodal fusion module, forming a closed-loop adaptive cycle from data acquisition to parameter optimization.

[0014] In a second aspect, a computing device includes: one or more processors; The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the system.

[0015] According to a third aspect, a computer-readable storage medium stores a program, which implements the system when executed by a processor.

[0016] The above solution of the present invention includes at least the following beneficial effects: Utilizing a multimodal sensor array, it comprehensively captures multi-dimensional data such as voice, gestures, eyes, and face. It also integrates a preset database to obtain key parameters of the teaching scenario, ensuring accurate perception of the user's status and teaching environment. Targeted processing of data from different modalities is performed, and a fused feature vector is generated through a dynamic weight allocation algorithm. This effectively integrates information from each modality, enabling a deep understanding of user intent and emotion, laying the foundation for precise interaction. Using a pre-trained teaching scenario classification model, the teaching scenario is judged based on the multimodal fusion feature vector. Appropriate character models and interaction rule sets are matched from a three-dimensional character template library, ensuring that the digital human can respond appropriately in different teaching scenarios, improving the adaptability and effectiveness of teaching interactions.

[0017] Based on a matching role model and rule set, targeted interactive responses are achieved. Whether it's a virtual teacher explaining knowledge, assigning collaborative tasks to students, or providing device operation guidance from an enterprise mentor, this provides users with an interactive experience tailored to their needs, enhancing the practicality of interactive teaching. By capturing user feedback data in real time and leveraging reinforcement learning agents to dynamically adjust system parameters and rule sets, this creates an adaptive optimization loop that continuously improves interactive performance and ensures the long-term efficiency and quality of interactive teaching. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a schematic diagram of a digital human three-dimensional reconstruction system for teaching interaction provided by an embodiment of the present invention.

[0019] Figure 2 This is a flowchart provided by an embodiment of the present invention for performing noise suppression and speech segmentation processing on a speech signal, extracting the time-frequency features of the speech signal, directly generating a semantic text sequence through an end-to-end speech recognition network, and calculating the emotional tendency intensity value based on the speech spectrum energy distribution. DETAILED DESCRIPTION

[0020] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0021] like Figure 1 As shown, an embodiment of the present invention provides a digital human 3D reconstruction system for teaching interaction, comprising: The data acquisition module is used to collect the user's voice signals, gesture data, eye tracking data and facial expression data in real time through a multimodal sensor array, and obtain the teaching type identifier, interaction stage state parameters and participant role identifier of the current teaching scene through a preset teaching scene database; The multimodal fusion module is used to convert speech signals into text, extract three-dimensional spatial trajectory features from gesture data, calculate the spatial coordinates of the gaze focus and encode expression features from eye tracking data and facial expression data, and generate a multimodal fusion feature vector through a dynamic weight allocation algorithm. The strategy matching module is used to input the multimodal fusion feature vector into the pre-trained teaching scenario classification model, dynamically match the corresponding role model and the associated interaction rule set from the three-dimensional role template library according to the output teaching type identifier, and obtain the response strategy matched at the current stage; An interactive response module, used to implement targeted interactive responses based on the corresponding role model and associated interaction rule set; The optimization loop module is used to capture the user's operation delay data and voice feedback accuracy in real time through feedback sensors, dynamically adjust the dynamic weight distribution algorithm parameters through reinforcement learning agents, and update the interaction rule set and device control protocol of the three-dimensional character template to form an adaptive optimization loop.

[0022] In an embodiment of the present invention, deep feature extraction and encoding are performed on a variety of data, and a fusion feature vector is generated through dynamic weight allocation, which improves the accurate understanding of user intentions and emotional states and enhances the intelligence of system interaction. By using pre-trained models and role template libraries, intelligent classification of teaching scenarios and dynamic matching of response strategies are achieved, and personalized services can be provided according to different teaching types and interaction stages. Targeted interactive responses are achieved based on matching role models and rule sets, supporting diverse forms of teaching interactions, such as voice questions and answers, action demonstrations, device control, etc., to enhance the immersiveness of the teaching experience. User feedback data is captured in real time through feedback sensors, and reinforcement learning is used to dynamically adjust system parameters and rules to form a closed-loop optimization mechanism to continuously improve system performance and teaching effectiveness.

[0023] In a preferred embodiment of the present invention, speech-to-text conversion is performed on speech signals, three-dimensional trajectory feature extraction is performed on gesture action data, gaze focus spatial coordinates are calculated and expression feature encoding is performed on eye tracking data and facial expression data, and a multimodal fusion feature vector is generated through a dynamic weight allocation algorithm, which may include: The speech signal is subjected to noise suppression and speech segmentation processing, the time-frequency features of the speech signal are extracted, and a semantic text sequence is directly generated through an end-to-end speech recognition network, and the emotional tendency intensity value is calculated based on the speech spectrum energy distribution; the joint point spatial coordinates of the gesture action data are calibrated, and a smooth three-dimensional spatial trajectory is generated based on the calibration results using the cubic spline interpolation algorithm, and the direction vector, curvature characteristics and instantaneous velocity parameters of the trajectory are extracted; the pupil's gaze focus spatial coordinates in the three-dimensional teaching scene are calculated based on the eye tracking data, and the core area and diffusion range of the user's attention distribution are determined using the density clustering algorithm; the facial expression data is subjected to facial muscle group movement amplitude detection, and combined with the predefined expression encoding rule library, expression category labels and emotional intensity level parameters are generated; among them, the speech signal is subjected to noise suppression and speech segmentation processing, the time-frequency features of the speech signal are extracted, and a semantic text sequence is directly generated through an end-to-end speech recognition network, and the emotional tendency intensity value is calculated based on the speech spectrum energy distribution. Specifically, it includes: The adaptive filtering algorithm is used to suppress background noise in speech signals, and the speech activity detection algorithm is used to segment effective speech segments and extract the time-frequency spectrum characteristics and energy distribution parameters of each speech segment. Input the time-frequency spectrum features into the end-to-end speech recognition network, output the semantic text sequence corresponding to the speech segment, and extract the recognition probability value of each word based on the word-level probability distribution output by the network; Based on the word-level probability values, the average probability value of all words in the text sequence is calculated as the basic confidence score, and the score is dynamically calibrated in combination with the energy distribution parameters of the speech segment to generate a global confidence score; Based on the predefined emotional keyword library, keyword matching is performed on the semantic text sequence, and the frequency and contextual association weight of emotional keywords are counted; Based on the speech energy distribution parameters, the high-frequency energy ratio and intonation fluctuation amplitude of the speech segment are analyzed, and the emotional tendency intensity value is generated by combining the frequency of emotional keywords and the contextual association weight; According to the teaching type identifier and the interaction stage state parameters, the semantic text sequence and emotional tendency intensity value, the extracted trajectory curvature features, the attention core area and the emotional intensity level parameters are input into the dynamic weight allocator, and a multimodal fusion feature vector is generated through weighted splicing and feature dimensionality reduction.

[0024] In this embodiment, within the first 1-2 seconds before a user speaks, ambient background signals (such as the sound of a classroom fan or the rubbing of desks and chairs) are collected, their energy distribution and frequency characteristics are calculated, and a background noise model is established. When the user begins speaking, an adaptive filtering algorithm compares the input signal with the noise model in real time: If the background noise is stable (such as a continuous low-frequency fan sound), the algorithm will slowly filter the noise with a smaller step size (0.01-0.03) to avoid accidentally deleting voice details; If high-frequency noise suddenly appears (such as the sound of a door closing), the algorithm immediately switches to a large step size (0.05-0.08) to quickly attenuate the noise frequency band.

[0025] Cut the continuous speech into short frames of 20-30 milliseconds, and calculate two indicators for each frame: Short-term energy: Calculates the sum of the squares of the signal amplitudes within a frame. If it is 1.5-3 times greater than the background noise energy (for example, if the background noise energy is 100, the speech energy must be 150-300), it is considered "likely to contain speech." Zero-crossing rate: Counts the number of positive and negative alternations in the signal within a frame. If it is within the range of 10-30 times / frame (which is consistent with human speech characteristics), it is determined to be a valid speech frame.

[0026] Only frames that meet both conditions will be merged into valid speech segments, and the rest will be treated as noise and removed.

[0027] Speech signal processing-time-frequency feature extraction: When performing frequency domain conversion on windowed speech segments, not only is it necessary to analyze the intensity of each frequency component, but it is also necessary to simulate the auditory characteristics of the human ear. The human ear is more sensitive to frequency changes in low-frequency sounds (such as bass) and less able to distinguish high-frequency sounds (such as whistles). For example, the human ear can easily distinguish between sounds at 100Hz and 200Hz, but has difficulty distinguishing between 4000Hz and 4100Hz. Mel-frequency cepstral coefficients (MFCCs) are feature parameters designed based on this characteristic. The specific analysis process is as follows: Convert linear frequency to Mel frequency: Convert the actual frequency of the speech signal (unit: Hz) to Mel frequency (unit: Mel). The conversion relationship is similar to "logarithmic scaling": The Mel frequency in the low frequency band (such as 0-1000Hz) is close to the linear relationship with the actual frequency, ensuring that low-frequency details are not lost; The Mel frequency in high frequency bands (such as greater than 1000Hz) grows slower, compressing high-frequency redundant information.

[0028] For example: The actual frequency of 100Hz corresponds to about 200Mel. The actual frequency of 1000Hz corresponds to about 1300Mel. The actual frequency of 4000Hz corresponds to about 2000Mel (only an increase of 700Mel, rather than a non-linear increase of 3000Hz).

[0029] Mel filter bank filtering: After conversion to Mel frequencies, a set of 20-40 triangular filters is used to cover the entire Mel frequency range. Each filter represents an "auditory perception unit" and is used to measure the total energy within that frequency range. The low-frequency filters are spaced narrowly (e.g., 100 Mel intervals) to ensure that fundamental frequency differences are captured (e.g., the fundamental frequencies of male and female voices are 150-300 Hz and 80-150 Hz, respectively, corresponding to Mel frequencies of approximately 300-600 Mel and 180-300 Mel, respectively, which are distinguished by narrowband filters). The high-frequency filters are spaced widely (e.g., 200 Mel intervals) to merge similar high-frequency components (e.g., 4000 Hz and 4500 Hz are processed by the same filter). Each filter outputs the energy value of that frequency band, forming a Mel spectrum.

[0030] Discrete Cosine Transform (DCT) Dimensionality Reduction: The Mel spectrum contains a lot of redundant information (such as the high correlation between the energy values ​​of adjacent filters). The most representative features are extracted through discrete cosine transform (DCT) to generate cepstral coefficients (take the 2nd to 13th coefficients, a total of 12 dimensions). These coefficients reflect the overall shape and trend of the Mel spectrum, for example: The low-order coefficients (such as the 2nd and 3rd coefficients) correspond to the low-frequency energy distribution of speech (related to voiced sounds and fundamental frequency); Higher-order coefficients (such as the 12th and 13th coefficients) correspond to high-frequency energy distribution (related to voiceless sounds and fricative sounds).

[0031] In the final MFCC feature vector, each value represents the intensity of the speech signal at a specific "auditory perception dimension": For female speech (higher fundamental frequency), the low-order MFCC coefficient values ​​are larger (reflecting a higher proportion of high-frequency energy); For male speech (low fundamental frequency), the low-order MFCC coefficient values ​​are smaller, but the energy distribution in the mid- and low-frequency bands is more concentrated.

[0032] For example, when a user says "hello", the third coefficient of the MFCC feature for women is 0.8 and the fifth coefficient is 0.6; the third coefficient of the male is 0.5 and the fifth coefficient is 0.4. These differences are used to distinguish speech features.

[0033] Speech Signal Processing - Speech Recognition and Confidence Calculation: The extracted MFCC features are fed into an end-to-end speech recognition network (such as a pre-trained DeepSpeech model). The network uses a multi-layer neural network to analyze the speech and predict the corresponding text frame by frame. For example, when inputting the speech features of "hello," the network first identifies phonemes such as "h," "e," "l," "l," and "o," then combines them into the word "hello" and outputs the recognition probability of each character (for example, the probability of "h" is 0.92 and the probability of "e" is 0.88).

[0034] The confidence calculation is divided into two steps: Basic score: Calculate the average recognition probability of all words. For example, if the average probability of "hello" is 0.90, it is judged as high confidence (greater than 0.8). If the average probability of a sentence is 0.7, it is judged as medium confidence (0.6-0.8).

[0035] Energy calibration: If the energy of a speech segment is higher than 1.2 times the average energy (e.g., a loud question), the base score is increased by 0.05. If the energy is lower than 0.8 times the average energy (e.g., a whisper), the score is decreased by 0.05. The final global confidence score is used to determine the reliability of the recognition result.

[0036] Speech signal processing-emotional tendency calculation: Load the sentiment keyword library (containing over 500 sentiment words, such as "awesome" and "so difficult") and perform word-by-word matching on the identified text. For example, in the text "This formula is too difficult to understand," "difficult" is marked as a negative keyword, and its weight is adjusted based on its position: If the keyword is at the beginning or end of the sentence (e.g., "too difficult", "I can't"), the weight is 0.8-1.0 (strong sentiment association); If it is located in a sentence (such as "This question is a bit difficult"), the weight is 0.3-0.5 (weak sentiment association).

[0037] The ratio of energy in the 3-8kHz frequency band to total energy is calculated. If it exceeds 0.4 (e.g., rising intonation during excitement), it is considered strong emotion; if it is below 0.2 (e.g., calm narration), it is considered weak emotion. The standard deviation of pitch variation is used to measure this. If the standard deviation is greater than 50Hz (e.g., large intonation fluctuations in interrogative sentences), it is considered high volatility; if it is less than 20Hz (e.g., steady intonation in declarative sentences), it is considered low volatility. Ultimately, the emotional intensity is determined by the keyword weight, the proportion of high-frequency energy, and intonation fluctuation. For example, "It's too difficult!" (high keyword weight + high high-frequency energy + high intonation fluctuation) is considered strong negative emotion.

[0038] Gesture action data processing-joint point calibration: When a user makes a gesture, the RGB-D camera synchronously captures the 2D joint coordinates (such as the (x, y) position of the index fingertip in the image), combines it with the camera's depth information (such as the distance measured by the ToF sensor), and calculates the 3D coordinates (x, y, z) of the joint through triangulation.

[0039] Gesture action data processing - trajectory generation and feature extraction: If the original joint coordinates are jittery (such as from slight hand tremors), a cubic spline interpolation algorithm is used to generate a smooth trajectory: multiple virtual points are inserted between two adjacent real coordinate points to ensure a continuous trajectory curve and smooth derivatives. For example, if the hand moves from point A (1, 2, 3) to point B (4, 5, 6), the interpolation algorithm will generate points (2, 3, 4), (3, 4, 5), and so on in the middle, forming a smooth motion path.

[0040] Three features are extracted from the smooth trajectory: Direction vector: Calculates the coordinate difference between two adjacent points (e.g., the vector from A to B is (3, 3, 3)), reflecting the direction of the gesture (e.g., swiping to the upper right); Curvature feature: The gesture type is determined by the curvature of the trajectory. For example, the curvature of a circle is always constant, while the curvature of a straight line is 0. Instantaneous speed: Calculates the displacement distance per unit time. If the speed exceeds 1.5 m / s (such as waving a hand quickly), it is considered an abnormal action and may require resampling.

[0041] Eye tracking data processing - gaze point calculation: Eye tracking devices (such as infrared cameras) record the positions of the pupil center and corneal reflection point in real time, and calculate the gaze point through a geometric mapping model: Assuming that the pupil center is P and the corneal reflection point is R, the path of light from the light source to R and then to P can determine the direction of vision; the direction of vision is intersected with the three-dimensional teaching scene model (such as a virtual blackboard, experimental equipment) to obtain the coordinates of the gaze point (x, y, z).

[0042] Eye tracking data processing-attention area analysis: Collect all fixations within a period of time (e.g. 5 seconds) and analyze them using a density clustering algorithm: Set a neighborhood radius of 5-10 cm (equivalent to a small area on the screen) and a minimum sample number of 5-10 points. If the number of fixations in a certain area reaches the minimum sample number and the distance between points is less than the neighborhood radius, it is determined to be the core attention area (such as the location of the formula explained by the teacher). The discrete points outside the core area constitute the diffusion range, reflecting the content of the user's peripheral vision (such as the teaching aids at the edge of the podium).

[0043] Facial expression data processing-muscle movement detection: By detecting facial key points (such as the inner corner of the left eye and the right corner of the mouth), the displacement of each key point is calculated: If the eyebrow keypoint moves upwards by more than 3 mm, it's considered a "raised eyebrow" (possibly indicating confusion); if the mouth corner keypoint moves sideways by more than 2 mm, it's considered a "smile" (possibly indicating understanding). All displacement values ​​are normalized to a range of 0-1 (0 represents no movement, 1 represents maximum detectable movement, such as a full frown).

[0044] Facial expression data processing - expression classification and intensity calculation: The normalized muscle movement data is input into an expression classification model (such as a FACS-based random forest model), which outputs the expression category (such as "confused" or "happy") and the confidence level: If the confidence level is higher than 0.7, it is considered a valid expression (e.g., the confidence level of “confused” is 0.85); If it is lower than 0.7, ignore the expression to avoid misjudgment (for example, a slight facial twitch may be mistakenly identified as "surprise").

[0045] Emotional intensity is calculated based on the weighted sum of key muscles, such as: Confused expression: The weight of the eyebrows raised is 0.6, and the weight of the corners of the eyes drooping is 0.4. The total of 0.5 indicates moderate confusion, and 0.8 indicates severe confusion.

[0046] Dynamic weight allocation and feature fusion: Assign weights to each modality based on the teaching scenario and stage: Classroom theoretical teaching: Voice (explaining knowledge): 0.5 (highest weight); gesture (blackboard demonstration): 0.2; gaze (student listening status): 0.2; expression (student understanding level): 0.1.

[0047] Experimental operation teaching: gesture (equipment operation): 0.5; voice (step instructions): 0.3; gaze (key parts of the equipment): 0.15; expression (operation proficiency): 0.05.

[0048] Adjust weights during the interaction phase: During the introductory phase at the beginning of the course, the weight of speech is increased by 0.1 (such as when the teacher introduces the title of a new course), and the weight of other modalities is reduced by 0.03 each. During the interactive phase of student discussion, the weight of facial expressions and gaze is increased by 0.1 each (focusing on students' communication emotions), and the weight of speech is reduced by 0.2.

[0049] Finally, the features of each modality (such as speech text, gesture curvature, gaze area, and expression intensity) are multiplied by weight and concatenated to form a long feature vector. Principal component analysis (PCA) is then used to remove redundant information and generate the final multimodal fusion feature vector.

[0050] By meticulously processing data such as voice, gestures, eye movements, and facial expressions, the system can deeply analyze users' intentions, emotions, and attention during teaching interactions. For example, it can accurately determine whether a student is frowning in confusion during a complex explanation, perhaps because they don't understand, or simply looking away due to lack of interest, thereby providing more targeted feedback. A dynamic weighting mechanism allows the digital human to flexibly adjust its focus on each modality based on different teaching scenarios and stages. During experimental instruction, it can quickly focus on students' gestures and promptly correct errors. In group discussions, it can pay more attention to students' expressions and eye contact to foster a positive and interactive atmosphere. Accurate emotion and attention analysis can promptly identify problems in students' learning, such as fatigue and confusion. The digital human can then adjust the teaching pace, add interesting content, or re-examine difficult points, effectively improving students' learning efficiency and knowledge mastery. This meticulous data processing ensures a more natural and fluid interaction between the digital human and the user. Smooth gesture recognition and precise gaze capture allow users to feel like they are communicating with real teachers or classmates when interacting with digital humans, enhancing the immersiveness of teaching interaction and user satisfaction.

[0051] In a preferred embodiment of the present invention, the multimodal fusion feature vector is input into a pre-trained teaching scene classification model, and the corresponding role model and associated interaction rule set are dynamically matched from the three-dimensional role template library according to the output teaching type identifier to obtain the response strategy matched at the current stage, which may include: A teaching scenario classification model is trained based on a multimodal training dataset containing multimodal feature vectors of speech text sequences, gesture trajectory features, attention core areas, and emotion intensity parameters, as well as corresponding teaching type labels. A probability distribution of teaching type identifiers is generated through a cross-modal feature fusion layer and a classifier. The multimodal fusion feature vector is input into the teaching scene classification model, the joint weight of each modal feature is calculated through the cross-modal feature fusion layer, the probability distribution of the current teaching type identifier is output, and the scene classification result is obtained; According to the teaching type identifier, the corresponding role model and associated interaction rule set are matched from the 3D role template library, including: When the teaching type identifier is classroom teaching module, the virtual teacher behavior model and the associated semantic guidance rule set, knowledge point association mapping table and visual dynamic demonstration strategy are extracted; When the teaching type identifier is the off-class practice team module, the student collaboration logic model and task priority allocation rules, abnormal operation detection protocol and collaboration status feedback mechanism are extracted; When the teaching type identifier is the enterprise practice module, the enterprise instructor scenario model and equipment operation instruction mapping table, process optimization constraints and production environment real-time data interface are extracted; Based on the state parameters of the interaction stage, dynamically extract the response strategies that match the current stage from the associated interaction rule set, including voice response delay compensation parameters, gesture action synchronization thresholds, and line-of-sight focus rendering priorities.

[0052] In the embodiments of the present invention, the training process of the teaching scenario classification model is as follows: In a real teaching environment, collect voice signals through a microphone array, capture gesture actions with a depth camera, record eye gaze data with an eye tracker, and collect facial expressions with a high-definition camera. Collect data covering 10 teaching types such as mathematics classroom lectures, chemistry experiment operations, and group discussions. For each type, collect 10,000 groups of samples. According to the teaching content, form, and objectives, label each group of data with the corresponding teaching type label, such as "classroom teaching - theoretical explanation" and "off-class practice - group experiment".

[0053] First, denoise the voice signal to remove environmental noise; then convert it into text, remove more than 2,000 stop words such as "de" and "le", and unify the word form, for example, change "running" to "run". Sample the gesture action data at an interval of 10 milliseconds. If the number of data points is less than 五百个, fill it with zeros; if it exceeds, intercept the first 500. Determine the three-dimensional coordinates of the attention core area of the eye gaze data every 500 milliseconds and organize them into a coordinate sequence. Quantify the facial expression intensity to the 0 - 1 interval by detecting the displacements of 12 key points (such as the corners of the eyes and the corners of the mouth), and refer to the FACS coding rule, where 0 represents no obvious expression and 1 represents strong emotion. Pretrain the BERT model using the Masked Language Model (Masked Language Model) and Next Sentence Prediction (NextSentence Prediction) tasks based on the processed voice text. In the Masked Language Model task, randomly replace some words in the text with "[MASK]" tags, and the model needs to predict the masked words based on the context. For example, for the sentence "[MASK] is a kind of fruit", the model should predict possible words such as "apple" and "banana". The Next Sentence Prediction task is to give two sentences and let the model determine whether the second sentence is the subsequent sentence of the first sentence to learn the logical relationship between sentences.

[0054] It should be noted that there is an unclear expression "五百个" in the original text which is translated as "五百个" here. You may need to clarify this content for a more accurate translation.During training, appropriate training parameters are set. The batch size is set to 128, meaning that 128 text samples are selected from the corpus and fed into the model for training each time. The initial learning rate is set to 0.0001. As training progresses, a learning rate decay strategy is implemented, with the learning rate multiplied by 0.9 after each training round (e.g., 100,000 steps). A total of approximately 1 million steps are trained to ensure that the model fully learns the semantic and grammatical information in the text. After training, the model is evaluated on a dedicated validation dataset containing text data that was not used in training. The effectiveness of the model training is assessed by calculating metrics such as the accuracy of masked word prediction and the accuracy of next sentence prediction. When the model's performance on the validation set meets the expected standards (e.g., masked word prediction accuracy exceeding 85% and next sentence prediction accuracy exceeding 90%), the model is considered successfully trained, resulting in a pre-trained BERT model that can be used to extract speech and text features.

[0055] Gesture data pre-training (spatiotemporal convolutional network): We collected a large amount of gesture data from a variety of people and scenarios. Using motion capture equipment, volunteers were asked to perform various common gestures (such as waving, giving a thumbs-up, and drawing a circle) in a laboratory setting. They also simulated gestures used in teaching scenarios (such as writing on a virtual blackboard and pointing at virtual laboratory instruments). The three-dimensional coordinate data of the joints associated with each gesture was recorded. Furthermore, gesture data from natural scenes was collected from public gesture datasets and video websites. After screening and annotation, a training set containing 100,000 gesture data sets was constructed. The gesture data was divided into fixed-length segments in chronological order, each containing joint coordinate data for 500 time points. The data was normalized, adjusting the joint coordinate values ​​to the range [0, 1] to ensure comparability across different scales.

[0056] A spatiotemporal convolutional network model was constructed, consisting of multiple convolutional, pooling, and fully connected layers. The convolutional layers extract spatial features of gestures, such as the relative positions of joints and the angles formed. The pooling layers reduce the data dimensionality and computational complexity while preserving key features. The fully connected layers integrate the extracted features and output a final feature vector. Supervised learning was used for training, with each gesture data segment labeled with a corresponding gesture category label (e.g., "wave" or "number 1 gesture"). During training, gesture data segments were fed into the network, which then output a predicted probability distribution for the gesture category. By calculating the cross-entropy loss between the predicted results and the true labels, the network parameters were adjusted using backpropagation to gradually approximate the model's predictions to the true labels. The training parameters were also set: a batch size of 64, an initial learning rate of 0.001, and a learning rate decay of 0.9 every 50 epochs. During training, the model's performance was regularly evaluated on a validation set consisting of 20,000 gesture data sets that were not used in training. The gesture classification accuracy is used as the evaluation indicator. When the classification accuracy of the model on the validation set exceeds 90%, the training is stopped and a pre-trained spatiotemporal convolutional network is obtained, which can be used to extract the gesture trajectory feature vector.

[0057] Eye Gaze Pre-training (Graph Neural Network): Eye gaze data from different individuals was collected during various visual tasks. Using an eye tracker, the coordinate sequences of eye fixations were recorded while participants were reading an article, watching an instructional video, and completing a visual search task. The corresponding scene information (e.g., paragraphs in an article, image content in a video) was annotated to construct a training set of 50,000 eye fixation data points. The eye fixation coordinate sequences were organized into a matrix, with each row representing the coordinates of a fixation point at a given time point, and the number of columns determined by the coordinate dimension (two-dimensional or three-dimensional). To better capture the relationships between fixations, a graph-structured data set was constructed, treating each fixation point as a node in a graph. Edges between nodes were established based on the temporal order and spatial position of the fixations. A graph neural network model was constructed, comprising components such as graph convolutional layers and graph pooling layers. The graph convolutional layer aggregates information between nodes and their neighbors to learn the relationship features between nodes; the graph pooling layer downsamples the graph structure to reduce computational complexity. The model was trained using a combination of unsupervised and supervised learning. During the unsupervised learning phase, a self-reconstruction loss function is designed to allow the model to learn how to reconstruct the original data from the input fixation map data, thereby extracting the data's latent feature representation. During the supervised learning phase, the model is trained to predict the scene category corresponding to the fixation point using annotated scene information as labels. Training parameters are set with a batch size of 32 and an initial learning rate of 0.0005. The learning rate is decayed to 0.9 after every 30 epochs. During training, model performance is evaluated using an independent validation set containing 10,000 data sets. Metrics such as scene category prediction accuracy and fixation feature reconstruction error are calculated. When the model's performance on the validation set meets the requirements (e.g., scene category prediction accuracy exceeding 80% and reconstruction error below a certain threshold), pre-training is completed, resulting in a graph neural network that can be used to extract attention feature vectors.

[0058] Emotional Intensity Pre-Training: A large amount of image and video data containing facial expressions was collected from sources such as public expression datasets, film clips, and social media videos. This data was annotated using the Facial Action Coding System (FACS). The images or video frames were labeled for 16 basic expression categories (such as smile, frown, surprise, and anger) and their intensity was assessed (categorized as weak, moderate, and strong). A training set of 200,000 images or video frames was constructed. The images or video frames were preprocessed by resizing them to a uniform size (e.g., 224×224 pixels) and normalizing them to keep pixel values ​​between 0 and 1. For video data, frames were extracted at a fixed frame rate (e.g., 10 frames per second). A deep learning model was constructed using a convolutional neural network (CNN) combined with fully connected layers. The CNN layer extracted facial features from the images, such as the texture and shape characteristics of the eyes, mouth, and eyebrows. The fully connected layer performed classification and regression on these extracted features, outputting confidence scores and emotion intensity ratings for the 16 basic expression categories. Training is performed using a multi-task learning approach, simultaneously optimizing both the expression category classification loss and the emotion intensity regression loss. During training, images or video frames are fed into the model, which then outputs predicted expression category probabilities and emotion intensity values. The model parameters are then updated using a backpropagation algorithm by calculating the cross-entropy loss (for expression classification) and mean squared error (for emotion intensity regression) between the predicted results and the true labels.

[0059] Training parameters were set, with a batch size of 64 and an initial learning rate of 0.001. The learning rate was decayed to its original value of 0.9 after every 40 epochs. The model was evaluated using a validation set consisting of 40,000 images or video frames. Metrics such as expression classification accuracy and mean absolute error (MAE) of emotion intensity prediction were calculated. Pre-training was completed when the model achieved expected performance on the validation set (e.g., expression classification accuracy exceeding 85% and MAE of emotion intensity prediction below 0.15). This resulted in a model that could scale emotion intensity values ​​to 32-dimensional vectors.

[0060] The features extracted from the above four dimensions (768-dimensional speech text features, 256-dimensional gesture trajectory feature vectors, 128-dimensional attention feature vectors, and 32-dimensional emotion intensity vectors) are spliced ​​together in sequence to form a 1280-dimensional multimodal feature vector.

[0061] A teaching scenario classification model was constructed using a four-layer cross-modal feature fusion layer and a softmax classifier. The first layer assigned speech modality weights based on the number of key knowledge points and the proportion of emotional vocabulary in the speech text, with important knowledge points weighed 0.7-0.9 and general narratives weighed 0.3-0.5. Gesture modality was weighted based on the magnitude of trajectory curvature changes and the number of speed mutations, with complex movements weighed 0.6-0.8 and regular movements weighed 0.2-0.4. Eye gaze was weighted based on the duration of gaze and the frequency of shifts, with areas of prolonged focused gaze weighed 0.7-0.9 and areas of dispersed gaze weighed 0.1-0.3. In the emotion intensity modality, strong emotions (0.8 and above) were weighted 0.8-1, while calm states (0.2 and below) were weighed 0.1-0.2.

[0062] The second layer calculates the interaction weights between modalities. When voice explains a knowledge point, gestures point in sync, and attention is focused, the interaction weights are each increased by 0.2. If the voice and gestures conflict, the weights are each reduced by 0.3. The third layer attenuates features below 0.4 based on the sum of the modal and interaction weights. The fourth layer concatenates the processed features. During training, the batch size is set to 64, the initial learning rate is 0.001, and the learning rate is decayed to 0.9 every five training cycles. Training is repeated for a total of 100 cycles, so that the model outputs probability distributions for 10 teaching types. Analysis of teaching scenario classification model: The multimodal fusion feature vector obtained in real time is input into the trained teaching scene classification model. In the cross-modal feature fusion layer, the self-attention weight is updated first: Real-time speech analysis: When knowledge point keywords are detected (matched by a vocabulary of 5,000+ knowledge points), the speech weight is set to 0.7-0.9; When the gesture involves complex operations (such as drawing circles continuously), the weight is adjusted to 0.6-0.8; If the eyes are fixed on a certain area for more than 3 seconds, the attention weight is increased to 0.8-1; When the facial expression shows strong emotions, the emotion weight is 0.8-1, and when it is calm, it is 0.1-0.2. The weights for intermodal interactions are then calculated. For example, if the experiment steps are explained by voice, gestures are directed at the device, and the student is looking at the device, the weights for each of these three interactions are increased by 0.2. If the student's expression is confused but the voice or gestures do not respond, the weights for each of these three interactions are decreased by 0.2. Finally, the Softmax classifier outputs a probability distribution of 10 teaching types. If the highest probability exceeds 0.6, it is determined to be that type. If it is less than 0.6, the top two probabilities are selected and weighted averaged according to the probability ratio to generate a mixed scenario label.

[0063] The role model matches the interaction rules: Based on the teaching type identifiers obtained through classification, a search is performed in the three-dimensional character template library. The template library is divided into 10 major categories based on teaching type, with 3-5 character models of different styles under each category. For example, for classroom teaching, there are rigorous professors and friendly lecturers. The similarity between the preset features of the candidate character models and the current multimodal features is calculated. Voice style is compared with speech speed (error ±10 words / minute) and intonation; gesture preference is compared with common gesture types; and facial expression features are compared with the frequency of smiling and frowning. The top three character models are selected using cosine similarity. Each character model is associated with a specific interaction rule set, which is divided into teaching stages. For example, the classroom teaching rule set includes rule groups for introduction, explanation, and questioning stages. The corresponding rule group is selected based on the teaching progress (percentage of course time, keywords in the explanation content). For example, the introduction rule group is selected at the beginning of the course.

[0064] Dynamic extraction of response strategies: The state parameters of the interactive phase are analyzed. The teaching progress is determined by the ratio of the current time to the total course duration. The knowledge point numbers are matched with the 1000+ knowledge point mapping table through semantic analysis. The student engagement is calculated by the time the student's attention stays in the core area (0.1 points per second) and the emotional intensity value, with a maximum score of 1 point. In the interactive rule set, each response strategy has trigger conditions and parameter adjustment rules: Voice response delay compensation: when the student engagement is lower than 0.4 and the emotion intensity is lower than 0.3, the delay is increased by 200 milliseconds, otherwise it remains at 100 milliseconds; The gesture synchronization threshold is relaxed from 5 cm to 5.75 cm when the gesture speed exceeds 1.2 m / s. The line of sight focus rendering priority is increased when gazing at a certain area for more than 3 seconds, and the model detail accuracy is improved by 2 levels.

[0065] Response strategies are extracted based on the current stage parameters and trigger conditions. In case of conflicts, they are handled according to the priority preset by teaching experts (voice > gesture > sight) and combined into the current stage response strategy set. By updating the template library and rule set, it can quickly adapt to new teaching modes, has strong scalability, and combines knowledge point mapping with visualization strategies to present complex content intuitively.

[0066] In a preferred embodiment of the present invention, based on the corresponding role model and the associated interaction rule set, implementing a targeted interactive response may include: Based on the role model and the associated interaction rule set, generate behavioral logic instructions corresponding to the teaching type identifier, including: When the role model is a virtual teacher behavior model, guided question-answering instructions and visual demonstration action sequences associated with knowledge points are generated based on the semantic guidance rule set; When the role model is a student collaboration logic model, collaborative task assignment instructions and abnormal operation detection signals are generated based on the task priority assignment protocol; When the role model is an enterprise mentor scenario model, device control signals and process optimization parameter adjustment instructions are generated based on the device operation instruction mapping relationship; The behavioral logic instructions are input into the three-dimensional virtual image parameterized driving engine, which drives the interactive response of the digital human by converting the guided question-and-answer instructions into playable voice waveform data of the speech synthesis model, converting the visual demonstration action sequence into skeletal driving parameters and corresponding limb motion trajectories, and mapping the device control signals into the device operation instruction flow of the virtual production environment.

[0067] In an embodiment of the present invention, when it is identified that the teaching type is a classroom teaching module, the virtual teacher behavior model is called. First, the semantic guidance rule set is parsed, which contains more than 5000 predefined rules, and each rule is associated with a specific knowledge point. For example, when it is detected that the knowledge point currently being explained is the "Pythagorean Theorem", the corresponding 3 guided question-and-answer strategies are extracted from the rule set (such as "The two right-angled sides of a right triangle are 3 and 4 respectively, what is the hypotenuse?", "How to find the other right-angled side through the hypotenuse and one right-angled side?") and 2 visual demonstration schemes (dynamic graphic display, algebraic derivation steps).

[0068] These strategies are matched to a knowledge point association mapping table that stores over 1,000 knowledge points and their corresponding teaching resource IDs. For example, the "Pythagorean Theorem" has ID KN0015, which is associated with five animation demonstration files and three example problem explanation videos. Based on the current teaching progress (e.g., the 20%-30% progress), the most suitable two animations and one video are selected from the associated resources. Guided question-and-answer instructions and a visual demonstration action sequence are generated, including the voice question text and the demonstration resource ID.

[0069] Student Collaboration Logic Model: In the off-class practice team module, the task priority assignment protocol is first analyzed. This protocol evaluates tasks based on eight dimensions, including task difficulty, time consumption, and resource requirements. For example, for a "chemistry experiment report writing" task, parameters such as the amount of experimental data (500+ data points are complex), the computational difficulty (involving multivariate equations is high difficulty), and the completion time requirement (within 24 hours is urgent) are analyzed. Based on these parameters, the "division of labor and collaboration" model is selected from 10 preset collaboration modes, and task assignment instructions are generated: the data sorting task is assigned to student A, who excels in data processing (based on historical collaboration records, his data processing accuracy rate is 95%), and the computational task is assigned to student B, who has strong mathematical skills (and has won awards in mathematics competitions). At the same time, the abnormal operation detection protocol is activated, setting the data entry error threshold to ±5% and the calculation result fluctuation threshold to ±3%, and generating corresponding monitoring signals.

[0070] Enterprise mentor scenario model: In the enterprise practice module, according to the mapping relationship of equipment operation instructions, the operation descriptions in the teaching content are converted into equipment control signals. For example, when explaining "CNC machine tool operation", the text instruction "Move the tool to X coordinate 150mm, Y coordinate 80mm, and Z coordinate 30mm" is converted into a control signal recognizable by the equipment (X: 150, Y: 80, Z: 30, Speed: 200). At the same time, analyze the current production process parameters (such as machining accuracy 0.05mm, production efficiency 80 pieces per hour), call the process optimization constraint conditions (such as cost ceiling 500 yuan per hour, energy consumption limit 80kW), and generate parameter adjustment instructions (adjust the feed rate from 200mm / min to 180mm / min, and the cutting depth from 2mm to 1.8mm) by comparing historical data (accuracy 0.03mm, efficiency 90 pieces per hour).

[0071] Parametric drive of 3D virtual avatar: Input the text content in the guided Q&A instruction into the speech synthesis model. First, perform text preprocessing, including polyphonic character judgment (such as "行" in "银行" is pronounced as háng), emotion annotation (such as marking interrogative sentences with a doubtful tone), and speech rate adjustment (set to 180 words per minute for the knowledge point explanation part and 150 words per minute for the key emphasis part). The speech synthesis model is based on the Tacotron2 architecture, first converts the text into a Mel spectrogram, and then converts the spectrogram into waveform data through the WaveGlow model. During the conversion process, according to the role setting of the virtual teacher (such as a female lecturer), adjust the fundamental frequency (set to 220Hz) and timbre parameters (increase brightness), and finally generate a playable speech file.

[0072] For visual demonstration action sequences, they are converted into skeletal drive parameters. For example, the "Writing a Formula on a Blackboard" action sequence is first parsed into keypoint coordinates (e.g., starting point X100, Y200, end point X300, Y200) and trajectory type (linear motion). Next, the action library, which stores skeletal parameters for over 500 basic actions (such as hand raising, pen waving, and turning), is queried. The "Writing with a Pen" action group (containing skeletal parameters for 12 consecutive frames) is selected. Based on the parsed coordinates and trajectory, the basic action parameters are scaled and translated (e.g., increasing the action amplitude by 1.2 times and translating 50 units along the X-axis) to generate the final skeletal drive parameters and limb motion trajectory. Device control signals are mapped into an operational command stream for the virtual production environment. For example, for the control signal "Start the centrifugal pump and set the speed to 1500 rpm," the device mapping table is first queried to determine that the corresponding device ID in the virtual environment is EQ007. Then, a command packet is generated containing the device ID, operation type (start), and parameter value (1500 rpm). To ensure operational safety, pre-check instructions (such as checking valve status and confirming power connection) and post-feedback instructions (monitoring pressure changes and recording startup time) are added to ultimately form a complete operational instruction flow.

[0073] Through semantic guidance rules and knowledge point mapping, interactive responses that are highly matched with teaching content can be achieved.

[0074] Precise mapping of device control signals and process optimization parameter adjustment reduce operational errors in virtual production environments. Personalized voice and motion control enhance the fit between the digital human and the teaching scenario, increasing learner immersion and engagement. Targeted interactive responses are supported for three core teaching scenarios, and by expanding the role model and rule set, new teaching needs can be quickly adapted.

[0075] In a preferred embodiment of the present invention, feedback sensors are used to capture user interaction response latency data and voice feedback accuracy in real time, and a reinforcement learning agent is used to dynamically adjust the parameters of the dynamic weight distribution algorithm and update the interaction rule set and device control protocol of the three-dimensional character template to form an adaptive optimization loop, which may include: The user's operational data on interactive responses is captured in real time through feedback sensors, including the response delay time and semantic matching accuracy of user voice input in voice interaction scenarios, the synchronization delay difference between user actions and digital human skeleton drive parameters in gesture interaction scenarios, and the offset between the user's gaze area and the rendering focus in visual focus scenarios. Quantify the difference between the interactive response operation data and the expected response target to generate optimization indicators, including voice response delay compensation deviation value, gesture synchronization error coefficient and visual focus deviation; Based on the optimization indicators, the weight distribution ratio of the multimodal fusion feature vector, the voice response delay compensation parameter, the gesture action synchronization threshold, and the associated interaction rule set are adjusted through the reinforcement learning agent to obtain the adjusted parameters and rule set; The adjusted parameters and rule sets are synchronized to the 3D character template library and multimodal fusion module, forming a closed-loop adaptive cycle from data acquisition to parameter optimization.

[0076] In an embodiment of the present invention, data is collected through a variety of feedback sensors deployed in the teaching environment. In the voice interaction scenario, the microphone array and the voice recognition module work together. When the digital human issues a guided question and answer instruction, the timer starts. Once the user's voice input is captured, the time difference from the instruction issuance to the voice reception is immediately recorded as the response delay time. At the same time, after the user's voice is converted into text, it is compared word by word with the correct semantic content expected by the system. The number of correctly matched words is counted and then divided by the total number of words to obtain the semantic matching accuracy. For example, if the question is "How to calculate the hypotenuse of a right triangle", if the user's answer contains key steps, each correct answer to a step corresponding to the word is recorded as a correct match.

[0077] In gesture interaction scenarios, the depth camera continuously tracks user gestures and acquires the digital human's skeletal drive parameters. The system calculates the spatial distance between the coordinates of the user's gesture joints and the corresponding joints of the digital human at the same moment in time, using this distance difference as the synchronization delay. If the user waves, the system compares the user's hand motion trajectory with the simulated wave trajectory of the digital human, accurately measuring the positional differences between the two at each point in time.

[0078] For visual focus scenarios, the eye tracker monitors the area where the user's gaze rests in real time, and the system knows the specific location of the digital human's rendering focal point. The visual focus offset is calculated by calculating the Euclidean distance in three-dimensional space between the user's gaze center coordinates and the rendering focal point coordinates. If the digital human is demonstrating an experimental device, the system determines whether the user's gaze is focused on a key part of the device. If not, the offset distance is calculated. The collected operational data is compared with the pre-set expected response target. For voice interaction, the expected voice response delay is set at 1-1.5 seconds, and the semantic matching accuracy target is 90%. The voice response delay compensation deviation value is calculated by subtracting the expected delay from the actual response delay. If the actual delay is 2 seconds, the deviation value is 0.5-1 second. For the semantic matching accuracy, the actual accuracy is subtracted from 90% to obtain the semantic matching accuracy difference.

[0079] In gesture interaction scenarios, the expected synchronization delay difference should be less than 10 cm. Compare the actual synchronization delay difference with 10 cm. If the actual difference is 15 cm, then the gesture synchronization error coefficient is the actual difference divided by 10 cm, which is 1.5.

[0080] In the visual focus scenario, the expected visual focus offset should be less than 5 cm. Simply compare the actual visual focus offset with 5 cm. If the offset is 8 cm, the visual focus offset is 8 cm. These calculated values ​​intuitively reflect the gap between the current interactive response and the ideal state.

[0081] The reinforcement learning agent adjusts system parameters and rules based on the generated optimization metrics. Regarding the weighting of multimodal fusion feature vectors, if the speech response delay compensation deviation is large and the semantic matching accuracy is low, the speech modality may be underweighted. The agent will then appropriately increase the weight of the speech-text feature in the multimodal fusion, for example, from 0.4 to 0.5. Regarding the speech response delay compensation parameter, if the deviation is positive (actual delay is too long), the agent will increase the speech response delay compensation time, for example, from 200 milliseconds to 300 milliseconds; if it is negative (actual delay is too short), the compensation time will be reduced. Regarding the gesture synchronization threshold, if the gesture synchronization error coefficient is greater than 1, indicating that the current threshold is too strict, the agent will relax the synchronization threshold from 5 cm to 6 cm. Regarding the associated interaction rule set, if users repeatedly experience low semantic matching accuracy for a certain knowledge point, the agent will add more examples and explanatory question-and-answer rules to the semantic guidance rule set for that knowledge point. If gesture synchronization issues are frequent, the agent will optimize the conversion rules between gestures and digital human skeletal drive parameters to achieve more accurate matching.

[0082] After the reinforcement learning agent completes parameter and rule adjustments, it synchronizes the new weight distribution ratios, adjusted voice response delay compensation parameters, gesture synchronization thresholds, and updated association interaction rule sets to the 3D character template library and multimodal fusion module. Each character model in the 3D character template library adjusts its behavioral logic based on the new rule set, such as the virtual teacher adopting a new question-and-answer strategy during lectures. The multimodal fusion module processes voice, gesture, and other data according to the new weight distribution ratios. New operational data is continuously collected, and the above process is repeated for continuous optimization, forming a complete closed loop from data collection, indicator calculation, parameter adjustment, to updated application.

[0083] Through real-time data feedback and dynamic adjustments, interaction accuracy is improved. Response strategies can be automatically optimized based on different users' interaction habits and scenario requirements, such as adding voice delay compensation for slower-responding users, providing a more personalized learning experience. This closed-loop adaptive system continuously learns and improves, reducing interaction lags caused by inappropriate parameters. Whether in classroom teaching, practical operations, or corporate training scenarios, it can quickly adapt to changing environments and optimize responses.

[0084] An embodiment of the present invention further provides a computing device comprising: a processor and a memory storing a computer program, wherein when the computer program is executed by the processor, the system described above is executed. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.

[0085] The embodiment of the present invention further provides a computer-readable storage medium storing instructions, which, when executed on a computer, cause the computer to execute the system described above. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.

[0086] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. A digital human 3D reconstruction system for teaching interaction, characterized by: include: The data acquisition module is used to collect the user's voice signals, gesture data, eye tracking data and facial expression data in real time through a multimodal sensor array, and obtain the teaching type identifier, interaction stage state parameters and participant role identifier of the current teaching scene through a preset teaching scene database; The multimodal fusion module is used to convert speech signals into text, extract three-dimensional spatial trajectory features from gesture data, calculate the spatial coordinates of the gaze focus and encode expression features from eye tracking data and facial expression data, and generate a multimodal fusion feature vector through a dynamic weight allocation algorithm. The strategy matching module is used to input the multimodal fusion feature vector into the pre-trained teaching scenario classification model, dynamically match the corresponding role model and the associated interaction rule set from the three-dimensional role template library according to the output teaching type identifier, and obtain the response strategy matched at the current stage; An interactive response module, used to implement targeted interactive responses based on the corresponding role model and associated interaction rule set; The optimization loop module is used to capture the user's operation delay data and voice feedback accuracy in real time through feedback sensors, dynamically adjust the dynamic weight distribution algorithm parameters through reinforcement learning agents, and update the interaction rule set and device control protocol of the three-dimensional character template to form an adaptive optimization loop.

2. The digital human 3D reconstruction system for teaching interaction according to claim 1 is characterized in that: The speech signal is converted into text, the gesture data is extracted into three-dimensional spatial trajectory features, the eye tracking data and facial expression data are calculated into the spatial coordinates of the gaze focus and the expression features are encoded, and a multimodal fusion feature vector is generated through a dynamic weight allocation algorithm, including: The speech signal is subjected to noise suppression and speech segmentation processing, the time-frequency features of the speech signal are extracted, and a semantic text sequence is directly generated through an end-to-end speech recognition network. At the same time, the emotional tendency intensity value is calculated based on the speech spectrum energy distribution; the gesture action data is calibrated for the joint point spatial coordinates, and a smooth three-dimensional spatial trajectory is generated based on the calibration results using the cubic spline interpolation algorithm, and the direction vector, curvature characteristics and instantaneous velocity parameters of the trajectory are extracted; the pupil's gaze focus spatial coordinates in the three-dimensional teaching scene are calculated based on the eye tracking data, and the core area and diffusion range of the user's attention distribution are determined through the density clustering algorithm; the facial expression data is subjected to facial muscle group movement amplitude detection, and combined with the predefined expression coding rule library, expression category labels and emotional intensity level parameters are generated; According to the teaching type identifier and the interaction stage state parameters, the semantic text sequence and emotional tendency intensity value, the extracted trajectory curvature features, the attention core area and the emotional intensity level parameters are input into the dynamic weight allocator, and a multimodal fusion feature vector is generated through weighted splicing and feature dimensionality reduction.

3. The interactive teaching digital human 3D reconstruction system according to claim 2 is characterized in that: Perform noise suppression and speech segmentation on speech signals, extract the time-frequency features of speech signals, and directly generate semantic text sequences through an end-to-end speech recognition network. At the same time, calculate the emotional tendency intensity value based on the speech spectrum energy distribution, including: The adaptive filtering algorithm is used to suppress background noise in speech signals, and the speech activity detection algorithm is used to segment effective speech segments and extract the time-frequency spectrum characteristics and energy distribution parameters of each speech segment. Input the time-frequency spectrum features into the end-to-end speech recognition network, output the semantic text sequence corresponding to the speech segment, and extract the recognition probability value of each word based on the word-level probability distribution output by the network; Based on the word-level probability values, the average probability value of all words in the text sequence is calculated as the basic confidence score, and the score is dynamically calibrated in combination with the energy distribution parameters of the speech segment to generate a global confidence score; Based on the predefined emotional keyword library, keyword matching is performed on the semantic text sequence, and the frequency and contextual association weight of emotional keywords are counted; Based on the speech energy distribution parameters, the high-frequency energy proportion and intonation fluctuation amplitude of the speech segment are analyzed, and the emotional tendency intensity value is generated by combining the frequency of emotional keywords and context-related weights.

4. The digital human 3D reconstruction system for teaching interaction according to claim 3 is characterized in that: The multimodal fusion feature vector is input into the pre-trained teaching scenario classification model. Based on the output teaching type identifier, the corresponding role model and associated interaction rule set are dynamically matched from the 3D role template library to obtain the response strategy matched at the current stage, including: A teaching scenario classification model is trained based on a multimodal training dataset containing multimodal feature vectors of speech text sequences, gesture trajectory features, attention core areas, and emotion intensity parameters, as well as corresponding teaching type labels. A probability distribution of teaching type identifiers is generated through a cross-modal feature fusion layer and a classifier. The multimodal fusion feature vector is input into the teaching scene classification model, the joint weight of each modal feature is calculated through the cross-modal feature fusion layer, the probability distribution of the current teaching type identifier is output, and the scene classification result is obtained; According to the teaching type identifier, the corresponding role model and the associated interaction rule set are matched from the three-dimensional role template library; Based on the interaction stage state parameters, the response strategy matching the current stage is dynamically extracted from the associated interaction rule set, including voice response delay compensation parameters, gesture action synchronization threshold and gaze focus rendering priority.

5. The teaching interactive digital human 3D reconstruction system according to claim 4 is characterized in that: According to the teaching type identifier, the corresponding role model and associated interaction rule set are matched from the 3D role template library, including: When the teaching type identifier is classroom teaching module, the virtual teacher behavior model and the associated semantic guidance rule set, knowledge point association mapping table and visual dynamic demonstration strategy are extracted; When the teaching type identifier is the off-class practice team module, the student collaboration logic model and task priority allocation rules, abnormal operation detection protocol and collaboration status feedback mechanism are extracted; When the teaching type identifier is the enterprise practice module, the enterprise instructor scenario model and equipment operation instruction mapping table, process optimization constraints and production environment real-time data interface are extracted.

6. The teaching interactive digital human 3D reconstruction system according to claim 5 is characterized in that: Based on the corresponding role model and associated interaction rule set, a targeted interactive response is achieved, including: Generate behavioral logic instructions corresponding to the teaching type identifier based on the role model and the associated interaction rule set; The behavioral logic instructions are input into the three-dimensional virtual image parameterized driving engine, which drives the interactive response of the digital human by converting the guided question-and-answer instructions into playable voice waveform data of the speech synthesis model, converting the visual demonstration action sequence into skeletal driving parameters and corresponding limb motion trajectories, and mapping the device control signals into the device operation instruction flow of the virtual production environment.

7. The teaching interactive digital human 3D reconstruction system according to claim 6 is characterized in that: Based on the role model and the associated interaction rule set, generate behavioral logic instructions that strictly correspond to the teaching type identifier, including: When the role model is a virtual teacher behavior model, guided question-answering instructions and visual demonstration action sequences associated with knowledge points are generated based on the semantic guidance rule set; When the role model is a student collaboration logic model, collaborative task assignment instructions and abnormal operation detection signals are generated based on the task priority assignment protocol; When the role model is an enterprise mentor scenario model, device control signals and process optimization parameter adjustment instructions are generated based on the device operation instruction mapping relationship.

8. The digital human 3D reconstruction system for teaching interaction according to claim 7 is characterized in that: Feedback sensors capture user interaction latency data and voice feedback accuracy in real time. A reinforcement learning agent dynamically adjusts the parameters of the dynamic weight distribution algorithm and updates the interaction rule set and device control protocol for the 3D character template, forming an adaptive optimization loop that includes: The user's operational data on interactive responses is captured in real time through feedback sensors, including the response delay time and semantic matching accuracy of user voice input in voice interaction scenarios, the synchronization delay difference between user actions and digital human skeleton drive parameters in gesture interaction scenarios, and the offset between the user's gaze area and the rendering focus in visual focus scenarios. Quantify the difference between the interactive response operation data and the expected response target to generate optimization indicators, including voice response delay compensation deviation value, gesture synchronization error coefficient and visual focus deviation; Based on the optimization indicators, the weight distribution ratio of the multimodal fusion feature vector, the voice response delay compensation parameter, the gesture action synchronization threshold, and the associated interaction rule set are adjusted through the reinforcement learning agent to obtain the adjusted parameters and rule set; The adjusted parameters and rule sets are synchronized to the 3D character template library and multimodal fusion module, forming a closed-loop adaptive cycle from data acquisition to parameter optimization.

9. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the system according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which, when executed by a processor, implements the system according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Virtual-real fusion teaching aid automatic generation method

    CN112230772A

  • Generative teaching resource system in virtual teaching scene and working method thereof

    CN117055724A

  • Intelligent simulation system and method based on multi-modal interaction and dynamic teaching strategy

    CN119624711A

  • Middle and primary school multi-person foreign language situational teaching method and system based on VR

    CN120031684A

  • Intelligent real-time interactive question-answering system based on virtual digital human

    CN120318388A

Cited By

  • Dynamic adjustment method, system and equipment for VR spacecraft assembly and storage medium

    CN120848777A

  • Dynamic adjustment method, system, device and storage medium for vr spacecraft assembly

    CN120848777B

  • Digital teacher personalized behavior modeling method based on multi-modal feature fusion

    CN121392076A