Teaching interaction method and device based on virtual digital human, equipment and medium
By partition coding and eye tracking of student video streams, combined with pupil coordinate mapping, driving virtual digital human head steering and 3D parallax rendering, the accuracy of multi-student interaction in the virtual digital human teaching system is solved, and the authenticity and resource utilization of immersive interaction are improved.
Patent Information
- Application Number
- CN202510499552.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing virtual digital human teaching system is difficult to achieve accurate interaction in multiple student scenarios, and the lack of fusion analysis of multi-dimensional data such as eye movement trajectory and micro-expression has led to imbalance in the allocation of attention resources, and the video encoding strategy is separated from the virtual digital human action control, resulting in serious loss of details in the lower part of high-concurrency scenarios and significant delay in action feedback, which limits the large-scale application of immersive educational technology.
By collecting multiple video streams of students, they perform facial, gesture and background area segmentation and differentiated encoding, combining eye tracking to generate attention weights, dynamically adjust the area code rate, and driving virtual digital human head steering and 3D parallax rendering based on pupil coordinate mapping, realizing adaptive slice distribution and optimizing network resource utilization.
It realizes accurate perception of multiple students' attention in teaching scenarios, and responds simultaneously with virtual image actions and real intentions, significantly improving the authenticity of immersive interactions and system resource utilization.
Smart Images

Figure CN120374812A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent interaction, and particularly relates to a teaching interaction method, device, equipment and medium based on a virtual digital human. Background Art
[0002] In a virtual digital human teaching system, how to achieve precise interaction in a multi-student scenario has always been a technical difficulty. Traditional teaching platforms usually adopt a video transmission with a fixed perspective and a one-way knowledge infusion mode, which is difficult to capture the real-time cognitive state differences of students, resulting in an unbalanced allocation of attention resources.
[0003] Existing technologies often rely on a single interaction signal (such as voice questions or gesture triggers) to judge students' needs, lacking the fusion analysis of multi-dimensional data such as eye movement trajectories and micro-expressions, and it is easy to misjudge the real attention focus of learners.
[0004] At the same time, the split design of the video coding strategy and the virtual digital human motion control makes the facial details seriously lost and the motion feedback significantly delayed in a high-concurrency scenario. Especially when the network fluctuates, the problems of video freezing and interaction out-of-sync are exacerbated.
[0005] The above defects lead to the difficulty of virtual teaching in achieving the interactive effect of a real classroom, restricting the large-scale application of immersive education technologies. Summary of the Invention
[0006] Based on this, it is necessary to provide a teaching interaction method, device, equipment and medium based on a virtual digital human for the above technical problems.
[0007] In a first aspect, the present application provides a teaching interaction method based on a virtual digital human, including:
[0008] S1: Collect the original video streams of multiple students, divide each original video stream into a face area, a gesture area and a background area; based on the face area, the gesture area and the background area, perform partition coding on the original video streams to generate a mixed-coded video stream;
[0009] S2: According to the eye movement characteristics of the students, screen out the high-concern areas from the face area, and mark the students corresponding to the high-concern areas as high-concern students;
[0010] S3: Generate an attention weight based on the duration of the high-concern area, and adjust the partition bit rate of the mixed-coded video stream according to the attention weight to generate an optimized video stream;
[0011] S4: Map the pupil center coordinates of the high-concern students to the facial animation parameters based on the QP value of the face area and the frame type of the gesture area in the optimized video stream;
[0012] S5: Generate the head turning angle of the virtual digital human based on the facial animation parameters; perform 3D parallax rendering on the head turning angle according to the multi-viewpoint encoding standard to generate a multi-viewpoint 3D video stream;
[0013] S6: Perform adaptive slicing on the multi-viewpoint 3D video stream to generate an adaptive bitrate video stream, and send the adaptive bitrate video stream to the student side; wherein, the adaptive slicing process includes dynamically adjusting the slice duration according to the network state.
[0014] In a second aspect, the present application also provides a teaching interaction device based on a virtual digital human, including:
[0015] A video stream processing module, configured to collect the original video streams of multiple students, divide each original video stream into a facial area, a gesture area, and a background area; based on the facial area, the gesture area, and the background area, perform partition encoding on the original video stream to generate a mixed encoded video stream;
[0016] A region of interest analysis module, configured to screen out the regions of high interest from the facial area according to the eye movement characteristics of the students, and mark the students corresponding to the regions of high interest as high-interest students;
[0017] A video stream optimization module, configured to generate an attention weight based on the duration of the regions of high interest, and adjust the partition bitrate of the mixed encoded video stream according to the attention weight to generate an optimized video stream;
[0018] An animation parameter generation module, configured to map the pupil center coordinates of the high-interest students to facial animation parameters based on the QP value of the facial area and the frame type of the gesture area in the optimized video stream;
[0019] A 3D video generation module, configured to generate the head turning angle of the virtual digital human based on the facial animation parameters; perform 3D parallax rendering on the head turning angle according to the multi-viewpoint encoding standard to generate a multi-viewpoint 3D video stream;
[0020] An adaptive video transmission module, configured to perform adaptive slicing on the multi-viewpoint 3D video stream to generate an adaptive bitrate video stream, and send the adaptive bitrate video stream to the student side; wherein, the adaptive slicing process includes dynamically adjusting the slice duration according to the network state.
[0021] In a third aspect, the present application also provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements a teaching interaction method based on a virtual digital human as in the first aspect.
[0022] Fourthly, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a teaching interaction method based on a virtual digital human as described in the first aspect.
[0023] The above-mentioned teaching interaction method, device, equipment and medium based on a virtual digital human perform semantic segmentation and differential encoding on the student video stream according to the face, gesture and background regions, dynamically adjust the bitrate allocation of each region by combining the attention weights generated by eye movement tracking, and drive the head turning and 3D parallax rendering of the virtual digital human based on the pupil coordinate mapping. At the same time, adaptive slice distribution is performed according to the network state, achieving the technical effects of accurately perceiving the attention of multiple students in the teaching scenario, synchronously responding the actions and real intentions of the virtual image, and optimizing the network resources as needed, significantly improving the authenticity of immersive interaction and the utilization rate of system resources. Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 It is a schematic flowchart of a teaching interaction method based on a virtual digital human provided by the present invention;
[0026] Figure 2 It is a schematic structural diagram of a teaching interaction device based on a virtual digital human provided by the present invention. Detailed Embodiments
[0027] In order to make the purpose, technical solutions and advantages of the present application clearer, the following further details the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0028] Refer to Figure 1 , which shows a schematic flowchart of a teaching interaction method based on a virtual digital human provided by the present application. The method includes the following steps:
[0029] S1: Collect the original video streams of multiple students, and divide each original video stream into a face region, a gesture region, and a background region; based on the face region, the gesture region, and the background region, perform partition encoding on the original video stream to generate a mixed encoded video stream.
[0030] Specifically, when collecting the original video stream, a high-resolution and low-latency camera array can be deployed to synchronously collect at a certain frame rate (such as 30fps or 60fps). The camera can have autofocus and light compensation functions to ensure the clarity and stability of the video stream under different lighting conditions.
[0031] The original video stream can be analyzed in real time based on a deep learning object detection model (such as YOLOv8 or MediaPipe), dividing the face area (covering key feature points such as eyes, nose, and mouth), gesture area (hand contour and action posture), and background area.
[0032] The H.265 / HEVC coding standard can be adopted, and different quantization parameters (QP) are set for different regions. The QP value of the face area is set to 20 - 25 to ensure the retention of key feature details; the QP value of the gesture area is 26 - 30 to balance action smoothness and coding efficiency; the QP value of the background area is 31 - 35 to reduce redundant data transmission. Through the region mask and the ROI (region of interest) mechanism of the encoder, a mixed-coded video stream is generated, and the coding delay can be controlled within 100ms.
[0033] S2: According to the eye movement characteristics of students, the high-attention areas are screened out from the face area, and the students corresponding to the high-attention areas are marked as high-attention students.
[0034] Specifically, an eye key point detection algorithm (a combined model based on OpenCV and Dlib) can be used to track 68 feature points of students' eyes in real time, calculating the horizontal and vertical saccade speeds, fixation duration, and blink frequency.
[0035] An attention evaluation model is constructed. Inputting the eye movement feature vector, it is determined whether there is a high-attention area within the face area through an SVM (support vector machine) classifier. For example, the model threshold can be set as the area where the fixation duration exceeds 2000ms and the saccade speed is lower than 50 degrees / second.
[0036] A unique ID label can be assigned to each student. When a high-attention area is detected, the ID of the high-attention student and the corresponding area coordinate information are broadcast to other modules of the system through the MQTT (Message Queuing Telemetry Transport) message bus, and the update frequency of the marking information can be set to 2 times per second.
[0037] S3: Generate an attention weight based on the duration of the high-attention area, and adjust the partition bit rate of the mixed-coded video stream according to the attention weight to generate an optimized video stream.
[0038] Specifically, based on the continuous detection duration of the high-attention area, an exponential decay function can be used to calculate the attention weight W(t) = e-λt where λ is the attenuation coefficient (the value can be 0.05 / s), and t is the duration. The weight value range can be set from 0.1 to 1.0, which reflects the degree of students' concentration.
[0039] A dynamic bitrate allocation matrix can be constructed to map the attention weight to the bitrate adjustment factor. For example, the bitrate priority of the facial area is increased by W(t)×50%, the gesture area is increased by W(t)×30%, and the background area is decreased by W(t)×20%. The real-time bitrate adjustment is achieved through the dynamic bitrate control module of the encoder (such as the vbv-bufsize mechanism of FFmpeg), and the adjustment period can be set to 500 ms.
[0040] The video stream after bitrate adjustment is encapsulated by RTP (Real-time Transport Protocol), and a custom payload header containing the attention weight and region information is added to ensure that the subsequent processing module can parse the optimization parameters.
[0041] S4: Map the pupil center coordinates of the highly concerned students to the facial animation parameters based on the QP value of the facial area and the frame type of the gesture area in the optimized video stream.
[0042] Specifically, the quantization parameter (QP) of the facial area directly affects the degree of detail retention in the facial area in H.265 / HEVC encoding. A lower QP value (such as 20 - 25) means less quantization loss and can retain fine features such as pupil edges and iris textures.
[0043] The frame types (I-frame, P-frame, B-frame) of the gesture area reflect the dynamic characteristics of the students' actions. The I-frame provides a reference for the complete gesture contour, while the P-frame and B-frame capture the changing trends of the gestures. By analyzing the inter-frame differences, it is judged whether the students' gestures are in a stable state (such as raising a hand to signal) or in a dynamic change (such as waving for interaction). The gesture state affects the response strategy of the virtual digital human to the attention direction. For example, when it is detected that the gesture area of three consecutive frames is an I-frame and the QP value is stable below 28, it is determined that the gesture is in a stable display state, and the pupil center coordinates in the corresponding direction are preferentially associated.
[0044] A pupil center coordinate mapping model weighted based on the QP value can be constructed. The inputs of the model are the QP value sequence of the facial area, the frame type encoding of the gesture area, and the original pupil center coordinates in the optimized video stream.
[0045] A 48-dimensional facial animation parameter vector can be defined, covering key features such as eye opening degree, eyeball rotation angle, and eyebrow displacement. By constructing a 3×48 parameter association matrix, a linear combination of the mapped pupil center coordinates and the facial animation parameters is performed to obtain new facial animation parameters.
[0046] A parameter correction mechanism based on student feedback can also be introduced. For example, when the student side detects that the deviation between the turning direction of the virtual digital human's head and its own gaze direction exceeds 10°, a correction request is sent to the server through the WebSocket channel. The server dynamically adjusts the weights of the corresponding rows in the parameter association matrix according to the deviation value, and the correction period can be set to 2 seconds to ensure that the system adapts to the visual habit differences of different students.
[0047] S5: Generate the turning angle of the virtual digital human's head based on the facial animation parameters; perform 3D parallax rendering on the turning angle based on the multi-viewpoint coding standard to generate a multi-viewpoint 3D video stream.
[0048] Specifically, based on the line-of-sight direction parameter in the facial animation parameters, the turning angle of the virtual digital human's head can be calculated through the quaternion interpolation algorithm. The turning angle range can be limited within ±45° to ensure a natural visual interaction effect.
[0049] The MPEG-MMT (Multimedia Multi-Viewpoint Transport) standard can be adopted to construct a multi-viewpoint coding framework that includes a base viewpoint (the front view of the virtual digital human) and auxiliary viewpoints (dynamically generated according to the student's attention direction). Each viewpoint is independently encoded, and the amount of redundant data is reduced through Inter-View Prediction.
[0050] The OpenGL ES 3.2 rendering engine can be used to calculate the parallax depth map based on the screen resolution and viewing distance parameters of the student side device. Through occlusion testing and depth compensation algorithms, a multi-viewpoint 3D video stream with a real parallax effect is generated, and the rendering frame rate can be maintained above 24fps.
[0051] S6: Perform adaptive slicing on the multi-viewpoint 3D video stream to generate an adaptive bitrate video stream, and send the adaptive bitrate video stream to the student side; among them, the adaptive slicing process includes dynamically adjusting the slice duration according to the network state.
[0052] Specifically, the HLS (HTTP Live Streaming) or DASH (Dynamic Adaptive Streaming over HTTP) standard can be adopted for slicing. The network state is evaluated in real time according to the network bandwidth monitoring module (calculated based on the TCP receive window and RTT delay), and the slice duration is dynamically adjusted.
[0053] Construct multiple bitrate ladders (for example, from 300 kbps to 8 Mbps), and each ladder corresponds to a different video quality level. The available bitrate options are provided to the student side through the MPD (Media Presentation Description) file, and the student side selects the optimal bitrate based on the local buffer situation and the network prediction algorithm.
[0054] An error correction strategy that combines forward error correction (FEC) and automatic repeat request (ARQ) can be adopted. For example, when the network packet loss rate is lower than 5%, FEC is preferentially used, and when the packet loss rate is higher than 5%, the ARQ mechanism is triggered to ensure the complete transmission of the video stream. At the same time, the transmission efficiency can also be optimized through the TCP BBR congestion control algorithm to reduce the average transmission delay.
[0055] The above teaching interaction method based on a virtual digital human, through semantic segmentation and differential encoding of the student video stream according to the face, gesture, and background regions, dynamically adjusts the bitrate allocation of each region by combining the attention weights generated by eye tracking, and drives the head turning and 3D parallax rendering of the virtual digital human based on the pupil coordinate mapping. At the same time, adaptive slice distribution is carried out according to the network state, achieving the technical effects of accurate perception of the attention of multiple students in the teaching scenario, synchronous response of the virtual image actions and real intentions, and on-demand optimization of network resources, significantly improving the authenticity of immersive interaction and the utilization rate of system resources.
[0056] In an optional embodiment, S1 includes the following steps:
[0057] S11: Process each frame of the original video stream using a semantic segmentation model to generate a face region mask, a gesture region mask, and a background region mask.
[0058] Specifically, a semantic segmentation model in deep learning (such as the U-Net architecture based on a convolutional neural network) can be used to perform pixel-level classification on each frame of the original video stream. The model learns the feature representations of the face, gesture, and background regions through training and can accurately distinguish and generate the corresponding masks. The input of the semantic segmentation model is each frame image of the original video stream, and the output is three binary masks: a face region mask, a gesture region mask, and a background region mask.
[0059] Among them, the model consists of an encoder (feature extraction) and a decoder (feature reconstruction). The encoder usually uses a pre-trained convolutional neural network (such as ResNet or VGG), and the decoder restores the spatial resolution of the image through upsampling and skip connections. The training data of the model includes video frame samples labeled with face, gesture, and background regions, and the training objective is to minimize the cross-entropy loss between the predicted mask and the true mask.
[0060] S12: Configure a lossless coding mode for the pixels covered by the face region mask, a motion compensation coding mode for the pixels covered by the gesture region mask, and a compression coding mode for the pixels covered by the background region mask.
[0061] Specifically, according to the pixel region covered by the mask, different coding modes are adopted for different regions:
[0062] 1) Lossless coding mode: Losslessly code the pixels covered by the facial region mask to ensure high fidelity of facial details. It can be implemented using Huffman coding or arithmetic coding.
[0063] 2) Motion compensation coding mode: Perform motion compensation coding on the pixels covered by the gesture region mask to capture the dynamic changes of gestures. It can be implemented using block matching algorithms (such as block matching and 3D search, BM3D) for motion estimation and compensation.
[0064] 3) Compression coding mode: Perform compression coding on the pixels covered by the background region mask to reduce the data volume. It can be implemented using DCT transform and quantization.
[0065] S13: Generate a hybrid coded video stream based on the pixels of each region mask after partition coding.
[0066] Specifically, reorganize the pixel data after partition coding into a video stream to ensure that the coded data of different regions are synchronized on the time axis. Package the coded data of different regions into video frames through a container format (such as MP4 or MKV), and add timestamps and region identifiers. The container format supports multi-stream multiplexing to ensure the integrity and transportability of the video stream.
[0067] In an alternative embodiment, S2 includes the following steps:
[0068] S21: Perform eye movement tracking on the video frames covered by the facial region to generate a sequence of pupil center coordinates.
[0069] Specifically, use eye movement tracking technology to analyze the video frames covered by the facial region in real time and extract the pupil center coordinates. Eye movement tracking is based on the principle of infrared reflection or visible light imaging technology. By identifying the pupil and corneal reflection points, the position of the pupil center is calculated. The generation of the sequence of pupil center coordinates is achieved by tracking the pupil center positions of consecutive frames.
[0070] S22: Calculate the standard deviation of the movement trajectories of a preset number of consecutive frames based on the sequence of pupil center coordinates as the eye movement feature.
[0071] Specifically, calculate the standard deviation of the movement trajectories of a preset number of consecutive frames based on the sequence of pupil center coordinates. The standard deviation reflects the change amplitude of the pupil center position. The smaller the standard deviation, the more stable the pupil position and the more concentrated the attention. The calculation formula is: where x i is the pupil center coordinate, is the average coordinate, and N is the preset number of consecutive frames.
[0072] S23: When the standard deviation of the movement trajectory is lower than the preset standard deviation threshold, mark the facial region as a high attention region and mark the student corresponding to the high attention region as a high attention student.
[0073] Specifically, when the standard deviation of the movement trajectory is lower than the preset standard deviation threshold, it indicates that the student's attention is highly concentrated, and the corresponding facial area is marked as a high-concern area. The preset standard deviation threshold is determined through experimental calibration to ensure accurate discrimination between high-concern and low-concern states. The marking of high-concern students is achieved by updating the student status table, which records the attention status of each student and the corresponding facial area information.
[0074] In an alternative embodiment, the calculation formula for the attention weight is:
[0075]
[0076] where W is the attention weight, σ is the standard deviation of the movement trajectory, and σ max is the maximum trajectory standard deviation threshold; α is dynamically adjusted according to the type of teaching stage, representing the trajectory stability index; β is negatively correlated with network latency, representing the time decay factor; T is the duration of the high-concern area, Δx is the horizontal difference between the pupil center coordinate and the screen center, and x max is the maximum horizontal offset, and γ is the spatial weight coefficient set according to the virtual digital human rendering field of view angle.
[0077] Specifically, σ (standard deviation of the movement trajectory) is obtained through eye movement tracking, representing the degree of dispersion of the pupil center coordinate in consecutive N frames; the smaller σ is, the more stable the student's line of sight focus (high-concern state), and the positive contribution of the weight is enhanced.
[0078] σ max (the maximum trajectory standard deviation threshold) can be dynamically calculated according to the screen resolution and camera focal length; when σ ≥ σ max , it is determined that the line of sight is wandering (low-concern), and the first term of the formula is zeroed, forcing the weight to decrease.
[0079] α (trajectory stability index) is dynamically adjusted according to the type of teaching stage. For example, α = 2 in the teaching stage and α = 1 in the practice stage; the influence of trajectory stability is strengthened in the teaching stage to suppress the interference of short-term distractions.
[0080] T (duration of the high-concern area) is the cumulative time of marking the high-concern area; through the exponential decay term [1 - e -βT , the non-linear cumulative effect of the duration on the weight is realized, avoiding the weight saturation delay caused by linear growth.
[0081] β (time decay factor) is negatively correlated with network latency (for example, ), the larger the RTT (poor network), the slower the decay; the contribution of the effective attention duration is extended when the network is poor to compensate for the impact of transmission latency on the interaction experience.
[0082] Δx (horizontal offset) is the difference between the pupil center coordinate and the screen center; the greater the offset, it indicates that the student may be looking at the content at the edge of the screen (such as blackboard notes), and the weight needs to be increased to trigger the virtual digital human's head to turn.
[0083] x max (Maximum horizontal offset) is used to normalize the offset and prevent deviation in weight calculation for devices with different resolutions.
[0084] γ (spatial weight coefficient) is set according to the rendering field of view (FOV) of the virtual digital human (for example, γ = 0.3 × FOV / 60°); the larger the FOV, the wider the viewing angle that the virtual digital human needs to cover, and the corresponding proportion of the weight influence of the offset increases.
[0085] In an optional embodiment, adjusting the partition bitrate of the hybrid encoded video stream according to the attention weight includes the following steps:
[0086] For students with an attention weight higher than the first weight threshold, reduce the QP value of the face region from the first QP value to the second QP value.
[0087] For students with an attention weight lower than the first weight threshold and higher than the second weight threshold, insert bidirectional prediction frames into the gesture region and shorten the GOP length to a preset value.
[0088] For students with an attention weight lower than the second weight threshold, increase the compression ratio of the background region from the first compression ratio to the second compression ratio;
[0089] wherein, the first weight threshold is greater than the second weight threshold.
[0090] Specifically, for students with an attention weight higher than the first weight threshold, reduce the quantization parameter (QP value) of the face region from the first QP value to the second QP value. The lower the QP value, the higher the video quality and the clearer the facial details.
[0091] For students with an attention weight lower than the first weight threshold and higher than the second weight threshold, insert bidirectional prediction frames into the gesture region and shorten the GOP (Group of Pictures) length to a preset value. Bidirectional prediction frames can capture the dynamic changes of gestures, and shortening the GOP length can reduce prediction errors and improve the video fluency.
[0092] For students with an attention weight lower than the second weight threshold, increase the compression ratio of the background region from the first compression ratio to the second compression ratio. The higher the compression ratio, the smaller the data volume, but the video quality will decrease to some extent. By increasing the compression ratio of the background region, the data volume can be effectively reduced, thus saving bandwidth.
[0093] The entire partition bitrate adjustment process adopts a dynamic adjustment mechanism to optimize the partition bitrate of the video stream in real time according to the student's attention weight. The dynamic adjustment mechanism ensures the transmission quality and interaction effect of the video stream by monitoring the network status and the student's attention status in real time.
[0094] In an alternative embodiment, S4 includes the following steps:
[0095] S41: Calculate the coordinate mapping error compensation amount according to the QP value of the face region. The calculation formula for the coordinate mapping error compensation amount is:
[0096]
[0097] where δ is the coordinate mapping error compensation amount, I QP is the QP value of the face region, X max is the horizontal screen resolution, and f is the camera focal length.
[0098] S42: If the frame type of the gesture region is a bidirectional prediction frame, set the value of the action coherence factor to the first value; if the frame type of the gesture region is a forward prediction frame, set the value of the action coherence factor to the second value.
[0099] S43: Calculate the horizontal head rotation angle based on the pupil center coordinates, the coordinate mapping error compensation amount, and the action coherence factor. The calculation formula for the horizontal head rotation angle is:
[0100]
[0101] where θ is the horizontal head rotation angle, x is the abscissa of the pupil center coordinates, x center is the horizontal center coordinate of the screen, and λ is the action coherence factor.
[0102] S44: Quantize the horizontal head rotation angle into facial animation parameters through a 3D rendering engine.
[0103] Specifically, the horizontal head rotation angle is quantized into facial animation parameters through a 3D rendering engine. The quantization process is implemented through an interpolation algorithm and skeletal animation technology. The facial animation parameters include the head turning angle, facial expressions, and eye gaze direction, and these parameters are updated in real time through the 3D rendering engine to ensure that the actions of the virtual digital human are synchronized with the student's attention status.
[0104] The above teaching interaction method based on a virtual digital human intelligently segments the student video stream into facial, gesture, and background regions and implements a differential coding strategy. It constructs a non-linear attention weight model based on the stability, duration, and spatial offset characteristics of eye movement trajectories, dynamically adjusts the coding parameters and network resource allocation. At the same time, it generates physiologically realistic virtual digital human head turning parameters based on the pupil coordinate positioning error compensation and frame type action coherence factor. Combining multi-viewpoint rendering and network-aware adaptive slicing technology, it realizes accurate recognition of the attention of multiple students in the teaching scenario, millisecond-level synchronization of virtual image actions and real intentions, and efficient coordinated allocation of cross-terminal resources, significantly improving the naturalness and system robustness of immersive teaching interaction.
[0105] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps in other steps.
[0106] Based on the same inventive concept, an embodiment of the present application also provides an apparatus for implementing the above-mentioned teaching interaction method based on a virtual digital human. The solution provided by this apparatus to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following teaching interaction apparatus based on a virtual digital human can refer to the limitations on a teaching interaction method based on a virtual digital human in the above text, and will not be repeated here.
[0107] In an exemplary embodiment, as Figure 2 shown, a teaching interaction apparatus 20 based on a virtual digital human is provided, including:
[0108] A video stream processing module 21, configured to collect the original video streams of multiple students, divide each original video stream into a facial region, a gesture region, and a background region; and perform partition coding on the original video stream based on the facial region, the gesture region, and the background region to generate a mixed coded video stream.
[0109] An attention area analysis module 22, configured to screen out high-attention areas from the facial region according to the eye movement characteristics of the students, and mark the students corresponding to the high-attention areas as high-attention students.
[0110] The video stream optimization module 23 is configured to generate attention weights based on the duration of the high - attention region, adjust the partition bitrate of the hybrid - encoded video stream according to the attention weights, and generate an optimized video stream.
[0111] The animation parameter generation module 24 is configured to map the pupil center coordinates of the highly - concerned student to facial animation parameters based on the QP value of the facial region and the frame type of the gesture region in the optimized video stream.
[0112] The 3D video generation module 25 is configured to generate the turning angle of the virtual digital human head based on the facial animation parameters; perform 3D parallax rendering on the turning angle according to the multi - view encoding standard to generate a multi - view 3D video stream.
[0113] The adaptive video transmission module 26 is configured to perform adaptive slicing on the multi - view 3D video stream to generate an adaptive bitrate video stream, and send the adaptive bitrate video stream to the student side; wherein, the adaptive slicing process includes dynamically adjusting the slice duration according to the network state.
[0114] Optionally, the video stream processing module 21 includes:
[0115] The semantic segmentation unit 211 is configured to process each frame of the original video stream using a semantic segmentation model to generate a facial region mask, a gesture region mask, and a background region mask.
[0116] The encoding mode configuration unit 212 is configured to configure a lossless encoding mode for the pixels covered by the facial region mask, configure a motion - compensated encoding mode for the pixels covered by the gesture region mask, and configure a compression encoding mode for the pixels covered by the background region mask.
[0117] The hybrid - encoding generation unit 213 is configured to generate a hybrid - encoded video stream according to the pixels of each region mask after partition encoding.
[0118] Optionally, the attention region analysis module 22 includes:
[0119] The eye movement tracking unit 221 is configured to perform eye movement tracking on the video frames covered by the facial region to generate a sequence of pupil center coordinates.
[0120] The feature calculation unit 222 is configured to calculate the standard deviation of the movement trajectories of a preset number of consecutive frames based on the sequence of pupil center coordinates as the eye movement feature.
[0121] The attention region determination unit 223 is configured to mark the facial region as a high - attention region and mark the student corresponding to the high - attention region as a highly - concerned student when the standard deviation of the movement trajectories is lower than a preset standard deviation threshold.
[0122] Optionally, the video stream optimization module 23 includes an attention weight calculation unit 231 for calculating the attention weight according to the following formula:
[0123]
[0124] where W is the attention weight, σ is the standard deviation of the movement trajectory, and σ max is the maximum trajectory standard deviation threshold; α is dynamically adjusted according to the type of teaching stage, representing the trajectory stability index; β is negatively correlated with the network delay, representing the time decay factor; T is the duration of the high attention area, Δx is the horizontal difference between the pupil center coordinate and the screen center, and x max is the maximum horizontal offset, and γ is the spatial weight coefficient set according to the virtual digital human rendering field of view angle.
[0125] Optionally, the video stream optimization module 23 further includes:
[0126] A high-attention student optimization unit 232 for reducing the QP value of the facial area from a first QP value to a second QP value for students with an attention weight higher than the first weight threshold.
[0127] A medium-attention student optimization unit 233 for inserting bidirectional prediction frames into the gesture area and shortening the GOP length to a preset value for students with an attention weight lower than the first weight threshold and higher than the second weight threshold.
[0128] A low-attention student optimization unit 234 for increasing the compression ratio of the background area from a first compression ratio to a second compression ratio for students with an attention weight lower than the second weight threshold.
[0129] where the first weight threshold is greater than the second weight threshold.
[0130] Optionally, the animation parameter generation module 24 includes:
[0131] An error compensation calculation unit 241 for calculating the coordinate mapping error compensation amount according to the QP value of the facial area. The calculation formula for the coordinate mapping error compensation amount is:
[0132]
[0133] where δ is the coordinate mapping error compensation amount, I QP is the QP value of the facial area, X max is the screen horizontal resolution, and f is the camera focal length.
[0134] An action coherence factor setting unit 242 is configured to set the value of the action coherence factor to a first value if the frame type of the gesture area is a bidirectional prediction frame; and set the value of the action coherence factor to a second value if the frame type of the gesture area is a forward prediction frame.
[0135] A head rotation angle calculation unit 243 is configured to calculate a head horizontal rotation angle based on the pupil center coordinates, the coordinate mapping error compensation amount, and the action coherence factor; the calculation formula for the head horizontal rotation angle is:
[0136]
[0137] where θ is the head horizontal rotation angle, x is the abscissa of the pupil center coordinates, x center is the screen horizontal center coordinate, and λ is the action coherence factor.
[0138] An animation parameter quantization unit 244 is configured to quantize the head horizontal rotation angle into facial animation parameters through a 3D rendering engine.
[0139] An embodiment of the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps in the foregoing method embodiments are implemented.
[0140] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the foregoing method embodiments are implemented.
[0141] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0142] The above embodiments only represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the application embodiments. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.
Claims
1. A teaching interaction method based on virtual digital humans, characterized in that, The method includes: S1: Collect the original video streams of multiple students, divide each of the original video streams into a face region, a gesture region, and a background region; based on the face region, the gesture region, and the background region, perform partition encoding on the original video stream to generate a mixed encoded video stream; S2: According to the eye movement characteristics of the students, screen out the high - attention regions from the face region, and mark the students corresponding to the high - attention regions as high - attention students; S3: Generate an attention weight based on the duration of the high - attention region, and adjust the partition bitrate of the mixed encoded video stream according to the attention weight to generate an optimized video stream; S4: Based on the QP value of the face region and the frame type of the gesture region in the optimized video stream, map the pupil center coordinates of the high - attention students to facial animation parameters; S5: Generate a virtual digital human head turning angle based on the facial animation parameters; perform 3D parallax rendering on the head turning angle according to the multi - viewpoint encoding standard to generate a multi - viewpoint 3D video stream; S6: Perform adaptive slicing processing on the multi - viewpoint 3D video stream to generate an adaptive bitrate video stream, and send the adaptive bitrate video stream to the student side; wherein, the adaptive slicing processing includes dynamically adjusting the slice duration according to the network state.
2. The method according to claim 1, wherein The S1 includes: S11: Use a semantic segmentation model to process each frame of the original video stream to generate a face region mask, a gesture region mask, and a background region mask; S12: Configure a lossless encoding mode for the pixels covered by the face region mask, configure a motion - compensated encoding mode for the pixels covered by the gesture region mask, and configure a compression encoding mode for the pixels covered by the background region mask; S13: Generate the mixed encoded video stream according to the pixels of each region mask after partition encoding.
3. The method according to claim 1, wherein The S2 includes: S21: Perform eye movement tracking on the video frames covered by the face region to generate a sequence of pupil center coordinates; S22: Calculate the standard deviation of the movement trajectories of a preset number of consecutive frames based on the sequence of pupil center coordinates as the eye movement characteristics; S23: When the standard deviation of the movement trajectory is lower than a preset standard deviation threshold, mark the face region as the high - attention region, and mark the student corresponding to the high - attention region as a high - attention student.
4. The method according to claim 3, characterized in that, The calculation formula of the attention weight is: Among them, W is the attention weight, σ is the standard deviation of the movement trajectory, and σ max is the maximum trajectory standard deviation threshold; α is dynamically adjusted according to the type of teaching stage, representing the trajectory stability index; β is negatively correlated with network latency, representing the time decay factor; T is the duration of the high-attention area, Δx is the horizontal difference between the pupil center coordinate and the screen center, and x max is the maximum horizontal offset, and γ is the spatial weight coefficient set according to the rendering field of view angle of the virtual digital human.
5. The method according to claim 1, wherein The adjustment of the partition bitrate of the mixed encoded video stream according to the attention weight includes: For students with an attention weight higher than the first weight threshold, reduce the QP value of the face region from the first QP value to the second QP value; For students with an attention weight lower than the first weight threshold and higher than the second weight threshold, insert a bidirectional prediction frame into the gesture region and shorten the GOP length to a preset value; For students with an attention weight lower than the second weight threshold, increase the compression ratio of the background region from the first compression ratio to the second compression ratio; Wherein, the first weight threshold is greater than the second weight threshold.
6. The method according to any one of claims 1 to 5, characterized in that, The S4 includes: S41: Calculate the coordinate mapping error compensation amount according to the QP value of the facial region. The calculation formula of the coordinate mapping error compensation amount is: where δ is the coordinate mapping error compensation amount, I QP is the QP value of the facial region, X max is the horizontal screen resolution, and f is the camera focal length; S42: If the frame type of the gesture region is a bidirectional prediction frame, set the value of the action coherence factor to the first value; if the frame type of the gesture region is a forward prediction frame, set the value of the action coherence factor to the second value; S43: Calculate the horizontal head rotation angle based on the pupil center coordinates, the coordinate mapping error compensation amount, and the action coherence factor. The calculation formula of the horizontal head rotation angle is: Where, θ is the horizontal rotation angle of the head, x is the abscissa of the pupil center coordinate, and x center is the horizontal center coordinate of the screen, and λ is the motion coherence factor; S44: Quantize the horizontal head rotation angle into the facial animation parameter through a 3D rendering engine.
7. A teaching interaction device based on a virtual digital human, characterized in that The device includes: A video stream processing module, configured to collect the original video streams of multiple students, divide each of the original video streams into a facial region, a gesture region, and a background region; based on the facial region, the gesture region, and the background region, perform partition encoding on the original video streams to generate a mixed encoded video stream; A region of interest analysis module, configured to screen out a region of high interest from the facial region according to the eye movement characteristics of the student, and mark the student corresponding to the region of high interest as a student with high interest; A video stream optimization module, configured to generate an attention weight based on the duration of the region of high interest, and adjust the partition bit rate of the mixed encoded video stream according to the attention weight to generate an optimized video stream; An animation parameter generation module, configured to map the pupil center coordinates of the student with high interest to the facial animation parameter based on the QP value of the facial region and the frame type of the gesture region in the optimized video stream; A 3D video generation module, configured to generate a virtual digital human head turning angle based on the facial animation parameter; perform 3D parallax rendering on the head turning angle according to the multi-viewpoint encoding standard to generate a multi-viewpoint 3D video stream; An adaptive video transmission module, configured to perform adaptive slicing processing on the multi-viewpoint 3D video stream to generate an adaptive bit rate video stream, and send the adaptive bit rate video stream to the student side; wherein, the adaptive slicing processing includes dynamically adjusting the slice duration according to the network state.
8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 6.
Citation Information
Cited By
Digital human coding method and device based on region of interest and long-term reference frame
CN121644814A
Digital human display and interaction system based on multiple clients
CN121711503A