Method for converting traditional videos into interactive automatic narration by artificial intelligence digital person

By performing multimodal deconstruction and temporal synchronization processing on traditional videos, a narration script is generated and dynamically expressed, solving the problems of inaccurate synchronization and insufficient interactivity of traditional video content, and achieving efficient user interaction and contextual supplementary narration.

CN121126039BActive Publication Date: 2026-02-03BEIJING MENGKE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511658063.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-03
Estimated Expiration
2045-11-13

AI Technical Summary

Technical Problem

Existing technologies and traditional video content cannot achieve precise synchronization between digital humans and video content, resulting in insufficient user interactivity and the inability to dynamically generate supplementary explanations based on user query intentions, leading to a fragmented and inconsistent interactive experience.

Method used

By performing multimodal deconstruction on the original video, extracting visual and semantic elements, generating a narration script and adding timestamps, the timing of video playback and narration content is synchronized; receiving user interruption requests, performing semantic matching and generating supplementary narration content, and driving the virtual avatar generator to dynamically express and output content.

Benefits of technology

It achieves precise synchronization between digital human narration and video content, improving the user viewing experience and content comprehension efficiency, supporting intelligent response to user query intent and contextual supplementation, and improving the interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121126039B_ABST
    Figure CN121126039B_ABST
Patent Text Reader

Abstract

The application provides an interactive automatic explanation method for converting traditional videos into artificial intelligence digital persons, relates to the technical field of artificial intelligence, and comprises the following steps: obtaining original video data and an audio track, performing semantic analysis on the audio track to obtain multimodal deconstruction data; generating an explanation script for each time period of the video based on the explanation text, and performing timestamp annotation on visual elements to form a time sequence synchronization data structure; driving a virtual image generator to synthesize a dynamic expression output of the digital person in real time according to a current playing time point during the playing process; after receiving a user interruption request, performing semantic matching on the query intention in the explanation script, positioning a target explanation segment and visual elements, and generating supplementary explanation content; driving the virtual image generator to synthesize a dynamic output synchronized with the supplementary explanation, and resuming playing or jumping to a specified time point according to a user instruction after the interaction is completed. The application realizes the conversion of traditional videos into interactive intelligent explanation videos, and improves the user watching experience and learning efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to artificial intelligence technology, and more particularly to a method for converting traditional videos into interactive, automatic explanations by an AI digital human. Background Technology

[0002] With the rapid development of artificial intelligence and digital media technologies, the intelligent presentation and interactive dissemination of video content have gradually become important needs in fields such as online education, corporate training, and knowledge sharing. Traditional video content is usually presented in a one-way playback format, where users can only passively receive information and cannot engage in in-depth interaction. In recent years, the maturity of virtual digital human technology has provided new possibilities for the intelligent explanation of video content. By combining digital humans with video content, users can be provided with a more vivid and personalized viewing experience. However, how to transform massive amounts of traditional video resources into interactive video content with intelligent explanation capabilities remains a significant challenge facing the current technological field.

[0003] Existing video processing technologies primarily focus on video compression, transcoding, or simple content tag extraction, lacking the precise establishment of deep semantic connections between video content and narration text. This results in difficulty in achieving accurate temporal synchronization between digital human narration and video content, impacting the user's viewing experience and comprehension. Current digital human narration systems typically play pre-recorded content, unable to dynamically generate supplementary narration based on the user's real-time query intent. When a user has questions about a specific segment or concept in the video, the system cannot provide targeted, in-depth answers, leading to severely insufficient interactivity. When handling user interruption requests, existing technologies often only pause video playback or jump to a preset time point, failing to intelligently identify the semantic relationship between the user's query intent and the video content, let alone automatically complete relevant knowledge points based on contextual logic. This results in a fragmented and incoherent interactive experience.

[0004] The technical problem to be solved by this invention is to provide a method for converting traditional videos into interactive and automatic explanations of artificial intelligence digital humans. By performing multimodal deconstruction on the original video and establishing a temporal synchronization relationship between visual elements and explanation content, the method achieves accurate synchronous explanation between the digital human and the video content. At the same time, it supports users to interrupt and query and intelligently generate supplementary explanation content, thereby effectively solving the problems of low temporal synchronization accuracy, insufficient intelligent interactive response, and lack of contextual supplementary explanation in the prior art. Summary of the Invention

[0005] This invention provides a method for converting traditional videos into interactive, automatically narrated AI digital humans, which can solve the problems in the prior art.

[0006] A first aspect of the present invention provides a method for converting traditional videos into interactive, automatically narrating AI-powered digital humans, comprising:

[0007] The original video data and its corresponding audio track are acquired. The original video data is decomposed temporally to extract the video frame sequence and identify the main objects and background areas. At the same time, the audio track is semantically parsed to obtain multimodal deconstructed data containing visual and semantic elements.

[0008] Based on the explanatory text in the multimodal deconstruction data, a corresponding explanatory script is generated for each time period of the video content, and the visual elements in the video frame sequence corresponding to each explanatory segment are timestamped to form a data structure that synchronizes video playback with the explanatory content in time.

[0009] During video playback, the virtual avatar generator synthesizes the dynamic expression output of the digital human in real time according to the narration script corresponding to the current playback time point, so that the digital human can narrate synchronously with the video content;

[0010] The system receives an interruption request from the user, which includes a query intent and control instructions for the content to be explained. The query intent is semantically matched in the explanation script to locate the target explanation segment and its associated visual elements related to the query intent, and supplementary explanation content is generated based on the context of the target explanation segment.

[0011] Based on the visual elements associated with the target explanation segment, the virtual avatar generator is driven to synthesize a dynamic expression output synchronized with the supplementary explanation content. After the interactive explanation is completed, the video playback and digital avatar synchronized explanation are resumed according to the user's instructions, or the playback is redirected to the video time point specified by the user to continue.

[0012] The original video data is temporally decomposed to extract video frame sequences and identify the main objects and background regions within them. Simultaneously, the audio track is semantically parsed to obtain multimodal deconstructed data containing both visual and semantic elements, including:

[0013] The original video data is segmented into a video frame sequence according to the timeline, and each frame in the video frame sequence is spatially divided to distinguish the foreground region and the background region; the main object in the foreground region is detected and its attributes are extracted. The attribute extraction includes identifying the contour boundary features, pose features and motion trajectory features of the main object, and establishing a cross-frame tracking identifier for the main object.

[0014] Acoustic signal processing is performed on the audio track to convert continuous speech signals into discrete text segments. Natural language understanding is performed on the discrete text segments to extract semantic units and their syntactic structures. The discrete text segments are timestamped with the video frame sequence. Based on the cross-frame tracking identifier, the main object is located in the aligned video frames. The visual state of the main object when it emits the corresponding semantic unit is determined, and the binding relationship between the semantic unit and the visual state is established.

[0015] The contour boundary features, pose features, motion trajectory features, and scene features of the background region are summarized as visual elements, the semantic units, syntactic structures, and timestamp information are summarized as semantic elements, and the binding relationship is used as the association index between the visual elements and the semantic elements to form multimodal deconstruction data.

[0016] Based on the explanatory text in the multimodal deconstruction data, a corresponding explanatory script is generated for each time segment of the video content, and the visual elements in the video frame sequence corresponding to each explanatory segment are timestamped to form a data structure that synchronizes video playback with the explanatory content in time, including:

[0017] The explanatory text is extracted from the multimodal deconstruction data, and the explanatory text is divided into multiple explanatory segments according to the timeline of video playback. Each explanatory segment corresponds to a continuous time period in the video, and each explanatory segment is marked with a start timestamp and an end timestamp.

[0018] Semantic analysis is performed on each of the explanatory segments to identify important information points and key points of explanation. Based on the important information points, extended explanatory content is generated, and the original explanatory text and the extended explanatory content are integrated to form a complete explanatory script.

[0019] The video frame sequence and its associated visual elements are obtained from the multimodal deconstruction data. The binding relationship in the multimodal deconstruction data is used to determine the timestamp range corresponding to each narration segment. The corresponding visual elements are extracted from the video frame sequence within the timestamp range.

[0020] The narration script, the visual elements, and the video frame sequence are aligned in three dimensions according to timestamps to construct a synchronized data structure that includes a video playback timeline, a narration content timeline, and a visual element timeline, ensuring that video playback and digital human narration are synchronized in time.

[0021] The system receives a user-input interruption request, which includes a query intent and control instructions regarding the content to be explained. It performs semantic matching of the query intent within the explanation script to locate the target explanation segment and its associated visual elements related to the query intent. Based on the contextual relationships of the target explanation segment, it generates supplementary explanation content, including:

[0022] During video playback, the system monitors the user input channel in real time. When an interruption request is detected, the system records the current video playback time as the interruption timestamp and sends a pause command to the video playback module and the digital human narration module, so that the video playback and digital human narration are paused synchronously.

[0023] The interruption request is semantically parsed to extract the core concepts in the query intent. The core concepts are then semantically similar to the explanatory segments in the explanatory script to locate the target explanatory segment related to the query intent and obtain the visual elements mapped to the target explanatory segment.

[0024] Based on the position of the target explanation segment on the timeline, the explanation script is expanded forward and backward to extract context-related explanation segments. By analyzing the semantic relationship between the target explanation segment and adjacent explanation segments, a context explanation sequence containing cause and effect is constructed.

[0025] The contextual explanation sequence is verified for completeness. If any missing information is detected, a bridging explanation segment that can fill in the missing information is searched from other time periods of the explanation script. The bridging explanation segment and its mapped visual elements are added to the contextual explanation sequence to form a logically complete supplementary explanation content.

[0026] Based on the position of the target explanation segment on the timeline, context-related explanation segments are extracted by extending the explanation script forward and backward. By analyzing the semantic relationships between the target explanation segment and adjacent explanation segments, a contextual explanation sequence containing cause and effect is constructed, including:

[0027] Using the start timestamp of the target explanation segment as a reference point, the explanation script is backtracked forward along the timeline to extract adjacent explanation segments with timestamps earlier than the target explanation segment. The semantic dependencies between the adjacent explanation segments and the target explanation segment are analyzed, and the preceding explanation segments that serve as prerequisites for the target explanation segment are identified to form a forward context sequence.

[0028] Using the end timestamp of the target explanation segment as a reference point, a forward search is performed in the explanation script to the back of the timeline to extract adjacent explanation segments with timestamps later than the target explanation segment. The semantic continuity relationship between the adjacent explanation segments and the target explanation segment is analyzed, and the subsequent explanation segments derived from the target explanation segment are identified to form a backward context sequence.

[0029] The forward context sequence and the backward context sequence are integrated in timestamp order to construct a contextual explanation sequence centered on the target explanation segment.

[0030] Based on the visual elements associated with the target explanation segment, the virtual avatar generator is driven to synthesize a dynamic expression output synchronized with the supplementary explanation content. After completing the interactive explanation, video playback and synchronized digital avatar explanation are restored according to user instructions, including:

[0031] The associated visual elements are obtained from the target explanation segment and the context explanation sequence. The pose features and facial expression features in the visual elements are modeled as inter-frame motion trajectories. The original feature frames and intermediate transition frames are combined in time sequence to form a smooth motion feature sequence.

[0032] The virtual avatar generator synthesizes a dynamic expression output based on the feature sequence. The dynamic expression output includes a continuous frame sequence of the virtual avatar. The energy function is calculated for the pose features of each frame in the continuous frame sequence, and the frames whose energy function values ​​reach a local maximum are marked as visual anchor points.

[0033] A temporal anchor point correspondence table is constructed. By performing prosodic intensity analysis on the supplementary explanation content, the syllable stress position and semantic key words are extracted as semantic anchor points. The semantic anchor points and the visual anchor points are matched with semantic relevance scores, and the time difference between the semantic anchor points and the visual anchor points is calculated.

[0034] Based on the time difference, the continuous frame sequence is segmented into nonlinear time mappings to align the visual anchor point with the semantic anchor point on the time axis. The aligned continuous frame sequence is then output synchronously with the supplementary explanation content to complete the interactive explanation.

[0035] After the interactive explanation is completed, the system receives a user's command to resume playback or jump to another location. If a command to resume playback is received, the system resumes video playback from the video position corresponding to the interrupted timestamp and drives the digital human to continue synchronous explanation according to the explanation script corresponding to the resumed position. If a command to jump to another location is received, the system extracts the target time point specified in the jump command, jumps the video playback position to the target time point, and drives the digital human to start synchronous explanation according to the explanation script corresponding to the target time point.

[0036] The virtual avatar generator synthesizes a dynamic expression output based on the feature sequence. The dynamic expression output includes a continuous frame sequence of the virtual avatar. Energy functions are calculated for the pose features of each frame in the continuous frame sequence, and frames whose energy function values ​​reach local maxima are marked as visual anchor points.

[0037] The virtual avatar generator synthesizes a dynamic expression output based on the feature sequence. The dynamic expression output includes a continuous frame sequence of the virtual avatar. Feature decomposition is performed on the posture feature data carried in each frame of the continuous frame sequence, and the joint angle vector and limb end displacement vector of the virtual avatar are extracted from the posture feature data of each frame.

[0038] A posture energy function is constructed based on the joint angle vector and the limb end displacement vector. The angle change rate of the joint angle vector and the displacement amplitude of the limb end displacement vector between adjacent frames are calculated based on the posture energy function. The angle change rate and the displacement amplitude are weighted and combined to obtain the motion intensity value corresponding to each frame.

[0039] The motion intensity value is searched for local extrema on the time axis. By comparing the motion intensity value of each frame with the motion intensity value of its neighboring frames, the frame position where the motion intensity value presents a local maximum value on the time series is identified, and the frame where the motion intensity value reaches a local maximum value is marked as a visual anchor frame.

[0040] A second aspect of the present invention provides an electronic device, comprising:

[0041] processor;

[0042] Memory used to store processor-executable instructions;

[0043] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0044] A third aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0045] The beneficial effects of this application are as follows:

[0046] This invention extracts video frame sequences by temporal decomposition of raw video data and identifies the main object and background areas. Simultaneously, it performs semantic parsing of the audio track to obtain multimodal deconstructed data. Furthermore, it timestamps the visual elements in the video frame sequences with each narration segment, forming a data structure that synchronizes video playback with the narration content in a timely manner. This multimodal deconstruction and temporal annotation mechanism ensures precise alignment between the digital human's narration and the video content on the timeline, effectively solving the problem of low temporal synchronization accuracy between narration content and visual elements in existing technologies. This significantly improves the efficiency of content comprehension and the quality of the immersive experience for users.

[0047] This invention receives user interruption requests and performs semantic matching of the query intent within the narration script. This allows for precise location of the target narration segment and its associated visual elements related to the query intent, and automatically generates supplementary narration content based on the context of the target narration segment. This intelligent semantic matching and context supplementation mechanism overcomes the limitation of existing technologies where digital humans can only play pre-recorded content. It enables the dynamic generation of targeted narrations based on actual user needs, significantly improving the system's interactive intelligence and the accuracy of user query responses.

[0048] This invention drives a virtual avatar generator to synthesize dynamic, synchronized outputs in real time based on supplementary explanations. After the interactive explanation is complete, users can resume video playback or jump to a specified time point to continue playback, thus constructing a complete interactive loop. This dynamic response and flexible control mechanism effectively solves the problems of fragmented interactive experience and poor contextual coherence in existing technologies when users interrupt their queries. It enables users to obtain in-depth knowledge supplementation without interrupting the learning process, significantly improving the user experience and knowledge dissemination effect of interactive video explanation systems. Attached Figure Description

[0049] Figure 1 This is a flowchart illustrating a method for converting traditional video into an interactive, automatically narrating AI digital human, according to an embodiment of the present invention.

[0050] Figure 2 This is a flowchart illustrating the generation of supplementary explanatory content based on semantic matching in embodiments of the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0053] Figure 1 This is a flowchart illustrating a method for converting traditional video into an interactive, automatically narrating AI digital human, as described in an embodiment of the present invention. Figure 1 As shown, the method includes:

[0054] The original video data and its corresponding audio track are acquired. The original video data is decomposed temporally to extract the video frame sequence and identify the main objects and background areas. At the same time, the audio track is semantically parsed to obtain multimodal deconstructed data containing visual and semantic elements.

[0055] Based on the explanatory text in the multimodal deconstruction data, a corresponding explanatory script is generated for each time period of the video content, and the visual elements in the video frame sequence corresponding to each explanatory segment are timestamped to form a data structure that synchronizes video playback with the explanatory content in time.

[0056] During video playback, the virtual avatar generator synthesizes the dynamic expression output of the digital human in real time according to the narration script corresponding to the current playback time point, so that the digital human can narrate synchronously with the video content;

[0057] The system receives an interruption request from the user, which includes a query intent and control instructions for the content to be explained. The query intent is semantically matched in the explanation script to locate the target explanation segment and its associated visual elements related to the query intent, and supplementary explanation content is generated based on the context of the target explanation segment.

[0058] Based on the visual elements associated with the target explanation segment, the virtual avatar generator is driven to synthesize a dynamic expression output synchronized with the supplementary explanation content. After the interactive explanation is completed, the video playback and digital avatar synchronized explanation are resumed according to the user's instructions, or the playback is redirected to the video time point specified by the user to continue.

[0059] After extracting the human voice text from the input video, the multimodal deconstruction data processing module initiates the synchronized narration process using a digital human. The narration text undergoes semantic analysis by a natural language processing engine to extract keywords, sentiment characteristics, and intonation features. Semantic analysis employs a pre-trained language model based on a transformer architecture, encoding the text sequence into a 768-dimensional semantic vector representation. The sentiment analysis module identifies the emotional state in the text, including three basic categories: positive, negative, and neutral, and sub-emotional labels such as excitement, calm, serious, and relaxed. Intonation analysis identifies declarative, interrogative, and exclamatory intonation types through fundamental frequency variations and energy distribution characteristics of the audio signal.

[0060] The synchronized narration audio generator synthesizes digital human speech based on the narration text and emotional tags. Speech synthesis employs an end-to-end neural network vocoder; the input text is converted into a phoneme sequence through a phoneme conversion module, with each phoneme corresponding to a specific pronunciation duration and pitch variation. Emotional information influences the prosodic parameters of the speech; positive emotions increase the average pitch by approximately 15%, while negative emotions decrease the speech rate by approximately 20%. The synthesized audio sampling rate is set to 22050Hz, and the quantization bit depth is 16 bits to ensure that the audio quality clarity meets the requirements of real-time playback. The duration of the narration audio is adjusted in real-time according to the video playback progress to maintain a synchronized narration rhythm with the video content.

[0061] The motion information extraction module analyzes the action descriptions and emotional states within the narration text to generate corresponding narration gesture sequences. The motion library contains narration motion elements at three levels: basic gestures, body posture, and movement trajectory. Basic gestures cover common narration gestures such as pointing, emphasizing, and describing; each gesture defines keyframe data for its start position, movement trajectory, and end position. Body posture controls parameters such as head tilt, shoulder height, and torso rotation, adjusting the formality of the posture according to the seriousness of the narration content. Narration trajectories are suitable for scenarios requiring pointing towards video content, defining the path of the digital human's positional changes within the narration space.

[0062] The facial expression generation module uses emotion tags to drive the interpretation of facial expression changes. Facial expressions are encoded using the FACS (Facial Action Coding System), which breaks down complex expressions into basic action unit combinations in areas such as eyebrows, eyes, nose, and mouth. Positive emotions activate the action units of raised corners of the mouth and smiling eyes, while negative emotions correspond to the combination of furrowed brows and downturned corners of the mouth. The intensity of expressions is adjusted based on the confidence level of the emotion analysis; when the confidence level exceeds 0.8, the expression intensity is set to full activation, and when it is below 0.5, the expression intensity is halved. Cubic Bézier curve interpolation is used for expression transitions to ensure the naturalness and smoothness of facial movements during interpretation.

[0063] The virtual avatar generator receives narration audio content, motion information, and facial expression information to synthesize a digital human narration animation synchronized with video playback. The rendering engine employs real-time graphics processing technology, supporting a narration animation playback speed of 30 frames per second. The skeletal animation system controls the narrator's body movements; the skeletal hierarchy includes major branches such as the root bones, spinal chain, limb chains, and finger chains. The skinning weight algorithm transmits skeletal transformations to mesh vertices, achieving realistic muscle deformation effects. Facial animation uses hybrid deformation technology, adjusting the weight values ​​of different expression targets to synthesize complex narration expressions.

[0064] The video playback and narration synchronization mechanism ensures consistency between the digital human's narration and the original video playback. The synchronization controller maintains a global timestamp, tracking video playback and narration progress with millisecond precision. Audio synchronization employs audio frame buffering technology, pre-loading 500 milliseconds of narration audio data into the playback buffer to prevent playback stuttering. Visual synchronization uses a frame rate adaptation algorithm to handle differences in display capabilities across different devices, automatically reducing rendering quality or frame rate on low-performance devices to ensure smooth video and narration playback. The synchronization error detection module continuously monitors the time deviation between video playback and narration, automatically triggering resynchronization when the deviation exceeds 100 milliseconds.

[0065] The narration interruption detection system supports both button operation and voice input for interruption. Button interaction provides control options such as pausing narration, asking questions, and replaying through a graphical user interface, with a button response time of less than 50 milliseconds. Voice interaction employs keyword detection technology, listening for preset wake words such as "pause narration," "question," and "repeat narration." Voice detection uses a lightweight neural network model with fewer than 1MB of parameters, supporting offline operation to reduce network latency. Upon detecting a valid interruption signal, the system immediately stops playing the current narration content, pauses video playback, and records the pause time.

[0066] The video frame and narration content matching module handles cases where the narration text and video content do not match. The content matching analyzer compares the semantic similarity between the narration text description and the corresponding video frame, employing a multimodal similarity calculation method. Text features are encoded into sentence-level vectors using a BERT model, while video features are extracted using a ResNet model to extract image semantic information. Similarity calculation uses a cosine distance metric with a threshold of 0.6; frames below this threshold are marked as having insufficient narration content.

[0067] A generative AI model receives insufficient narration text descriptions and generates video frames that meet the requirements. The generative model employs a diffusion model architecture, capable of generating high-quality image content based on narration text prompts. A text encoder converts the narration text into conditional vectors, guiding the image generation process. The generated images are set to a resolution of 1024x576 pixels, maintaining the same aspect ratio as the original video frames. Generation time is controlled within 3 seconds to meet the performance requirements of real-time narration. The generated video frames undergo style consistency processing to ensure visual style harmony and consistency with the original video.

[0068] The multi-narrator identification and separation system processes video content containing multiple narrators. The voiceprint recognition module uses a deep embedding network to extract the acoustic feature vector of each narrator, with the feature dimension set to 256. A clustering algorithm groups the acoustic features to identify different narrators. Video character detection uses an object detection algorithm to identify facial regions in the frame, and then associates voiceprints and visual identities through facial recognition technology. Each identified narrator is assigned an independent audio track and a corresponding digital narration role.

[0069] The multi-digital human coordinated narration system manages the simultaneous performance of multiple virtual narrator characters. The character scheduler arranges the timing of different digital humans' appearances based on the time distribution of the narration audio tracks. When a narrator character begins speaking, the corresponding digital human activates its narration animation, while other characters remain silent but maintain natural standby actions. The scene layout manager arranges the positional relationships of multiple digital humans within the virtual narration space, preventing character overlap or obstruction. Narration handover is handled with special character switching logic, enhancing the realism of multi-person narration through subtle movements such as eye contact and body turning.

[0070] The proactive interactive content generation module automatically inserts interactive segments based on the duration of the explanation. A duration analyzer tracks the duration of continuous explanations, triggering the interactive insertion mechanism when a single explanation segment exceeds 5 minutes. Interactive content includes Q&A sessions on knowledge points, comprehension level polls, and reinforced demonstrations of key concepts. The Q&A section extracts key information points from the explanation text to generate multiple-choice or short-answer questions. The voting function collects user feedback on the pacing and difficulty of the explanations, adjusting the presentation of subsequent explanation content accordingly.

[0071] The 3D virtual object explanation and demonstration module provides visual assistance for abstract concepts. The 3D model library contains common teaching props, geometric shapes, molecular structures, and other visualization elements. The model selection algorithm matches suitable 3D objects based on keywords in the explanation text, supporting fuzzy matching and semantic similarity search. 3D rendering uses physically based rendering technology to provide realistic lighting and material effects. Users can control the rotation, scaling, and dissection of 3D objects through gestures or voice commands, with a digital narrator providing corresponding explanations.

[0072] The panoramic content description and display system supports 360-degree environmental narration. Panoramic images are stored in an equidistant cylindrical projection format with a resolution of 4096x2048 pixels. The panoramic player supports mouse dragging or gyroscope-controlled viewpoint switching, with a response latency of less than 20 milliseconds. The hotspot annotation function adds interactive information points to the panoramic scene; users can click on hotspots to access detailed text or voice narration.

[0073] The knowledge graph construction module performs structured processing on the explanatory text, and the entity recognition algorithm extracts key concepts, people, locations, times, and other entity information from the text. The relation extraction module identifies semantic relationships between entities, including causal, temporal, and spatial relationships. The knowledge graph is stored in a graph database, where nodes represent entities and edges represent relationships, and each node and edge contains rich attribute information. The graph is updated incrementally, with new explanatory content automatically integrated into the existing knowledge structure.

[0074] The semantic matching engine processes user queries, while the query understanding module analyzes the user's intent, including factual queries, explanations, and comparative analyses. The entity linking algorithm maps keywords in the query to corresponding nodes in the knowledge graph. The path search algorithm searches the graph for subgraph structures related to the query, with a search depth limited to three levels to control response time. The relevance score comprehensively considers factors such as entity matching degree, path length, and node importance.

[0075] The supplementary explanation content generator synthesizes answer content based on the retrieved knowledge subgraph. The content organization employs a template-based generation method, selecting an appropriate answer structure according to different question types. Factual questions use a direct answer mode, explanatory questions use a causal analysis mode, and comparative questions use a comparative explanation mode. The generated content is kept under 200 characters to ensure conciseness and clarity. A content consistency check verifies the logical consistency between the generated answer and the original explanation content.

[0076] The RAG vector retrieval enhancement module expands the range of knowledge sources for explanations. User-uploaded digital materials are processed by a document parser, supporting common formats such as PDF, Word, and PPT. Text extraction utilizes OCR technology to process image content, achieving an accuracy rate of over 95%. Document content is segmented, with each segment limited to 512 characters for easy vectorization and storage. The vector database employs a FAISS index structure, supporting rapid retrieval of millions of documents. Retrieval results are integrated and sorted with knowledge graph content, providing a more comprehensive basis for explanations and answers.

[0077] The online search extension module obtains real-time commentary information via the internet. The search interface supports parallel queries from multiple search engines, improving information coverage. The search result filtering algorithm removes advertisements, duplicates, and low-quality content, retaining only authoritative information sources. The content summary generator extracts key information from search results, generating concise summary text. A real-time information caching mechanism avoids duplicate searches, with a cache validity period set to 24 hours.

[0078] The timing alignment and blending module ensures synchronized playback of video and narration. The video playback timeline serves as the baseline, aligning both the narration audio and animation to the video playback time. Alignment precision is controlled within one audio frame, approximately 23 milliseconds. Content blending employs layered rendering: video content as the background layer, the digital human narration animation as the foreground layer, supplementary visual materials as the middle layer, and text descriptions as an overlay layer. Rendering priority ensures that interactive interruptions are always on top, preventing them from being obscured by other content.

[0079] In one optional implementation, the original video data is temporally decomposed to extract video frame sequences and identify the main objects and background regions within them. Simultaneously, the audio track is semantically parsed to obtain multimodal deconstructed data containing both visual and semantic elements, including:

[0080] The original video data is segmented into a video frame sequence according to the timeline, and each frame in the video frame sequence is spatially divided to distinguish the foreground region and the background region; the main object in the foreground region is detected and its attributes are extracted. The attribute extraction includes identifying the contour boundary features, pose features and motion trajectory features of the main object, and establishing a cross-frame tracking identifier for the main object.

[0081] Acoustic signal processing is performed on the audio track to convert continuous speech signals into discrete text segments. Natural language understanding is performed on the discrete text segments to extract semantic units and their syntactic structures. The discrete text segments are timestamped with the video frame sequence. Based on the cross-frame tracking identifier, the main object is located in the aligned video frames. The visual state of the main object when it emits the corresponding semantic unit is determined, and the binding relationship between the semantic unit and the visual state is established.

[0082] The contour boundary features, pose features, motion trajectory features, and scene features of the background region are summarized as visual elements, the semantic units, syntactic structures, and timestamp information are summarized as semantic elements, and the binding relationship is used as the association index between the visual elements and the semantic elements to form multimodal deconstruction data.

[0083] When performing time-series decomposition on the original video data, the video file is loaded and the video stream and audio stream are separated by the decapsulation module. The video stream generates a video frame sequence in the decoding order. Each frame carries a millisecond-level timestamp and is stored in an index array. The array elements contain the frame index number, timestamp, image data pointer, and frame type identifier.

[0084] The spatial region segmentation of each frame in the video frame sequence employs a semantic segmentation model. The model takes a single-frame RGB image as input and outputs a segmentation mask of the same size as the input, with each pixel in the mask labeled with its category. The distinction between foreground and background regions is based on category assignment: moving objects such as humans, animals, and vehicles are classified as foreground regions, while stationary scenes such as the sky, buildings, and ground are classified as background regions. The segmentation model uses an encoder-decoder architecture. The encoder extracts multi-scale features, and the decoder restores spatial resolution through upsampling. The inference latency per frame is approximately fifty milliseconds. The segmentation mask undergoes morphological processing to remove isolated noise, and the smoothing window size is five pixels.

[0085] When detecting the main object in the foreground region, the object detection model takes a video frame image as input and outputs a list of detection results including category labels, confidence scores, and bounding box coordinates. The bounding box coordinates are the pixel coordinates of the top-left and bottom-right corners, normalized to the zero-to-one range, with a confidence threshold of 0.5. The detection model predicts at a multi-scale feature layer, with a single-frame inference latency of approximately 30 milliseconds. Detection results are sorted in descending order of confidence score, and non-maximum suppression is applied to overlapping bounding boxes of the same category, with an intersection-over-union (IoU) threshold of 0.45.

[0086] Contour boundary feature extraction is based on the foreground mask within the bounding box. Contour detection is performed on the mask edges, and the coordinate sequence of contour points is extracted and approximated as a polygon with an approximation accuracy of 2% of the contour perimeter. After approximation, the number of contour points is controlled between fifty and two hundred. Too many points are reduced by sampling at equal intervals, and too few points are supplemented by spline interpolation. The contour feature vector includes area, perimeter, circularity, aspect ratio, and convex hull area ratio, which are normalized and concatenated into a fixed-dimensional feature vector.

[0087] Pose feature extraction targets the human subject. The keypoint detection model takes a bounding box cropped image as input and outputs the coordinates and visibility confidence scores of seventeen human keypoints. Keypoints include the head, shoulders, elbows, wrists, hips, knees, and ankles, with coordinates normalized relative to the cropped image. The model uses a heatmap regression method to generate a Gaussian heatmap for each keypoint, with the peak position of the heatmap corresponding to the keypoint coordinates, and the standard deviation of the Gaussian kernel is three pixels. Keypoints with a visibility confidence score below 0.3 are marked as invisible. The pose feature vector consists of the relative distances between keypoints, angles, and bone length ratios, with a 50-dimensional dimension.

[0088] Motion trajectory features are based on the positional changes of the subject object across consecutive frames. The coordinates of the bounding box center points form a trajectory point sequence over time. Subtracting the center point coordinates of adjacent frames yields the displacement vector. Dividing the displacement vector by the inter-frame time interval yields the velocity vector. Dividing the inter-frame difference of the velocity vector by the time interval yields the acceleration vector. The trajectory point sequence is smoothed using a Kalman filter. The filter state vector contains both position and velocity. The standard deviation of the process noise is two pixels, and the standard deviation of the observation noise is five pixels. The motion trajectory feature vector includes trajectory length, average velocity, maximum velocity, velocity variance, and trajectory curvature.

[0089] When establishing cross-frame tracking identifiers, a multi-object tracking method is employed. For each subject object, an appearance feature vector is extracted, and a 128-dimensional embedding vector is extracted from the bounding box-cropped image using a convolutional neural network. The tracker maintains a list of active trajectories, with each trajectory recording the object's historical position, appearance features, and unique identifier. New frame detection results are matched against active trajectories. The matching cost is calculated by combining the bounding box intersection-union ratio (IUU) and the cosine similarity of appearance features, using the Hungarian algorithm to find the optimal match. Successfully matched detections update their corresponding trajectories, while unmatched detections are initialized as new trajectories. Unmatched trajectories are marked as terminated if they are not updated within ten consecutive frames. The tracking identifier uses a globally unique, incrementing integer.

[0090] In audio track acoustic signal processing, the continuous audio waveform is divided into audio segments with a duration of two seconds and an overlap of 50% between segments. Each segment undergoes pre-emphasis filtering to enhance high frequencies, with a filter coefficient of 0.97. The pre-emphasis signal is converted into a spectrum using a short-time Fourier transform with a Hamming window function, a window length of 25 milliseconds, and a frame shift of 10 milliseconds. The spectrum is input into the acoustic model for speech recognition. The model employs an attention mechanism architecture, where the encoder encodes the spectrum sequence and the decoder generates the text sequence. Decoding uses a beam search with a beam width of five, outputting the first five text hypotheses and their confidence scores. The recognition result is then corrected by a language model.

[0091] When performing natural language understanding on text fragments, word segmentation is performed first. For Chinese, a hybrid approach combining dictionary and statistical models is used. The segmented results are then tagged with parts of speech using a conditional random field (CRF) model. Features include the word itself and its preceding and following words. Syntactic structure is obtained through dependency parsing. The dependency model is based on a transition system, taking a word sequence and its part of speech as input, and outputting the dependency relationships and their types, including forty types such as subject-verb, verb-object, attributive-head, and adverbial.

[0092] Semantic unit extraction is based on dependency syntax trees, identifying core verbs and their governed noun phrases and prepositional phrases as semantic units. Each semantic unit carries a semantic role label, including agent, patient, instrument, location, and time, based on dependency relations and part-of-speech rules. Semantic units are represented as triples, containing the core predicate, arguments, and semantic roles. The syntactic structure is stored as a set of dependency relation edges, with each edge recording the source word index, target word index, and relation type.

[0093] The alignment of text fragments with video frame sequences is based on a forced alignment technique, mapping each word in the identified text to an audio waveform time interval. The alignment employs a Hidden Markov Model (HMM), where the model state corresponds to a phoneme sequence, and the observed values ​​are the audio frame spectra. The start and end times of each phoneme are decoded using the Viterbi algorithm with an accuracy of ten milliseconds. Phoneme times are aggregated into word-level times, and further aggregated into semantic unit time ranges. Timestamps include the start time and duration, in milliseconds.

[0094] When locating a subject object in aligned video frames based on tracking identifiers, the corresponding video frame index is queried according to the timestamp range of the text segment. Within the corresponding frame, the bounding box and attributes of the subject object are retrieved using the tracking identifier number. If the time range spans multiple frames, the frame corresponding to the midpoint of the time frame is selected as the representative frame, with keyframes being preferred. The contour boundary features, pose features, and position coordinates of the subject object located in the representative frame are extracted to represent the visual state of the object when emitting semantic units.

[0095] The semantic unit and visual state binding relationship is established as a key-value mapping. The key is a unique identifier for the semantic unit, generated by combining the text fragment index and the fragment's sequence number. The value is a visual state feature record, containing the subject object tracking identifier, video frame index, bounding box coordinates, contour feature vector, pose feature vector, and trajectory feature vector reference pointer. The binding relationship supports bidirectional queries.

[0096] When aggregating visual elements, contour boundary features, pose features, motion trajectory features, and background scene features are integrated into a visual element dataset. Background scene features are extracted using a scene classification model. The input is the image region corresponding to the background mask, and the output is the scene category probability distribution. The feature vector uses the second-to-last activation value of the model, with a dimension of 512. Visual elements are organized by timestamp, and each record contains a frame index, a list of main objects, and background scene features.

[0097] When aggregating semantic elements, semantic units, syntactic structures, and timestamp information are integrated into a semantic element dataset. Organized by text segment index, each segment record includes the original text, word segmentation results, part-of-speech tagging, dependency syntax tree, a list of semantic unit triples, and a timestamp range. Triples record the core predicate, a list of arguments, and semantic roles; arguments record the word index range and part-of-speech. Syntactic structures are stored as a list of dependency relation edges.

[0098] When using binding relationships as associated indexes, an inverted index structure is created. A list of associated visual element records is maintained for each semantic unit, and a list of associated semantic unit identifiers is maintained for each visual element. The index supports time range queries, implemented using a combination of hash tables and skip lists. Hash tables support constant-time exact queries, while skip lists support logarithmic time range queries. Index persistence uses a key-value database.

[0099] The resulting multimodal deconstructed data adopts a hierarchical structure. The top layer is video metadata, the second layer is a timeline index, and the third layer consists of visual element data tables and semantic element data tables, connected by associative indexes. The data uses columnar storage and lossless compression, with a compression ratio of approximately three to one. The data interface provides queries by time range, subject object identifier, semantic unit, and scene category, with query latency controlled within hundreds of milliseconds.

[0100] In one optional implementation, based on the explanatory text in the multimodal deconstruction data, a corresponding explanatory script is generated for each time segment of the video content, and the visual elements in the video frame sequence corresponding to each explanatory segment are timestamped to form a data structure that synchronizes video playback with the explanatory content in time, including:

[0101] The explanatory text is extracted from the multimodal deconstruction data, and the explanatory text is divided into multiple explanatory segments according to the timeline of video playback. Each explanatory segment corresponds to a continuous time period in the video, and each explanatory segment is marked with a start timestamp and an end timestamp.

[0102] Semantic analysis is performed on each of the explanatory segments to identify important information points and key points of explanation. Based on the important information points, extended explanatory content is generated, and the original explanatory text and the extended explanatory content are integrated to form a complete explanatory script.

[0103] The video frame sequence and its associated visual elements are obtained from the multimodal deconstruction data. The binding relationship in the multimodal deconstruction data is used to determine the timestamp range corresponding to each narration segment. The corresponding visual elements are extracted from the video frame sequence within the timestamp range.

[0104] The narration script, the visual elements, and the video frame sequence are aligned in three dimensions according to timestamps to construct a synchronized data structure that includes a video playback timeline, a narration content timeline, and a visual element timeline, ensuring that video playback and digital human narration are synchronized in time.

[0105] Extracting explanatory text from multimodal deconstructed data requires deep analysis of the data structure. The multimodal deconstructed data employs a layered storage architecture. The top layer is a metadata description layer that records basic information such as data type identifiers, encoding formats, and timestamp ranges. The middle layer is a content data layer that stores visual element data and semantic element data respectively. The bottom layer is a relationship layer that maintains the binding mapping between different data types. The text extraction module determines the storage location and format specifications of the semantic element data by parsing the metadata description layer. The semantic element data is stored in JSON format and contains structured information such as text content fields, timestamp fields, semantic tag fields, and confidence scores. During extraction, the root node of the semantic element data is read first, and all text content fields are traversed to obtain the original explanatory text. Each text record contains complete attributes such as start timestamp, end timestamp, text string, language type identifier, and phoneme alignment information. The text decoder performs character set conversion and format standardization on the extracted text string, supporting multiple character encodings such as UTF-8, GBK, and ASCII, while maintaining the integrity and accuracy of the original timestamp information during the conversion process.

[0106] The explanatory text is composed of multiple layers of semantic information. The base layer is the original transcribed text, recording the direct text-to-text conversion of human voices in the audio track. The transcription accuracy is required to reach over 95%, and it includes speech features such as punctuation marks, intonation markers, and pause markers. The semantic annotation layer adds part-of-speech tags, named entity recognition results, and semantic role annotations to the original text. Part-of-speech tags use a general part-of-speech tagging set containing eighteen basic types, including nouns, verbs, adjectives, adverbs, and prepositions. Named entity recognition covers seven major entity types, including personal names, place names, organization names, time, and quantity. The concept association layer establishes a mapping relationship between text content and domain knowledge. It identifies key information points such as professional terms, technical concepts, and operational steps through concept dictionary matching. Each concept association record includes attribute information such as concept identifier, definition description, association strength, and frequency of occurrence. The context dependency layer analyzes the logical relationships and semantic dependencies between sentences, identifying semantic connections such as causal relationships, temporal relationships, parallel relationships, and adversative relationships, providing a semantic foundation for subsequent script generation and expansion.

[0107] The relationship between text and semantic elements is established through a multi-dimensional mapping mechanism. The time dimension mapping ensures a precise correspondence between text fragments and audio timestamps, with mapping granularity reaching the vocabulary level. Each word records the start and end times of pronunciation with a precision of ten milliseconds. The semantic dimension mapping establishes the association between text content and conceptual knowledge. It calculates the similarity between text fragments and predefined concepts using semantic vectors, with a similarity threshold set at 0.6. Concepts exceeding this threshold are automatically associated with their corresponding text fragments. The sentiment dimension mapping analyzes the text's emotional tendency and intonation characteristics, identifying three emotional polarities: positive, negative, and neutral. Emotional intensity is evenly distributed across five levels from 0.2 to 1.0. The importance dimension mapping determines the importance of text fragments through information entropy calculation and keyword weight analysis. The information entropy threshold is set at 3.5, and keyword weights are calculated using the TF-IDF algorithm, with weight values ​​ranging from 0.1 to 10.0. The semantic element data structure maintains an index table of text identifiers and mapping values ​​for each dimension, supporting rapid retrieval and filtering operations based on arbitrary dimension conditions.

[0108] The text segmentation module intelligently segments the extracted narration text according to the video playback timeline. The segmentation algorithm comprehensively considers three factors: audio silence detection, semantic boundary recognition, and visual scene changes. Audio silence detection identifies natural pauses by analyzing the energy distribution of the audio track. The energy threshold is set to 10% of the average energy, and positions where the silence duration exceeds 300 milliseconds are marked as potential segmentation points. Semantic boundary recognition determines the boundaries of complete semantic units through syntactic analysis and semantic consistency checks. Syntactic analysis is based on dependency grammar to parse sentence structure, and semantic consistency is assessed through topic modeling to evaluate the topic relevance of adjacent sentences. Positions with a relevance below 0.4 are marked as semantic boundaries. Visual scene change detection identifies scene transition points by comparing the feature vectors of adjacent video frames. The feature vectors include visual descriptors such as color histograms, texture features, and edge density. The scene change threshold is set to 0.3, and frames exceeding the threshold are marked as visual segmentation points after corresponding to the audio timestamp.

[0109] The segment generator determines the final segment boundaries by integrating three segmentation factors, prioritizing semantic boundaries, audio silence, and visual changes. A weighted voting mechanism is used for segmentation decisions, with semantic boundaries weighted at 0.5, audio silence at 0.3, and visual changes at 0.2. Each segment's duration is controlled between eight and forty-five seconds. Segments that are too short are merged through semantic relevance analysis. Relevance calculation is based on word overlap rate and topic similarity, with an overlap rate threshold of 30% and a topic similarity threshold of 0.6. Longer segments are recursively segmented while maintaining semantic integrity, prioritizing grammatical boundaries such as clause boundaries, parallel structure divisions, and time adverbial positions. The timestamp annotation module precisely marks the start and end timestamps for each segment. The start timestamp corresponds to the beginning of the pronunciation of the first word in the segment, and the end timestamp corresponds to the end of the pronunciation of the last word plus a natural pause duration. The pause duration is determined based on speech rhythm analysis and ranges from one hundred to five hundred milliseconds.

[0110] The semantic analysis engine performs multi-level semantic parsing on each generated explanation segment. Lexical analysis identifies keywords in the segment through word frequency statistics, part-of-speech tagging, and word meaning disambiguation. Word frequency statistics use logarithmic frequency to avoid over-weighting high-frequency words, and word meaning disambiguation uses context windows and word co-occurrence patterns to determine the precise meaning of words in the current context. Syntactic analysis identifies grammatical information such as subject-verb-object structure, modification relations, and coordinate relations in sentences through dependency parsing. The parsing results are stored in a tree structure, with nodes recording lexical information and grammatical roles, and edges recording grammatical relationship types and dependency strengths. Semantic analysis identifies semantic roles such as action actor, action object, action mode, time, and place through semantic role labeling. The semantic role tag set includes twelve core role types, such as agent, patient, instrument, time, place, and mode.

[0111] The key information point identification module uses multi-dimensional feature fusion to determine the content that needs to be emphasized in a segment. Feature dimensions include indicators such as lexical importance, syntactic complexity, semantic novelty, and concept density. Lexical importance is calculated by matching TF-IDF weights with a domain dictionary, which includes high-value words such as technical terms, operational concepts, and key entities. The matching weights are set according to the importance of the words within the domain. Syntactic complexity is quantified using grammatical features such as sentence length, clause levels, and the number of modifiers. Sentences with high complexity usually contain more information and require detailed explanation. Semantic novelty assesses the degree of a concept's first appearance in the current context and its difficulty of comprehension. Concepts with high novelty require additional background explanation and examples. Concept density calculates the distribution density of technical concepts within a unit of text length. A density threshold is set to one technical concept per ten words; segments exceeding this threshold are marked as concept-dense regions.

[0112] The extended explanation content generator retrieves relevant supplementary information from a pre-built knowledge base based on identified key information points. The knowledge base uses a graph database architecture to store concept nodes and relational edges. Concept nodes contain complete descriptive information such as definitions, attributes, instances, and application scenarios, while relational edges record semantic connections such as hierarchical relationships, associations, and comparisons between concepts. The retrieval strategy employs a multi-hop graph traversal algorithm, starting from the concept node corresponding to the key information point and traversing adjacent nodes from one to three hops to obtain related concepts. The traversal depth is dynamically adjusted based on concept complexity and audience background. Relevance calculation comprehensively considers multiple factors such as semantic distance, co-occurrence frequency, and user feedback. Semantic distance is calculated using the cosine similarity of concept vectors; co-occurrence frequency is calculated by counting the number of times a concept jointly appears in historical explanations; and user feedback is learned from historical interaction data to assess user acceptance and satisfaction with different extended content.

[0113] The supplementary information is organized according to cognitive load theory and instructional design principles. The extended content is arranged in a logical order: definition and explanation, background introduction, examples, and application scenarios. The definition and explanation provide an accurate definition and description of the core characteristics of the concept, with a length of 30-50 characters, avoiding overly complex technical jargon. The background introduction supplements the concept's development history, theoretical basis, and applicable conditions, helping to understand the concept's origin and application context. Examples demonstrate the practical application and operational methods of the concept through specific cases, prioritizing scenarios relevant to the current video content. Application scenarios describe the specific uses of the concept in different fields and situations, expanding the audience's understanding of the concept's practical value.

[0114] The integration of the original explanatory text and expanded explanatory content employs an intelligent insertion algorithm. The insertion points are determined through semantic dependency analysis and cognitive process modeling. Semantic dependency analysis identifies key locations in the original text, such as the first appearance of a concept, logical turning points, and causal relationship nodes; these locations are typically the optimal times to insert expanded content. Cognitive process modeling simulates the audience's comprehension process, predicting where additional information is needed to support understanding. The model is trained based on cognitive psychology theory and teaching practice data. The insertion strategy adopts a progressive unfolding approach: basic definitions are inserted when a concept first appears, and deeper explanations and applications are inserted as needed based on the context. The integrated explanatory script maintains consistency and fluency in language style. Natural language generation technology is used to adjust the expression of the inserted content to ensure coordination with the tone, pace, and expression habits of the original text.

[0115] The video frame sequence processing module obtains preprocessed frame information from the visual element portion of the multimodal deconstruction data. Each frame is stored in a standardized format, containing complete metadata such as frame index number, absolute timestamp, relative timestamp, frame resolution, color space, and compression parameters. The frame index number uses zero-based incremental numbering to ensure the continuity and integrity of the frame sequence. The absolute timestamp records the frame's precise position within the entire video, while the relative timestamp records the frame's relative position within the current segment. Timestamp precision is uniformly set to millisecond levels. The visual element extractor performs multi-level visual analysis on each frame. Pixel-level analysis extracts basic visual features such as color distribution, brightness contrast, and saturation changes. The feature vector dimension is set to 128 dimensions, including various descriptors such as RGB color histogram, HSV color statistics, and grayscale gradient distribution. Object-level analysis identifies specific objects in the frame using a deep learning object detection model. The detection model supports the recognition of 80 common object categories, including people, vehicles, buildings, equipment, text, and charts. Each detection result includes detailed information such as bounding box coordinates, category label, confidence score, and attribute description.

[0116] Scene-level analysis identifies the scene type of a frame through a scene classification model. Scene classification includes fifteen main categories such as indoor, outdoor, meeting, teaching, demonstration, and experiment. The classification confidence threshold is set to 0.75; frames below the threshold are marked as having an unknown scene type. Semantic-level analysis identifies high-level semantic content within the frames, including text, symbols, and chart data. Text recognition uses optical character recognition technology to accurately identify Chinese, English, and common symbols, with an accuracy rate of over 92%. Symbol recognition includes commonly used symbols such as arrows, labels, highlights, and circles. Chart data recognition supports structured parsing of common chart types such as bar charts, line charts, pie charts, and flowcharts.

[0117] The binding relationship between visual elements and narration segments is established through precise association using timestamp information stored in the multimodal deconstruction data. The binding algorithm employs a time window matching strategy, determining a corresponding video frame range for each narration segment. The time window size is dynamically adjusted based on the segment duration: when the segment duration is less than ten seconds, the window size equals the segment duration; when the segment duration exceeds ten seconds, the window size is set to a buffer between the segment duration and two seconds. During the matching process, potential synchronization deviations between audio and video are considered, typically ranging from -100 milliseconds to +300 milliseconds. Cross-correlation analysis is used to detect and correct these deviations. The binding relationship record contains complete information including the narration segment identifier, start frame index, end frame index, keyframe list, and visual element summary. Keyframes are automatically selected through visual change detection and content importance assessment, with an average of three to eight keyframes per segment.

[0118] The 3D aligned data structure design employs a multi-level index and distributed storage architecture to ensure precise synchronization of the video playback timeline, the narration timeline, and the visual element timeline. The top-level index establishes coarse-grained positioning based on whole-second timestamps, with each index entry recording the data range and storage location within that second, supporting fast jumps and range queries. The second-level index establishes fine-grained positioning with 100-millisecond precision, covering specific boundary information for video frames, audio clips, and narration text. The index density ensures that data at any given time point can be accurately located within three queries. The tertiary index records specific data entity references, including frame numbers, text clip identifiers, and visual element identifiers—direct access information. Index updates use an incremental approach to avoid the performance overhead of full reconstruction.

[0119] The timing synchronization control mechanism ensures real-time synchronization during playback through multi-threaded concurrent processing and buffer queue management. The video rendering thread, audio playback thread, and narration synthesis thread run independently and are coordinated and synchronized via timestamps. The playback controller maintains a global clock reference, and the timestamps of all threads are calibrated relative to this global clock, with clock accuracy at the microsecond level and synchronization errors controlled within twenty milliseconds. The buffer queue maintains a three-second pre-loading buffer for each data stream. The buffer uses a circular queue structure to support efficient read and write operations. When the buffer occupancy rate falls below 30%, the pre-loading mechanism is triggered to replenish data. The synchronization correction algorithm monitors the time deviation between each data stream in real time. When the deviation exceeds a preset threshold, dynamic correction is performed through playback speed fine-tuning, frame skipping, and audio stretching. The correction intensity adaptively adjusts according to the degree of deviation, ensuring that the correction process minimizes the impact on the user experience.

[0120] In one optional implementation, receiving a user-input interruption request, the interruption request containing a query intent and control instructions regarding the content to be explained, performing semantic matching of the query intent within the explanation script, locating the target explanation segment related to the query intent and its associated visual elements, and generating supplementary explanation content based on the contextual relationships of the target explanation segment, including:

[0121] During video playback, the system monitors the user input channel in real time. When an interruption request is detected, the system records the current video playback time as the interruption timestamp and sends a pause command to the video playback module and the digital human narration module, so that the video playback and digital human narration are paused synchronously.

[0122] The interruption request is semantically parsed to extract the core concepts in the query intent. The core concepts are then semantically similar to the explanatory segments in the explanatory script to locate the target explanatory segment related to the query intent and obtain the visual elements mapped to the target explanatory segment.

[0123] Based on the position of the target explanation segment on the timeline, the explanation script is expanded forward and backward to extract context-related explanation segments. By analyzing the semantic relationship between the target explanation segment and adjacent explanation segments, a context explanation sequence containing cause and effect is constructed.

[0124] The contextual explanation sequence is verified for completeness. If any missing information is detected, a bridging explanation segment that can fill in the missing information is searched from other time periods of the explanation script. The bridging explanation segment and its mapped visual elements are added to the contextual explanation sequence to form a logically complete supplementary explanation content.

[0125] like Figure 2 As shown, the method includes:

[0126] During video playback, the user input monitoring module establishes a communication connection with the input devices, including keyboards, touchscreens, and voice input devices, while simultaneously monitoring the user input channel in real time. Voice input uses a microphone to collect audio signals at a sampling rate of 16 kHz and a sampling precision of 16 bits. Endpoint detection identifies speech activity segments based on an energy threshold and a zero crossover rate, with the energy threshold set to three times the average energy of the background noise. Upon detecting a speech activity segment, speech recognition is initiated, with audio spectral features as the model input and text transcription as the output. Text input monitors keyboard input events, capturing the input text upon detecting specific shortcut keys or text submission operations.

[0127] When an interruption request is detected, it is determined whether the input content conforms to the interruption request format. The interruption request format includes a query intent field and a control command field. The query intent includes user questions or keywords, and the control command includes operations such as pause and jump. Format validation is completed collaboratively by the regular expression matching and natural language understanding modules. Interruption requests that conform to the format trigger the interruption processing flow, recording the current video playback time as the interruption timestamp. The interruption timestamp is obtained by querying the playback status of the video playback module, with a time precision of milliseconds, and is stored in the interruption event log.

[0128] When a pause command is sent to the video playback module and the digital human narration module, control messages are transmitted via inter-process communication. The control messages use a structured data format, with the message header containing the message type, timestamp, and session identifier. Upon receiving the pause command, the video playback module stops rendering video frames and outputting audio, saving the current playback state to a memory cache. Upon receiving the pause command, the digital human narration module stops synthesizing and outputting narration audio, saving the current narration segment index and narration progress. The pause operation requires synchronized pausing of video playback and digital human narration; this synchronization is achieved through broadcast messages and acknowledgment responses. The controller broadcasts the pause command to both modules, waiting for a pause acknowledgment response. The acknowledgment response timeout is set to 200 milliseconds. If no response is received within the timeout period, the pause command is resent, with a maximum of three resentments.

[0129] When performing semantic parsing on interruption requests, the core concepts in the query intent are extracted. The semantic parsing module employs a combination of dependency parsing and named entity recognition. Dependency parsing identifies the subject-verb-object structure of the query intent and extracts the core verbs and noun phrases. Named entity recognition annotates proper nouns, terms, and key concepts in the query intent; entity types include personal names, place names, subject concepts, and technical terms. The core concept extraction rule is to select the core verb corresponding to the root node of the dependency parsing tree and its directly governed noun phrases, as well as all identified named entities. After deduplication of the core concept set, a list of query keywords is formed.

[0130] When calculating the semantic similarity between the core concepts and the explanatory segments in the script, a semantic representation vector is extracted for each segment. The semantic representation of the explanatory segment is generated based on a pre-trained language model. The model input is the explanatory segment text, and the output is a fixed-dimensional embedding vector with 768 dimensions. The core concepts of the query intent are also embedded using a language model. Multiple keyword embedding vectors are fused through a weighted average, with weights calculated based on word frequency inverse document frequency. The semantic similarity between the query intent embedding vector and the explanatory segment embedding vector is calculated using cosine similarity, with a value ranging from -1 to +1. The similarity calculation results are sorted in descending order, and explanatory segments with a similarity higher than a threshold of 0.6 are selected as candidate target segments, with a maximum of five candidate segments.

[0131] When locating a target explanatory segment related to the query intent, a secondary filtering process is performed on candidate target segments. This secondary filtering is based on the control instruction field of the query intent. If the control instruction specifies a particular time range or segment attribute, candidate segments that do not meet the conditions are filtered out. From the remaining candidate segments after filtering, the segment with the highest semantic similarity is selected as the target explanatory segment. After the target explanatory segment is determined, its segment identifier, timestamp range, and semantic similarity score are recorded.

[0132] When acquiring the visual elements mapped to the target narration segment, the timestamp range of the target narration segment is queried through a synchronized data structure, and the corresponding visual element record is retrieved from the visual element timeline table. Each visual element record contains a frame index, subject object tracking identifier, bounding box coordinates, contour feature vector, pose feature vector, and background scene features. The visual element records are sorted by frame index to form a sequence of visual elements corresponding to the target narration segment.

[0133] When extracting context-related explanatory segments based on their position on the timeline, the expansion range is determined. The expansion range is limited by both the time window and the number of segments. The forward expansion time window is set to 60 seconds before the start timestamp of the target explanatory segment, and the backward expansion time window is set to 60 seconds after the end timestamp. The number of segments is limited to a maximum of five segments for forward expansion and a maximum of five segments for backward expansion. Explanatory segments are extracted chronologically within the expansion range. Each extracted segment includes a segment identifier, a timestamp range, script text, and a list of semantic units.

[0134] When analyzing the semantic relationships between the target explanatory segment and adjacent explanatory segments, a semantic coherence score is calculated. The semantic coherence score comprehensively considers lexical overlap, topic consistency, and dependency continuity. Lexical overlap is calculated by dividing the number of co-occurring words in adjacent segments by the total number of words in both segments. Topic consistency is measured by the cosine similarity of the segment embedding vectors. Dependency continuity determines whether there are cross-segment dependency edges in the dependency syntactic trees of adjacent segments. The three indicators are weighted and summed to obtain the semantic coherence score, with lexical overlap weighted at 0.3, topic consistency weighted at 0.5, and dependency continuity weighted at 0.2. Adjacent segments with a coherence score higher than 0.5 are considered strongly related and included in the contextual explanatory sequence.

[0135] When constructing a contextual explanation sequence containing cause and effect, the target explanation segment is used as the core segment, the strongly related segments extending forward are used as the cause segment sequence, and the strongly related segments extending backward are used as the effect segment sequence. The cause segment sequence is arranged in chronological order, the effect segment sequence is arranged in chronological order, and the core segment is inserted between the end of the cause sequence and the beginning of the effect sequence. The contextual explanation sequence structure includes a sequence identifier, a core segment index, a segment list, and a total duration.

[0136] When verifying the completeness of the contextual explanation sequence, the system checks for missing information in the explanation logic. Information missing detection is based on causal chain integrity analysis and concept dependency graph verification. Causal chain integrity analysis extracts causal semantic units from the contextual explanation sequence to determine if the causal chain is broken. Causal semantic units are identified through causal edges in dependency parsing; causal relationship types include cause, effect, condition, and purpose. If a cause semantic unit in the causal chain has no corresponding effect semantic unit in the sequence, or a result semantic unit has no corresponding cause semantic unit, it is considered information missing. Concept dependency graph verification constructs all conceptual entities mentioned in the sequence and their dependencies, including definition dependencies, pre-dependencies, and derivation dependencies. If a pre-dependent concept of a conceptual entity does not appear in the sequence, it is considered information missing.

[0137] When missing information is detected, a bridging segment is searched from other time periods of the narration script to fill in the missing information. The bridging segment search strategy is based on semantic matching of the missing concept, using the missing concept as the query keyword, and calculating semantic similarity among all narration segments in the narration script. Segments with a similarity higher than 0.7 are selected as candidate bridging segments. Candidate bridging segments must meet the requirements of content independence and logical consistency. Content independence requires that the bridging segment does not depend on any missing prior concepts, and logical consistency requires that the bridging segment has internal semantic coherence and no contradictions. Candidate bridging segments are sorted in descending order of similarity, and the segment with the highest ranking is selected as the bridging narration segment.

[0138] When supplementing the bridging explanatory segments and their mapped visual elements into the contextual explanatory sequence, the insertion position of the bridging segment is determined. The insertion position is based on the location of missing information and the timeline order. If the missing information occurs between the cause segment sequence and the core segment, the bridging segment is inserted at the end of the cause sequence. If the missing information occurs between the core segment and the consequence segment sequence, the bridging segment is inserted after the core segment and before the beginning of the consequence sequence. After insertion, the segment list and the total sequence duration are updated. The visual elements of the bridging segment are retrieved through a synchronized data structure query, and the visual element records are added to the corresponding positions in the visual element sequence.

[0139] When generating logically complete supplementary explanation content, the supplemented contextual explanation sequence is used as the final output. The supplementary explanation content data structure includes sequence identifiers, core segment indexes, a segment list, a bridging segment index list, a visual element sequence, and a total duration. The segment list is arranged in chronological or logical order, and bridging segments are marked as inserted segments in the list. The supplementary explanation content supports rendering output; the rendering module plays the synthesized speech corresponding to the script text in segment order, and synchronously displays the video frames corresponding to the visual elements.

[0140] In one optional implementation, based on the position of the target explanatory segment on the timeline, context-related explanatory segments are extracted by extending the script forward and backward. By analyzing the semantic relationships between the target explanatory segment and adjacent explanatory segments, a contextual explanatory sequence containing cause and effect is constructed, including:

[0141] Using the start timestamp of the target explanation segment as a reference point, the explanation script is backtracked forward along the timeline to extract adjacent explanation segments with timestamps earlier than the target explanation segment. The semantic dependencies between the adjacent explanation segments and the target explanation segment are analyzed, and the preceding explanation segments that serve as prerequisites for the target explanation segment are identified to form a forward context sequence.

[0142] Using the end timestamp of the target explanation segment as a reference point, a forward search is performed in the explanation script to the back of the timeline to extract adjacent explanation segments with timestamps later than the target explanation segment. The semantic continuity relationship between the adjacent explanation segments and the target explanation segment is analyzed, and the subsequent explanation segments derived from the target explanation segment are identified to form a backward context sequence.

[0143] The forward context sequence and the backward context sequence are integrated in timestamp order to construct a contextual explanation sequence centered on the target explanation segment.

[0144] In one optional implementation, based on the visual elements associated with the target explanation segment, a virtual avatar generator is driven to synthesize a dynamic expressive output synchronized with the supplementary explanation content. After completing the interactive explanation, video playback and synchronized digital avatar explanation are restored according to user instructions, including:

[0145] The associated visual elements are obtained from the target explanation segment and the context explanation sequence. The pose features and facial expression features in the visual elements are modeled as inter-frame motion trajectories. The original feature frames and intermediate transition frames are combined in time sequence to form a smooth motion feature sequence.

[0146] The virtual avatar generator synthesizes a dynamic expression output based on the feature sequence. The dynamic expression output includes a continuous frame sequence of the virtual avatar. The energy function is calculated for the pose features of each frame in the continuous frame sequence, and the frames whose energy function values ​​reach a local maximum are marked as visual anchor points.

[0147] A temporal anchor point correspondence table is constructed. By performing prosodic intensity analysis on the supplementary explanation content, the syllable stress position and semantic key words are extracted as semantic anchor points. The semantic anchor points and the visual anchor points are matched with semantic relevance scores, and the time difference between the semantic anchor points and the visual anchor points is calculated.

[0148] Based on the time difference, the continuous frame sequence is segmented into nonlinear time mappings to align the visual anchor point with the semantic anchor point on the time axis. The aligned continuous frame sequence is then output synchronously with the supplementary explanation content to complete the interactive explanation.

[0149] After the interactive explanation is completed, the system receives a user's command to resume playback or jump to another location. If a command to resume playback is received, the system resumes video playback from the video position corresponding to the interrupted timestamp and drives the digital human to continue synchronous explanation according to the explanation script corresponding to the resumed position. If a command to jump to another location is received, the system extracts the target time point specified in the jump command, jumps the video playback position to the target time point, and drives the digital human to start synchronous explanation according to the explanation script corresponding to the target time point.

[0150] When extracting associated visual elements from the target narration segment and the context narration sequence, each narration segment contains a visual element field in its storage structure. This field stores pose feature data and facial expression feature data in a key-value mapping format. The pose feature data includes a sequence of 3D coordinates of skeletal nodes, with 23 nodes covering key positions of the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. Each node coordinate is represented by three floating-point numbers with single-precision floating-point format. The facial expression feature data includes a sequence of facial motion encoding unit parameters, with 52 parameters corresponding to the motion intensity values ​​of eyebrows, eyelids, nose, lips, and cheeks, ranging from zero to one normalized floating-point number. The visual element extraction operation traverses the segment identifier list of the context narration sequence, queries the visual element field corresponding to each segment identifier from the narration script storage system, and appends the query results to the visual element buffer queue in chronological order of the segment in the sequence. The buffer queue has a maximum capacity of 500 feature frames, each containing three fields: timestamp, pose feature array, and facial expression feature array. The timestamp accuracy is in milliseconds. The pose feature array stores the coordinates of twenty-three skeletal nodes, and the facial expression feature array stores the parameters of fifty-two motion coding units.

[0151] When modeling inter-frame motion trajectories for pose and facial expression features in visual elements, adjacent original feature frames are read from the visual element buffer queue in timestamp order, and the time interval and feature difference between adjacent frames are calculated. The time interval is the difference between the timestamp of the next frame and the timestamp of the previous frame, in milliseconds with integer precision. The feature difference is calculated separately for pose feature difference and facial expression feature difference. The pose feature difference is the Euclidean distance of the corresponding skeletal node coordinates, and the facial expression feature difference is the absolute difference of the parameters of the corresponding action coding unit. Motion trajectory modeling uses a cubic Hermitian interpolation strategy. The interpolation operation generates intermediate transition frames between adjacent frames with a time interval greater than 50 milliseconds. The number of intermediate transition frames is dynamically calculated based on the time interval. The calculation method is to divide the time interval by 25 milliseconds and round down. The rounded result minus one is the number of intermediate transition frames. During the interpolation process, the feature value of the previous frame is used as the starting point, the feature value of the next frame is used as the ending point, the feature difference between the previous frame and the frame before that is used as the starting tangent, and the feature difference between the next frame and the frame after that is used as the ending tangent. The tangent slope is calculated by dividing the feature difference by the time interval, with units of feature values ​​per millisecond. The cubic Hermitian interpolation function takes a normalized time parameter as input (ranging from zero to one) and outputs the interpolated feature values. The normalized time parameter is divided at equal intervals, with the number of intervals equal to the number of intermediate transition frames plus one. The interpolated intermediate transition frames are combined with the original feature frames in ascending order of timestamp and written into the motion-smoothed feature sequence buffer. The feature sequence buffer has the same storage structure as the visual element buffer queue, with a maximum capacity of two thousand feature frames.

[0152] When the virtual avatar generator synthesizes dynamic expression output based on feature sequences, it receives a data stream from the feature sequence buffer. The data stream transmission uses a shared memory mechanism with a shared memory block size of eight megabytes, supporting concurrent access in a producer-consumer model. The virtual avatar generator internally comprises three processing units: a skeleton skinning module, an expression blending module, and a rendering output module. The skeleton skinning module reads the pose feature array and applies the bone node coordinates to the virtual avatar's skeleton binding mesh. The mesh has 8,500 vertices, each associated with up to four bone nodes, and the sum of the association weights is normalized to one. The skeleton skinning transformation uses a linear blending skinning algorithm, and the final vertex position is the weighted average of the transformed positions of all associated bone nodes. The expression blending module reads the facial expression feature array and applies the motion encoding unit parameters to predefined expression base forms. There are 52 base forms, each storing the offset vector of a facial mesh vertex. The expression blending result is a linear superposition of the offset vectors of each base form multiplied by their corresponding parameter values. The rendering output module receives the mesh data after the skeleton skinning and facial expression are mixed, performs rasterization and shading calculations, and generates an image frame with a resolution of 1920 x 1080 pixels in 24-bit true-color bitmap format. The continuous frame sequence of dynamic expression output is generated in the order of the timestamps of the feature frames in the feature sequence. Each feature frame corresponds to one image output frame. The continuous frame sequence is stored in the video frame buffer queue, with a maximum queue capacity of two thousand frames.

[0153] When calculating the energy function for the pose features of each frame in a continuous frame sequence, the energy function is defined as a comprehensive measure of the amplitude of change in the temporal dimension and the distribution range in the spatial dimension of the pose features. The temporal amplitude is calculated by summing the squared differences between the pose features of the current frame and the previous frame, and then summing these squared differences over the 3D coordinates of all twenty-three skeletal nodes. The spatial distribution range is calculated by averaging the standard deviations of the coordinates of each skeletal node in the current frame over each of the three coordinate axes. The energy function value is a weighted sum of the temporal amplitude and the spatial distribution range, with a temporal weight of 0.7 and a spatial weight of 0.3. When the energy function value is higher than the energy function values ​​of other frames in the local neighborhood, that frame is marked as a visual anchor point. The local neighborhood is defined as a time window of five frames before and after the current frame, with a total window length of eleven frames. The marking operation for visual anchor points involves traversing the continuous frame sequence through a sliding window, calculating the energy function value for the center frame of each window, and comparing the magnitude of the energy function value of the center frame with the energy function values ​​of other frames within the window. If the energy function value of the center frame is greater than the energy function values ​​of all other frames within the window, then the center frame is marked as a visual anchor point. The visual anchor point marking results are stored as a list of key-value pairs between frame indices and energy function values, with the list sorted in ascending order of frame indices.

[0154] When constructing the temporal anchor point mapping table, a mapping data structure is created to store the pairing relationship between semantic anchor points and visual anchor points. Prosodic intensity analysis is performed on the supplementary explanation content to extract syllable stress positions and semantically important words as semantic anchor points. The supplementary explanation content is audio signal data with a sampling rate of 16 kHz and a quantization precision of 16 bits. Prosodic intensity analysis performs short-time energy calculation and fundamental frequency extraction on the audio signal. The short-time energy calculation uses the Hanning window function with a window length of 25 milliseconds and a frame shift length of 10 milliseconds. Fundamental frequency extraction uses an autocorrelation function algorithm, limiting the fundamental frequency range to 80 to 400 Hz. Syllable stress positions are determined as time points where the short-time energy is 1.5 times higher than the neighborhood mean and the fundamental frequency rise slope is greater than 5 Hz per frame. Semantically important word extraction involves performing part-of-speech tagging and keyword extraction on the text transcription of the supplementary explanation content. Keyword extraction is based on the inverse document frequency (IVF) algorithm, selecting the top ten words with the highest scores as semantically important words. The temporal positions of semantically important words in the audio signal are obtained through a forced alignment algorithm. The input to the forced alignment algorithm is the audio signal and the text transcription, and the output is the start and end timestamps of each word. Semantic anchors contain the syllable stress position and the temporal position of semantically important words, and are stored as a list of timestamps with millisecond precision and integer precision.

[0155] When pairing semantic anchors with visual anchors for semantic relevance scoring, each semantic anchor in the list is traversed, and a semantic relevance score is calculated between that semantic anchor and all visual anchors. The semantic relevance score considers both temporal proximity and content relevance. Temporal proximity is the reciprocal of the absolute value of the time difference between the timestamps of the semantic anchor and the visual anchor, measured in milliseconds with integer precision. Content relevance is determined by analyzing the semantic matching degree between the text content corresponding to the semantic anchor and the posture features corresponding to the visual anchor. The matching degree is based on a predefined action semantic mapping table, which stores the association between action type tags and text semantic tags. Action type tags are extracted through cluster analysis of the posture features of the visual anchor frames. The clustering algorithm uses the mean-shift algorithm, and the number of cluster centers is automatically determined. Text semantic tags are extracted by semantic role annotation of the text content corresponding to the semantic anchor. Semantic roles include four categories: action, object, attribute, and relationship. Content relevance is calculated by dividing the number of matches between action type tags and text semantic tags in the mapping table by the total number of text semantic tags. The semantic relevance score is a weighted sum of temporal proximity and content relevance, with temporal proximity weighted at 0.6 and content relevance weighted at 0.4. For each semantic anchor, the visual anchor with the highest semantic relevance score is selected for pairing. The pairing results are stored in a temporal anchor mapping table, which is a list of records consisting of four fields: semantic anchor timestamp, visual anchor frame index, semantic relevance score, and time difference.

[0156] When calculating the time difference between semantic anchors and visual anchors, the time difference is the timestamp of the feature frame corresponding to the visual anchor minus the timestamp of the semantic anchor. The unit of the time difference is milliseconds with integer precision, and the time difference can be positive or negative. A positive value indicates that the visual anchor lags behind the semantic anchor, while a negative value indicates that the visual anchor is ahead of the semantic anchor. The time difference value is stored in the time difference field of the time-series anchor mapping table.

[0157] When performing piecewise nonlinear temporal mapping on a continuous frame sequence based on time difference, the continuous frame sequence is divided into multiple time segments according to the paired visual anchor point positions. Each time segment begins at the previous visual anchor point frame and ends at the current visual anchor point frame. The number of time segments equals the number of paired visual anchor points. A temporal mapping transformation is performed on the frames within each time segment, with the goal of aligning the timestamps of the visual anchor points within that time segment with the timestamps of their corresponding semantic anchor points. The temporal mapping transformation employs a piecewise linear stretching strategy, with the stretching factor being the target duration of the time segment divided by the original duration of the time segment. The target duration of the time segment is the current semantic anchor point timestamp minus the previous semantic anchor point timestamp, and the original duration of the time segment is the current visual anchor point frame timestamp minus the previous visual anchor point frame timestamp. The mapped timestamp of each frame within a time segment is the previous semantic anchor point timestamp plus the frame's relative position within the time segment multiplied by the stretching factor. The relative position is the frame's original timestamp minus the previous visual anchor point frame timestamp. The mapped sequence of consecutive frames is reordered according to the mapped timestamps. Frames with uneven timestamp intervals are adjusted to the target frame rate through a strategy of repeating or discarding frames. The target frame rate is set to 25 frames per second, with a frame interval of 40 milliseconds. When the mapped timestamp interval is less than 40 milliseconds, the intermediate frame is discarded; when the mapped timestamp interval is greater than 40 milliseconds, the previous frame is repeated to fill the interval.

[0158] When synchronizing the aligned continuous frame sequence with the supplementary narration, the image frame data of the continuous frame sequence is written to the video encoder, and the audio signal data of the supplementary narration is written to the audio encoder. The video encoder uses a variable bitrate encoding mode with a target bitrate of 2 megabits per second and a keyframe interval of one frame per second. The audio encoder uses a constant bitrate encoding mode with a bitrate of 128 kilobits per second and a sampling rate of 16 kilohertz. The encoded streams output by the video encoder and audio encoder are multiplexed into a container format, which supports independent video and audio tracks, with a time base set to millisecond precision. The multiplexed media stream is written to an output buffer with a capacity of 16 megabytes, supporting streaming. The data in the output buffer is pushed to the display device and audio playback device through the rendering interface, with the push rate synchronized with the target frame rate to ensure that the time alignment error between the video frame and the audio signal does not exceed 20 milliseconds. The interactive narration is considered complete when the last frame of the continuous frame sequence has been pushed to the display device and the audio signal of the supplementary narration has finished playing.

[0159] Upon receiving a user's command to resume playback or skip to the next screen after the interactive explanation, the user's command is captured by an interactive event listener. The listener monitors three input channels: keyboard input, touch input, and voice input. The resume playback command is recognized as a specific keycode or voice keyword. The keycode is the scan code of the space bar, and the voice keyword is the voice recognition result of "continue" or "play". The skip to the next screen command is recognized as voice input or text input containing time information. The time information is extracted by matching the time format using regular expressions, supporting both minute / second and pure second representations. The command type is determined by comparing the captured input event data with a predefined command pattern. If a match is found, the command type identifier and command parameters are extracted.

[0160] If a resume playback command is received, video playback resumes from the video position corresponding to the interruption timestamp. The interruption timestamp is stored in a session state variable, which records the current video playback time position with millisecond-level precision and integer precision when the user inputs an interruption request. The video playback resumption operation calls the positioning method of the video playback control interface. The positioning method takes the interruption timestamp as input and executes the video decoder to jump to the keyframe corresponding to that timestamp. The keyframe index is obtained by querying the keyframe timestamp list in the video metadata. The decoder decodes from the keyframe to the precise position of the interruption timestamp, and the decoded frame is buffered to the rendering queue. The resumption operation of the digital human's synchronized narration queries the narration script storage system to retrieve the narration segment corresponding to the interruption timestamp. The segment retrieval condition is that the start timestamp is less than or equal to the interruption timestamp and the end timestamp is greater than the interruption timestamp. The script text of the retrieved narration segment is input to the speech synthesis module, which generates the digital human's narration audio. The audio start time offset is the interruption timestamp minus the start timestamp of the narration segment. The digital human motion generation module drives the virtual avatar generator based on the visual element data of the narration segment, generating a sequence of digital human motion frames. The starting frame index of the motion frame sequence is calculated based on the time offset. The digital human narration audio and the motion frame sequence are synchronously output to the audio playback device and display device, maintaining time alignment with the resumed video stream.

[0161] If a jump command is received, the target time point specified in the command is extracted. The target time point is parsed from the command parameter field, and the result is a millisecond-level timestamp with integer precision. The video playback position jump operation calls the positioning method of the video playback control interface. The positioning method takes the target time point as input and executes the video decoder to jump to the keyframe corresponding to that time point. The decoder decodes from the keyframe to the precise position of the target time point. The jump operation for the digital human's synchronized narration queries the narration script storage system to retrieve the narration segment corresponding to the target time point. The segment retrieval criteria are that the start timestamp is less than or equal to the target time point and the end timestamp is greater than the target time point. The script text of the retrieved narration segment is input to the speech synthesis module, which generates the digital human's narration audio. The audio start time offset is the target time point minus the narration segment start timestamp. The digital human motion generation module drives the virtual avatar generator based on the visual element data of the narration segment to generate a sequence of digital human motion frames. The start frame index of the motion frame sequence is calculated based on the time offset. The digital human's narration audio and motion frame sequence are synchronously output to the audio playback device and display device, maintaining time alignment with the video stream after the jump.

[0162] In one optional implementation, the virtual avatar generator synthesizes a dynamic expression output based on the feature sequence. The dynamic expression output includes a continuous frame sequence of the virtual avatar. Energy function calculations are performed on the pose features of each frame in the continuous frame sequence, and frames whose energy function values ​​reach local maxima are marked as visual anchor points.

[0163] The virtual avatar generator synthesizes a dynamic expression output based on the feature sequence. The dynamic expression output includes a continuous frame sequence of the virtual avatar. Feature decomposition is performed on the posture feature data carried in each frame of the continuous frame sequence, and the joint angle vector and limb end displacement vector of the virtual avatar are extracted from the posture feature data of each frame.

[0164] A posture energy function is constructed based on the joint angle vector and the limb end displacement vector. The angle change rate of the joint angle vector and the displacement amplitude of the limb end displacement vector between adjacent frames are calculated based on the posture energy function. The angle change rate and the displacement amplitude are weighted and combined to obtain the motion intensity value corresponding to each frame.

[0165] The motion intensity value is searched for local extrema on the time axis. By comparing the motion intensity value of each frame with the motion intensity value of its neighboring frames, the frame position where the motion intensity value presents a local maximum value on the time series is identified, and the frame where the motion intensity value reaches a local maximum value is marked as a visual anchor frame.

[0166] When the virtual avatar generator synthesizes dynamic expression output based on feature sequences, it reads the pose feature data stream from the feature sequence buffer. The feature sequence buffer is a shared memory structure with a capacity of four megabytes, supporting multi-process concurrent access. Each feature frame in the feature sequence contains a timestamp field and a pose feature data field. The timestamp has millisecond precision, and the pose feature data field stores a three-dimensional spatial coordinate array of the skeletal nodes. There are twenty-three skeletal nodes, covering the head and neck joints, shoulder, elbow, and wrist joints, spinal joints, hip, knee, and ankle joints, and finger and toe distal nodes. The coordinates of each node are represented by three single-precision floating-point numbers. The coordinate system adopts a right-handed Cartesian coordinate system, with the origin located at the center of the virtual avatar's pelvis, and the unit is centimeters.

[0167] The virtual avatar generator comprises three processing units: a skeletal inverse kinematics (IK) module, a mesh skinning and rendering module, and a frame buffer output module. The IK module receives an array of bone node coordinates, calculates the rotation angles of each joint and the position offset of the end effector, and employs a cyclic coordinate descent iterative algorithm. The iteration terminates when the end effector position error is less than 0.5 cm or the number of iterations reaches the maximum limit of twenty. The mesh skinning and rendering module applies the calculated joint angles to the 3D mesh model of the virtual avatar. The mesh has 12,000 vertices, each bound to a maximum of four bone nodes, with the sum of the binding weights normalized to one. The rendering output uses a physically based shading model, generating image frames with a resolution of 1920 x 1080 pixels, using the standard red-green-blue color space and a color depth of 8 bits per channel. The dynamic output of the continuous frame sequence is generated according to the timestamp order of each frame in the feature sequence. Each input feature frame corresponds to one output image frame. The continuous frame sequence is stored in a video frame queue with a maximum capacity of 1,000 frames, and the queue data structure is a circular buffer.

[0168] When performing feature decomposition on the pose feature data carried by each frame in a continuous frame sequence, image frames and their associated pose feature data are read from the video frame queue in timestamp order. The pose feature data contains the three-dimensional coordinates of twenty-three skeletal nodes. The feature decomposition operation converts these coordinate data into two forms: joint angle representation and limb end-effector position representation. The joint angle vector extraction process traverses the skeletal hierarchical tree structure, with the pelvis node as the root node, and each joint node is organized into a tree topology according to parent-child relationships. For each joint node, its local rotation angle relative to its parent node is calculated. The local rotation angle is obtained by subtracting the parent node's coordinates from the child node's coordinates to obtain the direction vector. The direction vector is then subjected to cross product and dot product operations with the parent node's local coordinate system basis vectors to obtain the rotation axis and rotation angle.

[0169] Rotation angles are represented using Euler angles, comprising rotational components around three coordinate axes. The rotation order is: first, yaw angle around the vertical axis; then, pitch angle around the lateral axis; and finally, roll angle around the forward axis. Euler angles range from -180 degrees to +180 degrees, with a precision of 0.1 degrees. The joint angle vector is a one-dimensional array composed of the Euler angles of all joint nodes, with a length equal to the number of joints multiplied by three, i.e., 69 floating-point elements. The limb end-effector displacement vector extraction process selects predefined end-effector nodes, including four locations: left wrist, right wrist, left ankle, and right ankle. For each end-effector node, its global displacement vector relative to the pelvic node is calculated. The global displacement vector is the end-effector node coordinates minus the pelvic node coordinates, yielding the three-dimensional displacement components. The limb end-effector displacement vector is a one-dimensional array composed of the displacement components of the four end-effector nodes, with a length of twelve floating-point elements.

[0170] When constructing a posture energy function based on joint angle vectors and limb displacement vectors, the posture energy function is defined as a numerical index quantifying the intensity of dynamic changes in the virtual avatar's posture. The energy function calculation involves two components: the rate of change of joint angle vectors between adjacent frames and the displacement amplitude of limb displacement vectors. The rate of change of angles is calculated as the difference between the joint angle vectors of the current frame and the joint angle vectors of the previous frame. This difference is represented by a difference array of corresponding elements, with a length of 69. The absolute values ​​of each element in the difference array are summed to obtain the total change in joint angles, in degrees. The total change in joint angles is divided by the time interval between adjacent frames to obtain the rate of change of angles, in degrees per second. The time interval is the current frame timestamp minus the previous frame timestamp, in milliseconds, with integer precision, typically ranging from 20 to 50 milliseconds. The displacement amplitude is calculated as the difference between the limb displacement vectors of the current frame and the limb displacement vectors of the previous frame. This difference is represented by a difference array of corresponding elements, with a length of 12. The Euclidean norm is calculated for every three elements in the difference array forming a three-dimensional displacement difference vector, yielding the displacement distances of the four end effectors in centimeters. The sum of the four displacement distances gives the total displacement amplitude, which is then divided by the time interval between adjacent frames to obtain the displacement velocity in centimeters per second.

[0171] After calculating the rate of change of joint angle vectors and the displacement amplitude of limb end-effector displacement vectors between adjacent frames based on the posture energy function, a weighted combination operation is performed on the rate of change of angle and the displacement amplitude. The weighted combination operation multiplies the rate of change of angle by an angle weighting coefficient and the displacement velocity by a displacement weighting coefficient, and then adds the two products to obtain the motion intensity value. The default value of the angle weighting coefficient is 0.4, with a range of 0 to 1 and a precision of 0.01. The default value of the displacement weighting coefficient is 0.6, with a range of 0 to 1 and a precision of 0.01. The sum of the angle weighting coefficient and the displacement weighting coefficient should equal one to ensure the consistency of the motion intensity value across different scenarios. The unit of the motion intensity value is a dimensionless value, and its range is related to the actual values ​​of the rate of change of angle and displacement velocity, typically ranging from 0 to 500. The motion intensity value is stored in a motion intensity array, with array indices corresponding one-to-one with the frame indices in the video frame queue. The array length is equal to the capacity of the video frame queue, consisting of one thousand floating-point elements. During the calculation of motion intensity, the motion intensity value of the first frame in the video frame queue is set to zero because there is no previous frame for reference.

[0172] When searching for local maxima in motion intensity values ​​on the time axis, a sliding window traversal strategy is used to scan the motion intensity array. The sliding window length is the number of local neighboring frames, with a default value of eleven frames and a range of five to twenty-one odd numbers with integer precision. Setting the window length to an odd number ensures that the center position of the window is unique and clear. The sliding window starts from the beginning of the motion intensity array and moves one element at a time until the right boundary of the window reaches the end of the array. For each window position, the motion intensity value of the center frame of the window is extracted as a candidate value, and the motion intensity values ​​of all other frames within the window are extracted as a reference value set. By comparing the motion intensity values ​​of each frame with the motion intensity values ​​of its neighboring frames, it is determined whether the candidate value is a local maximum. The determination condition is that the candidate value is greater than all elements in the reference value set. If the condition is met, the frame corresponding to the candidate value is identified as the local maximum frame position. The identification results of the local maximum frame positions are recorded in the visual anchor frame index list. The list elements are the index values ​​of the frames in the video frame queue, and the index values ​​are non-negative integers ranging from zero to the queue capacity minus one.

[0173] When identifying frame locations where motion intensity values ​​exhibit local maxima in the time series, a boundary handling strategy is applied to special cases near the sliding window boundaries. When the left boundary of the sliding window exceeds the starting position of the motion intensity array, the number of valid reference values ​​within the window is less than the number of local neighbor frames; in this case, only the reference values ​​actually existing within the window are used for comparison. Similarly, when the right boundary of the sliding window exceeds the end position of the motion intensity array, only the reference values ​​actually existing within the window are used for comparison. For adjacent frames with completely equal motion intensity values, local extremum search does not mark any of them as local maxima, avoiding the generation of redundant visual anchor frames in static or uniform motion segments.

[0174] When a frame whose motion intensity value reaches a local maximum is marked as a visual anchor frame, each index value in the visual anchor frame index list is associated with the corresponding image frame and pose feature data in the video frame queue. The visual anchor frame carries a flag field, which is a Boolean type. A true value indicates that the frame is a visual anchor frame, and a false value indicates that the frame is a normal frame. The marking operation iterates through the visual anchor frame index list, setting the flag field of the video frame queue element corresponding to each index value in the list to true. The marking result of the visual anchor frame is used in subsequent timing alignment and synchronization output processes, and the flag field is persistently stored as a frame attribute in the metadata area of ​​the video frame queue.

[0175] The method further includes:

[0176] The video file input module receives user-uploaded video files or video links, supporting various mainstream video formats including MP4, AVI, MOV, and WMV. The input module uses the FFmpeg decoding library to parse the video file format and separate the data stream, extracting the audio and video tracks into independent data buffers. Audio data is stored in PCM format, with a sampling rate normalized to 16kHz or 44.1kHz and a quantization bit depth of 16 bits to ensure compatibility and consistent quality in subsequent audio processing.

[0177] Multimodal AI models comprehensively analyze and process input videos. The visual processing branch employs Convolutional Neural Networks (CNNs) or Visual Transformers (ViTs) architectures to analyze video frame sequences frame by frame. CNN models use ResNet or EfficientNet backbones to extract image features, identifying scene elements, object instances, and actions through multi-layer convolutional operations. ViT models segment images into fixed-size patches, with each patch serving as a sequence element input to the transformer encoder, capturing global spatial relationships through a self-attention mechanism. After visual feature extraction, high-dimensional feature vectors are formed, typically with 512 or 1024 dimensions.

[0178] The audio processing branch converts the audio signal into a Mel spectrogram representation. The Mel spectrogram uses a short-time Fourier transform to convert the time-domain audio signal into a frequency-domain representation, with the frequency axis non-linearly mapped according to the Mel scale to simulate the human ear's perception of different frequencies. The time window length of the spectrogram is set to 25 milliseconds, the frame shift to 10 milliseconds, and the number of Mel filter banks to 80. The generated Mel spectrogram is used as input to a CNN model, which extracts the spectral feature patterns of the audio through a two-dimensional convolutional layer.

[0179] The voice recognition module employs a deep learning classifier to detect and classify audio content. The classifier uses a pre-trained audio classification model trained on a large-scale audio dataset, capable of distinguishing between different audio types such as human voices, music, and ambient noise. The model outputs the probability distribution of each audio segment's category. The probability threshold for the human voice category is set to 0.7; audio segments exceeding this threshold are marked as human voice segments. The time stamping of human voice segments is accurate to the millisecond level, recording the start and end times to establish a precise mapping between the audio timeline and the location of the human voice.

[0180] The voiceprint recognition algorithm is designed to distinguish speakers in multi-person speech scenarios. It employs a deep embedding network to extract speaker feature vectors. The network structure includes multiple one-dimensional convolutional layers and recurrent neural network layers. After pre-emphasis and frame segmentation, the input audio is processed to extract Mel-frequency cepstral coefficients (MFCC) features, with a feature dimension of 39. The deep network maps the MFCC features to 128-dimensional or 256-dimensional speaker embedding vectors, which characterize the speaker's acoustic properties. The embedding vectors of different speakers are compared using cosine distance, with a distance threshold of 0.6. Audio segments with similarity higher than the threshold are classified as belonging to the same speaker.

[0181] The digital human synthesis module drives the generation of synchronized animations for a 3D virtual character based on recognized human voice information. The 3D digital human model adopts a standard human skeletal binding structure, including major skeletal nodes such as the head, neck, torso, and limbs. Facial expression control uses the FACS facial motion coding unit, defining basic motion units for facial areas such as eyebrows, eyes, nose, and mouth. The lip-sync algorithm analyzes the phoneme sequence of the human voice, mapping each phoneme to a corresponding lip shape. The lip shape library contains standard lip shapes for vowels A, E, I, O, U, and consonants. Animation playback uses keyframe interpolation technology, setting the target state of the skeleton and expression at key time points on the audio timeline. Intermediate frames generate smooth animation transitions through linear interpolation or Bézier curve interpolation.

[0182] The interactive detection module continuously monitors ambient audio signals to enable user voice interruption. Audio sampling employs real-time streaming processing, acquiring audio segments every 50 milliseconds with a segment length set to 200 milliseconds. The Voice Activation Detection (VAD) algorithm analyzes the energy and spectral characteristics of audio segments to determine if they contain valid human voices. Energy detection calculates the short-time energy of the audio signal, setting an energy threshold of -40dB; segments exceeding this threshold are marked as potential speech. Spectral detection analyzes the frequency distribution of the audio; human voice frequencies are primarily concentrated between 85Hz and 255Hz, and segments with more than 60% energy within this frequency range are identified as human voice segments.

[0183] The speech recognition engine converts user speech input into text content. The engine employs an end-to-end deep learning model, combining acoustic and language models for joint optimization. The acoustic model uses an attention-based sequence-to-sequence network to directly map audio feature sequences to character or word sequences. The language model uses a large-scale pre-trained transformer network to provide contextual semantic information and grammatical constraints, improving recognition accuracy. The recognition process supports streaming decoding, enabling real-time output of partial recognition results while the user is speaking, reducing response latency.

[0184] The knowledge base construction module performs deep understanding and structured storage of video content. Video understanding employs multimodal fusion technology to analyze and correlate visual, audio, and textual features. Scene understanding identifies the environmental background, object layout, and spatial relationships in the video, generating scene description text. Action understanding analyzes characters' actions and interactions, extracting information such as action type, duration, and related objects. Semantic understanding, based on natural language processing technology, performs semantic parsing on audio-transcribed text, identifying key concepts, entity relationships, and logical structures.

[0185] The RAG vector retrieval enhancement module converts knowledge base content into vector representations, supporting semantic similarity retrieval. Text content is encoded into high-dimensional vectors using a pre-trained language model, typically with 768 or 1024 dimensions. Vector storage employs efficient approximate nearest neighbor search algorithms, such as FAISS or Annoy index structures, supporting fast retrieval of large-scale vectors. User query text is also encoded as a vector; the most relevant knowledge fragments are retrieved by calculating the cosine similarity between the query vector and the knowledge base vectors. Retrieval results are sorted by similarity score, and the top K most relevant fragments are selected as reference content for generating the answer.

[0186] The large language model generation module generates personalized answers based on retrieved knowledge content. The generation model employs pre-trained large-scale language models, such as the GPT or BERT series, possessing powerful text understanding and generation capabilities. Answer generation utilizes a context-aware strategy, combining user questions, retrieved knowledge, and dialogue history to generate coherent responses. The answer structure includes three parts: preceding context transitions, the main body of the answer, and following context transitions, ensuring a natural transition between the answer and video playback. The generation process employs beam search or kernel sampling algorithms to balance accuracy and diversity in the answers.

[0187] The text-to-speech module converts the generated response text into natural speech output. Speech synthesis employs neural network vocoder technology, such as WaveNet or Tacotron models, to generate high-quality audio that closely resembles human speech. The synthesis process comprises four stages: text preprocessing, phoneme conversion, prosodic modeling, and waveform generation. Text preprocessing normalizes the input text, handling numbers, abbreviations, and special symbols. Phoneme conversion converts text into corresponding phoneme sequences. Prosodic modeling predicts the pitch, stress, and pause patterns of the speech. Waveform generation synthesizes the final audio waveform based on phoneme and prosodic information.

[0188] The playback control module coordinates the timing synchronization of video playback, digital human animation, and voice interaction. The module maintains a global timeline, recording the playback progress of the original video, the playback status of the digital human animation, and the time points of user interactions. When a user interrupts voice input, the current playback time is recorded, and the original video playback is paused. After the digital human responds, playback resumes based on the recorded time point, ensuring playback continuity. The synchronization mechanism employs frame-precise time control, with playback errors controlled within the time range of one video frame, typically 33 milliseconds or 16 milliseconds.

[0189] A second aspect of the present invention provides an electronic device, comprising:

[0190] processor;

[0191] Memory used to store processor-executable instructions;

[0192] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.

[0193] A third aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0194] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0195] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for converting traditional videos into interactive, automatically narrating AI-powered digital humans, characterized in that: include: The original video data and its corresponding audio track are acquired. The original video data is decomposed temporally to extract the video frame sequence and identify the main objects and background areas. At the same time, the audio track is semantically parsed to obtain multimodal deconstructed data containing visual and semantic elements. Based on the explanatory text in the multimodal deconstruction data, a corresponding explanatory script is generated for each time period of the video content, and the visual elements in the video frame sequence corresponding to each explanatory segment are timestamped to form a data structure that synchronizes video playback with the explanatory content in time. During video playback, the virtual avatar generator synthesizes the dynamic expression output of the digital human in real time according to the narration script corresponding to the current playback time point, so that the digital human can narrate synchronously with the video content; The system receives an interruption request from the user, which includes a query intent and control instructions for the content to be explained. The query intent is semantically matched in the explanation script to locate the target explanation segment and its associated visual elements related to the query intent, and supplementary explanation content is generated based on the context of the target explanation segment. Based on the visual elements associated with the target explanation segment, the virtual avatar generator is driven to synthesize a dynamic expression output synchronized with the supplementary explanation content. After the interactive explanation is completed, the video playback and digital avatar synchronized explanation are resumed according to the user's instructions, or the playback is redirected to the video time point specified by the user to continue.

2. The method according to claim 1, characterized in that, The original video data is temporally decomposed to extract video frame sequences and identify the main objects and background regions within them. Simultaneously, the audio track is semantically parsed to obtain multimodal deconstructed data containing both visual and semantic elements, including: The original video data is segmented into a video frame sequence according to the timeline, and each frame in the video frame sequence is spatially divided to distinguish the foreground region and the background region; the main object in the foreground region is detected and its attributes are extracted. The attribute extraction includes identifying the contour boundary features, pose features and motion trajectory features of the main object, and establishing a cross-frame tracking identifier for the main object. Acoustic signal processing is performed on the audio track to convert continuous speech signals into discrete text segments. Natural language understanding is performed on the discrete text segments to extract semantic units and their syntactic structures. The discrete text segments are timestamped with the video frame sequence. Based on the cross-frame tracking identifier, the main object is located in the aligned video frames. The visual state of the main object when it emits the corresponding semantic unit is determined, and the binding relationship between the semantic unit and the visual state is established. The contour boundary features, pose features, motion trajectory features, and scene features of the background region are summarized as visual elements, the semantic units, syntactic structures, and timestamp information are summarized as semantic elements, and the binding relationship is used as the association index between the visual elements and the semantic elements to form multimodal deconstruction data.

3. The method according to claim 1, characterized in that, Based on the explanatory text in the multimodal deconstruction data, a corresponding explanatory script is generated for each time segment of the video content, and the visual elements in the video frame sequence corresponding to each explanatory segment are timestamped to form a data structure that synchronizes video playback with the explanatory content in time, including: The explanatory text is extracted from the multimodal deconstruction data, and the explanatory text is divided into multiple explanatory segments according to the timeline of video playback. Each explanatory segment corresponds to a continuous time period in the video, and each explanatory segment is marked with a start timestamp and an end timestamp. Semantic analysis is performed on each of the explanatory segments to identify important information points and key points of explanation. Based on the important information points, extended explanatory content is generated, and the original explanatory text and the extended explanatory content are integrated to form a complete explanatory script. The video frame sequence and its associated visual elements are obtained from the multimodal deconstruction data. The binding relationship in the multimodal deconstruction data is used to determine the timestamp range corresponding to each narration segment. The corresponding visual elements are extracted from the video frame sequence within the timestamp range. The narration script, the visual elements, and the video frame sequence are aligned in three dimensions according to timestamps to construct a synchronized data structure that includes a video playback timeline, a narration content timeline, and a visual element timeline, ensuring that video playback and digital human narration are synchronized in time.

4. The method according to claim 1, characterized in that, The system receives a user-input interruption request, which includes a query intent and control instructions regarding the content to be explained. It performs semantic matching of the query intent within the explanation script to locate the target explanation segment and its associated visual elements related to the query intent. Based on the contextual relationships of the target explanation segment, it generates supplementary explanation content, including: During video playback, the system monitors the user input channel in real time, or checks whether the user has pressed a button. When an interruption request is detected, the system records the current video playback time as the interruption timestamp and sends a pause command to the video playback module and the digital human narration module, so that the video playback and digital human narration are paused synchronously. The interruption request is semantically parsed to extract the core concepts in the query intent. The core concepts are then semantically similar to the explanatory segments in the explanatory script to locate the target explanatory segment related to the query intent and obtain the visual elements mapped to the target explanatory segment. Based on the position of the target explanation segment on the timeline, the explanation script is expanded forward and backward to extract context-related explanation segments. By analyzing the semantic relationship between the target explanation segment and adjacent explanation segments, a context explanation sequence containing cause and effect is constructed. The contextual explanation sequence is verified for completeness. If any missing information is detected, a bridging explanation segment that can fill in the missing information is searched from other time periods of the explanation script. The bridging explanation segment and its mapped visual elements are added to the contextual explanation sequence to form a logically complete supplementary explanation content.

5. The method according to claim 4, characterized in that, Based on the position of the target explanation segment on the timeline, context-related explanation segments are extracted by extending the explanation script forward and backward. By analyzing the semantic relationships between the target explanation segment and adjacent explanation segments, a contextual explanation sequence containing cause and effect is constructed, including: Using the start timestamp of the target explanation segment as a reference point, the explanation script is backtracked forward along the timeline to extract adjacent explanation segments with timestamps earlier than the target explanation segment. The semantic dependencies between the adjacent explanation segments and the target explanation segment are analyzed, and the preceding explanation segments that serve as prerequisites for the target explanation segment are identified to form a forward context sequence. Using the end timestamp of the target explanation segment as a reference point, a forward search is performed in the explanation script to the back of the timeline to extract adjacent explanation segments with timestamps later than the target explanation segment. The semantic continuity relationship between the adjacent explanation segments and the target explanation segment is analyzed, and the subsequent explanation segments derived from the target explanation segment are identified to form a backward context sequence. The forward context sequence and the backward context sequence are integrated in timestamp order to construct a contextual explanation sequence centered on the target explanation segment.

6. The method according to claim 1, characterized in that, Based on the visual elements associated with the target explanation segment, the virtual avatar generator is driven to synthesize a dynamic expression output synchronized with the supplementary explanation content. After completing the interactive explanation, video playback and synchronized digital avatar explanation are restored according to user instructions, including: The associated visual elements are obtained from the target explanation segment and the context explanation sequence. The pose features and facial expression features in the visual elements are modeled as inter-frame motion trajectories. The original feature frames and intermediate transition frames are combined in time sequence to form a smooth motion feature sequence. The virtual avatar generator synthesizes a dynamic expression output based on the feature sequence. The dynamic expression output includes a continuous frame sequence of the virtual avatar. The energy function is calculated for the pose features of each frame in the continuous frame sequence, and the frames whose energy function values ​​reach a local maximum are marked as visual anchor points. A temporal anchor point correspondence table is constructed. By performing prosodic intensity analysis on the supplementary explanation content, the syllable stress position and semantic key words are extracted as semantic anchor points. The semantic anchor points and the visual anchor points are matched with semantic relevance scores, and the time difference between the semantic anchor points and the visual anchor points is calculated. Based on the time difference, the continuous frame sequence is segmented into nonlinear time mappings to align the visual anchor point with the semantic anchor point on the time axis. The aligned continuous frame sequence is then output synchronously with the supplementary explanation content to complete the interactive explanation. After the interactive explanation is completed, the system receives a user's command to resume playback or jump to another location. If a command to resume playback is received, the system resumes video playback from the video position corresponding to the interrupted timestamp and drives the digital human to continue synchronous explanation according to the explanation script corresponding to the resumed position. If a command to jump to another location is received, the system extracts the target time point specified in the jump command, jumps the video playback position to the target time point, and drives the digital human to start synchronous explanation according to the explanation script corresponding to the target time point.

7. The method according to claim 6, characterized in that, The virtual avatar generator synthesizes a dynamic expression output based on the feature sequence. The dynamic expression output includes a continuous frame sequence of the virtual avatar. Energy functions are calculated for the pose features of each frame in the continuous frame sequence, and frames whose energy function values ​​reach local maxima are marked as visual anchor points. The virtual avatar generator synthesizes a dynamic expression output based on the feature sequence. The dynamic expression output includes a continuous frame sequence of the virtual avatar. Feature decomposition is performed on the posture feature data carried in each frame of the continuous frame sequence, and the joint angle vector and limb end displacement vector of the virtual avatar are extracted from the posture feature data of each frame. A posture energy function is constructed based on the joint angle vector and the limb end displacement vector. The angle change rate of the joint angle vector and the displacement amplitude of the limb end displacement vector between adjacent frames are calculated based on the posture energy function. The angle change rate and the displacement amplitude are weighted and combined to obtain the motion intensity value corresponding to each frame. The motion intensity value is searched for local extrema on the time axis. By comparing the motion intensity value of each frame with the motion intensity value of its neighboring frames, the frame position where the motion intensity value presents a local maximum value on the time series is identified, and the frame where the motion intensity value reaches a local maximum value is marked as a visual anchor frame.

8. A system for converting traditional videos into interactive, automated explanations using artificial intelligence digital humans, for implementing the method described in any one of claims 1-7, characterized in that, include: The first unit is used to acquire the original video data and its corresponding audio track, perform temporal decomposition on the original video data, extract the video frame sequence and identify the main object and background area in it, and perform semantic parsing on the audio track to obtain multimodal deconstruction data containing visual elements and semantic elements. The second unit is used to generate corresponding explanation scripts for each time period of the video content based on the explanation text in the multimodal deconstruction data, and to timestamp the visual elements in the video frame sequence corresponding to each explanation segment to form a data structure that synchronizes the video playback with the explanation content in time. The third unit is used to drive the virtual image generator to synthesize the dynamic expression output of the digital human in real time according to the narration script corresponding to the current playback time point during video playback, so that the digital human can narrate synchronously with the video content. The fourth unit is used to receive interruption requests input by the user. The interruption request includes a query intent and control instructions for the content to be explained. The query intent is semantically matched in the explanation script to locate the target explanation segment and its associated visual elements related to the query intent, and supplementary explanation content is generated based on the context of the target explanation segment. The fifth unit is used to drive the virtual avatar generator to synthesize a dynamic expression output that is synchronized with the supplementary explanation content based on the visual elements associated with the target explanation segment. After the interactive explanation is completed, the video playback and digital avatar synchronized explanation are restored according to the user's instructions, or the video playback is jumped to the time point specified by the user to continue.

9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Digital person teacher online teaching service device

    CN117333327A

  • Video courseware generation method and device based on artificial intelligence, equipment and medium

    CN119697457A