Digital human, PPT and picture frame level synchronization method and system
By extracting time-series data from PPT and/or PDF files and combining computer vision and speech recognition technologies, high-precision synchronization between digital humans and PPT and PDF files is achieved, solving the problems of low synchronization efficiency and occlusion in existing technologies. It is suitable for education, finance, or government scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WU HAN LUO JI ZHI HUI KE JI YOU XIAN GONG SI
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-01
AI Technical Summary
Existing digital human broadcasting systems cannot effectively map the semantics of image data back to the actions of digital humans. They suffer from occlusion problems, low processing efficiency, and large errors, making it difficult to meet the highly coupled needs of scenarios such as education, finance, or government affairs.
By acquiring time-series data from PPT and/or PDF files, computer vision technology is used to identify image regions. Combined with speech recognition and a pre-built gesture library, precise alignment of animations, images, and narration is achieved, and the final video stream is output.
It achieves high-precision synchronization between digital humans and PPT and/or PDF files, avoids occlusion, improves processing efficiency and reduces errors, and facilitates application and promotion.
Smart Images

Figure CN121967741A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multimedia information synchronization and digital human interaction technology, specifically relating to a method and system for frame-level synchronization of digital humans, PPT presentations, and images. Background Technology
[0002] Digital humans are digitized human figures created using digital technology, closely resembling human appearances. They encompass the penetration of digital technology into various levels and stages of human anatomy, physics, physiology, and intelligence. Their core value lies in providing anthropomorphic services by breaking physical boundaries through hyper-realism, tool-like features, and strong interactivity. The system framework typically consists of five modules: human image, voice generation, animation generation, audio-visual synthesis and display, and interaction. Existing digital human broadcasting systems can only achieve two-level alignment between text and lip movements or page turning and narration during broadcasting. They cannot reverse-map the semantics of image data to the digital human's actions, nor can they solve the problem of occlusion when the digital human's fingers or eyes overlap with image pixel areas. When a PowerPoint presentation contains multiple window images and has entry or exit animations, traditional timeline editing requires manual frame-by-frame adjustments, resulting in low processing efficiency and significant errors. Synchronization time is typically over 0.3 seconds, making it difficult to meet the high-coupling "pointing-explanation" requirements of scenarios such as education, finance, or government.
[0003] Therefore, how to provide an effective technical solution to address the problems in existing technologies, such as the inability to reverse map the semantics of image data to the actions of digital humans, the inability to solve occlusion problems, low processing efficiency, large errors, and low accuracy, has become an urgent problem to be solved in existing technologies. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for frame-level synchronization of digital humans, PPT presentations, and images, in order to solve the aforementioned problems existing in the prior art.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for frame-level synchronization of a digital human, a PPT presentation, and images, comprising: Obtain PPT and / or PDF files, and extract timing data from the PPT and / or PDF files. The timing data includes animation timing information and image data. Based on the animation timing information, obtain the animation timecode. The process involves acquiring the audio of the narration, performing speech recognition on the audio to obtain multiple word-level or phoneme-level timestamps, obtaining a word-level or phoneme-level timestamp sequence based on these timestamps, using the word-level or phoneme-level timestamp sequence as the narration timecode, and coarsely aligning the animation timecode with the narration timecode to obtain the aligned animation timetable. Computer vision technology is used to identify regions in image data, resulting in multiple ROI triples. Each ROI triple includes a key ROI region, a semantic label, and an importance score. Based on multiple ROI triples and an aligned timeline, an image ROI timecode is obtained. Semantic tags are matched based on a pre-built gesture library to obtain gesture data. Action timecodes are obtained based on gesture data and audio narration. Acquire the digital human video stream, mix it in real time with PPT and / or PDF files, and align the animation timecode, image ROI timecode, and motion timecode based on the narration timecode to output the final video stream.
[0006] In one possible design, timing data is extracted from PPT and / or PDF files, the timing data including animation timing information and image data. Based on the animation timing information, animation timecode is obtained, including: The PPT file is parsed using a preset PPT file parsing library to obtain the first PPT time series data, which includes the first initial image data. And / or use a preset PPT file parsing library to parse the timing definition of the PPT file to extract the triggering conditions, delay time and duration in the PPT file to obtain the second PPT timing data, which includes the initial animation timing information; And / or use a preset PDF file parsing library to parse the PDF file to obtain PDF time-series data, the PDF time-series data including second initial image data; The timing data of the first PPT, the timing data of the second PPT, and / or the timing data of the PDF are converted to obtain timing data, which includes animation timing information and image data. Based on the animation timing information, the start and end times of each animation in the preset PPT timeline are extracted from the PPT file, and then sorted according to the start and end times of each animation in the preset PPT timeline to obtain the animation timecode.
[0007] In one possible design, speech recognition is performed on the narration audio to obtain multiple word-level or phoneme-level timestamps. Based on these timestamps, a word-level or phoneme-level timestamp sequence is obtained. This sequence is then used as the narration timecode. The animation timecode is coarsely aligned with the narration timecode to obtain the aligned animation timetable, which includes: ASR technology is used to perform speech recognition on the narration audio to obtain text content and multiple timestamps. The timestamps are either word-level timestamps or phoneme-level timestamps. The word-level timestamps are used to represent the start and end times of each word in the text content in the narration audio, and the phoneme-level timestamps are used to represent the start and end times of pauses in the narration audio. The text content is combined with multiple timestamps to obtain a timestamp sequence, which is either a word-level timestamp sequence or a phoneme-level timestamp sequence. The word-level timestamp sequence or the phoneme-level timestamp sequence is used as the time code of the explanatory words. Obtain the total page display duration of PPT and / or PDF files, and extract the total audio duration of the narration audio; Based on the total audio duration and the total page display duration, a dynamic time warping algorithm is used to coarsely align the switching points in the animation timecode with the paragraph start points of the narration timecode, resulting in an aligned animation timetable.
[0008] In one possible design, computer vision techniques are used to perform region identification on the image data, resulting in multiple ROI triples. Based on these ROI triples and an aligned timeline, an image ROI timecode is obtained, including: Computer vision technology is used to identify regions in image data to obtain multiple key ROI regions, corresponding semantic labels, and original prediction confidence scores. The semantic labels are used to characterize the content semantics of the key ROI regions. Attention confidence is obtained by scoring key ROI regions based on a visual attention model. The original prediction confidence and attention confidence are weighted according to the preset weighting rules to obtain the importance score. The key ROI region, the corresponding semantic label and the importance score are used as the ROI triple. Based on the aligned timeline, the appearance and disappearance times of the image data corresponding to the ROI triples in the PPT and / or PDF files are extracted to obtain the maximum time window; Within the maximum time window, multiple ROI triples are semantically matched with the aligned time schedule to obtain the image ROI timecode.
[0009] In one possible design, within the maximum time window, multiple ROI triples are semantically matched with the aligned timeline to obtain the image ROI timecode, including: The text content of the commentary timecode is extracted based on the maximum time window to obtain the extracted text content; Semantic matching is performed between semantic tags and extracted text content to obtain multiple semantic similarities, and the timestamp of the explanatory words in the extracted text content corresponding to the highest semantic similarity is extracted; Based on the timestamps of the explanatory words in the extracted text corresponding to the highest semantic similarity, multiple ROI triples are sorted to obtain the image ROI timestamps.
[0010] In one possible design, a digital human video stream is acquired, mixed in real-time with PPT and / or PDF files, and simultaneously aligned with animation timecode, image ROI timecode, and motion timecode based on the narration timecode, outputting the final video stream, including: The process involves acquiring a real-time rendered digital human video stream, which includes a series of consecutive image frames. The RGB image of the current image frame is captured from the digital human video stream, and the RGB image is scaled based on a preset size to obtain a scaled RGB image. The scaled RGB image is then normalized to obtain a preprocessed RGB image. The preprocessed RGB image is input into a lightweight monocular depth estimation model, which outputs a depth map. Each pixel value in the depth map is used to represent the relative depth or relative distance. The depth map is post-processed to obtain the final depth map. The final depth map is then mixed in real time with PPT and / or PDF files. At the same time, the animation timecode, image ROI timecode, and motion timecode are aligned based on the narration timecode, and the final video stream is output.
[0011] In one possible design, the final depth map includes digital human limb depth values corresponding to the depth map; the real-time mixing of the final depth map with PPT and / or PDF files includes: Parse PPT and / or PDF files to obtain multi-layer RGBA images; The simulated depth values of each RGBA image layer are assigned based on the preset layout to obtain the depth values of multiple image layers. The depth values of the digital human limbs are compared with the depth values of multiple image layers. If the difference between the depth values of the digital human limbs and the depth values of the image layers is less than a preset depth threshold, the layer of the final depth map is moved to the top layer. The final depth map is blended with PPT and / or PDF files using a feathering algorithm.
[0012] Secondly, the present invention provides a digital human, PPT, and image frame-level synchronization system, comprising: The timing extraction module is used to acquire PPT and / or PDF files and extract timing data from the PPT and / or PDF files. The timing data includes animation timing information and image data. Based on the animation timing information, the animation timecode is obtained. The audio extraction module is used to acquire the audio of the narration, perform speech recognition on the audio of the narration to obtain multiple word-level timestamps or phoneme-level timestamps, obtain a word-level timestamp sequence or phoneme-level timestamp sequence based on the multiple word-level timestamps or phoneme-level timestamps, use the word-level timestamp sequence or phoneme-level timestamp sequence as the narration timecode, and coarsely align the animation timecode with the narration timecode to obtain the aligned animation timetable. The image recognition module is used to perform region recognition on image data using computer vision technology to obtain multiple ROI triples. The ROI triples include key ROI regions, semantic labels, and importance scores. Based on multiple ROI triples and an aligned timeline, the image ROI timecode is obtained. The gesture matching module is used to match semantic tags based on a pre-built gesture library to obtain gesture data, and to obtain action timecodes based on gesture data and audio narration. The mixed output module is used to acquire the digital human video stream, mix the digital human video stream with PPT and / or PDF files in real time, and align the animation timecode, image ROI timecode and motion timecode based on the narration timecode to output the final video stream.
[0013] Thirdly, the present invention provides a computer device comprising a memory, a processor, and a transceiver connected in sequence and communication, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the digital human, PPT, and image frame-level synchronization method as described in the first aspect above.
[0014] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, perform the digital human, PPT, and image frame-level synchronization method as described in the first aspect above.
[0015] Fifthly, the present invention provides a computer program product containing instructions that, when the instructions are executed on a computer, cause the computer to perform the digital human, PPT, and image frame-level synchronization method as described in the first aspect above.
[0016] The beneficial effects of this invention are as follows: This invention discloses a method and system for frame-level synchronization of digital humans, PPT presentations, and images. It extracts timing data from PPT and / or PDF files, including animation timing information and image data. Based on the animation timing information, an animation timecode is obtained. Then, audio narration is acquired, and speech recognition is performed on the narration audio to obtain a narration timecode. The animation timecode and the narration timecode are coarsely aligned to obtain an aligned animation timetable. Computer vision technology is used to perform region recognition on the image data to obtain multiple ROI triples. Based on the ROI triples and the aligned timetable, image ROI timecodes are obtained. Semantic tags in the ROI triples are matched using a pre-built gesture library to obtain gesture data. Based on the gesture data and the narration audio, action timecodes are obtained. Finally, a digital human video stream is acquired and mixed in real-time with the PPT and / or PDF files. Simultaneously, the animation timecode, image ROI timecode, and action timecode are aligned based on the narration timecode to output the final video stream. This invention achieves initial alignment between animation timecode and narration timecode by coarsely aligning them, eliminating the need for manual adjustments, thus improving processing efficiency and accuracy while reducing errors. It enables real-time mixing of digital human video streams with PPT and / or PDF files, aligning animation timecode, image ROI timecode, and motion timecode based on the narration timecode. Using the narration timecode as the alignment standard, it unifies the time scale and enables the semantic back-mapping of image data to the digital human's actions, achieving a correspondence between actions and animation. This avoids the digital human obscuring text, facilitating application and promotion. Attached Figure Description
[0017] Figure 1 A flowchart illustrating the frame-level synchronization method for digital humans, PPT presentations, and images provided in an embodiment of the present invention; Figure 2 This is a block diagram of a digital human, PPT, and image frame-level synchronization system provided in an embodiment of the present invention. Figure 3 A structural diagram of a computer device provided in an embodiment of the present invention; Figure 4 This is a comparison diagram with the final video stream provided in an embodiment of the present invention. Detailed Implementation
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.
[0019] It should be understood that although the terms first, second, etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit, without departing from the scope of the exemplary embodiments of the invention.
[0020] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.
[0021] Example: like Figure 1 As shown, the first aspect of this embodiment provides a method for frame-level synchronization of digital humans, PPT presentations, and images. This method can be executed, but is not limited to, by a computer device or virtual machine with certain computing resources, such as a personal computer or smartphone, or by a virtual machine. The frame-level synchronization method for digital humans, PPT presentations, and images includes, but is not limited to, the following steps: S1. Obtain PPT and / or PDF files, and extract timing data from the PPT and / or PDF files. The timing data includes animation timing information and image data. Based on the animation timing information, obtain the animation timecode. Specifically, in step S1, timing data is extracted from the PPT and / or PDF files. This timing data includes animation timing information and image data. Based on the animation timing information, the animation timecode is obtained, including: S11. Use a preset PPT file parsing library to parse the PPT file to obtain the first PPT timing data, which includes the first initial image data; S12. And / or use a preset PPT file parsing library to parse the timing definition of the PPT file to extract the triggering conditions, delay time and duration in the PPT file to obtain the second PPT timing data, the second PPT timing data including the initial animation timing information; S13. And / or use a preset PDF file parsing library to parse the PDF file to obtain PDF time-series data, wherein the PDF time-series data includes second initial image data; S14. Convert the format of the first PPT timing data, the second PPT timing data and / or PDF timing data to obtain timing data, wherein the timing data includes animation timing information and image data; S15. Based on the animation timing information, extract the start and end times of each animation in the preset PPT timeline from the PPT file, and sort them according to the start and end times of each animation in the preset PPT timeline to obtain the animation timecode.
[0022] It should be noted that the default PPT (PowerPoint) file parsing library in this embodiment includes the Open XML SDK parsing library. The Open XML SDK is a powerful tool library developed by Microsoft, specifically designed to manipulate Office documents that conform to the Office Open XML file format specification. For example, the python-pptx library in the Open XML SDK parsing library can directly read the XML structure in the PPT file to obtain the first PPT timing data. The first PPT timing data includes, but is not limited to, page sequence, shape list, embedded images, and text content. The second PPT timing data includes, but is not limited to, trigger conditions, delay time, and duration. Trigger conditions include "after the previous animation" or "simultaneously with the previous animation". The default PDF file parsing library in this embodiment includes, but is not limited to, Poppler, PDF.js, or iWork parsing library. The timing data includes, but is not limited to, page metadata and media content. Page metadata includes page size, background, and page number order. Media content includes cropping or locating all image elements on the page and recording the coordinates and display time of all image elements on each page. All image elements include images filled within shapes.
[0023] This embodiment performs format conversion on the first PPT timing data, the second PPT timing data, and / or PDF timing data, and uniformly converts the formats of the first PPT timing data, the second PPT timing data, and / or PDF timing data into an intermediate JSON description file containing hierarchy, position, and timing information, namely the timing data of S14.
[0024] S2. Obtain the audio of the narration, perform speech recognition on the audio of the narration to obtain multiple word-level timestamps or phoneme-level timestamps, obtain a word-level timestamp sequence or phoneme-level timestamp sequence based on the multiple word-level timestamps or phoneme-level timestamps, use the word-level timestamp sequence or phoneme-level timestamp sequence as the narration timecode, and coarsely align the animation timecode with the narration timecode to obtain the aligned animation timetable; Specifically, in step S2, speech recognition is performed on the narration audio to obtain multiple word-level or phoneme-level timestamps. Based on these timestamps, a word-level or phoneme-level timestamp sequence is obtained. This sequence is then used as the narration timecode. The animation timecode is coarsely aligned with the narration timecode to obtain the aligned animation timetable, including: S21. Use ASR technology to perform speech recognition on the narration audio to obtain text content and multiple timestamps. The timestamps are word-level timestamps or phoneme-level timestamps. The word-level timestamps are used to characterize the start and end times of each word in the text content in the narration audio. The phoneme-level timestamps are used to characterize the start and end times of pauses in the narration audio. S22. Combine the text content with multiple timestamps to obtain a timestamp sequence, wherein the timestamp sequence is a word-level timestamp sequence or a phoneme-level timestamp sequence, and use the word-level timestamp sequence or the phoneme-level timestamp sequence as the explanatory word timestamp code; S23. Obtain the total page display duration of the PPT file and / or PDF file, and extract the total audio duration of the narration audio; S24. Based on the total audio duration and the total page display duration, a dynamic time warping algorithm is used to coarsely align the switching points in the animation timecode with the paragraph start points of the narration timecode to obtain the aligned animation timetable.
[0025] It should be noted that ASR technology is a speech recognition technology, short for Automatic Speech Recognition. Its goal is to convert the lexical content of human speech into computer-readable input, such as binary codes or character sequences. By using ASR technology to perform speech recognition on the audio of the narration, the text content and the start and end times of the words corresponding to the text content or the start and end times of the pauses corresponding to the text content in the audio of the narration are obtained. Dynamic Time Warping (DTW) is a method to measure the similarity between two time series and to perform coarse alignment of the two time series by calculating the similarity between them.
[0026] In practice, ASR technology is used to perform speech recognition on the audio of the narration, obtaining the text content and multiple timestamps. The text content and multiple timestamps are then combined to output a timestamp sequence. At this time, the timestamp sequence is a structured list, for example, [{"word":"First Quarter","start":1250,"end":1850},...], with the time unit being milliseconds. The timestamp sequence generated at this time is used as the time code of the narration. The generated time code of the narration will serve as the basis for the display of subtitles and the main timeline for subsequent alignment.
[0027] S3. Use computer vision technology to perform region recognition on image data to obtain multiple ROI triples. The ROI triples include key ROI regions, semantic labels, and importance scores. Based on multiple ROI triples and the aligned timeline, obtain the image ROI timecode. Specifically, in step S3, computer vision technology is used to perform region identification on the image data to obtain multiple ROI triples. Based on the multiple ROI triples and the aligned timeline, the image ROI timecode is obtained, including: S31. Use computer vision technology to perform region recognition on image data to obtain multiple key ROI regions, corresponding semantic labels and original prediction confidence scores, wherein the semantic labels are used to characterize the content semantics of the key ROI regions; S32. Based on the visual attention model, perform attention scoring on key ROI regions to obtain attention confidence; S33. The original prediction confidence and attention confidence are weighted based on the preset weighting rules to obtain the importance score. The key ROI region, the corresponding semantic label and the importance score are used as the ROI triple. S34. Based on the aligned timeline, extract the appearance and disappearance times of the image data corresponding to the ROI triples in the PPT and / or PDF files to obtain the maximum time window; S35. Within the maximum time window, perform semantic matching between multiple ROI triples and the aligned time schedule to obtain the image ROI timecode.
[0028] It should be noted that the computer vision technology used in this embodiment is the Vision-Language Model (VLM), a multimodal model that can learn from both images and text. The key ROI region is used to represent the area in the image data that needs to be focused on or pointed to by the digital human. Specifically, it takes the form of a bounding box or a polygon mask. The coordinate values in the bounding box are normalized values relative to the width and height of the image data, with a value range of [0, 1], which facilitates compatibility with rendering layers of different resolutions. The polygon mask is used to indicate irregular shapes, such as map outlines or specific objects. It can be defined by a set of vertex coordinate sequences and may also include associated metadata, representing the page index of the key ROI region in the PPT file and its layer hierarchy information on that page. The visual attention model enhances the information processing of key regions and suppresses irrelevant backgrounds by simulating the selective focusing ability of the human visual system and dynamically allocating computing resources. By learning a weight distribution, it weights different regions, channels, or features of the input image to highlight important parts.
[0029] In this embodiment, semantic tags are one or more text tags that describe the semantics of the content of key ROI regions. They are mainly categories of key ROI regions predicted by computer vision technology, such as bar charts, rising arrows, faces, product logos, and key numbers. Optional attribute semantic tags, such as red bars and the largest sector, can be further output, which can transform visual content into logical information, thereby realizing the connection between vision and action. Furthermore, the importance score is a floating-point number, ranging from [0, 1], representing the prominence of the current key ROI region within the image data context. The original prediction confidence and attention confidence are weighted using preset weighting rules to obtain the final importance score. These preset weighting rules are set by the operator based on the domain or content. For example, in a business report, if the importance of bar charts and trend arrows is higher than that of decorative charts, then the weights of the original prediction confidence and attention confidence corresponding to those bar charts and trend arrows are increased. Similarly, if the importance of a larger key ROI region is higher than that of a smaller key ROI region, then the weights of the original prediction confidence and attention confidence of the larger key ROI region are increased to obtain the importance score. This facilitates the generation of corresponding gestures for key ROI regions with importance scores higher than the preset importance threshold, and allows for focused attention during time allocation.
[0030] In a preferred embodiment, within a maximum time window, multiple ROI triples are semantically matched with the aligned timeline to obtain the image ROI timecode, including: S35.1. Extract the text content from the time code of the commentary based on the maximum time window to obtain the extracted text content; S35.2. Perform semantic matching between semantic tags and extracted text content to obtain multiple semantic similarities, and extract the timestamp of the explanatory words in the extracted text content corresponding to the highest semantic similarity. S35.3. Sort multiple ROI triples based on the timestamps of the explanatory words in the extracted text content corresponding to the highest semantic similarity to obtain the image ROI timestamps.
[0031] It should be noted that in finding the most relevant segments to semantic tags in the extracted text content, natural language processing algorithms can be used to calculate the semantic similarity between the extracted text content and each semantic tag. The key ROI region with the highest semantic similarity is then aligned to the midpoint of the corresponding sentence in the extracted text content. Furthermore, the image ROI timecode is an event sequence ordered by time, not a single point in time, which facilitates synchronized explanation between the digital human and the content.
[0032] S4. Match semantic tags based on a pre-built gesture library to obtain gesture data, and obtain action timecodes based on gesture data and audio narration. In practice, a pre-built gesture library includes multiple gestures, such as pointing, circling, and pushing. By matching the pre-built gesture library with semantic tags, gesture data corresponding to the semantic tags is obtained, and the rhythm of the narration in the audio is extracted. Based on the gesture data corresponding to the semantic tags and the rhythm of the narration, action timecodes are obtained. Furthermore, a pre-built eye-tracking library is used to match the semantic tags to obtain eye-tracking data. Based on the eye-tracking data and the audio of the narration, action timecodes are obtained. The action timecodes, combined with the digital human's gestures and eye movements, enhance the narration effect.
[0033] S5. Acquire the digital human video stream, mix the digital human video stream with PPT and / or PDF files in real time, and align the animation timecode, image ROI timecode, and motion timecode based on the narration timecode to output the final video stream.
[0034] Specifically, in step S5, the digital human video stream is acquired, and then mixed in real time with the PPT and / or PDF files. Simultaneously, the animation timecode, image ROI timecode, and motion timecode are aligned based on the narration timecode, and the final video stream is output, including: S51. Obtain a real-time rendered digital human video stream, the digital human video stream including a series of continuous image frames, capture the RGB image of the current image frame from the digital human video stream, and scale the RGB image based on a preset size to obtain a scaled RGB image, and normalize the scaled RGB image to obtain a preprocessed RGB image. S52. Input the preprocessed RGB image into a lightweight monocular depth estimation model and output a depth map, wherein each pixel value in the depth map is used to represent relative depth or relative distance; S53. Post-process the depth map to obtain the final depth map, and mix the final depth map with the PPT and / or PDF files in real time. At the same time, align the animation timecode, image ROI timecode and motion timecode based on the narration timecode, and output the final video stream.
[0035] It should be noted that in this embodiment, the preset size is 256×256 or 384×384. Then, the scaled RGB image is normalized, and the pixel values of the scaled RGB image are normalized from [0, 255] to [0, 1]. The lightweight monocular depth estimation model is a lightweight version of MiDaS-small, FastDepth, or PINet. In this embodiment, the structure of the lightweight monocular depth estimation model includes an encoder and a decoder. In one possible design, a training dataset is obtained, which includes labeled RGB images with labels of relative depth or relative distance. The lightweight monocular depth estimation model is trained using the training dataset, and the parameters of the lightweight monocular depth estimation model are updated based on the loss function.
[0036] In practice, the output depth map is spatially aligned with the preprocessed RGB image. Smaller pixel values in the depth map indicate closer relative depth or distance, while larger values indicate greater relative depth or distance. In other words, this depth map is equivalent to the depth value distribution map of the digital human's limbs. Post-processing of the depth map includes size restoration, calibration, and filtering. Size restoration involves upsampling the depth map using bilinear interpolation to restore it to the resolution of the digital human video stream, ensuring a one-to-one correspondence with image frame pixels. Calibration involves adjusting the restored depth map... The depth map is mapped to a set Z-value range, such as 0.0-1.0 in this embodiment, where 0.0 represents the top layer of the screen, the position closest to the viewer, and 1.0 represents the bottom layer of the screen, the position farthest from the viewer. This ensures that the velocity values of the digital human and the PPT and / or PDF files are in the same comparable metric space. The filtering operation uses an edge-preserving filtering algorithm to smooth the depth map, reducing noise and preserving the digital human's outline. This avoids jagged edges in the depth information and prevents the digital human from being obscured by image layers in the final video stream. Figure 4 As shown, Figure 4 The image on the left shows the situation where the finger is obscured by the image layer before processing, while... Figure 4 The image on the right shows the final video stream output after remixing, without being obscured by the image layer.
[0037] The final depth map includes the digital human limb depth values corresponding to the depth map; the real-time mixing of the final depth map with PPT and / or PDF files includes: S53.1. Parse PPT and / or PDF files to obtain multi-layer RGBA images; S53.2. Assign values to each RGBA image layer based on the simulated depth values of the preset layout to obtain multiple image layer depth values; S53.3. Compare the depth value of the digital human limb with the depth values of multiple image layers. If the difference between the depth value of the digital human limb and the depth value of the image layers is less than a preset depth threshold, then move the layer of the final depth map to the top layer. S53.4. Blend the final depth map with PPT and / or PDF files based on the feathering algorithm.
[0038] It should be noted that the simulated depth values of the preset layout can be exemplified as follows: the simulated depth value of the title layer is 0.1, the simulated depth value of the image layer is 0.2, the simulated depth value of the background layer is 0.3, and the preset depth threshold is 0.05. The principle of the feathering algorithm is to create a gradient transition area of transparency in the selection area or image edge, so that the opacity of the edge pixels smoothly transitions from completely opaque to completely transparent, thereby reducing the edge contrast and achieving a natural blend with the background, avoiding harsh cutting.
[0039] like Figure 2 As shown, the second aspect of this embodiment provides a digital human, PPT, and image frame-level synchronization system, including: The timing extraction module is used to acquire PPT and / or PDF files and extract timing data from the PPT and / or PDF files. The timing data includes animation timing information and image data. Based on the animation timing information, the animation timecode is obtained. The audio extraction module is used to acquire the audio of the narration, perform speech recognition on the audio of the narration to obtain multiple word-level timestamps or phoneme-level timestamps, obtain a word-level timestamp sequence or phoneme-level timestamp sequence based on the multiple word-level timestamps or phoneme-level timestamps, use the word-level timestamp sequence or phoneme-level timestamp sequence as the narration timecode, and coarsely align the animation timecode with the narration timecode to obtain the aligned animation timetable. The image recognition module is used to perform region recognition on image data using computer vision technology to obtain multiple ROI triples. The ROI triples include key ROI regions, semantic labels, and importance scores. Based on multiple ROI triples and an aligned timeline, the image ROI timecode is obtained. The gesture matching module is used to match semantic tags based on a pre-built gesture library to obtain gesture data, and to obtain action timecodes based on gesture data and audio narration. The mixed output module is used to acquire the digital human video stream, mix the digital human video stream with PPT and / or PDF files in real time, and align the animation timecode, image ROI timecode and motion timecode based on the narration timecode to output the final video stream.
[0040] In a preferred embodiment, a three-timecode synchronization module is further included. This module uses a 48kHz audio sampling clock as a global reference and the animation timecode and image ROI timecode as slave clocks. A PID controller dynamically adjusts the playback rate, reducing the rendering resolution or decreasing the ROI detection frequency when network latency or excessive system load is detected. The three-timecode synchronization module ensures that the deviation converges to a specific value. Within a 1-frame error range, while ensuring real-time synchronization.
[0041] The working process, working details and technical effects of the digital human, PPT and image frame-level synchronization system provided in the second aspect of this embodiment can be found in the digital human, PPT and image frame-level synchronization method described in the first aspect, and will not be repeated here.
[0042] like Figure 3 As shown, the third aspect of this embodiment provides a computer device, including a memory, a processor, and a transceiver connected in sequence for communication. The memory stores a computer program, the transceiver sends and receives messages, and the processor reads the computer program and executes the digital human, PPT, and image frame-level synchronization method as described in the first aspect. Specifically, the memory may include, but is not limited to, random-access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or first-in-last-out (FILO) memory, etc.; the processor may include, but is not limited to, a microprocessor of the STM32F105 series. Furthermore, the computer device may also include, but is not limited to, a power module, a display screen, and other necessary components.
[0043] The working process, working details and technical effects of the aforementioned computer device provided in the third aspect of this embodiment can be found in the digital human, PPT and image frame-level synchronization method described in the first aspect, and will not be repeated here.
[0044] The fourth aspect of this embodiment provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the digital human, PPT, and image frame-level synchronization method described in the first aspect is performed. The computer-readable storage medium refers to a data storage medium, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0045] The working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment can be found in the digital human, PPT and image frame-level synchronization method described in the first aspect, and will not be repeated here.
[0046] The fifth aspect of this embodiment provides a computer program product, including a computer program or instructions, which, when executed by a computer, are used to implement the digital human, PPT, and image frame-level synchronization method as described in the first aspect.
[0047] The working process, working details, and technical effects of the aforementioned computer program product provided in this embodiment can be found in the digital human, PPT, and image frame-level synchronization method described in the first aspect, and will not be repeated here.
[0048] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for frame-level synchronization of digital humans, PPT presentations, and images, characterized in that, include: Obtain PPT and / or PDF files, and extract timing data from the PPT and / or PDF files. The timing data includes animation timing information and image data. Based on the animation timing information, obtain the animation timecode. The process involves acquiring the audio of the narration, performing speech recognition on the audio to obtain multiple word-level or phoneme-level timestamps, obtaining a word-level or phoneme-level timestamp sequence based on these timestamps, using the word-level or phoneme-level timestamp sequence as the narration timecode, and coarsely aligning the animation timecode with the narration timecode to obtain the aligned animation timetable. Computer vision technology is used to identify regions in image data, resulting in multiple ROI triples. Each ROI triple includes a key ROI region, a semantic label, and an importance score. Based on multiple ROI triples and an aligned timeline, an image ROI timecode is obtained. Semantic tags are matched based on a pre-built gesture library to obtain gesture data. Action timecodes are obtained based on gesture data and audio narration. Acquire the digital human video stream, mix it in real time with PPT and / or PDF files, and align the animation timecode, image ROI timecode, and motion timecode based on the narration timecode to output the final video stream.
2. The method for frame-level synchronization of digital human, PPT, and images according to claim 1, characterized in that, Extract timing data from PPT and / or PDF files, wherein the timing data includes animation timing information and image data; based on the animation timing information, obtain the animation timecode, including: The PPT file is parsed using a preset PPT file parsing library to obtain the first PPT time series data, which includes the first initial image data. And / or use a preset PPT file parsing library to parse the timing definition of the PPT file to extract the triggering conditions, delay time and duration in the PPT file to obtain the second PPT timing data, which includes the initial animation timing information; And / or use a preset PDF file parsing library to parse the PDF file to obtain PDF time-series data, the PDF time-series data including second initial image data; The timing data of the first PPT, the timing data of the second PPT, and / or the timing data of the PDF are converted to obtain timing data, which includes animation timing information and image data. Based on the animation timing information, the start and end times of each animation in the preset PPT timeline are extracted from the PPT file, and then sorted according to the start and end times of each animation in the preset PPT timeline to obtain the animation timecode.
3. The method for frame-level synchronization of digital human, PPT, and images according to claim 1, characterized in that, Speech recognition is performed on the narration audio to obtain multiple word-level or phoneme-level timestamps. Based on these timestamps, a word-level or phoneme-level timestamp sequence is obtained. This sequence is then used as the narration timecode. The animation timecode is coarsely aligned with the narration timecode to obtain the aligned animation timetable, including: ASR technology is used to perform speech recognition on the narration audio to obtain text content and multiple timestamps. The timestamps are either word-level timestamps or phoneme-level timestamps. The word-level timestamps are used to represent the start and end times of each word in the text content in the narration audio, and the phoneme-level timestamps are used to represent the start and end times of pauses in the narration audio. The text content is combined with multiple timestamps to obtain a timestamp sequence, which is either a word-level timestamp sequence or a phoneme-level timestamp sequence. The word-level timestamp sequence or the phoneme-level timestamp sequence is used as the time code of the explanatory words. Obtain the total page display duration of PPT and / or PDF files, and extract the total audio duration of the narration audio; Based on the total audio duration and the total page display duration, a dynamic time warping algorithm is used to coarsely align the switching points in the animation timecode with the paragraph start points of the narration timecode, resulting in an aligned animation timetable.
4. The method for frame-level synchronization of digital human, PPT, and images according to claim 1, characterized in that, Computer vision techniques are used to perform region identification on image data, resulting in multiple ROI triples. Based on these ROI triples and an aligned timeline, image ROI timecodes are derived, including: Computer vision technology is used to identify regions in image data to obtain multiple key ROI regions, corresponding semantic labels, and original prediction confidence scores. The semantic labels are used to characterize the content semantics of the key ROI regions. Attention confidence is obtained by scoring key ROI regions based on a visual attention model. The original prediction confidence and attention confidence are weighted according to the preset weighting rules to obtain the importance score. The key ROI region, the corresponding semantic label and the importance score are used as the ROI triple. Based on the aligned timeline, the appearance and disappearance times of the image data corresponding to the ROI triples in the PPT and / or PDF files are extracted to obtain the maximum time window; Within the maximum time window, multiple ROI triples are semantically matched with the aligned time schedule to obtain the image ROI timecode.
5. A method for frame-level synchronization of digital human, PPT, and images according to claim 4, characterized in that, Within the maximum time window, multiple ROI triples are semantically matched with the aligned timeline to obtain the image ROI timecode, including: The text content of the commentary timecode is extracted based on the maximum time window to obtain the extracted text content; Semantic matching is performed between semantic tags and extracted text content to obtain multiple semantic similarities, and the timestamp of the explanatory words in the extracted text content corresponding to the highest semantic similarity is extracted; Based on the timestamps of the explanatory words in the extracted text corresponding to the highest semantic similarity, multiple ROI triples are sorted to obtain the image ROI timestamps.
6. The method for frame-level synchronization of digital human, PPT, and images according to claim 1, characterized in that, Acquire the digital human video stream, blend it in real-time with PPT and / or PDF files, and align the animation timecode, image ROI timecode, and motion timecode based on the narration timecode to output the final video stream, including: The process involves acquiring a real-time rendered digital human video stream, which includes a series of consecutive image frames. The RGB image of the current image frame is captured from the digital human video stream, and the RGB image is scaled based on a preset size to obtain a scaled RGB image. The scaled RGB image is then normalized to obtain a preprocessed RGB image. The preprocessed RGB image is input into a lightweight monocular depth estimation model, which outputs a depth map. Each pixel value in the depth map is used to represent the relative depth or relative distance. The depth map is post-processed to obtain the final depth map. The final depth map is then mixed in real time with PPT and / or PDF files. At the same time, the animation timecode, image ROI timecode, and motion timecode are aligned based on the narration timecode, and the final video stream is output.
7. A method for frame-level synchronization of digital human, PPT, and images according to claim 6, characterized in that, The final depth map includes the digital human limb depth values corresponding to the depth map; the real-time mixing of the final depth map with PPT and / or PDF files includes: Parse PPT and / or PDF files to obtain multi-layer RGBA images; The simulated depth values of each RGBA image layer are assigned based on the preset layout to obtain the depth values of multiple image layers. The depth values of the digital human limbs are compared with the depth values of multiple image layers. If the difference between the depth values of the digital human limbs and the depth values of the image layers is less than a preset depth threshold, the layer of the final depth map is moved to the top layer. The final depth map is blended with PPT and / or PDF files using a feathering algorithm.
8. A digital human, PPT, and image frame-level synchronization system for implementing the method according to any one of claims 1 to 7, characterized in that, include: The timing extraction module is used to acquire PPT and / or PDF files and extract timing data from the PPT and / or PDF files. The timing data includes animation timing information and image data. Based on the animation timing information, the animation timecode is obtained. The audio extraction module is used to acquire the audio of the narration, perform speech recognition on the audio of the narration to obtain multiple word-level timestamps or phoneme-level timestamps, obtain a word-level timestamp sequence or phoneme-level timestamp sequence based on the multiple word-level timestamps or phoneme-level timestamps, use the word-level timestamp sequence or phoneme-level timestamp sequence as the narration timecode, and coarsely align the animation timecode with the narration timecode to obtain the aligned animation timetable. The image recognition module is used to perform region recognition on image data using computer vision technology to obtain multiple ROI triples. The ROI triples include key ROI regions, semantic labels, and importance scores. Based on multiple ROI triples and an aligned timeline, the image ROI timecode is obtained. The gesture matching module is used to match semantic tags based on a pre-built gesture library to obtain gesture data, and to obtain action timecodes based on gesture data and audio narration. The mixed output module is used to acquire the digital human video stream, mix the digital human video stream with PPT and / or PDF files in real time, and align the animation timecode, image ROI timecode and motion timecode based on the narration timecode to output the final video stream.
9. A computer device, characterized in that, The device includes a memory, a processor, and a transceiver that are sequentially and communicatively connected. The memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the digital human, PPT, and image frame-level synchronization method as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or the instructions are executed by the computer, they implement the digital human, PPT and image frame-level synchronization method as described in any one of claims 1 to 7.