Digital human mouth broadcast video generation method, system and device and medium
Through multimodal large model and deep learning technology, combined with three-dimensional modeling and audio-visual synchronization, personalized digital population videos are generated, solving the problems of insufficient flexibility and inefficiency in the existing technology, and achieving high-quality and fast video generation.
Patent Information
- Application Number
- CN202510800917.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-08-12
AI Technical Summary
The existing digital population video generation solution lacks flexibility and is difficult to meet highly personalized needs. The audio and video synchronization is poor, the workflow is complex, the professional skills are high, and the efficiency is low.
The multimodal large model is used for content analysis, the timestamp of copywriting in the video material is determined, and personalized digital people are generated by combining three-dimensional modeling and deep learning technology. High-quality oral video is generated through audio and video synchronization and pinching processing, and multi-threaded computing is used to accelerate the processing process.
It realizes highly customized digital population video generation, precise alignment of audio and video, reduces manual intervention, improves generation efficiency and video quality, and is suitable for diversified application scenarios.
Smart Images

Figure CN120475233A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular relates to a method, system, device and medium for generating spoken video of a digital human. Background Art
[0002] Digital human anchors are virtual characters generated using digital technology. Driven by artificial intelligence technology, they can simulate voice, expression, and movements, thereby acting as anchors to disseminate content such as live broadcasts, reports, and performances.
[0003] Currently, the production of digital human broadcast videos mostly relies on fixed-template video generation tools, independent speech synthesis software, and digital human image generation platforms. Users must use these tools separately for video editing, copywriting, speech synthesis, and subtitle creation, resulting in a fragmented workflow that relies heavily on manual intervention. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method, system, device and medium for generating a digital human oral broadcast video, so as to solve the problem in the related art that the production of the digital human oral broadcast video is cumbersome and inefficient.
[0005] In order to solve the above technical problems, this application provides the following technical solutions:
[0006] In a first aspect, the present application provides a method for generating a spoken video of a digital human, comprising:
[0007] Obtain oral copy and video material data;
[0008] Performing content analysis on the spoken copy and the video material using a multimodal large model to determine the timestamp information corresponding to each sentence in the copy in the video material;
[0009] Converting the spoken text into corresponding audio data, and preprocessing the audio data to obtain preprocessed voice data;
[0010] Merging the pre-processed audio data with the video data according to the timestamp information to generate first video data;
[0011] Generate digital humans based on user needs;
[0012] Performing matting processing on the digital human using image processing technology to obtain a digital human after matting processing;
[0013] The digital human after the matting process is combined with the first video to generate a spoken video of the digital human.
[0014] Furthermore, the preprocessing of the audio data to obtain preprocessed voice data includes:
[0015] The audio data is accelerated or decelerated so that the speech rhythm of the audio data matches the content of the video material.
[0016] Furthermore, the accelerating or decelerating the audio data includes:
[0017] The audio data is accelerated or decelerated using a phase vocoder or an overlap-add algorithm based on waveform similarity.
[0018] Furthermore, after performing matte processing on the digital human using image processing technology to obtain the matte-processed digital human, and before merging the matte-processed digital human with the first video to generate the spoken video of the digital human, the method further includes:
[0019] The edge of the digital human's outline after the matting process is optimized using an edge refinement algorithm.
[0020] Furthermore, the edge thinning algorithm specifically adopts the following calculation formula:
[0021]
[0022] Among them, P′ is the pixel value after expansion, P(i,j) represents the value at position (i,j), and K is the pixel coordinate set.
[0023] Furthermore, the digital human image after the matting process is merged with the first video to generate the digital human's spoken video, using the following calculation formula:
[0024]
[0025] A out =A layer +A in *(1-A layer )
[0026] Among them, C out is the output color; C in is the background image; C layer is the foreground image; A in is the transparency of the background image; A layer is the transparency of the foreground image; A out Output transparency.
[0027] In a second aspect, the present application further provides a system for generating a digital human oral video, comprising:
[0028] The acquisition module is used to obtain the oral copy and video material data;
[0029] A content analysis module, configured to perform content analysis on the oral text and the video material using a multimodal large model, and determine the timestamp information corresponding to each sentence in the text in the video material;
[0030] A preprocessing module, configured to convert the spoken text into corresponding audio data and preprocess the audio data to obtain preprocessed voice data;
[0031] a synthesis module, configured to combine the pre-processed audio data with the video data according to the timestamp information to generate first video data;
[0032] A digital human generation module is used to generate a digital human according to user needs;
[0033] A matting processing module, configured to perform matting processing on the digital human using image processing technology to obtain a digital human after matting processing;
[0034] The spoken video generation module is used to merge the digital human after the matting process with the first video to generate a spoken video of the digital human.
[0035] Furthermore, the system further comprises: a contour optimization module for optimizing the contour edge of the digital human after the matting process by using an edge thinning algorithm.
[0036] In a third aspect, the present application also provides a computer electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above-mentioned advertising bidding methods based on unified constraint bidding are implemented.
[0037] In a fourth aspect, the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method for generating a spoken video of a digital human as described above are implemented.
[0038] The present application provides a method, system, device, and medium for generating spoken video based on a digital human, which have the following beneficial effects:
[0039] First, it breaks the limitations of traditional template-based generation and achieves highly customized production. Digital human images are generated through user-definable input parameters and personalized adjustments are made using advanced 3D modeling and deep learning technologies. This meets users' multi-dimensional customization needs for digital human appearance, movement, style, and other aspects in different scenarios, effectively solving the problem of digital human image homogeneity. Furthermore, in the processing of spoken content and video materials, in-depth content analysis using a multimodal large model enables precise matching of text and video footage. Audio and video elements can be flexibly adjusted based on user needs to achieve customized content presentation, making digital human broadcast videos more suitable for diverse application scenarios such as e-commerce, education, and entertainment.
[0040] Secondly, the accuracy and quality of video generation have been significantly improved. By integrating natural language processing and computer vision technologies with a multimodal large model, the timestamp information of the text in the video material is accurately determined. Advanced audio and video synchronization algorithms are used to ensure precise alignment of audio and video images. In the audio pre-processing and digital human image processing, noise reduction, enhancement, and deep learning-based image processing are used to effectively improve audio clarity and digital human image quality, ultimately generating high-quality spoken-word videos with harmonious audio-visual effects and realistic and natural images.
[0041] Furthermore, video generation efficiency is improved. The solution utilizes multi-threading and distributed computing technologies to accelerate audio and video merging. Combined with automated data acquisition and processing, this reduces manual intervention and significantly shortens the production cycle for digital human broadcast videos. This meets user needs for rapid content iteration, reduces production costs for digital human videos, and enhances the solution's competitiveness and practicality in market applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0043] Figure 1 This is a flow chart of a method for generating a spoken video of a digital human in an embodiment of the present application;
[0044] Figure 2 This is a structural diagram of a system for generating spoken video of a digital human in an embodiment of the present application;
[0045] Figure 3 It is a structural diagram of a computer electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0047] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be an intermediate element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be an intermediate element. Conversely, when an element is referred to as being "directly on" another element, there is no intermediate element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only.
[0048] In this application, unless otherwise expressly specified or limited, terms such as "mounted," "connected," "connect," and "fixed" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integration; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components or interactions between two components. Those skilled in the art will understand the specific meanings of these terms in this application based on specific circumstances.
[0049] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0050] The terms used in one or more embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the", and "the" used in one or more embodiments of the present application are also intended to include the plural forms unless the context clearly indicates otherwise.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the template description herein are for the purpose of describing specific embodiments only and are not intended to limit this application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0052] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when..." or "when...".
[0053] Currently, existing solutions for generating spoken-word videos of digital humans have the following problems:
[0054] 1. Currently, most video generation tools are based on templates, which lack flexibility and are difficult to meet highly personalized needs.
[0055] 2. It does not support detailed time adjustment of the generated audio to match the video content, and the audio and video synchronization is relatively poor.
[0056] 3. High professional skills are required to effectively utilize all its functions.
[0057] 4. The workflow is complex and the learning curve is steep for non-professionals.
[0058] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes in certain embodiments will not be repeated. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0059] Please refer to Figure 1 The embodiment of the present application provides a method for generating a spoken video of a digital human, which comprises at least the following steps:
[0060] S10. Obtain the spoken copy and video material data.
[0061] Specifically, in actual application scenarios, the methods for obtaining spoken copy and video material data are diverse. Spoken copy comes from a wide range of sources. It can be custom text manually entered by users in the interactive interface, or it can be automatically synchronized from third-party platforms such as enterprise content management systems, e-commerce platform product details pages, and news information databases. The ways to obtain video material data are equally rich. In addition to importing from local storage devices (such as hard drives and USB flash drives), it can also be downloaded from video material libraries and cloud storage platforms (such as Alibaba Cloud OSS and Tencent Cloud COS) through network interfaces. At the same time, in order to ensure data quality, the system will perform syntax verification and semantic integrity checks on the acquired spoken copy, and eliminate content with obvious errors or irregular formats; it will also test parameters such as resolution, frame rate, and encoding format of video materials to ensure that they meet the requirements of subsequent processing.
[0062] S20: Perform content analysis on the oral text and the video material using a multimodal large model to determine the timestamp information corresponding to each sentence in the text in the video material.
[0063] Specifically, in this embodiment, the multimodal large model integrates technologies such as natural language processing (NLP) and computer vision (CV), and has powerful cross-modal semantic understanding capabilities. When processing oral copy, the large model extracts keywords, thematic information, and semantic logic from the copy through NLP technologies such as word segmentation, part-of-speech tagging, and named entity recognition. For video materials, CV technologies such as target detection, scene recognition, and action recognition are used to analyze visual features such as the video content, scene transitions, and character actions in the video. Then, through a cross-modal alignment algorithm, the semantic information of the copy is matched with the visual information of the video, and the correlation between the copy content and the video screen is analyzed. For example, when the copy describes the "appearance design of the product", the large model will search for a clip showing the product appearance in the video material and determine the start and end timestamps of the clip. In order to improve the accuracy of the timestamp information, the system will also introduce a time series analysis algorithm to analyze the time dimension features of the video, such as the audio waveform and the frequency of picture changes, to further optimize the positioning of the timestamp.
[0064] S30: convert the oral text into corresponding audio data, and pre-process the audio data to obtain pre-processed voice data.
[0065] Specifically, the spoken text is converted into audio data using advanced text-to-speech (TTS) technology, and mature TTS engines such as Baidu PaddlePaddle and iFlytek can be selected. During the conversion process, diversified voice output is achieved by setting different parameters such as speaking speed, intonation, timbre, and emotional style (cheerful, calm, friendly, etc.). The preprocessing stage includes noise reduction, which uses spectrum analysis and filtering algorithms to remove interference such as environmental noise and current noise; audio enhancement, which improves the clarity and intelligibility of the voice through dynamic range compression and equalizer adjustment; and voice segmentation, which divides long audio into sentences or semantic paragraphs to facilitate subsequent accurate matching with video materials.
[0066] In one embodiment of the present application, preprocessing the audio data to obtain preprocessed voice data includes:
[0067] The audio data is accelerated or decelerated so that the speech rhythm of the audio data matches the content of the video material.
[0068] Specifically, in this embodiment, a phase vocoder or an overlap-add algorithm based on waveform similarity may be used to accelerate or decelerate the audio data, as follows:
[0069] Use a phase vocoder or WSOLA (Waveform Similarity-based Over Lap-Add) method to change the time scale of an audio signal without affecting its pitch.
[0070] The goal is to make the audio clips match the duration of the video frames.
[0071] Time stretch factor Where T′ is the target duration and T is the original duration.
[0072] For each sampling point n, the new sampling point position is:
[0073] n′=round(α*n)
[0074] Interpolation method:
[0075] In practice, to avoid distortion, more complex interpolation methods are usually used, such as linear interpolation or cubic spline interpolation. For interpolation between two points, the following formula can be used:
[0076]
[0077] Linear interpolation:
[0078] Among them, x1, y1 and x2, y2 are the coordinates of two known data points, and x is the position of the point to be interpolated.
[0079] S40: Merge the preprocessed audio data and the video data according to the timestamp information to generate first video data.
[0080] Specifically, based on the timestamp information determined in step S20, audio and video synchronization technology is used to precisely align the preprocessed audio data with the video data. Algorithms such as time code synchronization and audio-video cross-correlation are used to ensure that every word and sentence in the audio accurately corresponds to the corresponding content in the video. During the merging process, parameters such as the volume, brightness, and contrast of the audio and video are uniformly adjusted to ensure that the generated first video data has a more coordinated and consistent audiovisual effect. Furthermore, to improve merging efficiency, the system uses multi-threading or distributed computing technology to parallelize the merging of multiple audio and video clips.
[0081] S50. Generate a digital human according to user needs.
[0082] Specifically, in this embodiment, a digital human can be generated based on the user's actual needs. For example, user input can be achieved through a visual interactive interface, such as using a slider to adjust the digital human's height and weight, a drop-down menu to select hairstyles and clothing styles, and an input box to customize facial feature parameters. The system has built a rich library of digital human models, covering basic digital human models of different genders, ages, races, and occupations. Based on the required parameters input by the user, the basic model is personalized and rendered using technologies such as 3D modeling, texture mapping, and skeletal animation. For example, by adjusting the facial muscle model and skeletal binding parameters, a digital human's unique expression and movement style can be achieved; using a deep learning generative adversarial network (GAN), a digital human image with specific appearance features is generated based on reference images or descriptions provided by the user.
[0083] S60: Perform matting processing on the digital human using image processing technology to obtain a digital human after matting processing.
[0084] Specifically, a variety of image processing algorithms can be used in combination during the matting process. First, color keying technology is used to separate the background color from the main body of the digital human by identifying the specific color of the digital human's background (such as the green of a green screen). Then, edge detection algorithms (such as Canny edge detection) are combined to accurately extract the contour edges of the digital human to avoid edge blur or jagged edges. For difficult-to-process areas such as complex hair and translucent objects, a deep learning-based matting network (such as Deep Image Matting) is used to learn the foreground and background information of the image through a large amount of training data to achieve high-quality matting effects. Finally, post-processing operations such as edge smoothing and color correction are performed on the matted digital human to make it more natural and realistic.
[0085] In one embodiment of the present application, after step S60 and before step S70, the following steps are further included:
[0086] S61 , optimizing the contour edge of the digital human after the matting process by using an edge thinning algorithm.
[0087] Specifically, after matting the digital human, it may be necessary to dilate its edges to enhance the visual effect. The dilation formula is as follows:
[0088] The dilation operation can be implemented through convolution operation. Assume that we have a structure element (kernel), such as a cross-shaped structure element:
[0089]
[0090] For each pixel, the expanded pixel value can be expressed as:
[0091]
[0092] That is, the maximum value of all pixels within the coverage of the structural element is selected as the result of expansion.
[0093] Among them, P ′ is the pixel value after expansion, P(i,j) represents the value at position (i,j), and K is the pixel coordinate set.
[0094] S70: Merge the digital human after the matting process with the first video to generate a spoken video of the digital human.
[0095] Specifically, during the merging phase, video synthesis technology is used to integrate the matted digital human into the primary video. Based on the scene and composition of the primary video, the digital human's position, size, and angle are adjusted to blend naturally with the video background. Depth estimation and occlusion processing algorithms are used to simulate the digital human's spatial relationship within the video scene, ensuring that the digital human does not appear to be "floating" or interspersed with background objects. Furthermore, to enhance the digital human's realism and expressiveness, lighting effects are added. Based on the light direction and intensity of the primary video, corresponding shadows and highlights are added to the digital human, ensuring that the lighting conditions of the digital human and the video environment are consistent. Ultimately, a high-quality, visually impactful digital human broadcast video is generated.
[0096] In one embodiment of the present application, digital human matting and digital human matting and video synthesis adopt
[0097] The Alpha Compositing formula has the specific goal of seamlessly blending the matted digital human image with the background image.
[0098] Alpha blending formula:
[0099] Assume there are two layers of images: background image C in and foreground image (i.e. digital human image) C layer , they all have corresponding transparency (alpha channel) A in and A layer .
[0100] Output color C out The calculation formula is:
[0101]
[0102] Output transparency A out The calculation formula is:
[0103] A out =A layer +A in *(1-A layer )
[0104] Among them, C out is the output color; C in is the background image; C layer is the foreground image; A in is the transparency of the background image; A layer is the transparency of the foreground image; A out Output transparency.
[0105] The present application provides a method for generating a spoken video based on a digital human, which has the following beneficial effects:
[0106] First, it breaks the limitations of traditional template-based generation and achieves highly customized production. Digital human images are generated through user-definable input parameters and personalized adjustments are made using advanced 3D modeling and deep learning technologies. This meets users' multi-dimensional customization needs for digital human appearance, movement, style, and other aspects in different scenarios, effectively solving the problem of digital human image homogeneity. Furthermore, in the processing of spoken content and video materials, in-depth content analysis using a multimodal large model enables precise matching of text and video footage. Audio and video elements can be flexibly adjusted based on user needs to achieve customized content presentation, making digital human broadcast videos more suitable for diverse application scenarios such as e-commerce, education, and entertainment.
[0107] Secondly, the accuracy and quality of video generation have been significantly improved. By integrating natural language processing and computer vision technologies with a multimodal large model, the timestamp information of the text in the video material is accurately determined. Advanced audio and video synchronization algorithms are used to ensure precise alignment of audio and video images. In the audio pre-processing and digital human image processing, noise reduction, enhancement, and deep learning-based image processing are used to effectively improve audio clarity and digital human image quality, ultimately generating high-quality spoken-word videos with harmonious audio-visual effects and realistic and natural images.
[0108] Furthermore, video generation efficiency is improved. The solution utilizes multi-threading and distributed computing technologies to accelerate audio and video merging. Combined with automated data acquisition and processing, this reduces manual intervention and significantly shortens the production cycle for digital human broadcast videos. This meets user needs for rapid content iteration, reduces production costs for digital human videos, and enhances the solution's competitiveness and practicality in market applications.
[0109] See also Figure 2 The embodiment of the present application further provides a digital human spoken video generation system 200, comprising:
[0110] Acquisition module 201, used to acquire spoken copy and video material data;
[0111] A content analysis module 202 is configured to perform content analysis on the oral text and the video material using a multimodal large model to determine the timestamp information corresponding to each sentence in the text in the video material;
[0112] The pre-processing module 203 is used to convert the oral text into corresponding audio data and pre-process the audio data to obtain pre-processed voice data;
[0113] A synthesis module 204 is configured to combine the pre-processed audio data with the video data according to the timestamp information to generate first video data;
[0114] The digital human generation module 205 is used to generate a digital human according to the user's needs;
[0115] The matting processing module 206 is used to perform matting processing on the digital human using image processing technology to obtain a digital human after matting processing;
[0116] The spoken video generation module 207 is used to merge the digital human after the matting process with the first video to generate a spoken video of the digital human.
[0117] In one embodiment of the present application, the system 200 further includes: a contour optimization module 2061, configured to optimize the contour edge of the digital human after the matting process by using an edge thinning algorithm.
[0118] See also Figure 3 The embodiment of the present application also provides a computer electronic device 300, including a memory 303 and a processor 302, wherein the memory 303 stores a computer program, and when the processor executes the computer program, the steps of the method for generating a digital human's spoken video as described above are implemented.
[0119] Specifically, the electronic device 300 includes: a transceiver 301, a bus interface and a processor 302, the processor 302 is used to obtain spoken text and video material data; use a multimodal large model to perform content analysis on the spoken text and the video material, and determine the timestamp information corresponding to each sentence in the text in the video material; convert the spoken text into corresponding audio data, and pre-process the audio data to obtain pre-processed voice data; according to the timestamp information, merge the pre-processed audio data with the video data to generate first video data; generate a digital human according to user needs; use image processing technology to perform image extraction processing on the digital human to obtain a digital human after image extraction processing; merge the digital human after image extraction processing with the first video to generate a spoken video of the digital human.
[0120] In the embodiment of the present application, the electronic device 300 further includes: a memory 303. Figure 3 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits such as one or more processors represented by processor 302 and memory represented by memory 303. The bus architecture may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 301 may be a plurality of components, i.e., a transmitter and a receiver, providing a unit for communicating with various other devices over a transmission medium. The processor 302 is responsible for managing the bus architecture and general processing, and the memory 303 may store data used by the processor 302 when performing operations.
[0121] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any one of the above-mentioned methods for generating a spoken-word video of a digital human are implemented.
[0122] In this embodiment, the computer-readable storage medium may be a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium may include, but is not limited to, a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0123] In all examples shown and described herein, any specific values should be interpreted as merely exemplary and not limiting, and thus other examples of the exemplary embodiments may have different values.
[0124] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0125] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or flowchart, and the combination of boxes in the structure diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0126] In addition, the functional modules or units in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0127] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a terminal device (which can be a smart phone, personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
[0128] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A method for generating a digital human oral video, characterized in that: include: Obtain oral copy and video material data; Performing content analysis on the spoken copy and the video material using a multimodal large model to determine the timestamp information corresponding to each sentence in the copy in the video material; Converting the spoken text into corresponding audio data, and preprocessing the audio data to obtain preprocessed voice data; Merging the pre-processed audio data with the video data according to the timestamp information to generate first video data; Generate digital humans based on user needs; Performing matting processing on the digital human using image processing technology to obtain a digital human after matting processing; The digital human after the matting process is combined with the first video to generate a spoken video of the digital human.
2. The method for generating oral video according to claim 1, wherein: The preprocessing of the audio data to obtain preprocessed speech data includes: The audio data is accelerated or decelerated so that the speech rhythm of the audio data matches the content of the video material.
3. The method for generating oral video according to claim 2, wherein: The accelerating or decelerating processing of the audio data includes: The audio data is accelerated or decelerated using a phase vocoder or an overlap-add algorithm based on waveform similarity.
4. The method for generating oral video according to claim 1, wherein: After performing matting processing on the digital human using image processing technology to obtain the matted digital human, and before merging the matted digital human with the first video to generate a spoken video of the digital human, the method further includes: The edge of the digital human's outline after the matting process is optimized using an edge refinement algorithm.
5. The method for generating spoken video according to claim 4, wherein: The edge thinning algorithm specifically adopts the following calculation formula: Among them, P ′ is the pixel value after expansion, P(i,j) represents the value at position (i,j), and K is the pixel coordinate set.
6. The method for generating spoken video according to claim 1, wherein: The digital human image after the matting process is merged with the first video to generate the digital human's spoken video, using the following calculation formula: A out =A layer +A in *(1-A layer ) Among them, C out is the output color; C in is the background image; C layer is the foreground image; A in is the transparency of the background image; A layer is the transparency of the foreground image; A out Output transparency.
7. A digital human oral video generation system, characterized in that: include: The acquisition module is used to obtain the oral copy and video material data; A content analysis module, configured to perform content analysis on the oral text and the video material using a multimodal large model, and determine the timestamp information corresponding to each sentence in the text in the video material; A preprocessing module, configured to convert the spoken text into corresponding audio data and preprocess the audio data to obtain preprocessed voice data; a synthesis module, configured to combine the pre-processed audio data with the video data according to the timestamp information to generate first video data; A digital human generation module is used to generate a digital human according to user needs; A matting processing module, configured to perform matting processing on the digital human using image processing technology to obtain a digital human after matting processing; The spoken video generation module is used to merge the digital human after the matting process with the first video to generate a spoken video of the digital human.
8. The spoken video generation system according to claim 7, characterized in that: The system further comprises: a contour optimization module for optimizing the contour edge of the digital human after the matting process by using an edge refinement algorithm.
9. A computer electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method for generating a spoken video of a digital human according to any one of claims 1 to 6 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for generating a spoken-word video of a digital human according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Virtual anchor live broadcast system and method, and related device
CN116668733A
Video processing method and electronic equipment
CN116668768A
Video synthesis method and device, equipment and medium
CN117014653A
Virtual human video clip synthesis method and system for smart video generation
CN118660117A
Video generation method and device, electronic equipment, storage medium and product
CN118803173A
Cited By
Long video generation method and device based on digital human, and storage medium
CN121397323A
Digital person-based long video generation method and device, and storage medium
CN121397323B
Multi-agent-based news generation and digital people broadcasting method and system
CN121815041A