Digital human video generation method and device, equipment, storage medium and program product

By automatically generating and configuring digital human video clips and background video clips, the problem of inconsistency between auxiliary video materials and main video materials in the existing technology is solved, and the efficiency of digital human video generation and content consistency are improved.

CN120602742APending Publication Date: 2025-09-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510773702.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

In the existing technology, during the digital human video generation process, the content progress of the auxiliary video material is inconsistent with that of the main video material, resulting in low generation efficiency and requiring manual editing and configuration to achieve a matching effect.

Method used

By acquiring spoken text and original video material, digital human video clips and background video clips are automatically generated to match their contents, and they are respectively configured into editing tracks, and the generated preview video is displayed, avoiding manual synchronization steps.

Benefits of technology

The content synchronization of digital human video clips and background video clips is achieved, which improves the generation efficiency and ensures the consistency of content progress between the main video material and the auxiliary video material.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602742A_ABST
    Figure CN120602742A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a digital human video generation method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring an oral broadcast text and an original video material; obtaining a digital human video clip and a background video clip based on the oral broadcast text and the original video material; wherein the digital human broadcast content of the digital human video clip is matched with the picture content in the background video clip; and respectively configuring the digital human video clip and the background video clip into the corresponding editing tracks, and displaying a preview video generated by the digital human video clip and the background video clip. The digital human video clip and the background video clip which are matched in content are generated based on the oral broadcast text and the original video material, so that the synchronization of the digital human video clip and the background video clip on the broadcast content is realized, the step of manual content synchronization is avoided, the generation efficiency of the digital human video is improved, and the user experience is improved. And the content progress consistency of the main video material and the auxiliary video material is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present disclosure relate to the field of image processing technology, and in particular to a method, apparatus, device, storage medium, and program product for generating a digital human video. Background Art

[0002] Currently, with the development and integration of technologies in multiple fields such as artificial intelligence (AI) and image processing, the relevant applications of virtual digital human technology have been realized. Among them, virtual digital humans, also known as digital humans, are visual objects with human appearance. By creating digital human videos to imitate human speech tone and behavioral movements to express information, such as news broadcasts and product introductions, it can achieve the purpose of replacing real-life recorded videos for information presentation, thereby effectively reducing the cost of producing such videos.

[0003] In the existing technology, relevant functions for creating digital human videos are provided in relevant video generation applications and platforms. Users can directly generate the main video material (A-roll) of the digital human video by inputting text, that is, the picture material of the digital human broadcasting the text input by the user.

[0004] However, for the auxiliary video materials (B-roll) of the video, such as the background display screen of the digital human in the video, users still need to manually edit and configure them in the video editing interface to achieve the effect of matching the digital human voice content with the background display screen, resulting in low efficiency in digital human video generation and inconsistent content progress between the main video materials and the auxiliary video materials. Summary of the Invention

[0005] The embodiments of the present disclosure provide a method, apparatus, device, storage medium, and program product for generating digital human videos, so as to overcome the problems of low efficiency in generating digital human videos and inconsistency between primary video materials and auxiliary video materials.

[0006] In a first aspect, an embodiment of the present disclosure provides a method for generating a digital human video, comprising:

[0007] Acquire spoken text and original video material; obtain digital human video clips and background video clips based on the spoken text and the original video material; wherein the digital human spoken content in the digital human video clips matches the picture content in the background video clips; respectively configure the digital human video clips and the background video clips into corresponding editing tracks, and display a preview video generated by the digital human video clips and the background video clips.

[0008] In a second aspect, an embodiment of the present disclosure provides a digital human video generation device, comprising:

[0009] Loading module, used to obtain spoken text and original video materials;

[0010] A generating module is configured to obtain a digital human video segment and a background video segment based on the spoken text and the original video material; wherein the spoken content of the digital human video segment matches the screen content of the background video segment;

[0011] The configuration module is used to configure the digital human video clip and the background video clip into corresponding editing tracks respectively, and display a preview video generated by the digital human video clip and the background video clip.

[0012] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;

[0013] The memory stores computer-executable instructions;

[0014] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the digital human video generation method as described in the first aspect and various possible designs of the first aspect.

[0015] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the digital human video generation method described in the first aspect and various possible designs of the first aspect is implemented.

[0016] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the digital human video generation method as described in the first aspect and various possible designs of the first aspect.

[0017] The digital human video generation method, apparatus, device, storage medium, and program product provided in this embodiment obtain a spoken text and original video material; based on the spoken text and the original video material, a digital human video clip and a background video clip are obtained; wherein the digital human spoken content of the digital human video clip matches the image content of the background video clip; the digital human video clip and the background video clip are respectively configured into corresponding editing tracks, and a preview video generated by the digital human video clip and the background video clip is displayed. By generating a digital human video clip and a background video clip with matching content based on the spoken text and the original video material, synchronization of the digital human video clip and the background video clip in terms of playback content is achieved, that is, the digital human spoken content matches the image content of the background video clip, avoiding the step of manual content synchronization, improving the efficiency of digital human video generation, and improving the consistency of content progress between the main video material and the auxiliary video material. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 A diagram of an application scenario of the digital human video generation method provided by an embodiment of the present disclosure;

[0020] Figure 2 Schematic diagram of the process of the digital human video generation method provided in the embodiment of the present disclosure Figure 1 ;

[0021] Figure 3 A schematic diagram of a process for generating a preview video provided by an embodiment of the present disclosure;

[0022] Figure 4 Schematic diagram of the process of the digital human video generation method provided in the embodiment of the present disclosure Figure 2 ;

[0023] Figure 5 for Figure 4 A flowchart of a specific implementation method of step S203 in the embodiment shown;

[0024] Figure 6 is a flowchart of a specific implementation method of step S204A;

[0025] Figure 7 for Figure 4 A flowchart of a specific implementation method of step S205 in the embodiment shown;

[0026] Figure 8 A schematic diagram of a process for generating background video clips provided by an embodiment of the present disclosure;

[0027] Figure 9 A structural block diagram of a digital human video generation device provided by an embodiment of the present disclosure;

[0028] Figure 10 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure;

[0029] Figure 11 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] The following explains the application scenarios of the embodiments of the present disclosure:

[0033] The digital human video generation method provided by the embodiments of the present disclosure can be applied to applications (APPs) with digital human video generation functions, such as video editing applications, short video applications, and live broadcast applications. More specifically, it can be applied to application scenarios such as generating marketing videos and news videos based on digital humans. The execution subject of this embodiment can be a terminal device running the aforementioned application with digital human video generation functions, or a server deploying the server corresponding to the aforementioned application, or other electronic devices that perform similar functions. When the execution subject is a terminal device, the terminal device executes the method provided by this embodiment by running the aforementioned application. When the execution subject is a server, the server of the aforementioned application with digital human video generation functions can be partially or entirely run on the server, and the method provided by this embodiment is executed on the server side, while the terminal device runs the client of the application, or runs a browser. The communication between the server and the terminal device is based on a client-server (CS) architecture, or the communication between the server and the terminal device is based on a browser-server (BS) architecture, so that the terminal device can obtain the execution results of the method provided by this embodiment and display them as needed.

[0034] Among them, in some embodiments, the terminal device or server can implement the digital human video generation method provided by the embodiment of the present disclosure by running various computer executable instructions or computer programs. For example, computer executable instructions can be program-level commands, machine instructions or software instructions. The computer program can be a native program or software module in the operating system; it can be a local application, that is, a program that needs to be installed in the operating system to run, or it can be a small program embedded in any APP, that is, a program that runs based on a browser environment. In summary, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form, and the specific implementation form can be configured as needed. Furthermore, in the process of implementing the digital human video generation method provided by the embodiment of the present disclosure, the terminal device can execute the method by running a computer executable instruction or computer program set locally, or it can execute the method by calling a computer executable instruction or computer program set in an external server. In some embodiments, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud storage, cloud communications, cloud databases, cloud computing, cloud functions, network services, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Among them, cloud services can be interactive processing services for terminal devices to call.

[0035] Figure 1 This is an application scenario diagram of the digital human video generation method provided by the embodiment of the present disclosure, referring to Figure 1As shown in the figure, a terminal device, for example, runs a target application with a digital human video generation function, such as a video editing program. Through the target application's interactive interface, the user configures spoken text, such as a news article or product introduction, within the editing track. Subsequently, by triggering the target application's video generation function, for example by clicking the "Digital Human" control shown in the figure, the terminal device invokes a cloud-based model and configures relevant parameters. Based on the spoken text, the terminal device generates a video of a digital human reciting the spoken text, i.e., a digital human video. For example, as shown in the figure, the spoken text includes: "Today, I'd like to introduce you to a newly released foldable screen mobile phone..." The generated digital human video plays the audio corresponding to the spoken content. Simultaneously, the humanoid virtual digital human's lip movements and body movements create the visual experience of a "real person speaking." For functional video scenes such as news manuscripts and product introductions, digital humans are usually used as the main video material, while news scene videos and product introduction videos are presented as auxiliary video materials in the video background (i.e., auxiliary video materials are configured) to further improve the richness of video content and information display effect.

[0036] However, in the existing technology, for the auxiliary video materials of the video, the user still needs to manually edit and configure them in the video editing interface after generating the digital human video according to the progress of the digital human's oral content in the digital human video, so as to achieve the effect of matching the digital human's oral content with the background display screen, which leads to the problems of low efficiency in digital human video generation and inconsistent progress of the content of the main video materials and the auxiliary video materials.

[0037] The embodiments of the present disclosure provide a method for generating a digital human video to solve the above problems.

[0038] refer to Figure 2 , Figure 2 Schematic diagram of the process of the digital human video generation method provided in the embodiment of the present disclosure Figure 1 The method of this embodiment can be applied in a terminal device or a server. In the case where the terminal device executes the method provided by this embodiment, in one possible implementation, the terminal device can implement the digital human video generation method provided by this embodiment by executing a locally deployed program code. In another possible implementation, the server can be used to deploy a functional service based on the digital human video generation method provided by this embodiment, and the terminal device can implement the digital human video generation method provided by this embodiment by accessing the above server and calling the corresponding functional service. Exemplarily, the digital human video generation method provided by this embodiment includes:

[0039] Step S101: Obtain spoken text and original video material.

[0040] Step S102: Based on the spoken text and the original video material, a digital human video clip and a background video clip are obtained; wherein the spoken content of the digital human in the digital human video clip matches the screen content in the background video clip.

[0041] Step S103: placing the digital human video clip and the background video clip into corresponding editing tracks respectively, and displaying a preview video generated by the digital human video clip and the background video clip.

[0042] refer to Figure 1 The application scenario diagram shown in the figure describes the provided digital human video generation method using a terminal device as the execution subject. For example, the terminal device runs a target application and provides the user with a digital human video generation function through the target application. The user triggers the digital human video generation function through the target application's interactive interface, which causes the digital human video generation function interface to appear. This interface then loads spoken text and original video material, and the subsequent process of generating digital human video clips and background video clips is performed based on the spoken text and original video material. The spoken text is the content text spoken by the digital human, such as a press release or product introduction. In subsequent steps, the corresponding audio and the digital human's lip movements will be generated based on this spoken text to achieve the effect of the digital human speaking the spoken text. The original video material is the material used to generate the digital human's background image, such as videos and pictures. Among them, for the video in the original video material, this embodiment does not limit the video length, but in a possible implementation method, the terminal device can analyze the matching degree of the video length and the number of words in the spoken text by detecting the two. For example, when the text has a large number of words and the video length is short, a prompt message will be displayed on the terminal device to prompt that the video length does not match the text word count, and to prompt the user to replace or modify the spoken text or the original video material, thereby improving the matching degree of the main video material and the auxiliary video material in the final generated digital human video.

[0043] After the terminal obtains the spoken text and original video material, in one possible implementation, the spoken text and original video material can be processed by a model deployed locally on the terminal device to generate a digital human video clip and a background video clip with matching content progress. In another possible implementation, the terminal device can call a cloud-based model to process the spoken text and original video material to generate a digital human video clip and a background video clip with matching content progress. The digital human video clip is a video material containing only the digital human image, and the background video clip is a video material serving as the video background. The digital human spoken content in the digital human video clip matches the visual content in the background video clip, meaning that the audio content progress of the digital human video clip matches the visual content progress of the background video clip. For example, when the digital human in the digital human video clip is "narrating" the appearance of a product, the visual content in the background video clip is the product's appearance; when the digital human in the digital human video clip is "narrating" the company's development, the visual content in the background video clip is the company's office space, group photos, and other images.

[0044] Afterwards, after the terminal device obtains the generated digital human video clip and background video clip, it will place them on track, for example, in the corresponding editing track of the target application, and display the preview video generated by the digital human video clip and background video clip in the video preview window. Different editing tracks correspond to different layers. In one possible implementation, the track layer corresponding to the digital human video clip is layered over the track layer corresponding to the background video clip, so that the digital human video clip is displayed over the background video clip, highlighting the priority display of the digital human in the digital human video and improving the authenticity of the digital human video.

[0045] Figure 3 A schematic diagram of a process for generating a preview video provided by an embodiment of the present disclosure is shown below. Figure 3 A more detailed description of the above process is given below: Figure 3As shown, for example, first, within the video editing interface of a video editing application (target application), multiple editing tracks are configured. The editing tracks within the video editing interface are used to configure materials, and corresponding preview videos are generated using the materials on each editing track. Within the video editing interface, a functional control for generating a digital human video is provided, such as "Control A" shown in the figure. After the user triggers "Control A," the user is redirected to the corresponding functional interface. Subsequently, the user can use the functional interface to input / load the spoken text and original video materials in the spoken text input area and the video material input area. After obtaining the spoken text and original video materials input by the user, the terminal device generates a matching digital human video clip video_1 and background video clip video_2 by calling the video generation model, and automatically places them into the editing tracks within the video editing interface. For example, the digital human video clip video_1 is configured to editing track #1, and the background video clip video_1 is configured to editing track #2. At the same time, a preview video consisting of the digital human video clip and the background video clip is displayed in the preview area of ​​the video editing interface.

[0046] Afterwards, in response to the user's modification and export operations, the material on the editing track can be further adjusted based on the preview video, or the digital human video corresponding to the preview video can be exported.

[0047] In this embodiment, a spoken text and original video material are obtained; based on the spoken text and original video material, a digital human video clip and a background video clip are generated; wherein the digital human spoken content of the digital human video clip matches the visual content of the background video clip; the digital human video clip and the background video clip are respectively configured into corresponding editing tracks, and a preview video generated by the digital human video clip and the background video clip is displayed. By generating a digital human video clip and a background video clip with matching content based on the spoken text and the original video material, synchronization of the digital human video clip and the background video clip in terms of playback content is achieved, that is, the digital human spoken content matches the visual content of the background video clip, avoiding the step of manual content synchronization, improving the efficiency of digital human video generation, and enhancing the consistency of content progress between the primary video material and the auxiliary video material.

[0048] refer to Figure 4 , Figure 4 Schematic diagram of the process of the digital human video generation method provided in the embodiment of the present disclosure Figure 2 In this embodiment Figure 2 Based on the embodiment shown, step S102 is further refined, and the digital human video generation method includes:

[0049] Step S201: Obtain spoken text and original video material.

[0050] Step S202: Generate a digital human video clip based on the spoken text.

[0051] Step S203: Based on the text content of the spoken text, a point sequence is obtained, where the point sequence includes at least two ordered play points, and each play point corresponds to a start timestamp and / or end timestamp of a text segment.

[0052] For example, after acquiring the spoken text and original video footage, the terminal device can first generate a corresponding digital human video clip based on the spoken text. Specifically, for example, using the digital human generation function provided by the target application, the spoken text is first added to the corresponding text editing track. In one possible implementation, the terminal device can also obtain the playback timestamp corresponding to the spoken text, such as the timestamp corresponding to each character or sentence in the spoken text. When adding the spoken text to the text editing track, the terminal device configures the playback timestamp based on this playback timestamp to control the spoken text's rhythm and duration. Subsequently, parameters such as the digital human's image, timbre, and video clip duration are set. Finally, a video generation model is invoked to process the text editing track based on these parameters, thereby generating a video clip containing the digital human reciting the spoken text on the text editing track, i.e., a digital human video clip. The method for generating the digital human video clip is conventional and will not be described in detail here.

[0053] After that, the terminal device generates a point sequence based on the text content of the spoken text, wherein the point sequence includes at least two ordered playback points, each playback point corresponds to the start timestamp and / or end timestamp of a text segment. Specifically, the point sequence is a collection of a series of time points. The playback point can be the timestamp of the start playback of the video segment (start timestamp), or the timestamp of the end playback (end timestamp), or a timestamp pair in the form of [p1, p2], where p1 is the start timestamp of the video segment, and p2 is the end timestamp of the video segment. Through the point sequence and the total playback time, the playback time interval corresponding to each video segment can be determined.

[0054] Among them, the above-mentioned point sequence is determined based on the text content of the spoken text. For example, according to the start and end timestamps of each natural paragraph of the above-mentioned spoken text in the text editing track, the corresponding playback points are determined, and then the point sequence is obtained; or, according to the semantic content of the above-mentioned spoken text, one or more semantic paragraphs are obtained, and then the corresponding playback points are determined based on the start and end timestamps of the semantic paragraphs in the text editing track, and then the point sequence is obtained. The spoken text paragraphs corresponding to the point sequence obtained based on the above method have content aggregation, that is, one playback point corresponds to one spoken text paragraph, and the content in the same spoken text paragraph is similar, but the content corresponding to different spoken text paragraphs is different. Correspondingly, the video clips generated based on the point sequence also have content aggregation. The specific method of dividing the text content of the spoken text based on the above-mentioned method to generate one or more spoken text paragraphs can be achieved through a large language model with semantic understanding capabilities, which will not be introduced in detail here. Afterwards, the corresponding playback points can be determined based on the start and end timestamps corresponding to the spoken text paragraphs, and then the point sequence can be determined.

[0055] In a possible implementation, the digital human video clip has a first playback duration, and the specific implementation of step S203 includes:

[0056] Based on the first playback time and a preset time ratio coefficient, a second playback time is obtained, and the second playback time is greater than the first playback time; based on the second playback time and the text content of the spoken text, a point sequence is obtained, wherein the total playback time of the video clips corresponding to each playback point in the point sequence is greater than or equal to the first playback time and less than or equal to the second playback time.

[0057] For example, in actual applications, the duration of the digital human video clips and background video clips is not necessarily completely consistent. Typically, the background video clips are longer, while the digital human video clips are shorter. For example, at the beginning and end of the video, only the background video clips are displayed, without the digital human video clips (i.e., the final generated digital human video does not contain the digital human voice at the beginning and end), thereby achieving a better video display effect. However, at the same time, there are still certain duration restrictions between the video clips of the digital human and the background video clips. When the duration of the digital human video clips is too short and the duration of the background video clips is too long, it will also cause a large amount of the background video clips to have "no introduction" problem, affecting the efficiency and effect of information display. To address this problem, in this embodiment, since the digital human video clip is generated first, its playback time is also determined first. Therefore, after the digital human video clip is generated, the first playback time corresponding to the digital human video clip is first obtained. Then, based on the first playback time and the preset time ratio coefficient, a second playback time is obtained. The second playback time is the maximum permitted time of the background video clip. Then, based on the second playback time and the text content of the spoken text, a point sequence is jointly generated, so that the total playback time of the video clip corresponding to each playback point in the point sequence is greater than or equal to the first playback time and less than or equal to the second playback time, so that the video time of the background video clip generated in the subsequent steps can achieve the purpose of being slightly greater than the video time of the digital human video clip, thereby achieving better information display efficiency and display effect.

[0058] In another possible implementation, Figure 5 As shown, the specific implementation of step S203 includes:

[0059] Step S2031: Obtain the text block timestamp corresponding to each text block in the spoken text;

[0060] Step S2032: clustering the text blocks according to their text semantics to obtain at least one text segment;

[0061] Step S2033: The text segment obtains a corresponding playback point according to the text block timestamp of the text block in the text segment;

[0062] Step S2034: Obtain a point sequence based on the set of playback points corresponding to each text segment.

[0063] Exemplarily, a text block can be a set of characters of fixed or non-fixed length, for example, one Chinese character corresponds to one text block, and for another example, a paragraph separated by punctuation marks can also correspond to one text block. After dividing the above-mentioned oral text of the text editing track into text blocks, the text block timestamp corresponding to each text block can be obtained. Afterwards, by extracting the text semantics corresponding to each text block and clustering the text blocks based on the text semantics, one or more text segments can be obtained, and each text segment is a set of one or more text blocks. Afterwards, the text block timestamp (start timestamp or end timestamp) corresponding to the first text block in the text segment and / or the text block timestamp (start timestamp or end timestamp) corresponding to the last text block in the text segment is used to determine the playback point corresponding to the text segment. By analogy, after obtaining the playback point corresponding to each text segment, based on the playback point corresponding to each of the above-mentioned text segments, a point sequence corresponding to the text segment based on text semantic aggregation can be obtained.

[0064] Furthermore, in another possible implementation, before step S2032, the following steps are further included:

[0065] Step S2030: Obtain a video generation template, where the video generation template is used to indicate a target number of video clips.

[0066] Accordingly, the specific implementation of step S2032 includes:

[0067] Step S2032A: Cluster the text blocks according to their text semantics and target number to obtain the target number of text segments.

[0068] For example, in another implementation, during the process of generating background video clips, the background video clips are generated based on a video generation template. The video generation template is a template used to limit parameters such as the layout, transitions, and special effects of the video screen. The video generation template usually reserves a position for the material to be inserted. By inserting / linking the material into the video generation template, the video can be quickly generated and the generated video can integrate the screen layout, transitions, special effects, and other parameters in the video generation template. Video generation templates are a common method in the field of video generation technology and will not be described in detail here.

[0069] In the case of using a video generation template, in this embodiment, in addition to clustering based on the text semantics of the text blocks, the video generation template is further obtained, and the target number of video clips indicated by the video generation template is obtained, that is, the number of video materials to be inserted reserved in the video generation template. For example, the video generation template M indicates a target video consisting of three video clips and two transition effects, then the target number corresponding to the video generation template M is 3. Afterwards, in the process of constructing the text segments, semantic aggregation will be performed based on the target number to generate the corresponding target number of text segments. For example, if the target number corresponding to the video generation template M is 3, all the text blocks will be aggregated into 3 text segments; and if the target number corresponding to the video generation template M is 6, all the text blocks will be aggregated into 6 text segments, thereby achieving the matching of the subsequently generated video segments based on semantic aggregation with the video generation template.

[0070] Step S204: Based on the picture content of the original video material, the original video material is segmented and extracted to generate at least one video material slice.

[0071] Step S205: Slices of the video material that match the text segment are combined based on the point sequence to generate a background video segment.

[0072] Exemplarily, in another aspect, the terminal device performs semantic segmentation and extraction of the original video material based on its visual content, for example, by invoking an image processing model, to generate at least one video material slice. For example, based on the visual content of the original video material, the terminal device divides the original video material into three video material slices, where the first video material slice corresponds to the visual content of "product appearance"; the second video material slice corresponds to the visual content of "core functions"; and the third video material slice corresponds to the visual content of "after-sales policy." Subsequently, based on the text segments corresponding to each playback point in the point sequence, the terminal device searches for semantically matching video material slices and combines them to generate a background video segment.

[0073] In a possible implementation, the specific implementation of step S204 includes:

[0074] Step S204A: obtaining picture features of the original video material, and segmenting the original video material based on the picture features to obtain at least one video material slice.

[0075] For example, the process of generating video material slices can be executed in an asynchronous thread after obtaining the original video material. Specifically, for example, the image processing model is called to extract the picture features of the original video material, and based on the changes between the picture features, the original video material is segmented to obtain at least one video material slice. In a possible implementation, Figure 6 As shown, the specific implementation of step S204A includes:

[0076] Step S204A-1: dividing the original video material into at least two initial material slices according to the picture features of the original video material, wherein the feature distance between the picture features of the initial material slices is greater than a first feature threshold.

[0077] Step S204A-2: Obtain the confidence level of the picture semantics corresponding to each initial material slice, and filter out the initial material slices with a confidence level less than a confidence threshold to obtain at least one video material slice.

[0078] Exemplarily, first, the image processing model is called to extract the picture features of each frame (or every N frames) in the original video material. Then, based on the changes in the picture features between frames, the video frames belonging to different picture scenes are determined, so that the video frames corresponding to different picture scenes are combined into corresponding initial material slices. For example, the original video material consists of 1024 video frames. According to the changes in the picture features corresponding to each video frame of the original video material, the 1st to 512th video frames are divided into initial material slices P1, the 513th to 768th video frames are divided into initial material slices P2, and the 769th to 1024th video frames are divided into initial material slices P3. In the above-mentioned initial material slices P1, initial material slices P2, and initial material slices P3, the picture features of each video frame in each initial material slice have similar picture semantics, that is, the picture features are similar and there is no feature jump; while the video frames between the initial material slices have different picture semantics, that is, the picture features are quite different and there is feature jump, that is, the feature distance between the picture features of the initial material slices is greater than the first feature threshold.

[0079] On this basis, after the image processing model divides the original video material into multiple video material slices based on the picture features, it will simultaneously output the confidence of the picture semantics corresponding to each video material slice. Among them, the picture semantics, for example, represent the specific categories of picture content such as "product appearance", "core functions", and "after-sales policy". When the confidence is high, it means that the video material slice has a relatively definite content category, that is, the credibility of its unique content category is high, and it is retained as a video material slice; on the other hand, for the initial material slice with low confidence, it means that the video material slice has a relatively vague content category, and it is filtered out, so that the video material slices used in the subsequent composition of the background video clip have clear and high-credibility semantic categories, thereby improving the matching success and accuracy of the text clip and the video material slice, and improving the content progress consistency of the digital human video clip and the background video clip in the final preview video.

[0080] Correspondingly, such as Figure 7 As shown, the specific implementation of step S205 includes:

[0081] Step S2051: obtaining the image semantics of each video material slice, and obtaining the video material slice corresponding to each playback point based on the similarity between the image semantics of each video material slice and the text semantics of the text segment corresponding to each playback point in the point sequence;

[0082] Step S2052: Generate background video clips based on the video material slices corresponding to each playback point.

[0083] For example, based on the picture semantics of the video material slices obtained in the previous step, the similarity between the picture semantics of each video material slice and the text semantics of the text segment corresponding to each playback point in the point sequence is compared to obtain multiple video material slices corresponding to each playback point, and for each playback point, the video material slices are sorted based on the degree of semantic matching (matching value), and the video material slice with the greatest degree of semantic matching is used as the video material slice corresponding to the playback point, and then the video material slices corresponding to each playback point are combined, and combined with the video parameters configured by the video generation template to generate a background video segment.

[0084] Figure 8 A schematic diagram of a process for generating background video segments provided by an embodiment of the present disclosure is shown below. Figure 8 The above process is further introduced as follows: Figure 8As shown, after obtaining the spoken text and original video material, first, based on the spoken text and the video generation template, multiple text segments based on semantic aggregation and corresponding point sequences are generated. For example, as shown in the figure, they include text segment T1, text segment T2, and text segment T3. Each text segment corresponds to a playback point, such as playback point t1, playback point t2, and playback point t3. Then, based on the above text segments, corresponding digital human video segments are generated. The digital human video segments will orally recite the above text segments T1, T2, and T3 at the corresponding playback points. On the other hand, based on the picture features of the original video material, the original video material is divided into multiple initial material slices, such as initial material slice p1, initial material slice p2, initial material slice p3, initial material slice p4, and initial material slice p5 as shown in the figure. Afterwards, the above initial material slices are screened to obtain video material slices p1, video material slice p3, video material slice p4, and video material slice p5. Finally, semantic matching is performed between the text segments and the video slices, determining that playback point t1 corresponds to video slice p3, playback point t2 corresponds to video slice p4, and playback point t3 corresponds to video slice p1. Next, video slices p3, p4, and p1 are combined to generate a background video segment. If any video slices overlap, the overlapping portion of the video slice belonging to the previous playback point is removed.

[0085] In this embodiment, a point sequence is generated by text fragments based on semantic aggregation, and then the original video material is sliced ​​and screened to obtain multiple video material slices. Finally, the video material slices and text fragments are semantically matched to obtain multiple video material slices arranged based on the point sequence, which are combined into background content fragments. The content display progress of the background content fragments can match the oral progress of the digital human content generated based on the oral text, thereby achieving accurate correspondence between the main video material and the auxiliary video material, and improving the information display effect of the digital human video.

[0086] Step S206: Determine the positional relationship between the digital human video segment and the background video segment based on the screen content of the background video segment.

[0087] Step S207: arranging the digital human video clip and the background video clip into corresponding editing tracks respectively.

[0088] Step S208: Determine video reference information based on the screen content of the background video clip, where the video reference information includes transition effects and / or subtitles of the preview video.

[0089] Step S209: Generate a preview video based on the digital human video clip, the background video clip, and the position relationship and / or video reference information.

[0090] Step S210: displaying a preview video generated by the digital human video clip and the background video clip.

[0091] For example, after obtaining the digital human video clip and the background video clip, the positional relationship between them can be determined by default configuration parameters. For example, based on a preset video screen size, the background video clip is displayed full screen, while the digital human video clip is displayed below the background video clip at a fixed size. In another possible implementation, the positional relationship between the digital human video clip and the background video clip can also be determined based on the screen content of the background video clip. Specifically, the positional relationship between the digital human video clip and the background video clip can be determined based on the screen content of the background video clip. For example, the position of the digital human video clip can be dynamically adjusted based on the screen content of the background video clip, so that the digital human is always located in a blank or non-important position in the background video clip, avoiding the digital human from blocking important information. This improves the information display efficiency of the preview video (and the ultimately generated digital human video).

[0092] At the same time, in a possible implementation, in this embodiment, video reference information can also be determined based on the screen content of the background video clip, wherein the video reference information includes the transition effects and / or subtitles of the preview video. Through the transition effects and / or subtitles, the information display effect of the generated preview video (digital human video) can be further improved and enriched. At the same time, the video reference information can also be determined based on the screen content of the background video clip, by setting transition effects and / or subtitles that match the screen content of the background video clip, such as configuring the subtitle color to be different from the screen content color, the duration of the transition effect to be proportional to the background video clip, etc. This allows the transition effects and / or subtitles to have a better display effect, or to have better consistency with the screen content of the background video clip.

[0093] In this embodiment, the implementation of steps S201, S207, and S210 is the same as that of the present disclosure. Figure 2 The implementation methods of step S101 and step S103 in the illustrated embodiment are the same and will not be described in detail here.

[0094] Corresponding to the digital human video generation method of the above embodiment, Figure 9This is a block diagram of the structure of a digital human video generation device provided in an embodiment of the present disclosure. The methods described in the above embodiments can be executed by this digital human video generation device, which can be implemented using software and / or hardware and integrated into electronic devices with certain data processing capabilities. These electronic devices may include, but are not limited to, mobile terminals with big data processing capabilities, as well as fixed terminals with big data processing capabilities, such as desktop computers and supercomputers.

[0095] For ease of explanation, only the parts related to the embodiments of the present disclosure are shown. Figure 9 , the digital human video generating device 3 includes:

[0096] The loading module 31 is used to obtain the spoken text and the original video material;

[0097] A generating module 32 is configured to obtain a digital human video segment and a background video segment based on the spoken text and the original video material; wherein the spoken content of the digital human video segment matches the image content of the background video segment;

[0098] The configuration module 33 is used to configure the digital human video clip and the background video clip into corresponding editing tracks respectively, and display the preview video generated by the digital human video clip and the background video clip.

[0099] According to one or more embodiments of the present disclosure, the generation module 32 is specifically used to: generate a digital human video clip based on the spoken text; obtain a point sequence based on the text content of the spoken text, the point sequence including at least two ordered playback points, each playback point corresponding to the start timestamp and / or end timestamp of a text clip; based on the picture content of the original video material, segment and extract the original video material to generate at least one video material slice, and generate a background video clip based on the point sequence combination of the video material slices that match the text clip.

[0100] According to one or more embodiments of the present disclosure, when the generation module 32 obtains a point sequence based on the text content of the spoken text, it is specifically used to: obtain the text block timestamp corresponding to each text block in the spoken text; cluster the text blocks according to the text semantics of the text blocks to obtain at least one text segment; according to the text block timestamps of the text blocks within the text segment, the text segment obtains the corresponding playback point; according to the set of playback points corresponding to each text segment, obtain a point sequence.

[0101] According to one or more embodiments of the present disclosure, the generation module 32 is further used to: obtain a video generation template, which is used to indicate a target number of video segments; when the generation module 32 clusters the text blocks according to the text semantics of the text blocks to obtain at least one text segment, it is specifically used to: cluster the text blocks according to the text semantics of the text blocks and the target number to obtain the target number of text segments.

[0102] According to one or more embodiments of the present disclosure, when the generation module 32 segments and extracts the original video material based on the picture content of the original video material to generate at least one video material slice, it is specifically used to: obtain the picture features of the original video material, and segment the original video material based on the picture features to obtain at least one video material slice; when the generation module 32 generates background video segments based on the video material slices that match the text segments based on the point sequence combination, it is specifically used to: obtain the picture semantics of each video material slice, and obtain the video material slice corresponding to each playback point based on the similarity between the picture semantics of each video material slice and the text semantics of the text segment corresponding to each playback point in the point sequence; generate a background video segment based on the video material slices corresponding to each playback point.

[0103] According to one or more embodiments of the present disclosure, when the generation module 32 divides the original video material based on the picture features to obtain at least one video material slice, it is specifically used to: divide the original video material into at least two initial material slices based on the picture features of the original video material, and the feature distance between the picture features of each initial material slice is greater than a first feature threshold; obtain the confidence of the picture semantics corresponding to each initial material slice, and filter out the initial material slices with a confidence less than the confidence threshold to obtain at least one video material slice.

[0104] According to one or more embodiments of the present disclosure, the digital human video clip has a first playback duration. When the generation module 32 obtains a point sequence based on the text content of the spoken text, it is specifically used to: obtain a second playback duration based on the first playback duration and a preset duration ratio coefficient, and the second playback duration is greater than the first playback duration; obtain a point sequence based on the second playback duration and the text content of the spoken text, wherein the total playback duration of the video clip corresponding to each playback point in the point sequence is greater than or equal to the first playback duration, and less than or equal to the second playback duration.

[0105] According to one or more embodiments of the present disclosure, the configuration module 33 is further used to: determine the positional relationship between the digital human video clip and the background video clip, and / or video reference information based on the screen content of the background video clip, where the video reference information includes transition effects and / or subtitles of the preview video; and generate a preview video based on the digital human video clip, the background video clip, and the positional relationship and / or video reference information.

[0106] The loading module 31, the generating module 32 and the configuring module 33 are connected in sequence. The digital human video generating device 3 provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0107] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure is shown in FIG. Figure 10 As shown, the electronic device 4 includes:

[0108] A processor 41, and a memory 42 communicatively connected to the processor 41;

[0109] Memory 42 stores computer-executable instructions;

[0110] The processor 41 executes the computer execution instructions stored in the memory 42 to implement the following Figure 2-Figure 8 The digital human video generation method in the illustrated embodiment.

[0111] Optionally, the processor 41 and the memory 42 are connected via a bus 43 .

[0112] For related instructions, please refer to Figure 2-Figure 8 The relevant descriptions and effects corresponding to the steps in the corresponding embodiments can be understood, and no further details are given here.

[0113] The present invention provides a computer-readable storage medium that stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the computer-executable instructions are used to implement the present invention. Figure 2-Figure 8 The digital human video generation method provided by any one of the corresponding embodiments.

[0114] The present invention provides a computer program product, including a computer program, which implements the present invention when executed by a processor. Figure 2-Figure 8 The digital human video generation method provided by any one of the corresponding embodiments.

[0115] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.

[0116] refer to Figure 11, which shows a schematic structural diagram of an electronic device 900 suitable for implementing an embodiment of the present disclosure. The electronic device 900 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 11 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0117] like Figure 11 As shown, the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0118] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 11 The electronic device 900 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0119] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0120] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0121] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0122] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0123] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0125] The units or modules involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit or module does not, in some cases, limit the unit itself.

[0126] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0127] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0128] In a first aspect, according to one or more embodiments of the present disclosure, a method for generating a digital human video is provided, comprising:

[0129] Acquire spoken text and original video material; obtain digital human video clips and background video clips based on the spoken text and the original video material; wherein the digital human spoken content in the digital human video clips matches the picture content in the background video clips; respectively configure the digital human video clips and the background video clips into corresponding editing tracks, and display a preview video generated by the digital human video clips and the background video clips.

[0130] According to one or more embodiments of the present disclosure, based on the spoken text and the original video material, a digital human video clip and a background video clip are obtained, including: generating a digital human video clip based on the spoken text; obtaining a point sequence based on the text content of the spoken text, the point sequence including at least two ordered playback points, each of the playback points corresponding to the start timestamp and / or end timestamp of a text clip; based on the picture content of the original video material, the original video material is segmented and extracted to generate at least one video material slice, and based on the point sequence, the video material slices matching the text clip are combined to generate the background video clip.

[0131] According to one or more embodiments of the present disclosure, the point sequence is obtained based on the text content of the spoken text, including: obtaining the text block timestamp corresponding to each text block in the spoken text; clustering the text blocks according to the text semantics of the text blocks to obtain at least one text segment; according to the text block timestamps of the text blocks within the text segment, the text segment obtains the corresponding playback point; according to the set of playback points corresponding to each text segment, a point sequence is obtained.

[0132] According to one or more embodiments of the present disclosure, the method further includes: obtaining a video generation template, wherein the video generation template is used to indicate a target number of video segments; clustering the text blocks according to the text semantics of the text blocks to obtain at least one text segment, including: clustering the text blocks according to the text semantics of the text blocks and the target number to obtain the target number of text segments.

[0133] According to one or more embodiments of the present disclosure, the segmentation and extraction of the original video material based on the picture content of the original video material to generate at least one video material slice includes: obtaining the picture features of the original video material, and segmenting the original video material based on the picture features to obtain at least one video material slice; the generation of the background video segment based on the video material slices that match the text segment based on the point sequence combination includes: obtaining the picture semantics of each of the video material slices, and obtaining the video material slice corresponding to each playback point based on the similarity between the picture semantics of each of the video material slices and the text semantics of the text segment corresponding to each playback point in the point sequence; generating the background video segment based on the video material slices corresponding to each playback point.

[0134] According to one or more embodiments of the present disclosure, the segmenting of the original video material based on the picture features to obtain at least one video material slice includes: segmenting the original video material into at least two initial material slices based on the picture features of the original video material, wherein the feature distance between the picture features of each of the initial material slices is greater than a first feature threshold; obtaining the confidence of the picture semantics corresponding to each initial material slice, and filtering out the initial material slices whose confidence is less than the confidence threshold to obtain at least one video material slice.

[0135] According to one or more embodiments of the present disclosure, the digital human video clip has a first playback duration, and the point sequence is obtained based on the text content of the spoken text, including: obtaining a second playback duration based on the first playback duration and a preset duration ratio coefficient, and the second playback duration is greater than the first playback duration; obtaining a point sequence based on the second playback duration and the text content of the spoken text, wherein the total playback duration of the video clip corresponding to each playback point in the point sequence is greater than or equal to the first playback duration, and less than or equal to the second playback duration.

[0136] According to one or more embodiments of the present disclosure, the method further includes: determining the positional relationship between the digital human video clip and the background video clip, and / or video reference information based on the screen content of the background video clip, wherein the video reference information includes transition effects and / or subtitles of the preview video; and generating the preview video based on the digital human video clip, the background video clip, and the positional relationship and / or the video reference information.

[0137] In a second aspect, according to one or more embodiments of the present disclosure, a digital human video generation device is provided, comprising:

[0138] Loading module, used to obtain spoken text and original video materials;

[0139] A generating module is configured to obtain a digital human video segment and a background video segment based on the spoken text and the original video material; wherein the spoken content of the digital human video segment matches the screen content of the background video segment;

[0140] The configuration module is used to configure the digital human video clip and the background video clip into corresponding editing tracks respectively, and display a preview video generated by the digital human video clip and the background video clip.

[0141] According to one or more embodiments of the present disclosure, the generation module is specifically used to: generate a digital human video clip based on the spoken text; obtain a point sequence based on the text content of the spoken text, and the point sequence includes at least two ordered playback points, each of the playback points corresponds to the start timestamp and / or end timestamp of a text clip; based on the picture content of the original video material, the original video material is segmented and extracted to generate at least one video material slice, and based on the point sequence, the video material slice that matches the text clip is combined to generate the background video clip.

[0142] According to one or more embodiments of the present disclosure, when the generation module obtains a point sequence based on the text content of the spoken text, it is specifically used to: obtain the text block timestamp corresponding to each text block in the spoken text; cluster the text blocks according to the text semantics of the text blocks to obtain at least one text segment; according to the text block timestamps of the text blocks within the text segment, the text segment obtains the corresponding playback point; according to the set of playback points corresponding to each text segment, obtain a point sequence.

[0143] According to one or more embodiments of the present disclosure, the generation module is further used to: obtain a video generation template, which is used to indicate a target number of video segments; when the generation module clusters the text blocks according to the text semantics of the text blocks to obtain at least one text segment, it is specifically used to: cluster the text blocks according to the text semantics of the text blocks and the target number to obtain the target number of text segments.

[0144] According to one or more embodiments of the present disclosure, when the generation module segments and extracts the original video material based on the picture content of the original video material to generate at least one video material slice, it is specifically used to: obtain the picture features of the original video material, and segment the original video material based on the picture features to obtain at least one video material slice; when the generation module generates the background video segment based on the video material slices that match the text segment based on the point sequence combination, it is specifically used to: obtain the picture semantics of each of the video material slices, and obtain the video material slice corresponding to each playback point based on the similarity between the picture semantics of each of the video material slices and the text semantics of the text segment corresponding to each playback point in the point sequence; generate the background video segment based on the video material slices corresponding to each playback point.

[0145] According to one or more embodiments of the present disclosure, when the generation module divides the original video material based on the picture features to obtain at least one video material slice, it is specifically used to: divide the original video material into at least two initial material slices based on the picture features of the original video material, and the feature distance between the picture features of each of the initial material slices is greater than a first feature threshold; obtain the confidence of the picture semantics corresponding to each initial material slice, and filter out the initial material slices whose confidence is less than the confidence threshold to obtain at least one video material slice.

[0146] According to one or more embodiments of the present disclosure, the digital human video clip has a first playback duration. When the generation module obtains a point sequence based on the text content of the spoken text, it is specifically used to: obtain a second playback duration based on the first playback duration and a preset duration ratio coefficient, and the second playback duration is greater than the first playback duration; obtain a point sequence based on the second playback duration and the text content of the spoken text, wherein the total playback duration of the video clip corresponding to each playback point in the point sequence is greater than or equal to the first playback duration, and less than or equal to the second playback duration.

[0147] According to one or more embodiments of the present disclosure, the configuration module is further used to: determine the positional relationship between the digital human video clip and the background video clip, and / or video reference information based on the screen content of the background video clip, wherein the video reference information includes the transition effects and / or subtitles of the preview video; and generate the preview video based on the digital human video clip, the background video clip, and the positional relationship and / or the video reference information.

[0148] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;

[0149] The memory stores computer-executable instructions;

[0150] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the digital human video generation method as described in the first aspect and various possible designs of the first aspect.

[0151] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the digital human video generation method described in the first aspect and various possible designs of the first aspect is implemented.

[0152] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the digital human video generation method as described in the first aspect and various possible designs of the first aspect.

[0153] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present disclosure.

[0154] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0155] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating a digital human video, characterized in that: include: Obtain spoken text and original video material; Based on the spoken text and the original video material, a digital human video clip and a background video clip are obtained; wherein the spoken content of the digital human video clip matches the screen content in the background video clip; The digital human video clip and the background video clip are respectively configured into corresponding editing tracks, and a preview video generated by the digital human video clip and the background video clip is displayed.

2. The method according to claim 1, characterized in that Based on the spoken text and the original video material, a digital human video clip and a background video clip are obtained, including: Generating a digital human video clip based on the spoken text; Based on the text content of the spoken text, a point sequence is obtained, wherein the point sequence includes at least two ordered playback points, each of the playback points corresponding to a start timestamp and / or an end timestamp of a text segment; Based on the screen content of the original video material, the original video material is segmented and extracted to generate at least one video material slice, and based on the point sequence, the video material slices that match the text segment are combined to generate the background video segment.

3. The method according to claim 2, characterized in that The step of obtaining a point sequence based on the text content of the spoken text includes: Obtaining a text block timestamp corresponding to each text block in the spoken text; Clustering the text blocks according to text semantics of the text blocks to obtain at least one text segment; According to the text block timestamp of the text block in the text segment, the text segment obtains a corresponding playback point; According to the set of playback points corresponding to each text segment, a point sequence is obtained.

4. The method according to claim 3, characterized in that The method further comprises: Obtaining a video generation template, wherein the video generation template is used to indicate a target number of video clips; The clustering of the text blocks according to the text semantics of the text blocks to obtain at least one text segment includes: The text blocks are clustered according to the text semantics of the text blocks and the target number to obtain the target number of text segments.

5. The method according to claim 3, characterized in that The step of segmenting and extracting the original video material based on the screen content of the original video material to generate at least one video material slice includes: Acquiring picture features of the original video material, and segmenting the original video material based on the picture features to obtain at least one video material slice; The step of combining the video material slices that match the text segment based on the point sequence to generate the background video segment includes: Obtaining the picture semantics of each of the video material slices, and obtaining the video material slice corresponding to each playback point based on the similarity between the picture semantics of each of the video material slices and the text semantics of the text segment corresponding to each playback point in the point sequence; The background video clips are generated based on the video material slices corresponding to each playback point.

6. The method according to claim 5, characterized in that The segmenting of the original video material based on the picture feature to obtain at least one video material slice includes: Slicing the original video material into at least two initial material slices according to the picture features of the original video material, wherein a feature distance between the picture features of the initial material slices is greater than a first feature threshold; The confidence level of the picture semantics corresponding to each initial material slice is obtained, and the initial material slices with the confidence level less than a confidence threshold are filtered out to obtain at least one video material slice.

7. The method according to claim 2, characterized in that The digital human video clip has a first playback duration, and the point sequence is obtained based on the text content of the spoken text, including: Obtaining a second playback duration based on the first playback duration and a preset duration ratio coefficient, wherein the second playback duration is greater than the first playback duration; Based on the second playback duration and the text content of the spoken text, a point sequence is obtained, wherein the total playback duration of the video clips corresponding to each playback point in the point sequence is greater than or equal to the first playback duration and less than or equal to the second playback duration.

8. The method according to claim 1, characterized in that The method further comprises: Determining the positional relationship between the digital human video clip and the background video clip, and / or video reference information based on the screen content of the background video clip, wherein the video reference information includes transition effects and / or subtitles of the preview video; The preview video is generated according to the digital human video clip, the background video clip, the positional relationship and / or the video reference information.

9. A digital human video generation device, characterized in that: include: Loading module, used to obtain spoken text and original video materials; A generating module is configured to obtain a digital human video segment and a background video segment based on the spoken text and the original video material; wherein the spoken content of the digital human video segment matches the screen content of the background video segment; The configuration module is used to configure the digital human video clip and the background video clip into corresponding editing tracks respectively, and display a preview video generated by the digital human video clip and the background video clip.

10. An electronic device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the digital human video generation method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the digital human video generation method according to any one of claims 1 to 8 is implemented.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for generating a digital human video according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Video editing method and apparatus, device, and storage medium

    WO2026157444A1