Video generation method and device, equipment, medium and product

By extracting feature descriptions and speech content from the target text, generating image diagrams and audio recordings, the high cost and time consumption of traditional actor-driven performances are solved, enabling fast and efficient book video processing and improving video quality.

CN121924322APending Publication Date: 2026-04-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-10-24
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, generating video versions of books through actor performances is costly and time-consuming, and cannot meet users' needs for video versions of various books.

Method used

By extracting feature descriptions and speech content of the target object from the target text, an image and audio recording of the speech are generated. The audio recording is then used to drive the image and construct a video corresponding to the target text.

Benefits of technology

It enables fast and automatic book video processing, improving video production efficiency and quality, and meeting users' needs for video production of various books.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121924322A_ABST
    Figure CN121924322A_ABST
Patent Text Reader

Abstract

The invention discloses a video generation method and device, equipment, a medium and a product, and the method comprises the steps: obtaining a target text, so as to enable the target text to describe stories among a plurality of objects; feature description information of the target object and speaking content of the target object are determined from the target text; then, according to the feature description information, generating an image picture of the target object, and according to the feature description information and the speaking content, generating a speaking audio of the target object; secondly, driving the image picture by using the speaking audio to obtain a speaking video of the target object; and finally, according to the speaking video, constructing a video corresponding to the target text, so that the video can represent the states of different objects under the speaking content thereof, thereby quickly realizing automatic video processing for the book, effectively overcoming the defects caused by video processing realized through the deduction of an actor, and improving the efficiency of the video processing. And the video effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a video generation method, apparatus, device, medium, and product. Background Technology

[0002] With the widespread use of electronic devices, more and more users tend to learn about certain books, such as novels, by watching videos on electronic devices, making the video processing of books a pressing technical problem to be solved. Summary of the Invention

[0003] This application provides a video generation method, apparatus, device, medium, and product, which are beneficial for improving video quality.

[0004] To achieve the above objectives, the technical solution provided in this application is as follows:

[0005] This application provides a video generation method, the method comprising: acquiring target text, the target text being used to describe multiple objects; determining feature description information of a target object and the speech content of the target object from the target text, the multiple objects including the target object; generating an image of the target object based on the feature description information, and generating audio of the target object's speech based on the feature description information and the speech content; using the speech audio to drive the image to obtain a speech video of the target object; and constructing a video corresponding to the target text based on the speech video.

[0006] In one possible implementation, the process of determining the target object includes: simplifying the target text to obtain simplified text and annotation information for each segment in the simplified text, wherein the simplified text has fewer characters than the target text, and for any segment, the annotation information includes the type of the segment and / or the affiliation of the segment, wherein the type is used to indicate whether the segment belongs to narration, and the affiliation is used to identify the speaker of the segment; and determining the target object based on the annotation information, wherein the simplified text includes the speech content of the target object.

[0007] In one possible implementation, the method further includes: obtaining simplification constraints, the simplification constraints indicating the degree of simplification for the target text, so as to simplify the target text according to the simplification constraints.

[0008] In one possible implementation, the method further includes: determining image style information based on the subject matter of the target text, so as to generate an image of the target object based on the image style information and the feature description information.

[0009] In one possible implementation, the image style information is further determined based on at least one item extracted from the target text; the at least one item includes scene information and / or text fragments used to indicate the image style.

[0010] In one possible implementation, the plurality of objects further includes a reference object, the image of which is generated earlier than the image of the target object; the method further includes: determining relationship description information between the reference object and the target object from the target text, so as to generate an image of the target object based on the feature description information, the relationship description information, and the image of the reference object.

[0011] In one possible implementation, the process of determining the spoken audio includes: determining the timbre of the target object based on the feature description information; and generating the spoken audio of the target object based on the timbre and the spoken content.

[0012] In one possible implementation, the method further includes: determining scene information corresponding to the speech content from the target text; generating a scene graph based on the scene information; and processing the speech video based on the scene graph to obtain a processed video, wherein the background of each frame in the processed video is determined based on the scene graph, so as to construct a video corresponding to the target text based on the processed video.

[0013] In one possible implementation, the method further includes: extracting narration content from the target text; determining the timbre of the narration content based on the subject matter of the target text; generating narration audio based on the timbre and the narration content, so as to construct a video corresponding to the target text based on the speaking video and the narration audio.

[0014] In one possible implementation, the method further includes: determining scene information corresponding to the narration content from the target text; generating a scene diagram based on the scene information; and generating a narration video based on the scene diagram and the narration audio, so as to construct a video corresponding to the target text based on the speaking video and the narration video.

[0015] In one possible implementation, the scene information is determined from the narration content.

[0016] This application provides a video generation apparatus, comprising: an acquisition unit for acquiring target text, the target text describing multiple objects; a determination unit for determining feature description information of a target object and speech content of the target object from the target text, the multiple objects including the target object; a generation unit for generating an image of the target object based on the feature description information, and generating speech audio of the target object based on the feature description information and the speech content; a driving unit for driving the image using the speech audio to obtain a speech video of the target object; and a construction unit for constructing a video corresponding to the target text based on the speech video.

[0017] This application provides an electronic device, the device comprising: a processor and a memory; the memory for storing instructions or computer programs; the processor for executing the instructions or computer programs in the memory to cause the electronic device to perform the video generation method provided in this application.

[0018] This application provides a computer-readable medium storing instructions or a computer program that, when executed on a device, causes the device to perform the video generation method provided in this application.

[0019] This application provides a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the video generation method provided in this application.

[0020] Compared with related technologies, this application has at least the following advantages:

[0021] The technical solution provided in this application first acquires target text, such as all chapters of a novel, to describe the story between multiple objects. Then, it determines the characteristic description information and speech content of the target objects from the target text, so that the characteristic description information represents the characteristics of the target objects, such as height and other physical features, and the speech content represents the speech content of the target objects in each story. Next, based on the characteristic description information, it generates an image of the target objects, so that the image can better describe the characteristics of the target objects visually. Finally, based on the characteristic description information and the speech content, it generates audio recordings of the target objects' speech, so that... The audio recording can represent the target object's state under the given speech content, such as pronunciation. Secondly, the audio recording drives the image to obtain a video of the target object speaking, allowing the video to better represent the target object's speaking state, such as lip movements. Finally, based on the video, a video corresponding to the target text is constructed, enabling the video to represent the state of different objects under their given speech content. This allows for rapid and automatic video processing of books, effectively overcoming the shortcomings of video processing achieved through actor performances, and thus improving video processing effects, such as increasing video processing efficiency.

[0022] In addition, since the feature description information of the target object is extracted from the target text, this feature description information can accurately and comprehensively represent the characteristics of the target object. As a result, the image generated based on this feature description information can better represent the image of the target object in the target text. Consequently, the video obtained based on the image can better represent the target object's speaking state, such as lip movements. This is beneficial for improving the video quality and other effects. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating a video generation method provided in this application embodiment;

[0025] Figure 2 A schematic diagram of a video generation process provided in an embodiment of this application;

[0026] Figure 3This is a schematic diagram of the structure of a video generation device provided in an embodiment of this application;

[0027] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0028] Research has found that in some scenarios, videos of specific books can be created through actor performances. However, this method has drawbacks such as high video production costs and long production times, limiting the selection of books for video production. Consequently, users can only learn about a limited number of books through video, failing to meet their desire to explore a diverse range of books via video.

[0029] Based on the above research, and to better overcome the aforementioned shortcomings, this application provides a video generation method, which includes: first, acquiring target text, such as all chapters of a novel, so that the target text can describe the story between multiple objects; then, determining the feature description information and speech content of the target objects from the target text, so that the feature description information can represent the characteristics of the target objects, such as physical features like height, and the speech content can represent the speech content of the target objects in each story; then, generating an image of the target objects based on the feature description information, so that the image can better describe the characteristics of the target objects through images, and generating a video based on the feature description information and the speech content. The process involves three main steps: First, the audio of the speech is used to represent the target object's state in relation to the given speech content, such as pronunciation. Second, the audio drives the image to generate a video of the target object's speech, which better represents the object's speaking state, such as lip movements. Finally, based on the video, a video corresponding to the target text is constructed, showing the different states of different objects within the given speech content. This allows for rapid and automatic video processing of books, effectively overcoming the shortcomings of video processing achieved through actor performances and improving video quality and efficiency, thus better meeting user needs.

[0030] Furthermore, this application does not limit the entity executing the video generation method. For example, the method can be applied to a terminal device or a server. Alternatively, the method can be implemented through data interaction between the terminal device and the server. The terminal device can be a smartphone, computer, personal digital assistant (PDA), tablet computer, etc. The server can be a standalone server, a cluster server, or a cloud server.

[0031] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0032] To better understand the technical solution provided in this application, the video generation method provided in this application will be explained below with reference to some accompanying drawings. For example... Figure 1 As shown, the video generation method provided in this application includes S1-S5 below.

[0033] S1: Get the target text, which is used to describe multiple objects.

[0034] The target text refers to the text that needs to be processed into a video, such as the entire text of a novel or other book.

[0035] Therefore, in one possible implementation, if the goal is to digitize a target book (such as a novel), the target text can include all chapters of the target book (e.g., ...). Figure 2 (The text of all chapters shown) is used to enable the target text to describe all the storylines described in the target book in a textual way, so that the target text can describe the relevant content of each object in the target book, such as appearance characteristics, speech content, and speech scenes.

[0036] It should be noted that this application does not limit the object. For example, the object can refer to a character appearing in a book, such as a character who can speak. As another example, since characters in a book are usually uniquely identified by their names, the object can be implemented using a character's name so that it can be used to uniquely identify a specific character in the book. Thus, different objects are used to represent different characters appearing in the target text.

[0037] Furthermore, this application does not limit the implementation of S1. For example, it can specifically be: after detecting a user's video request triggered for a target book, all chapter texts of the target book are determined as the target text.

[0038] S2: Determine the characteristic description information of the target object and the speech content of the target object from the target text. The target text contains multiple objects, including the target object.

[0039] The target object is used to represent the person who needs to appear in the video version of the target text; and this application does not limit the number of target objects.

[0040] Additionally, when all stories described by the target text need to be video-ified, the target object can be used to represent each of the multiple objects described by the target text, so that the number of characters appearing in the video-ified result of the target text is the same as the number of characters appearing in the target text; however, when only a portion of the story described by the target text needs to be video-ified, the target object can be used to represent the objects involved in that portion of the story, so that the number of characters appearing in the video-ified result is less than the number of characters appearing in the target text.

[0041] As can be seen, in some scenarios, such as the generation of novel advertisement videos or the videoization of the core story of a novel, the process of determining the target object may include steps 11-12 below.

[0042] Step 11: Simplify the target text to obtain the simplified text and the annotation information of each segment in the simplified text. The simplified text has fewer characters than the target text. For any segment, the annotation information includes the type of the segment and / or the affiliation of the segment. The type is used to indicate whether the segment belongs to narration, and the affiliation is used to identify the speaker of the segment.

[0043] The simplification process is used to extract the core storyline from a book; however, this application does not limit the simplification process. For example, it can employ any method capable of summarizing a book, such as using a machine learning model with summarization capabilities. Furthermore, this application does not limit the implementation method of the model; for example, it can be implemented using a Large Language Model (LLM).

[0044] The simplified text includes a portion of the target text, making the simplified text fewer in number than the target text. This results in the simplified text describing fewer story details than the target text, thus allowing the simplified text to be used to describe the core storyline of the target text. Here, "core" refers to the main part.

[0045] Therefore, in one possible implementation, when the target text comprises multiple chapter texts, the simplified text may include a story summary of each chapter text, so that the simplified text can describe the main part of the story described in each chapter text with fewer characters. Here, the story summary refers to the main content of a chapter text.

[0046] Furthermore, this application does not limit the simplified text; for example, it may include dialogue and narration. Therefore, in some scenarios, the simplified text can be implemented in the form of {speech content 1, speech content 2, narration content 1, speech content 3, narration content 2, ...}.

[0047] Furthermore, for any segment in the simplified text, the segment refers to a string existing in the simplified text, such as speech content 1 or narration content 1; and the annotation information of the segment is used to describe the attributes of the segment, so that the annotation information can indicate some characteristics of the segment in the target text. Specifically, the annotation information is automatically annotated for the segment based on the target text, so that the annotation information can indicate the state of the segment in the target text, such as the speaker, whether it is a dialogue, the order of appearance, and the scene.

[0048] In addition, for any segment in the simplified text, the tagging information of that segment can at least include the type of the segment. This type indicates whether the segment is a narration, allowing the type to take the values ​​of either dialogue or narration, thus indicating whether the segment exists in the target text as dialogue or narration.

[0049] Furthermore, for any segment in the simplified text, if the segment belongs to a dialogue, the tagging information for that segment may also include its attribution. This attribution identifies the speaker of the segment, with the attribution value being the name of a character appearing in the target text. This allows the attribution to indicate which character's speech the segment represents within the target text.

[0050] Furthermore, for any segment in the simplified text, the annotation information for that segment may also include its position within the text. This position indicates the order in which the segment appears in the final video output.

[0051] Furthermore, for any segment in the simplified text, the annotation information for that segment may also include the scene information corresponding to that segment. This scene information describes the context in which the segment occurs, enabling it to indicate the specific context within the target text in which the segment appears.

[0052] It should be noted that this application does not limit the above-mentioned scene information; for example, it can be implemented using text such as "A gloomy night, without a moon,..." Furthermore, this application does not limit the implementation method of the scene; for example, it can be implemented using a storyboard approach.

[0053] Research has found that any scene in a novel can be described using narration to indicate the environment in which the speakers are located when dialogue takes place, such as a market, winter, or night.

[0054] Based on the above research, it is known that the scene information in different segments of the simplified text may be the same or different, and this application does not impose any limitations on this. In addition, for any scene, the scene information used to describe the scene can be determined based on the narration content appearing in the scene, so that the scene information can accurately and comprehensively represent the environment in which each speaker in the scene is speaking.

[0055] Therefore, for any given scenario, if at least one spoken content and one narration appear in the scenario, the scenario information used to describe the scenario can be determined first based on the narration; then, this scenario information can be used as the scenario information corresponding to each spoken content and the scenario information corresponding to the narration in the scenario.

[0056] It should be noted that this application does not limit the method of determining scene information. For example, it can be implemented by any method that can extract scene information from text, such as by using a pre-built machine learning model (such as LLM) with scene information extraction function.

[0057] Research has revealed that the required degree of simplification varies across different application scenarios. Therefore, to better meet simplification requirements, step 11 above can specifically be implemented as follows: After obtaining the simplification constraints, the target text is simplified based on these constraints to obtain the simplified text and the annotation information for each segment within it. This ensures that the simplified text satisfies the simplification constraints, thereby improving video quality while meeting the video requirements of the current application scenario and enhancing flexibility. The simplification constraints indicate the degree of simplification for the target text, representing the required simplification level in the current application scenario. Furthermore, this application does not limit the implementation method of the simplification constraints; for example, it can be implemented using the constraint of "no more than 100 characters".

[0058] Based on the content of step 11, after obtaining the target text, the text is input into a machine learning model with core story extraction capabilities. This allows the model to summarize the text and extract the core story's wording (e.g., ...). Figure 2 (Chinese text), indicate the type and affiliation of each text in the text, and indicate the order of different texts in the text.

[0059] Step 12: Based on the above annotation information, determine the target object, and the simplified text includes the target object's speech content.

[0060] It should be noted that this application does not limit the implementation of step 12. For example, for any segment in the simplified text, when the annotation information of the segment includes the type of the segment, step 12 includes: if the type indicates that the segment belongs to a dialogue, then find the speaker of the segment from the target text as the target object, so as to filter the target object from multiple objects described in the target text according to the annotation information, so that the target object can represent the objects appearing in the simplified text, thereby enabling the target object to represent the speaker involved in video generation based on the simplified text.

[0061] For example, for any segment in the simplified text, when the annotation information of the segment includes at least the attribution of the segment, step 12 includes: determining the target object based on the attribution, so that the person described by the target object is the same person identified by the attribution, so that the target object can represent the object appearing in the simplified text, so that the target object can represent the speaker involved in the video generation based on the simplified text.

[0062] Based on the relevant content of steps 11 to 12, it can be seen that for some scenarios, such as those with simplification requirements, after obtaining the target text, the text can be summarized first; then the target object can be determined based on the summary results, so that the target object can represent the object that appears in the text and needs to be described through video.

[0063] Additionally, in some scenarios, such as Figure 2 In the scenario shown, the process of determining the target object can be as follows: First, extract all character names from the target text; then, based on these character names, perform disambiguation processing on the target text to obtain processed text, so that the processed text can more accurately represent the relevant content of each character; next, count the frequency of occurrence of each character name in the processed text; secondly, sort these character names according to these frequencies; finally, based on the sorting results, determine the target object so that the target object can represent the more important characters described by the target text, such as the characters involved in the core storyline.

[0064] The feature description information of the target object is used to describe the characteristics of the target object presented in the target text, such as height, appearance, clothing style, personality and other characteristics; moreover, the feature description information is extracted from the target text (or the text processed above) so that the feature description information can represent the content in the target text that describes the characteristics of the target object, so that the feature description information can describe the characteristics of the target object as comprehensively and accurately as possible.

[0065] It should be noted that this application does not limit the method of determining the above-mentioned feature description information. For example, it can be implemented using any method that can extract the characteristics of a person from the text, such as by using a machine learning model (such as LLM) with the function of extracting characteristics of a person.

[0066] The target subject's speech content refers to the spoken content that needs to appear in the video output. Therefore, in some scenarios, such as when a complete novel needs to be video-encoded, the target subject's speech content can include the spoken content present in the target text. In other scenarios, such as when a portion of a novel needs to be video-encoded, the target subject's speech content can include fragments belonging to the target subject that exist in the simplified text described above.

[0067] Based on the relevant content of S2, after obtaining the target text, some target objects, feature description information of each target object, and speech content of each target object are determined from the target text so that video processing can be performed based on this information.

[0068] S3: Generate an image of the target object based on its feature description information, and generate audio of the target object's speech based on its feature description information and speech content.

[0069] Among them, the image of the target object is used to describe the characteristics of the target object in the target text in a visual way.

[0070] Furthermore, this application does not limit the method of obtaining the image; for example, it can employ any method capable of generating character images based on some content, such as using a machine learning model with image generation capabilities. It should be noted that this application does not limit the implementation method of the model; for example, it can employ a diffusion model.

[0071] As can be seen, in one possible implementation, the process of generating the image of the target object can be as follows: first, generate conditional text (prompt) based on the feature description information of the target object, so that the semantic content described by the text includes the characteristics of the person described by the feature description information; then, input the text into a machine learning model with text-to-image function, and have the model perform image generation processing based on the text to obtain and output the image of the target object, so that the image satisfies the constraints described by the text.

[0072] Research has found that the visual effect of the same person varies under different image styles (such as art style). Therefore, in order to improve the video effect, the process of generating the image of the target object may include steps 21-22 below.

[0073] Step 21: Determine the image style information based on the subject matter of the target text.

[0074] The subject matter of the target text is used to indicate the genre to which the story described by the target text belongs, such as fantasy, urban, horror, etc.

[0075] Image style information is used to indicate the style of the characters in the target text. For example, the style of characters in fantasy novels includes ethereal and exquisite makeup; the style of characters in urban novels includes simple and elegant makeup; and the style of characters in horror novels includes ordinary, timid and cowardly appearance.

[0076] Furthermore, this application does not limit the implementation of step 21. For example, it can be: searching for image style information that best matches the subject matter of the target text from a pre-constructed first mapping relationship, thus enabling the image style to be determined through a retrieval method. The first mapping relationship is used to record image style information corresponding to a large number of subjects.

[0077] Research has revealed that some novel chapters contain image style constraints, such as dynasty and setting. Therefore, to improve accuracy, step 21 above can be specifically described as follows: based on the subject matter of the target text and at least one element extracted from the target text, determine the image style information so that the image style information satisfies both the image style constraints corresponding to the subject matter and the image style constraints described by the at least one element. Here, the at least one element refers to the content existing in the target text used to constrain the image style; however, this application does not limit the scope of this at least one element. For ease of understanding, three examples are provided below.

[0078] Example 1: At least one of the above-mentioned contents may include scene information extracted from the target text, so that the image style information determined based on the content at least satisfies the image style constraints described by the scene information. Therefore, if the target text describes a story in multiple scenes (such as stories in different seasons, stories in different places, etc.), the image style information is used to describe the character image style corresponding to each scene, so that the image style information can be used to obtain the image of the same character in different scenes. This effectively overcomes the inconsistency caused by using the same image style in different scenes.

[0079] Example 2: At least one of the above-mentioned contents may include text fragments extracted from the target text that indicate the image style, such as the term "XXX dynasty," so that the image style information determined based on this content at least satisfies the image style constraints indicated by the text fragment. Here, the text fragment refers to content in the target text that directly indicates the image style through textual means, such as words, phrases, paragraphs, etc.

[0080] Example 3: at least one of the above may include scene information extracted from the target text, and text fragments extracted from the target text to indicate the image style.

[0081] Based on the relevant content of step 21 above, after obtaining the target text and its subject matter, the image style information is determined according to the subject matter and the text, so that the image style information can represent the style characteristics of each character in the text in each scene as comprehensively and accurately as possible, such as appearance characteristics.

[0082] Step 22: Based on the above image style information and feature description information of the target object, generate an image of the target object.

[0083] It should be noted that this application does not limit the implementation of step 22. For example, it can specifically be: first, generate conditional text based on the above-mentioned image style information and feature description information of the target object, so that the semantic content described by the text includes the characteristics of the person described by the feature description information and the image style constraints described by the image style information; then input the text into a machine learning model with text-to-image generation function, and have the model perform image generation processing based on the text to obtain and output an image of the target object, so that the image satisfies the constraints described by the text.

[0084] Based on the relevant content of steps 21 to 22, for any target object, an image of the object can be generated according to the image style and some characteristics of the object, so that the image can better describe these characteristics while satisfying the image style, thereby making the video obtained based on the image have better quality.

[0085] Research has found that the relationships between different individuals influence the similarities and differences in their appearance. For example, if Person 1 and Person 2 are related by blood, their faces will show some similarities. Similarly, if Person 3 and Person 4 belong to the same department, their clothing will show some similarities. Furthermore, if Person 5 and Person 6 are unrelated, they will have almost no similarities.

[0086] Based on the above research, in order to improve accuracy, when the multiple objects described by the target text also include a reference object, and the image of the reference object is generated earlier than the image of the target object, the generation process of the image of the target object may include steps 31-32 below.

[0087] Step 31: Determine the relationship description information between the reference object and the target object from the target text, so that the relationship description information can accurately and comprehensively describe what kind of relationship exists between the two objects.

[0088] It should be noted that this application does not limit the above-mentioned relationship description information. For example, if the relationship description information is empty, it can be determined that the reference object and the target object have no relationship. If the relationship description information includes at least one string extracted from the target text, it can be determined that there is an association relationship between the two objects described by these strings.

[0089] It should also be noted that this application does not limit the method of determining the aforementioned relationship description information. For example, it can be implemented using any method capable of extracting relationships between different people from text, such as using a machine learning model with character relationship extraction capabilities. It should also be noted that this application does not limit the implementation method of this model.

[0090] Step 32: Generate an image of the target object based on the feature description information of the target object, the above relationship description information, and the image of the reference object.

[0091] It should be noted that this application does not limit the implementation of step 32 above. For example, it can specifically be: first, based on the feature description information of the target object and the above-mentioned relationship description information, determine the conditional text so that the semantic content described by the text includes the characteristics of the person described by the feature description information and the relationship constraints described by the relationship description information, so that the text is at least used to indicate the relationship between the reference object and the target object, and thus the text can, to a certain extent, represent the similarities and differences between the image of the reference object and the image of the target object; then, input the text and the image of the reference object into a machine learning model with text-to-image generation function, and have the model perform image generation processing based on these two types of data to obtain and output the image of the target object, so that the image satisfies the constraints described by the text, and that the similarities and differences between the image of the target object and the image of the reference object satisfy the relationship constraints, which is beneficial to improving the generation effect.

[0092] It should also be noted that this application does not limit the method of determining the "conditional text" in the above paragraph. For example, in order to better improve the generation effect, it can be: determining the conditional text based on the feature description information of the target object, the above-mentioned relationship description information and image style information, so that the semantic content described by the text includes the constraints described by these three types of information, so that the image generated based on the text satisfies the image constraints described by these three types of information.

[0093] Based on the relevant content of steps 31 to 32, for any target object, an image of the target object can be generated at least based on the relationship between the object and other objects, as well as the image of the other objects, so that the differences between the image images of different objects satisfy the appearance differences constraints represented by the relationship between these objects, thereby making the video obtained based on these image images have better quality.

[0094] In addition, to improve the generation effect, the process of determining the image of the target object can include at least the following: first, generating multiple images based on the feature description information of the target object (such as images obtained by any of the image generation methods mentioned above); then determining the scoring results of each image; and finally, selecting one image from the multiple images as the image of the target object based on these scoring results. This is beneficial to improving the generation effect.

[0095] It should be noted that, for any image, the rating result is used to indicate the degree of fit between the image depicted by the image and the image depicted by the target text for the target object; and this application does not limit the method of determining the rating result. For example, it may be determined based on the feature description information of the target object, the aforementioned relationship description information, and the image style information, so that the rating result can at least indicate the degree of satisfaction of the image with the image constraints described by these three types of information.

[0096] The target object's speech audio is used to describe the target object's speech content through audio; moreover, the speech audio is generated based on the target object's feature description information and the target object's speech content, so that the speech characteristics presented by the speech audio are as consistent as possible with the characteristics of the person described by the feature description information, and that the semantic information carried by the speech audio is the same as the semantic information carried by the speech content.

[0097] Furthermore, this application does not limit the method of generating the aforementioned audio recordings. For example, it can be implemented using any method that can generate audio based on the characteristics of the person and the content of the speech, such as a machine learning model with audio generation capabilities.

[0098] In addition, to improve the generation effect, the process of generating the target object's speech audio may include steps 41-42 below.

[0099] Step 41: Based on the characteristic description information of the target object, determine the timbre of the target object so that the timbre matches the characteristics of the person described by the characteristic description information to the highest degree.

[0100] It should be noted that this application does not limit the implementation of step 41. For example, it can specifically be: searching for the timbre that best matches the feature description information of the target object from a pre-constructed second mapping relationship, so as to determine the timbre through a retrieval method. The second mapping relationship includes some timbres and tags for each timbre. For any timbre, the tag of the timbre is used to describe the characteristics of the person to whom the timbre is applicable, such as personality traits.

[0101] For example, to improve flexibility, step 41 can specifically be: first, based on the feature description information of the target object, determine multi-dimensional timbre parameters to maximize the compatibility between the timbre characteristics described by the parameters and the character characteristics described by the feature description information; then, drive the timbre adjustment tool based on the parameters to obtain the timbre of the target object. The timbre adjustment tool is used to adjust the timbre based on the input timbre parameters; moreover, this application does not limit the implementation method of the tool. For example, it can be implemented using any tool capable of adjusting timbre based on parameters, such as a voice changer or a machine learning model.

[0102] Step 42: Generate the target's audio based on the target's timbre and speech content.

[0103] It should be noted that this application does not limit the implementation of step 42. For example, it can be: based on the timbre of the target object and the speech content of the target object, calling a text-to-speech (TTS) model to synthesize the speech audio of the target object.

[0104] Based on the relevant content of steps 41 to 42, for any target object, the appropriate timbre for the object can be determined first based on the object's characteristics; then, according to the timbre, the object's speech content can be converted into audio so that the speaking characteristics presented by the audio are highly compatible with the object, thereby making the video obtained based on the audio better.

[0105] Furthermore, this application does not limit the relationship between the execution time of the audio generation process and the execution time of the image generation process; for example, the former may precede the latter. Or, the latter may precede the former. Or, they may be the same.

[0106] Based on the relevant content of S3, for any target object, after extracting the characteristics of the object from the target text, an image of the object can be generated based on these characteristics, and the object's speech content can be converted into audio based on these characteristics, so that the speaking characteristics presented by the audio are highly compatible with the character characteristics presented by the image. This helps to overcome the abruptness caused by the mismatch between the two, thereby improving the generation effect.

[0107] S4: Use the target object's audio to drive the target object's image, and obtain the target object's video.

[0108] It should be noted that this application does not limit the implementation of S4. For example, it can be: inputting the target object's speech audio and the target object's image into a machine learning model with single-image-driven function, so that the model can use the audio to drive the image to generate a speech video of the target object, so that the image of the person presented in the video is consistent with the image of the person described in the image, and the facial expressions (such as lip movements) presented in the video are consistent with the pronunciation characteristics presented in the audio, so that the video can represent the state of the target object under the audio.

[0109] It should also be noted that this application does not limit the implementation of the model in the above paragraph. For example, it can be implemented using any machine learning model that can generate spoken video based on audio and human images.

[0110] S5: Based on the target subject's speaking video, construct the video corresponding to the target text.

[0111] It should be noted that this application does not limit the implementation of S5. For example, it can specifically be: splicing the video messages corresponding to each message according to their order to obtain the video corresponding to the target text. Here, for any given message, the video message corresponding to that message refers to the video generated based on that message.

[0112] Based on the content of S1 to S5 above, the video generation method in this application first obtains target text, such as all chapters of a novel, so that the target text can describe the story between multiple objects; then, it determines the feature description information and speech content of the target objects from the target text, so that the feature description information can represent the characteristics of the target objects, such as height and other physical features, and that the speech content can represent the speech content of the target objects in each story; then, based on the feature description information, it generates an image of the target objects so that the image can better describe the characteristics of the target objects in an image manner, and based on the feature description information and the speech content, it generates the target... The process involves three main steps: First, the target object's audio is used to represent its state (e.g., pronunciation) within the given speech content. Second, the audio drives the image to generate a video representation of the target object's speech, showcasing its speaking state (e.g., lip movements). Finally, based on this video, a corresponding video for the target text is constructed, demonstrating the different states of various objects within the given speech content. This method enables rapid and automatic video processing of books, effectively overcoming the limitations of video processing relying solely on actors' performances and improving overall video quality and efficiency.

[0113] Research has shown that different stories may take place in different settings. Therefore, in order to improve the video quality, the video generation method provided in this application may include at least steps 51-54 below.

[0114] Step 51: For the target's speech content, determine the scene information corresponding to the speech content from the target text, so that the scene information can describe the environment in which the target is speaking, such as a market.

[0115] Step 52: Based on the above scene information, generate a scene diagram so that the scene diagram can describe the environment described by the scene information in an image format.

[0116] It should be noted that this application does not limit the implementation of step 52. For example, it can be: first, generate conditional text based on scene information so that the semantic content carried by the text includes the scene characteristics described by the scene information; then input the text into a machine learning model with text-to-image function, and have the model perform image generation processing based on the text to obtain and output a scene graph so that the scene graph satisfies the constraints described by the text.

[0117] Step 53: Based on the above scene diagram, process the target object's speaking video to obtain the processed video. The background of each frame in the processed video is determined based on the scene diagram.

[0118] It should be noted that this application does not limit the implementation method of step 53. For example, if the target's speaking video does not have a background, the scene image can be directly superimposed on the speaking video as a background to obtain the processed video. Alternatively, if the speaking video does have a background, the scene image can be used to replace the background in the speaking video to obtain the processed video.

[0119] Step 54: Based on the processed video above, construct the video corresponding to the target text.

[0120] It should be noted that this application does not limit the implementation of step 54. For example, it can specifically be: splicing the processed videos corresponding to the various statements in the order they are arranged to obtain the video corresponding to the target text. Specifically, for any statement, the processed video corresponding to that statement is obtained by processing the video corresponding to that statement based on the aforementioned scene diagram.

[0121] Based on the relevant content of steps 51 to 54, for any given speech content, the speech video corresponding to that content can be adjusted according to the scene information to obtain the processed video corresponding to that content. This ensures that the background presented in the processed video is consistent with the environment described by the scene information, thereby enabling the processed video to better present the speech state of the content, which is beneficial to improving the video quality.

[0122] In addition, in order to better improve the video quality, the video generation method provided in this application may include at least steps 61-64 below.

[0123] Step 61: Extract narration content from the target text to help users better understand relevant information about the speech.

[0124] Step 62: Determine the timbre of the narration content based on the subject matter of the target text, so that the timbre is as compatible as possible with the subject matter.

[0125] It should be noted that this application does not limit the implementation of step 62. For example, it can be: searching for the timbre that best matches the subject matter of the target text from a pre-constructed third mapping relationship, so as to determine the narration timbre through retrieval. The third mapping relationship is used to record timbres suitable for certain subjects in a key-value pair manner, such as a more subdued narration timbre for urban novels and a more low and husky narration timbre for horror novels.

[0126] Step 63: Generate narration audio based on the timbre and content of the narration.

[0127] It should be noted that this application does not limit the implementation of step 63. For example, it can be: based on the timbre and content of the narration, calling a TTS model to synthesize and obtain the narration audio of the target object.

[0128] Step 64: Based on the target subject's speaking video and the aforementioned narration audio, construct the video corresponding to the target text.

[0129] It should be noted that this application does not limit the implementation of step 64. For example, it can specifically be: after obtaining the order of all spoken content and all narration content, inputting the order, the spoken video (or processed video) corresponding to each spoken content, and the narration audio corresponding to each narration content into a video synthesis tool, so that the tool can construct a video corresponding to the target text based on this information. Specifically, for any narration content, the narration audio corresponding to that narration content is generated based on that narration content.

[0130] It should also be noted that the video compositing tool is used to combine some videos and some audio in a certain order to obtain a new video; and this application does not limit the tool, for example, it can be implemented using the Fast Forward Mpeg (ffmpeg) tool.

[0131] For example, in order to improve the generation effect, step 64 may include steps 641-644 below.

[0132] Step 641: Determine the scene information corresponding to the narration content from the target text.

[0133] It should be noted that this application does not limit the implementation of step 641. For example, it can be: after extracting the narration content from the target text, determining scene information based on the narration content, so that the scene information can represent the environment described by the narration content, so that the scene information can be used to constrain the speaking state of the speech content corresponding to the narration content. Wherein, for narration content and speech content with a corresponding relationship, both occur in the same scene.

[0134] Step 642: Generate a scene diagram based on the above scene information.

[0135] It should be noted that the relevant content of step 642 can be found in step 52 above.

[0136] Step 643: Based on the above scene diagram and narration audio, generate a narration video so that the background described in the video is consistent with the scene in which the narration content described in the audio takes place.

[0137] It should be noted that this application does not limit the implementation of step 643. For example, it can be: for any narration content, superimposing the narration audio corresponding to the narration content onto the scene graph corresponding to the narration content to obtain the narration video corresponding to the narration content. The scene graph is generated based on the scene information corresponding to the narration content. Furthermore, this application does not limit the implementation method of this superposition. For example, it can be implemented using any method capable of generating video based on audio and background graphs, such as a machine learning model or tool (e.g., ffmpeg) with video generation capabilities.

[0138] Step 644: Based on the target subject's speaking video and the aforementioned narration video, construct the video corresponding to the target text.

[0139] It should be noted that this application does not limit the implementation of step 644. For example, it can specifically be: after obtaining the arrangement order between all spoken content and all narration content, according to the arrangement order, splicing the spoken video (or processed video) corresponding to these spoken content and the narration audio corresponding to these narration content to obtain the video corresponding to the target text.

[0140] Based on the relevant content of steps 641 to 644, for spoken content and narration content in the same scene, a scene diagram of the scene can be constructed first; then, based on the scene diagram, videos corresponding to these two types of content can be constructed respectively, so that the backgrounds described by these videos are consistent with the scene diagram, thus making these videos have the same background; then, based on these videos, videos corresponding to the target text can be constructed, so that the videos can at least describe the dialogue and narration that occur in the scene, which is conducive to improving the video effect.

[0141] In addition, to further improve the video quality, the video generation method provided in this application may also include video post-processing as shown in steps 71-72 below.

[0142] Step 71: For any scene image, determine the scene sound corresponding to the scene image based on the scene information corresponding to the scene image, so as to maximize the degree of adaptation between the scene sound and the scene image.

[0143] It should be noted that this application does not limit the implementation of step 71. For example, it can specifically be: for any scene information corresponding to a scene graph, find the background sound that best matches the scene information from the pre-constructed fourth mapping relationship, and use it as the scene sound corresponding to the scene graph, so as to realize the determination of scene sound through retrieval. The fourth mapping relationship is used to record some background sounds and the scene characteristics used by these background sounds.

[0144] Step 72: For any scene image, add the scene sound corresponding to the scene image as background sound to the video segment in the video corresponding to the target text that has the scene image as background, so as to obtain the processed video corresponding to the target text, so that the background of each frame in the processed video is more compatible with its corresponding background sound.

[0145] Based on the relevant content of steps 71 to 72, after obtaining the video corresponding to the target text, background sound can be added to the video so that the background of each frame in the video is highly adapted to its corresponding background sound, thereby enabling the video to better represent the scene where each story takes place, which is beneficial to improving the video effect.

[0146] In addition, to further improve the video rendering effect, when the video corresponding to the target text is used to describe a story in multiple sequentially arranged scenes, the following processing can be performed on the video: For any two adjacent scenes, randomly select an effect from a pre-built transition effect library as the transition effect between the two adjacent scenes, and use this transition effect to process the video segments of the two adjacent scenes in the video corresponding to the target text to obtain the processed video corresponding to the target text, so that the processed video can present a more natural and smooth scene transition effect, which is beneficial to improving the video rendering effect.

[0147] Based on the video generation method provided in this application, this application also provides a video generation apparatus, such as... Figure 3 As shown, the video generation apparatus 300 provided in this application includes:

[0148] Acquisition unit 301 is used to acquire target text, which is used to describe multiple objects;

[0149] The determining unit 302 is used to determine the feature description information of the target object and the speech content of the target object from the target text, wherein the plurality of objects includes the target object;

[0150] The generation unit 303 is configured to generate an image of the target object based on the feature description information, and to generate the audio of the target object's speech based on the feature description information and the speech content;

[0151] The driving unit 304 is used to drive the image using the spoken audio to obtain the spoken video of the target object;

[0152] The construction unit 305 is used to construct the video corresponding to the target text based on the spoken video.

[0153] In one possible implementation, the process of determining the target object includes: simplifying the target text to obtain simplified text and annotation information for each segment in the simplified text, wherein the number of characters in the simplified text is less than the number of characters in the target text, and for any segment, the annotation information includes the type of the segment and / or the affiliation of the segment, wherein the type is used to indicate whether the segment belongs to narration, and the affiliation is used to identify the speaker of the segment; and determining the target object based on the annotation information, wherein the simplified text includes the speech content of the target object.

[0154] In one possible implementation, the process of determining the target object further includes: obtaining simplification constraints, which indicate the degree of simplification for the target text, so as to simplify the target text according to the simplification constraints.

[0155] In one possible implementation, the determining unit 302 is further configured to: determine image style information based on the subject matter of the target text;

[0156] The generation unit 303 is specifically used to: generate an image of the target object based on the image style information and the feature description information.

[0157] In one possible implementation, the image style information is further determined based on at least one item extracted from the target text; the at least one item includes scene information and / or text fragments used to indicate the image style.

[0158] In one possible implementation, the plurality of objects further includes a reference object, the image of which is generated earlier than the image of the target object.

[0159] The determining unit 302 is further configured to: determine the relationship description information between the reference object and the target object from the target text;

[0160] The generation unit 303 is specifically used to: generate an image of the target object based on the feature description information, the relationship description information, and the image of the reference object.

[0161] In one possible implementation, the generation unit 303 is specifically used to: determine the timbre of the target object based on the feature description information; and generate the speech audio of the target object based on the timbre and the speech content.

[0162] In one possible implementation, the process of determining the video corresponding to the target text includes: determining scene information corresponding to the speech content from the target text; generating a scene map based on the scene information; processing the speech video based on the scene map to obtain a processed video, wherein the background of each frame in the processed video is determined based on the scene map; and constructing the video corresponding to the target text based on the processed video.

[0163] In one possible implementation, the process of determining the video corresponding to the target text includes: extracting narration content from the target text; determining the timbre of the narration content based on the subject matter of the target text; generating narration audio based on the timbre and the narration content; and constructing the video corresponding to the target text based on the speaking video and the narration audio.

[0164] In one possible implementation, the process of determining the video corresponding to the target text includes: determining scene information corresponding to the narration content from the target text; generating a scene diagram based on the scene information; generating a narration video based on the scene diagram and the narration audio; and constructing the video corresponding to the target text based on the speaking video and the narration video.

[0165] In one possible implementation, the scene information is determined from the narration content.

[0166] Based on the relevant content of the video generation device 300, its working principle includes: first, acquiring target text, such as all chapters of a novel, so that the target text can describe the story between multiple objects; then, determining the characteristic description information and speech content of the target objects from the target text, so that the characteristic description information can represent the characteristics of the target objects, such as height and other physical features, and that the speech content can represent the speech content of the target objects in each story; then, generating an image of the target objects based on the characteristic description information, so that the image can better describe the characteristics of the target objects visually, and generating a video generator based on the characteristic description information and the speech content. The process involves three main steps: First, the target object's audio is used to represent its state (e.g., pronunciation) within the given speech content. Second, the audio drives the image to generate a video representation of the target object's speech, showcasing its speaking state (e.g., lip movements). Finally, based on this video, a corresponding video for the target text is constructed, demonstrating the different states of various objects within the given speech content. This method enables rapid and automatic video processing of books, effectively overcoming the limitations of video processing relying solely on actors' performances and improving overall video quality and efficiency.

[0167] It should be noted that for the technical details of the video generation apparatus provided in this application embodiment, please refer to the relevant content of the video generation method above. For the sake of brevity, it will not be repeated here.

[0168] In addition, this application also provides an electronic device, the device including a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device performs any implementation of the video generation method provided in this application.

[0169] See Figure 4 This diagram illustrates a structural schematic of an electronic device 400 suitable for implementing embodiments of the present disclosure. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0170] like Figure 4As shown, electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. The processing device 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.

[0171] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0172] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 409, or installed from a storage device 408, or installed from a ROM 402. When the computer program is executed by the processing device 401, it performs the functions defined in the methods of embodiments of this disclosure.

[0173] The electronic device provided in this embodiment belongs to the same inventive concept as the method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0174] This application also provides a computer-readable medium storing instructions or a computer program that, when executed on a device, causes the device to perform any implementation of the video generation method provided in this application.

[0175] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0176] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0177] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0178] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, enable the electronic device to perform the aforementioned methods.

[0179] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0180] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0181] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units / modules do not necessarily limit the specific unit itself.

[0182] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0183] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0184] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.

[0185] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0186] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0187] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0188] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video generation method, characterized in that, The method includes: Obtain the target text, which is used to describe multiple objects; The feature description information of the target object and the speech content of the target object are determined from the target text, wherein the plurality of objects includes the target object; Based on the feature description information, an image of the target object is generated, and based on the feature description information and the speech content, the audio of the target object's speech is generated; The speech audio is used to drive the image, resulting in a speech video of the target object. Based on the spoken video, construct the video corresponding to the target text.

2. The method according to claim 1, characterized in that, The process of determining the target object includes: The target text is simplified to obtain simplified text and annotation information for each segment in the simplified text. The simplified text has fewer characters than the target text. For any segment, the annotation information includes the type of the segment and / or the affiliation of the segment. The type is used to indicate whether the segment belongs to narration, and the affiliation is used to identify the speaker of the segment. Based on the annotation information, the target object is determined, and the simplified text includes the speech content of the target object.

3. The method according to claim 2, characterized in that, The method further includes: Obtain simplification constraints, which indicate the degree of simplification for the target text; The simplification process for the target text includes: The target text is simplified based on the simplification constraints.

4. The method according to claim 1, characterized in that, The method further includes: Based on the subject matter of the target text, determine the image style information; The step of generating an image of the target object based on the feature description information includes: Based on the image style information and the feature description information, an image of the target object is generated.

5. The method according to claim 4, characterized in that, The image style information is also determined based on at least one item extracted from the target text; The at least one of the contents includes scene information and / or text fragments used to indicate the style of the image.

6. The method according to claim 1, characterized in that, The plurality of objects also includes a reference object, wherein the image of the reference object is generated earlier than the image of the target object. The method further includes: Determine the relationship description information between the reference object and the target object from the target text; The step of generating an image of the target object based on the feature description information includes: Based on the feature description information, the relationship description information, and the image of the reference object, an image of the target object is generated.

7. The method according to claim 1, characterized in that, The process of determining the spoken audio includes: Based on the feature description information, the timbre of the target object is determined; Based on the timbre and the content of the speech, the speech audio of the target object is generated.

8. The method according to claim 1, characterized in that, The method further includes: Determine the scene information corresponding to the speech content from the target text; Based on the scene information, generate a scene diagram; Based on the scene diagram, the speaking video is processed to obtain a processed video, wherein the background of each frame in the processed video is determined based on the scene diagram; The step of constructing the video corresponding to the target text based on the spoken video includes: Based on the processed video, construct the video corresponding to the target text.

9. The method according to claim 1, characterized in that, The method further includes: Extract the narration content from the target text; Based on the subject matter of the target text, determine the timbre of the narration content; Based on the timbre and the narration content, generate narration audio; The step of constructing the video corresponding to the target text based on the spoken video includes: Based on the spoken video and the narration audio, a video corresponding to the target text is constructed.

10. The method according to claim 9, characterized in that, The method further includes: Determine the scene information corresponding to the narration content from the target text; Based on the scene information, generate a scene diagram; Based on the scene diagram and the narration audio, generate a narration video; The step of constructing the video corresponding to the target text based on the spoken video and the narration audio includes: Based on the spoken video and the narration video, a video corresponding to the target text is constructed.

11. The method according to claim 10, characterized in that, The scene information is determined from the narration content.

12. A video generation apparatus, characterized in that, include: An acquisition unit is used to acquire target text, which is used to describe multiple objects; A determining unit is configured to determine feature description information of a target object and the speech content of the target object from the target text, wherein the plurality of objects includes the target object; The generation unit is configured to generate an image of the target object based on the feature description information, and to generate audio of the target object's speech based on the feature description information and the speech content; The driving unit is used to drive the image using the spoken audio to obtain the spoken video of the target object; The construction unit is used to construct the video corresponding to the target text based on the spoken video.

13. An electronic device, characterized in that, The device includes: a processor and a memory; The memory is used to store instructions or computer programs; The processor is configured to execute the instructions or computer program in the memory to cause the electronic device to perform the method according to any one of claims 1-11.

14. A computer-readable medium, characterized in that, The computer-readable medium stores instructions or computer programs that, when executed on the device, cause the device to perform the method according to any one of claims 1-11.

15. A computer program product, characterized in that, It includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the method according to any one of claims 1-11.