Video generation method and device, equipment and storage medium
By extracting the environmental and plot descriptions from literary works and combining them with elements from film and television content to generate video images, the inconsistency problem of unadapted content in film and television works is solved, enabling the generation of highly conflict-driven videos and enhancing the viewing experience.
Patent Information
- Application Number
- CN202511546786.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-01-23
AI Technical Summary
When adapting literary works into film and television works, existing technologies often result in inaccurate feature matching during the video generation process of the unadapted content. This leads to poor consistency between the video and the unadapted content, a lack of conflict and appeal, and an impact on the viewing experience.
Based on the ungenerated film and television text of the target literary work, the environmental description text and plot description text of each scene are extracted, semantic description information is generated according to the set format, and combined with the content element set of the film and television work, a set of descriptive images is generated, and the images are generated through the text-to-image model.
Accurately generate video images associated with unadapted content, improve the consistency between the video and the unadapted content, increase the video's conflict and appeal, enhance the visual presentation, and improve the viewing experience.
Smart Images

Figure CN121397265A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to the technical field of artificial intelligence, and provides a video generation method and device, equipment and a storage medium. BACKGROUND
[0002] With the development of artificial intelligence technology, the entertainment function of mobile devices is increasingly enhanced. Adapting literary works into video works has become a popular cross-media integration method. However, in the process of adapting literary works into video works, due to reasons such as page limit, plot selection or censorship risk, selective script adaptation and shooting of literary works into video works are usually performed, that is, only part of the chapters in the literary works are shot, resulting in incomplete original plot of the literary works and affecting the viewing experience.
[0003] In order to improve the viewing experience, an implementation manner for generating a video for unadapted content in a literary work is given in the related art. In the process of generating the video, first, the already shot video work is subjected to frame extraction processing to obtain corresponding frame images, second, image feature extraction is performed on the frame images by a cross-modal model to obtain corresponding frame image features, and the frame image features are stored in an image feature library, third, a paragraph is extracted from the literary work for unadapted content, and text feature extraction is performed on the paragraph by the cross-modal model to obtain corresponding text features, fourth, the text features are matched with the frame image features stored in the image feature library, and the frame images corresponding to the frame image features that are successfully matched are taken as video images associated with the paragraph, and finally, the matched multiple video images are spliced to generate continuous playable images for the unadapted content, and a voice-over is added to form a video.
[0004] When the video images are obtained by feature matching, the features cannot be completely matched, at which time there will be a lot of content in the video images that does not match the paragraph, affecting the viewing experience. At the same time, the already shot video images are directly applied to the unshot content without adaptation or modification, resulting in poor consistency between the generated video and the unadapted content, and the video images in the generated video are all video images in the already shot video work, which will result in lack of conflict and lack of highlights in the generated video, further affecting the viewing experience.
[0005] Therefore, how to accurately generate video images associated with unadapted content and further ensure that the generated video has high conflict and improves the visual presentation effect is a technical problem to be solved at present. SUMMARY
[0006] The embodiments of the present application provide a video generation method, device, equipment and storage medium to accurately generate video images associated with unadapted content and further ensure that the generated video has high conflict and improves the visual presentation effect.
[0007] The embodiment of the application provides a video generation method, which comprises the following steps: Based on the to-be-converted text of the target literary work, environment description text and plot description text of at least one scene are extracted; wherein the to-be-converted text is a text part of the target literary work which has not generated a video work; For each scene, the following is performed: based on the environment description text and the plot description text of the scene, at least one semantic description information is generated in a set format; wherein each semantic description information is used for describing at least one of the environment layout and the subject object in the scene; For each generated semantic description information, the following is performed: based on the semantic description information, a corresponding description image set is generated by combining a content element set in the video work which has been generated based on the target literary work; wherein each content element is a reference layout or a reference object in the video work; Based on the description image set corresponding to each semantic description information, a target video is obtained.
[0008] The embodiment of the application provides a video generation device, which comprises the following steps: An extraction unit is configured to extract environment description text and plot description text of at least one scene based on to-be-converted text of a target literary work; wherein the to-be-converted text is a text part of the target literary work which has not generated a video work; A first generation unit is configured to perform the following for each scene: based on the environment description text and the plot description text of the scene, at least one semantic description information is generated in a set format; wherein each semantic description information is used for describing at least one of the environment layout and the subject object in the scene; A second generation unit is configured to perform the following for each generated semantic description information: based on the semantic description information, a corresponding description image set is generated by combining a content element set in the video work which has been generated based on the target literary work; wherein each content element is a reference layout or a reference object in the video work; An obtaining unit is configured to obtain a target video based on the description image set corresponding to each semantic description information.
[0009] In a possible implementation manner, the extraction unit is specifically configured to: The to-be-converted text is split to obtain at least one scene information; wherein the scene information is part of the text content in the to-be-converted text; For each scene information, the following is performed: based on pre-constructed visual presentation prompt information, corresponding environment description text and plot description text are extracted from the scene information.
[0010] In a possible implementation manner, the extraction unit is specifically configured to: Identify chapter titles from the text to be converted, and split the text to be converted into multiple subtexts according to the chapter titles; each subtext corresponds to a chapter title; For each subtext, the following is performed: based on the set game split condition, the subtext is split into at least one game information.
[0011] In a possible implementation, the extraction unit is specifically configured to: Split the text to be converted based on the pre-constructed dramatic event template to obtain at least one game information; The dramatic event template is used to describe the interactive plot characteristics between at least two subject objects in the reference text.
[0012] In a possible implementation, the first generation unit is specifically configured to: Based on the environment description text, generate at least one initial description information in a set format; the set format includes external environment parameters and subject feature parameters; the external environment parameters are used to describe the visual context of the environment layout, and the subject feature parameters are used to describe the visual performance of the subject object; For each initial description information, the following is performed: subject object matching between the plot description text and the environment description text is performed to obtain a corresponding matching result, and based on the matching result, the initial description information is combined to generate corresponding semantic description information.
[0013] In a possible implementation, the first generation unit is specifically configured to: When the matching result represents a missing subject object, the initial description information is supplemented based on the plot description text to generate corresponding semantic description information; When the matching result represents no missing subject object, the initial description information is used as the semantic description information.
[0014] In a possible implementation, the content element set is obtained in the following manner: Frame extraction is performed on the video work to obtain at least one key video frame; For each key video frame, the following is performed: when a reference object is identified from the key video frame, a content element of a target size is extracted from the key video frame with the reference object as the center, or when a reference layout is identified from the key video frame, a content element of a target size is extracted from the key video frame; The target size meets the vertical screen display condition.
[0015] In a possible implementation, the second generation unit is specifically configured to: Perform word segmentation processing on the semantic description information to obtain at least one keyword; retrieve, for each keyword, a corresponding content element from the content element set; splice the retrieved at least one content element to obtain at least one initial image; generate a corresponding description image set based on the at least one initial image.
[0016] In a possible implementation, the second generation unit is specifically configured to: perform visual evaluation on the at least one initial image respectively to obtain a respective visual evaluation value of the at least one initial image; select, based on the at least one visual evaluation value, an initial image that meets an evaluation condition from the at least one initial image; compose the selected initial image into the description image set.
[0017] In a possible implementation, the second generation unit is specifically configured to: perform word segmentation processing on the semantic description information to obtain at least one keyword; retrieve, for each keyword, a corresponding content element from the content element set; perform denoising processing on at least one noise image feature associated with the retrieved content element based on the text feature extracted from the semantic description information to obtain a corresponding denoised image feature; perform description image prediction on each denoised image feature to generate a corresponding description image; generate a corresponding description image set based on the generated at least one description image.
[0018] In a possible implementation, the steps of performing denoising processing on at least one noise image feature associated with the retrieved content element based on the text feature extracted from the semantic description information to obtain a corresponding denoised image feature, and performing description image prediction on each denoised image feature to generate a corresponding description image are performed by a target text-to-image model; The target text-to-image model is obtained by performing cyclic iteration training on a training set based on a text-image sample pair. The text-image sample pair includes a sample image and a description text of the sample image extracted from a film and television work, and the description text is generated in a set format.
[0019] In a possible implementation, the obtaining unit is specifically configured to: perform, for each semantic description information, the following operations respectively: extracting subtitle information of the corresponding description image set from plot description information corresponding to the semantic description information, and converting the subtitle information into voice information; For each session, the following is performed: based on at least one semantic description information associated with the session, the arrangement order in the part of text content associated with the session, and the video synthesis of at least one description image set and corresponding voice information by combining a preset video special effect, to obtain a sub-video corresponding to the session; Based on the arrangement order of at least one session in the to-be-converted text, the sub-videos corresponding to the at least one session are spliced to obtain a target video.
[0020] The electronic device provided by the embodiment of the present application includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of any one of the video generation methods described above.
[0021] In a possible implementation, when the computer program is executed by the processor, the processor executes the following process: Based on the to-be-converted text of the target literary work, the environment description text and the plot description text of at least one session are extracted; wherein the to-be-converted text is a text part of the target literary work that has not generated a video work; For each session, the following is performed: based on the environment description text and the plot description text of the session, at least one semantic description information is generated in a set format; wherein each semantic description information is used to describe at least one of the environment layout and the subject object in the session; For each generated semantic description information, the following is performed: based on the semantic description information, a corresponding description image set is generated by combining a content element set in the video work generated based on the target literary work; wherein each content element is a reference layout or a reference object in the video work; Based on the description image set corresponding to each of the at least one semantic description information, a target video is obtained.
[0022] The computer readable storage medium provided by the embodiment of the present application includes a computer program, and when the computer program runs on the electronic device, the computer program is used to make the electronic device execute the steps of any one of the video generation methods described above.
[0023] The computer program product provided by the embodiment of the present application includes a computer program, and the computer program is stored in a computer readable storage medium; when the processor of the electronic device reads the computer program from the computer readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of the video generation methods described above.
[0024] The beneficial effects of the present application are as follows: The embodiment of the present application provides a video generation method, device and equipment and a storage medium, relates to the technical field of data processing, in particular to the field of artificial intelligence. In the implementation manner of the video generation provided in the embodiment of the present application: Based on the text part (i.e. non-adapted content) of the target literary work for which no film and television work is generated, at least one of environment description text and plot description text of a scene is extracted, and further, semantic description information for describing at least one of environment layout and subject object in the scene is accurately generated according to a set format, and then a description image set is generated in combination with a content element set in the film and television work. As can be seen, the generation process of the image is no longer limited to the feature matching of the text paragraph and the video image that has been shot, but rather, the image with more pertinence and uniqueness can be generated according to the characteristics such as plot and environment of the non-adapted content, and the scene and plot in the non-adapted content can be more accurately presented in the form of an image, the generated image can better reflect the characteristics of the non-adapted content, accurately show the connotation of the non-adapted content, and avoid the problem that many contents do not match the paragraph in the feature matching process.
[0025] In summary, in the embodiment of the present application, the video image associated with the non-adapted content can be accurately generated, and then the corresponding target video can be generated. As can be seen, the target video is constructed according to the specific description of the non-adapted content in combination with the content element set in the film and television work, can better match the non-adapted content, improves the consistency of the video and the non-adapted content, increases the conflict and highlights of the video, ensures that the generated video has high conflict, thereby obtaining better visual feeling, improving the visual presentation effect, and improving the viewing experience.
[0026] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and achieved by the structure particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0027] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The illustrative embodiments of the present application and their description are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings: Figure 1 An application scenario schematic diagram provided by the embodiment of the present application; Figure 2 A video generation and playing scenario schematic diagram provided by the embodiment of the present application; Figure 3 A video generation method flowchart provided by the embodiment of the present application; Figure 4 A schematic diagram for extracting environment description text and plot description text provided by the embodiment of the present application; Figure 5 Another schematic diagram of extracting environment description text and plot description text provided by an embodiment of the present application; Figure 6 A method flowchart of generating at least one semantic description information provided by an embodiment of the present application; Figure 7 A schematic diagram of obtaining a reference object provided by an embodiment of the present application; Figure 8 A schematic diagram of obtaining a reference layout provided by an embodiment of the present application; Figure 9 A method flowchart of generating a description image set provided by an embodiment of the present application; Figure 10 A schematic diagram of generating a description image set provided by an embodiment of the present application; Figure 11 A structural schematic diagram of a text-to-image model provided by an embodiment of the present application; Figure 12 A schematic diagram of a denoising network provided by an embodiment of the present application; Figure 13 A specific structural schematic diagram of a text-to-image model provided by an embodiment of the present application; Figure 14 A schematic diagram of obtaining a text-image sample pair provided by an embodiment of the present application; Figure 15 A flowchart of a text-to-image model training method provided by an embodiment of the present application; Figure 16 Another method flowchart of generating a description image set provided by an embodiment of the present application; Figure 17 A schematic diagram of generating a description image set by a text-to-image model provided by an embodiment of the present application; Figure 18 A specific implementation schematic diagram of video generation provided by an embodiment of the present application; Figure 19 A structural diagram of a video generation apparatus provided by an embodiment of the present application; Figure 20 A structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0028] To make the purposes, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments described in the present application document, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the technical solutions of the present application.
[0029] Some concepts involved in the embodiments of the present application will be introduced below.
[0030] 1. Literary works: Literary works are artistic creations that express thoughts and emotions, exhibit aesthetic values and cultural connotations through fiction or non-fiction using language as a medium. They include various categories such as poetry, novels, essays, and reportage.
[0031] A literary work usually contains: theme, characters (such as characters in a story, embodying personality and fate), plot (such as the development of events), environment (such as the time, place, and background of the story), language style, narrative perspective, etc.
[0032] 2. Content elements: Content elements are the contents appearing in the video frames of a film or television work, which can be characters, animals, buildings, landscapes, etc.
[0033] 3. Text-to-image model: Text-to-image model, also known as text-to-image diffusion model, is a deep learning model used for text-to-image generation tasks. After being trained through the reverse process of natural image diffusion, a text-to-image model can gradually generate new natural images from a completely random noise image under the guidance of text. Noise images are generated when shooting or transmitting is disturbed by random signals, resulting in random changes in image information or pixel brightness.
[0034] 4. Variational AutoEncoder (VAE): Variational AutoEncoder is a probabilistic model based on variational inference, belonging to generative models. Its architecture design includes encoder and decoder.
[0035] Encoder is used to map the original high-dimensional data to a low-dimensional feature space, which is generally smaller than the original data dimension, to compress or reduce the dimension, and the low-dimensional feature is often an intermediate latent representation; Decoder is used to reconstruct the original data based on the compressed low-dimensional feature.
[0036] 5、Low-Rank Adaptation (LoRA): Low-rank adaptation is the low-rank adaptation of large language models. It freezes the weights of the pre-trained model and injects trainable rank decomposition matrices into the model architecture, greatly reducing the number of trainable parameters for downstream tasks. In the embodiments of the present application, lora mainly injects trainable network parameters into the denoising network in the text-to-image model, and the denoising network layer is used to associate the image with the description text, while the lora weight affects the network parameters corresponding to the denoising network layer, such as the weight matrix part of the denoising network layer.
[0037] 6、Text-to-Speech (TTS): Text-to-speech is an artificial intelligence technology that enables computers to "read" text and generate speech that sounds like a real person talking. Text-to-speech involves an encoder, a synthesizer, and a vocoder; among them, the encoder extracts the audio features of the input text, the synthesizer generates a speech spectrogram based on the audio features, and the vocoder converts the spectrogram into the final audio. Text-to-speech is mainly used in: electronic dictionaries, audio materials in the education field, voice prompts in car navigation systems in navigation systems, audio books, video dubbing in content creation, and voice interaction in games and virtual humans.
[0038] The word "exemplary" used in the following means "used as an example, embodiment or illustrative". Any embodiment described as "exemplary" does not necessarily mean that it is superior or better than other embodiments.
[0039] The terms "first", "second" in the text are only for descriptive purposes, and cannot be understood as explicitly or implicitly indicating relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more features; such as "first generation unit" and "second generation unit" in the embodiments of the present application are used to represent two different generation units, and the operations performed by the two generation units are different.
[0040] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more.
[0041] At present, adapting literary works into video works has become a popular cross-media integration method. However, in the process of adapting literary works into video works, due to reasons such as page limit, plot selection or censorship risk, selective script adaptation and shooting of literary works into video works are usually performed, that is, only part of the chapters in the literary works are shot, resulting in incomplete original plot of the literary works and affecting the viewing experience.
[0042] In order to improve the viewing experience, an implementation manner for generating a video for unadapted content in a literary work is given in the related art. In the process of generating the video, first, the already shot video work is subjected to frame extraction processing to obtain corresponding frame images, second, image feature extraction is performed on the frame images by a cross-modal model to obtain corresponding frame image features, and the frame image features are stored in an image feature library, third, for unadapted content in the literary work, a paragraph is extracted from the literary work, and text feature extraction is performed on the paragraph by the cross-modal model to obtain corresponding text features, fourth, the text features are matched with the frame image features stored in the image feature library, and the frame images corresponding to the frame image features that are successfully matched are taken as video images associated with the paragraph, and finally, the matched multiple video images are spliced to generate continuous playable images for the unadapted content, and a voiceover is added to form a video.
[0043] When the video images are obtained by feature matching, the features cannot be completely matched, at which time there will be a lot of content in the video images that does not match the paragraph, affecting the viewing experience. At the same time, the already shot video images are directly applied to the unshot content without adaptation or modification, resulting in poor consistency between the generated video and the unadapted content, and the video images in the generated video are all video images in the already shot video work, which will result in lack of conflict and lack of highlights in the generated video, further affecting the viewing experience.
[0044] Therefore, how to accurately generate video images associated with unadapted content and further ensure that the generated video has high conflict and improves the visual presentation effect is a technical problem to be solved at present.
[0045] Therefore, how to accurately generate video images associated with unadapted content and further ensure that the generated video has high conflict and improves the visual presentation effect is a technical problem to be solved at present.
[0046] The video generation method provided in the embodiments of the present application is different from the method in the related art of directly applying the obtained shot video image to the unadapted content without adaptation or modification. The video generation method provided in the embodiments of the present application extracts at least one kind of environment description text and plot description text based on the text part of the target literary work in which no video work is generated, and then accurately generates semantic description information for describing at least one of the environment layout in the scene and the subject object in the scene according to the set format based on the environment description text and the plot description text of the scene. Then, the corresponding description image set is generated based on the semantic description information and in combination with the content element set in the video work. As can be seen, the generation process of the image is no longer limited to the feature matching between the text paragraph and the shot video image, but can generate images with more pertinence and uniqueness according to the characteristics such as the plot and the environment of the unadapted content, more accurately present the scene and the plot in the unadapted content in the form of images, and generate images that can better reflect the characteristics of the unadapted content and accurately present the connotation of the unadapted content. The accuracy and efficiency of generating the video image are improved.
[0047] In summary, the video image associated with the unadapted content can be accurately generated, and then the corresponding target video is generated. In this way, the target video is constructed according to the specific description of the unadapted content in combination with the content element set in the video work, which can better match the unadapted content, improve the consistency of the video and the unadapted content, increase the conflict and highlights of the video, ensure that the generated video has high conflict, thereby obtaining a better visual experience, improving the visual presentation effect, and improving the viewing experience.
[0048] In the video generation process of the embodiments of the present application, various models are also proposed to respectively realize the automatic conversion of the text to be converted into a video dialogue and the automatic conversion of the text to be converted into an image, thereby reducing the time-consuming link in the entire video generation process, improving the video generation efficiency, and reducing the production cost of the target video.
[0049] The application scenarios provided in the present application will be briefly described below. It should be noted that the following scenarios are only used to illustrate the embodiments of the present application and are not limited. In specific implementation, the technical solutions provided in the embodiments of the present application can be flexibly applied according to actual needs.
[0050] Referring to Figure 1 , Figure 1 is a schematic diagram of an application scenario in the embodiments of the present application. The application scenario diagram includes a terminal device 110 and a server 120.
[0051] In an optional implementation, the terminal device 110 and the server 120 can communicate through a communication network. The communication network is a wired network or a wireless network.
[0052] Therefore, the terminal device 110 and the server 120 can be connected directly or indirectly through wired or wireless communication. For example, the terminal device 110 can be indirectly connected to the server 120 through a wireless access point, or the terminal device 110 can be directly connected to the server 120 through the Internet, which is not limited in the present application.
[0053] In the embodiments of the present application, the terminal device 110 includes but is not limited to a mobile phone, a tablet computer, a notebook computer, a desktop computer, an electronic book reader, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, and the like. The terminal device can be installed with a video playing client, which can be software, such as a browser with a video playing function, application software (such as social software, a browser, video software, etc.) with a video playing function, a webpage with a video playing function, an applet, and the like.
[0054] In the embodiments of the present application, the server 120 is a background server corresponding to the client, which is not limited in the present application. The server 120 can be a stand-alone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.
[0055] It should be noted that the video generation method in each embodiment of the present application can be executed by an electronic device, which can be the terminal device 110 or the server 120, that is, the method can be executed by the terminal device 110 or the server 120 alone, or by the terminal device 110 and the server 120 together.
[0056] Referring to Figure 2 , Figure 2 A video generation and playing scene schematic diagram is provided in the embodiments of the present application. In the video generation scene, the terminal device 110 transmits a text to be converted to the server 120, the server 120 generates a target video according to the text to be converted, and presents the generated target video through the terminal device 110. Specifically: The terminal device 110 sends a text to be converted, which needs to generate a target video, to the server 120; wherein the text to be converted is a text part of a target literary work which does not generate a video work; The server 120 extracts at least one of the environment description text and the plot description text based on the text to be converted, and generates semantic description information for describing at least one of the environment layout and the subject object in the game session according to a set format based on the environment description text and the plot description text of the game session, and generates a corresponding description image set in combination with a content element set in a video work associated with the target literary work, and obtains a target video based on the description image set. The server 120 receives the presentation instruction of the target video sent by the terminal device 110, and presents the target video in the terminal device 110.
[0057] The presentation instruction can be triggered in at least one of the following ways: Triggered by a web link during web browsing; Trigger the corresponding video playback by associating the paragraph content on the reading software; Triggered by recommendation information on the reading software, the recommendation information being a part of the target video, a part of the video work, etc. After the video work is played on the video software, trigger the watching entrance provided by the video work page, such as the watching entrance "AI sequel xxx".
[0058] It should be noted that Figure 1 The number of servers 120 is not limited in practice, and is not specifically limited in the embodiments of the present application. In the embodiments of the present application, when the number of servers 120 is multiple, the multiple servers 120 can be composed of a blockchain, and the server 120 is a node on the blockchain.
[0059] In addition, the embodiments of the present application can be applied to various scenes, including but not limited to cloud technology, artificial intelligence, comics, animation, radio drama, audio book, game, etc.
[0060] Taking the game scene as an example, in order to introduce the game in detail, so that the game player can better understand the game content, the game introduction video can be generated according to the game introduction text combined with the game characters in the created game.
[0061] It should be emphasized that in the specific embodiments of the present application, user-related data such as target literary works, video works, etc. are involved. When the above embodiments of the present application are applied to specific products or technologies, the permission or consent of the object is required, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0062] The video generation method provided by the exemplary embodiments of the present application is described below in combination with the application scenarios described above with reference to the accompanying drawings. It should be noted that the above-mentioned application scenarios are only shown for the purpose of facilitating the understanding of the spirit and principles of the present application, and the embodiments of the present application are not limited in this respect. In the following description of the embodiments of the present application, the video generation process is described in detail from the perspective of the device for generating the video.
[0063] Referring to Figure 3 , Figure 3 A flowchart of a video generation method provided by an embodiment of the present application is applied to an electronic device and includes the following steps: Step S300: Based on the to-be-converted text of the target literary work, environment description text and plot description text of at least one scene are extracted; wherein the to-be-converted text is a text part of the target literary work that has not generated a video work.
[0064] In the embodiments of the present application, in order to generate more accurate images, the way of extracting a paragraph from the to-be-converted text and generating a video image based on the extracted paragraph is no longer used, but the environment description text and the plot description text of the scene are extracted based on the to-be-converted text of the target text work. The environment description text contains environment information for determining the presentation effect and content of the picture, and the plot description text contains interaction information between characters, such as dialogues, actions, etc., so as to generate more accurate and vivid images based on the environment description text and the plot description text, and ensure the video conflict.
[0065] In the embodiments of the present application, based on the to-be-converted text of the target text work, at least one scene of environment description text and plot description text is extracted by setting rules; wherein the setting rules can be pre-constructed extraction conditions, or can also be a large language model. Next, the extraction of the environment description text and the plot description text is described in detail through specific ways.
[0066] Method one: environment description text and plot description text are extracted through pre-selected extraction conditions. For details, see steps A1-A2: Step A1: The to-be-converted text is split to obtain at least one scene information; wherein the scene information is part of the text content in the to-be-converted text.
[0067] In one possible implementation, when the to-be-converted text is split to obtain at least one scene information, first, the chapter title is identified from the to-be-converted text, and the to-be-converted text is split into multiple subtexts according to the chapter title, each subtext corresponding to a chapter title; then, for each subtext, the subtext is split into at least one scene information based on the set scene splitting condition; the scene splitting condition is at least one of a set number of words and a set dialogue discussion.
[0068] Step A2, for each game information, respectively: based on the pre-constructed visual presentation prompt information, extract the corresponding environment description text and plot description text from the game information.
[0069] The visual presentation prompt information includes but is not limited to: close-up, full-body, expression, action, background, and other visual presentation related information.
[0070] In the embodiments of the present application, visual presentation prompt information is set for environment description text and plot description text, and then when the corresponding environment description text and plot description text are extracted from the game information, the game information is subjected to word segmentation processing, the segmented sentences or words are classified and matched with the visual presentation prompt information of the environment description text, and the environment description text is constructed based on the matched sentences, and the segmented sentences or words are classified and matched with the visual presentation prompt information of the plot description text, and the plot description text is constructed based on the matched sentences.
[0071] It should be noted that the classification and matching method can be realized by a classification model; the constructed description text can be a simple sentence or a combination of segmented words, or a logical and clear description text generated based on the matched sentences or segmented words by a text generation model.
[0072] Referring to Figure 4 , Figure 4 An example of extracting environment description text and plot description text is provided in the embodiments of the present application; from Figure 4 It can be seen that: The to-be-converted text is: Chapter 1 Dispute Du Nv Yi felt that she would not live long.
[0073] Wei Mou laughed coldly: "He made such a big mistake, and he was only punished for three years of salary."
[0074] Chapter 2 Disagreement
[0075] Song Mou, the eldest son of Song XX, the Prince's mansion.
[0076] The mother-in-law asked anxiously, "What you said is true?"
[0077] Chapter 3 Bitterness
[0078] Du Nv Yi laughed and said, "Just a few days ago, Wang Fu asked someone to pass a message to me.
[0079] At this time, based on the to-be-converted text, at least one environment description text and plot description text of a game are extracted: First, according to the chapter title, the text to be converted is divided into multiple subtexts. For example, subtext 1: The woman in the Du family felt that she would not live long... Wei laughed: "He made such a big mistake, but was only punished for three years of salary"; subtext 2: Song, the eldest son of Song XX of the National Palace... The mother-in-law has already said "Is what you said true?" ; subtext 3: The woman in the Du family laughed, "Just the other day, the queen sent someone to tell me"... It was just an emotional outburst.
[0080] Then, according to the set number of words (such as 100 words), each chapter content (subtext) is divided into at least one session information. For example, session information 1 (100 words): The woman in the Du family felt that she would not live long... ; Session information 2 (100 words):... Wei laughed: "He made such a big mistake, but was only punished for three years of salary"; Session information 3 (100 words): Song, the eldest son of Song XX of the National Palace... ; Session information 4 (100 words):... The mother-in-law has already said "Is what you said true?" ; Session information 5 (100 words): The woman in the Du family laughed, "Just the other day, the queen sent someone to tell me"... ; Session information 6 (100 words):... It was just an emotional outburst.
[0081] Next, each session information is subjected to word segmentation processing, and the segmented sentences are matched with the visual presentation prompt information corresponding to the environment description text and the visual presentation prompt information corresponding to the plot description text, and the corresponding description text is constructed according to the respective matching results.
[0082] Method two: Extract environment description text and plot description text through large language model. Specifically, through the large language model, the following steps B1-step B2 are executed to realize the extraction of environment description text and plot description text.
[0083] Step B1, in combination with the pre-constructed dramatic event template, the text to be converted is subjected to splitting processing to obtain at least one session information; wherein the dramatic event template is used to describe: the interactive plot characteristics between at least two subject objects in the reference text.
[0084] Step B2, for each session information, respectively execute: based on the pre-constructed visual presentation prompt information, extract the corresponding environment description text and plot description text from the session information.
[0085] Reference Figure 5 , Figure 5 Another schematic diagram for extracting environment description text and plot description text provided by the embodiments of the present application can be known from Figure 5 . The text to be converted is: Chapter 1 Dispute Dong Nv Yi feel that they will not live long......
[0086] Wei coldly: "he made such a mistake, but is punished for three years salary."
[0087] Chapter II divergence
[0088] Song, the eldest son of the Duke of Song XX.....
[0089] Mother-in-law has been in a hurry to say "you said it is true?"
[0090] Chapter III bitter
[0091] Dong Nv Yi laughed, "just a few days ago, Wang Fuan entrusted me to pass on the words"..... but it is difficult to control emotions.
[0092] Dramatic event template is: Little sister was pushed by her mother and hit the wall, the porcelain plate broke the sound of explosion.
[0093] Mother: Li is your stepfather, you dare to steal his old-age money...
[0094] Little sister: I don't... Mom... I really don't...
[0095] She stumbled to the ground and clutched the corner of her mother's dress: don't be pitiful! It's been taken clearly!
[0096] Blood beads seep out of the palm, leather shoes into the entrance. Little sister saw the stepfather's gloomy eyes and suddenly rushed to pull his trouser legs.
[0097] Little sister: uncle, you speak...
[0098] Li laughed and spread her fingers: you sneaked into the study last night. Do you think I can't see it?
[0099] Little sister's pupils shrank. The mirror in the corridor reflected the distorted faces of the three: Mother / Li: thief... thief... She curled up in the broken porcelain and screamed: the money is for her brother's treatment!
[0100] At this time, the text to be converted, the dramatic event template, and the scene generation prompt information, visual presentation prompt information are input into the large language model, and the large language model can output the environment description text and the plot description text of multiple scenes. It should be noted that the large language model in the present application embodiment is based on the currently open source large language model, and the parameters are adjusted for the application scenario of the present application embodiment.
[0101] The number of episodes to be generated and episode information conditions are given in the episode generation prompt information; for example, the episode generation prompt information is: please refer to the dramatic event template to write 10 episodes of the same number of characters, and require 10 episodes to form 10 consecutive plots (each episode refers to the dramatic event template, and needs 5 to 10 rounds of role dialogue to explain a story detail), and each episode information is selected or transformed from the to-be-converted text and is mostly highly conflictual / dramatic / passionate content of the male and female leads.
[0102] In a possible implementation, the output environment description text and plot description text of multiple episodes can be presented in the form of a table. As shown in Table 1: Table 1
[0103] By chapter title and split conditions, episode information is accurately extracted, and visual presentation prompt information is combined to realize granular content recognition, more accurately extract environment description text and plot description text; by using a large language model and combining a pre-constructed dramatic event template, environment description text and plot description text can be quickly and accurately extracted, so that the generated environment description text and plot description text have more conflict; and based on the environment description text and plot description text, the scene restoration and immersion of text-to-picture and text-to-audio conversion can be improved, and the audio-visual expressiveness is increased.
[0104] In step S301, for each episode, the following is performed: based on the environment description text and the plot description text of the episode, at least one semantic description information is generated in a set format; wherein each semantic description information is used to describe at least one of the environment layout and the subject object in the episode.
[0105] Referring to Figure 6 , Figure 6 The method flowchart for generating at least one semantic description information provided by the embodiments of the present application is applied to an electronic device, and includes the following steps: In step S600, at least one initial description information is generated in a set format based on the environment description text; the set format includes external environment parameters and subject feature parameters.
[0106] The external environment parameters are used to describe the visual context of the environment layout; the external environment parameters include but are not limited to: time (day, night, dusk), weather (cloudy, sunny, rainy, yellow haze, white haze), light (bright, general, dim, dark), light and shadow, background description, and color tone (cool or warm).
[0107] The subject feature parameters are used to describe the visual performance of the subject object; the subject feature parameters include but are not limited to: person (name), clothing, expression, and action.
[0108] Step S601, for each initial description information, respectively perform: subject object matching of the plot description text and the environment description text to obtain the corresponding matching result.
[0109] When subject object matching of the plot description text and the environment description text is performed, the subject object of each speech is found from the plot description text, each subject object is found from the environment description text, and the found subject objects are compared and matched to determine whether the matching result represents a missing subject object.
[0110] Step S602, determine whether the matching result represents a missing subject object, if yes, perform step S603, otherwise perform step S604.
[0111] Step S603, when the matching result represents a missing subject object, supplement the initial description information based on the plot description text to generate corresponding semantic description information.
[0112] Step S604, when the matching result does not represent a missing subject object, take the initial description information as the semantic description information.
[0113] In one possible implementation, the semantic description information can also be generated by a large language model.
[0114] For example, by a large language model, first generate an initial description text based on the environment description text combined with a set format, and then generate a description generation prompt information and supplement the subject object description on the initial description text based on the plot description text. For example, the set format is: time (day, night, dusk), weather (cloudy, sunny, rainy, yellow haze, white fog), light (bright, general, dim, dark), character (name), clothing, expression, action, light and shadow, background description, color tone (warm or cool); the description generation prompt information is: according to the subject object dialogue in the plot description text, generate corresponding description information for each speaking subject object, and require to include: time (day, night, dusk), weather (cloudy, sunny, rainy, yellow haze, white fog), light (bright, general, dim, dark), character (name), clothing, expression, action, etc.
[0115] In the embodiments of the present application, the initial description information is first simply generated based on the environment description information, and then the subject object matching and environment-plot linkage are performed, and intelligent supplementation is performed when missing to generate accurate semantic description information, ensuring the integrity of the semantic description information, and further improving the context restoration degree and narrative coherence of the audio, and increasing the audio-visual expressiveness.
[0116] It should be noted that in the embodiments of the present application, the environment description text and the plot description text of the game are extracted, and the semantic description information is generated based on the environment description text and the plot description text. The same large language model can be used to achieve this. The environment description text and the plot description text are intermediate products of the large language model, and the semantic description information is the final output result of the large language model. When the large language model directly outputs the semantic description information, the input includes the text to be converted, the dramatic event template, and various generation prompt information related to generating the semantic description information.
[0117] In step S302, for each semantic description information generated, the following is performed: based on the semantic description information, the content element set in the generated video work is combined to generate a corresponding description image set; each content element is a reference layout or a reference object in the video work.
[0118] To ensure that the generated description image inherits part of the content of the video work and improves the correlation between the target video and the video work, thereby improving the viewing experience. In the embodiments of the present application, an implementation manner of generating a corresponding description image set based on semantic description information and combining a content element set in a video work is proposed.
[0119] Since the content element set needs to be combined when generating the description image set, the content element is first introduced in detail, and then the generation of the corresponding description image set based on the semantic description information and the combination of the content element set in the video work is introduced in detail.
[0120] In the embodiments of the present application, the content element set includes a reference object and a reference layout, and the determination manners of the reference object and the reference layout are different, which are introduced in detail below.
[0121] When the content element is a reference object, the content element is obtained by the following steps C1-C3: Step C1, frame extraction processing is performed on the video work to obtain at least one key video frame.
[0122] For example, independent shots are cut from the multi-episode video of the video work, each shot represents a sub-scene, and each sub-scene includes a plurality of video frames. For each sub-scene, a video frame with a set data is extracted as a key video frame.
[0123] Step C2, an object detection model is used to detect objects in each key video frame, and when a reference object is determined to be included, a cropping region corresponding to the reference object is determined.
[0124] For example, the reference object detected is taken as the center to expand the range, and the cropping region is determined so that the determined cropping region meets the target size.
[0125] Assuming the target video is a short video displayed in portrait mode, the target size must meet the requirements for portrait screen display, meaning the target size needs to fit the portrait screen size, such as 240x360. Furthermore, since a portrait screen image is being generated, the vertical portion of the image should be preserved as much as possible, and unnecessary horizontal or horizontal portions should be cropped.
[0126] Step C3: Based on the cropped area, obtain the reference object from the corresponding key video frame.
[0127] See Figure 7 , Figure 7 This is a schematic diagram illustrating the acquisition of a reference object according to an embodiment of this application; from Figure 7 From this, we can know that: The key video frame is input into the object detection model. After the object detection model detects that the key video frame contains a reference object, the location of the reference object is determined. The cropping region is determined with the reference object as the center. The reference object is extracted from the key video frame according to the cropping region and the extracted reference object is used as the content element.
[0128] When the content element is a reference layout, obtain the content element through steps D1-D4: Step D1: Perform frame extraction on the film or television work to obtain at least one key video frame.
[0129] Step D2: When it is determined that the key video frame does not contain a reference object, extract at least one reference layout from the key video frame according to the target size.
[0130] In one possible implementation, after detecting that the key video frame does not contain a reference object, a reference layout is extracted based on the center and target size of the key video frame. The reference layout is then extracted by moving the center left and right and using the target size. Alternatively, the reference layout can be extracted by moving the center up and down and using the target size. Content can also be extracted from the key video frame at intervals and spliced together according to the target size to generate a reference layout.
[0131] For example, when extracting a reference layout by moving it left and right from the center, the left and right directions are each moved by half the width of the target size; when extracting a reference layout by moving it up and down from the center, the up and down directions are each moved by half the height of the target size. Sometimes, if the up and down space is limited and cannot be moved, then extraction in the up and down direction from the center is abandoned.
[0132] Step D3: Perform a visual evaluation on at least one extracted reference layout to obtain a visual evaluation value for each reference layout.
[0133] For example, each reference layout is input into an open-source aesthetics model, visually evaluated using the open-source aesthetics model, and the visual evaluation value of the reference layout is output.
[0134] Step D4: Among the extracted reference layouts, select the reference layout whose visual evaluation value meets the visual conditions as the content element.
[0135] See Figure 8 , Figure 8 This is a schematic diagram illustrating an embodiment of obtaining a reference layout, from... Figure 8 From this, we can know that: The key video frame is input into the object detection model. After the object detection model detects that the key video frame does not contain a reference object, six reference layouts are extracted from the key video frame based on the target size. Among them, reference layout 1 is the image content in the central region that meets the target size, reference layout 2 is the image content in the left region of the center that meets the target size, reference layout 3 is the image content in the right region of the center line that meets the target size, reference layout 4 is the image content in the upper region of the center that meets the target size, reference layout 5 is the image content in the lower region of the center that meets the target size, and reference layout 6 is a combination of multiple image contents stitched together according to the target size. Multiple extracted reference layouts are input into the aesthetics model to obtain the visual evaluation value of each reference layout. Then, based on the visual evaluation value, the reference layout with the largest visual evaluation value is selected as the content element from at least one reference layout.
[0136] By intelligently extracting key video frames from multiple video frames in film and television works and performing object detection on the key videos, reference objects and reference layouts are accurately obtained. Combined with visual evaluation, those with high visual presentation quality are selected as content elements. In this process, it is ensured that the content elements meet the content display requirements, thereby improving the composition rationality and visual appeal of short videos.
[0137] In this embodiment of the application, when generating a corresponding descriptive image set based on semantic description information and the content element set in the film and television works, the content elements that match the semantic description information are first selected from the content element set, and then the descriptive image set is generated based on the selected content elements.
[0138] In the process of generating a descriptive image set based on selected content elements, it can be done by splicing content elements or by combining text-generated image models with semantic descriptive information. The different methods are explained in detail below.
[0139] Method 1: Generate a descriptive image set by splicing content elements.
[0140] See Figure 9 , Figure 9 A flowchart of a method for generating a descriptive image set, provided in an embodiment of this application and applied to an electronic device, includes the following steps: Step S900, the semantic description information is segmented and processed to obtain at least one keyword.
[0141] Exemplarily, when the semantic description information is segmented and processed, meaningless stop words are filtered first, and then keywords are extracted by using word frequency, TextRank, etc. Bottom logic is used to ensure that at least one keyword is returned. This way takes into account efficiency and accuracy, and is suitable for most keyword extraction scenarios of semantic description information.
[0142] Step S901, for each keyword, the corresponding content element is searched in the content element set.
[0143] Exemplarily, the text features of the keywords are extracted by the cross-modal model, and the image features of the content elements are extracted, and then the text features and the image features are matched, and the content element with the highest matching degree is selected as the content element corresponding to the keyword.
[0144] Step S902, at least one content element obtained by searching is spliced to obtain at least one initial image.
[0145] In a possible implementation, based on the semantic description information, at least one content element obtained by searching is spliced according to a target size to obtain at least one initial image conforming to the target size.
[0146] Step S903, at least one initial image is respectively visually evaluated to obtain a visual evaluation value corresponding to each of the at least one initial image.
[0147] Exemplarily, each initial image is input into an open-source aesthetic degree model, and the initial image is visually evaluated by the open-source aesthetic degree model to output a visual evaluation value of the initial image.
[0148] Step S904, based on at least one visual evaluation value, an initial image satisfying an evaluation condition is selected from at least one initial image.
[0149] In a possible implementation, an initial image with a visual evaluation value greater than an evaluation threshold value is selected, or an initial image with a visual evaluation value in the top K is selected.
[0150] Step S905, the initial image obtained by screening is composed into a description image set.
[0151] Referring to Figure 10 , Figure 10 An exemplary diagram for generating a description image set provided by an embodiment of the present application, from Figure 10 It can be known that: The semantic description information is "Zhang Mou is walking a dog", the semantic description information is segmented, the keywords "Zhang Mou" and "dog" are obtained, and the content elements matching "Zhang Mou" and the content elements matching "dog" are retrieved based on the keywords, and then the content elements containing "Zhang Mou" and the content elements containing "dog" are spliced to generate at least one initial image, and then the at least one initial image is input into an open-source aesthetic degree model to obtain a visual evaluation value of each initial image, and the initial images with the top K visual evaluation values are selected based on the visual evaluation values, and the selected initial images are combined to form a description image set.
[0152] By extracting keywords through segmentation and retrieving matching elements in the content element set, the automatic association of semantics and visual elements is realized. After splicing to generate initial images, a visual evaluation mechanism is introduced to quantitatively score the image quality, and high-quality images that meet the aesthetic or semantic expression requirements are selected, effectively avoiding the use of low-quality or semantically deviated images. The final description image set not only accurately reflects the context of the original text, but also has high visual expressiveness, realizing efficient and controllable generation from text to image, and improving the consistency and reliability of multi-modal content generation.
[0153] Method two: generating a description image set through a text-to-image model.
[0154] In the embodiments of the present application, in order to ensure the accuracy of the description image set generated by the text-to-image model, the text-to-image model needs to be trained. The training process is a process of repeatedly training the to-be-trained text-to-image model using training samples, which mainly includes a model design stage, a data preparation stage and an iterative training stage, which will be introduced below.
[0155] I. Model design stage.
[0156] Referring to Figure 11 , Figure 11 is a structural schematic diagram of a text-to-image model provided in the embodiments of the present application. The text-to-image model includes a noise adding network, a noise removing network, a text encoding network, an image encoding network and an image decoding network, wherein: The image encoding network is configured to perform image encoding on the obtained random image (which can also be referred to as a random seed) to obtain corresponding image features. In a possible implementation manner, the image encoding network can adopt, but is not limited to, a variational autoencoder (VAE), which maps the random image to a latent feature space to obtain corresponding image features.
[0157] The noise adding network is configured to diffuse noise to the image feature to obtain a corresponding image feature after noise adding; in a possible implementation manner, the noise adding network is configured to randomly add Gaussian features to the image feature, and the process can be a fixed Markov chain process, and the original data distribution is changed into a normal distribution by continuously adding Gaussian noise.
[0158] The text encoding network is configured to encode the obtained description text (which can be a set of keywords, referred to as multi-label information) to obtain corresponding text features; in a possible implementation manner, the text encoding network can adopt, but is not limited to, a contrastive text-image pre-training model (CLIP).
[0159] The denoising network is configured to denoise the obtained image feature after noise adding according to the obtained text features to obtain a denoised image feature; in a possible implementation manner, the denoising network converts Gaussian noise into content of a known data distribution through an iterative denoising process, such as using a neural network to restore data from a normal distribution to an original data distribution, so that the generated image has better diversity and realism.
[0160] In a possible implementation manner, the denoising network can adopt, but is not limited to, a U-Net network, and the U-Net network can adopt, but is not limited to, an attention mechanism.
[0161] Reference is made to Figure 12 , Figure 12 FIG. 1 is a schematic diagram of a denoising network in an embodiment of the present application; in an example, the denoising network is a U-Net network, and the U-Net network includes a plurality of cross-attention (QKV) modules (also referred to as cross-attention layers), and the cross-attention (QKV) modules included in the U-Net network are also named as denoising network layers according to the functions of the network, and are used to model the relationship between text and image and to denoise the image conditioned on the text.
[0162] In which, Figure 12 The structures of different layers of the U-Net network are listed, and due to the limitation of the length, Figure 12 only a part is shown. It can be known from Figure 12 that the denoising network layer is divided into three parts: an input part (IN) 121, a middle part (MID) 122 and an output part (OUT) 123. In addition, a text encoder (BASE) 124 is additionally provided.
[0163] In addition to the text encoder (BASE), the input part (IN), the middle part (MID) and the output part (OUT) can be understood as the denoising network layer in the text-to-image model.
[0164] As Figure 12 shown, wherein the input part 121 simply exemplifies 4 layers, respectively residual module 1211, attention module 1212, residual module 1213 and attention module 1214; the middle part 122 simply exemplifies 3 layers, respectively residual module 1221, attention module 1222, residual module 1223; the output part 123 simply exemplifies 4 layers, respectively residual module 1231, attention module 1232, residual module 1233 and attention module 1234. It should be noted that, Figure 12 The Unet structure listed is only a simple example.
[0165] Among them, the U-Net network can also contain a skip connection structure, and each downsampling can have a skip connection cascaded with the corresponding upsampling, so that the U-Net network fuses the features at the corresponding positions of the encoder in the channel at each upsampling, and improves the accuracy through the fusion of features of different sizes.
[0166] The image decoding network is used to decode the obtained denoised image features to obtain a predicted image corresponding to a group of keywords, that is, an output image.
[0167] Referring to Figure 13 , Figure 13 is a specific structure diagram of a text-to-image model provided by an embodiment of the present application; as Figure 13 shown: for a random image x, an image feature (Z) is obtained through an image encoding network (E), then the image feature (Z) is diffused and added with noise through a noise adding network to project into a latent space to obtain a latent space vector, that is, a noise image feature (Z T ), at the same time, a description text (Text) is obtained through a text encoding network (τ); then the text feature and the noise image feature (Z T ) are input into a denoising network, and the noise image feature (Z T ) is denoised for T times under the constraint of the text feature to finally generate a latent space prediction vector (Z'), that is, a predicted image feature; finally, the latent space prediction vector (Z') is decoded through an image decoding network (D) to output an image (Y ), the image (Y ) is a predicted image.
[0168] In the noise adding network, the image feature (Z) undergoes a T-time diffusion process to generate a noise image feature (Z T ), and Z TThis represents the latent space value at time T. Correspondingly, in the denoising network, the noisy image features (Z) are denoised through a denoising process. T Perform T denoising predictions to obtain the predicted image features (Z'). Taking the first denoising process as an example: the text features are used as the KV in the QKV module, and the noisy image features (Z') are used as the predicted features. T As Q in the QKV module, text features are used to constrain noisy image features (Z). T The denoising process of the QKV module enables it to output predicted image features (Z') that are related to the input descriptive text after T denoising operations.
[0169] It should be noted that, Figure 13 Only one possible hierarchical relationship is shown. In actual applications, the number of QKV modules and their connection relationships can be designed based on the actual situation.
[0170] In this embodiment, the features of a random image encoded by VAE are mapped to the latent space vector at time T through a diffusion network. Subsequently, a noise representation (i.e., image prediction noise) is learned and fitted through a denoising network, thereby subtracting the image prediction noise to obtain the image representation that is actually needed. Then, the image is obtained by a decoder.
[0171] II. Data Preparation Stage.
[0172] Data collection is of paramount importance in machine learning, arguably the most crucial step. The receipt preparation stage in this application embodiment mainly includes the preparation process of image and text sample pairs.
[0173] In this embodiment of the application, the image-text sample pair includes: a sample image extracted from a film or television work and a descriptive text for the sample image, wherein the descriptive text is generated according to a set format.
[0174] See Figure 14 , Figure 14 This is a schematic diagram illustrating the acquisition of image-text sample pairs provided in an embodiment of this application. Figure 14 From this, we can know that: Frame extraction is performed on the film / video work to obtain at least one key video frame. An object detection model is used to perform object detection on each key video frame. If a reference object is determined to be present, a reference object of the target size is extracted from the key video frame based on the vertical screen orientation, centered on the reference object. If no reference object is determined to be present, a reference layout of the target size is extracted from the key video frame based on the vertical screen orientation. The reference object or reference layout is then used as an image in an image-text sample pair. It should be noted that the method for extracting reference objects and reference layouts from key video frames is described above. Figure 7 and Figure 8 As shown, it will not be repeated here.
[0175] Using a large language model, generate a description text for the reference layout, which needs to include the following content: time (day, night, dusk), weather (cloudy, sunny, rainy, yellow haze, white haze), light (bright, general, dim, dark), color tone (warm, cool), scene description; generate a description text for the reference object, which needs to include the following content: time (day, night, dusk), weather (cloudy, sunny, rainy, yellow haze, white haze), light (bright, general, dim, dark), character (name), clothing, expression, action, light and shadow, background description, color tone (warm and cool); the description text generated by the large language model is used as the description text in the picture-text sample pair, and the picture-text sample pair is constructed in this way.
[0176] For example, for Figure 7 The reference object is obtained from the key video frame, and the generated description text is: daytime, sunny, sufficient light, the woman wears a light-colored traditional costume (light beige / light blue long gown), expression calm and thoughtful, side head gazes at the right front of the picture, one side of the face is obviously light, and there is a natural shadow gradient in the folds of the clothes. The background is blurred, and it is an indoor environment with wooden structure carvings and other furnishings, warm brown background, and the character forms a low-contrast soft color tone.
[0177] For Figure 8 The reference layout is obtained from the key video frame, and the generated description text is: daytime, sunny, bright sunlight penetrating the leaves forms a light spot, warm color tone (sunlight + red ribbon main color) and cool color tone (green leaf background) form a contrast, a tree is hung with red ribbons, red ribbons are densely distributed on the branches and dance with the wind, dense green leaves, sunlight penetrates the gaps between leaves and falls golden light spots.
[0178] After obtaining the picture-text sample pair, the picture-text model to be trained is iteratively trained based on the picture-text sample pair. For details, see "III. Iterative training phase".
[0179] III. Iterative training phase.
[0180] In the embodiment of the present application, the text-to-image model to be trained is trained by cyclic iteration based on the image-text sample pairs in the training set, and a target text-to-image model is obtained. In the model training process, the full set of image-text sample pairs (i.e., the image-text sample pair training set) is cyclically iterated for multiple rounds (e.g., 10 rounds). In each round of iteration, due to the limited memory resources of the training machine, the full set of image-text sample pairs cannot be input into the model at one time for training, so all image-text sample pairs are trained in batches (batch), and the model parameters are updated once for each batch. For example, each batch of samples is generated by random division, and each batch of samples is input into the model for forward calculation, backward calculation, model parameter update, and other training.
[0181] Before the first round of training, the parameters of the text-to-image model to be trained need to be initialized. For example, the model parameters of the pre-trained model are used for the image encoding network, the text encoding network, the denoising network, and the like, that is, the open source trained model parameters are used; and the LoRA weights of the injection network are randomly initialized and updated during each round of training. Further, the batch, the number of iterations (epoch), and the learning rate (learning rate) and other hyperparameters are set. After setting, training is started to obtain the target text-to-image model. The learning rate is initialized to 0.0004, and after every 5 rounds of learning, the learning rate becomes 0.1 times the original.
[0182] In a possible implementation, since the operations performed in each cyclic iteration are consistent, the training of the text-to-image model to be trained is described by taking one cyclic iteration as an example.
[0183] Referring to Figure 15 , Figure 15 A flowchart of a text-to-image model training method provided by the embodiment of the present application is applied to an electronic device, which includes the following steps: Step S1500, selecting an image-text sample pair from the image-text sample pair training set; wherein the image-text sample pair includes a sample image and a description text of the sample image, and the description text is generated in a set format.
[0184] For example, a batch of image-text sample pairs is extracted, and the extracted image-text sample pairs are input into the text-to-image model to be trained, so as to train the text-to-image model to be trained based on the extracted image-text sample pairs.
[0185] Step S1501, inputting the image-text sample pair into the text-to-image model to be trained to obtain a predicted image output by the text-to-image model to be trained for the description text in the image-text sample pair.
[0186] Exemplarily, each sample image is image encoded by an image encoding network to obtain an image feature, and then the image feature is diffused and added with noise by a noise adding network to project into a latent space to obtain a latent space vector, i.e., a noisy image feature; each description text is text encoded by a text encoding network to obtain a text feature; the text feature and the noisy image feature are input into a denoising network, and the noisy image feature is denoised for T times by the denoising network under the constraint of the text feature to generate a predicted image feature; finally, the predicted image feature is decoded by an image decoding network to obtain a predicted image.
[0187] In step S1502, a loss function is constructed according to the image label in the image-text sample pair and the predicted image.
[0188] Exemplarily, the loss function adopts a mean squared error (MSE), i.e., is determined by the mean squared error of the predicted image and the sample image. The loss function is: ; wherein, is each pixel value in the sample image, is each pixel value in the predicted image.
[0189] In step S1503, the model parameters of the to-be-trained text-to-image model are adjusted according to the loss function to obtain a trained target text-to-image model.
[0190] Exemplarily, the loss is calculated according to the loss function, and the loss is the total loss of the current batch of image-text sample pairs.
[0191] In a possible implementation, when the model parameters of the to-be-trained text-to-image model are adjusted according to the loss function to obtain a trained target text-to-image model, the text encoding network in the to-be-trained text-to-image model and the LoRA module of the CrossAttention linear layer in the CrossAttention module of the denoising network are adjusted according to the loss function to obtain the Text Embeddings and the LoRA weight of the text, and the trained target text-to-image model is obtained according to the Text Embeddings and the LoRA weight.
[0192] In a possible implementation, the stochastic gradient descent (SGD) is used for parameter adjustment, i.e., the gradient of the model parameters is obtained by returning the loss direction to the model and the parameters are updated.
[0193] It should be noted that after the training of a batch is completed, an iteration process is ended; in each iteration training process, in addition to LoRA training, full parameter training can also be performed. The model used to generate images in the embodiments of the present application can also be fine-tuned using a DiT generation model, such as a mixed meta text-to-image open source large model, or an SDXL model.
[0194] In the embodiments of the present application, before the model parameter adjustment is performed, it can also be judged whether the model convergence condition is met. Exemplarily, the model convergence condition can include at least one of the following conditions: the model loss is not greater than a preset loss value threshold; the number of iterations reaches a preset upper limit value.
[0195] After the target text-to-image model is trained, based on the target text-to-image model, the content elements retrieved in combination with the semantic description information are used to generate a description image set.
[0196] Referring to Figure 16 , Figure 16 Another method flowchart for generating a description image set provided by the embodiments of the present application is applied to an electronic device, and includes the following steps: Step S1600, performing word segmentation processing on the semantic description information to obtain at least one keyword.
[0197] Step S1601, for each keyword, retrieving corresponding content elements in the content element set.
[0198] Step S1602, inputting the retrieved content elements and the semantic description information into the target text-to-image model, and performing the following steps through the target text-to-image model: Step S16021, based on the text features extracted from the semantic description information, performing denoising processing on at least one noise image feature associated with the retrieved content elements respectively to obtain corresponding denoised image features.
[0199] Step S16022, for each denoised image feature, respectively performing description image prediction to generate a corresponding description image.
[0200] Step S1603, based on the generated description images, composing a description image set.
[0201] Referring to Figure 17 , Figure 17 A schematic diagram of generating a description image set by a text-to-image model provided by the embodiments of the present application is shown in Figure 17 It can be known from the above that: The semantic description information is "Zhang Mou is walking a dog", the keyword "Zhang Mou" and "dog" are obtained by performing word segmentation processing on the semantic description information, and the content elements matching "Zhang Mou" and the content elements matching "dog" are retrieved based on the keywords in the content elements. Then, the content elements containing "Zhang Mou", the content elements containing "dog", and the semantic description information "Zhang Mou is walking a dog" are input into the target text-to-image model, and the target text-to-image model outputs a set of description images.
[0202] The content elements are retrieved by word segmentation, and are input into the target text-to-image model trained in cooperation with the semantic description information. The text features are used to guide the denoising process of the noise image features, so that the generated description images are more accurate and consistent with the original text context. In this process, the retrieved content elements are introduced, which can improve the accuracy and consistency of the key visual elements. The denoising process combines semantic context to enhance the detail restoration capability, so that the finally constructed description image set has high quality and high semantic fidelity.
[0203] In step S303, a target video is obtained based on the description image set corresponding to each of the at least one semantic description information.
[0204] In one possible implementation, when the target video is obtained based on the description image set corresponding to each of the at least one semantic description information, a silent video is generated based on the description image set corresponding to each game, and the silent video is dubbed and subtitled based on the plot description information associated with the game to generate a corresponding sub-video. Then, the multiple sub-videos are spliced to obtain the target video.
[0205] Next, steps E1-E3 are used to describe in detail the process of obtaining the target video based on the description image set corresponding to each of the at least one semantic description information.
[0206] In step E1, for each semantic description information, the following steps are performed: extracting subtitle information of the corresponding description image set from the plot description information corresponding to the semantic description information, and converting the subtitle information into voice information.
[0207] The subtitle information is the speaking content of each subject object. For example, the plot description information includes: Du Nv Yi (steps on the pearl necklace): "What did Wang Mou take out in the past… was it Liaodong aconite powder?" At this time, "What did Wang Mou take out in the past… was it Liaodong aconite powder?" is the subtitle content.
[0208] In one possible implementation, a speech synthesis tool is used to convert the subtitle content into voice information. The speech synthesis tool can be a text-to-audio model.
[0209] In the embodiments of the present application, the text-to-audio generation model can be set for different subject objects, and each text-to-audio generation model is obtained by training according to the audio data of a subject object in a film and television work, that is, the subject object and the text-to-audio generation model correspond one by one. At this time, when there is speaking content corresponding to a certain subject object, the text-to-audio generation model of the subject object is used to generate voice information.
[0210] Step E2, for each game, respectively: based on at least one semantic description information associated with the game, the arrangement order in the part of the text content associated with the game, combining the preset video special effect, video synthesis is carried out on at least one description image set and corresponding voice information, and the sub-video corresponding to the game is obtained.
[0211] Exemplarily, for at least one description image set and corresponding voice information of each game, the display duration of the corresponding picture display is determined, and the display duration is the voice duration of the dialogue content contained in the plot description information. Then, according to the description image set, the display duration, and the corresponding voice information, a tool (such as opencv) is used to synthesize a video, so as to obtain a sub-video corresponding to a game.
[0212] Step E3, based on the arrangement order of at least one game in the to-be-converted text, the sub-videos corresponding to the at least one game are spliced to obtain a target video.
[0213] Exemplarily, a tool (such as opencv) is used to concatenate the sub-videos of all games into a target video, and picture conversion modes such as dissolve conversion, rotation, fade-in and fade-out can be added in the concatenation process.
[0214] The cooperation of text and voice is realized by semantic association, which ensures that the subtitles and voice content are consistent. The sub-videos are synthesized by combining the text order and video special effects, which enhances the narrative coherence and visual expressiveness. Finally, the complete video is spliced according to the original text structure, which ensures the unity of the overall logic and rhythm, and realizes the automatic and high-quality conversion from literary text to multi-modal audio-visual content.
[0215] In the embodiments of the present application, the target video can be a micro-short drama. The micro-short drama is a new network literary style. The micro-short drama refers to a drama with a single set duration of tens of seconds to about 15 minutes, a relatively clear theme and main line, and a relatively continuous and complete plot. It has the characteristics of short duration and compact plot. Most of such dramas are adapted from network novels and are distributed on new media short video platforms such as Douyin and Kuaishou. A considerable part of them is directly made into vertical screen for easy mobile viewing. It has high conflict plot and produces continuous climax experience through fast-paced picture information / dialogue.
[0216] Referring to Figure 18 ,Figure 18 A video generation specific implementation schematic diagram is provided for the embodiments of the present application, from Figure 18 It can be known from The text to be converted in the target literary work which has not generated a film and television work is obtained, at least one semantic description information corresponding to each scene is generated by a large language model, and then a description image set is generated by a target text generation model based on the semantic description information and the content element set.
[0217] For each description image set, the corresponding subtitle information is extracted from the plot description information corresponding to the semantic description information associated with the description image set, and the subtitle information is converted into voice information, so as to obtain the subtitle information and voice information associated with each description image set; the associated description image set is added with subtitles based on the subtitle information, and the associated description image set is configured with audio based on the voice information and in combination with the subtitle information.
[0218] For each scene, at least one description image set and corresponding voice information are video synthesized based on the arrangement order of the at least one semantic description information associated with the scene in the part of text content associated with the scene, in combination with a preset video special effect, to obtain a sub-video corresponding to the scene; and then the sub-videos corresponding to at least one scene are spliced based on the arrangement order of the at least one scene in the text to be converted, to obtain a target video.
[0219] It should be noted that since the target video is generated by splicing multiple sub-videos, and each sub-video is generated in the same manner, and the implementation manner of generating part of the sub-videos based on each semantic description information is the same, Figure 18 An example of generating a sub-video corresponding to a description image set associated with one semantic description information is given in
[0220] In the embodiments of the present application, based on the text part (i.e., the unadapted content) in the target literary work which has not generated a film and television work, the environment description text and the plot description text of at least one scene are extracted, and further according to a set format, the semantic description information for describing at least one of the environment layout and the subject object in the scene is accurately generated, and then the description image set is generated in combination with the content element set in the film and television work; it can be seen that the generation process of the image is no longer limited to the feature matching of the text paragraph and the video image that has been shot, but rather, according to the characteristics such as the plot and the environment of the unadapted content, images more targeted and unique can be generated, and the scenes, plots, etc. in the unadapted content can be more accurately presented in the form of images, the generated images can better reflect the characteristics of the unadapted content, accurately show the connotation of the unadapted content, and avoid the problem that many contents do not match the paragraphs in the feature matching process.
[0221] In the embodiments of the present application, the video image associated with the unadapted content can be accurately generated, and then the corresponding target video is generated. It can be seen that the target video is constructed according to the specific description of the unadapted content in combination with the content element set in the film and television work, which can better match the unadapted content, improve the consistency of the video and the unadapted content, increase the video conflict and highlights, ensure that the generated video has high conflict, thereby obtaining better visual experience, improving the visual presentation effect, and improving the viewing experience.
[0222] Based on the same inventive concept, the embodiments of the present application also provide a video generation device. Referring to Figure 19 , Figure 19 FIG. 1 is a structural schematic diagram of a video generation device 1900, which comprises: An extraction unit 1901 is configured to extract at least one environmental description text and plot description text of each scene based on the to-be-converted text of the target literary work; wherein the to-be-converted text is a text part of the target literary work that has not generated a film and television work; A first generation unit 1902 is configured to, for each scene, respectively perform: generating at least one semantic description information according to a set format based on the environmental description text and the plot description text of the scene; wherein each semantic description information is used to describe at least one of the environmental layout and the subject object in the scene; A second generation unit 1903 is configured to, for each semantic description information, respectively perform: generating a corresponding description image set based on the semantic description information in combination with a content element set in a film and television work that has been generated based on the target literary work; wherein each content element is a reference layout or a reference object in the film and television work; An obtaining unit 1904 is configured to obtain a target video based on the description image set corresponding to each semantic description information.
[0223] In a possible implementation manner, the extraction unit 1901 is specifically configured to: perform splitting processing on the to-be-converted text to obtain at least one scene information; wherein the scene information is part of the text content in the to-be-converted text; for each scene information, respectively perform: extracting the corresponding environmental description text and plot description text from the scene information based on the pre-constructed visual presentation prompt information.
[0224] In a possible implementation manner, the extraction unit 1901 is specifically configured to: identify chapter titles from the to-be-converted text, and split the to-be-converted text into a plurality of subtexts according to the chapter titles; wherein each subtext corresponds to a chapter title; for each subtext, respectively perform: splitting the subtext into at least one scene information based on the set scene splitting condition.
[0225] In a possible implementation, the extraction unit 1901 is specifically configured to: split the to-be-converted text according to a pre-constructed dramatic event template to obtain at least one performance information; The dramatic event template is used to describe an interactive plot feature between at least two subject objects in the reference text.
[0226] In a possible implementation, the first generation unit 1902 is specifically configured to: generate at least one initial description information according to a set format based on the environment description text; the set format includes an external environment parameter and a subject feature parameter; the external environment parameter is used to describe a visual context of an environment layout, and the subject feature parameter is used to describe a visual performance of a subject object; For each initial description information, the following is performed: subject object matching of the plot description text and the environment description text is performed to obtain a corresponding matching result, and corresponding semantic description information is generated based on the matching result and the initial description information.
[0227] In a possible implementation, the first generation unit 1902 is specifically configured to: When the matching result indicates that a subject object is missing, the initial description information is supplemented based on the plot description text to generate corresponding semantic description information; When the matching result indicates that a subject object is not missing, the initial description information is taken as the semantic description information.
[0228] In a possible implementation, the set of content elements is obtained in the following manner: frame extraction is performed on the video work to obtain at least one key video frame; For each key video frame, the following is performed: when a reference object is identified from the key video frame, a content element of a target size is extracted from the key video frame with the reference object as the center, or when a reference layout is identified from the key video frame, a content element of a target size is extracted from the key video frame; The target size meets a vertical screen display condition.
[0229] In a possible implementation, the second generation unit 1903 is specifically configured to: perform word segmentation processing on the semantic description information to obtain at least one keyword; For each keyword, a corresponding content element is searched in the set of content elements; At least one initial image is obtained by splicing the at least one searched content element; A corresponding set of description images is generated based on the at least one initial image.
[0230] In a possible implementation, the second generation unit 1903 is specifically configured to: respectively perform visual evaluation on the at least one initial image to obtain a respective visual evaluation value corresponding to each of the at least one initial image; based on the at least one visual evaluation value, screen the initial image that meets the evaluation condition from the at least one initial image; compose the initial image obtained by screening into the description image set.
[0231] In a possible implementation, the second generation unit 1903 is specifically configured to: perform word segmentation processing on the semantic description information to obtain at least one keyword; for each keyword, search for a corresponding content element in the content element set; based on the text feature obtained by extracting the semantic description information, respectively perform denoising processing on at least one noise image feature associated with the content element searched to obtain a corresponding denoised image feature; for each denoised image feature, respectively perform description image prediction to generate a corresponding description image; based on the at least one generated description image, generate a corresponding description image set.
[0232] In a possible implementation, the steps of based on the text feature obtained by extracting the semantic description information, respectively performing denoising processing on at least one noise image feature associated with the content element searched to obtain a corresponding denoised image feature, and for each denoised image feature, respectively performing description image prediction to generate a corresponding description image are performed by a target text-to-image model; wherein, the target text-to-image model is obtained by performing cyclic iteration training on a training set based on a text-image sample pair; the text-image sample pair includes a sample image and a description text of the sample image extracted from a film and television work, and the description text is generated in a set format.
[0233] In a possible implementation, the obtaining unit 1904 is specifically configured to: for each semantic description information, respectively perform: extracting subtitle information of the corresponding description image set from the plot description information corresponding to the semantic description information, and converting the subtitle information into voice information; for each session, respectively perform: based on at least one semantic description information associated with the session, arranging the order of the part of the text content associated with the session, combining a preset video special effect, and performing video synthesis on at least one description image set and corresponding voice information to obtain a sub-video corresponding to the session; The target video is obtained by splicing the sub-videos corresponding to the at least one scene based on the arrangement order of the at least one scene in the text to be converted.
[0234] In the embodiment of the application, based on the text part (i.e. the non-adapted content) of the target literary work for which no film and television work has been generated, the environment description text and the plot description text of at least one scene are extracted, and further, semantic description information for describing at least one of the environment layout and the subject object in the scene is accurately generated according to a set format, and then the description image set is generated in combination with the content element set in the film and television work. As can be seen, the generation process of the image is no longer limited to the feature matching between the text paragraph and the video image that has been shot, but rather, the image with more pertinence and uniqueness can be generated according to the characteristics such as the plot and the environment of the non-adapted content, and the scene and the plot in the non-adapted content can be more accurately presented in the form of an image, the generated image can better reflect the characteristics of the non-adapted content, accurately show the connotation of the non-adapted content, and avoid the problem that many contents do not match the paragraph in the feature matching process.
[0235] In the embodiment of the application, the video image associated with the non-adapted content can be accurately generated, and then the corresponding target video can be generated. As can be seen, the target video is constructed according to the specific description of the non-adapted content in combination with the content element set in the film and television work, can better match the non-adapted content, improves the consistency between the video and the non-adapted content, increases the conflict and the highlight of the video, ensures that the generated video has high conflict, and thus better visual experience is obtained, the visual presentation effect is improved, and the viewing experience is improved.
[0236] For the convenience of description, each part is described as a module (or unit) according to the function. Of course, the functions of the modules (or units) can be implemented in the same or multiple software or hardware in the implementation of the application.
[0237] In the embodiment of the application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.
[0238] After introducing the video generation method and device of the example embodiment of the application, next, an electronic device according to another example embodiment of the application is introduced.
[0239] Those skilled in the art can understand that each aspect of the present application can be implemented as a system, a method or a computer program product. Therefore, each aspect of the present application can be embodied in a form of entirely hardware, entirely software (including firmware, microcode, etc.), or a combination of hardware and software, which can be collectively referred to as "circuitry", "module" or "system".
[0240] Based on the same inventive concept as the method embodiments described above, the electronic device is also provided in the embodiments of the present application. In an embodiment, the electronic device is taken as an example of a server, and the structure of the electronic device can be as shown in Figure 20 The electronic device includes a memory 2001, a communication module 2003, and one or more processors 2002.
[0241] The memory 2001 is used to store computer programs executed by the processor 2002. The memory 2001 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and programs required for running instant messaging functions, etc.; and the data storage area can store various instant messaging information and operation instruction sets, etc.
[0242] The memory 2001 can be a volatile memory such as a random-access memory (RAM), and the memory 2001 can also be a non-volatile memory such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), or any other medium capable of carrying or storing desired computer programs in the form of instructions or data structures and capable of being accessed by a computer, but is not limited thereto. The memory 2001 can be a combination of the above memories.
[0243] The processor 2002 can include one or more central processing units (CPUs) or digital processing units, etc. The processor 2002 is used to realize the above-mentioned video generation method when calling the computer programs stored in the memory 2001.
[0244] The communication module 2003 is used to communicate with terminal devices and other servers.
[0245] The specific connection medium between the above-mentioned memory 2001, communication module 2003 and processor 2002 is not limited in the embodiments of the present application. The embodiments of the present application can be implemented in various forms, such as a bus, a point-to-point connection, a shared bus, a multi-drop bus, a crossbar switch, etc. Figure 20The bus 2004 interconnects the memory 2001 and the processor 2002, and the bus 2004 is in Figure 20 The connection between the components is described by thick lines, and the connection between other components is only schematically described and is not limited. The bus 2004 can be divided into an address bus, a data bus, a control bus, and the like. For the convenience of description, Figure 20 Only one thick line is described, but only one bus or only one type of bus is not described.
[0246] In some possible implementation, the memory 2001 stores a computer storage medium, and the computer storage medium stores a computer program. The computer program is used to implement the steps of the video generation method of the embodiments of the present application. The processor 2002 is used to execute the above-mentioned video generation method.
[0247] In some possible implementation, each aspect of the video generation method provided by the present application can also be implemented in the form of a computer program product, which includes a computer program. When the computer program product is run on an electronic device, the computer program is used to make the electronic device execute the steps of the video generation method according to the various exemplary embodiments of the present application described in the specification.
[0248] The computer program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0249] The computer program product of the embodiments of the present application can adopt a portable compact disk read-only memory (CD-ROM) and include a computer program, and can be run on an electronic device. However, the computer program product of the present application is not limited to this. In this document, the readable storage medium can be any tangible medium containing or storing a program, which can be used or combined with a command execution system, device or apparatus.
[0250] A readable signal medium can be any medium that can be read by a machine. The machine or its controller can include one or more processors for processing associated with the subject matter. A readable signal medium can include one or more machine-readable storage media. A machine- readable storage medium can include any medium that stores digital information including, without limitation, magnetic, optical, and solid-state storage mediums. A readable signal medium can include any medium that can transmit information that can include procedures, rules, protocols, or codes. Many aspects are disclosed with reference to various methods and devices. It is understood that where such methods and devices are disclosed in the present disclosure, additional methods and devices pertaining to such generally described methods and devices can likewise be used. Such additional methods and devices are believed to be within the scope of the disclosure and are contemplated at being within the scope of equivalents and claims thereof.
[0251] A computer program included on a readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, and the like, or any suitable combination of the foregoing.
[0252] Computer programs for implementing operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, and the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer programs can execute entirely on the user's electronic device, partly on the user's electronic device, as a stand-alone software package, partly on the user's electronic device and partly on a remote electronic device or entirely on the remote electronic device or server. In the latter scenario, the remote electronic device can be connected to the user's electronic device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external electronic device (for example, through the Internet using an Internet Service Provider).
[0253] It should be noted that although several units or sub-units of the apparatus are mentioned in the above detailed description, such division is merely exemplary and not mandatory. Indeed, features and functions of two or more units described above can be embodied in one unit, according to the embodiments of the present application. Conversely, features and functions of one unit described above can be further divided into units embodied by several units.
[0254] Although preferred embodiments of the application have been described herein, changes and modifications can be suggested to one skilled in the art, and it is intended that the scope of the application be limited only by the scope of the appended claims, including the equivalents thereof.
[0255] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the application, the application can be practiced otherwise than as specifically described herein.
Claims
1. A video generation method, characterized in that, The method includes: Based on the text to be converted from the target literary work, extract at least one scene's environmental description text and plot description text; wherein, the text to be converted is the text portion of the target literary work that has not been used to generate a film or television work; For each of the aforementioned scenes, the following steps are performed: based on the environment description text and plot description text of the scene, at least one semantic description information is generated according to a set format; wherein, each of the semantic description information is used to describe at least one of the environment layout and the main object in the scene; For each type of semantic description information generated, the following steps are performed: Based on the semantic description information, and combined with the set of content elements in the film and television works already generated based on the target literary work, a corresponding set of descriptive images is generated; wherein each content element is a reference layout or reference object in the film and television works; The target video is obtained based on the description image set corresponding to each of the at least one semantic description information.
2. The method as described in claim 1, characterized in that, The text to be converted based on the target literary work extracts at least one scene's environmental description text and plot description text, including: The text to be converted is split to obtain at least one session information; wherein, the session information is a portion of the text content in the text to be converted; For each scene information, the following steps are performed: Based on pre-constructed visual presentation prompts, extract the corresponding environmental description text and plot description text from the scene information.
3. The method as described in claim 2, characterized in that, The step of splitting the text to be converted to obtain at least one session information includes: Identify chapter titles from the text to be converted, and divide the text into multiple sub-texts based on the chapter titles; wherein each sub-text corresponds to a chapter title; For each subtext, the following steps are performed: based on the set session splitting conditions, the subtext is split into at least one session information.
4. The method as described in claim 2, characterized in that, The step of splitting the text to be converted to obtain at least one session information includes: By combining a pre-built dramatic event template, the text to be converted is split to obtain at least one scene information; The dramatic event template is used to describe: the interactive plot features between at least two main objects in the reference text.
5. The method as described in claim 1, characterized in that, The environmental description text and plot description text based on the scene are used to generate at least one semantic description information according to a set format, including: Based on the environmental description text, at least one initial description information is generated according to a set format; the set format includes: external environment parameters and subject feature parameters; the external environment parameters are used to describe: the visual context of the environmental layout, and the subject feature parameters are used to describe: the visual representation of the subject object; For each initial description, the following steps are performed: subject object matching is performed between the plot description text and the environment description text to obtain the corresponding matching results, and corresponding semantic description information is generated based on the matching results and the initial description information.
6. The method as described in claim 5, characterized in that, The step of generating corresponding semantic description information based on the matching result and the initial description information includes: When the matching result represents a missing subject object, the initial description information is supplemented based on the plot description text to generate corresponding semantic description information; When the matching result indicates that the subject object is not missing, the initial description information is used as the semantic description information.
7. The method as described in claim 1, characterized in that, The set of content elements is obtained in the following way: Frame extraction is performed on the film and television work to obtain at least one key video frame; For each key video frame, the following actions are performed: when a reference object is identified from the key video frame, extract content elements of the target size from the key video frame with the reference object as the center; or when a reference layout is identified from the key video frame, extract content elements of the target size from the key video frame. The target size meets the requirements for vertical screen display.
8. The method according to any one of claims 1-7, characterized in that, The step of generating a corresponding descriptive image set based on the semantic description information and in combination with the content element set in the film and television work includes: The semantic description information is segmented to obtain at least one keyword; For each keyword, retrieve the corresponding content element from the content element set; At least one content element obtained from the retrieval is concatenated to obtain at least one initial image; Based on the at least one initial image, a corresponding descriptive image set is generated.
9. The method as described in claim 8, characterized in that, The step of generating a corresponding descriptive image set based on the at least one initial image includes: Perform visual evaluation on each of the at least one initial image to obtain a visual evaluation value corresponding to each of the at least one initial image; Based on at least one visual evaluation value, select initial images that meet the evaluation criteria from the at least one initial image; The initial images obtained through filtering are used to form the descriptive image set.
10. The method according to any one of claims 1-7, characterized in that, The step of generating a corresponding descriptive image set based on the semantic description information and in combination with the content element set in the film and television work includes: The semantic description information is segmented to obtain at least one keyword; For each keyword, retrieve the corresponding content element from the content element set; Based on the text features extracted from the semantic description information, at least one noisy image feature associated with the retrieved content element is denoised to obtain the corresponding denoised image feature. For each of the denoised image features, a descriptive image prediction is performed to generate a corresponding descriptive image; Based on at least one generated descriptive image, a corresponding descriptive image set is generated.
11. The method as described in claim 10, characterized in that, The steps of extracting text features based on the semantic description information, denoising at least one noisy image feature associated with the retrieved content element to obtain corresponding denoised image features, and predicting a description image for each denoised image feature to generate a corresponding description image are performed by the target text-to-image model. The target text-to-image model is obtained by performing iterative training on the text-to-image model to be trained based on a training set of text-to-image sample pairs. The text-to-image sample pairs include: sample images extracted from the film and television works and descriptive text of the sample images, wherein the descriptive text is generated according to the set format.
12. The method according to any one of claims 1-7, characterized in that, Based on the description image set corresponding to each of the at least one semantic description information, the target video is obtained, including: For each type of semantic description information, the following steps are performed: extract subtitle information of the corresponding description image set from the plot description information corresponding to the semantic description information, and convert the subtitle information into speech information; For each session, the following steps are performed: based on at least one semantic description information associated with the session, the order of arrangement in the partial text content associated with the session, and combined with preset video effects, at least one set of descriptive images and corresponding audio information are synthesized into a video to obtain a sub-video corresponding to the session. Based on the arrangement order of the at least one scene in the text to be converted, the sub-videos corresponding to each of the at least one scene are spliced together to obtain the target video.
13. A video generation apparatus, characterized in that, The device includes: An extraction unit is used to extract at least one scene's environmental description text and plot description text based on the text to be converted from the target literary work; wherein, the text to be converted is the text portion of the target literary work that has not been used to generate a film or television work; The first generation unit is configured to perform the following for each of the aforementioned scenes: based on the environment description text and plot description text of the scene, generate at least one semantic description information according to a set format; wherein each of the semantic description information is used to describe at least one of the environment layout and the main object in the scene; The second generation unit is used to perform the following for each type of semantic description information: based on the semantic description information, and combined with the set of content elements in the film and television works generated based on the target literary work, generate a corresponding set of descriptive images; wherein each content element is a reference layout or reference object in the film and television works; The obtaining unit is used to obtain the target video based on the description image set corresponding to each of the at least one semantic description information.
14. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of any of the methods described in claims 1-12.
15. A computer-readable storage medium, characterized in that, It includes a computer program that, when run on an electronic device, causes the electronic device to perform the steps of any of the methods described in claims 1-12.
16. A computer program product, characterized in that, The method includes a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any one of claims 1-12.