Video generation method and device, equipment, medium and product
In the video generation method in the field of education, the interactive interface is used to determine the story text and divide narrative clips to generate videos that match the story text, which solves the problem of a single method of displaying idiom stories and improves the richness of user experience and interactive scenes.
Patent Information
- Application Number
- CN202510662466.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-01
AI Technical Summary
In the field of education, idiom stories are presented in a single way and inefficient manner, and there is a lack of automated video generation technology to improve the richness of user experience and interactive scenarios.
Interact with users through a preset interactive interface, determine the target story text, and divide it into narrative clip text, combine the time series relationship to generate scene images and dubbing audio, and automatically generate display videos matching the story text.
It realizes the vivid display of idiom stories, improves the richness of user experience and interactive scenes, and improves the automation and efficiency of video generation.
Smart Images

Figure CN120416622A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video synthesis, and in particular to a method, apparatus, device, medium and product for video generation. Background Art
[0002] With the continuous development of video synthesis technology, it has been applied in many fields. However, in fields such as education, the display of some story texts still mainly relies on traditional text and voice narration. Especially for the explanation and display of some idiom stories, there are problems such as a single presentation method and low production efficiency.
[0003] Therefore, how to automatically parse story texts, generate images and dub during the process of interacting with users for video display, so as to generate a story video that matches the story text, improve the user experience, and enhance the richness of the interaction scenario is an urgent problem to be solved at present. Summary of the Invention
[0004] The present invention provides a method, apparatus, device, medium and product for video generation, so as to automatically parse story texts, generate images and dub during the process of interacting with users for video display, so as to generate a story video that matches the story text, improve the user experience, and enhance the richness of the interaction scenario.
[0005] According to one aspect of the present invention, a method for video generation is provided, including:
[0006] Responding to a video generation request sent by a target user, interacting with the target user through a preset interaction interface to determine the target story text selected by the target user;
[0007] Dividing the target story text into at least two narrative segment texts, and determining the temporal relationship between the narrative segment texts;
[0008] Determining the scene images, sequence images and dubbing audio corresponding to each narrative segment text, and generating a display video corresponding to the target story text in combination with the temporal relationship between the narrative segment texts, so as to play and display the video to the target user.
[0009] According to another aspect of the present invention, a video generation apparatus is provided, including:
[0010] A text determination module, configured to respond to a video generation request sent by a target user, and interact with the target user through a preset interaction interface to determine the target story text selected by the target user;
[0011] A relationship determination module, configured to divide the target story text into at least two narrative segment texts, and determine the temporal relationship between the narrative segment texts;
[0012] A generation module, configured to determine the scene images, sequence images, and dubbed audio corresponding to each narrative segment text, and generate a display video corresponding to the target story text in combination with the temporal relationship between the narrative segment texts, so as to play the display video to the target user.
[0013] According to another aspect of the present invention, there is provided an electronic device, the electronic device including:
[0014] At least one processor; and
[0015] A memory communicatively connected to the at least one processor; wherein,
[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the video generation method according to any embodiment of the present invention.
[0017] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the video generation method according to any embodiment of the present invention when executed.
[0018] According to another aspect of the present invention, there is also provided a computer program product including a computer program that implements the video generation method according to any embodiment of the present invention when executed by a processor.
[0019] The technical solution of the embodiment of the present invention responds to a video generation request sent by a target user, interacts with the target user through a preset interaction interface to determine the target story text selected by the target user; divides the target story text into at least two narrative segment texts, and determines the temporal relationship between the narrative segment texts; determines the scene images, sequence images, and dubbed audio corresponding to each narrative segment text, and generates a display video corresponding to the target story text in combination with the temporal relationship between the narrative segment texts, so as to play the display video to the target user. By automatically parsing the story text, generating images, and dubbing during the process of interacting with the user for video display, a story video matching the story text can be generated, improving the user experience and enhancing the richness of the interaction scenario.
[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0022] Figure 1 It is a flowchart of a video generation method provided in Embodiment 1 of the present invention;
[0023] Figure 2 It is a flowchart of a video generation method provided in Embodiment 2 of the present invention;
[0024] Figure 3 It is a structural block diagram of a video generation device provided in Embodiment 3 of the present invention;
[0025] Figure 4 It is a schematic structural diagram of an electronic device provided in Embodiment 4 of the present invention. Detailed implementation manners
[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0027] It should be noted that the terms "first", "second", "target", "candidate", "alternative", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices. The acquisition, storage, use, processing, etc. of data in the technical solutions of this application all comply with the relevant regulations of laws and regulations.
[0028] Embodiment 1
[0029] Figure 1The figure is a flowchart of a video generation method provided in the first embodiment of the present invention. This embodiment is applicable to the situation where the story text selected by the user is automatically parsed to generate an appropriate display video. This method can be executed by a video generation device, which can be implemented in the form of hardware and / or software. The video generation device can be configured in an electronic device, such as an interactive device that can interact with the user, and is executed by the video synthesis system in the interactive device. As Figure 1 shown, the video generation method includes:
[0030] S101. In response to a video generation request issued by a target user, interact with the target user through a preset interaction interface to determine the target story text selected by the target user.
[0031] Among them, the target user refers to the user who interacts with the video synthesis system of the interactive device. The target user can be, for example, a student or a teaching teacher. The interactive device can be, for example, a user terminal device with a display function. The preset interaction interface refers to the interface through which the target user issues a request and the video is displayed to the target user. The preset interaction interface can be, for example, a touchable interface, and the target user can issue a video generation request through a preset touch method. The video generation request refers to a request to automatically generate a display video for the target story text selected by the target user. The target story text refers to the story text to be generated into a corresponding video selected by the target user. The target story text can be, for example, an idiom story text.
[0032] Optionally, interacting with the target user through the preset interaction interface to determine the target story text selected by the target user includes: displaying candidate idioms to the target user through the preset interaction interface, and determining the target idiom from the candidate idioms according to the selection operation of the target user; matching in the preset idiom story library according to the target idiom to determine the target story text corresponding to the target idiom.
[0033] Among them, the candidate idioms refer to the idioms in the idiom library pre-stored in the video generation system. The candidate idioms stored in the idiom library and the story texts stored in the preset idiom story library can be in one-to-one correspondence. The target idiom refers to the idiom selected by the target user from the candidate idioms for video generation.
[0034] Optionally, when the video generation system detects that the target user touches the corresponding area of the preset interaction interface, it is considered that the video generation request issued by the target user is detected, and further, the candidate idioms in the idiom library are displayed to the target user to indicate that the target user selects the idiom for video generation from them to determine the target idiom.
[0035] S102. Divide the target story text into at least two narrative segment texts, and determine the temporal relationship between the narrative segment texts.
[0036] Among them, the narrative segment text refers to at least two story segments included in the target story text. Each narrative segment text can correspond to a story scene in the target story text. For example, for the idiom story text of "Drawing a snake and adding feet", the narrative segments it contains can include the narrative segments corresponding to key scenes such as "the characters start drawing a snake", "the snake takes shape initially", and "extra feet appear". The temporal relationship refers to the relationship of the time before and after of each narrative segment text appearing in the target story text. By combining at least two narrative segment texts divided from the target story text based on the temporal relationship, the target story text can be restored.
[0037] Optionally, the target story text can be divided according to the title information and chapter distribution in the target story text to obtain at least two narrative segment texts, and the temporal relationship between the narrative segment texts can be determined according to the front-back position relationship of each narrative segment text in the target story text.
[0038] S103. Determine the scene images, sequence images, and dubbed audio corresponding to each narrative segment text, and combine the temporal relationship between the narrative segment texts to generate a display video corresponding to the target story text for playing the display video to the target user.
[0039] Among them, one narrative segment text corresponds to one scene image or a group of sequence images. If the narrative segment text corresponds to dynamic association information, video generation is performed based on a group of sequence images corresponding to the narrative segment text. If there is no dynamic association information in the association information corresponding to the narrative segment text, that is, only static association information is included, video generation is performed based on the scene image corresponding to the narrative segment text. The scene image refers to an image that can represent the scene depicted by the narrative segment text. The sequence image refers to a group of scene images that can represent the scene depicted by the narrative segment text and at the same time reflect the dynamic effect. The dubbed audio refers to the audio generated after dubbing the text in the narrative segment text. The display video refers to the video shown to the target user for vividly describing the target story text. The display video can represent the scene depicted by the target story text in video form.
[0040] Optionally, according to the temporal relationship between the narrative segment texts, the scene images corresponding to the narrative segment texts, at least two frames of sequence images corresponding to the scene images, and the dubbed audio can be spliced in sequence to generate a display video corresponding to the target story text and play the display video to the target user.
[0041] Optionally, determining the scene images and sequence images corresponding to each narrative segment text includes: determining the entity words included in each narrative segment text based on a preset entity extraction component and dependency syntactic analysis strategy, and determining the association information between the entity words and other grammatical components; using a pre-trained language model to determine the scene emotion label corresponding to each narrative segment text, and based on the association information and the scene emotion label, determining the scene image corresponding to each narrative segment text based on a preset visual template library and a preset image library; generating sequence images corresponding to at least two frame scene images according to the dynamic association information in the association information to reflect the dynamic effect corresponding to the dynamic association information.
[0042] Among them, the preset entity extraction component refers to a component used to extract entity words in the narrative segment text. The entity extraction component can be, for example, a component in natural language processing (NLP) used to identify predefined object categories in the text body, that is, a named entity recognition (NER) component. Entity words can be words representing people, time, place, or dialogue. The dependency syntactic analysis strategy refers to a strategy used to perform syntactic analysis on the narrative sentences in the narrative segment text. The dependency syntactic analysis strategy can analyze and identify the grammatical components such as "subject-predicate-object" and "attribute-adverbial-complement" in the sentence and analyze the relationships between the components.
[0043] The pre-trained language model refers to a preset model used to identify the dominant emotion in the text. The pre-trained language model can be, for example, a BERT model (Bidirectional Encoder Representations from Transformers). The scene emotion label can be, for example, tense, warm, or lyrical, etc.
[0044] Optionally, each narrative segment text can be input into the pre-trained language model respectively to obtain the scene emotion label corresponding to each narrative segment text. Exemplarily, when it is detected that the narrative segment text appears the phrase "everyone is waiting anxiously" based on the pre-trained language model, the output scene emotion label can be "tense". [[ID=?]]
[0045] Optionally, a preset image variant component can be used to generate sequence images corresponding to at least two frame scene images according to the dynamic association information in the association information to reflect the dynamic effect corresponding to the dynamic association information. For example, for the text of the idiom story of "borrowing arrows with thatched boats", the dynamic association information in its narrative segment text can be "the river is flowing" and "arrows are flying". At this time, the preset image variant component can be used to generate image variants for the river surface and the arrows respectively to obtain a multi-frame image sequence, thereby realizing the corresponding dynamic effect.
[0046] Optionally, based on a preset entity extraction component and dependency syntactic analysis strategy, determine the entity words included in each narrative segment text, and determine the association information between the entity words and other grammatical components, including: for each narrative segment text, based on the preset entity extraction component, determine the entity words included in the narrative segment text; determine the narrative sentences to which the entity words belong, and based on the dependency syntactic analysis strategy, perform component analysis on the narrative sentences to which the entity words belong to determine the association information between the entity words and other grammatical components in each narrative sentence.
[0047] Among them, a narrative sentence refers to the sentence in which the entity word is located in the narrative segment text. Other grammatical components refer to other segments such as predicates, subjects, or objects in the sentence when the entity word is the subject or object. Association information refers to information that can represent the association relationship such as the subordinate relationship or affiliated relationship between the entity word and other grammatical components. Association information includes: dynamic association information and static association information. For example, if the entity word is a person, the corresponding dynamic association information is the person's action information, and the static association information is the person's speech information and / or person description information. Person description information can be, for example, hair color, skin color, old person, doctor, and young person, etc.
[0048] Exemplarily, for the text of the idiom story "Drawing a snake and adding feet", the corresponding entity words can be "snake", "foot", "draw", etc., and the corresponding association information can be "a person draws a snake", "a person adds feet", etc. For the text of the idiom story "Borrowing arrows with thatched boats", the corresponding entity words can be "thatched boat", "arrow", "warship", "river surface", etc., and the corresponding association information can be "warship comparison", "river surface flowing", "arrow flying", "thatched boat arrow inventory is low", etc.
[0049] Optionally, based on the association information and scene emotion labels, and based on a preset visual template library and a preset image library, determine the scene images corresponding to each narrative segment text, including: match in the preset visual template library according to the static association information in the association information to determine the composition rules corresponding to each narrative segment text; match from the preset image library the scene images corresponding to each narrative segment text according to the composition rules and the color matching scheme corresponding to the scene emotion labels.
[0050] Among them, the visual template library refers to a library of composition rules with preset configurations that can achieve different visual effects, and the color matching scheme refers to the color scheme corresponding to the narrative segment text.
[0051] Exemplarily, if the scene emotion label is tense, the corresponding color matching scheme can be a cool color system, and if the scene emotion label is cheerful, the corresponding color matching scheme can be a warm color system.
[0052] Optionally, determine the dubbing audio corresponding to each narrative segment text, including: dividing each narrative segment text into a dialogue part and a narration part, and respectively determining the dialogue voice corresponding to the dialogue part and the narration voice corresponding to the narration part; according to the positional structure relationship between the dialogue part and the narration part in each narrative segment text, allocate the narration voice and the dialogue voice to generate the dubbing audio corresponding to each narrative segment text.
[0053] Optionally, if the entity word associated with the dialogue part is the target person, then according to the static association information of the entity words in each narrative segment text, the target role to which the target person belongs can be matched from the preset role library; further, the dialogue voice corresponding to the dialogue part can be generated according to the voice characteristics of the target role.
[0054] Optionally, the narration speed and narration intonation of the narration part can be determined according to the scene emotion label corresponding to each narrative segment text; and according to the narration speed and narration intonation, the preset text-to-speech technology (Text To Speech, TTS) is used to generate the narration voice corresponding to the narration part.
[0055] Exemplarily, if the scene emotion label is tense, the corresponding narration speed can be fast, and the corresponding narration intonation can be compact; if the scene emotion label is relaxed, the corresponding narration speed can be slow, and the corresponding narration intonation can be slow, and the timbre can be soft.
[0056] Exemplarily, corresponding timbre styles can be set for different scene emotion labels in the preset template. When the scene emotion label is "tense", the background low drumbeat can be increased or the speed can be increased; when the scene emotion label is "warm", a soft background music can be selected, the speed can be reduced and the transition can be smoothed.
[0057] The technical solution of the embodiment of the present invention responds to the video generation request sent by the target user, interacts with the target user through the preset interaction interface to determine the target story text selected by the target user; divides the target story text into at least two narrative segment texts, and determines the chronological relationship between the narrative segment texts; determines the scene image, sequence image and dubbing audio corresponding to each narrative segment text, and combines the chronological relationship between the narrative segment texts to generate the display video corresponding to the target story text, so as to play the display video for the target user. By automatically parsing the story text, generating images and dubbing during the process of interacting with the user for video display, a story video matching the story text can be generated, improving the user experience and enhancing the richness of the interaction scenario.
[0058] Embodiment 2
[0059] Figure 2It is a flowchart of a video generation method provided in the second embodiment of the present invention. On the basis of the above embodiment, a preferred example of generating a display video corresponding to the target story text is provided. Specifically, as Figure 2 shown, the method includes:
[0060] S201. In response to a video generation request sent by a target user, display candidate idioms to the target user through a preset interaction interface, and determine a target idiom from the candidate idioms according to the selection operation of the target user.
[0061] S202. Match in a preset idiom story library according to the target idiom to determine a target story text corresponding to the target idiom.
[0062] S203. Divide the target story text into at least two narrative segment texts, and determine the temporal sequence relationship between the narrative segment texts.
[0063] S204. Based on a preset entity extraction component and a dependency syntax analysis strategy, determine the entity words included in each narrative segment text, and determine the association information between the entity words and other grammatical components.
[0064] S205. Use a pre-trained language model to determine the scene emotion labels corresponding to each narrative segment text, and based on the association information and the scene emotion labels, determine the scene images corresponding to each narrative segment text based on a preset visual template library and a preset image library.
[0065] S206. Generate a sequence of images corresponding to at least two scene images according to the dynamic association information in the association information to reflect the dynamic effect corresponding to the dynamic association information.
[0066] S207. Divide each narrative segment text into a dialogue part and a narration part, and respectively determine the dialogue voice corresponding to the dialogue part and the narration voice corresponding to the narration part.
[0067] S208. According to the positional structure relationship between the dialogue part and the narration part in each narrative segment text, allocate the narration voice and the dialogue voice to generate a dubbed audio corresponding to each narrative segment text.
[0068] S209. Combine the scene images, sequence images and dubbed audio corresponding to each narrative segment text, and generate a display video corresponding to the target story text according to the temporal sequence relationship between the narrative segment texts, so as to play the display video to the target user.
[0069] Embodiment Three
[0070] Figure 3It is a structural block diagram of a video generation device provided in Embodiment 3 of the present invention; this embodiment is applicable to the situation of automatically parsing the story text selected by the user to generate an appropriate display video. The video generation device provided in the embodiments of the present invention can execute the video generation method provided in any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method; the video generation device can be implemented in the form of hardware and / or software, and is configured in an electronic device with video generation function, such as an interactive device that can interact with the user, and is executed by the video synthesis system in the interactive device, such as Figure 3 As shown, the video generation device may specifically include:
[0071] A text determination module 301, configured to interact with the target user through a preset interaction interface in response to a video generation request issued by the target user, so as to determine the target story text selected by the target user;
[0072] A relationship determination module 302, configured to divide the target story text into at least two narrative segment texts, and determine the temporal relationship between the narrative segment texts;
[0073] A generation module 303, configured to determine the scene images, sequence images, and dubbed audio corresponding to each narrative segment text, and generate a display video corresponding to the target story text in combination with the temporal relationship between the narrative segment texts, so as to play the display video to the target user.
[0074] The technical solution of the embodiment of the present invention, in response to a video generation request issued by the target user, interacts with the target user through a preset interaction interface to determine the target story text selected by the target user; divides the target story text into at least two narrative segment texts, and determines the temporal relationship between the narrative segment texts; determines the scene images, sequence images, and dubbed audio corresponding to each narrative segment text, and generates a display video corresponding to the target story text in combination with the temporal relationship between the narrative segment texts, so as to play the display video to the target user. By automatically parsing the story text, generating images, and dubbing during the process of interacting with the user for video display, a story video matching the story text can be generated, improving the user experience and enhancing the richness of the interaction scenario.
[0075] Further, the generation module 303 may include:
[0076] An information determination unit, configured to determine the entity words included in each narrative segment text based on a preset entity extraction component and dependency syntax analysis strategy, and determine the association information between the entity words and other grammatical components;
[0077] An image determination unit, which uses a pre-trained language model to determine the scene emotion labels corresponding to the text of each narrative segment, and based on the association information and the scene emotion labels, determines the scene images corresponding to the text of each narrative segment based on a preset visual template library and a preset image library;
[0078] An effect determination unit, which generates a sequence of images corresponding to at least two frame scene images according to the dynamic association information in the association information to reflect the dynamic effect corresponding to the dynamic association information.
[0079] Furthermore, the information determination unit is specifically used for:
[0080] For each narrative segment text, based on a preset entity extraction component, determine the entity words included in the narrative segment text;
[0081] Determine the narrative sentences to which the entity words belong, and based on the dependency syntactic analysis strategy, perform constituent analysis on the narrative sentences to which the entity words belong to determine the association information between the entity words and other grammatical components in each narrative sentence; the association information includes: dynamic association information and static association information.
[0082] Furthermore, the image determination unit is specifically used for:
[0083] Match in the preset visual template library according to the static association information in the association information to determine the composition rules corresponding to the text of each narrative segment;
[0084] Match the scene images corresponding to the text of each narrative segment from the preset image library according to the composition rules and the color matching scheme corresponding to the scene emotion labels.
[0085] Furthermore, the generation module 303 is also used for:
[0086] Divide the text of each narrative segment into a dialogue part and a narration part, and respectively determine the dialogue voice corresponding to the dialogue part and the narration voice corresponding to the narration part;
[0087] According to the positional structure relationship between the dialogue part and the narration part in the text of each narrative segment, allocate the narration voice and the dialogue voice to generate the dubbing audio corresponding to the text of each narrative segment.
[0088] Furthermore, the text determination module 301 is specifically used for:
[0089] Display candidate idioms to the target user through a preset interaction interface, and determine the target idiom from the candidate idioms according to the selection operation of the target user;
[0090] Match in the preset idiom story library according to the target idiom to determine the target story text corresponding to the target idiom.
[0091] Example 4
[0092] Figure 4 FIG. is a schematic structural diagram of an electronic device provided in Example 4 of the present invention. Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0093] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0094] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0095] The processor 11 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the video generation method.
[0096] In some embodiments, the video generation method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the video generation method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the video generation method by any other suitable means (e.g., by means of firmware).
[0097] The various implementations of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0098] The computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0099] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0100] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0101] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0102] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs that run on respective computers and have a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to address the defects of high management difficulty and weak business scalability existing in traditional physical hosts and VPS services.
[0103] In one embodiment, the embodiment of the present invention further includes a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the video generation method of any embodiment of the present invention.
[0104] In the process of implementing the computer program product, the computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages and also conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (e.g., connected via the Internet using an Internet service provider).
[0105] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0106] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A video generation method, characterized in that, including: In response to a video generation request issued by a target user, interact with the target user through a preset interaction interface to determine the target story text selected by the target user; Divide the target story text into at least two narrative segment texts and determine the temporal relationship between the narrative segment texts; Determine the scene images, sequence images, and dubbed audio corresponding to each narrative segment text, and generate a display video corresponding to the target story text in combination with the temporal relationship between the narrative segment texts, so as to play the display video for the target user.
2. The method according to claim 1, wherein Determine the scene images and sequence images corresponding to each narrative segment text, including: Based on a preset entity extraction component and dependency syntax analysis strategy, determine the entity words included in each narrative segment text, and determine the association information between the entity words and other grammatical components; Adopt a pre-trained language model to determine the scene emotion label corresponding to each narrative segment text, and based on the association information and scene emotion label, determine the scene image corresponding to each narrative segment text based on a preset visual template library and a preset image library; Generate sequence images corresponding to at least two frames of scene images according to the dynamic association information in the association information to reflect the dynamic effect corresponding to the dynamic association information.
3. The method according to claim 2, wherein Based on a preset entity extraction component and dependency syntax analysis strategy, determine the entity words included in each narrative segment text, and determine the association information between the entity words and other grammatical components, including: For each narrative segment text, based on a preset entity extraction component, determine the entity words included in the narrative segment text; Determine the narrative sentences to which the entity words belong, and based on the dependency syntax analysis strategy, perform component analysis on the narrative sentences to which the entity words belong to determine the association information between the entity words and other grammatical components in each narrative sentence; the association information includes: dynamic association information and static association information.
4. The method according to claim 2, characterized in that, Based on the association information and scene emotion label, determine the scene image corresponding to each narrative segment text based on a preset visual template library and a preset image library, including: Match according to the static association information in the association information in the preset visual template library to determine the composition rules corresponding to each narrative segment text; Match the scene images corresponding to each narrative segment text from the preset image library according to the composition rules and the color matching scheme corresponding to the scene emotion label.
5. The method according to claim 1, characterized in that Determine the dubbed audio corresponding to each narrative segment text, including: Divide each narrative segment text into a dialogue part and a narration part, and respectively determine the dialogue voice corresponding to the dialogue part and the narration voice corresponding to the narration part; According to the positional structure relationship between the dialogue part and the narration part in each narrative segment text, allocate the narration voice and the dialogue voice to generate the dubbed audio corresponding to each narrative segment text.
6. The method according to claim 1, wherein Interact with the target user through a preset interaction interface to determine the target story text selected by the target user, including: Display candidate idioms to the target user through a preset interaction interface, and determine the target idiom from the candidate idioms according to the selection operation of the target user; Match in a preset idiom story library according to the target idiom to determine the target story text corresponding to the target idiom.
7. A video generation device, characterized in that, including: A text determination module, configured to interact with a target user through a preset interaction interface in response to a video generation request issued by the target user, so as to determine target story text selected by the target user; A relationship determination module, configured to divide the target story text into at least two narrative segment texts and determine the temporal relationship between the narrative segment texts; A generation module, configured to determine scene images, sequence images, and dubbed audio corresponding to the narrative segment texts, and generate a display video corresponding to the target story text in combination with the temporal relationship between the narrative segment texts, so as to play the display video to the target user.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute the video generation method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the video generation method according to any one of claims 1-6 when executed by a processor.
10. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program implements the video generation method according to any one of claims 1-6 when executed by a processor.
Citation Information
Patent Citations
Video generation method, video generation system, electronic equipment and storage medium
CN114827752A
Story video generation corresponding to user input using generative models
CN118212328A
Video generation method and device, electronic equipment and storage medium
CN118283288A
Video generation method and device, equipment, storage medium and program product
CN119031201A
Automated creation of storyboards from screenplays
US9106812B1
Cited By
Video processing method and device, electronic equipment, storage medium and program product
CN121418635A
Data video generation method and device, equipment and medium
CN121482220A