Video generation method, device, machine, memory medium and program product

The video generation method automates the creation of text and video based on user-input keywords, addressing the inefficiency in video creation by providing an efficient one-stop solution for video creation.

JP2025518428AActive Publication Date: 2025-06-17BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023578865
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-19
Filing Date
2023-12-12
Publication Date
2025-06-17
Estimated Expiration
2043-12-12

AI Technical Summary

Technical Problem

The inefficiency in video creation due to the time-consuming process of searching for numerous video materials such as images, texts, and music, which hinders the productivity of users sharing videos on platforms.

Method used

A video generation method and apparatus that automatically generates text based on user-input keywords and subsequently creates a video using the generated text, providing an efficient one-stop solution for video creation.

Benefits of technology

This approach significantly enhances video creation efficiency by automating the text and video generation process, reducing the time and effort required for users to find and assemble video materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025518428000001_ABST
    Figure 2025518428000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video generation method, apparatus, device, storage medium, and program product. The method includes obtaining first text information for describing the creation requirements of video captions, generating second text information based on the first text information, where the second text information is caption information that meets the creation requirements described in the first text information, generating multimedia editing data based on third text information, and generating a target video based on the multimedia editing data. In an embodiment of the present invention, based on the creation requirements of video captions, caption information that meets the described creation requirements is generated, and further, a video is created using the generated caption information, thereby providing an efficient one-stop video creation solution and improving the video creation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of video processing, and in particular to a video generation method, apparatus, device, storage medium, and program product.

Background Art

[0002] With the rapid development of computer technology and mobile communication technology, various video platforms based on electronic devices have become generally available, greatly enriching people's daily lives. The number of users who enjoy sharing their video works on video platforms so that other users can view them is increasing.

[0003] Before sharing a video work, the user needs to edit it himself to create a video. When creating a video, it is necessary to search for a large number of video materials by oneself, such as images, texts, music, etc. Searching for materials takes a long time, leading to inefficiency in video creation.

Summary of the Invention

Problems to be Solved by the Invention

[0004] To solve the above technical problems, embodiments of the present invention provide a video generation method, apparatus, device, storage medium, and program product that automatically generate a text based on keywords input by a user and automatically generate a video based on the generated text, providing an efficient one-stop video creation solution and improving video creation efficiency.

Means for Solving the Problems

[0005] In a first aspect, embodiments of the present disclosure provide a video generation method. The video generation method includes: obtaining first text information for describing the creation requirements of video language; generating second text information based on the first text information, where the second text information is language information that meets the creation requirements described in the first text information; Generating multimedia editing data based on third text information obtained based on the second text information, wherein the multimedia editing data includes at least one video track clip and at least one audio track clip, the at least one video track clip and the at least one audio track clip respectively correspond to at least one text clip partitioned by the third text information, the target audio track clip is used to fill in a voiceover that matches the target text clip, and the target video track clip in the at least one video track clip and the target audio track clip occupy the same timeline position on the video editing timeline, and Generating a target video based on the multimedia editing data, and include.

[0006] In a second aspect, an embodiment of the present disclosure provides a video generation device. The video generation device is A first text information acquisition module for acquiring first text information, the first text information acquisition module for generating multimedia editing data based on third text information, and A second text information generation module for generating second text information based on the first text information, wherein the second text information is a second text information generation module that is literal information that meets the creation requirements described in the first text information, and A multimedia editing data generation module for generating multimedia editing data based on third text information obtained based on the second text information, wherein the multimedia editing data includes at least one video track clip and at least one audio track clip, and the at least one video track clip and the at least one audio track clip respectively correspond to at least one text clip partitioned by the third text information, the target audio track clip is used to fill in the voiceover that matches the target text clip, and the target video track clip in the at least one video track clip and the target audio track clip occupy the same timeline position on the video editing timeline, and a multimedia editing data generation module, A target video generation module for generating a target video based on the multimedia editing data, Comprising.

[0007] In a fourth aspect, an embodiment of the present disclosure provides an electronic device. The electronic device Comprises at least one processor and A storage device for storing at least one program. When the at least one program is executed by the at least one processor, the at least one processor realizes the video generation method described in any one of the first aspects above.

[0008] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the video generation method described in any one of the first aspects above is realized.

[0009] In a fifth aspect, embodiments of the present disclosure provide a computer program product. The computer program product includes a computer program or instructions. When the computer program or instructions are executed by a processor, the video generation method according to any one of the first aspects is implemented.

[0010] Embodiments of the present disclosure provide a video generation method, apparatus, device, storage medium, and program product. The method includes obtaining first text information for describing the creation requirements of video captions, and generating second text information based on the first text information, where the second text information is caption information that meets the creation requirements described in the first text information, and generating multimedia editing data based on third text information obtained based on the second text information, where the multimedia editing data includes at least one video track clip and at least one audio track clip, the at least one video track clip and the at least one audio track clip respectively correspond to at least one text clip partitioned by the third text information, the target audio track clip is used to fill in the voiceover that matches the target text clip, the target video track clip in the at least one video track clip and the target audio track clip occupy the same timeline position on the video editing timeline, and generating a target video based on the multimedia editing data. In embodiments of the present invention, based on the creation requirements of video captions, caption information that meets the described creation requirements is generated, and further, a target video is created using the generated caption information, thereby providing an efficient one-stop video creation solution and improving the video creation efficiency.

Brief Description of the Drawings

[0011] Referring to the following specific embodiments in conjunction with the drawings, the above and other features, advantages, and aspects of each example of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original and elements are not necessarily drawn to scale.

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7a

Figure 7b

Figure 7c

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0013] Hereinafter, with reference to the drawings, embodiments of the present disclosure will be described in more detail. Although some embodiments of the present disclosure are shown in the drawings, the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Rather, it should be understood that these embodiments are provided to more clearly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are used only for illustration and not for limiting the scope of protection of the present disclosure.

[0014] It should be understood that each step described in the method embodiments of the present disclosure may be executed in a different order and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this regard.

[0015] As used herein, the term "comprising" and its variations are non-limiting inclusion, that is, "including... but not limited to...". The term "based on" means "at least partially based on...". The term "one embodiment" means "at least one embodiment", the term "another embodiment" means "at least one another embodiment", and the term "some embodiments" means "at least some embodiments". Related definitions of other terms are given in the following description.

[0016] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and do not limit the order or interdependence of the functions executed by these devices, modules, or units.

[0017] It should be noted that the modifications of "one" and "a plurality" mentioned in the present disclosure are not limiting but exemplary, and those skilled in the art should understand that, unless otherwise explicitly stated in the context, it should be understood as "at least one".

[0018] The names of messages or information exchanged between multiple devices in the embodiments of the present disclosure are used merely for the purpose of description and are not used to limit the scope of these messages or information.

[0019] Hereinafter, with reference to the drawings, embodiments of the present disclosure will be described in detail. Note that the same reference numerals in different drawings are used to refer to the same described elements.

[0020] FIG. 1 is a flowchart of a video generation method according to an embodiment of the present disclosure. This embodiment is applicable when generating a video from keywords. The method is executable by a video generation device. The video generation device can be implemented in software and / or hardware. The video generation method is applicable to electronic devices.

[0021] The above-mentioned electronic device may be any other type of electronic device capable of executing data processing, including but not limited to mobile phones, websites, units, devices, multimedia computers, multimedia tablets, Internet nodes, communication devices, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, and it is understood that it may include accessories and peripheral devices of these devices or any combination thereof.

[0022] As shown in FIG. 1, the video generation method according to the embodiment of the present invention mainly includes steps S101 to S104.

[0023] In S101, first text information for describing the creation requirements of video language is obtained.

[0024] In one embodiment of the present disclosure, the creation requirements of the video caption may be the key content of the target video that the user wants to generate. Specifically, the first text information may be one or more keywords describing the video caption, or one or more topic words.

[0025] In one exemplary illustration, the user wants to create a video related to "movie recommendation". The first text information may include "movie", "2023", "award", "good reviews", etc. The user wants to create a video related to "sales points of mobile phones". The first text information may include "models of mobile phones", "ultra-large screen", "excellent battery life", "affordable price", etc.

[0026] In one embodiment of the present disclosure, obtaining the first text information includes obtaining the first text information corresponding to the input operation in response to the user's input operation. Specifically, a text creation video control is presented on the front page of video creation, and a video creation screen is displayed in response to a trigger operation on the text creation video control. As shown in FIG. 2, the video creation screen includes a text creation area 21, a video category selection area 22, and a video generation control 23. The text creation area 21 includes a text display area and an intelligent text generation control. The text display area is for displaying the third text information for generating multimedia editing data. The intelligent text generation control is for displaying a caption input screen in response to the user's trigger operation.

[0027] In one embodiment of the present disclosure, as shown in FIG. 3, the caption input page includes a caption input area 31 for displaying the text information input by the user in response to the user's input operation.

[0028] As shown in FIG. 3, the text input area 31 includes an input box for acquiring first text information according to a user's input operation.

[0029] In S102, second text information is generated based on the first text information. The second text information is text information that conforms to the creation requirements described in the first text information.

[0030] In an embodiment of the present disclosure, the second text information means text information that conforms to the creation requirements. Further, the second text information has a connection before and after and is a sentence that makes sense. The second text information may have one paragraph or may have a plurality of paragraphs. The second text information may be generated based on the first text information by an intelligent text generation algorithm. Specific embodiments of the intelligent text generation algorithm will not be further described in the embodiments of the present disclosure.

[0031] As shown in FIG. 3, in response to a confirmation operation on the input box in the text input area 31, second text information is generated based on the first text information, and the generated second text information is displayed in the text input area 31.

[0032] In S103, multimedia editing data is generated based on the third text information obtained based on the second text information. The multimedia editing data includes at least one video track clip and at least one audio track clip. The at least one video track clip and the at least one audio track clip respectively correspond to at least one text clip partitioned by the third text information. The target audio track clip in the at least one audio track clip is used to fill in the voiceover that matches the target text clip. The target video track clip in the at least one video track clip and the target audio track clip occupy the same timeline position on the video editing timeline.

[0033] In one embodiment of the present disclosure, the third text information may include one or a combination of one or more second text information, edited second text information, and other text information edited by the user.

[0034] In one embodiment of the present disclosure, steps S101 to S102 may be executed multiple times to generate multiple second text information, or the generated second text information may be edited and changed, or text information may be manually input into the text input area 31.

[0035] In one embodiment of the present disclosure, as shown in FIG. 3, in response to a trigger operation on the "Complete" control in the text input area 31, the screen switches to the video creation screen shown in FIG. 2, and the third text information is displayed on the video creation screen.

[0036] In one embodiment of the present disclosure, when the third text information is displayed on the video creation screen, in response to a trigger operation on the video generation control 23, multimedia editing data is generated based on the third text information, and the multimedia editing data is displayed on the multimedia editing screen.

[0037] In one embodiment of the present disclosure, the multimedia editing data further includes at least one subtitle track clip for filling subtitle information that matches the target text clip.

[0038] In one embodiment of the present disclosure, the multimedia editing data further includes at least one background music track clip for filling background music.

[0039] In one embodiment of the present disclosure, as shown in FIG. 4, the multimedia editing screen may include a video preview area 41 for previewing a target video generated by the multimedia editing data, and a multimedia track area 42. The multimedia editing data displayed in the multimedia track area 42 may include a video track clip, an audio track clip, and a background music track clip.

[0040] In one embodiment of the present disclosure, the target audio track clip of the at least one audio track clip is used to fill the voiceover that matches the target text clip. The target video track clip in the at least one video track clip and the target audio track clip occupy the same timeline position on the video editing timeline.

[0041] In one embodiment of the present disclosure, the target text clip may be a text clip partitioned by third text information. The text clip may be a single sentence, an incomplete sentence constructed according to language delimiter rules, or multiple sentences, and is not particularly limited in the embodiments of the present disclosure. The target video track clip refers to a video clip corresponding to the target text clip. The target audio track clip refers to an audio clip corresponding to the target text clip.

[0042] Note that the background music in the background music track clip is not partitioned by text clips and is filled with a complete audio file.

[0043] In one embodiment of the present disclosure, the target video track clip is an empty clip.

[0044] In an embodiment of the present disclosure, as shown in FIG. 2, in response to a trigger operation on the "first video category" control in the video creation screen, the video track clip in the multimedia editing data is set to an empty clip. That is, no picture or video image is added to the video track clip.

[0045] Furthermore, in response to an operation by the user on the empty clip, a picture selected by the user is added to the empty clip. In other words, the user may fill the empty clip with a picture selected by the user so that the created and completed target video matches the user's expectations.

[0046] In one embodiment of the present disclosure, the target video track clip is used to fill a video image that matches the target text clip.

[0047] In an embodiment of the present disclosure, as shown in FIG. 2, in response to a trigger operation on the "Second Video Category" control on the video creation screen, a video image matching the text clip is filled into the video track clip in the multimedia editing data. The video image can be obtained based on matching the text clip within a preset picture database according to an image matching algorithm.

[0048] In one embodiment of the present disclosure, the video image may be automatically matched within the picture database based on the text content. This avoids the user spending time searching for materials and further improves the video creation efficiency.

[0049] In one embodiment of the present disclosure, the target video track clip is used to fill an expression image matching the target text clip.

[0050] In an embodiment of the present disclosure, as shown in FIG. 2, in response to a trigger operation on the "Third Video Category" control on the video creation screen, an expression image matching the text clip is filled into the video track clip in the multimedia editing data. The expression image can be obtained based on matching the text clip within a preset expression image database according to a preset algorithm.

[0051] In one embodiment of the present disclosure, the expression image may be automatically matched within the expression image database based on the text content. Thereby, a more personalized video may be created based on the text content.

[0052] In an embodiment of the present disclosure, by separating the expression image database and the picture database into two databases, the problem of inconsistent video styles caused by the presence of both normal pictures and expression images in one video can be avoided.

[0053] In S104, a target video is generated based on the multimedia editing data.

[0054] In one embodiment of the present disclosure, in response to a trigger operation for video completion, a target video is generated based on multimedia editing data.

[0055] In one embodiment of the present disclosure, the trigger operation for "completion" of the video may refer to a trigger operation for the "Export" control on the multimedia editing screen. The export method may be to locally save the target video, or to share the target video on another video sharing platform or website. In the embodiments of the present disclosure, it is not specifically limited.

[0056] In one embodiment of the present disclosure, in response to a trigger operation for the "Import and Edit" control on the multimedia editing screen, the multimedia editing data is imported into a video editor, and subsequent editing is performed on the multimedia editing data.

[0057] In addition to the above embodiments, the embodiments of the present disclosure further optimize the video generation method. As shown in FIG. 5, the optimized video generation method mainly includes the following steps.

[0058] In S201, first text information for describing the creation requirements of video language is obtained.

[0059] Step S201 according to the embodiment of the present disclosure is the same as the specific execution flow in step S101 according to the above embodiment. Specifically, reference may be made to the description in the above embodiment, so the details thereof are omitted in the embodiments of the present disclosure.

[0060] In S202, based on the first text information, at least one candidate language information is generated. All of the at least one candidate language information meets the creation requirements indicated by the first text information.

[0061] In one embodiment of the present disclosure, a target language category is identified from a plurality of language categories. The candidate language information is generated based on the target language category and the first text information. The language category of the second text information is the target language category.

[0062] In one embodiment of the present disclosure, the language category refers to the category to which the formed language belongs. Specifically, the language category may include a first language category and a second language category. The first language category may be understood as a language category generally applied by users to various topics. Specifically, various topics include science and technology, economy, entertainment, etc. The second language category may be understood as a language category applied to product planning or product marketing. Specifically, it may be an introduction of the sales points of a certain mobile phone, or an introduction of the reasons for recommending a certain item.

[0063] In one embodiment of the present disclosure, the target language category may be determined based on the user's selection. Each target language category corresponds to an intelligent language generation algorithm model.

[0064] In one embodiment of the present disclosure, the language input screen includes a "first language category" control and a "second language category" control. The "first language category" control uses the language category corresponding to the first language category control as the target language category in response to the user's trigger operation. The "second language category" control uses the language category corresponding to the second language category control as the target language category in response to the user's trigger operation.

[0065] As shown in FIG. 3, a "first text category" control and a "second text category" control are displayed on the text input screen. In an embodiment of the present disclosure, different text category controls correspond to different prompt information, and different text category controls correspond to different intelligent text generation algorithm models.

[0066] In one embodiment of the present disclosure, in response to a selection operation on the "first text category" control, prompt information "Please write one paragraph of text in the first category. The topic is," is displayed in the "input box" in the text input area 31. The user may insert first text information after the prompt information. In response to a confirmation operation on the "input box", the first text information is obtained, and at least one candidate text information is generated based on the first text category and the first text information.

[0067] In one embodiment of the present disclosure, generating at least one candidate text information based on the first text category and the first text information includes calling a first intelligent text generation algorithm corresponding to the first text category based on the first text category, and using the first intelligent text generation algorithm to process the first text information to generate a plurality of candidate text information.

[0068] In one embodiment of the present disclosure, in response to a selection operation on the "second text category" control, prompt information "Please write one paragraph of text in the second category. The product and selling points are," is displayed in the input box in the text input area 31. The user may input the first text information after the prompt information. In response to a confirmation operation on the "input box", the first text information is obtained, and at least one candidate text information is generated based on the second text category and the first text information.

[0069] In one embodiment of the present disclosure, generating at least one candidate text information based on a second text category and first text information includes calling a second intelligent text generation algorithm corresponding to the second text category based on the second text category, and using the second intelligent text generation algorithm to process the first text information to generate a plurality of candidate text information.

[0070] It should be noted that the first intelligent text generation algorithm and the second intelligent text generation algorithm are two different intelligent text generation algorithms.

[0071] In one embodiment of the present disclosure, the basic network models used by the above two intelligent text generation algorithms may be the same or different. The training methods of the above two intelligent text generation algorithms may be the same or different. It should be noted that the training samples of the first intelligent text generation algorithm and the second intelligent text are different. The training sample of the first intelligent text generation algorithm is the text information of the first text category. The training sample of the second intelligent text generation algorithm is the text information of the second text category.

[0072] In S203, in response to a switching operation triggered by the user, different candidate text information among the at least one candidate text information is switched and displayed on the text input screen.

[0073] In one embodiment of the present disclosure, the text input screen includes a candidate text area for displaying candidate text information. Specifically, the candidate text area is displayed in the text input area 31 in an inserted form.

[0074] In one embodiment of the present disclosure, on the text input screen, a text "switch" control for triggering the switching operation is included so that different candidate text information among the at least one candidate text information is switched and presented in the candidate text area.

[0075] In one embodiment of the present disclosure, as shown in FIG. 6, the candidate text area is displayed in the text input area 31 in an inserted form. The candidate text area 61 includes a "first text switch" control 62 and a "second text switch" control 63.

[0076] In one embodiment of the present disclosure, a plurality of candidate text information are arranged in a set order. The first text switch control 62 is used to display, in the candidate text area, candidate text information whose arrangement order is before the current candidate text information in response to a trigger operation by the user. The second text switch control 63 is used to display, in the candidate text area, candidate text information whose column order is after the current candidate text information in response to a trigger operation by the user.

[0077] In one embodiment of the present disclosure, an example will be described in which there are five pieces of candidate text information. The five pieces of candidate text information are, in order, candidate text information A, candidate text information B, candidate text information C, candidate text information D, and candidate text information E. In the candidate text area, the candidate text information A with the highest display rank is displayed. In response to a trigger operation on the second text switch control 63, the candidate text information B is displayed in the candidate text area. At this time, in response to a trigger operation on the first text switch control 62, the candidate text information A is displayed in the candidate text area.

[0078] In S204, in response to a confirmation operation triggered by the user, among the at least one candidate text information, the candidate text information that is switched and displayed on the text input screen is determined as the second text information.

[0079] In one embodiment of the present disclosure, the text input screen includes a text confirmation control. The text confirmation control is used to trigger the confirmation operation so that, among the at least one candidate text information, the candidate text information that is switched and displayed in the candidate text area is determined as the second text information, and the second text information is displayed in the text input area.

[0080] In one embodiment of the present disclosure, as shown in FIG. 6, the candidate text area 61 includes a text confirmation control. In response to a trigger operation on the text confirmation control, the candidate text information displayed in the candidate text area is determined as the second text information, and the second text information is displayed in the text input area 31.

[0081] In one embodiment of the present disclosure, after the at least one candidate text information is generated, the candidate text information displayed in the candidate text area is inserted into the user input position in the text input area. When the confirmation operation is responded to, the candidate text area is deleted from within the text input area.

[0082] In one embodiment of the present disclosure, as shown in FIG. 6, the candidate text information (AAAAA) displayed in the candidate text area 61 is inserted into the user input position in the text input area 31. The user input position is the position where the cursor was before the candidate text information was generated. Further, in response to a trigger operation by the user on the text confirmation control, the candidate text area is deleted.

[0083] In one embodiment of the present disclosure, as shown in FIG. 7a, when there is no other text information in the text input area before the at least one candidate text information is generated, after the second text information is determined, the second text information is displayed in the text input area 31, and the candidate text area 61 is deleted.

[0084] In S205, a third text is generated based on the second text information.

[0085] In one embodiment of the present disclosure, when fourth text information is displayed in the text input area before the at least one candidate phrase information is generated, after the second text information is determined, fifth text information in which the second text information and the fourth text information are merged is displayed in the text input area.

[0086] In an embodiment of the present disclosure, the fourth text information may be text information manually input by the user, may be the second text information determined in steps S201 to S204, or may be text information edited and changed by the user.

[0087] In an embodiment of the present invention, as shown in FIG. 7b, before the at least one candidate phrase information is generated, fourth text information (#) is displayed in the text input area 31, and when the user input position is at the end of the fourth text information, the candidate phrase area 61 is displayed at the end of the fourth text information. In response to a trigger operation by the user on the text confirmation control, the candidate phrase area is deleted, and the second text information (AAAAAA) is joined to the end of the fourth text information (#), so that, as shown in FIG. 7b, fifth text information (#AAAAAA) is formed. In FIG. 7b, the case where the user input position is at the end of the fourth text information is taken as an example for explanation.

[0088] In one embodiment of the present disclosure, when the user input position is at an intermediate position of the fourth text information, in the text input area, the fourth text information is separated from the intermediate position by the candidate phrase area and displayed on both sides of the candidate phrase area. The second text information in the fifth text information is inserted at the intermediate position of the fourth text information.

[0089] In an embodiment of the present disclosure, as shown in FIG. 7c, before the at least one candidate phrase information is generated, fourth text information (#) is displayed in the phrase input area 31, and when the user input position is at the middle position of the fourth text information, the candidate phrase area 61 separates the fourth text information in the phrase input area 31 from the user input position, and the two separated portions of the fourth text information are respectively displayed on both sides of the candidate phrase area 61. Further, in response to a trigger operation by the user on the phrase confirmation control, the candidate phrase area is deleted, and the second text information (AAAAAA) is inserted at the middle position of the fourth text information (#), and as shown in FIG. 7c, fifth text information (AAAAAA#) is formed. In FIG. 7c, the case where the user input position is at the middle position of the fourth text information is taken as an example for explanation.

[0090] In one embodiment of the present disclosure, it further includes obtaining the third text information by editing the fifth text information in response to a user input operation.

[0091] In one embodiment of the present disclosure, as shown in FIGS. 7a, 7b, and 7c, an intelligent phrase generation control is included in the phrase input area 31. In response to a trigger operation on the intelligent phrase generation control, a page as shown in FIG. 3 including confirmed text information is displayed. In response to an operation on the phrase editing screen as shown in FIG. 3, generation of new second text information is started.

[0092] In one embodiment of the present disclosure, in response to a user editing operation on the phrase input area 31, the fifth text information is edited to obtain the third text information. The editing includes operations such as input, deletion, copy, and paste.

[0093] In one embodiment of the present disclosure, in response to a trigger operation on the "complete" control in the phrase input area 31, the phrase input screen is closed, and the third text information is displayed in the text creation area 21 (as shown in FIG. 2).

[0094] In S206, multimedia editing data is generated based on the third text information.

[0095] In S207, a target video is generated based on the multimedia editing data.

[0096] Steps S206 to S207 according to the embodiments of the present disclosure are the same as the specific execution flows in steps S103 to S104 according to the above embodiments. Therefore, specifically, reference may be made to the description in the above embodiments, and the details thereof are omitted in the embodiments of the present disclosure.

[0097] FIG. 8 is a schematic configuration diagram of a video generation device according to an embodiment of the present disclosure. This embodiment is applicable when generating a video from input text. The video generation device can be realized in a software and / or hardware manner.

[0098] As shown in FIG. 8, the video generation device 80 according to the embodiment of the present disclosure mainly includes a first text information acquisition module 81, a second text information generation module 82, a multimedia editing data generation module 83, and a target video generation module 84.

[0099] The first text information acquisition module 81 is used to acquire the first text information. It generates multimedia editing data based on the third text information. The second text information generation module 82 is used to generate the second text information based on the first text information. The second text information is literal information that conforms to the creation requirements described in the first text information. The multimedia editing data generation module 83 is used to generate multimedia editing data based on the third text information obtained based on the second text information. The multimedia editing data includes at least one video track clip and at least one audio track clip. The at least one video track clip and the at least one audio track clip respectively correspond to at least one text clip partitioned by the third text information. The target audio track clip is used to fill in the voiceover that matches the target text clip. The target video track clip in the at least one video track clip and the target audio track clip occupy the same timeline position on the video editing timeline. The target video generation module 84 is used to generate a target video based on the multimedia editing data.

[0100] In one embodiment of the present disclosure, the second text information generation module 82 includes a candidate phrase information generation unit, a candidate phrase information switching unit, and a candidate phrase information confirmation unit. The candidate phrase information generation unit is used to generate at least one candidate phrase information based on the first text information. All of the at least one candidate phrase information conforms to the creation requirements indicated by the first text information. The candidate phrase information switching unit is used to switch and display different candidate phrase information among the at least one candidate phrase information on the phrase input screen in response to a switching operation triggered by a user. The candidate phrase information confirmation unit determines, in response to a confirmation operation triggered by a user, the candidate phrase information that is switched and displayed on the phrase input screen among the at least one candidate phrase information as the second text information.

[0101] In one embodiment of the present disclosure, the phrase input screen includes a phrase input area and a candidate phrase area. The phrase input screen includes a phrase switching control. The phrase switching control is used to trigger the switching operation so that different candidate phrase information among the at least one candidate phrase information is switched and displayed in the candidate phrase area. The phrase input screen includes a phrase confirmation control. The phrase confirmation control is used to trigger the confirmation operation so that the candidate phrase information presented in the candidate phrase area is determined as the second text information and the second text information is displayed in the phrase input area.

[0102] In one embodiment of the present disclosure, after the at least one candidate phrase information is generated, the candidate phrase information displayed in the candidate phrase area is inserted into the user input position in the phrase input area. When the confirmation operation is responded to, the candidate phrase area is deleted from within the phrase input area.

[0103] In one embodiment of the present disclosure, when fourth text information is displayed in the text input area before the at least one candidate text information is generated, after the second text information is determined, fifth text information in which the second text information and the fourth text information are fused is displayed in the text input area.

[0104] In one embodiment of the present disclosure, when the user input position is at an intermediate position of the fourth text information, the fourth text information is separated from the intermediate position by the candidate text area and displayed on both sides of the candidate text area in the text input area, and the second text information in the fifth text information is inserted at the intermediate position of the fourth text information.

[0105] In one embodiment of the present disclosure, in response to a user input operation, the fifth text information is edited to obtain the third text information.

[0106] In one embodiment of the present disclosure, the apparatus further includes a target text category determination module. The target text category determination module is used to identify a target text category from a plurality of text categories. The candidate text information is generated based on the target text category and the first text information. The text category of the second text information is the target text category.

[0107] In one embodiment of the present disclosure, the text input screen includes a first text category control and a second text category control. The first text category control is used to set the text category corresponding to the first text category control as the target text category in response to a trigger operation by the user. The second text category control is used to set the text category corresponding to the second text category control as the target text category in response to a trigger operation by the user.

[0108] In one embodiment of the present disclosure, the target video track clip is an empty clip. Alternatively, the target video track clip is used to fill a video image that matches the target text clip. Alternatively, the target video track clip is used to fill an expression image that matches the target text clip.

[0109] The video generation device according to an embodiment of the present disclosure can execute the steps executed in the video generation method according to the method embodiment of the present disclosure, and has the execution steps and beneficial effects, and the description thereof is omitted here.

[0110] FIG. 9 is a schematic configuration diagram of an electronic device in an embodiment of the present disclosure. Hereinafter, FIG. 9 showing a schematic configuration diagram suitable for realizing the electronic device 900 in the embodiment of the present disclosure will be specifically referred to. The electronic device 900 in the embodiment of the present disclosure may include, for example, portable terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (tablets), PMPs (Portable Multimedia Players), in-vehicle terminals (for example, car navigation terminals), wearable terminal devices, etc., and fixed terminals such as digital TVs, desktop computers, intelligent home devices, etc., but is not limited thereto. The electronic device shown in FIG. 9 is only an example and does not limit the functions and usage ranges of the embodiments of the present disclosure.

[0111] As shown in FIG. 9, the electronic device 900 may include a processing device (e.g., a CPU, a graphics processor, etc.) 901 that executes various appropriate operations and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903 to implement the picture rendering method of the embodiments described in the present disclosure. Various programs and data necessary for the operation of the terminal device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0112] Normally, an input device 906 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc., an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc., a storage device 908 including, for example, a magnetic tape, a hard disk, etc., and a communication device 909 may be connected to the I / O interface 905. The communication device 909 may allow the terminal device 900 to communicate with other devices wirelessly or wiredly to exchange data. FIG. 9 shows a terminal device 900 having various devices, but it should be understood that it is not necessarily required to implement or include all of the shown devices. More or fewer devices may alternatively be implemented or included.

[0113] In particular, according to the embodiments of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product including a computer program mounted on a non-transitory computer-readable medium including program code for executing the method shown in the flowchart, thereby implementing the video generation method as described above. In such an embodiment, the computer program may be downloaded and installed from a network by the communication device 909, or may be installed from the storage device 908, or may be installed from the ROM 902. When the computer program is executed by the processing device 901, the above functions limited to the method according to the embodiments of the present disclosure are executed.

[0114] Note that the computer-readable medium in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above, but is not limited thereto. More specific examples of the computer-readable storage medium may include an electrical connection having at least one wire, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical memory device, a magnetic memory device, or any suitable combination of the above, but is not limited thereto. In the present disclosure, the computer-readable storage medium may be any tangible medium that includes or stores a program that can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or propagated as part of a carrier wave on which a computer-readable program code is carried. Such a propagated data signal may take various forms, including electromagnetic signals, optical signals, or any suitable combination of the above, but is not limited thereto. Also, the computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code included in the computer-readable medium may be transmitted via any suitable medium, including wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above, but is not limited thereto.

[0115] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be connected to digital data communication of any form or medium (e.g., communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), extranets (e.g., the Internet), end-to-end networks (e.g., ad hoc end-to-end networks), and networks currently known or future-developed.

[0116] The computer-readable medium may be included in the electronic device or may exist separately without being incorporated into the electronic device.

[0117] The computer-readable medium stores one or more programs, and when the one or more programs are executed by the terminal device, the terminal device is caused to: obtain first text information for describing requirements for creating video language; generate second text information based on the first text information, where the second text information is language information that meets the creation requirements described in the first text information; generate multimedia editing data based on third text information obtained based on the second text information, where the multimedia editing data includes at least one video track clip and at least one audio track clip, the at least one video track clip and the at least one audio track clip respectively correspond to at least one text clip partitioned by the third text information, the target audio track clip in the at least one audio track clip is used to fill in a voiceover that matches the target text clip, the target video track clip in the at least one video track clip and the target audio track clip occupy the same timeline position on the video editing timeline; and generate a target video based on the multimedia editing data. When the one or more programs are executed by the terminal device, the terminal device may also execute other steps described in the above embodiments.

[0118] The computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computer, partially on the user computer, executed as an independent software package, partially executed on the user computer and partially on a remote computer, or entirely executed on a remote computer or server. When a remote computer is involved, the remote computer may be connected to the user computer via any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, connected via the Internet using an Internet service provider).

[0119] Flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation that can be implemented according to the systems, methods, and computer program products in various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent one module, program segment, or part of code that includes at least one executable instruction for implementing a given logic function. It should also be noted that in some realizations as an alternative, the functions represented within a block may occur in a different order than that shown in the drawings. For example, two blocks shown consecutively may actually be executed substantially in parallel or may sometimes be executed in the reverse order depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that executes a given function or operation, or may be implemented by a combination of dedicated hardware and computer instructions.

[0120] The units according to the embodiments of the present disclosure described above may be implemented in software or in hardware. The name of a unit may not be a limitation to the unit itself in some cases.

[0121] The functions described above in this specification may be executed, at least in part, by at least one hardware logic component. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0122] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or stores a program used by or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium may include, but are not limited to, electrical connections based on at least one wire, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] The above description is only for the better embodiments of the present disclosure and the applicable technical principles. Those skilled in the art will understand that the scope of the disclosure according to the present disclosure is not limited to the technical solutions consisting of specific combinations of the above technical features, and other technical solutions consisting of any combination of the above technical features or their equivalent features should also be included without departing from the idea of the above disclosure. For example, a technical solution formed by replacing each other the above features with technical features having similar functions (but not limited thereto) disclosed in the present disclosure.

[0124] Also, although the operations are depicted in a particular order, they should not be construed as requiring that the operations be performed in the particular order shown or sequentially. In some environments, multitasking or parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Some features described in the context of individual embodiments may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented separately or in any suitable sub-combination in multiple embodiments.

[0125] The subject matter has been described in language specific to structural features and / or methodological and logical acts. It is understood, however, that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely exemplary forms of implementing the claims.

Claims

1. Obtaining first text information for describing the creation requirements of video language; Generating second text information based on the first text information, where the second text information is language information that conforms to the creation requirements described in the first text information; Generating multimedia editing data based on third text information obtained based on the second text information, where the multimedia editing data includes at least one video track clip and at least one audio track clip, and the at least one video track clip and the at least one audio track clip respectively correspond to at least one text clip partitioned by the third text information. The target audio track clip in the at least one audio track clip is used to fill in the reading voice that matches the target text clip, and the target video track in the at least one video track clip and the target audio track clip occupy the same timeline position on the video editing timeline; Generating a target video based on the multimedia editing data; A video generation method characterized by including the above.

2. Generating the second text information based on the first text information includes: Generating at least one candidate language information based on the first text information, where all of the at least one candidate language information conforms to the creation requirements indicated by the first text information; In response to a switching operation triggered by the user, switching and displaying different candidate language information among the at least one candidate language information on the language input screen; In response to a confirmation operation triggered by the user, determining, among the at least one candidate language information, the candidate language information that is switched and displayed on the language input screen as the second text information. The method according to claim 1, characterized in that it includes

3. The text input screen includes a text input area and a candidate text area, The text input screen includes a text switching control, and the text switching control is used to trigger the switching operation so that different candidate text information among the at least one candidate text information is switched and displayed in the candidate text area. The text input screen includes a text confirmation control, and the text confirmation control is used to trigger the confirmation operation so that the candidate text information displayed in the candidate text area is determined as the second text information and the second text information is displayed in the text input area. The method according to claim 2, characterized in that

4. After the at least one candidate text information is generated, the candidate text information displayed in the candidate text area is inserted at the user input position in the text input area. When the confirmation operation is responded to, the candidate text area is deleted from within the text input area. The method according to claim 3, characterized in that

5. When fourth text information is displayed in the text input area before the at least one candidate text information is generated, after the second text information is determined, fifth text information in which the second text information and the fourth text information are fused is displayed in the text input area. The method according to claim 4, characterized in that

6. When the user input position is at an intermediate position of the fourth text information, in the text input area, the fourth text information is separated from the intermediate position by the candidate text area and displayed on both sides of the candidate text area, and the second text information in the fifth text information is inserted at the intermediate position of the fourth text information. The method according to claim 5, characterized in that.

7. Further comprising editing the fifth text information according to a user input operation to obtain the third text information. The method according to claim 5, characterized in that.

8. Identifying a target language category from a plurality of language categories, wherein the candidate language information is generated based on the target language category and the first text information, and the language category of the second text information is the target language category. The method according to claim 2, characterized in that.

9. The language input screen includes a first language category control and a second language category control. The first language category control is used to set the language category corresponding to the first language category control as the target language category in response to a user trigger operation. The second language category control is used to set the language category corresponding to the second language category control as the target language category in response to a user trigger operation. The method according to claim 8, characterized in that.

10. The target video clip is an empty clip, or The target video clip is used to fill a video image matching the target text clip, or The target video clip is used to fill an expression image matching the target text clip. The method according to claim 1, characterized in that.

11. A first text information acquisition module for acquiring first text information for describing the creation requirements of video language. A second text information generation module for generating second text information based on the first text information, wherein the second text information is a second text information generation module that is text information that conforms to the creation requirements described in the first text information, A multimedia editing data generation module for generating multimedia editing data based on third text information obtained based on the second text information, wherein the multimedia editing data includes at least one video track clip and at least one audio track clip, and the at least one video track clip and the at least one audio track clip respectively correspond to at least one text clip partitioned by the third text information. The target audio track clip in the at least one audio track clip is used to fill in the voiceover that matches the target text clip, and the target video track in the at least one video track clip and the target audio track clip occupy the same timeline position on the video editing timeline. Multimedia editing data generation module, A target video generation module for generating a target video based on the multimedia editing data, A video generation device, characterized by comprising the above.

12. At least one processor, A storage device storing at least one program, and comprising, When the at least one program is executed by the at least one processor, the at least one processor realizes the method according to any one of claims 1 to 10, An electronic device, characterized by the above.

13. A computer-readable storage medium storing a computer program, wherein when the program is executed by a processor, the method according to any one of claims 1 to 10 is realized, A computer-readable storage medium. **Claim 14** A computer program product comprising a computer program or instructions which, when executed by a processor, implement the method according to any one of claims 1 to 10, characterized in that.

Citation Information

Patent Citations

  • Video generation method and device, computer equipment and storage medium

    CN114513706A