Video generation method, device, equipment, storage medium, and program product

The method and device automate video creation by generating text and multimedia data from user keywords, addressing inefficiencies in video production and enhancing the creation process through synchronized audio and video tracks.

JP7782941B2Active Publication Date: 2025-12-09BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023578865
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-04-19
Filing Date
2023-12-12
Publication Date
2025-12-09
Estimated Expiration
2043-12-12

AI Technical Summary

Technical Problem

Users face inefficiency in creating videos due to the time-consuming process of searching for video materials such as images and music before sharing their work on video platforms.

Method used

A method and device that automatically generate text and videos based on user input keywords, using intelligent algorithms to create multimedia editing data with synchronized audio and video tracks, enabling efficient one-stop video creation.

Benefits of technology

Enhances video creation efficiency by automating the process, reducing the time required for material search and editing, and allowing for personalized video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007782941000001
    Figure 0007782941000001
  • Figure 0007782941000002
    Figure 0007782941000002
  • Figure 0007782941000003
    Figure 0007782941000003
Patent Text Reader

Abstract

The present disclosure relates to a video generation method, apparatus, device, storage medium, and program product. The method includes obtaining first text information for describing the creation requirements of video captions, generating second text information based on the first text information, where the second text information is caption information that meets the creation requirements described in the first text information, generating multimedia editing data based on third text information, and generating a target video based on the multimedia editing data. In an embodiment of the present invention, based on the creation requirements of video captions, caption information that meets the described creation requirements is generated, and further, a video is created using the generated caption information, thereby providing an efficient one-stop video creation solution and improving the video creation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the technical field of video processing, and in particular to a video generating method, apparatus, device, storage medium and program product. [Background technology]

[0002] With the rapid development of computer technology and mobile communication technology, various video platforms based on electronic devices have become commonplace and have greatly enriched people's daily lives. More and more users enjoy sharing their own video works on video platforms for other users to view.

[0003] Before sharing their video work, users must edit and create the video themselves. When creating a video, they must search for a large amount of video materials, such as images, text, and music. This takes a long time and leads to inefficiency in video creation. Summary of the Invention [Problem to be solved by the invention]

[0004] In order to solve the above technical problems, embodiments of the present invention provide a video generation method, device, equipment, storage medium and program product that automatically generates text based on keywords entered by a user and automatically generates a video based on the generated text, thereby providing an efficient one-stop video creation solution and improving video creation efficiency. [Means for solving the problem]

[0005] In a first aspect, an embodiment of the present disclosure provides a video generation method, the video generation method comprising: Obtaining first text information describing video copy creation requirements; generating second text information based on the first text information, wherein the second text information is wording information that meets creation requirements described in the first text information; generating multimedia editing data based on third text information obtained based on the second text information, the multimedia editing data including at least one video track clip and at least one audio track clip, the at least one video track clip and the at least one audio track clip respectively corresponding to at least one text clip partitioned by the third text information, the target audio track clip being used to fill in a reading audio matching the target text clip, and the target video track clip and the target audio track clip in the at least one video track clip occupying the same timeline position on a video editing timeline; generating a target video based on the multimedia editing data; Includes.

[0006] In a second aspect, an embodiment of the present disclosure provides an animation production device, the animation production device comprising: a first text information acquisition module for acquiring first text information, the first text information acquisition module generating multimedia editing data based on the third text information; a second text information generation module for generating second text information based on the first text information, the second text information being wording information that meets creation requirements described in the first text information; a multimedia editing data generation module for generating multimedia editing data based on third text information obtained based on the second text information, the multimedia editing data including at least one video track clip and at least one audio track clip, the at least one video track clip and the at least one audio track clip each corresponding to at least one text clip partitioned by the third text information, the target audio track clip being used to fill in a reading audio that matches the target text clip, and the target video track clip and the target audio track clip in the at least one video track clip occupying the same timeline position on a video editing timeline; a target video generation module for generating a target video based on the multimedia editing data; Equipped with.

[0007] In a third aspect, an embodiment of the present disclosure provides an electronic device, the electronic device comprising: at least one processor; and a storage device that stores at least one program. When the at least one program is executed by the at least one processor, the at least one program causes the at least one processor to realize the moving image generating method according to any one of the first aspects.

[0008] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, the computer program implementing the video generation method according to any one of the first aspects when executed by a processor.

[0009] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, the computer program product including computer programs or instructions that, when executed by a processor, implement the video generation method according to any one of the first aspects.

[0010] An embodiment of the present disclosure provides a video generation method, an apparatus, a storage medium, and a program product, the method including: acquiring first text information describing production requirements for video text; generating second text information based on the first text information, the second text information being text information that matches the production requirements described in the first text information; generating multimedia editing data based on third text information obtained based on the second text information, the multimedia editing data including at least one video track clip and at least one audio track clip, the at least one video track clip and the at least one audio track clip each corresponding to at least one text clip defined by the third text information, the target audio track clip being used to fill in a reading audio that matches the target text clip, the target video track clip and the target audio track clip of the at least one video track clip occupying the same timeline position on a video editing timeline; and generating a target video based on the multimedia editing data. In an embodiment of the present invention, based on the production requirements of the video text, text information of the described production requirements is generated, and then the target video is created according to the generated text information, thereby providing an efficient one-stop video production method and improving the efficiency of video production. [Brief explanation of the drawings]

[0011] These and other features, advantages, and aspects of each embodiment of the present disclosure will become more apparent with reference to the following specific embodiments in conjunction with the drawings. Identical or similar reference numerals represent identical or similar elements throughout the drawings. It should be understood that the drawings are schematic and that objects and elements are not necessarily drawn to scale.

[0012] [Figure 1] FIG. 1 is a flow diagram of a video generation method according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a schematic diagram of a video creation page according to an embodiment of the present disclosure. [Figure 3] FIG. 10 is a schematic diagram of a text input screen according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a schematic diagram of a multimedia editing page according to an embodiment of the present disclosure. [Figure 5] FIG. 1 is a flow diagram of a video generation method according to an embodiment of the present disclosure. [Figure 6] FIG. 10 is a schematic diagram of a text input screen according to an embodiment of the present disclosure. [Figure 7a] FIG. 10 is a schematic diagram of a text input screen according to an embodiment of the present disclosure. [Figure 7b] FIG. 10 is a schematic diagram of a text input screen according to an embodiment of the present disclosure. [Figure 7c] FIG. 10 is a schematic diagram of a text input screen according to an embodiment of the present disclosure. [Figure 8] 1 is a schematic diagram illustrating the configuration of a moving image generating device according to an embodiment of the present invention. [Figure 9] FIG. 1 is a schematic diagram illustrating the configuration of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the drawings show several embodiments of the present disclosure, it should be understood that the present disclosure can be realized in various forms and should not be construed as being limited to the embodiments described herein, but rather, these embodiments are provided to provide a clearer and more complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are used for illustration purposes only and are not used to limit the scope of protection of the present disclosure.

[0014] It should be understood that the steps described in the method embodiments of the present disclosure may be performed in a different order and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the performance of steps shown. The scope of the present disclosure is not limited in this respect.

[0015] As used herein, the term "comprises" and variations thereof are open-ended inclusions, i.e., "including, but not limited to." The term "based on" means "based at least in part on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one other embodiment," and the term "some embodiments" means "at least some embodiments." Relevant definitions of other terms are provided in the description below.

[0016] It should be noted that the concepts of "first," "second," etc. referred to in this disclosure are used only to distinguish between different devices, modules, or units, and do not limit the order or interdependence of functions performed by these devices, modules, or units.

[0017] It will be appreciated by those skilled in the art that the modifications "one" and "multiple" referred to in this disclosure are exemplary rather than limiting, and should be understood as "at least one" unless the context clearly indicates otherwise.

[0018] The names of messages or information exchanged between devices in the embodiments of the present disclosure are used for explanation purposes only and are not used to limit the scope of these messages or information.

[0019] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings, in which it should be noted that the same reference numerals in different drawings are used to refer to the same elements depicted.

[0020] FIG. 1 is a flowchart of a video generation method according to an embodiment of the present disclosure. This embodiment is applicable to generating videos from keywords. The method can be executed by a video generation device. The video generation device can be implemented in software and / or hardware. The video generation method can be applied to electronic devices.

[0021] It will be understood that the above electronic devices may be any other type of electronic device capable of performing data processing, and may include, but are not limited to, a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof.

[0022] As shown in FIG. 1, the moving image generating method according to the embodiment of the present invention mainly includes steps S101 to S104.

[0023] In S101, first text information for describing requirements for creating video text is acquired.

[0024] In one embodiment of the present disclosure, the creation requirements for the video text may be key content of the target video that the user wants to generate. Specifically, the first text information may be one or more keywords that describe the video text, or one or more topic words.

[0025] In one exemplary illustration, a user wants to create a video about "Movie Recommendations." The first text information may include "Movie," "2023," "Award Winning," "Popular Reviews," etc. A user wants to create a video about "Mobile Phone Selling Points." The first text information may include "Mobile Phone Model," "Super Large Screen," "Excellent Battery Life," "Affordable Price," etc.

[0026] In one embodiment of the present disclosure, acquiring first text information includes acquiring first text information corresponding to a user's input operation in response to the user's input operation. Specifically, a text creation video control is presented on a video creation front page, and a video creation screen is displayed in response to a trigger operation on the text creation video control. As shown in FIG. 2, the video creation screen includes a text creation area 21, a video category selection area 22, and a video creation control 23. The text creation area 21 includes a text display area and an intelligent text generation control. The text display area is for displaying third text information for generating multimedia editing data. The intelligent text generation control is for displaying a wording input screen in response to a user's trigger operation.

[0027] In one embodiment of the present disclosure, as shown in FIG. 3, the text input page includes a text input area 31 that displays text information input by the user in response to the user's input operation.

[0028] As shown in FIG. 3, the text input area 31 includes an input box for acquiring first text information in response to an input operation by the user.

[0029] In S102, second text information is generated based on the first text information, and the second text information is wording information that matches the creation requirements described in the first text information.

[0030] In the embodiments of the present disclosure, the second text information refers to wording information that meets production requirements. Furthermore, the second text information is a sentence that has context and is understandable. The second text information may have one paragraph or multiple paragraphs. The second text information may be generated based on the first text information by an intelligent wording generation algorithm. A specific embodiment of the intelligent wording generation algorithm will not be further described in the embodiments of the present disclosure.

[0031] As shown in FIG. 3, in response to a confirmation operation on the input box in the text input area 31, second text information is generated based on the first text information, and the generated second text information is displayed in the text input area 31.

[0032] At S103, multimedia editing data is generated based on third text information obtained based on the second text information. The multimedia editing data includes at least one video track clip and at least one audio track clip. The at least one video track clip and the at least one audio track clip each correspond to at least one text clip defined by the third text information. A target audio track clip in the at least one audio track clip is used to fill in a reading audio that matches the target text clip. The target video track clip and the target audio track clip in the at least one video track clip occupy the same timeline position on a video editing timeline.

[0033] In one embodiment of the present disclosure, the third text information may include one or more combinations of one or more second text information, edited second text information, and other text information edited by a user.

[0034] In one embodiment of the present disclosure, steps S101 to S102 may be performed multiple times to generate multiple pieces of second text information, the generated second text information may be edited and changed, or the text information may be manually entered into the text input area 31.

[0035] In one embodiment of the present disclosure, as shown in FIG. 3, in response to a trigger operation on the “Complete” control in the text input area 31, the video creation screen is switched to the video creation screen shown in FIG. 2, and third text information is displayed on the video creation screen.

[0036] In one embodiment of the present disclosure, when third text information is displayed on the video creation screen, in response to a trigger operation on the video generation control 23, multimedia editing data is generated based on the third text information, and the multimedia editing data is displayed on the multimedia editing screen.

[0037] In one embodiment of the present disclosure, the multimedia editing data further includes at least one subtitle track clip for filling subtitle information matching the target text clip.

[0038] In one embodiment of the present disclosure, the multimedia editing data further includes at least one background music track clip for filling in background music.

[0039] 4, the multimedia editing screen may include a video preview area 41 for previewing the target video generated by the multimedia editing data, and a multimedia track area 42. The multimedia editing data displayed in the multimedia track area 42 may include video track clips, audio track clips, and background music track clips.

[0040] In one embodiment of the present disclosure, a target audio track clip of the at least one audio track clip is used to fill in a reading audio that matches the target text clip, and the target video track clip and the target audio track clip of the at least one video track clip occupy the same timeline position on a video editing timeline.

[0041] In one embodiment of the present disclosure, the target text clip may be a text clip segmented by third text information. The text clip may be a sentence, an incomplete sentence constructed according to language segmentation rules, or multiple sentences, and is not particularly limited in the embodiments of the present disclosure. The target video track clip refers to a video clip corresponding to the target text clip. The target audio track clip refers to an audio clip corresponding to the target text clip.

[0042] Note that the background music in the background music track clip is not separated by text clips but is filled in as a complete audio file.

[0043] In one embodiment of the present disclosure, the target video track clip is a blank clip.

[0044] In an embodiment of the present disclosure, as shown in Figure 2, in response to a trigger operation on the "First Video Category" control on the video creation screen, a video track clip in the multimedia editing data is set to an empty clip, i.e., no picture or video image is added to the video track clip.

[0045] Furthermore, a picture selected by the user is added to the blank clip according to the user's operation on the blank clip. In other words, the user may fill the blank clip with a picture selected by the user so that the created and completed target video better matches the user's wishes.

[0046] In one embodiment of the present disclosure, the target video track clip is used to fill the target text clip with matching video images.

[0047] In an embodiment of the present disclosure, as shown in Figure 2, in response to a trigger operation on the "Second Video Category" control on the video creation screen, a video image matching the text clip is filled into a video track clip in the multimedia editing data, which can be obtained based on matching the text clip in a pre-set picture database according to an image matching algorithm.

[0048] In one embodiment of the present disclosure, video images may be automatically matched in a picture database based on text content, which avoids users from having to spend time searching for materials and further improves video creation efficiency.

[0049] In one embodiment of the present disclosure, the target video track clip is used to fill the target text clip with matching facial imagery.

[0050] In an embodiment of the present disclosure, as shown in Figure 2, in response to a trigger operation on the "Third Video Category" control on the video creation screen, a facial expression image matching a text clip is filled into a video track clip in the multimedia editing data. The facial expression image can be obtained based on matching the text clip in a preset facial expression image database according to a preset algorithm.

[0051] In one embodiment of the present disclosure, facial images may be automatically matched in a facial image database based on text content, thereby creating more personalized videos based on text content.

[0052] In an embodiment of the present disclosure, by separating the facial expression image database and the picture database into two databases, the problem of inconsistent video style caused by the presence of both normal pictures and facial expression images in a single video can be avoided.

[0053] At S104, a target moving image is generated based on the multimedia editing data.

[0054] In one embodiment of the present disclosure, a target movie is generated based on the multimedia editing data in response to a movie completion trigger operation.

[0055] In one embodiment of the present disclosure, the trigger operation of "complete" the video may refer to a trigger operation on an "export" control on a multimedia editing screen. The export method may be to save the target video locally or to share the target video on another video sharing platform or website. The embodiment of the present disclosure is not specifically limited.

[0056] In one embodiment of the present disclosure, in response to triggering an "import and edit" control on a multimedia editing screen, multimedia editing data is imported into a video editor for subsequent editing of the multimedia editing data.

[0057] In addition to the above embodiments, the embodiments of the present disclosure further optimize the video generation method. As shown in Figure 5, the optimized video generation method mainly includes the following steps:

[0058] In S201, first text information for describing requirements for creating video text is acquired.

[0059] Step S201 in the embodiment of the present disclosure is the same as the specific execution flow of step S101 in the above embodiment, and for specific details, please refer to the description in the above embodiment, so details will be omitted in the embodiment of the present disclosure.

[0060] In S202, at least one piece of candidate wording information is generated based on the first text information, and each of the at least one piece of candidate wording information meets the creation requirements indicated by the first text information.

[0061] In one embodiment of the present disclosure, a target word category is identified from a plurality of word categories, the candidate word information is generated based on the target word category and the first text information, and the word category of the second text information is the target word category.

[0062] In one embodiment of the present disclosure, the phrase category refers to a category to which the formed phrase belongs. Specifically, the phrase category may include a first phrase category and a second phrase category. The first phrase category may be understood as a phrase category that a user generally applies to various topics. Specifically, the various topics include science and technology, economy, entertainment, etc. The second phrase category may be understood as a phrase category applied to product planning or product marketing. Specifically, the second phrase category may be an introduction to the selling points of a certain mobile phone or an introduction to the reasons for recommending a certain item.

[0063] In one embodiment of the present disclosure, the target wording categories may be determined based on a user's selection, and each target wording category corresponds to one intelligent word generation algorithm model.

[0064] In one embodiment of the present disclosure, the text input screen includes a "first text category" control and a "second text category" control. The "first text category" control sets the text category corresponding to the first text category control as a target text category in response to a trigger operation by the user. The "second text category" control sets the text category corresponding to the second text category control as a target text category in response to a trigger operation by the user.

[0065] As shown in Figure 3, a "first text category" control and a "second text category" control are displayed on the text input screen. In the embodiment of the present disclosure, different text category controls correspond to different prompt information, and different text category controls correspond to different intelligent text generation algorithm models.

[0066] In one embodiment of the present disclosure, in response to a selection operation on the "first wording category" control, presented information such as "Please scribble one line of words from the first category. The topic is..." is displayed in the "input box" in the wording input area 31. The user may insert first text information after the presented information. In response to a confirmation operation on the "input box", the first text information is acquired, and at least one candidate wording information is generated based on the first wording category and the first text information.

[0067] In one embodiment of the present disclosure, generating at least one candidate wording information based on a first wording category and first text information includes invoking a first intelligent wording generation algorithm corresponding to the first wording category based on the first wording category, and using the first intelligent wording generation algorithm to process the first text information to generate a plurality of candidate wording information.

[0068] In one embodiment of the present disclosure, in response to a selection operation on the "second wording category" control, suggested information such as "Please write one line of words from the second category. The product and its selling points are..." is displayed in the input box in the wording input area 31. The user may input first text information after the suggested information. In response to a confirmation operation on the "input box", the first text information is acquired, and at least one candidate wording information is generated based on the second wording category and the first text information.

[0069] In one embodiment of the present disclosure, generating at least one candidate wording information based on the second wording category and the first text information includes invoking a second intelligent wording generation algorithm corresponding to the second wording category based on the second wording category, and using the second intelligent wording generation algorithm to process the first text information to generate a plurality of candidate wording information.

[0070] It should be noted that the first intelligent word generation algorithm and the second intelligent word generation algorithm are two different intelligent word generation algorithms.

[0071] In one embodiment of the present disclosure, the basic network models used by the two intelligent word generation algorithms may be the same or different. The training methods of the two intelligent word generation algorithms may be the same or different. The training samples of the first intelligent word generation algorithm and the second intelligent word generation algorithm are different. The training samples of the first intelligent word generation algorithm are word information of a first word category. The training samples of the second intelligent word generation algorithm are word information of a second word category.

[0072] In S203, different candidate wording information from the at least one candidate wording information is switched and displayed on the wording input screen in response to a switching operation triggered by the user.

[0073] In one embodiment of the present disclosure, the text input screen includes a candidate text area for displaying candidate text information. Specifically, the candidate text area is displayed in the text input area 31 in an inserted form.

[0074] In one embodiment of the present disclosure, the text input screen includes a text "switch" control for triggering the switching operation so that different candidate text information from the at least one candidate text information is switched and presented in the candidate text area.

[0075] In one embodiment of the present disclosure, as shown in Fig. 6, the candidate text area is displayed in an inserted form in the text input area 31. The candidate text area 61 includes a "first text switch" control 62 and a "second text switch" control 63.

[0076] In one embodiment of the present disclosure, multiple candidate wording information are arranged in a set order. The first text switching control 62 is used to display candidate wording information that precedes the current candidate wording information in the arrangement order in the candidate text area in response to a user's trigger operation. The second text switching control 63 is used to display candidate wording information that follows the current candidate wording information in the column order in the candidate text area in response to a user's trigger operation.

[0077] In one embodiment of the present disclosure, an example will be described in which there are five pieces of candidate wording information. The five pieces of candidate wording information are, in order, candidate wording information A, candidate wording information B, candidate wording information C, candidate wording information D, and candidate wording information E. The candidate wording information A, which has the highest display order, is displayed in the candidate wording area. In response to a trigger operation on the second text switching control 63, candidate wording information B is displayed in the candidate wording area. At this time, in response to a trigger operation on the first text switching control 62, candidate wording information A is displayed in the candidate wording area.

[0078] In S204, in response to a confirmation operation triggered by the user, the candidate wording information that is switched and displayed on the wording input screen from among the at least one candidate wording information is determined as the second text information.

[0079] In one embodiment of the present disclosure, the text input screen includes a text confirmation control, which is used to trigger the confirmation operation so that one of the at least one candidate text information items that is switched and displayed in the candidate text area is determined as the second text information item and the second text information item is displayed in the text input area.

[0080] 6, the candidate text area 61 includes a text confirmation control. In response to a trigger operation on the text confirmation control, the candidate text information displayed in the candidate text area is determined as the second text information, and the second text information is displayed in the text input area 31.

[0081] In one embodiment of the present disclosure, after the at least one candidate wording information is generated, the candidate wording information displayed in the candidate wording area is inserted into the user input position in the word input area, and when the confirmation operation is responded to, the candidate wording area is deleted from the word input area.

[0082] 6, the candidate text information (AAAAA) displayed in the candidate text area 61 is inserted at a user input position in the text input area 31. The user input position is the position where the cursor was located before the candidate text information was generated. Furthermore, the candidate text area is deleted in response to a trigger operation by the user on the text confirmation control.

[0083] In one embodiment of the present disclosure, as shown in FIG. 7a, if there is no other text information in the text input area before the at least one candidate text information is generated, after the second text information is determined, the second text information is displayed in the text input area 31 and the candidate text area 61 is deleted.

[0084] In S205, a third text is generated based on the second text information.

[0085] In one embodiment of the present disclosure, if a fourth text information is displayed in the text input area before the at least one candidate text information is generated, after the second text information is determined, a fifth text information that is a fusion of the second text information and the fourth text information is displayed in the text input area.

[0086] In an embodiment of the present disclosure, the fourth text information may be text information manually entered by the user, may be the second text information determined in steps S201 to S204, or may be text information edited or changed by the user.

[0087] In an embodiment of the present invention, as shown in Fig. 7b, when fourth text information (#######) is displayed in the text input area 31 before the at least one candidate text information is generated and the user input position is at the end of the fourth text information, a candidate text area 61 is displayed at the end of the fourth text information. In response to a trigger operation by the user on the text confirmation control, the candidate text area is deleted and the second text information (AAAAAA) is spliced ​​to the end of the fourth text information (#######), thereby forming fifth text information (########AAAAAA), as shown in Fig. 7b. In Fig. 7b, a case where the user input position is at the end of the fourth text information will be described as an example.

[0088] In one embodiment of the present disclosure, when the user input position is at an intermediate position of the fourth text information, the fourth text information is separated from the intermediate position by the candidate text area in the text input area and displayed on both sides of the candidate text area, and the second text information in the fifth text information is inserted at the intermediate position of the fourth text information.

[0089] In an embodiment of the present disclosure, as shown in FIG. 7c, when fourth text information (#######) is displayed in the text input area 31 before the at least one candidate text information is generated and the user input position is located in the middle of the fourth text information, the candidate text area 61 separates the fourth text information in the text input area 31 from the user input position, and displays the two separated parts of the fourth text information on both sides of the candidate text area 61. Furthermore, in response to a trigger operation by the user on the text confirmation control, the candidate text area is deleted, and the second text information (AAAAAA) is inserted in the middle of the fourth text information (#######), forming fifth text information (###AAAAAA####) as shown in FIG. 7c. An example will be described in FIG. 7c where the user input position is located in the middle of the fourth text information.

[0090] In one embodiment of the present disclosure, the method further includes editing the fifth text information to obtain the third text information in response to an input operation by a user.

[0091] In one embodiment of the present disclosure, as shown in Figures 7a, 7b, and 7c, an intelligent text generation control is included in the text input area 31. In response to a trigger operation on the intelligent text generation control, a page including confirmed text information is displayed, as shown in Figure 3. In response to an operation on the text editing screen, as shown in Figure 3, generation of new second text information is initiated.

[0092] In one embodiment of the present disclosure, the fifth text information is edited to obtain the third text information in accordance with a user's editing operation on the text input area 31. The editing includes operations such as input, deletion, copy, and paste.

[0093] In one embodiment of the present disclosure, in response to a trigger operation on the "Done" control in the text input area 31, the text input screen is closed and the third text information is displayed in the text creation area 21 (as shown in FIG. 2).

[0094] In S206, multimedia editing data is generated based on the third text information.

[0095] In S207, a target moving image is generated based on the multimedia editing data.

[0096] Steps S206 to S207 in the embodiments of the present disclosure are the same as the specific execution flow of steps S103 to S104 in the above embodiments, so for specific details, please refer to the description in the above embodiments, and therefore details thereof will be omitted in the embodiments of the present disclosure.

[0097] 8 is a schematic diagram of a configuration of a video generation device according to an embodiment of the present disclosure. This embodiment is applicable to generating video from input text. The video generation device can be realized in software and / or hardware.

[0098] As shown in FIG. 8, the video generation device 80 according to an embodiment of the present disclosure mainly includes a first text information acquisition module 81, a second text information generation module 82, a multimedia editing data generation module 83, and a target video generation module 84.

[0099] The first text information acquisition module 81 is used to acquire first text information. Multimedia editing data is generated based on third text information. The second text information generation module 82 is used to generate second text information based on the first text information. The second text information is wording information that matches the production requirements described in the first text information. The multimedia editing data generation module 83 is used to generate multimedia editing data based on third text information obtained based on the second text information. The multimedia editing data includes at least one video track clip and at least one audio track clip. The at least one video track clip and the at least one audio track clip each correspond to at least one text clip defined by the third text information. The target audio track clip is used to fill in a reading audio that matches the target text clip. The target video track clip and the target audio track clip in the at least one video track clip occupy the same timeline position on the video editing timeline. The target video generation module 84 is used to generate a target video based on the multimedia editing data.

[0100] In one embodiment of the present disclosure, the second text information generation module 82 includes a candidate wording information generation unit, a candidate wording information switching unit, and a candidate wording information confirmation unit. The candidate wording information generation unit is used to generate at least one candidate wording information based on the first text information. The at least one candidate wording information all meets the creation requirements indicated by the first text information. The candidate wording information switching unit is used to switch between different candidate wording information of the at least one candidate wording information and display them on the text input screen in response to a switching operation triggered by a user. The candidate wording information confirmation unit determines, as the second text information, the candidate wording information of the at least one candidate wording information that is switched and displayed on the text input screen in response to a confirmation operation triggered by a user.

[0101] In one embodiment of the present disclosure, the text input screen includes a text input area and a candidate text area. The text input screen includes a text switching control. The text switching control is used to trigger the switching operation so that different candidate text information among the at least one candidate text information is switched and displayed in the candidate text area. The text input screen includes a text confirmation control. The text confirmation control is used to trigger the confirmation operation so that the candidate text information presented in the candidate text area is determined as the second text information and the second text information is displayed in the text input area.

[0102] In one embodiment of the present disclosure, after the at least one candidate wording information is generated, the candidate wording information displayed in the candidate wording area is inserted into the user input position in the word input area, and when the confirmation operation is responded to, the candidate wording area is deleted from the word input area.

[0103] In one embodiment of the present disclosure, if a fourth text information is displayed in the text input area before the at least one candidate text information is generated, after the second text information is determined, a fifth text information that is a fusion of the second text information and the fourth text information is displayed in the text input area.

[0104] In one embodiment of the present disclosure, when the user input position is at the intermediate position of the fourth text information, the fourth text information is separated from the intermediate position by the candidate text area in the text input area and displayed on both sides of the candidate text area, and the second text information in the fifth text information is inserted at the intermediate position of the fourth text information.

[0105] In one embodiment of the present disclosure, the fifth text information is edited in response to an input operation by a user to obtain the third text information.

[0106] In one embodiment of the present disclosure, the device further includes a target wording category determination module. The target wording category determination module is used to identify a target wording category from a plurality of wording categories. The candidate wording information is generated based on the target wording category and the first text information. The wording category of the second text information is the target wording category.

[0107] In one embodiment of the present disclosure, the text input screen includes a first text category control and a second text category control. The first text category control is used to set a text category corresponding to the first text category control as a target text category in response to a trigger operation by a user. The second text category control is used to set a text category corresponding to the second text category control as a target text category in response to a trigger operation by a user.

[0108] In one embodiment of the present disclosure, the target video track clip is an empty clip, or is used to fill the target text clip with matching video images, or is used to fill the target text clip with matching facial expression images.

[0109] The video generation device according to the embodiment of the present disclosure can execute the steps performed in the video generation method according to the method embodiment of the present disclosure, and has the execution steps and beneficial effects, the description of which will be omitted here.

[0110] FIG. 9 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. Hereinafter, specific reference will be made to FIG. 9 , which illustrates a schematic diagram of a configuration suitable for implementing electronic device 900 according to an embodiment of the present disclosure. Electronic device 900 according to an embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (PDAs), tablets (PADs), portable multimedia players (PMPs), in-vehicle terminals (e.g., car navigation terminals), and wearable terminals, as well as fixed terminals such as digital TVs, desktop computers, and intelligent home devices. The electronic device illustrated in FIG. 9 is merely an example and does not limit the functionality and scope of use of the embodiment of the present disclosure.

[0111] 9, electronic device 900 may include a processing unit (e.g., a CPU, a graphics processor, etc.) 901 that can perform various appropriate operations and processes in accordance with a program stored in read-only memory (ROM) 902 or a program loaded from a storage device 908 into random access memory (RAM) 903 to realize the picture rendering method of the embodiments described in this disclosure. RAM 903 also stores various programs and data necessary for the operation of terminal device 900. Processing unit 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.

[0112] Typically, input devices 906 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc., output devices 907 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc., storage devices 908 including, for example, a magnetic tape, hard disk, etc., and communication devices 909 may be connected to the I / O interface 905. The communication devices 909 may allow the terminal device 900 to communicate wirelessly or via wires with other devices to exchange data. While FIG. 9 shows the terminal device 900 having various devices, it should be understood that it is not necessary to implement or include all of the devices shown. More or fewer devices may alternatively be implemented or included.

[0113] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product including a computer program stored on a non-transitory computer-readable medium, the computer program including program code for executing the methods illustrated in the flowcharts, thereby implementing the video generation method described above. In such embodiments, the computer program may be downloaded and installed from a network via the communication device 909, installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the functions defined by the methods according to the embodiments of the present disclosure are performed.

[0114] It should be noted that the computer-readable medium in this disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above. The computer-readable storage medium may be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to, an electrical connection having at least one wire, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical memory device, a magnetic memory device, or any suitable combination of the above. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system, apparatus, or device. In this disclosure, the computer-readable signal medium may include a data signal, propagated in baseband or as part of a carrier wave, in which computer-readable program code is embedded. Such propagated data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Also, a computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained in a computer-readable medium may be transmitted over any suitable medium, including, but not limited to, wire, optical cable, RF (radio frequency), etc., or any suitable combination of the above.

[0115] In some embodiments, clients and servers may communicate using any network protocol now known or later developed, such as HTTP (Hyper Text Transfer Protocol), and may be connected to each other via any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), extranets (e.g., the Internet), end-to-end networks (e.g., ad-hoc end-to-end networks), and other networks now known or later developed.

[0116] The computer-readable medium may be included in the electronic device, or may exist separately from the electronic device.

[0117] The computer-readable medium stores one or more programs, which, when executed by the terminal device, cause the terminal device to perform the following operations: acquire first text information describing video text creation requirements; generate second text information based on the first text information, wherein the second text information is text information that matches the creation requirements described in the first text information; generate multimedia editing data based on third text information obtained based on the second text information, wherein the multimedia editing data includes at least one video track clip and at least one audio track clip, wherein the at least one video track clip and the at least one audio track clip each correspond to at least one text clip defined by the third text information; a target audio track clip in the at least one audio track clip is used to fill in reading audio that matches the target text clip; and the target video track clip and the target audio track clip in the at least one video track clip occupy the same timeline position on a video editing timeline; and generate a target video based on the multimedia editing data. When the one or more programs are executed by the terminal device, the terminal device may perform other steps described in the above embodiments.

[0118] Computer program code for carrying out the operations of the present disclosure may be written in one or more programming languages, or a combination thereof. Such programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, etc., as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may run entirely on the user computer, partially on the user computer, as a separate software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. When a remote computer is involved, the remote computer may be connected to the user computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0119] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operations that can be implemented according to systems, methods, and computer program products in various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, program segment, or portion of code, including at least one executable instruction for implementing a given logical function. It should also be noted that in some implementations, the functions noted in the blocks may occur in a different order than that shown in the figures. For example, two blocks shown in succession may actually be executed substantially in parallel or may be executed in the reverse order depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs a given function or operation, or in a combination of dedicated hardware and computer instructions.

[0120] The units according to the embodiments of the present disclosure described may be implemented in software or hardware, and the names of the units are not necessarily limitations on the units themselves in some cases.

[0121] The functions described herein above may be performed, at least in part, by at least one hardware logic component. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.

[0122] In the context of this disclosure, a machine-readable medium may be a tangible medium that can contain or store a program used by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of machine-readable storage media include an electrical connection based on at least one wire, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0123] The above description merely describes preferred embodiments and applied technical principles of the present disclosure. Those skilled in the art will understand that the scope of the present disclosure is not limited to a technical solution consisting of a specific combination of the above technical features, but should also include other technical solutions consisting of any combination of the above technical features or their equivalent features without departing from the idea of ​​the above disclosure. For example, a technical solution formed by mutually replacing the above features with technical features having functions similar to those disclosed in the present disclosure (but not limited thereto).

[0124] Additionally, although operations are depicted in a particular order, this should not be understood as requiring these operations to be performed in the particular order shown, or to be performed sequentially. In certain environments, multitasking or parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Some features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination.

[0125] Although the present subject matter has been described in language specific to structural features and / or methodological and logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. Obtaining first text information for describing video text creation requirements; generating second text information based on the first text information, the second text information being wording information that meets creation requirements described in the first text information; generating multimedia editing data based on third text information obtained based on the second text information, the multimedia editing data including at least one video track clip and at least one audio track clip, the at least one video track clip and the at least one audio track clip each corresponding to at least one text clip partitioned by the third text information, a target audio track clip in the at least one audio track clip being used to fill in a reading audio matching the target text clip, and the target video track clip and the target audio track clip in the at least one video track clip occupying the same timeline position on a video editing timeline; generating a target video based on the multimedia editing data; A moving image generating method comprising:

2. Generating second text information based on the first text information includes: generating at least one candidate wording information based on the first text information, wherein the at least one candidate wording information all meets the creation requirements indicated by the first text information; Switching and displaying different candidate wording information from the at least one candidate wording information on the wording input screen in response to a switching operation triggered by a user; determining, as the second text information, candidate wording information that is switched and displayed on the wording input screen from among the at least one candidate wording information in response to a confirmation operation triggered by a user; 2. The method of claim 1, comprising:

3. the text input screen includes a text input area and a candidate text area, the text input screen includes a text switching control, the text switching control being used to trigger the switching operation so that different candidate text information among the at least one candidate text information is switched and displayed in the candidate text area; the text input screen includes a text confirmation control, and the text confirmation control is used to trigger the confirmation operation so that candidate text information displayed in the candidate text area is determined as the second text information and the second text information is displayed in the text input area.

3. The method of claim 2.

4. After the at least one piece of candidate text information is generated, the candidate text information displayed in the candidate text area is inserted into a user input position in the text input area; When the confirmation operation is responded, the candidate text area is deleted from the text input area.

4. The method of claim 3.

5. When fourth text information is displayed in the text input area before the at least one candidate text information is generated, fifth text information in which the second text information and the fourth text information are combined is displayed in the text input area after the second text information is determined.

5. The method of claim 4.

6. When the user input position is at an intermediate position of the fourth text information, the fourth text information is separated from the intermediate position by the candidate text area in the text input area and displayed on both sides of the candidate text area, and the second text information in the fifth text information is inserted at the intermediate position of the fourth text information.

6. The method of claim 5.

7. and further including: editing the fifth text information in response to an input operation by a user to obtain the third text information.

6. The method of claim 5.

8. identifying a target wording category from a plurality of wording categories, wherein the candidate wording information is generated based on the target wording category and the first text information, and the wording category of the second text information is the target wording category; 3. The method of claim 2.

9. The text input screen includes a first text category control and a second text category control, the first text category control is used to set a text category corresponding to the first text category control as a target text category in response to a trigger operation by a user, and the second text category control is used to set a text category corresponding to the second text category control as a target text category in response to a trigger operation by a user.

9. The method of claim 8.

10. the target video track clip is an empty clip, or the target video track clip is used to fill the target text clip with matching video images; or the target video track clip is used to fill the target text clip with matching facial expression images; 2. The method of claim 1 .

11. a first text information acquisition module for acquiring first text information for describing video text creation requirements; a second text information generation module for generating second text information based on the first text information, the second text information being wording information that matches creation requirements described in the first text information; a multimedia editing data generation module for generating multimedia editing data based on third text information obtained based on the second text information, the multimedia editing data including at least one video track clip and at least one audio track clip, the at least one video track clip and the at least one audio track clip each corresponding to at least one text clip partitioned by the third text information, a target audio track clip in the at least one audio track clip being used to fill in a reading audio matching the target text clip, and the target video track clip and the target audio track clip in the at least one video track clip occupying the same timeline position on a video editing timeline; a target video generation module for generating a target video based on the multimedia editing data; A moving image generating device comprising:

12. at least one processor; a storage device that stores at least one program; said at least one program, when executed by said at least one processor, causing said at least one processor to implement the method of any one of claims 1 to 10; An electronic device characterized by:

13. A computer-readable storage medium having stored thereon a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 10. A computer-readable storage medium.

14. When executed by a processor, it implements the method of any one of claims 1 to 10. A computer program characterized by:

Citation Information

Patent Citations

  • Video generation method and device, computer equipment and storage medium

    CN114513706A