Video generation method, device, equipment, storage medium and program product

By automatically generating videos, based on the keywords input by the user, copywriting information that meets the writing requirements is generated, and combined with multimedia editing data, efficient one-stop video production is achieved, solving the problem of low video production efficiency caused by users' search for materials.

CN118828105BActive Publication Date: 2025-09-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310424794.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2025-09-30
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

Before sharing their video works, users need to find a large amount of video materials themselves, resulting in low video production efficiency.

Method used

Based on the keywords entered by the user, articles are automatically generated, and then videos are automatically generated based on the generated articles, providing an efficient one-stop video production solution.

Benefits of technology

It improves the efficiency of video production, reduces the time of searching for materials, and realizes efficient one-stop video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118828105B_ABST
    Figure CN118828105B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video generation method, apparatus, device, storage medium, and program product. The method comprises: obtaining first text information, wherein the first text information is used to describe the writing requirements of a video copy; generating second text information based on the first text information, wherein the second text information is copy information that meets the writing requirements described in the first text information; generating multimedia editing data based on third text information; and generating a target video based on the multimedia editing data. In an embodiment of the present disclosure, by generating copy information describing the writing requirements based on the description of the video copy, and then producing a video based on the generated copy information, an efficient one-stop video production solution is provided to improve video production efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing technology, and in particular to a video generation method, apparatus, device, storage medium, and program product. Background Art

[0002] With the rapid development of computer technology and mobile communication technology, various video platforms based on electronic devices have been widely used, greatly enriching people's daily lives. More and more users are willing to share their videos on video platforms for other users to watch.

[0003] Before sharing their videos, users need to edit and create them themselves. This process requires searching for a large amount of video material, such as images, articles, and music. This search can waste significant time and lead to low video production efficiency. Summary of the Invention

[0004] In order to solve the above technical problems, the embodiments of the present disclosure provide a video generation method, device, equipment, storage medium and program product, which automatically generate articles based on keywords input by users, and automatically generate videos based on the generated articles, providing an efficient one-stop video production solution to improve video production efficiency.

[0005] In a first aspect, an embodiment of the present disclosure provides a video generation method, comprising:

[0006] Obtaining first text information, wherein the first text information is used to describe writing requirements for a video copy;

[0007] generating second text information based on the first text information, wherein the second text information is copywriting information that meets the writing requirements described in the first text information;

[0008] Generate multimedia editing data based on third text information; wherein the third text information is obtained based on the second text information; the multimedia editing data includes at least one video track segment and at least one audio track segment, the at least one video track segment and the at least one audio track segment respectively corresponding to at least one text segment divided by the third text information, the target audio track segment is used to fill the reading voice matching the target text segment, and the target video track in the at least one video track segment and the target audio track segment occupy the same timeline position on the video editing timeline;

[0009] A target video is generated based on the multimedia editing data.

[0010] In a second aspect, an embodiment of the present disclosure provides a video generation device, including:

[0011] a first text information acquisition module, configured to acquire first text information, wherein multimedia editing data is generated based on third text information;

[0012] A second text information generating module, configured to generate second text information based on the first text information, wherein the second text information is copywriting information that meets the writing requirements described in the first text information;

[0013] a multimedia editing data generation module, configured to generate multimedia editing data based on the second text information; wherein the third text information is obtained based on the second text information; the multimedia editing data includes at least one video track segment and at least one audio track segment, wherein the at least one video track segment and the at least one audio track segment respectively correspond to at least one text segment divided from the third text information; the target audio track segment is used to fill in the reading voice matching the target text segment; and the target video track in the at least one video track segment and the target audio track segment occupy the same timeline position on the video editing timeline;

[0014] The target video generation module is used to generate a target video based on the multimedia editing data.

[0015] In a third aspect, an embodiment of the present disclosure provides an electronic device, the electronic device comprising:

[0016] at least one processor;

[0017] a storage device for storing at least one program;

[0018] When the at least one program is executed by the at least one processor, the at least one processor implements the video generation method as described in any one of the first aspects above.

[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video generation method as described in any one of the above-mentioned first aspects.

[0020] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, implements the video generation method as described in any one of the first aspects above.

[0021] The present disclosure provides a video generation method, apparatus, device, storage medium, and program product, the method comprising: obtaining first text information, wherein the first text information is used to describe the writing requirements of a video copy; generating second text information based on the first text information, wherein the second text information is copy information that meets the writing requirements described in the first text information; generating multimedia editing data based on third text information; wherein the third text information is obtained based on the second text information; the multimedia editing data includes at least one video track segment and at least one audio track segment, wherein the at least one video track segment and the at least one audio track segment respectively correspond to at least one text segment divided by the third text information, the target audio track segment is used to fill the reading voice matching the target text segment, and the target video track in the at least one video track segment and the target audio track segment occupy the same timeline position on the video editing timeline; and generating a target video based on the multimedia editing data. The present disclosure provides an efficient one-stop video production solution by generating copy information describing the writing requirements based on the description of the video copy, and then producing a target video according to the generated copy information, thereby improving video production efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0023] Figure 1 A flowchart of a video generation method provided by an embodiment of the present disclosure;

[0024] Figure 2 A schematic diagram of a video creation page provided in an embodiment of the present disclosure;

[0025] Figure 3 A schematic diagram of a document input interface provided in an embodiment of the present disclosure;

[0026] Figure 4 A schematic diagram of a multimedia editing page provided in an embodiment of the present disclosure;

[0027] Figure 5 A flowchart of a video generation method provided by an embodiment of the present disclosure;

[0028] Figure 6 A schematic diagram of a document input interface provided in an embodiment of the present disclosure;

[0029] Figure 7aA schematic diagram of a text input interface in an embodiment of the present disclosure;

[0030] Figure 7b A schematic diagram of a text input interface in an embodiment of the present disclosure;

[0031] Figure 7c A schematic diagram of a text input interface in an embodiment of the present disclosure;

[0032] Figure 8 is a structural diagram of a video generating device according to an embodiment of the present disclosure;

[0033] Figure 9 Schematic diagram of the structure of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0034] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0035] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0036] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0037] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0038] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "at least one".

[0039] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0040] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that the same reference numerals in different drawings will be used to refer to the same elements described above.

[0041] Figure 1 This is a flowchart of a video generation method in an embodiment of the present disclosure. This embodiment is applicable to the situation where videos are generated based on keywords. The method can be executed by a video generation device, and the video generation device can be implemented in software and / or hardware. The video generation method can be applied to electronic devices.

[0042] It will be understood that the above-mentioned electronic devices may be any other type of electronic device capable of performing data processing, which may include but are not limited to: mobile phones, stations, units, devices, multimedia computers, multimedia tablets, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication systems (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices or any combination thereof, including accessories and peripherals of these devices or any combination thereof.

[0043] like Figure 1 As shown, the video generation method provided by the embodiment of the present disclosure mainly includes steps S101-S104.

[0044] S101. Obtain first text information, where the first text information is used to describe writing requirements for a video copy.

[0045] In one embodiment of the present disclosure, the writing requirements for describing the video copy may be the key content of the target video that the user wants to generate. Specifically, the first text information may be one or more keywords or one or more subject words describing the video copy.

[0046] In an exemplary description, a user wants to make a video about "movie recommendations". The first text information may include "movie", "2023", "award-winning", "good reviews", etc. A user wants to make a video about "mobile phone selling points". The first text information may include "mobile phone model", "super large screen", "super long battery life", "reasonable price", etc.

[0047] In one embodiment of the present disclosure, obtaining the first text information includes: in response to the user's input operation, obtaining the first text information corresponding to the input operation. Specifically, presenting a text production video control on the video creation homepage, and displaying the video creation interface in response to the triggering operation of the text production video control. Figure 2 As shown, the video creation interface includes a text creation area 21, a video category selection area 22, and a video generation control 23. The text creation area 21 includes a text display area and an intelligent text generation control. The text display area 21 is used to display the third text information generated as multimedia editing data. The intelligent text generation control is used to display the text input interface in response to a user trigger operation.

[0048] In one embodiment of the present disclosure, Figure 3 As shown, the text input page includes a text input area 31. The text input area 31 displays text information input by the user in response to the user's input operation.

[0049] like Figure 3 As shown, the text input area 31 includes an input box, which is used to obtain first text information in response to the user's input operation.

[0050] S102: Generate second text information based on the first text information, wherein the second text information is copywriting information that meets the writing requirements described in the first text information.

[0051] In the embodiment of the present disclosure, the second text information refers to copywriting information that meets the writing requirements. Furthermore, the second text information is an article with coherent sentences and smooth sentences. The second text information can have one paragraph or multiple paragraphs. The second text information can be generated based on the first text information by an intelligent copywriting generation algorithm. The specific implementation of the intelligent copywriting generation algorithm will not be repeated in the embodiment of the present disclosure.

[0052] like Figure 3 As shown, in response to a confirmation operation on an input box in the text input area 31 , second text information is generated based on the first text information, and the generated second text information is displayed in the text input area 31 .

[0053] S103. Generate multimedia editing data based on the third text information; wherein, the third text information is obtained based on the second text information; the multimedia editing data includes at least one video track segment and at least one audio track segment, the at least one video track segment and the at least one audio track segment respectively correspond to at least one text segment divided from the third text information, the target audio track segment of the at least one audio track segment is used to fill in the reading voice matching the target text segment, and the target video track in the at least one video track segment and the target audio track segment occupy the same timeline position on the video editing timeline.

[0054] In one embodiment of the present disclosure, the third text information may include one or more of the following combinations: one or more second text information, edited second text information, and other text information edited by the user.

[0055] In one embodiment of the present disclosure, steps S101 - S102 may be performed multiple times to generate multiple second text messages. The generated second text messages may be edited or modified, and text information may be manually input in the text input area 31 .

[0056] In one embodiment of the present disclosure, Figure 3 As shown, in response to the trigger operation of completing the control in the text input area 31, switching to Figure 2 The video creation interface shown in FIG. 1 is displayed, and the third text information is displayed in the video creation interface.

[0057] In one embodiment of the present disclosure, when third text information is displayed in the video creation interface, in response to a trigger operation on the video generation control 23, multimedia editing data is generated based on the third text information, and the multimedia editing data is displayed in the multimedia editing interface.

[0058] In one embodiment of the present disclosure, the multimedia editing data further includes at least one subtitle track segment, and the target subtitle track segment is used to fill in subtitle information matching the target text segment.

[0059] In one embodiment of the present disclosure, the multimedia editing data further includes at least one background music track segment, and the background music track segment is used to fill in the background music.

[0060] In one embodiment of the present disclosure, Figure 4As shown, the multimedia editing interface may include a video preview area 41 and a multimedia track area 42, wherein the video preview area 41 is used to preview the target video generated by the multimedia editing data. The multimedia editing data displayed in the multimedia track area 42 includes video track segments, audio track segments, and background music track segments.

[0061] In one embodiment of the present disclosure, the target audio track segment of the at least one audio track segment is used to fill the reading voice matching the target text segment, and the target video track in the at least one video track segment occupies the same timeline position as the target audio track segment on the video editing timeline.

[0062] In one embodiment of the present disclosure, the target text segment may be a text segment segmented from the third text information. The text segment may be a single sentence, an incomplete sentence constructed according to Chinese segmentation rules, or multiple sentences, which is not specifically limited in the present embodiment. The target video segment refers to the video segment corresponding to the target text segment, and the target audio segment refers to the audio segment corresponding to the target text segment.

[0063] It should be noted that the background music in the background music track clip is not divided according to the text clip, but fills a complete audio file.

[0064] In one embodiment of the present disclosure, the target video segment is an idle segment.

[0065] In the embodiments of the present disclosure, Figure 2 As shown, in response to the triggering operation on the first video type control in the video creation interface, the video track segment in the multimedia editing data is set to an empty segment, that is, no picture or video image is added to the video track segment.

[0066] Furthermore, in response to the user's operation on the vacant segment, the picture selected by the user is added to the vacant segment. In other words, the user can fill the vacant segment with the picture selected by the user, so that the target video produced is more in line with the user's needs.

[0067] In one embodiment of the present disclosure, the target video segment is used to fill in a video image that matches the target text segment.

[0068] In the embodiments of the present disclosure, Figure 2 As shown, in response to the triggering operation of the second video type control in the video creation interface, the video track segment in the multimedia editing data is filled with a video image that matches the text segment. The video image can be obtained based on the text segment matching in a preset image database based on an image matching algorithm.

[0069] In one embodiment of the present disclosure, video images can be automatically matched in an image database based on text content, avoiding users from spending time searching for materials and further improving video production efficiency.

[0070] In one embodiment of the present disclosure, the target video segment is used to fill in an expression image that matches the target text segment.

[0071] In the embodiments of the present disclosure, Figure 2 As shown, in response to the triggering operation of the third video type control in the video creation interface, the video track segment in the multimedia editing data is filled with an expression image that matches the text segment. The expression image can be obtained based on the text segment matching in a preset expression package database based on a preset algorithm.

[0072] In one embodiment of the present disclosure, emoticon images can be automatically matched in an emoticon package database based on text content, and a more personalized video can be produced based on the text content.

[0073] In the disclosed embodiment, the emoticon package database and the picture database are divided into two databases, which can avoid the problem of inconsistent video style caused by the presence of both regular pictures and emoticon package pictures in one video.

[0074] S104: Generate a target video based on the multimedia editing data.

[0075] In one embodiment of the present disclosure, in response to a trigger operation of video completion, a target video is generated based on multimedia editing data.

[0076] In one embodiment of the present disclosure, the triggering operation for video completion may be triggering an export control in a multimedia editing interface. The exporting method may be to save the target video locally or to share the target video on another video sharing platform or website. This is not specifically limited in the present embodiment.

[0077] In one embodiment of the present disclosure, in response to a triggering operation of importing an editing control in a multimedia editing interface, multimedia editing data is imported into a video editor, and subsequently edited.

[0078] Based on the above embodiments, the embodiment of the present disclosure further optimizes the video generation method. Figure 5 As shown, the optimized video generation method mainly includes the following steps:

[0079] S201. Obtain first text information, where the first text information is used to describe writing requirements for a video copy.

[0080] Step S201 provided in the embodiment of the present disclosure is the same as the specific execution process of step S101 provided in the above embodiment. For details, please refer to the description in the above embodiment and will not be specifically limited in the embodiment of the present disclosure.

[0081] S202: Generate at least one candidate copywriting information based on the first text information, where the at least one candidate copywriting information meets the writing requirements indicated by the first text information.

[0082] In one embodiment of the present disclosure, a target document category is determined from a plurality of document categories; the candidate document information is generated based on the target document category and the first text information, and the document category of the second text information is the target document category.

[0083] In one embodiment of the present disclosure, the copy category refers to the category to which the generated copy belongs. Specifically, the copy category may include: a first copy category and a second text category. The first text category can be understood as a copy category that is generally applicable to users on various topics, specifically, various topics include: technology, finance, entertainment, etc. The second text category can be understood as a text category applicable to product planning or product marketing. Specifically, it can be an introduction to the selling points of a certain mobile phone, or the reason for recommending a certain item, etc.

[0084] In one embodiment of the present disclosure, the target copy category can be determined based on the user's selection, and each target copy category corresponds to an intelligent copy generation algorithm model.

[0085] In one embodiment of the present disclosure, the copy input interface has a first copy category control and a second copy category control. The first copy category control is used to respond to a user's trigger operation and use the copy category corresponding to the first copy category control as the target copy category; the second copy category control is used to respond to a user's trigger operation and use the copy category corresponding to the second copy category control as the target copy category.

[0086] like Figure 3 As shown, a first text category control and a second text category control are presented in the text input interface. In the embodiment of the present disclosure, different text category controls correspond to different prompt information, and different text category controls correspond to different intelligent text generation algorithm models.

[0087] In one embodiment of the present disclosure, in response to a selection operation on the first copy category control, a prompt message "Write a copy of the first category, the subject is:" is displayed in the input box in the copy input area 31, and the user can insert the first text message after the prompt message. In response to a confirmation operation on the input box, the first text message is obtained, and at least one candidate copy message is generated based on the first copy category and the first text message.

[0088] In one embodiment of the present disclosure, at least one candidate copy information is generated based on a first copy category and first text information, including: calling a first intelligent copy generation algorithm corresponding to the first copy category based on the first copy category, using the first intelligent copy generation algorithm to process the first text information to generate multiple candidate copy information.

[0089] In one embodiment of the present disclosure, in response to a selection operation on the second copy category control, a prompt message "Write a copy for the second category. The product and selling point are:" is displayed in the input box in the copy input area 31. The user can enter the first text message after the prompt message. In response to a confirmation operation on the input box, the first text message is obtained, and at least one candidate copy message is generated based on the second copy category and the first text message.

[0090] In one embodiment of the present disclosure, at least one candidate copy information is generated based on the second copy category and the first text information, including: calling a second intelligent copy generation algorithm corresponding to the second copy category based on the second copy category, and using the second intelligent copy generation algorithm to process the first text information to generate multiple candidate copy information.

[0091] It should be noted that the first intelligent copy generation algorithm and the second intelligent copy generation algorithm are two different intelligent copy generation algorithms.

[0092] In one embodiment of the present disclosure, the basic network models used by the above-mentioned two intelligent copy generation algorithms may be the same or different. The training methods of the above-mentioned two intelligent copy generation algorithms may be the same or different. It should be noted that the training samples of the first intelligent copy generation algorithm and the second intelligent copy are different. The training samples of the first intelligent copy generation algorithm are copy information of the first copy category, and the training samples of the second intelligent copy generation algorithm are copy information of the second copy category.

[0093] S203: In response to a switching operation triggered by the user, different candidate copy information among the at least one candidate copy information is switched and presented in the copy input interface.

[0094] In one embodiment of the present disclosure, a candidate text area is provided in the text input interface, and the candidate text area is used to display candidate text information. Specifically, the candidate text area is displayed in the text input area 31 in an inserted form.

[0095] In one embodiment of the present disclosure, the text input interface has a text switching control, which is used to trigger the switching operation so as to switch and present different candidate text information in the at least one candidate text information in the candidate text area.

[0096] In one embodiment of the present disclosure, Figure 6 As shown, the candidate text area is displayed in the text input area 31 in an inserted form, and the candidate text area 61 includes a first text switching control 62 and a second text switching control 63 .

[0097] In one embodiment of the present disclosure, multiple candidate copy information is arranged in a set order, and the first text switch control 62 is used to respond to the user's trigger operation to display the candidate copy information that is arranged before the current candidate copy information in the candidate text area, and the second text switch control 63 is used to respond to the user's trigger operation to display the candidate copy information that is arranged after the current candidate copy information in the candidate text area.

[0098] In one embodiment of the present disclosure, five candidate text information items are described as follows: candidate text information A, candidate text information B, candidate text information C, candidate text information D, and candidate text information E. The text candidate area displays candidate text information A, which is ranked highest. In response to a triggering operation on the second text switching control 63, the text candidate area displays candidate text information B. Meanwhile, in response to a triggering operation on the first text switching control 62, the text candidate area displays candidate text information A.

[0099] S204: In response to a confirmation operation triggered by the user, determine the candidate copy information that is switched to the copy input interface and presented among the at least one candidate copy information as the second text information.

[0100] In one embodiment of the present disclosure, the text input interface has a text confirmation control, which is used to trigger the confirmation operation so that the candidate text information switched to the candidate text area for presentation in the at least one candidate text information is determined as the second text information and the second text information is presented in the text input area.

[0101] In one embodiment of the present disclosure, Figure 6As shown, the candidate copy area 61 has a copy confirmation control. In response to a trigger operation on the copy confirmation control, the candidate copy information presented in the candidate copy area is determined as the second text information and the second text information is presented in the copy input area 31.

[0102] In one embodiment of the present disclosure, after the at least one candidate copy information is generated, the candidate copy information presented in the candidate copy area is inserted into the user input position in the copy input area; in response to the confirmation operation, the candidate copy area is deleted from the copy input area.

[0103] In one embodiment of the present disclosure, Figure 6 As shown, the candidate text information (AAAAA) presented in the candidate text area 61 is inserted into the user input position within the text input area 31. The user input position is the cursor position before the candidate text information is generated. Furthermore, in response to the user triggering the text confirmation control, the candidate text area is deleted.

[0104] In one embodiment of the present disclosure, Figure 7a As shown, if there is no other text information in the text input area before the at least one candidate text information is generated, the second text information is presented in the text input area 31 after the second text information is determined, and the candidate text area 61 is deleted.

[0105] S205: Generate a third text based on the second text information.

[0106] In one embodiment of the present disclosure, if fourth text information is presented in the text input area before the at least one candidate text information is generated, fifth text information formed by merging the second text information and the fourth text information is presented in the text input area after the second text information is determined.

[0107] In the embodiment of the present disclosure, the fourth text information may be text information manually input by a user, or may be the second text information determined in steps S201 - S204 , or may be text information edited and modified by a user.

[0108] In the embodiment of the present disclosure, Figure 7bAs shown, if a fourth text message (#######) is presented in the text input area 31 before the at least one candidate text message is generated, and the user input position is at the end of the fourth text message, the candidate text area 61 is displayed at the end of the fourth text message. In response to the user triggering the text confirmation control, the candidate text area is deleted, and the second text message (AAAAAA) is spliced ​​at the end of the fourth text message (#######), forming a fifth text message (#######AAAAAA), as shown in FIG. Figure 7b shown. Figure 7b The description is made by taking the case where the user input position is at the end of the fourth text message as an example.

[0109] In one embodiment of the present disclosure, if the user input position is located in the middle position of the fourth text information, then in the text input area, the fourth text information is cut from the middle position by the candidate copy area and presented on both sides of the candidate copy area, and the second text information in the fifth text information is inserted into the middle position of the fourth text information.

[0110] In the embodiment of the present disclosure, Figure 7c As shown, if the fourth text information (#######) is presented in the text input area 31 before the at least one candidate text information is generated, and the user input position is located in the middle of the fourth text information, the candidate text area 61 splits the fourth text information in the text input area 31 from the user input position, and displays the two parts of the fourth text information on both sides of the candidate text area 61. Furthermore, in response to the user triggering the text confirmation control, the candidate text area is deleted, and the second text information (AAAAAA) is inserted in the middle of the fourth text information (#######), forming the fifth text information (###AAAAAA####), as shown in FIG. Figure 7c shown. Figure 7c Here, the user input position is in the middle of the fourth text information as an example for explanation.

[0111] In one embodiment of the present disclosure, the method further includes: in response to a user input operation, editing the fifth text information to obtain the third text information.

[0112] In one embodiment of the present disclosure, Figure 7a 、 7b As shown in FIG7c, there is an intelligent copy generation control in the text input area 31. In response to the triggering operation for the intelligent copy generation control, the following is displayed: Figure 3 The page shown in FIG, wherein the page includes the confirmed text information. Figure 3The operation of the text editing interface shown starts to generate new second text information.

[0113] In one embodiment of the present disclosure, in response to a user's editing operation on the text input area 31, the fifth text information is edited to obtain the third text information, wherein the editing includes operations such as input, deletion, copy, and paste.

[0114] In one embodiment of the present disclosure, in response to the triggering operation of completing the control in the text input area 31, the copy input interface is closed, and the third text information is placed in the text creation area 21 (such as Figure 2 shown).

[0115] S206: Generate multimedia editing data based on the third text information.

[0116] S207: Generate a target video based on the multimedia editing data.

[0117] Steps S206-S207 provided in the embodiment of the present disclosure are similar to the specific execution process of steps S103-S104 provided in the above embodiment. For details, please refer to the description in the above embodiment and will not be further limited in the embodiment of the present disclosure.

[0118] Figure 8 Schematic diagram of the structure of a video generating device in an embodiment of the present disclosure. This embodiment is applicable to the case of generating a video based on input text. The video generating device can be implemented in software and / or hardware.

[0119] like Figure 8 As shown, the video generating device 80 provided by the embodiment of the present disclosure mainly includes: a first text information acquisition module 81, a second text information generation module 82, a multimedia editing data generation module 83 and a target video generation module 84.

[0120] Among them, the first text information acquisition module 81 is used to acquire the first text information, wherein multimedia editing data is generated based on the third text information; the second text information generation module 82 is used to generate the second text information based on the first text information, wherein the second text information is copywriting information that meets the writing requirements described in the first text information; the multimedia editing data generation module 83 is used to generate multimedia editing data based on the second text information; wherein the third text information is obtained based on the second text information; the multimedia editing data includes at least one video track segment and at least one audio track segment, the at least one video track segment and the at least one audio track segment respectively correspond to the at least one text segment divided by the third text information, the target audio track segment is used to fill the reading voice matching the target text segment, and the target video track in the at least one video track segment and the target audio track segment occupy the same timeline position on the video editing timeline; the target video generation module 84 is used to generate a target video based on the multimedia editing data.

[0121] In one embodiment of the present disclosure, the second text information generation module 82 includes: a candidate copy information generation unit, which is used to generate at least one candidate copy information based on the first text information, and the at least one candidate copy information meets the writing requirements indicated by the first text information; a candidate copy information switching unit, which is used to switch and present different candidate copy information in the at least one candidate copy information in the copy input interface in response to a switching operation triggered by the user; and a candidate copy information confirmation unit, which is used to determine the candidate copy information in the at least one candidate copy information that is switched to the copy input interface for presentation as the second text information in response to a confirmation operation triggered by the user.

[0122] In one embodiment of the present disclosure, the copy input interface has a copy input area and a candidate copy area; the copy input interface has a copy switching control, and the copy switching control is used to trigger the switching operation so that different candidate copy information among the at least one candidate copy information is switched and presented in the candidate copy area; the copy input interface has a copy confirmation control, and the copy confirmation control is used to trigger the confirmation operation so that the candidate copy information presented in the candidate copy area is determined as the second text information and the second text information is presented in the copy input area.

[0123] In one embodiment of the present disclosure, after the at least one candidate copy information is generated, the candidate copy information presented in the candidate copy area is inserted into the user input position in the copy input area; in response to the confirmation operation, the candidate copy area is deleted from the copy input area.

[0124] In one embodiment of the present disclosure, if fourth text information is presented in the text input area before the at least one candidate text information is generated, fifth text information formed by merging the second text information and the fourth text information is presented in the text input area after the second text information is determined.

[0125] In one embodiment of the present disclosure, if the user input position is located in the middle position of the fourth text information, then in the text input area, the fourth text information is cut from the middle position by the candidate copy area and presented on both sides of the candidate copy area, and the second text information in the fifth text information is inserted into the middle position of the fourth text information.

[0126] In one embodiment of the present disclosure, in response to an input operation of a user, the fifth text information is edited to obtain the third text information.

[0127] In one embodiment of the present disclosure, the device also includes: a target copy category confirmation module, used to determine a target copy category among multiple copy categories; the candidate copy information is generated based on the target copy category and the first text information, and the copy category of the second text information is the target copy category.

[0128] In one embodiment of the present disclosure, the copy input interface has a first copy category control and a second copy category control. The first copy category control is used to respond to a user's trigger operation and use the copy category corresponding to the first copy category control as the target copy category; the second copy category control is used to respond to a user's trigger operation and use the copy category corresponding to the second copy category control as the target copy category.

[0129] In one embodiment of the present disclosure, the target video segment is a vacant segment; or the target video segment is used to fill in a video image that matches the target text segment; or the target video segment is used to fill in an expression image that matches the target text segment.

[0130] The video generation device provided in the embodiment of the present disclosure can execute the steps executed in the video generation method provided in the embodiment of the method of the present disclosure, and the execution steps and beneficial effects are not repeated here.

[0131] Figure 9 This is a schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 9, which shows a schematic structural diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure. The electronic device 900 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable terminal devices, and the like, as well as fixed terminals such as digital TVs, desktop computers, smart home devices, and the like. Figure 9 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0132] like Figure 9 As shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903 to implement the picture rendering method of the embodiment as described in the present disclosure. In the RAM 903, various programs and data required for the operation of the terminal device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0133] Typically, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the terminal device 900 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 9 The terminal device 900 is shown as having various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0134] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart, thereby implementing the video generation method described above. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0135] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection having at least one wire, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0136] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0137] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0138] The computer-readable medium carries one or more programs. When the one or more programs are executed by the terminal device, the terminal device: obtains first text information, wherein the first text information is used to describe the writing requirements of the video copy; generates second text information based on the first text information, wherein the second text information is copy information that meets the writing requirements described in the first text information; generates multimedia editing data based on third text information; wherein the third text information is obtained based on the second text information; the multimedia editing data includes at least one video track segment and at least one audio track segment, wherein the at least one video track segment and the at least one audio track segment respectively correspond to at least one text segment divided by the third text information, the target audio track segment of the at least one audio track segment is used to fill the reading voice matching the target text segment, and the target video track in the at least one video track segment and the target audio track segment occupy the same timeline position on the video editing timeline; and generates a target video based on the multimedia editing data. Optionally, when the one or more programs are executed by the terminal device, the terminal device may also perform other steps described in the above embodiment.

[0139] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code includes at least one executable instruction for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0141] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0142] The functions described above herein may be performed, at least in part, by at least one hardware logic component. For example, and without limitation, exemplary types of hardware logic components that may be used include: a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a system on a chip (SOC), a complex programmable logic device (CPLD), and the like.

[0143] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on at least one line, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0144] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0145] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0146] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video generation method, characterized in that: include: Obtaining first text information, wherein the first text information is used to describe writing requirements for a video copy; generating second text information based on the first text information, wherein the second text information is copywriting information that meets the writing requirements described in the first text information; Generate multimedia editing data based on third text information; wherein the third text information is obtained based on the second text information; the multimedia editing data includes at least one video track segment and at least one audio track segment, the at least one video track segment and the at least one audio track segment respectively corresponding to at least one text segment divided by the third text information, a target audio track segment of the at least one audio track segment is used to fill in the reading voice matching the target text segment, and the target video track in the at least one video track segment and the target audio track segment occupy the same timeline position on the video editing timeline; A target video is generated based on the multimedia editing data.

2. The method according to claim 1, characterized in that The generating the second text information based on the first text information includes: generating at least one candidate copywriting information based on the first text information, wherein the at least one candidate copywriting information meets the writing requirements indicated by the first text information; In response to a switching operation triggered by the user, switching and presenting different candidate copy information among the at least one candidate copy information in the copy input interface; In response to a confirmation operation triggered by the user, the candidate copy information that is switched to the copy input interface and presented among the at least one candidate copy information is determined as the second text information.

3. The method according to claim 2, characterized in that The text input interface has a text input area and a candidate text area; The text input interface includes a text switching control, and the text switching control is used to trigger the switching operation so as to switch and present different candidate text information among the at least one candidate text information in the candidate text area; The text input interface includes a text confirmation control, which is used to trigger the confirmation operation, so that the candidate text information presented in the candidate text area is determined as the second text information and the second text information is presented in the text input area.

4. The method according to claim 3, characterized in that After the at least one candidate copy information is generated, the candidate copy information presented in the candidate copy area is inserted into the user input position in the copy input area; In response to the confirmation operation, the candidate text area is deleted from the text input area.

5. The method according to claim 4, characterized in that If fourth text information is presented in the text input area before the at least one candidate text information is generated, fifth text information formed by merging the second text information and the fourth text information is presented in the text input area after the second text information is determined.

6. The method according to claim 5, characterized in that If the user input position is located in the middle of the fourth text information, the fourth text information is cut from the middle position by the candidate text area in the text input area and presented on both sides of the candidate text area, and the second text information in the fifth text information is inserted into the middle position of the fourth text information.

7. The method according to claim 5, characterized in that Also includes: In response to the user's input operation, the fifth text information is edited to obtain the third text information.

8. The method according to claim 2, characterized in that Also includes: Identify the target copywriting category among multiple copywriting categories; The candidate text information is generated based on the target text category and the first text information, and the text category of the second text information is the target text category.

9. The method according to claim 8, characterized in that The text input interface has a first text category control and a second text category control, wherein the first text category control is used to respond to a user's triggering operation and set the text category corresponding to the first text category control as the target text category; The second text category control is used to respond to a user's triggering operation and use the text category corresponding to the second text category control as a target text category.

10. The method according to claim 1, characterized in that The target video segment is an empty segment; or The target video segment is used to fill in a video image that matches the target text segment; or The target video segment is used to fill in the expression image that matches the target text segment.

11. A video generating device, characterized in that: include: A first text information acquisition module is used to acquire first text information, wherein the first text information is used to describe the writing requirements of the video copy; A second text information generating module, configured to generate second text information based on the first text information, wherein the second text information is copywriting information that meets the writing requirements described in the first text information; A multimedia editing data generation module is configured to generate multimedia editing data based on third text information; wherein the third text information is obtained based on the second text information; the multimedia editing data includes at least one video track segment and at least one audio track segment, wherein the at least one video track segment and the at least one audio track segment respectively correspond to at least one text segment divided from the third text information; a target audio track segment of the at least one audio track segment is used to fill in a reading voice matching the target text segment; and a target video track in the at least one video track segment and the target audio track segment occupy the same timeline position on the video editing timeline; The target video generation module is used to generate a target video based on the multimedia editing data.

12. An electronic device, characterized in that: The electronic device comprises: at least one processor; a storage device for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

14. A computer program product, comprising a computer program or instructions, which implements the method according to any one of claims 1 to 10 when executed by a processor.

Citation Information

Patent Citations

  • Video generation method and device

    CN112291614A

  • Video generation method and device, storage medium and electronic equipment

    CN112929746A