Video generation method, apparatus, device, storage medium, and program product
The video generation method addresses the challenge of meeting individualized video production needs by allowing users to select video materials based on input text, resulting in more accurate and personalized video content.
Patent Information
- Application Number
- JP2023578852
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-09-02
- Filing Date
- 2023-09-04
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-09-04
AI Technical Summary
Existing video production methods fail to meet individualized video production needs as they rely on intelligent matching algorithms that may not accurately match video images with user preferences.
A video generation method that generates first video editing data based on input text, allowing users to freely select video materials for editing, including embedding reading voices and enabling users to insert target videos into empty segments.
Enables personalized video production by allowing users to select video materials according to their preferences, improving the accuracy of video content matching and enhancing user experience.
Smart Images

Figure 0007684446000001 
Figure 0007684446000002 
Figure 0007684446000003
Abstract
Description
Technical Field
[0001] [Cross - reference to Related Applications] This application claims the priority of Chinese Patent Application No. 202211075071.9 filed on September 2, 2022, and the entire content disclosed in the above - mentioned Chinese Patent Application is incorporated herein by reference in its entirety.
[0002] This disclosure relates to a video generation method, apparatus, device, storage medium, and program product.
Background Art
[0003] With the rapid development of computer technology and mobile communication technology, various video platforms based on electronic devices have been widely applied, greatly enriching people's daily lives. The number of users who share their video works on video platforms and let other users view them is increasing.
[0004] The video production process is to obtain the input text entered by the user, match the corresponding video image to the input text through an intelligent matching algorithm, and synthesize the corresponding target video based on the input text and the video image.
[0005] In the above - mentioned video production process, the video image is obtained by an intelligent matching algorithm, and there may be a possibility that it cannot meet the individualized video production needs of users.
Summary of the Invention
Means for Solving the Problems
[0006] To solve the above technical problems, embodiments of the present disclosure provide a video generation method, apparatus, device, storage medium, and program product, which generate first video editing data based on input text. In the first video editing data, a user can freely select video materials according to their preferences to perform video editing, meeting the needs of personalized video production.
[0007] Embodiments of the present disclosure responding to a first instruction triggered for the input text, generating first video editing data based on the input text, wherein the first video editing data includes at least one first video segment and at least one audio segment, the at least one first video segment and the at least one audio segment respectively correspond to at least one divided text segment of the input text, a first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the target audio segment is used to be embedded by a reading voice that matches the target text segment, and the first target video segment is an empty segment; importing the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor, wherein an interval between time lines of tracks of the first target video segment and the target audio segment is the same; In response to triggering a second command for the first target video segment with the video editor, based on the first video editing data, embedding a first target video into the first target video segment to obtain second video editing data, wherein the first target video is a video obtained based on first target video material indicated by the second command; A video generation method is provided, including: generating a first target video based on second edited video data.
[0008] Another embodiment of the present disclosure is A first video editing data determination module for generating first video editing data based on input text in response to a first command triggered for the input text, wherein the first video editing data includes at least one first video segment and at least one audio segment, the at least one first video segment and the at least one audio segment respectively correspond to at least one divided text segment of the input text, a first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the target audio segment is used to be embedded by a reading voice matching the target text segment, and the first target video segment is an empty segment With A certain first video editing data determination module; A first video editing data import module for importing the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor, wherein an interval between timelines of tracks of the first target video segment and the target audio segment is the same. In response to triggering a second command for the first target video segment by the video editor, based on the first video editing data, a second video editing data determination module for embedding a first target video into the first target video segment to obtain second video editing data, wherein the first target video is a second video editing data determination module for a video obtained based on first target video material indicated by the second command, Second 2-bit Video Editing Based on the data, a second 1-bit Video generation module for generating a first target video, and a video generation device including the same is provided.
[0009] Another embodiment of the present disclosure is One or more processors and A storage device for storing one or more programs, and includes When the one or more programs are executed by the one or more processors, an electronic device is provided that causes the one or more processors to implement the video generation method according to any one of the first aspects.
[0010] Yet another embodiment of the present disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the video generation method according to any one of the first aspects.
[0011] Yet another embodiment of the present disclosure provides a computer program product including a computer program or instructions that, when executed by a processor, implement the video generation method according to any one of the first aspects.
[0012] Embodiments of the present disclosure provide a video generation method, apparatus, device, storage medium, and program product. The method includes: in response to a first command triggered for an input text, generating first video editing data based on the input text, where the first video editing data includes at least one first video segment and at least one audio segment, the at least one first video segment and the at least one audio segment respectively correspond to at least one segmented text segment of the input text, a first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the target audio segment is used to be embedded with a spoken voice matching the target text segment, and the first target video segment is an empty segment; importing the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor, where an interval between timelines of tracks of the first target video segment and the target audio segment is the same; in response to triggering a second command for the first target video segment in the video editor, based on the first video editing data, embedding a first target video into the first target video segment to obtain second video editing data, where the first target video is a video obtained based on a first target video material indicated by the second command; and generating a video based on the second video editing data. Embodiments of the present disclosure generate first video editing data based on an input text, whereby in the first video editing data, a user can freely select video materials according to their preferences for video editing, meeting the needs of individualized video production. 2-bit video Editing Based on the data, the 1-bit video is generated.
Brief Description of the Drawings
[0013] The above features, other features, advantages and aspects of each embodiment of the present disclosure will become clearer by referring to the following specific embodiments in conjunction with the drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the components and elements are not necessarily drawn to scale.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Modes for Carrying Out the Invention
[0014] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the drawings. Although several embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the protection scope of the present disclosure.
[0015] It should be understood that each step described in the embodiments of the method of the present disclosure may be executed in various orders and / or executed in parallel. In addition, the embodiments of the method may include additional steps and / or omit the execution of the steps shown. The scope of the present disclosure is not limited in this regard.
[0016] As used herein, the term "including" and its variations mean non-limiting inclusion, that is, "including but not limited to...". The term "based on" means "at least partially based on". The term "one embodiment" represents "at least one embodiment", the term "another embodiment" represents "at least one other embodiment", and the term "some embodiments" represents "at least some embodiments". Related definitions of other terms will be described below.
[0017] It should be noted that the concepts such as "first", "second", etc. referred to in the present disclosure are only for distinguishing different devices, modules or units, and do not limit the order of functions executed by these devices, modules or units or their mutual dependencies.
[0018] It should be noted that the modifications of "one" and "a plurality" referred to in the present disclosure are not restrictive but illustrative. Those skilled in the art should understand that, in the context and unless otherwise specified, it should be understood as "one or a plurality".
[0019] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are merely for the purpose of explanation and are not intended to limit the scope of these messages or information.
[0020] Before further elaborating on the embodiments of the present disclosure, the nouns and terms related to the embodiments of the present disclosure will be explained, and the nouns and terms related to the embodiments of the present disclosure are applicable to the following interpretations.
[0021] In related technologies, users can produce videos using mobile terminals such as mobile phones, tablet computers, notebook computers, or other electronic devices. Currently, in the commonly used video production method, although the user creates input text in advance, there is no appropriate image or video material. In this case, usually, the user inputs the text, and the video production client matches the pixel material corresponding to the input text through an intelligent matching algorithm, and synthesizes the corresponding target video based on the input text and the pixel material.
[0022] In the above video production process, the video images are obtained by an intelligent matching algorithm, and the user cannot interfere with the pixel material matched by the intelligent matching algorithm. Therefore, the matched pixel material may not meet the individualized video production needs of the user. For example, when the user wants to produce a video of the cooking process, the user has pre-created a recipe on what materials are needed and what operations are to be performed at each step, and for each material and each step, the user has taken corresponding photos or short videos. According to the current video production method, it is necessary to intelligently match the pixel material based on the recipe. However, for example, in the step of stir-frying, there is a certain consistency and context regarding what to put in first and what to put in later, and the current intelligent matching algorithm may not be able to accurately match the image that meets the user's needs.
[0023] To solve the above technical text, an embodiment of the present disclosure includes the step of generating first video editing data based on input text in response to a first command triggered for the input text, where the first video editing data includes at least one first video segment and at least one audio segment, the at least one first video segment and the at least one audio segment respectively correspond to at least one divided text segment of the input text, a first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the target audio segment is used to be embedded by a reading voice that matches the target text segment, and the first target video segment is an empty segment; the step of importing the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor, where the time line interval between the tracks of the first target video segment and the target audio segment is the same; the step of obtaining second video editing data by embedding a first target video into the target video segment based on the first video editing data in response to triggering a second command for the first target video segment in the video editor, where the first target video is a video obtained based on a first target video material indicated by the second command; and the step of generating a video based on the second video editing data. 1-bit A video generation method is provided, which includes the above steps.
[0024] Embodiments of the present disclosure generate first video editing data based on input text. As a result, in the first video editing data, users can freely select video materials according to their preferences for video editing, meeting the needs of individualized video production.
[0025] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. It should be noted that the same reference numerals in different figures are used to refer to the same components that have been described.
[0026] FIG. 1 is a Video generation method of system Schematic diagram of provided by an embodiment of the present disclosure. As shown in FIG. 1, the system 100 may include a plurality of user terminals 110, a network 120, a server 130, and a database 140. For example, the system 100 can implement the video generation method described in any of the embodiments of the present disclosure.
[0027] It can be understood that the user terminal 110 may be any other type of electronic device capable of executing data processing, including but not limited to mobile phones, sites, units, devices, multimedia computers, multimedia tablets, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, accessories and peripheral devices including these devices, or any combination thereof.
[0028] The user can operate through an application program installed on the user terminal 110. The application program transmits the user's behavior data to the server 130 via the network 120, and the user terminal 110 can further receive the data transmitted from the server 130 via the network 120. The embodiments of the present disclosure do not limit the hardware system and software system of the user terminal 110. For example, the user terminal 110 may be based on a processor such as ARM or X86, and may be equipped with input / output devices such as a camera, a touch panel, and a microphone, and may execute an operating system such as Windows, iOS, Linux (registered trademark), Android, or HarmonyOS.
[0029] For example, the application program on the user terminal 110 may be a video production application program, for example, a video production application program based on multimedia resources such as videos, photos, and texts. Taking a video production application program based on multimedia resources such as videos, photos, and texts as an example, the user can perform video shooting, script creation, video production, video clips, etc. on the user terminal 110 through the video production application program. At the same time, the user can view or browse videos and other content posted by other users, and can perform operations such as liking, commenting, and reposting.
[0030] The user terminal 110 can implement the video generation method provided by the embodiments of the present disclosure by executing a process or a thread. In some examples, the user terminal 110 can execute the video generation method by its built-in application program. In some other examples, the user terminal 110 can execute the video generation method by calling an application program stored outside the user terminal 110.
[0031] Network 120 may be a single network or a combination of at least two different networks. For example, network 120 can include, but is not limited to, one or a combination of multiple types such as a local area network, a wide area network, a public network, a private network, etc. Network 120 may be a computer network such as the Internet and / or various electronic communication networks (e.g., 3G / 4G / 5G mobile communication networks, WIFI, Bluetooth (registered trademark), ZigBee, etc.), and the embodiments of the present disclosure do not limit this.
[0032] Server 130 may be a single server, a server group, or a cloud server, and each server in the server group is connected via a wired or wireless network. One server group may be centralized such as a data center or decentralized. Server 130 may be local or remote. Server 130 can communicate with user terminal 110 via a wired or wireless network. The embodiments of the present disclosure do not limit the hardware system and software system of server 130.
[0033] Database 140 may generally refer to a device having a storage function. Database 140 is mainly used to store various data utilized, generated, or output by user terminal 110 and server 130 during operation. For example, taking the video production application program on user terminal 110 based on multimedia resources such as the above-mentioned videos, photos, texts, etc. as an example, the data stored in database 140 may include resource data such as videos and texts uploaded by the user via user terminal 110, and interactive operation data such as likes and comments.
[0034] The database 140 may be local or remote. The database 140 may include various memories such as Random Access Memory (RAM) and Read Only Memory (ROM). The above-mentioned storage devices are only some of the examples, and the storage devices available for the system 100 are not limited to these. The embodiments of the present disclosure do not limit the hardware system and software system of the database 140. For example, it may be a relational database or a non-relational database.
[0035] The database 140 may be interconnected or communicated with the server 130 or a part thereof via the network 120, directly interconnected or communicated with the server 130, or a combination of the above two methods.
[0036] In some examples, the database 140 may be an independent device. In some other examples, the database 140 may be integrated into at least one of the user terminal 110 and the server 130. For example, the database 140 may be installed in the user terminal 110 or in the server 130. Also, for example, the database 140 may be distributed such that a part of it is installed in the user terminal 110 and another part is installed in the server 130.
[0037] FIG. 2 is a flowchart of a video generation method according to an embodiment of the present disclosure. This embodiment is applicable when generating a video based on input text. The method can be executed by a video generation device, which can be realized in a software and / or hardware manner. The video generation method can be realized in the System as described in FIG. 1.
[0038] As shown in FIG. 2, the video generation method provided by the embodiment of the present disclosure mainly includes steps S101 to S104.
[0039] S101: In response to a first instruction triggered by the input text, generate first video editing data based on the input text. The first video editing data includes at least one first video segment and at least one audio segment. The at least one first video segment and the at least one audio segment respectively correspond to at least one segmented text segment of the input text. The first target video segment in the at least one first video segment and the target audio segment in the at least one audio segment correspond to the target text segment in the at least one text segment. The target audio segment is used to be embedded with the synthesized voice that matches the target text segment, and the first target video segment is an empty segment.
[0040] A response is used to represent a condition or state on which the operation to be executed depends. When the dependent condition or state is satisfied, one or more operations to be executed may be in real time or may have a set delay. Unless otherwise specified, the order of execution of multiple operations to be executed is not limited.
[0041] The video editing data may be understood as a video editing draft or an editing project file and is used to record and play back the user's video editing process. Specifically, it includes the audio material and video material to be edited, and the instruction information of the editing operations performed on the audio material and video material.
[0042] In one embodiment of the present disclosure, the first instruction only instructs the client to generate first video editing data based on the input text, and can be understood as an instruction that does not require intelligent matching of pixel materials based on the input text. The first instruction triggered by the input text may be in response to a trigger operation on a first control on the video production page by the user.
[0043] In one embodiment of the present disclosure, when a user wants to create a video based on input text, one video production application program can be launched in advance, and one of the subroutines included in the video production application program has a function of creating a video based on the input text. A video production application program for creating a video based on the input text may be launched.
[0044] In one embodiment of the present disclosure, the method further includes a step of displaying a video production page in response to a trigger operation on the video production control, where the video production page includes a first control, a second control, and a text editing area. The first control is used to trigger the first instruction in response to the user's trigger operation, the second control is used to trigger the third instruction in response to the user's trigger operation, and the text editing area is used to obtain input text in response to the user's editing operation.
[0045] In an embodiment of the present disclosure, a video production application program is launched, the application program interface is displayed, the application program interface includes a video production control, and a video production page is displayed in response to a trigger operation on the video production control by the user. Here, the trigger operation may be one or a combination of operations such as clicking, long pressing, hovering, touching, etc., and the embodiments of the present disclosure do not limit this.
[0046] In an embodiment of the present disclosure, as shown in FIG. 3, the video production page 30 includes a first control 301, a second control 302, and a text editing area 303. The first control 301 is used to trigger the first command in response to a trigger operation by the user. The second control 302 is used to trigger the third command in response to a trigger operation by the user. The text editing area 303 is used to obtain input text in response to an editing operation by the user.
[0047] In an embodiment of the present disclosure, in response to a trigger operation by the user on the first control, the first command triggered for the input text is triggered. In response to the first command, first video editing data is generated based on the input text. The input text refers to the text saved and displayed in the text editing area when responding to the first command.
[0048] Furthermore, in response to a trigger operation by the user on the first control, the first control is selected. Next, the video production control included in the video production page 30 to In response, the first command triggered for the input text is triggered. In response to the first command, first video editing data is generated based on the input text. The input text refers to the text saved and displayed in the text editing area when responding to the first command.
[0049] In one embodiment of the present disclosure, the input text is divided into at least one text segment, and for each target text segment, its corresponding spoken voice is obtained by intelligent matching. The target text segment and its corresponding spoken voice are aligned on the timeline of the track to obtain its corresponding target audio segment, and an empty segment is obtained as the corresponding first target video segment of the text segment. The first target video segment and the target audio segment are synthesized to obtain a first video segment. The above operations are performed for each target text segment to obtain a plurality of first video segments, and the plurality of first video segments are synthesized according to the order of the target text segments in the input text to obtain first video editing data.
[0050] It can be understood that the first target video segment and the target audio segment corresponding to the target text segment mean that the timelines of the above three segments are aligned and the expression contents correspond. For example, their timelines are all segments between 1 minute 55 seconds and 1 minute 57 seconds. For example, the target text segment is "stir-fry quickly over high heat", and the target audio segment is the spoken voice of the four characters "stir-fry quickly over high heat".
[0051] The first target video segment can be understood as any video segment among at least one first video band, and the fact that the first target video segment is empty can be understood as setting the first target video segment to be empty in the first video editing data. This empty space may empty the video track without setting any pixel material, or may set the pixel material preset on the video track. The preset pixel material is set by the system and cannot be arbitrarily changed by the user. That is, no matter what the input text is, its corresponding first target video segment is always the set pixel material. For example, the pixel material may be a black image. In other words, no matter what the input text is, its corresponding first target video segments are all black photos.
[0052] In one embodiment of the present disclosure, the first video editing data includes at least one subtitle segment, and the target subtitle segment in the at least one subtitle segment corresponds to the target text segment in the at least one text segment. The target subtitle segment is used to embed text subtitles that match the target text segment.
[0053] By adding text subtitles that match the target text segment to the first target video segment, the user can intuitively view the subtitles corresponding to the read-aloud voice while watching the video, improving the user's viewing experience.
[0054] S102: Import the first video editing data into the video editor so that the video editor displays the at least one first video segment and the at least one audio segment on the video editing track of the video editor, and the time interval between the tracks of the first target video segment and the target audio segment is the same.
[0055] In an embodiment of the present disclosure, as shown in FIG. 4, on the video editing page 40 of the video editor, mainly a video preview area 401 and a video editing area 402 are included. A video editing track is displayed in the video editing area 402, and the video editing track includes a video track 403, an audio track 404, and a subtitle track 405. Further, a first video segment is imported into the video track 403, an audio segment is imported into the audio track 404, and a text subtitle is imported into the subtitle track 405. The content imported into the track can be edited in response to an operation on the above video editing track. For example, in the audio track 404, an audio style such as a sweet style or a serious style can be selected. In the audio track 404, the timbre, contrast, etc. can also be edited. That is, all parameters related to the audio can be edited in the audio track 404. Similarly, in the subtitle track 405, all parameters related to the text subtitle such as the color of the subtitle, the display mode of the subtitle, the font of the subtitle, etc. can be edited.
[0056] It should be noted that in the process of editing the above tracks, all the audio segments may be edited, or one of the target audio segments may be edited, which is not specifically limited in the embodiments of the present disclosure. In the process of editing the above tracks, all the subtitle segments may be edited, or one of the target subtitle segments may be edited, which is not specifically limited in the embodiments of the present disclosure.
[0057] In one embodiment of the present disclosure, the interval between the timelines of the tracks of the first target video segment and the target audio segment is the same. For example, their timelines are both segments between 1 minute 55 seconds and 1 minute 57 seconds.
[0058] S103: In response to triggering a second command for the first target video segment with the video editor, based on the first video editing data, embed a first target video into the first target video segment to obtain second video editing data, where the first target video is a video obtained based on first target video material indicated by the second command.
[0059] The first target video segment may be any one of at least one first video segment. Further, the first target video segment may be a video segment that the current user needs to edit.
[0060] In an embodiment of the present disclosure, as shown in FIG. 4, in response to a trigger operation for the first target video segment 406, jump to an image selection page. As shown in FIG. 5, the image selection page 50 includes an image preview area 501 and an image selection area 502. The image selection area 502 further includes a local image control 503, a network image control 504, a stamp control 505, and an image viewing area 506.
[0061] The local image control 503 is used to obtain images in the local album or display videos in the image viewing area in response to a user's trigger operation. The network image control 504 is used to retrieve images or videos from a network or a corresponding database of the client and display them in the image viewing area in response to a user's trigger operation. The stamp control 505 is used to obtain commonly used stamps or popular stamps and display them in the image viewing area in response to a user's trigger operation. The image viewing area 506 is used to display multiple images or short videos in a vertical movement manner in response to a user's vertical slide operation. Further, the image viewing area 506 is used to determine the corresponding image of the trigger operation as the first target pixel material in response to a trigger operation on an image by the user, and display the first target pixel material in the image preview area 501 so that the user can preview it.
[0062] Furthermore, the image selection page 50 further includes a shooting control, and the shooting control is used to call the camera in the terminal to take a picture and obtain the first target pixel material in response to a user's operation.
[0063] Furthermore, the first target pixel material may be a photo or a video. If the first target pixel material is a photo, a video can be obtained by processing the photo using a technology for generating a video from a photo, such as a camera movement effect or a freeze-frame video of the photo. When the first target pixel material is a video, if the time length of the video does not match the time length of the target video segment, the video can be cropped into a video with the same time length as the target video segment. If the time length of the video matches the time length of the target video segment, the video can be directly cropped and embedded into the first target video segment.
[0064] S104: The2-bit DEO Editing Based on the data, the 1-bit DEO is generated.
[0065] In an embodiment of the present disclosure, in response to a trigger operation for video generation, based on the second video editing data, the 1-bit DEO is generated. The trigger operation for video generation can refer to a trigger operation for the export control 407 within the video editing page 40. As an export method, the 1-bit DEO can be saved locally or shared with other video sharing platforms or websites. It is not specifically limited in the embodiments of the present disclosure.
[0066] Based on the above embodiments, the method includes a step of generating third video editing data based on the input text in response to a third command triggered for the input text, where the third video editing data includes at least one second video segment and the at least one audio segment, the at least one second video segment and the at least one audio segment respectively correspond to at least one text segment obtained by splitting the input text, a second target video segment within the at least one second video segment and a target audio segment within the at least one audio segment correspond to a target text segment within the at least one text segment, the second target video segment is a video obtained based on second target pixel materials, and the second target pixel materials match the target text segment; a step of importing the third video editing data into the video editor so as to display the at least one second video segment and the at least one audio segment on a video editing track of the video editor, where an interval between timelines of tracks of the second target video segment and the target audio segment is the same; and the 3-bit DEO Editing Based on the data, the2-bit further includes the step of generating a DEO.
[0067] The third command refers to a command that intelligently matches pixel materials for the input text and further generates first video data. The second target pixel material is matched based on the target text segment and is the pixel material matched to the target text segment by an intelligent matching algorithm.
[0068] In one embodiment of the present disclosure, after obtaining the third video editing data, in response to a video generation trigger operation, a target video is generated based on the edited third video editing data.
[0069] In one embodiment of the present disclosure, after obtaining the third video editing data, the third video editing data can be further imported into a video editor for editing. The specific editing method can refer to the above description of the implementation and will not be described in detail in the embodiments of the present disclosure. After editing the third video editing data, in response to a video generation trigger operation, a target video is generated based on the edited third video editing data.
[0070] In an embodiment of the present disclosure, in response to a first command triggered for an input text, generating first video editing data based on the input text, the first video editing data including at least one first video segment and at least one audio segment, the at least one first video segment and the at least one audio segment respectively corresponding to at least one segmented text segment of the input text, a first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment corresponding to a target text segment in the at least one text segment, the target audio segment being used to be embedded with a spoken voice matching the target text segment, and the first target video segment being an empty segment; importing the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor, the time line interval between tracks of the first target video segment and the target audio segment being the same; in response to triggering a second command for the first target video segment in the video editor, embedding a first target video into the first target video segment based on the first video editing data to obtain second video editing data, the first target video being a video obtained based on first target video material indicated by the second command; and generating a video based on the second video editing data. An embodiment of the present disclosure generates first video editing data based on an input text, whereby in the first video editing data, a user can freely select video material according to their preferences for video editing, meeting the needs of individualized video production. 2-bit Video Editing Based on the data, the 1-bit Video is generated. A video generation method is provided that includes the above steps. An embodiment of the present disclosure generates first video editing data based on an input text, whereby in the first video editing data, a user can freely select video material according to their preferences for video editing, meeting the needs of individualized video production.
[0071] Based on the above embodiments, embodiments of the present disclosure provide several methods for obtaining input text in response to a user's editing operation, specifically as follows. 30 In one embodiment of the present disclosure, the step of obtaining input text in response to a user's editing operation includes: displaying a text input page in response to a trigger operation on the text editing area; and obtaining the corresponding input text of the input operation in response to an input operation on the text input page.
[0072]
[0073] In an embodiment of the present disclosure, as shown in FIG. 3, in response to a trigger operation on the text editing area, a text input page is displayed. As shown in FIG. 6, the text input page 60 includes a title editing area 601, a content editing area 602, and an editing completion control 603.
[0074] Both the title editing area 601 and the content editing area 602 can obtain the text input by the user in response to the user's operation on a virtual keyboard, a physical keyboard, etc. on the terminal. Embodiments of the present disclosure do not limit the above input method. Any method for obtaining input text, such as identifying text in a photo by OTC, identifying the voice input by the user by voice recognition to obtain the input text, etc., is within the protection scope of the present disclosure.
[0075] Furthermore, in response to a trigger operation on the editing completion control 603, the input text in the title editing area 601 and the content editing area 602 is obtained as the input text of the text editing area, and the video production page 30 is jumped to.
[0076] In one embodiment of the present disclosure, the text editing area includes a network address copy control, and the step of obtaining input text in response to a user's editing operation includes displaying a network address input area in response to a trigger operation on the network address copy control, and Input receiving the corresponding network address of the input operation in response to an operation on the network address input area, and obtaining the corresponding input text of the network address.
[0077] In an embodiment of the present disclosure, as shown in FIG. 3, the text editing area 303 includes a network address copy control 304. In response to a trigger operation on the network address copy control, the mask layer area 701 shown in FIG. 7 is displayed, and the mask layer area 701 includes a network address input area 702. Within the mask layer area 701, the user can input an address into the network address input area 702 by means of keyboard input, or can also input an address into the network address input area 702 by means of copy-and-paste.
[0078] Furthermore, in response to a trigger operation on the text acquisition control 703 included in the mask layer area 701, the network address within the network address input area 702 is received, and the content of the corresponding web page of the homepage address is used as the input text within the text editing area 303.
[0079] In one embodiment of the present disclosure, the step of obtaining the corresponding input text of the network address includes determining whether there is original input text within the text editing area, and if there is original input text within the text editing area, deleting the original input text and obtaining the corresponding input text of the network address.
[0080] In an embodiment of the present disclosure, the original input text refers to the input text entered within the text editing area before extracting the content of the corresponding web page of the homepage address. It is determined whether the original input text exists within the text editing area. If the original input text exists within the text editing area, a prompt floating box is displayed, and the prompt floating box is used to prompt the user whether to delete the original input text existing within the text editing area. In response to a trigger operation on the input text deletion control, the original input text is deleted, and the input text corresponding to the network address is obtained. In response to a trigger operation on the cancel control, the original input text is not processed, that is, the original input text remains in the text editing area.
[0081] In an embodiment of the present disclosure, the original input text refers to the input text entered within the text editing area before extracting the content of the corresponding web page of the homepage address. It is determined whether the original input text exists within the text editing area. If the original input text exists within the text editing area, the user is prompted to select the insertion position of the input text corresponding to the network address, and according to the user's selection operation, the input text corresponding to the network address is inserted at the corresponding position of the original input text.
[0082] In one embodiment of the present disclosure, after the step of responding to a trigger operation on the video production control, the steps include obtaining the network address carried on the clipboard, obtaining the input text corresponding to the network address, and displaying a video production page, wherein the video production page includes a text editing area, and the text editing area is used to display the input text corresponding to the network address.
[0083] In one embodiment of the present disclosure, in response to a trigger operation on video production control, the network address carried in the clipboard is detected and acquired. As shown in FIG. 8, when a network address is carried, a prompt box is displayed, and the prompt box is used to display the carried network address and prompt the user whether to identify the content of the corresponding web page of the network address. In response to a trigger operation on the confirmation control 801 in the prompt box, the network address is received, and the content of the web page corresponding to the home page address is used as the input text in the text editing area 303 inside. In response to a trigger operation on the cancel control 802 in the prompt box, the network address is directly ignored, and the text acquisition interface to jumps, and at this time, there is no input text in the text editing area 303 of the text acquisition interface within .
[0084] In the embodiments of the present disclosure, in order to facilitate the user's selection, input methods for various input texts are set according to different situations.
[0085] FIG. 9 is a schematic structural diagram of a video generation device in an embodiment of the present disclosure. This embodiment can be applied to the case of generating a video based on input text, and the video generation device can be realized in the form of software and / or hardware.
[0086] As shown in FIG. 9, the video generation device 90 provided by the embodiment of the present disclosure mainly includes a first video editing data determination module 91, a first video editing data import module 92, a second video editing data determination module 93, and a 1-bit video generation module 94.
[0087] The first video editing data determination module 91 is used to generate first video editing data based on the input text in response to a first command triggered by the input text. The first video editing data includes at least one first video segment and at least one audio segment. The at least one first video segment and the at least one audio segment respectively correspond to at least one divided text segment of the input text. The first target video segment in the at least one first video segment and the target audio segment in the at least one audio segment correspond to the target text segment in the at least one text segment. The target audio segment is used to be embedded with a reading voice that matches the target text segment. The first target video segment is the input text of an empty segment. The first video editing data import module 92 is used to import the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on the video editing track of the video editor. The interval of the time lines of the tracks between the first target video segment and the target audio segment is the same. The second video editing data determination module 93 is used to obtain second video editing data by embedding a first target video into the first target video segment based on the first video editing data in response to triggering a second command for the first target video segment in the video editor. The first target video is a video obtained based on the first target video material indicated by the second command. 1-bit The video generation module 94 is used to generate a video based on the 2-bit video Editing data. 1-bit It is used to generate a video based on the
[0088] In one embodiment of the present disclosure, the first video editing data includes at least one subtitle segment, and a target subtitle segment in the at least one subtitle segment corresponds to a target text segment in the at least one text segment. The target subtitle segment is used to embed a text subtitle that matches the target text segment.
[0089] In one embodiment of the present disclosure, the apparatus is a third video editing data generation module for generating third video editing data based on input text in response to a third command triggered for the input text. The third video editing data includes at least one second video segment and the at least one audio segment. The at least one second video segment and the at least one audio segment respectively correspond to at least one text segment obtained by splitting the input text. A second target video segment in the at least one second video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment. The second target video segment is a video obtained based on a second target video material, and the second target video material matches the target text segment. The third video editing data generation module, and a third video editing data import module for importing the third video editing data into the video editor so as to display the at least one second video segment and the at least one audio segment on a video editing track of the video editor. The third video editing data import module is such that an interval between timelines of tracks of the second target video segment and the target audio segment is the same. And a third 3-bit video Editing Based on the data, a third 2-bit video for generating a 2-bit video generation module, and further includes.
[0090] In one embodiment of the present disclosure, the apparatus further includes a video production page display module for displaying a video production page in response to a trigger operation for video production control. The video production page includes a first control, a second control, and a text editing area. The first control is used to trigger the first instruction in response to a user's trigger operation. The second control is used to trigger the third instruction in response to a user's trigger operation. The text editing area is used to obtain input text in response to a user's editing operation.
[0091] In one embodiment of the present disclosure, the step of obtaining input text in response to a user's editing operation includes: displaying a text input page in response to a trigger operation on the text editing area; and obtaining the corresponding input text of the input operation in response to an input operation on the text input page.
[0092] In one embodiment of the present disclosure, the text editing area includes a network address copy control. The step of obtaining input text in response to a user's editing operation includes: displaying a network address input area in response to a trigger operation on the network address copy control; receiving the corresponding network address of the input operation in response to an operation on the network address input area; and obtaining the corresponding input text of the network address. Input The step of obtaining the corresponding input text of the network address includes: determining whether there is original input text in the text editing area; and if there is original input text in the text editing area, deleting the original input text and obtaining the corresponding input text of the network address.
[0093] In one embodiment of the present disclosure, the step of obtaining the corresponding input text of the network address includes: determining whether there is original input text in the text editing area; and if there is original input text in the text editing area, deleting the original input text and obtaining the corresponding input text of the network address.
[0094] In one embodiment of the present disclosure, after the step of responding to a trigger operation for video production control, the method further includes: obtaining a network address carried on the clipboard; obtaining corresponding input text of the network address; and displaying a video production page, where the video production page includes a text editing area, and the text editing area is used to display the corresponding input text of the network address.
[0095] The video generation device provided by the embodiments of the present disclosure can execute the steps executed in the video generation method provided by the method embodiments of the present disclosure. Specific execution steps and beneficial effects are not described in detail here.
[0096] FIG. 10 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure. Hereinafter, with specific reference to FIG. 10, a schematic structural diagram of an electronic device 1000 applicable to the implementation of the embodiments of the present disclosure is shown. The electronic device 1000 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable terminal devices, and fixed terminals such as digital TVs, desktop computers, and smart home devices. The electronic device shown in FIG. 10 is only an example and does not impose any limitations on the functions and usage ranges of the embodiments of the present disclosure.
[0097] As shown in FIG. 10, the electronic device 1000 may include a processing device (e.g., a central processor, a graphics processor, etc.) 1001, and based on a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003, perform various appropriate operations and processes, and the Video generation The method can be implemented. The RAM 1003 further stores various programs and data necessary for the operation of the terminal device 1000. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0098] Normally, for example, an input device 1006 including a touch panel, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc., an output device 1007 including a liquid crystal display (LCD), a speaker, an oscillator, etc., a storage device 1008 including a magnetic tape, a hard disk, etc., and a communication device 1009 can be connected to the I / O interface 1005. The communication device 1009 Electronic can enable the device 1000 to communicate with other devices wirelessly or wiredly to exchange data. FIG. 10 shows the Electronic device 1000 having various devices, but it should be understood that it is not necessary to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.
[0099] In particular, according to an embodiment of the present disclosure, the process described with reference to the above flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product including a computer program carried on a non-transitory computer-readable medium, the computer program including program code for executing the method shown in the flowchart, thereby implementing the video generation method described above. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 1009, installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions limited to the method of the embodiment of the present disclosure are executed.
[0100] As what needs to be explained, the computer-readable medium described in this disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may be an electrical connection having one or more leads, a portable computer magnetic disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact magnetic disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above, but not limited to these. In this disclosure, the computer-readable storage medium may be any tangible medium that includes or stores a program, and the program may be used in or in combination with an instruction execution system, apparatus, or device. In this disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, and carry computer-readable program code. Such a propagated data signal can use various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may further be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can transmit, propagate, or transmit a program for use in or in combination with an instruction execution system, apparatus, or device. The program code included in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0101] In some embodiments, the client and the server can communicate using any network protocol, such as HTTP (Hyper Text Transfer Protocol), that is currently known or will be developed in the future, and can be connected to each other for digital data communication in any form or medium (e.g., communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the World Wide Web (e.g., the Internet), end-to-end networks (e.g., ad hoc end-to-end networks), and any network that is currently known or will be developed in the future.
[0102] The computer-readable medium may be included in the electronic device or may exist alone and not be assembled within the electronic device.
[0103] One or more programs are carried on the computer-readable medium. When the one or more programs are executed by the terminal device, the terminal device responds to a first instruction triggered by input text, generates first video editing data based on the input text. The first video editing data includes at least one first video segment and at least one audio segment. The at least one first video segment and the at least one audio segment respectively correspond to at least one text segment obtained by splitting the input text. A first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment. The target audio segment is used to be embedded with a synthesized voice that matches the target text segment. The first target video segment is an empty segment. Import the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor. The time line interval between tracks of the first target video segment and the target audio segment is the same. In response to triggering a second instruction for the first target video segment in the video editor, based on the first video editing data, embed a first target video into the first target video segment to obtain second video editing data. The first target video is a video obtained based on a first target video material instructed by the second instruction. Second 2-bit video Editing Based on the data, second 1-bit video is generated.
[0104] Optionally, when the one or more programs are executed by the terminal device, the terminal device can further execute other steps described in the above embodiments.
[0105] Computer program code for performing the operations of the present disclosure can be created in one or more programming languages or combinations thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, C++; and further including conventional procedural programming languages such as the "C" language, or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user computer via any type of network including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., connected via the Internet using an Internet service provider).
[0106] Flowcharts and block diagrams in the drawings illustrate the architecture, functionality, and operation of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram can represent a module, a program section, or a portion of code, and the module, program section, or portion of code includes one or more executable instructions for implementing a given logical function. Note that in some alternative implementations, the functions shown in the blocks may occur in a different order than shown in the figures. For example, two blocks shown consecutively may actually be executed substantially in parallel depending on the relevant functions, or in the reverse order. Also, each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated system based on hardware for performing a given function or operation, or may be implemented by a combination of dedicated hardware and computer instructions.
[0107] The units according to the embodiments of the present disclosure may be implemented in the form of software or in the form of hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0108] In this specification, the functions described above may be executed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used include, but are not limited to, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), etc.
[0109] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or stores a program for use in or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include electrical connections by one or more cables, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact magnetic disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0110] According to one or more embodiments of the present disclosure, the present disclosure includes generating first video editing data based on the input text in response to a first command triggered for the input text, where the first video editing data includes at least one first video segment and at least one audio segment, the at least one first video segment and the at least one audio segment respectively correspond to at least one segmented text segment of the input text, a first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the target audio segment is used to be embedded with a voiceover that matches the target text segment, and the first target video segment is an empty segment; importing the first video editing data into the video editor to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor, where the time line interval between the tracks of the first target video segment and the target audio segment is the same; in response to triggering a second command for the first target video segment in the video editor, embedding a first target video into the first target video segment based on the first video editing data to obtain second video editing data, where the first target video is a video obtained based on first target video material instructed by the second command; and generating a video based on the second video editing data. 2-bit Video Editing Based on the data, the 1-bit step of generating a video is included. A video generation method is provided.
[0111] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, which is used to embed text subtitles in which a target subtitle segment in at least one subtitle segment included in the first video editing data corresponds to a target text segment in at least one text segment and the target subtitle segment matches the target text segment.
[0112] According to one or more embodiments of the present disclosure, the present disclosure includes the step of generating third video editing data based on the input text in response to a third command triggered for the input text, where the third video editing data includes at least one second video segment and the at least one audio segment, the at least one second video segment and the at least one audio segment respectively correspond to at least one text segment obtained by splitting the input text, a second target video segment in the at least one second video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the second target video segment is a video obtained based on a second target video material, and the second target video material matches the target text segment; and the step of importing the third video editing data into the video editor so as to display the at least one second video segment and the at least one audio segment on a video editing track of the video editor, where the time interval between the timelines of the tracks of the second target video segment and the target audio segment is the same; and a step of generating a video based on the video data. 3-bit video Editing Based on the video 2-bit to generate a video.
[0113] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, which includes a step of displaying a video production page in response to a trigger operation for video production control, wherein the video production page includes a first control, a second control, and a text editing area, the first control is used to trigger the first command in response to a user's trigger operation, the second control is used to trigger the third command in response to a user's trigger operation, and the text editing area is used to obtain input text in response to a user's editing operation.
[0114] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method, wherein the step of obtaining input text in response to a user's editing operation includes a step of displaying a text input page in response to a trigger operation for the text editing area, and a step of obtaining the corresponding input text of the input operation in response to an input operation for the text input page.
[0115] According to one or more embodiments of the present disclosure, the text editing area includes a network address copy control in the present disclosure, and the step of obtaining input text in response to a user's editing operation includes a step of displaying a network address input area in response to a trigger operation for the network address copy control, and Input a step of receiving the corresponding network address of the input operation in response to an operation for the network address input area, and a step of obtaining the corresponding input text of the network address.
[0116] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation method including: determining whether the original input text exists in the text editing area before the step of obtaining the corresponding input text of the network address; and when the original input text exists in the text editing area, deleting the original input text and obtaining the corresponding input text of the network address.
[0117] According to one or more embodiments of the present disclosure, the present disclosure further provides a video generation method including: after the step of responding to a trigger operation on video production control, obtaining a network address carried in a clipboard; obtaining the corresponding input text of the network address; and displaying a video production page, where the video production page includes a text editing area, and the text editing area is used to display the corresponding input text of the network address.
[0118] According to one or more embodiments of the present disclosure, the present disclosure provides a first video editing data determination module for generating first video editing data based on input text in response to a first command triggered for the input text, where the first video editing data includes at least one first video segment and at least one audio segment, the at least one first video segment and the at least one audio segment respectively correspond to at least one text segment obtained by splitting the input text, a first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the target audio segment is used to be embedded by a reading voice matching the target text segment, and the first target video segment is an empty segment With A first video editing data determination module, and a first video editing data import module for importing the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor, the first video editing data import module having the same interval between the timelines of the tracks of the first target video segment and the target audio segment, and a second video editing data determination module for embedding a first target video into the first target video segment based on the first video editing data in response to triggering a second command for the first target video segment in the video editor to obtain second video editing data, the first target video being a video obtained based on first target video material indicated by the second command, and a 2-bit video Editing Based on the data, a 1-bit video for generating a 1-bit video generation module, and a video generation apparatus including the same is provided.
[0119] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation apparatus, wherein the first video editing data includes at least one subtitle segment, a target subtitle segment in the at least one subtitle segment corresponds to a target text segment in the at least one text segment, and the target subtitle segment is used to embed a text subtitle that matches the target text segment.
[0120] According to one or more embodiments of the present disclosure, the present disclosure provides a third video editing data generation module for generating third video editing data based on the input text in response to a third command triggered for the input text, wherein the third video editing data includes at least one second video segment and the at least one audio segment, the at least one second video segment and the at least one audio segment respectively correspond to at least one segmented text segment of the input text, a second target video segment in the at least one second video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the second target video segment is a video obtained based on a second target pixel material, and the second target pixel material matches the target text segment; a third video editing data import module for importing the third video editing data into the video editor so as to display the at least one second video segment and the at least one audio segment on a video editing track of the video editor, wherein the time line interval between tracks of the second target video segment and the target audio segment is the same; and a third video generation module for generating a video based on the data. 3-bit ideo Editing Based on the data, a 2-bit ideo for generating a 2-bit ideo generation module, and further provides a video generation device including the same.
[0121] According to one or more embodiments of the present disclosure, the present disclosure further includes a video production page display module for displaying a video production page in response to a trigger operation for video production control. The video production page includes a first control, a second control, and a text editing area. The first control is used to trigger the first command in response to a user's trigger operation. The second control is used to trigger the third command in response to a user's trigger operation. The text editing area is used to obtain input text in response to a user's editing operation. A video generation device is provided.
[0122] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device, and the step of obtaining input text in response to a user's editing operation includes: displaying a text input page in response to a trigger operation on the text editing area; and obtaining the corresponding input text of the input operation in response to an input operation on the text input page.
[0123] According to one or more embodiments of the present disclosure, the text editing area includes a network address copy control in the present disclosure. The step of obtaining input text in response to a user's editing operation includes: displaying a network address input area in response to a trigger operation on the network address copy control; Input receiving the corresponding network address of the input operation in response to an operation on the network address input area; and obtaining the corresponding input text of the network address. A video generation device is provided.
[0124] According to one or more embodiments of the present disclosure, the present disclosure provides a video generation device including: a step of determining whether original input text exists in the text editing area before the step of obtaining the corresponding input text of the network address; and a step of deleting the original input text and obtaining the corresponding input text of the network address when the original input text exists in the text editing area.
[0125] According to one or more embodiments of the present disclosure, the present disclosure further provides a video generation device including: a step of obtaining a network address carried in a clipboard after the step of responding to a trigger operation for video production control; a step of obtaining the corresponding input text of the network address; and a step of displaying a video production page, where the video production page includes a text editing area, and the text editing area is used to display the corresponding input text of the network address.
[0126] According to one or more embodiments of the present disclosure, the present disclosure one or more processors; a memory for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement any of the video generation methods provided by the present disclosure. The present disclosure provides an electronic device including
[0127] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements any of the video generation methods provided by the present disclosure.
[0128] Embodiments of the present disclosure further provide a computer program product including a computer program or instructions that, when executed by a processor, implement the above video generation method.
[0129] The above description is only an explanation of the preferred embodiments of the present disclosure and the technical principles used. As can be understood by those skilled in the art, the scope of the disclosure according to the present disclosure is not limited to the technical solutions formed by specific combinations of the above technical features, and without departing from the above disclosed ideas, other technical solutions formed by any combination of the above technical features or their equivalent features, for example, the above features and technical features having similar functions (but not limited thereto) disclosed in the present disclosure are replaced with each other. The formed technical solution should also be included.
[0130] In addition, although each operation has been described in a specific order, it should not be construed as requiring these operations to be performed in sequence according to the specific order or sequence shown. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although the above discussion includes details of some specific implementations, it should not be construed as a limitation on the scope of the present disclosure. Some features described in the context of a single embodiment may be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may be implemented separately or in any suitable sub-combination in a plurality of embodiments.
[0131] The subject matter has been described in terms specific to structural features and / or logical operations of methods, but it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or operations described above. Rather, the specific features and operations described above are disclosed as exemplary forms for implementing the claims.
Claims
1. A video generation method, comprising: responding to a first instruction triggered by an input text, generating first video editing data based on the input text, wherein the first video editing data includes at least one first video segment and at least one audio segment, the at least one first video segment and the at least one audio segment respectively correspond to at least one divided text segment of the input text, a first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the target audio segment is used to be embedded by a reading voice matching the target text segment, and the first target video segment is an empty segment; importing the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor, wherein an interval between timelines of tracks of the first target video segment and the target audio segment is the same; responding to triggering a second instruction for the first target video segment by the video editor, based on the first video editing data, embedding a first target video into the first target video segment to obtain second video editing data, wherein the first target video is a video obtained based on a first target video material indicated by the second instruction; generating a first video based on the second video editing data. The step of generating first video editing data based on the input text in response to a first instruction triggered by the input text includes: dividing the input text into the at least one text segment; generating at least one first synthesis segment respectively corresponding to the at least one text segment; synthesizing the at least one first synthetic segment according to the order of the at least one text segment in the input text to generate the first video editing data, generating at least one first synthetic segment respectively corresponding to the at least one text segment includes: for each target text segment in the at least one text segment, obtaining a spoken audio corresponding to the target text segment by intelligent matching, aligning the target text segment and the spoken audio on the time line of the track, obtaining a target audio segment corresponding to the target text segment, and obtaining an empty segment as a first target video segment corresponding to the target text segment; A video generation method, including: synthesizing the first target video segment corresponding to the target text segment and the target audio segment corresponding to the target text segment to obtain a first synthetic segment corresponding to the target text segment. **Claim 2** The method according to claim 1, wherein the first video editing data includes at least one subtitle segment, a target subtitle segment in the at least one subtitle segment corresponds to a target text segment in the at least one text segment, and the target subtitle segment is used to embed a text subtitle matching the target text segment. Step 3: In response to a third instruction triggered for the input text, generating third video editing data based on the input text, where the third video editing data includes at least one second video segment and the at least one audio segment, the at least one second video segment and the at least one audio segment respectively corresponding to the at least one text segment into which the input text is divided, a second target video segment within the at least one second video segment and a target audio segment within the at least one audio segment corresponding to a target text segment within the at least one text segment, the second target video segment being a video obtained based on second target pixel material, and the second target pixel material matching the target text segment; Step 4: Importing the third video editing data into the video editor so as to display the at least one second video segment and the at least one audio segment on the video editing track of the video editor, where an interval between time lines of tracks of the second target video segment and the target audio segment is the same; Step 5: Further including generating a second video based on the third video editing data. The method according to claim 1 or 2.
4. Step 6: In response to a trigger operation on video production control, displaying a video production page, where the video production page includes a first control, a second control, and a text editing area, the first control being used to trigger the first instruction in response to the trigger operation, the second control being used to trigger the third instruction in response to the trigger operation, and the text editing area being used to obtain input text in response to an editing operation. The method according to claim 3, further including this step.
5. The step of obtaining input text in response to the editing operation includes: Step 7: In response to a trigger operation on the text editing area, displaying a text input page; The method according to claim 4, comprising: obtaining corresponding input text of the input operation in response to an input operation on the text input page.
6. The text editing area includes a network address copy control, and the step of obtaining input text in response to the editing operation includes: displaying a network address input area in response to a trigger operation on the network address copy control; receiving a corresponding network address of the input operation in response to an input operation on the network address input area; The method according to claim 4, further comprising: obtaining corresponding input text of the network address.
7. The step of obtaining corresponding input text of the network address includes: determining whether original input text exists in the text editing area; The method according to claim 6, further comprising: when the original input text exists in the text editing area, deleting the original input text and obtaining corresponding input text of the network address.
8. After the step of responding to a trigger operation on the video production control, the method further comprises: obtaining a network address carried in the clipboard; obtaining corresponding input text of the network address; displaying the video production page, where the video production page includes the text editing area, and the text editing area is used to display corresponding input text of the network address. The method according to claim 6.
9. A video generation device, A first video editing data determination module for generating first video editing data based on the input text in response to a first command triggered for the input text, wherein the first video editing data includes at least one first video segment and at least one audio segment, the at least one first video segment and the at least one audio segment respectively correspond to at least one segmented text segment of the input text, a first target video segment in the at least one first video segment and a target audio segment in the at least one audio segment correspond to a target text segment in the at least one text segment, the target audio segment is used to be embedded with a spoken voice matching the target text segment, and the first target video segment is an empty segment, and the first video editing data determination module, A first video editing data import module for importing the first video editing data into the video editor so as to display the at least one first video segment and the at least one audio segment on a video editing track of the video editor, and the interval of the time lines of the tracks of the first target video segment and the target audio segment is the same, and the first video editing data import module, A second video editing data determination module for obtaining second video editing data by embedding a first target video into the first target video segment based on the first video editing data in response to triggering a second command for the first target video segment in the video editor, wherein the first target video is a video obtained based on first target video material indicated by the second command, and the second video editing data determination module, A first video generation module for generating a first video based on the second video editing data, and including, When executing to generate first video editing data based on the input text in response to a first command triggered for the input text, the first video editing data determination module, dividing the input text into the at least one text segment; generating at least one first composite segment respectively corresponding to the at least one text segment; used for synthesizing the at least one first composite segment according to the order before and after the at least one text segment in the input text to generate the first video editing data; When executing to generate at least one first composite segment respectively corresponding to the at least one text segment, the first video editing data determination module for each target text segment in the at least one text segment, obtaining a reading voice corresponding to the target text segment by intelligent matching, aligning the target text segment and the reading voice on the time line of the track, obtaining a target audio segment corresponding to the target text segment, and obtaining an empty segment as the first target video segment corresponding to the target text segment; A video generation device used for synthesizing the first target video segment corresponding to the target text segment and the target audio segment corresponding to the target text segment to obtain a first composite segment corresponding to the target text segment.
10. An electronic device, including one or more processors, and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the method according to Claim 1 is realized by the one or more processors.
11. A computer-readable storage medium storing a computer program which, when executed by a processor, realizes the method according to Claim 1.
Citation Information
Patent Citations
Animation video generation method and related device
CN114390220A
System, method, and program for content generation
JP2005062420A
System and method for recording
JP2007328849A
Voice-attached animation production / distribution service system
JP2011082789A
Video editing output control device using text data, video editing output method using text data, and program
JP2021077432A