Video editing methods, apparatus, devices, media, and programs
The video editing method and apparatus facilitate user-friendly integration of virtual objects by aligning text and object segments, using AI-generated content, and displaying real-time editing effects, addressing the inflexibility of existing tools and enhancing video creation with virtual objects.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-09-27
- Publication Date
- 2026-06-24
AI Technical Summary
Existing video editing technologies require professional tools for creating digital human animations, which are inflexible and lack user-friendly interfaces for integrating virtual objects into video editing processes.
A video editing method and apparatus that includes a user-friendly interface with editing tracks for text and virtual object materials, allowing users to align and edit text and virtual object segments, and display the editing effects in real-time through a preview area, using AI-generated text and virtual objects.
Enables users to easily integrate virtual objects for reading text aloud, enhancing video editing flexibility and reducing the barrier to entry for creating professional-quality videos with virtual objects.
Smart Images

Figure 2026520763000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - reference to Related Applications] This application claims priority based on Chinese Patent Application No. 202311278289.9, titled "Video Editing Method, Apparatus, Device, and Medium", filed on September 28, 2023, and the entire disclosure of this application is incorporated herein by reference in its entirety.
[0002] [Technical Field] This disclosure relates to the field of multimedia technology, and particularly to video editing methods, apparatuses, devices, and media.
Background Art
[0003] With the development of computer technology, virtual objects such as digital humans are becoming increasingly popular and have a certain appeal to users. However, in related technologies generally, in order to create digital human animations and directly export them as videos, independent digital human construction tools are required, and the required professional hurdles are high and lack flexibility.
Summary of the Invention
[0004] To solve or at least partially solve the above technical problems, this disclosure provides a video editing method, apparatus, device, and medium.
[0005] Embodiments of the present disclosure provide a video editing method which includes displaying a target video editing interface including an editing area and a video preview area; acquiring target text and acquiring a target virtual object; displaying a text segment formed based on the target text and a video segment formed based on the target virtual object in the editing area, wherein the text segment and the video segment have aligned time axis positions; performing video editing based on the text segment and the video segment; and displaying the video editing effect via the video preview area, wherein the video editing effect is used to display at least a screen and audio of the target virtual object reading the target text, and the screen displays subtitles determined based on the text segment.
[0006] Optionally, the editing area includes a plurality of editing tracks, each of which includes at least a first editing track corresponding to text material and a second editing track corresponding to virtual object material, wherein the text segment is displayed via the first editing track and the video segment is displayed via the second editing track.
[0007] Optionally, the plurality of editing tracks further include a third editing track corresponding to at least one target type material, the target type material being different from the text material and the virtual object material, the track segments formed based on the target type material being displayed via the third editing track, and the aforementioned video editing based on the text segment and the video segment including the track segments corresponding to the text segment, the video segment and the target type material.
[0008] Selectively performing video editing based on the text segment and video segment includes performing video editing based on the editing operation information, the text segment, and the video segment when editing operation information corresponding to the editing area is obtained.
[0009] Optionally, obtaining the target text includes, in response to receiving user-uploaded text, setting the user-uploaded text as the target text, or, in response to receiving user-uploaded information, generating recommended text based on the user-uploaded information and setting the recommended text as the target text.
[0010] Selectively acquiring the target virtual object includes, in response to a virtual object trigger operation, displaying multiple candidate virtual objects, where different candidate virtual objects have different images and / or behaviors, and, in response to a selection operation on the multiple candidate virtual objects, designating the virtual object corresponding to the selection operation as the target virtual object.
[0011] The candidate virtual objects are generated based on a pre-configured object generation algorithm, and the object generation algorithms corresponding to different candidate virtual objects are either the same or different.
[0012] Optionally, displaying video editing effects via the video preview area includes displaying a still image corresponding to the target virtual object in the video preview area for the duration until rendering of the target virtual object is complete, displaying a video image corresponding to the target virtual object in the video preview area after rendering of the target virtual object is complete, and converting the target text into target audio, playing the target audio during the screen display period of the target virtual object, and displaying a screen and sound of the target virtual object reading the target text via the video preview area.
[0013] Optionally, the static screen of the target virtual object may be a pre-set screen, or the static screen of the target virtual object may be a rendering screen of the first frame of the target virtual object, and the video screen of the target virtual object may be a rendering screen of a pre-set action sequence corresponding to the target virtual object.
[0014] Optionally, displaying the target video editing interface includes displaying the target video editing interface in response to a selection operation for a target subtask in a pre-configured video editing task, wherein the video editing task includes multiple subtasks, and different video editing interfaces correspond to different subtasks.
[0015] Optionally, the method further includes obtaining an editing template to be used in the video editing task, and determining and displaying a number of subtasks included in the video editing task based on the editing template.
[0016] Embodiments of the present disclosure further provide a video editing apparatus, the apparatus comprising: an editing interface display module for displaying a target video editing interface including an editing area and a video preview area; a segment display module having aligned time axis positions for acquiring target text and a target virtual object, and for displaying text segments formed based on the target text and video segments formed based on the target virtual object in the editing area, wherein the text segments and video segments are aligned time axis positions; and a video editing module used for performing video editing based on the text segments and video segments and for displaying video editing effects via the video preview area, wherein the video editing effects are used to display at least a screen and audio of the target virtual object reading the target text, and the screen displays subtitles determined based on the text segments.
[0017] Embodiments of the present disclosure further provide electronic equipment comprising a processor and a memory for storing instructions executable by the processor, wherein the processor is used to read the executable instructions from the memory and execute the instructions to implement a video editing method according to embodiments of the present disclosure.
[0018] Embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored therein, which is used to perform a video editing method according to embodiments of the present disclosure.
[0019] Embodiments of the present disclosure further provide a computer program product, which includes a computer program that, when executed by a processor, implements the video editing method.
[0020] The content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following description.
[0021] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments suitable for the present disclosure, and are used together with the specification to interpret the principles of the present disclosure.
[0022] To more clearly explain the technical solutions in the embodiments of the present disclosure or the prior art, the drawings necessary for the description of the embodiments or the prior art will be briefly described below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative efforts.
Brief Description of the Drawings
[0023] [Figure 1] It is a flowchart of a video editing method according to an embodiment of the present disclosure. [Figure 2] It is a schematic diagram of a target video editing interface according to an embodiment of the present disclosure. [Figure 3] It is a schematic diagram of a video preview area according to an embodiment of the present disclosure. [Figure 4] It is a schematic diagram of the configuration of a video editing device according to an embodiment of the present disclosure. [Figure 5] It is a schematic diagram of the configuration of an electronic device according to an embodiment of the present disclosure.
Modes for Carrying Out the Invention
[0024] To more clearly understand the above objects, features, and advantages of the present disclosure, the technical solutions of the present disclosure will be further described below. Unless there is no conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0025] In the following description, many specific details are set forth in order to provide a thorough understanding of the present disclosure. However, the present disclosure may be implemented in other forms different from those described herein. Clearly, the embodiments in the specification are only a part of the embodiments of the present disclosure, not all of them.
[0026] FIG. 1 is a flowchart of a video editing method according to an embodiment of the present disclosure. The method is executable by a video editing device, where the device can be implemented by adopting software and / or hardware and can generally be integrated into an electronic device. As shown in FIG. 1, the method mainly includes the following steps S102 to S106.
[0027] Step S102: Display a target video editing interface including an editing area and a video preview area.
[0028] In actual applications, the target video editing interface may be an initial editing interface directly displayed by video editing software, or may be a unique interface displayed only when the video editing software reaches certain conditions. For example, when monitoring a creation request for a voice broadcast video creation task, display a target video editing interface for realizing the voice broadcast video creation task.
[0029] In some embodiments, the editing area of the target video editing interface includes multiple editing tracks, each including at least a first editing track corresponding to text material and a second editing track corresponding to virtual object material, where text segments are displayed via the first editing track and video segments are displayed via the second editing track. This arrangement facilitates the introduction of virtual objects into the video editing process as material, not only facilitating creation of virtual objects but also enriching the video editing formats. In some specific applications, the virtual object may be a 2D object or a 3D object, and is not limited by the embodiments of this disclosure; for example, the virtual object may be a digital human. In embodiments of this disclosure, the virtual object introduced into the video editing software may be used for reading text aloud, that is, by employing the virtual object to read a specified text on behalf of a real person, a voice broadcast task can be completed quickly and easily.
[0030] The area in the editing area to which multiple editing tracks belong may be called the editing track area, and in addition to the editing track area, the target video editing interface may further include an editing tools area and a material selection area, the editing tools area may be used to display editing controls, and the material selection area may be used to display materials to be selected and / or selected materials. Editing controls include, but are not limited to, operational controls for editing tools such as track editing controls and tool editing controls, and are used to trigger video editing processes in response to user operations. Embodiments of this disclosure do not limit the arrangement of the above-mentioned areas of the target video editing interface, and for example, the video preview area may be set above the editing track area, and the operational controls for specified editing tools may be set below the editing track area, etc.
[0031] In step S104, the target text is obtained and the target virtual object is obtained, and the text segment formed based on the target text and the video segment formed based on the target virtual object are displayed in the editing area, where the text segment and the video segment have aligned time axis positions, in other words, the text segment and the video segment formed based on the target virtual object are aligned on the editing time axis.
[0032] The target text may be text entered by the user, or it may be text automatically generated according to the user's needs, but is not limited thereto. In practical applications, the user may be provided with multiple optional virtual objects from which they can select the target virtual object they need. After the user has determined the target text, a text segment formed based on the target text may be displayed on a first editing track, and after the user has determined the target virtual object, a video segment formed based on the target virtual object may be displayed on a second editing track. In embodiments of this disclosure, since it is desirable to obtain the effect of the target virtual object reading the target text, the text segment and the video segment formed based on the target virtual object have aligned time axis positions, and the time axis positions of the text segment and the video segment formed based on the target virtual object may all overlap.
[0033] In step S106, video editing is performed based on the text segment and video segment, and the video editing effect is displayed via the video preview area. Here, the video editing effect is used to display at least the screen and audio of the target virtual object reading the target text, and the screen displays subtitles determined based on the text segment. This method effectively ensures the user's viewing experience.
[0034] In actual applications, users may only select target text and target virtual objects and not perform any further editing operations. Therefore, if editing operation information corresponding to the editing area is not obtained, video editing can be performed directly based on the text segment and the video segment. On the other hand, some users may perform further editing operations on materials such as text and virtual objects or the corresponding segments as needed. Therefore, if editing operation information corresponding to the editing area is obtained, video editing can be performed based on the editing operation information, text segment, and video segment. Editing operation information includes, but is not limited to, text modification information and / or object modification information. For example, text modification information includes one or more of text content modification information, text style modification information, text font modification information, text size modification information, and text position modification information. For example, object modification information includes one or more of object size modification information, object position modification information, and object image modification information. The above is merely an illustrative description, and any method of editing text or virtual objects is possible and is not limited thereto. In specific embodiments, video editing may be performed based on editing operation information, text segments, and video segments, and the resulting video editing result may be a target video draft, and the user may save the target video draft directly or export a target video based on the target video draft, and the embodiments of this disclosure do not limit the video editing results.
[0035] In practical applications, video editing can be performed based on editing operation information, text segments, and video segments. The video editing effects can be intuitively and directly displayed to the user through the video preview area. Users can clearly see the current edited effect without having to wait to export the video and view the effect, significantly improving the user's video editing experience.
[0036] The user may edit only the text read aloud by the virtual object, or they may add other materials other than text and virtual objects, and fuse the other materials with the virtual object to create rich video editing effects. Therefore, in some embodiments, the multiple editing tracks further include a third editing track corresponding to at least one target type material, and unlike the text material and virtual object material, the target type material may include, but is not limited to, any material necessary for video editing, such as video material, image material, effect material, sticker material, and music material. The track segments formed based on the target type material are displayed via the third editing track. Accordingly, performing video editing based on text segments and video segments includes performing video editing based on the text segments, video segments, and track segments corresponding to the target type material. For example, the user provides background video material and, through multi-track editing, overlays a virtual object on the background video and reads text content related to the background video, thereby giving the user a good viewing experience. To facilitate understanding, please refer to the schematic diagram of the target video editing interface shown in Figure 2, which shows the editing tool area, material selection area, video preview area, and editing track area. Furthermore, the editing track area may show the first editing track, the second editing track, and the third editing track. For example, if the target type material is background video material uploaded by the user, the video segment formed based on this video material is displayed on the third editing track. For ease of distinction, in Figure 2, this video segment is abbreviated as the background video segment, and the video segment formed based on the target virtual object is abbreviated as the object video segment. Based on this, the second editing track may be a picture-in-picture track.Through the interface described above, the effect of spoken text in a scenario corresponding to the background video of a virtual object can be displayed to a person by editing text segments, object video segments, and background video segments. Note that Figure 2 is merely an illustrative example, and in actual applications, the video editing interface may include more or fewer areas than the video editing interface in Figure 2, and the area layout may also differ, without limitation. In actual applications, a second editing track corresponding to the target virtual object can be made a picture-in-picture track, and multiple editing operations can be performed on the target virtual object based on the picture-in-picture function, including, but not limited to, one or more operations such as split, change speed, volume, mixing mode, animation, delete, replace, edit, basic attributes, level, keyframe, cut, copy, filter, adjust, beautify / body care, mask, audio separation, opacity, reverse playback, freeze frame, and voice change, greatly enriching the variety of editing the user can do on the virtual object.
[0037] In summary, the above-described technical solution according to the embodiments of this disclosure provides a user-friendly video editing interface (a target video editing interface including an editing track area and a video preview area, wherein the editing track area includes multiple editing tracks, including at least a first editing track corresponding to text material and a second editing track corresponding to virtual object material). Users can flexibly and conveniently use the tracks in the editing track area to edit the necessary virtual objects and text, and can introduce virtual objects for reading text aloud as material into video editing. Furthermore, the effect of the virtual object reading text aloud can be intuitively displayed through the video preview area. This contributes to users clearly knowing and flexibly adjusting the editing effects before finally generating the video, significantly lowering the barrier to creation using virtual objects, enabling more users to flexibly and conveniently create videos using virtual objects, and enriching the video creation formats.
[0038] The embodiments of this disclosure provide two embodiments for obtaining target text, and the target text can be obtained by referring to Method 1 or Method 2 below.
[0039] Method 1: In response to receiving user-uploaded text, the user-uploaded text is used as the target text.
[0040] For example, text uploaded by the user may be retrieved via a text input control and directly used as target text to be read aloud by the target virtual object required by the user. Such a method effectively ensures that the text content read aloud by the target virtual object perfectly matches the user's needs.
[0041] Method 2: In response to receiving user upload information, a recommendation text is generated based on the user upload information, and this recommendation text is used as the target text.
[0042] For example, information uploaded by the user may be retrieved through an information upload control. User-uploaded information includes, but is not limited to, text, voice, image, and video formats. Here, user-uploaded information in text and voice formats only needs to be expressed using keyword phrases. Based on the user-uploaded information, embodiments of this disclosure may employ an artificial intelligence model, such as a language model, to automatically generate recommendation text that matches the user-uploaded information. For example, if the user-uploaded information is a picture of a cup, the recommendation text may be an advertising slogan corresponding to the cup in the picture. Alternatively, if the user-uploaded information is keywords for a landscape, the recommendation text may be a detailed description of the landscape. Alternatively, if the user-uploaded information is a knowledge point, the recommendation text may be a detailed introduction to the knowledge point. This method can appropriately save the user's time and effort and provide the user with text that meets their needs. In actual applications, users may be provided with text modification controls to modify the recommendation text.
[0043] Furthermore, embodiments of the present disclosure provide embodiments for obtaining a target virtual object, which can be performed by referring to the following steps (1) and (2).
[0044] Step (1) In response to a virtual object trigger operation, multiple candidate virtual objects are displayed, with different candidate virtual objects having different images and / or behaviors.
[0045] In actual applications, an object selection control may be set up, and when the object selection control is triggered, it is determined that a virtual object trigger operation has been monitored, and at this time, multiple candidate virtual objects may be displayed on the interface. The candidate virtual objects are generated based on a pre-set object generation algorithm, and the object generation algorithms corresponding to different candidate virtual objects may be the same or different, and even if the same object generation algorithm is used, the parameters corresponding to different candidate virtual objects may be different, so there may be differences in the resulting image or operation.
[0046] Step (2) In response to a selection operation on multiple candidate virtual objects, the candidate virtual object corresponding to the selection operation is set as the target virtual object.
[0047] The user may apply the target virtual object to video editing by selecting one or more virtual objects from a group of candidate virtual objects as the target virtual object, depending on their needs.
[0048] To effectively ensure the user's video editing experience, embodiments of the present disclosure further provide embodiments that display video editing effects through a video preview area, which can be carried out with reference to steps A and B below.
[0049] Step A: During the period until the rendering of the target virtual object is complete, a still image corresponding to the target virtual object is displayed in the video preview area. After the rendering of the target virtual object is complete, a video image corresponding to the target virtual object is displayed in the video preview area.
[0050] Given that rendering virtual objects generally takes a long time, to avoid user waiting times, a still image corresponding to the target virtual object may be displayed to the user preferentially. Rendering may be completed quickly during the display period of the still image, and once rendering of the target virtual object is complete, the still image may be switched to a video image corresponding to the target virtual object. For example, the still image of the target virtual object may be a pre-configured screen, or the still image of the target virtual object may be the rendering screen of the first frame of the target virtual object, while the video image of the target virtual object may be the rendering screen of a pre-configured action sequence corresponding to the target virtual object. The pre-configured action sequence may be automatically generated based on an object generation algorithm corresponding to the target virtual object, or it may be manually configured and input into the object generation algorithm; however, this is not limited to this method. This method allows the user to view the editing effects of the virtual object in a timely manner during the video editing period without waiting or exporting the video, thus ensuring a proper user experience.
[0051] Step B: The target text is converted to target audio, and the target audio is played during the screen display period of the target virtual object, displaying the screen and audio of the target virtual object reading the target text via the video preview area, and the screen also displays subtitles determined based on the text segment.
[0052] In practical applications, by employing TTS (Text To Speech) technology to quickly convert target text into target audio and playing it back in sync with the target virtual object's screen, the effect of the target virtual object reading the target text aloud can be displayed. Note that the execution order of steps A and B is not limited, and in practical applications, steps A and B may be executed simultaneously.
[0053] For easier understanding, please refer to the schematic diagram of the video preview area shown in Figure 3, which illustrates that the virtual object is a digital human and the target type material is a picture, and that the text read aloud in real time by the digital human is simultaneously displayed on the interface. If the rendering of the digital human is not yet complete, a still image corresponding to the digital human is used and superimposed on the background image to show the user the effect of a still digital human. However, since the display time of the still digital human is generally short, once the rendering of the digital human is complete, a video image corresponding to the digital human may be used and superimposed on the background image to show the user the effect of a dynamic digital human. In all of the above cases, playback of audio converted from the target text is performed, and as a whole, the user is shown the effect of video playback in which the digital human explains the product in real time.
[0054] To facilitate quick editing by the user, embodiments of the present disclosure can provide the user with an editing template, which can divide video editing into multiple submodules, specifically, a video editing task can be divided into multiple subtasks, each with a different video editing interface corresponding to a different subtask, allowing the user to conveniently edit each subtask video and combine the video editing results corresponding to each subtask to ultimately obtain the desired video. Based on this, the method according to embodiments of the present disclosure further includes obtaining an editing template to be used in a video editing task, and determining and displaying the multiple subtasks included in the video editing task based on the editing template.
[0055] The video editing task may be a user-initiated video editing task, or it may be, for example, an advertising editing task. Embodiments of this disclosure may provide multiple editing templates in advance, and different editing templates may be used to perform different video editing tasks. The editing templates may specify a particular way of subdividing the video editing task into subtasks and provide the corresponding editing components. Exemplaryly, subtasks may be a voice-over video editing subtask, a voice-broadcast video editing subtask, a video display subtask, etc. Here, the main commentary content in the voice-over video editing subtask is derived from narration audio, and the video screen can jump according to the audio content. In the voice-broadcast video editing subtask, a virtual object or real person must appear and read text, and a picture or video material may be further attached to the background. The video display subtask may display only the video without narration or commentary by a virtual object / real person, and all of the above are examples, and embodiments of this disclosure do not limit the type of subtask. The above method allows for the breakdown of video editing tasks as needed and the rapid editing of each subtask. Different video editing interfaces may correspond to different subtasks, and the editing components provided may also differ. Based on this, in some embodiments of displaying a target video editing interface, the target video editing interface may be displayed in response to a selection operation on a target subtask in a pre-configured video editing task. Exemplarily, the target subtask may be the aforementioned voice broadcast video editing subtask. That is, only when the voice broadcast video editing subtask is triggered, a target video editing interface including a first editing track corresponding to text material and a second editing track corresponding to virtual object material is displayed to the user, allowing the user to edit through the interface, thereby realizing a voice broadcast effect in which a virtual object reads text aloud.
[0056] In summary, the video editing method according to the embodiments of this disclosure provides a user-friendly video editing interface, allowing users to flexibly and conveniently edit necessary virtual objects and text, introduce virtual objects for reading text aloud as material into video editing, and intuitively display the effect of virtual objects reading text aloud through the video preview area. This also contributes to users clearly knowing and flexibly adjusting the editing effect before finally generating the video, significantly lowering the barrier to creation using virtual objects, enabling more users to flexibly and conveniently use virtual objects to create videos, and enriching the video creation formats.
[0057] In accordance with the video editing method described above, embodiments of the present disclosure further provide a video editing apparatus, Figure 4 being a schematic diagram of the configuration of a video editing apparatus according to an embodiment of the present disclosure, which can be implemented by software and / or hardware and can generally be integrated into electronic equipment, as shown in Figure 4, the video editing apparatus is An editing interface display module 402 for displaying a target video editing interface including an editing area and a video preview area, It is used to acquire target text and target virtual object, and to display text segments formed based on the target text and video segments formed based on the target virtual object in the editing area, where the text segments and video segments are segment display modules 404 having aligned time axis positions, The system includes a video editing module 406 that performs video editing based on text segments and video segments, and displays the video editing effects via a video preview area, wherein the video editing effects are used to display at least a screen and audio of a target virtual object reading target text, and the screen displays subtitles determined based on the text segments.
[0058] The apparatus according to the embodiments of this disclosure provides a user-friendly video editing interface (a target video editing interface including an editing area and a video preview area, the editing area being capable of displaying text segments formed based on target text and video segments formed based on target virtual objects), allowing users to flexibly and conveniently edit necessary virtual objects and text, introduce virtual objects for reading text aloud as material into video editing, and intuitively display the effect of virtual objects reading text aloud via the video preview area. This also contributes to users clearly knowing and flexibly adjusting the editing effect before finally generating the video, significantly lowering the barrier to creation using virtual objects, enabling more users to flexibly and conveniently use virtual objects to create videos, and enriching video creation formats.
[0059] In some embodiments, the editing area includes a plurality of editing tracks, the plurality of editing tracks including at least a first editing track corresponding to text material and a second editing track corresponding to virtual object material, wherein the text segment is displayed via the first editing track and the video segment is displayed via the second editing track.
[0060] In some embodiments, the plurality of editing tracks further include a third editing track corresponding to at least one target type material, the target type material being different from the text material and the virtual object material, the track segments formed based on the target type material being displayed via the third editing track, and the video editing module 406 being used specifically to perform video editing based on the text segments, the video segments and the track segments corresponding to the target type material.
[0061] In some embodiments, the video editing module 406 is specifically used to perform video editing based on the editing operation information, the text segment, and the video segment when it obtains editing operation information corresponding to the editing area.
[0062] In some embodiments, the segment display module 404 is specifically used to target user-uploaded text in response to receiving user-uploaded text, or to generate recommendation text based on user-uploaded information in response to receiving user-uploaded information, and to target the recommendation text.
[0063] In some embodiments, the segment display module 404 is used to display a plurality of candidate virtual objects in response to a virtual object trigger operation, where different candidate virtual objects have different images and / or behaviors, and to set the virtual object corresponding to a selection operation as the target virtual object in response to a selection operation on the plurality of candidate virtual objects.
[0064] In some embodiments, the candidate virtual objects are generated based on a pre-configured object generation algorithm, and the object generation algorithms corresponding to different candidate virtual objects are the same or different.
[0065] In some embodiments, the video editing module 406 is used to display a still image corresponding to the target virtual object in the video preview area during the period until the rendering of the target virtual object is complete, to display a video image corresponding to the target virtual object in the video preview area after the rendering of the target virtual object is complete, to convert the target text into target audio, and to play the target audio during the screen display period of the target virtual object, thereby displaying a screen and sound of the target virtual object reading the target text via the video preview area.
[0066] In some embodiments, the static screen of the target virtual object is a pre-configured screen, or the static screen of the target virtual object is a rendering screen of the first frame of the target virtual object, and the video screen of the target virtual object is a rendering screen of a pre-configured action sequence corresponding to the target virtual object.
[0067] In some embodiments, the editing interface display module 402 is specifically used to display a target video editing interface in response to a selection operation for a target subtask in a pre-configured video editing task, where the video editing task includes multiple subtasks, and different video editing interfaces correspond to different subtasks.
[0068] In some embodiments, the apparatus further includes a subtask determination module for obtaining an editing template used in the video editing task, and determining and displaying a plurality of subtasks included in the video editing task based on the editing template.
[0069] A video editing apparatus according to an embodiment of the present disclosure can perform a video editing method according to any embodiment of the present disclosure and has corresponding functional modules and beneficial effects for performing the method.
[0070] As will be obvious to those skilled in the art, for the sake of ease and conciseness of explanation, the specific operating processes of the embodiments of the apparatus described above can be referenced to the corresponding processes in the embodiments of the method and will not be described again here.
[0071] Embodiments of the present disclosure provide electronic devices comprising a storage device in which a computer program is stored, and a processing device for executing the computer program in the storage device to perform a step of any one of the methods of the present disclosure.
[0072] Referring below to Figure 5, which shows a schematic diagram of the configuration of an electronic device 500 suitable for realizing an embodiment of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital voice broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 5 is merely an example and does not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0073] As shown in Figure 5, the electronic device 500 may include a processing unit (e.g., a central processing unit, an image processing unit, etc.) 501 which can perform various appropriate operations and processes based on programs stored in read-only memory (ROM) 502 or programs loaded from storage device 508 into random access memory (RAM) 503. RAM 503 further stores various programs and data necessary for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0074] Generally, devices such as input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, and gyroscopes; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, and vibrators; storage devices 508 including, for example, magnetic tape and hard disks; and communication devices 509 can be connected to the I / O interface 505. The communication device 509 can allow the electronic device 500 and other devices to exchange data via wireless or wired communication. Figure 5 shows an electronic device 500 with various devices, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have all of the devices instead.
[0075] In particular, according to embodiments of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product which includes a computer program placed on a non-temporary computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such embodiments, the computer program may be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, the above-described functions, limited by the methods of embodiments of the present disclosure, are performed.
[0076] In addition to the methods and apparatus described above, embodiments of the present disclosure may also be computer program products including computer program instructions, which, when executed by a processor, cause the processor to execute the image processing method according to the embodiments of the present disclosure. The computer program product can be composed of program code for executing the operations of the embodiments of the present disclosure in any combination of one or more programming languages, and the programming languages include object-oriented programming languages such as Java® and C++, and further include conventional procedural programming languages such as the C language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or fully on a remote computing device or server.
[0077] Furthermore, the embodiments of this disclosure may also be computer-readable storage media on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the processor is made to execute the video editing method according to the embodiments of this disclosure.
[0078] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any combination of more than these. More specific examples of readable storage media (a non-exhaustive list) include electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0079] Embodiments of the present disclosure further provide a computer program product which includes a computer program / instruction that, when executed by a processor, realizes the video editing method of the embodiment of the present disclosure.
[0080] Before using any of the technical proposals disclosed in each embodiment of this disclosure, it is understood that the user should be informed in an appropriate manner, in accordance with applicable laws and regulations, of the type of personal information related to this disclosure, the scope of use, and the usage scenarios, and permission should be obtained from the user.
[0081] For example, in response to receiving a voluntary request from the user, prompt information is sent to the user to clearly prompt the user that the operation requested to be performed requires the acquisition and use of the user's personal information. This allows the user to choose whether or not to provide personal information to software or hardware such as electronic devices, application programs, servers, or storage media that perform the operation of the proposed technical method of this disclosure based on the prompt information.
[0082] As a selective but non-restrictive implementation, a method for sending prompt information to a user in response to receiving a voluntary request from the user may be, for example, a pop-up window, and the prompt information may be presented in text format within the pop-up window. Furthermore, the pop-up window may include a selection control that allows the user to choose whether to "agree" or "disagree" to providing personal information to the electronic device.
[0083] The above notice and user permission acquisition process are merely schematic and do not limit the forms in which this disclosure may be implemented. It should be understood that other methods that comply with applicable laws and regulations may also be applicable to the implementation of this disclosure.
[0084] In this specification, relational terms such as “first” and “second” are merely used to distinguish one entity or operation from another, and do not necessarily require or suggest that any such actual relationship or order exists between these entities or operations. Furthermore, the terms “include,” “incorporate,” or any other variations thereof are intended to cover non-exclusive inclusion, thereby meaning that a process, method, article, or apparatus that includes a set of elements includes not only those elements but also other elements not explicitly listed, or elements specific to such a process, method, article, or apparatus. Unless further limited, the elements limited by the phrase “include one…” do not preclude the existence of other identical elements in a process, method, article, or apparatus that includes such elements.
[0085] The above description is merely a set of specific embodiments of the Disclosure intended to enable those skilled in the art to understand or implement the Disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can also be implemented in other embodiments without departing from the spirit or scope of the Disclosure. Therefore, the Disclosure is not limited to these embodiments described herein, but rather conforms to the broadest extent to which the principles and novel features disclosed herein are consistent.
Claims
1. A video editing method, To display the target video editing interface, including the editing area and the video preview area, The process involves acquiring target text and target virtual objects, and displaying text segments formed based on the target text and video segments formed based on the target virtual objects in the editing area, wherein the text segments and video segments have aligned time axis positions. A video editing method comprising performing video editing based on the text segment and the video segment, and displaying the video editing effect through the video preview area, wherein the video editing effect is used to display at least a screen and audio of the target virtual object reading the target text, and the screen displays subtitles determined based on the text segment.
2. The method according to claim 1, wherein the editing area includes a plurality of editing tracks, the plurality of editing tracks including at least a first editing track corresponding to text material and a second editing track corresponding to virtual object material, wherein the text segment is displayed via the first editing track and the video segment is displayed via the second editing track.
3. The plurality of editing tracks further include a third editing track corresponding to at least one target type material, and unlike the text material and the virtual object material, the track segments formed based on the target type material are displayed via the third editing track. The method according to claim 2, wherein performing video editing based on the text segment and the video segment includes performing video editing based on the text segment and the video segment and the track segment corresponding to the target type material.
4. Performing video editing based on the text segment and the video segment is, The method according to claim 1, further comprising performing video editing based on the editing operation information, the text segment, and the video segment when editing operation information corresponding to the editing area is obtained.
5. Obtaining the aforementioned target text means In response to receiving user-uploaded text, the user-uploaded text is set as the target text. Or, The method according to claim 1, comprising: in response to receiving user-uploaded information, generating recommendation text based on the user-uploaded information, and setting the recommendation text as the target text.
6. Obtaining the aforementioned target virtual object means In response to a virtual object trigger operation, multiple candidate virtual objects are displayed, and the image and / or behavior of different candidate virtual objects differs. The method according to claim 1, further comprising setting a candidate virtual object corresponding to a selection operation as a target virtual object in response to a selection operation on the plurality of candidate virtual objects.
7. The method according to claim 6, wherein the candidate virtual object is generated based on a pre-configured object generation algorithm, and the object generation algorithms corresponding to different candidate virtual objects are the same or different.
8. Displaying the video editing effects via the aforementioned video preview area means that During the period until the rendering of the target virtual object is completed, a still image corresponding to the target virtual object is displayed in the video preview area, and after the rendering of the target virtual object is completed, a video image corresponding to the target virtual object is displayed in the video preview area. The method according to claim 1, further comprising converting the target text into target audio, playing the target audio during the screen display period of the target virtual object, and displaying a screen and sound of the target virtual object reading the target text via the video preview area.
9. The static screen of the target virtual object is a pre-configured screen, or the static screen of the target virtual object is the rendering screen of the first frame of the target virtual object. The method according to claim 8, wherein the video screen of the target virtual object is a rendering screen of a pre-set operation sequence corresponding to the target virtual object.
10. Displaying the aforementioned target video editing interface means The method according to claim 1, comprising displaying a target video editing interface in response to a selection operation for a target subtask in a pre-configured video editing task, wherein the video editing task comprises a plurality of subtasks, and different video editing interfaces corresponding to different subtasks are different.
11. The aforementioned method, Obtaining the editing template used in the aforementioned video editing task, The method according to claim 10, further comprising determining and displaying the plurality of subtasks included in the video editing task based on the editing template.
12. A video editing device, An editing interface display module for displaying a target video editing interface including an editing area and a video preview area, A segment display module used to acquire target text and target virtual objects, and to display text segments formed based on the target text and video segments formed based on the target virtual object in the editing area, wherein the text segments and video segments have aligned time axis positions, A video editing apparatus comprising: a video editing module used to perform video editing based on the text segment and the video segment, and to display the video editing effect via the video preview area, wherein the video editing effect is used to display at least a screen and sound of the target virtual object reading the target text, and the screen displays subtitles determined based on the text segment.
13. It is an electronic device, A memory device in which a computer program is stored, Electronic device comprising a processing unit for executing the computer program in the storage device to realize the steps of the video editing method described in any one of claims 1 to 11.
14. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and the computer program is used to perform the video editing method described in any one of claims 1 to 11.
15. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, realizes the video editing method described in any one of claims 1 to 11.