Video editing method, apparatus, device, and medium

By displaying target videos captured by multiple user devices in the video editing interface and creating video track segments on the editing track based on editing templates, the high cost of online collaborative video production is solved, and convenient video editing is achieved.

CN118741243BActive Publication Date: 2026-03-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing online collaborative video production methods are costly, require users to use multiple devices, and are complex to operate, making editing difficult.

Method used

A video editing method is provided, which acquires target videos captured by multiple user devices and displays them in a video editing interface based on an editing template. The video editing interface includes a preview playback area and an editing track area. Images from multiple target videos are respectively formed into video track segments on the editing track, and the playback state can be adjusted using video operation controls.

Benefits of technology

It reduces the production cost of online collaborative videos, allowing users to easily edit videos from multiple users and improving operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118741243B_ABST
    Figure CN118741243B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video editing method, device, equipment and medium, the method comprising: obtaining a plurality of target videos, the plurality of target videos comprising a first video obtained by different user devices respectively shooting and recording for a same recording task; obtaining an editing template corresponding to the plurality of target videos; the editing template comprising display position information of the plurality of target videos and track information of the plurality of target videos; based on the plurality of target videos and the editing template, a video editing interface is displayed; the video editing interface comprises a preview playing area and an editing track area, the editing track area comprises a plurality of video editing tracks, images of the plurality of target videos are respectively displayed in display areas indicated by respective display position information of the plurality of target videos in the preview playing area, and each target video in the plurality of target videos forms a video track segment on one video editing track in the plurality of video editing tracks based on the track information. The present disclosure can effectively reduce the cost of online co-creation videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of video processing technology, and in particular to a video editing method, apparatus, device and medium. Background Technology

[0002] Nowadays, more and more users are no longer satisfied with conventional video creation methods, and online collaborative video creation is gradually emerging. Methods such as simultaneous game recording, online duets, and livestream battles can all be considered online collaborative video creation. While online collaborative video creation does not require multiple users to gather in the same place to shoot, it is costly to produce and lacks convenience. Summary of the Invention

[0003] To solve the above-mentioned technical problems, or at least partially solve them, this disclosure provides a video editing method, apparatus, device, and medium.

[0004] This disclosure provides a video editing method, the method comprising: acquiring multiple target videos, the multiple target videos including first videos recorded by different user devices for the same recording task; acquiring editing templates corresponding to the multiple target videos; the editing templates including display position information and track information of the multiple target videos; and displaying a video editing interface based on the multiple target videos and the editing templates; wherein the video editing interface includes a preview playback area and an editing track area, the editing track area including multiple video editing tracks, the images of the multiple target videos being respectively presented in the preview playback area in the display area indicated by the display position information of the multiple target videos, and each of the multiple target videos forming a video track segment on one of the multiple video editing tracks based on the track information, the timeline positions of the video track segments of the multiple target videos partially or completely overlapping.

[0005] Optionally, the target video may further include a third video obtained based on the second video, wherein the second video is a video played by the different user devices during the shooting and recording process for the recording task, and the third video is used to present the playback status of the second video during the shooting and recording process.

[0006] Optionally, the different user devices include a target user device; the user interface provided by the target user device for the recording task displays video operation controls in a triggerable state; the user interfaces provided by the other user devices besides the target user device for the recording task do not display the video operation controls in a triggerable state; when the video operation controls are in a triggerable state, in response to detecting a user trigger operation on the target user device for the video operation controls, the playback state of the second video is adjusted according to the user trigger operation.

[0007] Optionally, the different user devices include a first user device and a second user device, wherein the first user device is a device that imports the second video for the recording task; when the recording task is not closed and the first user device has not exited the recording task, the target user device is the first user device; when the recording task is not closed and the first user device has exited the recording task, the target user device is the second user device.

[0008] Optionally, the video editing track corresponding to the third video is the main track, and the video editing track corresponding to the first video is the picture-in-picture track.

[0009] Optionally, if the plurality of target videos contain only the first video, the video editing track corresponding to the first video captured by the third user device among the different user devices is the main track, and the video editing tracks corresponding to the first videos captured by other user devices among the plurality of target videos are picture-in-picture tracks; wherein, the third user device is the device that initiates the recording task.

[0010] Optionally, in response to receiving an editing request from a fourth user device among the different user devices, a template selection page is displayed; wherein the template selection page includes multiple video layout effect images, each video layout effect image includes multiple areas, each of the multiple areas corresponds to a video display position and a video editing track; in response to detecting a selection operation for a target effect image among the multiple video layout effect images, the editing template corresponding to the target effect image is determined as the editing template corresponding to the multiple target videos.

[0011] Optionally, determining the editing template corresponding to the plurality of target videos based on the editing template corresponding to the target effect image includes: displaying a layout preview image of the plurality of target videos based on the target effect image selected by the target user among the plurality of video layout effect images and the plurality of target videos; and determining the editing template corresponding to the plurality of target videos based on the target user's adjustment operation on the display position of the target videos in the layout preview image.

[0012] Optionally, before acquiring multiple recorded videos, the method further includes: in response to receiving a duet request from a target user device, creating a recording task and displaying a user interface corresponding to the recording task on the target user device; wherein the user interface displays a video adding control and a user invitation control; in response to detecting that the video adding control is triggered, acquiring video information uploaded by the target user device and obtaining a second video based on the video information; wherein the video information includes a local video file and / or a network video link; in response to detecting that the user invitation control is triggered, generating invitation information for the recording task, so that the target user device can send the invitation information to a designated user device, the invitation information being used to prompt the designated user device to join the recording task.

[0013] Optionally, the user interface of both the target user device and other user devices performing the recording task displays a first area and a second area; wherein, the first area is used to display the image of the first video captured and recorded by each of the user devices, and the second area is used to display the image of the second video.

[0014] Optionally, the user interface of the target user device further displays a start recording control; if the target user device does not trigger the video add control or the user invite control, the start recording control is in an untriggerable state; wherein, the start recording control in the untriggerable state cannot be triggered to record video; if the target user device triggers the video add control and / or the target user device triggers the user invite control, the start recording control is in a triggerable state; wherein, the start recording control in the triggerable state is used to start recording video when triggered.

[0015] Optionally, acquiring multiple recorded videos includes: in response to detecting that a start recording control in a triggerable state is triggered, recording based on frame images captured by the front-facing camera of the user device performing the recording task to obtain a first video; if a second video is acquired through the target user device, synchronously playing the second video during the video recording process through the user device performing the recording task, and recording the playback process of the second video to obtain a third video.

[0016] This disclosure also provides a video editing device, comprising: a video acquisition module for acquiring multiple target videos, the multiple target videos including first videos recorded by different user devices for the same recording task; a template acquisition module for acquiring editing templates corresponding to the multiple target videos; the editing templates including display position information and track information of the multiple target videos; and a video editing module for displaying a video editing interface based on the multiple target videos and the editing templates; wherein the video editing interface includes a preview playback area and an editing track area, the editing track area including multiple video editing tracks, the images of the multiple target videos are respectively presented in the preview playback area in the display area indicated by the display position information of the multiple target videos, and each of the multiple target videos forms a video track segment on one of the multiple video editing tracks based on the track information, and the timeline positions of the video track segments of the multiple target videos partially or completely overlap.

[0017] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the video editing method provided in this disclosure.

[0018] This disclosure also provides a computer-readable storage medium storing a computer program for performing the video editing method provided in this disclosure.

[0019] The technical solution provided in this disclosure acquires multiple target videos (including first videos recorded by different user devices for the same recording task) and obtains editing templates corresponding to the multiple target videos (including display position information and track information of the multiple target videos). Then, a video editing interface can be directly displayed based on the multiple target videos and editing templates. This video editing interface includes a preview playback area and an editing track area. The editing track area contains multiple video editing tracks. The images of the multiple target videos are respectively presented in the display areas indicated by the display position information of each target video in the preview playback area. Each target video forms a video track segment on one of the multiple video editing tracks based on the track information. The timeline positions of the video track segments of the multiple target videos partially or completely overlap. Through this method, videos recorded by multiple users for the same recording task can be edited conveniently and quickly, effectively reducing the cost of online co-creation of videos.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0022] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 A flowchart illustrating a video editing method provided in an embodiment of this disclosure;

[0024] Figure 2 A schematic diagram of a user interface provided for an embodiment of this disclosure;

[0025] Figure 3 A schematic diagram of a user interface provided for an embodiment of this disclosure;

[0026] Figure 4 These are schematic diagrams illustrating various layout effects provided in the embodiments of this disclosure;

[0027] Figure 5 A flowchart of a video editing method provided in this disclosure embodiment;

[0028] Figure 6 This is a schematic diagram of the structure of a video editing device provided in an embodiment of the present disclosure;

[0029] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0030] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0031] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0032] The inventors discovered through research that existing online collaborative video creation methods are costly to produce. For example, if multiple users need to record their online interactions or different users' viewing of the same video, each user needs at least two devices (such as a mobile phone and a computer simultaneously) to record, project, and interact online simultaneously. Furthermore, multiple software tools need to be running concurrently, which is not only time-consuming and labor-intensive, but also only produces a screen recording, making further editing difficult. Additionally, directly recording each user's video and then integrating multiple videos requires each user to transmit their recording to a designated user for editing, which is also time-consuming, labor-intensive, and costly. To address at least one of these problems, this disclosure provides a video editing method, apparatus, device, and medium, which are described in detail below.

[0033] Figure 1 This is a flowchart illustrating a video editing method provided in an embodiment of the present disclosure. The method can be executed by a video editing device, which can be implemented using software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method mainly includes the following steps S102 to S106:

[0034] Step S102: Acquire multiple target videos, which include first videos captured and recorded by different user devices for the same recording task.

[0035] User devices can be mobile phones, computers, wearable devices, etc., and are not limited herein. In practical applications, there can be multiple user devices, and each user can use only one user device. When multiple users perform the same recording task, they can use their respective devices to shoot and record, thereby obtaining a first video corresponding to each device. This disclosure does not limit the recording task. For example, the recording task can be to use the front-facing camera of the user device to shoot the user's behavior, for example, to shoot the user's behavior when interacting online; or the recording task can be to use the user device to shoot the scene in which the user is located, etc. The specific recording task can be flexibly set according to the needs, and this disclosure does not impose any restrictions.

[0036] It should be noted that in practical applications, the target video may include not only the first video. For example, the target video may also include a third video derived from the second video, where the second video is played on different user devices during the recording process for the recording task, and the third video is used to present the playback status of the second video during the recording process. That is, the target video may also include a third video derived from the second video played synchronously on different user devices during the recording process. This third video can present the playback status of the second video during the recording process, so that it can be used in video editing together with the first video in later stages.

[0037] For example, multiple users can simultaneously watch a second video, meaning each user's device plays the second video synchronously, while each user's viewing behavior is recorded using their front-facing camera. The second video can be determined by the initiator of the recording task (also known as the creator), for example, by obtaining a second video file or link uploaded by the creator, and then played synchronously on all user devices. During the playback of the second video, its playback status can be adjusted and controlled, such as speeding up playback, slowing down the video, or rewinding the video; there are no restrictions on this. The resulting third video can simply be the second video, and can be further associated with the playback control information of the second video so that the playback status of the second video can be restored based on the playback control information when the third video is played later; the third video can also be a recording of the second video, for example, recording the entire playback process of the second video. Specific settings can be flexibly configured according to needs, and there are no restrictions on this. Later, the third video obtained from the second video watched by multiple users can be combined with the first video obtained by recording each user's viewing behavior to generate a collaborative video of multiple people watching the same video content.

[0038] In practical applications, a target user device can be defined. Different user devices record a first video for the same recording task. Other user devices can send their recorded first videos to the target user device and obtain a third video through the target user device. Both the first and third videos are then integrated on the target user device for subsequent processing. Alternatively, each user device can upload both its recorded first video and the third video recorded by the target user device to a server for further processing; this is not a limitation. In some specific implementation examples, the different user devices include a first user device and a second user device. The first user device is the device that imports the second video for the recording task. Furthermore, the first user device can also be the device that initiates the recording task; that is, the first user device can both initiate the recording task and import the second video for that task. This can be flexibly configured according to requirements. When the recording task is not closed and the first user device has not exited the recording task, the target user device can be the first user device; when the recording task is not closed and the first user device has exited the recording task, the target user device can be the second user device. The second user device can be any device other than the first user device among multiple user devices performing the recording task. It can be manually designated, randomly determined, or determined according to a preset method, such as designating a user device invited by the first user device as the second user device. This approach effectively ensures the smooth execution of the recording task and video editing processing.

[0039] Step S104: Obtain editing templates corresponding to multiple target videos; the editing templates include the display position information and track information of the multiple target videos. The track information refers to the editing track information of the multiple target videos in the multi-track editor, such as the type of editing track (main track, picture-in-picture track), etc.

[0040] In practical applications, multiple editing templates can be pre-set, and users can then select the desired template from these preset options. The selected template can then be directly used as the editing template for multiple target videos. Alternatively, users (such as the initiator of a recording task) can further adjust and modify the selected template according to their needs, thereby obtaining editing templates for multiple target videos. In this embodiment, by obtaining the editing templates, the display position and corresponding editing track of each target video can be directly determined, enabling convenient and quick subsequent editing of multiple target videos.

[0041] For example, when the target video includes a third video, the video editing track corresponding to the third video is the main track, and the video editing track corresponding to the first video is the picture-in-picture track. The picture-in-picture track is the secondary track. When multiple target videos contain only the first video, the video editing track corresponding to the first video captured by a third user device in different user devices is the main track, and the video editing tracks corresponding to the first videos captured by other user devices (excluding the third user device) in the multiple target videos are the picture-in-picture tracks; the third user device is the device that initiates the recording task. In practical applications, the track correspondence for each target video can be flexibly set, such as allowing the user to specify the track for each video, etc., without limitation. In practical applications, the aforementioned third user device can not only initiate a recording task but also upload a second video. The third user device can also be the same user device as the first user device that imports the second video for the recording task, such as the creator's device; furthermore, the third user device can be different from the first user device, for example, the first user device and the third user device are devices of two different users, and the two users can upload different second videos according to their needs, without limitation.

[0042] Step S106: Based on multiple target videos and editing templates, display the video editing interface.

[0043] In some implementation examples, the video editing interface includes a preview playback area and an editing track area. The editing track area contains multiple video editing tracks, which may include a main track and one or more picture-in-picture tracks (sub-tracks). These tracks can be arranged from top to bottom, with the main track at the top or bottom, without restriction. Each target video is assigned a video track segment on one of the multiple video editing tracks based on the track information. The timeline positions of the video track segments from multiple target videos may partially or completely overlap. It is understood that the track information can indicate the track type, track position, or track identifier for each target video. Based on the track information of each target video, the corresponding video editing track can be directly determined, and the target video can be treated as a video track segment on the video editing track, facilitating editing of the video track segment composed of multiple target videos. The timeline positions of multiple target videos may partially or completely overlap. For example, if all users' recording start and end times are the same, then the timelines can be considered to completely overlap. However, if some users' recording start and end times are different from other users', such as ending recording early, then their target videos may partially overlap with the timeline positions of other users' target videos. The specific settings can be flexibly configured according to needs, and there are no restrictions here.

[0044] The images of multiple target videos are displayed in the preview playback area according to the display position information of each target video. Specifically, for each target video, its display area in the preview playback area is determined based on its display position information. In this way, the display areas of the images of multiple target videos in the preview playback area can be directly determined based on the display position information in the editing template, allowing users to clearly understand the presentation effect of each target video through the preview playback area.

[0045] In addition, the video editing interface may also include operation controls for editing tools, used to respond to user operations and trigger video editing processes. These operation controls may include text editing controls, texture editing controls, animation effects controls, and track editing controls, and are not limited thereto. This embodiment does not limit the arrangement of the above areas of the video editing interface; for example, the preview playback area can be placed above the editing track area, and the operation controls for the specified editing tools can be placed below the editing track area.

[0046] The above method allows for convenient and quick editing of videos recorded by multiple users for the same recording task, effectively reducing the cost of online collaborative video creation. The device executing this video editing method can be either a user device or a server, and it can interact with other devices during execution; no restrictions are imposed.

[0047] In some specific implementation examples, before the step of acquiring multiple recorded videos, the video editing method provided in this disclosure embodiment further includes the following steps (1) to (3):

[0048] (1) In response to receiving a video recording request from a target user device, a recording task is created, and the user interface corresponding to the recording task is displayed on the target user device; wherein, the user interface displays video addition controls and user invitation controls. For example, the recording task can be a task where multiple people watch a video together, simultaneously recording each user's viewing performance and the playback status of the video (the aforementioned second video) watched by multiple users simultaneously. This disclosure does not limit the specific implementation of creating the recording task; for example, an online chat room can be created for multiple users to enter the online chat room for online interaction and / or to watch the video played in the chat room, while simultaneously recording each user's online interaction performance and video playback status.

[0049] (2) In response to the detection that the video addition control has been triggered, the system obtains the video information uploaded by the target user device and obtains the second video based on the video information; wherein, the video information includes local video files and / or network video links. The second video is the video that multiple users watch together. In practical applications, if the target video contains the second video, the target user device can be a device that simultaneously initiates a recording task and imports the second video for the recording task. In some specific implementation examples, the target user device can be equivalent to the aforementioned first user device, which can both initiate a recording task and import the second video for the recording task.

[0050] In practical applications, various types of video adding controls can be provided to users, such as controls for uploading local video files and / or controls for entering web video links. Users can choose the desired video adding control according to their needs. Furthermore, the number of second videos can be one or more; for example, a user can simultaneously upload multiple second video files from their local machine and / or enter multiple web links for second videos. In practical applications, the uploaded video information can also be parsed. For example, when a user uploads a web video link, the link can be parsed to obtain the corresponding second video. If parsing fails, a prompt can be sent to the user to modify the web video link. Furthermore, to ensure information security, video information can be reviewed to prevent second videos from containing illegal content or from sources that do not meet preset source requirements.

[0051] (3) In response to the detection that the user invitation control is triggered, an invitation message for the recording task is generated so that the target user device can send an invitation message to the specified user device. The invitation message is used to prompt the specified user device to join the recording task.

[0052] Specifically, the user of the target user device (the user who initiates the recording task, which can be referred to as the creator) can invite a specified user by triggering a user invitation control. For example, when the creator triggers the user invitation control, an invitation password or chat room link can be displayed on the user interface so that the creator can forward the invitation information to other users. Other users can directly enter the chat room created by the creator to perform the recording task based on the invitation information.

[0053] For easier understanding, please refer to Figure 2The diagram illustrates a user interface where, after a creator initiates a duet request through a target user's device, the user interface corresponding to the recording task is displayed on the target user's device. This interface displays video addition controls and user invitation controls. Users can upload video files by triggering local controls, enter video links by triggering link controls, or directly upload video files by dragging and dropping. Furthermore, pre-set sample videos can be provided to users so they can directly use them as a second video for testing. In addition... Figure 2 The text also indicates the main creators ( Figure 2 The front-facing camera recording window of user A) can be used to display real-time recording footage including the creator's face. A user invitation control is located to the side of the creator's front-facing camera recording window. By triggering this control, the creator can obtain invitation information and invite specified users to join the chat room. At this time, the front-facing camera recording windows of the invited users will be displayed, and the position of the user invitation control will be adjusted to the outside of the front-facing camera recording windows of the most recently invited users. This continues until the number of users invited by the creator reaches a preset threshold, at which point the user invitation control will no longer be displayed. It should be noted that if the creator neither adds video nor invites users, the start recording control on the user interface will be in an untriggerable state, such as being grayed out and unable to respond to user triggering actions.

[0054] In practical applications, both the target user device and the user interfaces of other user devices performing the recording task display a first area and a second area. The first area displays images of the first video recorded by each user device, while the second area displays images of the second video. That is, each user participating in the recording task can simultaneously view the second video and also simultaneously view the user's performance recorded by the front-facing cameras of all users.

[0055] For example, see Figure 3 The diagram illustrates a user interface on a target user device, showing front-facing camera recording windows for multiple users (User A to User D) displayed on the user interface. These recording windows display the video images of each user's reactions (i.e., the first video). Figure 3 The user interface provided also displays a second video display window to show the images of the second video. During this time, the creator (let's assume it's user A) has playback control permissions for the second video and can adjust the playback status such as progress or speed of the second video at any time as needed. Figure 3The user interface is designed for the creator's video, and therefore also shows the creator the start recording control and the playback adjustment control for the second video. For participants such as users B through D, their user interface is basically the same as the creator's, allowing them to view the images of all users' reaction videos and the second video. The main difference is that the start recording control and the second video playback adjustment control are no longer displayed; that is, only the creator has control over the recording task. As can be seen from the above user interface, each user in the chat room can simultaneously watch the second video and each other's reactions, thus engaging in online interaction and further enabling multi-person collaborative video creation.

[0056] As mentioned earlier, the target user device's user interface displays a "Start Recording" control so that the user (creator) can trigger recording via this control. If neither the target user device's video addition control nor user invitation control is detected, the "Start Recording" control remains in an untriggerable state. In this untriggerable state, the "Start Recording" control cannot be used to record video. That is, if the creator has neither uploaded a video nor invited other users, the recording conditions are not met, and recording cannot begin. Therefore, the "Start Recording" control can be set to an untriggerable state, preventing users from initiating the recording process.

[0057] The "Start Recording" control becomes triggerable when the target user device triggers the video addition control and / or the user invitation control. This triggerable "Start Recording" control is used to begin video recording upon being triggered. In other words, when the creator meets the recording conditions, they can initiate the recording process by issuing a recording command through the "Start Recording" control. In practical applications, the "Start Recording" control can be displayed only on the user interface of the target user device, and not on the user interfaces of other user devices; that is, only the creator is given control over starting recording.

[0058] Based on the above, when acquiring multiple recorded videos, you can refer to steps a and b below:

[0059] Step a: In response to the detection that the start recording control in a triggerable state has been triggered, recording is performed based on the frame images captured by the front-facing camera of the user device performing the recording task to obtain the first video. Specifically, after the start recording control is triggered, each user device performing the recording task can start recording the performance video of its own user, such as capturing video frames containing the face of the user through the front-facing camera to obtain the first video corresponding to each user device.

[0060] Step b: Having obtained the second video through the target user device, the user device performing the recording task simultaneously plays the second video during the recording process and records the playback of the second video to obtain the third video. For example, each user device can simultaneously play the second video while recording the user's viewing behavior through the front-facing camera. The target user device or server can record the complete playback process of the second video, such as recording the content area of ​​the second video, including pauses, playback, progress bar dragging, audio tracks, etc., that occur during playback to obtain the third video.

[0061] In practical applications, different user devices, including the target user device, provide a user interface with triggerable video operation controls on the recording task. These video operation controls, also known as the second video operation controls, are used to adjust the playback status of the second video. Other user devices besides the target user device do not have triggerable video operation controls on their user interfaces for the recording task. When the video operation controls are triggerable, in response to detecting a user-triggered operation on the target user device, the playback status of the second video is adjusted according to the user's trigger operation. That is, the target user device initiating the recording task can be granted operation permissions for the second video. The user of the target user device has the right to adjust the playback status of the second video during playback. This playback status includes, but is not limited to, pause, rewind, fast forward, and slow motion. The user of the target user device can flexibly control the playback status of the second video as needed, such as slowing down the playback speed of important content or rewinding that important content multiple times.

[0062] In specific implementation, different user devices can include a first user device and a second user device. The first user device is the device that imports the second video for the recording task. If the recording task is not closed and the first user device has not exited the recording task, the target user device is the first user device; if the recording task is not closed and the first user device has exited the recording task, the target user device is the second user device. In practical applications, the creator also has the authority to end the recording. Based on this, the user interface of the first user device also includes an end-of-recording control. The method further includes: if the end-of-recording control of the first user device is not detected, but the first user device stops executing the recording task, it indicates that the creator may have unexpectedly lost connection or needs to go offline midway, specifically corresponding to the situation where the recording task is not closed and the first user device has not exited the recording task. In this case, a substitute user device (i.e., the second user device mentioned above) can be determined among the different user devices, and the user interface displayed on the substitute user device when executing the recording task will present video operation controls and an end-of-recording control that can be triggered by the user. For example, the recording control can be transferred to a participant designated by the creator, or it can be transferred to the first-priority participant. The specific settings can be flexibly configured according to needs and are not limited here. In other words, by using the above method, when the main creator goes offline in advance, a substitute user can be found to continue as the new main creator to perform subsequent video playback control and end recording operations, so as to fully ensure that the recording task can be executed normally.

[0063] Furthermore, in the case where multiple target videos are obtained through the aforementioned method, this embodiment of the disclosure also provides an implementation example for obtaining editing templates corresponding to the multiple target videos, which can be performed with reference to the following steps A and B:

[0064] Step A: In response to receiving an editing request from a fourth user device among different user devices, a template selection page is displayed. Exemplarily, the fourth user device can be the user device initiating the recording task, the device uploading the second video, or a specific user device; no limitation is imposed here. That is, the fourth user device can be the same as the aforementioned first user device and / or third user device. Furthermore, the fourth user device can also be other devices specified among multiple user devices; no limitation is imposed here. In other words, in this embodiment of the disclosure, the device initiating the recording task, importing the second video, and performing video editing can be the same device or different devices, and can be flexibly set according to requirements. The template selection page includes various video layout effect diagrams. Each video layout effect diagram includes multiple areas, and each area corresponds to a video display position and a video editing track. That is, the video layout effect diagram can display the window display area of ​​each target video, thereby presenting the relative positional relationship between the various target videos. The display position of the target video is different in different video layout effect diagrams. In this embodiment, the video editing track for each area can be preset. Users do not need to set the video track. By selecting the video layout effect diagram, the display position and video editing track corresponding to each target video can be automatically determined, which helps to further improve video editing efficiency.

[0065] Step B: In response to the detection of a selection operation for a target effect image among multiple video layout effect images, the editing template corresponding to the target effect image is determined for multiple target videos.

[0066] The video layout mockups visually present the display position of each video on the interface and the relative positions of different videos. Therefore, users can easily and quickly select the desired target mockup from multiple video layout mockups. Each video layout mockup corresponds to an editing template. After selecting a target mockup, the user can directly use the editing template corresponding to the target mockup as the editing template for multiple target videos. Users can also further personalize the editing template of the target mockup according to their needs and use the adjusted template as the editing template for multiple target videos.

[0067] In some implementation examples, steps B1 to B2 can be performed as follows:

[0068] Step B1: Based on the target image selected by the target user from various video layout renderings and multiple target videos, display layout preview images of multiple target videos.

[0069] For easier understanding, please refer to Figure 4The illustrations show various layout effects, specifically six simplified video layout effects. Each effect indicates the display area and relative position of different target videos. Users can select the desired target effect image and, based on the video display positions indicated by the target effect image and multiple target videos, easily and quickly generate layout preview images for multiple target videos. This clearly and intuitively presents the display effect of multiple target videos during post-editing. Furthermore, users can adjust the proportion of each video display area indicated by each layout effect while maintaining the relative position of each video display area. Alternatively, users can further adjust the relative position of each video display area based on a selected target effect image for personalized settings; there are no restrictions on this.

[0070] Step B2: Based on the target user's adjustment of the display position of the target video in the layout preview image, determine the editing templates corresponding to multiple target videos.

[0071] In practice, each video layout mockup corresponds to an initial editing template. First, based on the target mockup selected by the user, the corresponding initial editing template, and multiple target videos, layout preview images for multiple target videos are generated. The display position of each target video in the layout preview image is determined based on the display position information of each target video indicated by the initial editing template corresponding to the target mockup. Based on the user's adjustments to the display positions of the target videos in the layout preview images, the display position information of the initial editing template corresponding to the target mockup is adjusted to obtain the target editing template, which is the editing template corresponding to the aforementioned multiple target videos. The main difference between the target editing template and the initial editing template corresponding to the target mockup lies in the different display positions of the target videos adjusted by the user.

[0072] This disclosure also provides a method for users to adjust the display position of target videos. For example, users can adjust the display position of target videos by dragging their floating windows. Furthermore, horizontal and vertical layouts can be preset. If one or more floating windows of target videos are dragged to a first designated area, a horizontal layout is displayed to the user; for example, multiple first videos are arranged horizontally above a third video. If one or more floating windows are dragged to a second designated area, a vertical layout is displayed to the user; for example, multiple first videos are arranged vertically to the side of the third video. Through the above methods, the layout of multiple target videos can be conveniently and quickly adjusted, thereby quickly determining the corresponding editing template for multiple target videos.

[0073] After identifying the corresponding editing templates for multiple target videos (i.e., the target editing templates mentioned above), importing the target videos into the target editing templates generates a video editing draft. This video editing draft can be considered a project file, containing materials such as videos, as well as editing operation information for those materials. The video editing draft can then be imported into the multitrack editor. In other words, the editing draft displayed in the multitrack editor is based on the target videos imported from the target editing templates. This presents the user with a video editing interface, allowing them to directly edit multiple target videos, ultimately generating a pre-edited fourth video, also known as a collaborative video or co-created video.

[0074] The above method significantly reduces the editing costs of online collaborative video creation. Multiple users are not restricted by geographical location, and each only needs to use one device and open one software application capable of executing the video editing method provided in this disclosure embodiment to conveniently and quickly achieve the purpose of video collaborative creation. For ease of understanding, this disclosure embodiment also... Figure 5 The flowchart illustrates a video editing method, showing the operation flow of the creator (corresponding to the target user device that initiated the recording task) and participants (corresponding to the invited user devices) before, during, and after recording. Before recording, the creator can trigger the creation of a chat room through the duet function, import the video (the second video mentioned above) that all users need to watch together during the recording task into the chat room interface, and then invite participants by triggering user invitation controls. Participants can directly enter the chat room after receiving the invitation. During recording, the creator has permissions to start recording, play / pause the video everyone is watching, and confirm completion of recording. Participants simply need to watch the video with the creator and complete the recording accordingly. After recording, the creator can select a template and download the participants' reaction videos (i.e., the videos recorded by the participants' devices showing their reactions to the video). Participants only need to upload their reaction videos directly. The creator can download the videos from each participant in the background, import all the recorded videos into a multi-track editor for editing, and finally submit the final edited video. The above methods enable convenient and quick online video co-creation, greatly saving the cost of multi-person video co-creation and effectively improving the user experience.

[0075] Corresponding to the aforementioned video editing methods, Figure 6 This is a schematic diagram of a video editing device provided in an embodiment of the present disclosure. The device can be implemented by software and / or hardware, and is generally integrated into an electronic device, such as... Figure 6 As shown, the video editing device includes:

[0076] The video acquisition module 602 is used to acquire multiple target videos, including first videos captured and recorded by different user devices for the same recording task.

[0077] The template acquisition module 604 is used to acquire editing templates corresponding to multiple target videos; the editing templates include the display position information and track information of multiple target videos;

[0078] The video editing module 606 is used to display a video editing interface based on multiple target videos and editing templates. The video editing interface includes a preview playback area and an editing track area. The editing track area contains multiple video editing tracks. The images of the multiple target videos are respectively presented in the display area indicated by the display position information of the multiple target videos in the preview playback area. Each of the multiple target videos forms a video track segment on one of the multiple video editing tracks based on the track information. The timeline positions of the video track segments of the multiple target videos partially or completely overlap.

[0079] The aforementioned device allows for convenient and quick centralized editing of videos recorded by multiple users, effectively reducing the cost of online collaborative video creation.

[0080] In some implementations, the target video further includes a third video derived from the second video, wherein the second video is a video played by the different user devices during the shooting and recording process for the recording task, and the third video is used to present the playback status of the second video during the shooting and recording process.

[0081] In some embodiments, the different user devices include a target user device; the user interface provided by the target user device for the recording task displays video operation controls in a triggerable state; the user interfaces provided by the other user devices besides the target user device for the recording task do not display the video operation controls in a triggerable state; when the video operation controls are in a triggerable state, in response to detecting a user trigger operation on the target user device for the video operation controls, the playback state of the second video is adjusted according to the user trigger operation.

[0082] In some implementations, the different user devices include a first user device and a second user device, wherein the first user device is a device that imports the second video for the recording task; when the recording task is not closed and the first user device has not exited the recording task, the target user device is the first user device; when the recording task is not closed and the first user device has exited the recording task, the target user device is the second user device.

[0083] In some implementations, the video editing track corresponding to the third video is the main track, and the video editing track corresponding to the first video is the picture-in-picture track.

[0084] In some implementations, when the plurality of target videos contain only the first video, the video editing track corresponding to the first video captured by the third user device among the different user devices is the main track, and the video editing tracks corresponding to the first videos captured by other user devices among the plurality of target videos are picture-in-picture tracks; wherein, the third user device is the device that initiates the recording task.

[0085] In some implementations, the template acquisition module 604 is specifically used to: in response to receiving an editing request from a fourth user device among the different user devices, display a template selection page; wherein the template selection page includes multiple video layout effect images, each video layout effect image includes multiple areas, each of the multiple areas corresponds to a video display position and a video editing track; in response to detecting a selection operation for a target effect image among the multiple video layout effect images, determine the editing template corresponding to the target effect image as the editing template corresponding to the multiple target videos.

[0086] In some implementations, the template acquisition module 604 is specifically used to: display layout preview images of multiple target videos based on the target effect image selected by the target user from multiple video layout effect images and multiple target videos; and determine the editing templates corresponding to the multiple target videos based on the target user's adjustment operation on the display position of the target videos in the layout preview images.

[0087] In some embodiments, the above-described apparatus further includes:

[0088] The task creation module is used to respond to a collaborative recording request received from a target user's device, create a recording task, and display the corresponding user interface on the target user's device; the user interface displays video addition controls and user invitation controls;

[0089] The second video acquisition module is used to acquire video information uploaded by the target user device in response to the detection that the video addition control has been triggered, and to obtain the second video based on the video information; wherein, the video information includes local video files and / or network video links;

[0090] The task invitation module is used to generate invitation information for a recording task in response to the detection that a user invitation control has been triggered, so that the target user device can send the invitation information to the designated user device. The invitation information is used to prompt the designated user device to join the recording task.

[0091] In some implementations, both the target user device and the other user devices performing the recording task display a first area and a second area on their user interfaces; wherein, the first area is used to display the image of the first video captured and recorded by each user device, and the second area is used to display the image of the second video.

[0092] In some implementations, the user interface of the target user device also displays a start recording control; if the target user device does not trigger the video add control or the user invite control, the start recording control is in an untriggerable state; wherein, the start recording control in the untriggerable state cannot be triggered to record video; if the target user device triggers the video add control and / or the target user device triggers the user invite control, the start recording control is in a triggerable state; wherein, the start recording control in the triggerable state is used to start recording video when triggered.

[0093] In some implementations, the video acquisition module 602 is specifically used to: in response to detecting that a start recording control in a triggerable state is triggered, record based on frame images captured by the front-facing camera of the user device performing the recording task to obtain a first video; if a second video is obtained through the target user device, play the second video synchronously during the video recording process through the user device performing the recording task, and record the playback process of the second video to obtain a third video.

[0094] The video editing apparatus provided in this disclosure can execute the video editing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0095] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device embodiments can be referred to the corresponding process in the method embodiments, and will not be repeated here.

[0096] This disclosure also provides an electronic device, which includes: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the video editing method described above.

[0097] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Figure 7 As shown, the electronic device 700 includes one or more processors 701 and memory 702.

[0098] The processor 701 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 700 to perform desired functions.

[0099] The memory 702 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 701 may execute the program instructions to implement the video editing method of the embodiments of this disclosure described above and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.

[0100] In one example, the electronic device 700 may also include an input device 703 and an output device 704, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0101] In addition, the input device 703 may also include, for example, a keyboard, a mouse, etc.

[0102] The output device 704 can output various information to the outside, including determined distance information, direction information, etc. The output device 704 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0103] Of course, for the sake of simplicity, Figure 7 Only some of the components of the electronic device 700 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 700 may include any other suitable components depending on the specific application.

[0104] In addition to the methods and devices described above, embodiments of this disclosure may also be computer program products, including computer program instructions that, when executed by a processor, cause the processor to perform the video editing methods provided in embodiments of this disclosure.

[0105] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0106] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the video editing method provided in embodiments of this disclosure.

[0107] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0108] This disclosure also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the video editing method of this disclosure.

[0109] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0110] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video editing method characterized by, The method comprises: obtaining a plurality of target videos, the plurality of target videos comprising a first video obtained by different user devices respectively shooting for a same recording task, and a third video obtained based on a second video, the second video being a same video played by the different user devices in a shooting process for the recording task, and the third video being used to present a playing state of the second video in the shooting process; obtaining an editing template corresponding to the plurality of target videos; the editing template comprising display position information of the plurality of target videos and track information of the plurality of target videos, the track information being used to represent track types and track positions corresponding to the plurality of target videos respectively; based on the plurality of target videos and the editing template, displaying a video editing interface; wherein the video editing interface comprises a preview playing area and an editing track area, the editing track area comprising a plurality of video editing tracks, images of the plurality of target videos being respectively presented in display areas indicated by the display position information of the plurality of target videos in the preview playing area, and each target video of the plurality of target videos forming a video track segment on one video editing track of the plurality of video editing tracks based on the track information, time line positions of the video track segments of the plurality of target videos being partially or totally overlapped.

2. The method of claim 1, wherein, The different user devices comprise a target user device; the target user device providing a user interface for the recording task, the user interface presenting a video operation control in a triggerable state; other user devices of the different user devices providing user interfaces for the recording task, the user interfaces not presenting the video operation control in the triggerable state; in a case where the video operation control is in the triggerable state, in response to detecting a user trigger operation on the target user device for the video operation control, adjusting a playing state of the second video according to the user trigger operation.

3. The method of claim 2, wherein, The different user devices comprise a first user device and a second user device, the first user device being a device importing the second video for the recording task; in a case where the recording task is not closed and the first user device does not exit the recording task, the target user device being the first user device; in a case where the recording task is not closed and the first user device exits the recording task, the target user device being the second user device.

4. The method of claim 1, wherein, A video editing track corresponding to the third video is a main track, and a video editing track corresponding to the first video is a picture-in-picture track.

5. The method of claim 1, wherein, in a case where the plurality of target videos only comprise the first video, a video editing track corresponding to a first video obtained by a third user device of the different user devices being a main track, and video editing tracks corresponding to first videos obtained by other user devices of the plurality of target videos except the third user device being picture-in-picture tracks; wherein the third user device is a device initiating the recording task.

6. The method of claim 1, wherein, The obtaining of the editing template corresponding to the plurality of target videos comprises: In response to receiving an editing request of a fourth user equipment in the different user equipments, a template selection page is displayed; wherein the template selection page comprises a plurality of video layout effect diagrams, and each of the plurality of video layout effect diagrams corresponds to a video display position and a video editing track; In response to detecting a selection operation on a target effect diagram in the plurality of video layout effect diagrams, an editing template corresponding to the target effect diagram is determined as an editing template corresponding to the plurality of target videos.

7. The method of claim 6, wherein, The determination of the editing template corresponding to the target effect diagram as the editing template corresponding to the plurality of target videos comprises: According to the target effect diagram selected by the target user in the plurality of video layout effect diagrams and the plurality of target videos, a layout preview image of the plurality of target videos is displayed; Based on an adjustment operation of the target user on a display position of a target video in the layout preview image, an editing template corresponding to the plurality of target videos is determined.

8. The method of claim 1, wherein, Before the plurality of recorded videos is obtained, the method further comprises: In response to receiving a joint shooting request of a target user equipment, the recording task is created, and a user interface corresponding to the recording task is displayed on the target user equipment; wherein the user interface displays a video adding control and a user invitation control; In response to detecting that the video adding control is triggered, video information uploaded by the target user equipment is obtained, and a second video is obtained based on the video information; wherein the video information comprises a local video file and / or a network video link; In response to detecting that the user invitation control is triggered, invitation information of the recording task is generated, so that the target user equipment sends the invitation information to a specified user equipment, and the invitation information is used to prompt the specified user equipment to join the recording task.

9. The method of claim 8, wherein, The user interface of the target user equipment and the user interface of other user equipments performing the recording task both display a first area and a second area; wherein the first area is used to display an image frame of a first video recorded by each of the user equipments, and the second area is used to display an image frame of the second video.

10. The method of claim 8, wherein, The user interface of the target user equipment further displays a start recording control; In a case where it is not detected that the target user equipment triggers the video adding control and it is not detected that the target user equipment triggers the user invitation control, the start recording control is in an untriggerable state; wherein the start recording control in the untriggerable state cannot be triggered to record a video; In a case where it is detected that the target user equipment triggers the video adding control and / or it is detected that the target user equipment triggers the user invitation control, the start recording control is in a triggerable state; wherein the start recording control in the triggerable state is used to start recording a video when triggered.

11. The method of claim 10, wherein, The obtaining of the plurality of recorded videos comprises: In response to detecting that the start recording control in the triggerable state is triggered, a frame image captured by a front camera of a user equipment performing the recording task is recorded to obtain a first video; In a case where the second video is obtained by the target user device, the user device performing the recording task synchronously plays the second video in a video recording process and records a playing process of the second video to obtain a third video.

12. A video editing apparatus characterized by comprising: Comprise: A video obtaining module is configured to obtain a plurality of target videos, the plurality of target videos comprising a first video obtained by different user devices respectively shooting and recording for a same recording task, and a third video obtained based on a second video, the second video being a same video played by the different user devices in a shooting and recording process for the recording task, the third video being used to present a playing state of the second video in the shooting and recording process; A template obtaining module is configured to obtain an editing template corresponding to the plurality of target videos; The editing template comprises display position information of the plurality of target videos and track information of the plurality of target videos, the track information being used to represent track types and track positions corresponding to the plurality of target videos respectively; A video editing module is configured to display a video editing interface based on the plurality of target videos and the editing template, wherein the video editing interface comprises a preview playing area and an editing track area, the editing track area comprising a plurality of video editing tracks, images of the plurality of target videos being respectively displayed in display areas indicated by display position information of the plurality of target videos in the preview playing area, and each target video of the plurality of target videos forming a video track segment on one video editing track of the plurality of video editing tracks based on the track information, time line positions of video track segments of the plurality of target videos being partially or totally overlapped.

13. An electronic device, comprising: The electronic device comprises: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the video editing method in any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the video editing method in any one of claims 1-11.

Citation Information

Patent Citations

  • Video shooting method, apparatus, electronic device, and computer-readable storage medium

    CN108989691A

  • Video co-shooting method, video editing method, video co-shooting device, video editing device and electronic equipment

    CN111866434A

  • Multi-channel video stream synthesis method and device

    CN111901572A