Video compositing method, device, electronic device and computer-readable medium
The video compositing method merges user-shot videos with original videos by determining foreground and background from either source, addressing the poor interaction in existing technologies and enhancing the shooting experience.
Patent Information
- Application Number
- JP2022575908
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-06-12
- Filing Date
- 2021-06-09
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-06-09
AI Technical Summary
Existing composite video technologies result in a poor sense of interaction between user-shot videos and original videos, with obvious gaps and a diminished interactive effect.
A video compositing method that activates a video capture device based on a composite shooting request, merges a first video with a second video to create a target video, where the foreground and background are derived from either the first or second video, enhancing interactivity by eliminating the sense of separation.
Improves user interactivity and enhances the shooting experience by seamlessly integrating user-shot videos with original videos, eliminating gaps and improving the interactive effect.
Smart Images

Figure 0007787105000001 
Figure 0007787105000002 
Figure 0007787105000003
Abstract
Description
[Technical Field]
[0001] For all purposes, this application claims priority to Chinese Patent Application No. 202010537842.6, filed on June 12, 2020, the entire contents of which are incorporated herein by reference.
[0002] The present disclosure relates to video compositing methods, devices, electronic equipment, and computer-readable media. [Background technology]
[0003] With the development of network technology, many social applications have become capable of distributing videos, and it is now popular for users to engage in social activities by distributing videos. Summary of the Invention [Means for solving the problem]
[0004] At least one embodiment of the present disclosure provides a video compositing method, the method comprising: receiving a composite shooting request input by a user based on the first video; activating a video capture device in response to the composite capture request and capturing a second video by the video capture device; The method includes a step of fusing the first video and the second video to obtain a target video, wherein the foreground of the target video is obtained from one of the first video and the second video, and the background of the target video is obtained from the other of the first video and the second video.
[0005] At least one embodiment of the present disclosure further provides a video composite camera, the camera comprising: a composite shooting request receiving module for receiving a composite shooting request input by a user based on the first video; a video capture module for activating a video capture device in response to the composite capture request and capturing a second video with the video capture device; and a video fusion module for fusing the first video and the second video to obtain a target video, wherein the foreground of the target video is obtained from one of the first video and the second video, and the background of the target video is obtained from the other of the first video and the second video.
[0006] At least one embodiment of the present disclosure further provides an electronic device, the electronic device comprising: one or more processors; Memory and and one or more application programs, the one or more application programs stored in memory and configured to be executed by one or more processors, the one or more programs configured to perform the video compositing method.
[0007] At least one embodiment of the present disclosure further provides a computer-readable medium, the readable medium storing at least one instruction, at least one program segment, code set, or instruction set, the at least one instruction, at least one program segment, code set, or instruction set being uploaded and executed by a processor to realize the above-mentioned video compositing method.
[0008] In order to more clearly describe the technical solutions in the embodiments of the present disclosure, the following briefly describes the drawings that need to be used to describe the embodiments of the present disclosure. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a flowchart of a video compound shooting method according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a schematic diagram of a target video according to an embodiment of the present disclosure. [Figure 3] FIG. 3 is a schematic diagram of yet another target video according to an embodiment of the present disclosure. [Figure 4]FIG. 4 is a flowchart of a method for adding background music to a target video according to an embodiment of the present disclosure. [Figure 5] FIG. 5 is a flowchart of another method for adding background music to a target video according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a flowchart of yet another method for adding background music to a target video according to an embodiment of the present disclosure. [Figure 7] FIG. 7 is a schematic diagram of a background music adding interface according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a flowchart of a video distribution method according to an embodiment of the present disclosure. [Figure 9] FIG. 9 is a structural schematic diagram of a video composite camera according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure.
[0010] These and other features, advantages, and aspects of each embodiment of the present disclosure will become more apparent by reference to the following specific embodiments in conjunction with the drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It is understood that the drawings are illustrative, and that actual objects and elements are not necessarily drawn to scale. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the drawings show several embodiments of the present disclosure, it should be understood that the present disclosure may be realized in various forms and should not be understood as being limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are merely illustrative and do not limit the scope of protection of the present disclosure.
[0012] It should be understood that the steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel, and that method embodiments may include additional steps and / or omit performing illustrated steps, and the scope of the present disclosure is not limited in this respect.
[0013] As used herein, the term "comprises" and variations thereof are intended to be open-ended, i.e., "including, but not limited to." The term "based on" means "based at least in part on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one other embodiment," and the term "some embodiments" means "at least some embodiments." Relevant definitions of other terms are provided in the description below.
[0014] It should be noted that the concepts of "first," "second," etc. referred to in this disclosure are merely intended to distinguish between devices, modules, or units, and do not limit these devices, modules, or units to being necessarily different devices, modules, or units, nor do they limit the order or interdependence of functions performed by these devices, modules, or units.
[0015] It should be noted that the modifications "one" and "multiple" referred to in this disclosure are intended to be exemplary rather than limiting, and should be understood as "one or more" unless the context expressly dictates otherwise, as would be understood by one of ordinary skill in the art.
[0016] The names of messages or information interacted between devices in the embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0017]
[0023] The following specific examples will be used to describe in detail the technical solutions of the present disclosure and how the technical solutions of the present disclosure solve the above technical problems. Some of the following specific examples can be combined with each other, and detailed descriptions of the same or similar concepts or processes may be omitted in some examples.
[0024] The following examples will be described with reference to the drawings.
[0018] In existing video social technologies, composite video technology has emerged, allowing users to display their own videos on the same screen as others' videos to achieve an interactive effect. However, with existing composite video technology, the videos shot by the user and others' videos must be split into left and right sections or top and bottom sections, resulting in a poor sense of interaction between the user and the original video. The videos shot by the user are separated from the original video, resulting in obvious gaps between the videos and a poor interactive effect.
[0019] As can be seen from the above, the existing composite shooting technology has the problem that the sense of interaction between the user and the original video is poor, the video shot by the user is different from the original video, there is an obvious gap between the videos, and the interactive effect is poor.
[0020] In an embodiment of the present disclosure, a video capture device is activated based on a composite shooting request input by a user based on a first video, a second video is acquired by the video capture device, and the first video and the second video are merged to obtain a target video. After the first video and the second video are merged, there is no sense of separation between the two videos. When a user shoots a video using the video composite shooting method of an embodiment of the present disclosure, the user's interactivity between the videos can be improved and the enjoyment of shooting can be enhanced.
[0021] In an embodiment of the present disclosure, a video compositing method is provided, as shown in FIG. 1, the method includes: Step S101: receiving a composite shooting request input by a user based on a first video; Step S102: activating a video capture device in response to the composite shooting request and capturing a second video by the video capture device; and step S103 of fusing the first video and the second video to obtain a target video.
[0022] The video composite shooting method according to the embodiment of the present disclosure can be applied to any terminal device, and the terminal device may be a terminal device equipped with a video capture device, such as a smartphone or a tablet computer. If the terminal device does not have a video capture device, a video capture device can be connected externally. In the embodiment of the present disclosure, the video capture device is activated based on a composite shooting request input by a user based on a first video, and a second video is acquired by the video capture device. The first video and the second video are merged to obtain a target video merging the two video contents. After the first video and the second video are merged, there is no sense of separation between the two videos. When a user shoots a video using the video composite shooting method according to the embodiment of the present disclosure, the user's interactivity between the videos can be improved, and the user's shooting experience can be enhanced.
[0023] The steps of the above video compound shooting method will now be described in detail.
[0024] In step S101, a composite shooting request input by a user based on a first video is received.
[0025] In the embodiment of the present disclosure, the solution is applied to an APP (application, application program) as an example. The first video is a video uploaded and distributed by a user in the APP. When the current user watches the video on the APP, he or she can start composite shooting based on the video.
[0026] For the embodiments of the present disclosure, taking the above implementation as an example, when a current user watches a video and wants to take a composite photo with the video, he can input a composite photo request based on the video. The manner of inputting the composite photo request can be to trigger the composite photo control of the APP interface, for example, to start the composite photo request by clicking the composite photo button of the video interface in the APP interface, or to start the composite photo request by clicking the composite photo transfer button of the video. When the current user triggers the composite photo control of the APP interface, the terminal device receives the composite photo request.
[0027] In step S102, in response to the composite shooting request, the video capture device is started and the second video is acquired by the video capture device.
[0028] In the embodiment of the present disclosure, the second video refers to a video captured by the current user based on a composite capturing request, and the subject may be a person, a scene, etc.
[0029] Regarding the embodiments of the present disclosure, as in the above embodiment, the smart terminal to which the solution according to the embodiments of the present disclosure is applied may include at least one video capture device such as a camera, and after receiving a composite shooting request input by a user based on a first video, the terminal device activates the video capture device in response to the composite shooting request and obtains a second video through the video capture device.
[0030] In step S103, the first video and the second video are fused to obtain a target video, where the foreground of the target video is obtained from the first video and the background is obtained from the second video, or the foreground of the target video is obtained from the second video and the background is obtained from the first video.
[0031] In an embodiment of the present disclosure, after the terminal device acquires the first video and the second video, it fuses the first video and the second video to form a target video, and the fusion method may be fusing some content of the first video with the second video, or fusing all content of the first video with all content of the second video, or fusing some content of the first video with some content of the second video, or fusing some content of the second video with all content of the first video, and the specific fusion method is not limited to the embodiment of the present disclosure.
[0032] In an embodiment of the present disclosure, a video capture device is activated based on a composite shooting request input by a user based on a first video, a second video is acquired by the video capture device, and the first video and the second video are merged to obtain a target video. After the first video and the second video are merged, there is no sense of separation between the two videos. When a user shoots a video using the video composite shooting method of an embodiment of the present disclosure, the user's interactivity between the videos can be improved and the enjoyment of shooting can be enhanced.
[0033] In the embodiments of the present disclosure, a possible implementation is provided, in which the step of fusing the first video and the second video to obtain the target video comprises: Extracting a first target content from the first video and fusing the first target content with the second video to obtain a target video; extracting second target content from the second video and fusing the second target content with the first video to obtain a target video; The method includes at least one of steps of extracting third target content from the first video, extracting fourth target content from the second video, and fusing the third target content and the fourth target content to obtain the target video.
[0034] In the embodiments of the present disclosure, the first target content may be a portion of content in the first video, and the second target content may be a portion of content in the second video. The portion of content may be, for example, the background or foreground of the video, or some, some, or all of the target objects in the video, including, but not limited to, people. Optionally, the first target content and the second target content may be objects in the corresponding videos, for example, people, buildings, animals, etc. The third target content may be the foreground or background of the first video, and the fourth target content may be the foreground or background of the second video. Merging the third target content and the fourth target content may be blending the foreground of the first video with the background of the second video, blending the foreground of the first video with the foreground of the second video, or even blending the background of the first video with the foreground of the second video, or blending the background of the first video with the background of the second video. The specific blending manner is not limited to the embodiments of the present disclosure.
[0035] In the embodiments of the present disclosure, there are various ways to fuse the first video and the second video. In one embodiment of the embodiments of the present disclosure, when fusing the first video and the second video, a first target content is extracted from the first video, and then the first target content is fuse with the second video to obtain the target video, specifically, the background in the first video is used as the first target content, and then the background of the first video is used as the background of the target video and fuse with the second video to obtain the target video. In yet another embodiment of the present disclosure, when fusing the first video and the second video, a second target content is extracted from the second video, and then the second target content is fuse with the first video to obtain the target video, specifically, the foreground of the second video is used as the second target content, and then the foreground of the second video is used as the foreground of the target video and fuse with the first video to obtain the target video.
[0036] An embodiment of the present disclosure extracts target content from a first video and / or a second video, fuses the extracted target content with the first video and / or the second video, and cross-fuse the contents of the first video and the second video to obtain a target video, so that there is no gap between the first video and the second video in the target video, and the first video and the second video are displayed through one video, thereby improving the interactive feeling of users interacting through the video.
[0037] In the embodiments of the present disclosure, a possible implementation is provided, in which the step of extracting a first target content from a first video and fusing the first target content with a second video includes: When the composite shooting request is a first composite shooting request, the method includes a step of: setting the first target content as the foreground of the target video, setting the second video as the background of the target video, and fusing the first target content with the second video.
[0038] In the embodiments of the present disclosure, there are various methods for fusing the first video and the second video, and different fusing methods can be selected based on the type of fusing request input by the user, where the first fusing request refers to a fusing request for fusing the first target content of the first video with the second video. Optionally, when the first target content of the first video is the background of the first video, the background of the first video can be used as the background of the target video to fusing with the second video to obtain the target video.
[0039] In an embodiment of the present disclosure, when the composite shooting request input by the user is a first composite shooting request, the first target content is used as the foreground of the target video, the second video is used as the background of the target video, and the first target content is merged with the second video. In one embodiment of the present disclosure, the first target content may be the foreground of the first video, or a person with certain features, a landscape, etc. After extracting the first target content, as shown in FIG. 2, the first target content is merged with the foreground 201 of the target video and the background 202 of the target video to obtain the target video. Specifically, if the first video is a live video of a singer, the singer's body can be used as the first target content when extracting the first target content from the first video, and the singer's body can be incorporated into the foreground of the second video when merging the first target content with the second video to obtain the target video, resulting in a visual effect of the singer singing in a scene from the second video captured by the user.
[0040] An embodiment of the present disclosure can incorporate a first target content in a first video into the foreground of a second video, allowing a user to include the first target content in the first video in a video scene shot by the user, eliminating the sense of disconnection between the first video and the second video, and providing better interactivity between the second video shot by the user and the first video.
[0041] The embodiments of the present disclosure provide another possible implementation, in which the step of extracting second target content from the second video and fusing the second target content with the first video includes: When the composite shooting request is a second composite shooting request, the method includes a step of: setting the second target content as the foreground of the target video, setting the first video as the background of the target video, and fusing the second target content with the first video.
[0042] In the embodiments of the present disclosure, there are various ways to combine the first video and the second video, and different combination shooting ways can be selected based on the request type of the combination shooting request input by the user, and the second combination shooting request refers to the combination shooting request for combining the first video and the second target content of the second video.
[0043] In an embodiment of the present disclosure, when the composite shooting request input by the user is a second composite shooting request, the second target content is used as the foreground of the target video, the first video is used as the background of the target video, and the second target content is merged with the first video. In one embodiment of the present disclosure, the second target content may be the foreground of the second video, or a person or landscape with certain features. After extracting the second target content, as shown in FIG. 3, the second target content is merged with the first video as the foreground 301 of the target video and the background 302 of the target video to obtain the target video. Specifically, if the first video is a video of a singer's live concert, the captured second video may show the user singing. When extracting the second target content from the second video, the user's body is used as the second target content. When the second target content is merged with the first video, the user's body is incorporated into the foreground of the first video to obtain the target video, resulting in a visual effect of the user and the singer singing on the same stage.
[0044] The embodiments of the present disclosure can incorporate the second target content in the second video into the foreground of the first video, allowing the user to include the second target content in the second video in the video scene of the first video, eliminating the sense of disconnection between the first video and the second video, and providing better interactivity between the second video and the first video shot by the user.
[0045] As yet another embodiment of the present disclosure, when fusing a first video and a second video, a first target content of the first video and a second target content of the second video can be selected and fusing. For example, the first target content can be fusing as the background and the second target content as the foreground, or the first target content can be fusing as the foreground and the second target content as the background to obtain a target video. Alternatively, without extracting target content from the first video and the second video, one of the first video and the second video can be directly fusing as the background and the other of the first video and the second video as the foreground to obtain a target video. All of the above solutions are within the scope of protection of the present disclosure.
[0046] The embodiments of the present disclosure further provide a possible implementation, as shown in FIG. 4, in which the method includes: Step S401: activating an audio capture device based on a composite capture request; The method further includes step S402 of obtaining the audio captured by the audio capture device and using the audio captured by the audio capture device as background music for the target video.
[0047] In the embodiments of the present disclosure, the terminal device may be a terminal device equipped with an audio capture device. When the terminal device does not have an audio capture device, an audio capture device can be connected externally, and the user can add background music to the target video through the terminal device. There are various ways to add background music. In one embodiment, the user inputs a composite shooting request and simultaneously activates the audio capture device, and the audio captured by the audio capture device is used as background music of the target video. The adding of background music will be described in detail below.
[0048] In step S401, the audio capture device is started based on a composite imaging request.
[0049] In an embodiment of the present disclosure, when a user inputs a composite shooting request based on the first video, the composite shooting request further includes a request to activate an audio capture device, and the audio capture device may be a device such as a microphone. In one embodiment of the present disclosure, when a user clicks a control to initiate a composite shooting request, the user simultaneously clicks a control to activate the audio capture device, and the terminal device activates the audio capture device based on the click, or the user independently clicks the control to activate the audio capture device to activate the audio capture device.
[0050] In step S402, the audio captured by the audio capture device is obtained, and the audio captured by the audio capture device is used as background music for the target video.
[0051] In an embodiment of the present disclosure, after the terminal device starts the audio capture device, it obtains the audio captured by the audio capture device and uses the audio as background music of the target video, and the audio captured by the audio capture device is the background music of the second video.
[0052] In an embodiment of the present disclosure, an audio capture device is activated based on a composite shooting request input by a user, and the audio captured by the audio capture device is used as background music for the target video, thereby enhancing the listening experience of the target video.
[0053] Another possible implementation is provided in the embodiments of the present disclosure, and as shown in FIG. 5 , in this implementation, the method includes: A step S501 of extracting audio from a first video; The method further includes a step S502 of using the audio of the first video as background music of the target video.
[0054] In the embodiments of the present disclosure, there are various ways to add background music. In the previous embodiment, the background music of the second video is used as the background music of the target video, but in this embodiment, the background music of the first video can be used as the background music of the target video. For example, when a current user wants to match the dance moves he or she has filmed with the music rhythm of the first video, he or she can select the background music of the first video to use as the background music of the target video. A specific embodiment is as follows:
[0055] In an embodiment of the present disclosure, if a user does not input an audio capture request when inputting a composite shooting request based on a first video, for example, when the user selects the background music of the first video to use as the background music of the target video without clicking the audio capture control in the APP interface, the audio of the first video can be extracted based on the user's composite shooting request, and optionally, the audio of the first video can be extracted, and after the fusion of the first video and the second video is completed, the audio of the first video can be used as the background music of the target video, or the audio of the first video can be used as the background music of the target video during the shooting of the second video or during the fusion of the videos.
[0056] In the embodiment of the present disclosure, the background music of the first video is used as the background music of the target video, so that when the user takes a composite shot, the target video taken by the user and the first video use the same background music, thereby improving the interactivity between the target video and the first video.
[0057] In the embodiments of the present disclosure, yet another possible implementation is provided, as shown in FIG. 6 , in which the method includes: Step S601: receiving a background music addition request input by a user based on a target video; Step S602 of displaying a background music adding interface in response to the background music adding request; Step S603: accepting an audio selection operation by a user based on a background music adding interface; The method further includes step S604 of using the music corresponding to the audio selection operation as background music for the target video.
[0058] In the embodiments of the present disclosure, there are various methods for adding background music to the target video. In the above two embodiments, the background music of the first video and the second video are respectively used as the background music of the target video. When the user does not want to use the background music of the first video and the second video, the user can choose to add his / her desired background music as the background music of the target video. The solution will be described in detail below.
[0059] In step S601, a request to add background music input by a user based on a target video is received.
[0060] In the embodiments of the present disclosure, when a user chooses to add music other than the music in the first video and the second video as background music for the target video, a request for adding background music is required, and the terminal device can easily add background music to the target video based on the request for adding background music.
[0061] For the embodiments of the present disclosure, the user's operation of inputting a request to add background music based on the target video may be that the user clicks a control to add background music in the APP interface, and the terminal device receives the request to add background music based on the operation.
[0062] In step S602, a background music addition interface is displayed in response to a background music addition request.
[0063] In an embodiment of the present disclosure, the terminal device displays a background music adding interface in response to the background music adding request. As shown in FIG. 7, the APP display interface includes a target video display area 701 and a background music adding control 702. When the background music adding control 702 is clicked, a background music adding interface 703 is displayed, and the user selects and adds background music based on the background music adding interface.
[0064] In step S603, an audio selection operation by the user based on the background music adding interface is accepted.
[0065] In the embodiments of the present disclosure, a user can perform an audio selection operation on a background music adding interface displayed on a terminal device, and the audio selection operation may be performed by clicking a background music icon or music name on the background music adding interface to select background music.
[0066] In step S604, the music corresponding to the audio selection operation is used as background music for the target video.
[0067] In an embodiment of the present disclosure, the device terminal uses corresponding music as background music for the target video based on the user's audio selection operation, and the music may be locally cached music or music downloaded from a network.
[0068] In the embodiments of the present disclosure, a request to add background music input by a user is received, an interface for adding background music is displayed based on the request to add background music, and corresponding music is determined as the background music of the target video based on the user's operation to add background music through the interface for adding background music, so that the user can add background music to the target video according to their own preferences, and the user experience is better.
[0069] In the embodiments of the present disclosure, a possible implementation is provided, as shown in FIG. 8, in the implementation, the video compositing method includes: Step S801: receiving a video distribution request input by a user based on a target video; a step S802 of determining a similarity between the target video and the first video in response to the video delivery request; The method further includes a step S803 of sending the target video to a server when the similarity between the target video and the first video does not exceed a preset threshold.
[0070] In the embodiment of the present disclosure, after a user completes video composition, he or she can choose to distribute the video. However, before distributing the video, it is necessary to detect whether the video meets the distribution requirements, that is, it is necessary to ensure that the similarity between the target video and the first video is not too high, thereby preventing the user from directly stealing and distributing other people's videos. The above solution will be described in detail below.
[0071] In step S801, a video distribution request input by a user based on a target video is received.
[0072] In an embodiment of the present disclosure, after completing the video composition shooting, the user can choose to distribute the target video, and optionally, the user can click the video distribution control on the interface of the target video to initiate a video distribution request, and the terminal device receives the video distribution request input by the user based on the target video.
[0073] In step S802, in response to a video distribution request, a similarity between the target video and the first video is determined.
[0074] In an embodiment of the present disclosure, when calculating the similarity between the target video and the first video, the calculation may be performed using a similarity algorithm, or the similarity between the first video and the target video may be calculated by identifying whether the target video contains content that is not present in the first video.
[0075] In step S803, if the similarity between the target video and the first video does not exceed a preset threshold, the target video is sent to the server.
[0076] In an embodiment of the present disclosure, when the similarity between the target video and the first video does not exceed a preset threshold, it indicates that the target video and the first video are significantly different, and the target video can be sent to a server for distribution; when it is detected that the target video contains content that is not in the first video, it similarly indicates that the target video and the first video are significantly different, and the target video can be sent to a server for distribution; when it is detected that the target video contains content that is not in the first video, and the time that the content that is not in the first video occupies the target video exceeds a preset percentage, it similarly indicates that the target video and the first video are significantly different, and the target video can be sent to a server for distribution. The content that is not in the first video may be people, animals, scenery, etc. For ease of explanation, take a specific application scenario as an example: the first video is a live video of a singer's concert, and when a current user wants to composite film themselves based on the live video of the concert, if the current user's image is detected in the target video and the current user's image exists in the target video for a predetermined time, for example, more than 3 seconds, it is determined that the target video and the first video are significantly different, and the target video can be sent to a server for distribution; if the current user's image is not detected in the target video, or the current user's image appears in the target video for less than 3 seconds, it is determined that the target video and the first video are too similar, and the target video cannot be distributed. Optionally, when the target video is distributed, a composite filming link can be automatically generated, and the composite filming link may include a homepage link of the first video creator, which facilitates other users to learn more about the first video creator's related works and also plays a certain promotional role for the first video creator.
[0077] An embodiment of the present disclosure calculates the similarity between the target video and the first video, and only sends the target video to a server for distribution if the similarity does not exceed a preset threshold, thereby preventing users from directly using other people's videos and causing copyright infringement.
[0078] In an embodiment of the present disclosure, a video capture device is activated based on a composite shooting request input by a user based on a first video, a second video is obtained by the video capture device, the first video and the second video are merged to obtain a target video, and after the first video and the second video are merged, some or all of the content of the two videos is merged, so that there is no sense of separation between the two videos. When a user shoots a video using the video composite shooting method of an embodiment of the present disclosure, the user's interactivity between the videos can be improved and the enjoyment of shooting can be enhanced.
[0079] An embodiment of the present disclosure provides a video composite camera, and as shown in FIG. 9, the video composite camera 90 may include a composite camera request receiving module 901, a video acquisition module 902, and a video fusion module 903.
[0080] The composite shooting request receiving module 901 is used for receiving a composite shooting request input by a user based on the first video.
[0081] The video capture module 902 is used to activate a video capture device in response to a composite shooting request and capture a second video through the video capture device.
[0082] The video fusion module 903 is used to fuse the first video and the second video to obtain a target video, where the foreground of the target video is obtained from the first video and the background is obtained from the second video, or the foreground of the target video is obtained from the second video and the background is obtained from the first video.
[0083] Optionally, when the video fusion module 903 fuses the first video and the second video to obtain the target video, Extracting a first target content from the first video and fusing the first target content with the second video to obtain a target video; and / or Extracting second target content from the second video and fusing the second target content with the first video to obtain a target video; and / or It may be used to extract a third target content from the first video, extract a fourth target content from the second video, and fuse the third target content with the fourth target content to obtain a target video.
[0084] Optionally, the video fusion module 903 extracts the first target content from the first video, and fuses the first target content with the second video to obtain the target video: When the composite shooting request is a first composite shooting request, the first target content may be used as the foreground of the target video and the second video as the background of the target video, and the first target content may be used to blend the first target content with the second video.
[0085] Optionally, the video fusion module 903 extracts a second target content from the second video, and fuses the second target content with the first video to obtain a target video: When the composite shooting request is a second composite shooting request, the second target content may be used as the foreground of the target video and the first video as the background of the target video, and the second target content may be used to merge with the first video.
[0086] Optionally, the video fusion module 903 further comprises: activating an audio capture device based on the composite capture request; The audio may be captured by an audio capture device and used as background music for the target video.
[0087] Optionally, the video fusion module 903 further comprises: Extracting audio from the first video; It may also be used to use the audio of the primary video as background music for the target video.
[0088] Optionally, the video fusion module 903 further comprises: receiving a background music addition request input by a user based on a target video; displaying a background music adding interface in response to a background music adding request; Accepting an audio selection operation by a user based on a background music adding interface; The music corresponding to the audio selection operation may be used as background music for the target video.
[0089] Optionally, the video composite camera according to the embodiment of the present disclosure further includes a video distribution module, and the video distribution module: receiving a video delivery request input by a user based on a target video; determining a similarity between the target video and the first video in response to the video delivery request; and sending the target video to the server when the similarity between the target video and the first video does not exceed a preset threshold.
[0090] The modules may be implemented as software components running on one or more general-purpose processors, or as hardware that performs certain functions or a combination thereof, such as programmable logic devices and / or application-specific integrated circuits. In some embodiments, the modules may be embodied in the form of a software product, which may be stored on a non-volatile storage medium that causes a computing device (e.g., a personal computer, a server, a network device, a mobile terminal, etc.) to implement the methods described in the embodiments of the present disclosure. In one embodiment, the modules may be implemented on a single device or distributed across multiple devices. The functionality of the modules may be combined with each other or further divided into multiple sub-modules.
[0091] The video composite shooting device of this embodiment can implement the video composite shooting method shown in the above embodiments of the present disclosure, and the realization principle thereof is similar, so detailed description is omitted here.
[0092] In an embodiment of the present disclosure, a video capture device is activated based on a composite shooting request input by a user based on a first video, a second video is acquired by the video capture device, and the first video and the second video are merged to obtain a target video. After the first video and the second video are merged, there is no sense of separation between the two videos. When a user shoots a video using the video composite shooting method of an embodiment of the present disclosure, the user's interactivity between the videos can be improved and the enjoyment of shooting can be enhanced.
[0093] 10 shows a structural schematic diagram of an electronic device for implementing an embodiment of the present disclosure. Electronic devices in the embodiment of the present disclosure include, but are not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG. 10 is merely an example and does not limit the functionality and scope of use of the embodiment of the present disclosure.
[0094] The electronic device includes a memory and a processor, which may be referred to as a processing unit 1001 below, and the memory may include at least one of a read-only memory (ROM) 1002, a random access memory (RAM) 1003, and a storage device 1008, specifically as follows:
[0095] 10, the electronic device 1000 may include a processing unit (also referred to as a "processor", e.g., a central processor, a graphics processor, etc.) 1001, which can perform various appropriate operations and processes based on programs stored in a read-only memory (ROM) 1002 or programs uploaded from a storage device 1008 to a random access memory (RAM) 1003. The RAM 1003 further stores various programs and data necessary for the operation of the electronic device 1000. The processing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0096] Typically, devices such as input devices 1006, including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 1007, including, for example, a liquid crystal display (LCD), speakers, oscillators, etc.; storage devices 1008, including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009 may be connected to the I / O interface 1005. The communication devices 1009 enable the electronic device 1000 to communicate wirelessly or via wires with other devices to exchange data. While FIG. 10 shows the electronic device 100 with various devices, it should be understood that it is not required to implement or include all of the illustrated devices. More or fewer devices may alternatively be implemented or included.
[0097] In particular, according to embodiments of the present disclosure, the processes described with reference to the flowcharts may be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the methods shown in the flowcharts. In such embodiments, the computer program may be downloaded and installed from a network via the communication device 1009, or may be installed from the storage device 1008, or may be installed from the ROM 1002. When the computer program is executed by the processing device 1001, it performs the functions defined in the methods of the embodiments of the present disclosure.
[0098] It should be noted that the computer-readable medium of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more conductors, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which is capable of transmitting, propagating, or transmitting a program used by or in connection with an instruction execution system, apparatus, or device. Program code contained in a computer-readable medium may be transmitted over any suitable medium, including, but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0099] In some embodiments, clients and servers may communicate using any now known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and may be connected to one another via any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., international networks), and end-to-end networks (e.g., ad hoc end-to-end networks), and any now known or later developed networks.
[0100] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated in the electronic device.
[0101] The computer-readable medium is equipped with one or more programs, and when the one or more programs are executed by the electronic device, the electronic device receives a composite shooting request input by a user based on a first video, activates a video capture device in response to the composite shooting request, acquires a second video through the video capture device, and fuses the first video and the second video to obtain a target video.
[0102] Computer program code for carrying out the operations of the present disclosure can be written using one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may run entirely on the user computer, partially on the user computer, as a separate software package, partially on the user computer and partially on a remote computer, or entirely on a remote computer or server. When referring to a remote computer, the remote computer may be connected to the user computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0103] The flowcharts and block diagrams in the drawings illustrate possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in a flowchart or block diagram may represent a module, program segment, or portion of code, which includes one or more executable instructions for implementing the specified logical function. It should be noted that in some alternative implementations, the functions marked in the boxes may be executed in an order different from the order marked in the drawings. For example, two boxes shown in succession may actually be executed substantially in parallel or in the reverse order, depending on the functionality involved. It should be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented in a dedicated hardware-based system that performs the specified functions or operations, or in a combination of dedicated hardware and computer instructions.
[0104] The functionality described herein may be implemented, at least in part, in one or more hardware logic components. For example, but not limited to, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), etc.
[0105] In the context of this disclosure, a machine-readable medium may be a tangible medium that can contain or store a program used by or in connection with an instruction execution system, device, or apparatus. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. Further specific examples of machine-readable storage media include one or more wire-based electrical connections, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0106] According to one or more embodiments of the present disclosure, there is provided a method for video compositing, comprising: receiving a composite shooting request input by a user based on the first video; activating a video capture device in response to the composite capture request and capturing a second video by the video capture device; The method includes a step of fusing the first video and the second video to obtain a target video, wherein the foreground of the target video is obtained from the first video and the background is obtained from the second video, or the foreground of the target video is obtained from the second video and the background is obtained from the first video.
[0107] Furthermore, the step of fusing the first video and the second video to obtain the target video includes: Extracting a first target content from the first video and fusing the first target content with the second video to obtain a target video; and / or Extracting a second target content from the second video and fusing the second content with the first video to obtain a target video; and / or The method includes obtaining a third target content of the first video, extracting a fourth target content of the second video, and fusing the third target content and the fourth target content to obtain a target video.
[0108] Further, the step of extracting the first target content from the first video and fusing the first target content with the second video includes: When the composite shooting request is a first composite shooting request, the method includes a step of: setting the first target content as the foreground of the target video, setting the second video as the background of the target video, and fusing the first target content with the second video.
[0109] Further, the step of extracting second target content from the second video and fusing the second content with the first video includes: When the composite shooting request is a second composite shooting request, the method includes a step of: setting the second target content as the foreground of the target video, setting the first video as the background of the target video, and fusing the second target content with the first video.
[0110] Furthermore, the video compound shooting method includes: activating an audio capture device based on the composite capture request; The method further includes obtaining the audio captured by the audio capture device and using the audio captured by the audio capture device as background music for the target video.
[0111] Furthermore, the video compound shooting method includes: Extracting audio from the first video; and using the audio of the first video as background music for the target video.
[0112] Furthermore, the video compound shooting method includes: receiving a request to add background music input by a user based on a target video; displaying a background music adding interface in response to a background music adding request; receiving an audio selection operation by a user based on a background music adding interface; and using the music corresponding to the audio selection operation as background music for the target video.
[0113] Furthermore, the video compound shooting method includes: receiving a video delivery request input by a user based on a target video; determining a similarity between the target video and the first video in response to the video delivery request; The method further includes sending the target video to a server when the similarity between the target video and the first video does not exceed a preset threshold.
[0114] According to one or more embodiments of the present disclosure, there is provided a video composite camera, comprising: a composite shooting request receiving module for receiving a composite shooting request input by a user based on the first video; a video capture module for activating a video capture device in response to the composite capture request and capturing a second video with the video capture device; and a video fusion module for fusing the first video and the second video to obtain a target video, wherein the foreground of the target video is obtained from the first video and the background is obtained from the second video, or the foreground of the target video is obtained from the second video and the background is obtained from the first video.
[0115] Optionally, the video fusion module, when fusing the first video and the second video to obtain the target video, Extracting a first target content from the first video and fusing the first target content with the second video to obtain a target video; and / or Extracting second target content from the second video and fusing the second target content with the first video to obtain a target video; and / or It may be used to extract a third target content from the first video, extract a fourth target content from the second video, and fuse the third target content with the fourth target content to obtain a target video.
[0116] Optionally, the video fusion module extracts a first target content from the first video, and fuses the first target content with the second video to obtain a target video: When the composite shooting request is a first composite shooting request, the first target content may be used as the foreground of the target video and the second video as the background of the target video, and the first target content may be used to blend the first target content with the second video.
[0117] Optionally, the video fusion module according to the embodiment of the present disclosure extracts a second target content from the second video, and fuses the second target content with the first video to obtain a target video, When the composite shooting request is a second composite shooting request, the second target content may be used as the foreground of the target video and the first video as the background of the target video, thereby fusing the second target content with the first video.
[0118] Optionally, the video fusion module further comprises: activating an audio capture device based on the composite capture request; The audio may be captured by an audio capture device and used as background music for the target video.
[0119] Optionally, the video fusion module further comprises: Extracting audio from the first video; It may also be used to use the audio of the primary video as background music for the target video.
[0120] Optionally, the video fusion module further comprises: receiving a background music addition request input by a user based on a target video; displaying a background music adding interface in response to a background music adding request; Accepting an audio selection operation by a user based on a background music adding interface; The music corresponding to the audio selection operation may be used as background music for the target video.
[0121] Optionally, the video composite camera according to the embodiment of the present disclosure further includes a video distribution module, and the video distribution module: receiving a video delivery request input by a user based on a target video; determining a similarity between the target video and the first video in response to the video delivery request; and sending the target video to the server when the similarity between the target video and the first video does not exceed a preset threshold.
[0122] According to one or more embodiments of the present disclosure, there is provided an electronic device including one or more processors, a memory, and one or more application programs, the one or more application programs being stored in the memory and configured to be executed by the one or more processors, and the one or more programs being configured to perform the above-described video compositing method.
[0123] According to one or more embodiments of the present disclosure, there is provided a computer-readable medium having stored thereon at least one instruction, at least one program segment, code set, or instruction set, the computer-readable medium having stored thereon at least one instruction, at least one program segment, code set, or instruction set, the computer-readable medium being uploaded and executed by a processor to implement the above-described video compositing method.
[0124] The above description is merely a description of the preferred embodiments of the present disclosure and the technical principles used. As will be understood by those skilled in the art, the scope of the present disclosure is not limited to the technical solution consisting of a specific combination of the above technical features, but should also include other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the concept of the above disclosure. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (not limited to) disclosed in the present disclosure may be included.
[0125] Also, although operations are described in a particular order, this should not be understood as requiring that these operations be performed in the particular order shown, or sequentially. In some cases, multitasking and parallel processing may be advantageous. Similarly, although the above discussion includes several specific implementation details, these should not be construed as limiting the scope of the present disclosure. Some features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination.
[0126] Although the present subject matter has been described in language specific to structural features and / or logical operations of a method, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or operations described above. Rather, the specific features and operations described above are merely example forms of implementing the claims.
Claims
1. 1. A video compositing method, comprising: receiving a composite shooting request input by a user based on a first video, the first video being a video uploaded and distributed by another user of the application; activating a video capture device in response to the composite capture request and capturing a second video by the video capture device; fusing the first video and the second video to obtain a target video, wherein a foreground of the target video is obtained from one of the first video and the second video, and a background of the target video is obtained from the other of the first video and the second video; receiving a video delivery request input by the user based on the target video; determining a similarity between the target video and the first video in response to the video delivery request; When the similarity between the target video and the first video does not exceed a preset threshold, sending the target video to a server of the application, and the server distributes the target video and generates a composite shooting link including a homepage link of the creator of the first video; The step of determining a similarity between the target video and the first video includes: determining a similarity between the target video and the first video based on whether an image of the user is detected in the target video and a time at which the image of the user appears in the target video.
2. The step of fusing the first video and the second video to obtain the target video includes: extracting a first target content from the first video and fusing the first target content with the second video to obtain the target video; extracting second target content from the second video and fusing the second target content with the first video to obtain the target video; 2. The video compositing method of claim 1, further comprising at least one of the steps of: extracting a third target content from the first video; extracting a fourth target content from the second video; and fusing the third target content and the fourth target content to obtain the target video.
3. The step of extracting a first target content from the first video and fusing the first target content with the second video includes:
3. The video composite shooting method of claim 2, further comprising: when the composite shooting request is a first composite shooting request, fusing the first target content with the second video by using the first target content as a foreground of the target video and the second video as a background of the target video.
4. The step of extracting second target content from the second video and fusing the second target content with the first video includes:
4. The video composite shooting method according to claim 2 or 3, further comprising the step of: when the composite shooting request is a second composite shooting request, setting the second target content as the foreground of the target video, setting the first video as the background of the target video, and fusing the second target content with the first video.
5. activating an audio capture device based on the composite capture request; 5. The video composition method according to claim 1, further comprising the steps of: obtaining the audio captured by the audio capture device; and using the audio captured by the audio capture device as background music for the target video.
6. extracting audio from the first video; 5. The video composition method of claim 1, further comprising the step of: using the audio of the first video as background music for the target video.
7. receiving a request to add background music input by a user based on the target video; displaying a background music adding interface in response to the background music adding request; receiving an audio selection operation by a user based on the background music adding interface; 5. The video compounding method according to claim 1, further comprising the step of: using the music corresponding to the audio selection operation as background music for the target video.
8. A video composite camera, comprising: a composite photography request receiving module for receiving a composite photography request input by a user based on a first video, the first video being a video uploaded and distributed by another user in the application; a video capture module for activating a video capture device in response to the composite capture request and capturing a second video by the video capture device; a video fusion module for fusing the first video and the second video to obtain a target video, wherein a foreground of the target video is obtained from one of the first video and the second video, and a background of the target video is obtained from the other of the first video and the second video; a video distribution module, receiving a video delivery request input by the user based on the target video; determining a similarity between the target video and the first video in response to the video delivery request; a video distribution module used for: sending the target video to a server of the application when the similarity between the target video and the first video does not exceed a preset threshold, and the server distributes the target video and generates a composite shooting link including a homepage link of the creator of the first video; Determining a similarity between the target video and the first video includes: determining a similarity between the target video and the first video based on whether an image of the user is detected from the target video and a time at which the image of the user appears in the target video.
9. An electronic device, one or more processors; Memory and and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the video composite shooting method according to any one of claims 1 to 7.
10. 8. A computer-readable medium having stored thereon at least one instruction, at least one program segment, code set or instruction set, the at least one instruction, the at least one program segment, code set or instruction set being uploaded and executed by a processor to implement the video compositing method of any one of claims 1 to 7.
Citation Information
Patent Citations
Image photographing apparatus
JP2005223513A
Video camera
JP2010193155A
Image synthesis device and image synthesis method
JP2011029947A
Image printer, image printing system and image printing method
JP2016116117A
Image processing device, image processing method, and program
WO2019163558A1