A video stitching method, apparatus, device and medium

By processing video clips with both images and audio to ensure consistency between the visuals and audio, the problem of disjointedness caused by direct splicing is solved, and the overall visual effect of video splicing is improved.

CN115766973BActive Publication Date: 2026-08-04BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2021-09-02
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing technologies, directly splicing two video clips can easily result in a disjointed feel, leading to a poor overall visual experience.

Method used

By performing image processing on the video clips to be spliced ​​to ensure consistent visual presentation, and processing the audio to ensure consistent background sound, the processed video clips are finally spliced ​​together.

Benefits of technology

It achieves natural transitions and coherence between video clips, enhancing the overall visual appeal of spliced ​​videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766973B_ABST
    Figure CN115766973B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a video splicing method, device, equipment and medium, wherein the method comprises: obtaining a first video segment and a second video segment to be spliced; performing image processing on the first video segment and the second video segment, so that the image-processed first video segment and the image-processed second video segment have the same picture display effect; the picture display effect includes image quality and / or picture style; performing audio processing on the first video segment and the second video segment, so that the audio-processed first video segment and the audio-processed second video segment have the same background sound; and splicing the image-processed and audio-processed first video segment and the image-processed and audio-processed second video segment. Embodiments of the present disclosure can make the splicing transition of two video segments more natural, the spliced video more coherent, and effectively improve the overall sensory effect of the spliced video for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of video processing technology, and in particular to a video splicing method, apparatus, device and medium. Background Technology

[0002] In many applications, it is necessary to splice specific segments of two videos to create a new video. Current technology typically involves directly splicing two video segments together. However, the inventors discovered that because the shooting conditions or post-processing techniques of the two videos are often different, directly splicing the two video segments together results in a noticeable sense of disjointedness in the final video, leading to a poor overall viewing experience for the user. Summary of the Invention

[0003] To solve the above-mentioned technical problems, or at least partially solve them, this disclosure provides a video stitching method, apparatus, device, and medium.

[0004] This disclosure provides a video splicing method, the method comprising: acquiring a first video segment and a second video segment to be spliced; performing image processing on the first video segment and the second video segment to make the image-processed first video segment and the image-processed second video segment have the same display effect; the display effect includes image quality and / or image style; performing audio processing on the first video segment and the second video segment to make the audio-processed first video segment and the audio-processed second video segment have the same background sound; and splicing the image-processed and audio-processed first video segment and the image-processed and audio-processed second video segment together.

[0005] Optionally, the step of performing image processing on the first video segment and the second video segment includes: determining the target display effect; and converting the original display effect of the first video segment and the original display effect of the second video segment into the target display effect.

[0006] Optionally, the step of determining the target screen display effect includes: using a pre-set screen display effect as the target screen display effect; or, determining the target screen display effect based on the original screen display effect of the first video segment and the original screen display effect of the second video segment.

[0007] Optionally, the display effect includes image quality and display style; the step of determining the target display effect based on the original display effect of the first video segment and the original display effect of the second video segment includes: selecting one of the original image quality of the first video segment and the original image quality of the second video segment as the target image quality; selecting one of the original display style of the first video segment and the original display style of the second video segment as the target display style; and determining the target display effect based on the target image quality and the target display style.

[0008] Optionally, the step of selecting one of the original image quality of the first video segment and the original image quality of the second video segment as the target image quality includes: selecting one of the original image quality of the first video segment and the original image quality of the second video segment as the target image quality according to a preset quality selection strategy; wherein the quality selection strategy includes: performing quality selection based on user instructions, or performing quality selection based on the image quality comparison results between the first video segment and the second video segment.

[0009] Optionally, the step of selecting one of the original image styles of the first video segment and the second video segment as the target image style includes: selecting one of the original image styles of the first video segment and the second video segment as the target image style according to a preset style selection strategy; wherein the style selection strategy includes: style selection based on user instructions, style selection based on video source, or style selection based on segment sorting position.

[0010] Optionally, the step of converting the original display effects of the first video segment and the second video segment into the target display effect includes: determining original display effects that are inconsistent with the target display effect based on the original display effects of the first video segment and the second video segment, and using the inconsistent original display effects as display effects to be converted; using a preset image quality conversion algorithm to convert the original image quality in the display effects to be converted into the target image quality in the target display effect; wherein the image quality conversion algorithm includes a conversion algorithm between LDR and HDR; and using a preset style transfer algorithm to transfer the target image style in the target display effect to the display effects to be converted, so as to adjust the original image style of the display effects to be converted to match the target image style.

[0011] Optionally, the step of performing audio processing on the first video segment and the second video segment includes: obtaining the original background audio of the first video segment and the original background audio of the second video segment; determining the target background audio; and converting both the original background audio of the first video segment and the original background audio of the second video segment into the target background audio.

[0012] Optionally, the step of obtaining the original background audio of the first video segment and the original background audio of the second video segment includes: extracting a first specified type of sound contained in the first video segment, and using all other sounds except the first specified type of sound as the original background audio of the first video segment; extracting a second specified type of sound contained in the second video segment, and using all other sounds except the second specified type of sound as the original background audio of the second video segment.

[0013] Optionally, the step of determining the target background sound includes: using a pre-set background sound as the target background sound; or, determining the target background sound based on the original background sound of the first video segment and the original background sound of the second video segment.

[0014] Optionally, the step of determining the target background sound based on the original background sound of the first video segment and the original background sound of the second video segment includes: selecting one of the original background sound of the first video segment and the original background sound of the second video segment as the target background sound; or, fusing the original background sound of the first video segment and the original background sound of the second video segment to obtain the target background sound.

[0015] Optionally, the step of converting the original background audio of the first video segment and the original background audio of the second video segment into the target background audio includes: deleting the original background audio of the first video segment and the original background audio of the second video segment; and uniformly adding the target background audio to the first video segment and the second video segment.

[0016] This disclosure also provides a video splicing device, comprising: a segment acquisition module for acquiring a first video segment and a second video segment to be spliced; an image processing module for performing image processing on the first video segment and the second video segment to make the image-processed first video segment and the image-processed second video segment have the same display effect; the display effect includes image quality and / or image style; an audio processing module for performing audio processing on the first video segment and the second video segment to make the audio-processed first video segment and the audio-processed second video segment have the same background sound; and a segment splicing module for splicing the image-processed and audio-processed first video segment and the image-processed and audio-processed second video segment together.

[0017] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the video stitching method provided in this disclosure.

[0018] This disclosure also provides a computer-readable storage medium storing a computer program for performing the video stitching method provided in this disclosure.

[0019] The technical solution provided in this disclosure first obtains a first video segment and a second video segment to be spliced. Then, image processing and audio processing are performed on the first and second video segments respectively, so that the image-processed first and second video segments have the same visual display effect (image quality and / or visual style); the audio-processed first and second video segments have the same background sound; finally, the image-processed and audio-processed first and second video segments are spliced ​​together. Through this method, the visual display effect and background sound of the two video segments to be spliced ​​are unified, making the splicing transition more natural and the spliced ​​video more coherent. This effectively improves the obvious disjointedness of spliced ​​videos in the prior art and enhances the overall visual experience for users.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0022] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 A flowchart illustrating a video stitching method provided in an embodiment of this disclosure;

[0024] Figure 2 This is a schematic diagram of the structure of an HDR network model provided in an embodiment of the present disclosure;

[0025] Figure 3 This is a schematic diagram of the structure of a style transfer model provided in an embodiment of the present disclosure;

[0026] Figure 4 This is a schematic diagram of video stitching provided in an embodiment of the present disclosure;

[0027] Figure 5 A flowchart illustrating a video stitching method provided in an embodiment of this disclosure;

[0028] Figure 6 This is a schematic diagram of the structure of a video splicing device provided in an embodiment of the present disclosure;

[0029] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0030] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0031] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0032] The inventors discovered through research that the shooting conditions (such as the precision of the shooting equipment, the shooting environment, and the professionalism of the filming personnel) or post-processing techniques (such as image editing and filter processing) of two videos are often different. For example, the differences in both visual and audio performance between clips from movies and TV series and clips from personally shot videos are significant. If these clips are directly spliced ​​together, there will be a noticeable sense of disjointedness. Similarly, videos with different shooting conditions or post-processing techniques will also often exhibit varying degrees of disjointedness when spliced ​​together, resulting in a poor overall viewing experience for the user. To improve this problem, this disclosure provides a video splicing method, apparatus, device, and medium, which will be described in detail below.

[0033] Figure 1 This is a flowchart illustrating a video stitching method provided in an embodiment of the present disclosure. The method can be executed by a video stitching device, which can be implemented using software and / or hardware, and is generally integrated into an electronic device. Figure 1 As shown, the method mainly includes the following steps S102 to S108:

[0034] Step S102: Obtain the first video segment and the second video segment to be spliced.

[0035] In practical applications, the first and second video clips can originate from different videos, such as one from a film or television work and the other from a personally filmed video. This disclosure does not limit the source or filming conditions of the first and second video clips; any two video clips that need to be spliced ​​can be used. By splicing different video clips, a better dramatic effect can be achieved. For example, by cutting a film or television work into multiple segments, and by having the user film matching the content of the segments, and finally splicing all the segments together in chronological order (or the order of events), a storyline with contrast and vitality can be created. It is understood that any two video clips to be spliced ​​can serve as the aforementioned first and second video clips.

[0036] Step S104: Perform image processing on the first video clip and the second video clip to make the image-processed first video clip and the image-processed second video clip have the same display effect; the display effect includes image quality and / or image style.

[0037] Considering that the two main factors influencing the display effect are image quality and image style, in some embodiments, the display effect can be considered to include image quality and / or image style. Image quality (also referred to as picture quality) can be directly characterized by HDR (High Dynamic Range) or LDR (Low Dynamic Range), or by data that directly affects image quality, such as resolution. Image style (also referred to as picture style) is the overall effect expressed by one or more factors such as color tone, brightness, color contrast, and sharpness. Different style names can be preset for different effects, such as Hong Kong and Taiwan style, fresh style, retro style, and everyday life style. In practical applications, the shooting conditions and image processing methods (such as filter processing) of different videos are mostly different, so the final image style also varies. For example, taking filters as an example, videos processed with different filters have different image styles. In the embodiments of this disclosure, the display effect can be characterized by image quality and image style.

[0038] In some implementations, the above image processing includes image quality unification processing and / or screen style unification processing. For example, the target screen display effect can be determined first; then the original screen display effect of the first video segment and the original screen display effect of the second video segment can be converted into the target screen display effect, thereby achieving a unified screen display effect for the two video segments.

[0039] Step S106: Perform audio processing on the first video segment and the second video segment so that the audio-processed first video segment and the audio-processed second video segment have the same background sound.

[0040] Considering that the visual disjointness caused by splicing two video clips is due not only to inconsistent image quality and style, but also to the difference in background audio between the two clips, which is a major reason for the disjointed and unnatural transitions in the resulting video. In some implementations, background audio can be understood as any sound other than a specified type (such as human voice) (e.g., ambient noise). For example, if one video clip has a noisy background audio while the other has a simpler background audio, splicing the two clips together directly often creates discomfort. Taking the above into account, this disclosure embodiment performs audio processing on the first and second video clips. In some implementations, this audio processing includes background audio unification. For example, the original background audio of the first and second video clips can be obtained first; a target background audio can be determined; and then both the original background audio of the first and second video clips can be converted into the target background audio, thereby achieving the effect of unifying the background audio of the first and second video clips.

[0041] Step S108: The first video segment after image and audio processing and the second video segment after image and audio processing are spliced ​​together.

[0042] In some implementations, the visual effects of the first video clip and the second video clip can be unified to the target visual effect, and the background audio of the first video clip and the second video clip can be unified to the target background audio, so that the visual effects and background audio of the processed first video clip and the second video clip are consistent.

[0043] By using the above method, the visual presentation and background sound of the two video clips to be spliced ​​can be unified, making the splicing transition between the two video clips more natural and the spliced ​​video more coherent. This effectively improves the obvious sense of disjointedness that exists in spliced ​​videos in existing technologies and enhances the overall visual experience of spliced ​​videos for users.

[0044] In practical applications, before steps S104 and S106, the images and audio of the first and second video segments can be separated to facilitate unification of the images and audio of the first and second video segments separately. After unification, the unified images and audio can be synthesized to obtain the final merged video.

[0045] In some implementations, the present disclosure provides the following two methods for determining the target screen display effect:

[0046] (1) Use a pre-set display effect as the target display effect. That is, the target display effect can be pre-set according to needs or preferences, such as pre-setting the target image quality and target display style, and finally unify the two video clips to the pre-set target display effect. The advantage of this method is that it is relatively simple to implement. Regardless of the display effect of the first and second video clips, in practical applications, it is only necessary to pre-set the target display effect to unify the two video clips to be spliced ​​according to the target display effect.

[0047] (2) Determine the target display effect based on the original display effect of the first video clip and the original display effect of the second video clip. The advantage of this method is that it is more flexible. It can combine the actual situation of the first video clip and the second video clip to determine the corresponding target display effect. That is, the determined target display effect is related to the original display effect of the first video clip and the original display effect of the second video clip, which is more easily accepted by users and provides a better user experience.

[0048] In some embodiments, taking the display effect as including image quality and display style as an example, the original display effect includes the original image quality and the original display style; the target display effect includes the target image quality and the target display style. In the above steps, the target image quality and target display style can be determined based on the original image quality and original display style of the first video segment and the second video segment. The target image quality can be one of the original image qualities of the two video segments, or it can be different from both of the original image qualities of the two video segments. Similarly, the target display style can be one of the original display styles of the two video segments, or it can be different from both of the original display styles of the two video segments. The specific choice can be determined according to the actual situation and is not limited here.

[0049] In some specific implementation examples, the steps for determining the target display effect based on the original display effect of the first video segment and the original display effect of the second video segment can be performed as follows: steps a to c:

[0050] Step a: Select one of the original image quality of the first video segment and the original image quality of the second video segment as the target image quality.

[0051] In some implementations, a preset quality selection strategy can be used to select one of the original image quality of the first video segment and the original image quality of the second video segment as the target image quality. The quality selection strategy includes: selecting quality based on user instructions, or selecting quality based on a comparison of the image quality between the first and second video segments. For ease of understanding, the following explanation is provided:

[0052] When the quality selection strategy is based on user instructions, prompts can be sent to the user to select the desired image quality from the first video segment and the second video segment, and the target image quality can be determined based on the user's selection.

[0053] When the quality selection strategy is based on a comparison of image quality between the first and second video clips, the target image quality can be pre-set to select the image with better quality from the two clips to provide a better viewing experience for the user. For example, if the first video clip has HDR quality and the second has LDR quality, HDR is superior to LDR, so HDR can be selected as the target image quality. Alternatively, the target image quality can be selected based on factors such as bandwidth / processing speed. Specific settings can be configured according to actual needs and are not limited here.

[0054] Step b: Select one of the original image styles of the first video segment and the second video segment as the target image style.

[0055] In some implementations, a target image style can be selected from the original image style of a first video segment and the original image style of a second video segment, according to a preset style selection strategy. The style selection strategy includes: style selection based on user instructions, style selection based on video source, or style selection based on segment order. For ease of understanding, the following explanation is provided:

[0056] When the quality selection strategy is to select a style based on user instructions, a prompt can be issued to the user, who can then select the desired visual style from the first and second video clips, and the target visual style can be determined based on the user's selection.

[0057] When the quality selection strategy is to select a style based on the video source, the preferred video source can be preset, and the picture style corresponding to the video clip from the preferred video source can be used as the target picture style. For example, the video source includes movies and TV series and user personal works. Assuming that the source of the first video clip is a movie or TV series and the source of the second video clip is a user personal work, and the preferred source of the movie or TV series is preset, then the picture style of the first video clip can be used as the target picture style.

[0058] When the quality selection strategy is based on the segment ranking position for style selection, the selection criterion for ranking position can be preset. For example, the visual style corresponding to the video segment ranked earlier can be prioritized as the target visual style. For instance, assuming the first video segment is before the second video segment (i.e., the first video segment is played first, followed by the second video segment), the visual style corresponding to the first video segment can be prioritized as the target visual style. Of course, the visual style corresponding to the video segment ranked later can also be prioritized as the target visual style. The specific settings can be flexibly configured according to actual needs and are not limited here.

[0059] Step c: Determine the target image display effect based on the target image quality and the target image style. In some implementations, the target image display effect includes both the target image quality and the target image style.

[0060] Through the above steps a to c, the target screen display effect can be determined reasonably. The target image quality and target screen style in the target screen display effect are related to the original image quality and original screen style of the first video segment and the second video segment, which makes the subsequent unified processing of the first video segment and the second video segment smoother and more easily accepted by users.

[0061] After determining the target display effect, the original display effects of both the first and second video clips can be converted to the target display effect. In other words, the processed first and second video clips will both display as the target display effect.

[0062] In some implementations, steps 1 to 3 can be referred to below:

[0063] Step 1: Based on the original display effects of the first and second video clips, identify the original display effects that are inconsistent with the target display effect, and use these inconsistent original display effects as the display effects to be converted. It is understood that the target display effect may be either the original display effect of the first or the original display effect of the second video clip; therefore, only the inconsistent original display effect needs to be selected as the object to be processed.

[0064] Step 2: A preset image quality conversion algorithm is used to convert the original image quality in the display effect to the target image quality in the display effect. The image quality conversion algorithm includes a conversion algorithm between LDR and HDR. In this embodiment, LDR and HDR are mainly used as representation methods of image quality. The LDR and HDR conversion algorithm includes a conversion algorithm for converting LDR to HDR and a conversion algorithm for converting HDR to LDR.

[0065] In some implementations, to present a better display effect to the user, it is assumed that the target image quality is HDR. If the original image quality contains LDR, then the LDR needs to be converted to HDR. For ease of understanding, embodiments of this disclosure provide a conversion algorithm for converting LDR to HDR, which can be implemented using an HDR algorithm network model.

[0066] like Figure 2The diagram illustrates the structure of an HDR network model, primarily comprising parallel local branch networks, extended branch networks, and global branch networks, as well as a stitching and fusion network connected to these networks. The LDR image is input into each of the local, extended, and global branch networks. The local branch network extracts features from the LDR image to obtain first local features; the extended branch network extracts features to obtain second local features; and the second local features are more specific than the first local features. The global branch network then extracts features from the LDR image to obtain global features. Finally, the first, second, and global features are all input into the stitching and fusion network. By stitching and fusing these three features, the final HDR image is obtained. In practical implementation, the local branch network, extended branch network, and global branch network can all be constructed using fully convolutional modules. For example, the input to the global branch network is a 256*256 image. After processing by multiple convolutional modules, it extracts 1*1*64 features. These features contain the global features of the input image. The global branch network needs to perform downsampling when extracting global features, while the local branch network and extended branch network do not perform downsampling, thus better preserving the local features of the image. The final size of the generated local features is consistent with the input image. The stitching and fusion network can include a stitching and fusion layer and convolutional layers. The stitching and fusion layer can be used to stitch and fuse the features output from the three network branches, and the convolutional layers can be used to restore the stitched and fused features to an HDR image through convolutional operations.

[0067] Furthermore, this disclosure also provides a training method for an HDR network model, which can be implemented using supervised learning. For example, a batch of HDR image training samples can be obtained first. For instance, a batch of original HDR images can be collected first, and during the training process, the original HDR images can be randomly sampled and randomly cropped to achieve the effect of sample size expansion, resulting in multiple HDR image samples. Then, a single-frame exposure operator can be used to convert the final HDR image samples into LDR images, thereby establishing HDR and LDR image sample pairs. The HDR network model to be trained is used to convert the LDR image samples to obtain HDR images. Based on a preset loss function, the loss value between the HDR image output by the HDR network model and the HDR image samples (real HDR images) is calculated. This loss value characterizes the degree of difference between the HDR image output by the HDR network model and the HDR image samples. Based on the loss value, gradient descent is used to optimize the parameters of the HDR network model until the loss value meets the preset conditions, at which point the training ends. At this point, the HDR network model can effectively convert LDR images into HDR images that meet expectations.

[0068] It should be noted that the above HDR network model is only an illustrative example and should not be considered a limitation. In practical applications, any algorithm or model that can convert LDR images into HDR images is acceptable.

[0069] Understandably, when splicing clips from films and television shows with clips from personal videos, the difference in image quality between the two can be smoothed out by converting the clips to HDR, since the images from films and television shows are usually in HDR and the images from personal videos are usually in LDR.

[0070] Step 3: Use a preset style transfer algorithm to transfer the style of the target image to the image to be converted, so as to adjust the original style of the image to be converted to match the style of the target image. The matching of the adjusted original style with the target style can be understood as the similarity reaching a preset level.

[0071] In some implementations, the style transfer algorithm includes a color transfer algorithm or a style feature transfer algorithm based on a neural network model. For ease of understanding, exemplary descriptions are given below:

[0072] It is understandable that color is a major factor influencing the style of an image. Therefore, style transfer can be achieved through color migration. Color migration algorithms refer to transferring the colors in the target image to the image to be transformed. For example, for a brief overview, let's assume that the colors on the reference image are transferred to the target image. In practice, the reference image and the target image can first be converted to the LAB color space (also known as the Lab color space). Then, the mean and standard deviation of the pixels in the reference image and the target image in the LAB color space are obtained. For each pixel value in the target image, the mean of the target image can be subtracted. Then, the difference is multiplied by a pre-calculated ratio (that is, the ratio between the standard deviations of the reference image and the target image). Finally, the mean of the reference image is added back. In this way, the original color of the target image can be adjusted, and the overall color performance of the target image after adjustment is similar to that of the reference image.

[0073] The aforementioned color transfer method involves relatively low computational complexity, is easy to implement, and can roughly align the colors of two video clips. It is well-suited for devices with limited data processing capabilities, such as mobile phones. To achieve better style transfer results, a style feature transfer algorithm based on a neural network model, i.e., a deep learning algorithm, can be used. Exemplarily, this disclosure also provides an implementation method for a style transfer model.

[0074] See Figure 3The diagram illustrates the structure of a style transfer model, primarily comprising a VGG encoder, a Transformation network, and a decoder. Further, in... Figure 3 The diagram also illustrates the internal structure of the Transformation network. The following section combines... Figure 3 The principles of style transfer models are explained:

[0075] The first image Ic and the second image Is are input into the VGG encoder to transfer the style of the second image Is to the first image Ic. For example, the first image Ic can be a video frame from a user-captured video, and the second image Is can be an image extracted from a film or television show. The VGG encoder extracts features from the first image Ic and the second image Is respectively, obtaining features Fc and Fs. Then, a Transformation network is used to fuse features Fc and Fs to obtain a new feature Fd. Feature Fd contains both the content features of the first image Ic and the style features of the second image Is. Finally, feature Fd is decoded to recover an RGB image (i.e., ...). Figure 3 (The output image in the image). Furthermore... Figure 3 The diagram also illustrates the internal workings of the Transformation network. Fc undergoes feature extraction via a convolutional module (containing multiple convolutional layers) to obtain Fc'. Fc' is then multiplied by itself to obtain cov(Fc). ′ ), cov(Fc ′ After passing through a fully connected (FC) layer, the first extracted feature is obtained. Similarly, Fs undergoes feature extraction via a convolutional module to obtain Fs'. Fs' is then multiplied by itself to obtain cov(Fs). ′ ), cov(Fs ′ After passing through a fully connected (FC) layer, the second extracted feature is obtained. The first and second extracted features are then multiplied by a matrix to obtain the matrix transpose T. Furthermore, Figure 5 In this context, 'c' represents the compression operation, and 'u' represents the decompression operation.

[0076] The output image of the style transfer model is expected to be consistent with the first image Ic in content (achieving a specified level of similarity) and consistent with the second image Is in style (achieving a specified level of similarity). To achieve this, the loss function required to train the style transfer model consists of two components (see...). Figure 3The VGG loss unit (in the model) includes content loss and style loss. In practice, the output image can be input back into the VGG encoder to extract content features and style features respectively. By comparing the loss between the content features of the output image and the content features of the first image Ic, and comparing the loss between the style features of the output image and the style features of the second image Is, the network parameters of the style transfer model are trained. After training, the resulting style transfer model can output images whose content features are consistent with the content features of the first image Ic, and whose style features are consistent with the style features of the second image Is.

[0077] It should be noted that the style transfer model above is only an illustrative example and should not be regarded as a limitation. In practical applications, any algorithm or model that can achieve style transfer is acceptable.

[0078] Through steps 1 to 3 above, the original display effects of the first video clip and the second video clip can be converted into the target display effect, achieving the goal of unified display effect. This makes the transition between the two video clips after splicing more natural and the overall feel stronger.

[0079] In some embodiments, this disclosure provides specific implementation methods for audio processing of the first video segment and the second video segment, which can be achieved by referring to the following steps A to C:

[0080] Step A: Obtain the original background audio of the first video clip and the original background audio of the second video clip.

[0081] In some implementations, a first specified type of sound can be extracted from a first video segment, and all other sounds besides the first specified type can be used as the original background audio of the first video segment; and a second specified type of sound can be extracted from a second video segment, and all other sounds besides the second specified type can be used as the original background audio of the second video segment. In practical applications, the first specified type of sound and the second specified type of sound can be the same or different. For example, both the first specified type of sound and the second specified type of sound can be human voices, or both can be instrumental sounds, or one can be human voices and the other can be instrumental sounds. The above is only an illustrative example and should not be considered as a limitation. In addition, the first specified type of sound can contain one or more types of sound, and the second specified type of sound can also contain one or more types of sound, and then all other sound types besides the specified type (such as ambient noise, ambient noise, etc.) are used as the original background audio.

[0082] In practical applications, taking the audio of the first video segment as an example, the audio can be separated into its own audio tracks based on a first specified sound type. All other sounds are then considered as the original background audio of the first video segment. For instance, if the first specified sound type is human voice, then human voice is separated from the audio of the first video segment, while other environmental noise is considered as the original background audio.

[0083] Step B: Determine the target background sound. In some embodiments, this disclosure further provides the following two methods for determining the target background sound:

[0084] (1) Use a pre-set background sound as the target background sound. That is, the target background sound can be pre-set according to needs or preferences. The target background sound can be background music, uniform ambient noise, or even blank (mute). This application does not limit the specific form of the target background sound. Finally, the background sounds of both video clips are unified to the pre-set target background sound. The advantage of this method is that it is relatively simple to implement. Regardless of the background sound of the first and second video clips, in practical applications, only the target background sound needs to be pre-set to unify the audio effects of the two video clips to be spliced ​​according to the target background sound.

[0085] Taking background music as the target background sound as an example, in practical applications, default background music can be added automatically, or user-selected background music can be added; there are no restrictions here. By adding background music, the background sound of the two video clips can be unified, and the spliced ​​video can be made more engaging and dramatic. Furthermore, if the target background sound is blank, only the desired sound type (such as only human voice) will be retained in the two video clips. By removing environmental noise from each clip, the audio playback will be cleaner. Additionally, if the target background sound is preset environmental noise, the audio playback effect will be more natural and realistic. The specific target background sound can be set according to actual needs; the above are just examples and should not be considered limitations.

[0086] (2) Determine the target background sound based on the original background sound of the first video segment and the original background sound of the second video segment. The advantage of this method is that it is more flexible and can combine the actual situation of the first and second video segments to determine the corresponding target background sound. That is, the determined target background sound is related to the original background sound of the first video segment and the original background sound of the second video segment, which is more easily accepted by users and provides a better user experience.

[0087] In some specific implementation examples, determining the target background audio based on the original background audio of the first video segment and the original background audio of the second video segment can be done in the following two ways:

[0088] Method 1: Select one of the original background audio from the first video clip and the original background audio from the second video clip as the target background audio. Specifically, a preset background audio selection strategy can be used to select one of the original background audio from the first video clip and the original background audio from the second video clip. This background audio selection strategy includes: selecting background audio based on user quality, selecting background audio based on video source, selecting background audio based on clip order, or selecting background audio based on the comparison results between the first and second video clips. For example, prioritizing the background audio with lower noise level between the two video clips as the target background audio. The implementation methods for other background audio selection strategies can refer to the aforementioned style selection strategies, and will not be elaborated here.

[0089] Method 2: Merge the original background audio from the first video clip and the original background audio from the second video clip to obtain the target background audio. In this method, the background audio from the two video clips can be directly merged into the target background audio, ensuring that the target background audio contains all the background audio elements from both video clips. Any audio fusion algorithm can be used; no specific limitation is made here.

[0090] It should be understood that the above is only an illustrative example, and in practical applications, any method that can determine the target background sound is acceptable.

[0091] Step C: Convert the original background audio of the first video clip and the original background audio of the second video clip into the target background audio.

[0092] For example, this disclosure provides a relatively simple implementation method: deleting the original background audio of the first video segment and the original background audio of the second video segment; and uniformly adding the target background audio to both the first and second video segments. Through this method, a rapid change of background audio can be achieved, resulting in a unified and natural transition between the two video segments.

[0093] In summary, by unifying the visual and audio effects of the first and second video segments and splicing them together, the spliced ​​videos can be presented to users in a consistent manner according to the target visual and background audio effects. The spliced ​​videos are coherent and natural in both visual and audio effects, effectively improving the obvious sense of disjointedness in spliced ​​videos in existing technologies and enhancing the overall visual experience for users.

[0094] The video splicing method provided in this disclosure can be flexibly applied to any two video segments that need to be spliced. For example, two independent videos can be spliced ​​directly according to the above video splicing method, or two independent videos can be divided into multiple video segments and then spliced ​​alternately according to the above video splicing method. In addition, multiple video segments from different sources can be spliced ​​sequentially in a certain order. Regardless of the method, the video splicing method provided in this disclosure can be used to splice the two video segments to be spliced, and finally obtain the spliced ​​merged video (also known as the fused video).

[0095] For ease of understanding, this disclosure provides an application scenario for the above-described video stitching method. See [link to relevant documentation]. Figure 4 The diagram illustrates a video splicing process, showing video A and video B. Video A is divided into video segments A1, A2, and A3, and video B is divided into video segments B1, B2, and B3. Videos A and B are spliced ​​alternately, resulting in a spliced ​​video of A1B1A2B2A3B3. It can be understood that the above video splicing method can be used to splice any two adjacent video segments. The resulting merged video has good overall coherence and consistency, making the video splicing transition more natural and effectively alleviating the disjointed feeling caused by splicing in existing technologies.

[0096] This disclosure does not limit the source of the two video clips to be spliced. In some embodiments, video A is a segment from a film or television drama, and video B is a personal creation. The target visual style is the same as that of video A, and the target audio track is a human voice track. Multiple video clips from video A and video B are then spliced ​​together using the format A1B1A2B2A3B3, achieving the effect of dialogue between characters in a film or television drama and real-life individuals, thus achieving a better dramatic effect. The method of segmenting video clips (segmentation nodes, clip length, etc.) can be determined according to actual needs, and this disclosure does not impose any limitations.

[0097] Furthermore, this disclosure also provides one implementation method for the above-described video stitching method, which can be found in [reference needed]. Figure 5The diagram illustrates a video splicing method. It shows how video {Ai} is split into video and audio components, resulting in video V-Ai and audio A-Ai; video {Bi} is also split into video and audio components, resulting in video V-Bi and audio A-Bi; video V-Ai and video V-Bi are combined to form the video frame to be spliced; and audio A-Ai and audio A-Bi are combined to form the audio frame to be spliced. By normalizing video V-Ai and video V-Bi (i.e., unifying the display effect), processed video V'-Ai and video V'-Bi are obtained. These processed video V'-Ai and video V'-Bi undergo a video transition (which can be understood as a video splicing method), allowing them to be seamlessly spliced ​​together according to a specified transition pattern. By normalizing the audio A-Ai and A-Bi (i.e., unifying the background sound), we can obtain the processed audio A'-Ai and A'-Bi. Applying audio transitions (which can be understood as a type of audio splicing) to these processed audio A'-Ai and A'-Bi allows them to be seamlessly spliced ​​together according to a specified transition pattern. Afterwards, the resulting video and audio can be combined to produce the final output video. Additionally, in... Figure 5 The text briefly illustrates the specific implementation methods of video normalization and audio normalization. In video normalization, normalization can be applied to one or more factors affecting the image display, such as resolution, HDR (i.e., the aforementioned image quality), style (the aforementioned image style), and color. It's understandable that style typically includes color, but... Figure 5 The separate listing of color indicates that in practical applications, normalization can also be based solely on color. In audio normalization processing, normalization can be applied to one or more factors affecting audio playback, such as gain, vocals, and noise. Specifically, this can include gain adjustment, vocal extraction, and noise reduction. Vocals refer to the aforementioned types of sound, while noise can be considered background noise other than vocals, thus requiring noise reduction / denoising. It's understandable that different videos are shot in different environments, resulting in significant variations in environmental noise. Direct splicing would create a noticeable disjointedness and incongruity. Therefore, the audio tracks of the videos to be spliced ​​can be separated, such as separating the vocal track and the environmental noise track. In some implementations, only the vocals from both video segments are retained. By removing environmental noise, the resulting spliced ​​video transitions more realistically and naturally.

[0098] also, Figure 5 This only briefly illustrates a few influencing factors in the normalization processing of audio and video, and does not list them all, so it should not be regarded as a limitation.

[0099] Furthermore, in order to create an atmosphere, Figure 5 Furthermore, background music was added. By removing environmental noise from each video and adding background music uniformly, not only was the background sound of the two video clips unified, but a better artistic effect was also created.

[0100] In summary, the video splicing method provided in this disclosure can make the splicing transition between two video segments more natural, and the spliced ​​video more coherent, effectively improving the overall visual experience of the spliced ​​video for the user.

[0101] Corresponding to the aforementioned video stitching method, this disclosure provides a video stitching device. Figure 6 This is a schematic diagram of a video splicing device provided in an embodiment of this disclosure. The device can be implemented by software and / or hardware and is generally integrated into an electronic device. Figure 6 As shown, the device includes:

[0102] The segment acquisition module 602 is used to acquire the first video segment and the second video segment to be spliced;

[0103] The image processing module 604 is used to perform image processing on the first video segment and the second video segment so that the image-processed first video segment and the image-processed second video segment have the same display effect; the display effect includes image quality and / or image style.

[0104] The audio processing module 606 is used to perform audio processing on the first video segment and the second video segment so that the first video segment and the second video segment after audio processing have the same background sound.

[0105] The segment splicing module 608 is used to splice the first video segment after image processing and audio processing and the second video segment after image processing and audio processing.

[0106] The above-mentioned device can unify the visual display and background sound of the two video clips to be spliced, making the splicing transition between the two video clips more natural and the spliced ​​video more coherent. It effectively improves the obvious sense of disjointedness in spliced ​​videos in the existing technology and enhances the overall visual experience of spliced ​​videos for users.

[0107] In some implementations, the image processing module 604 is specifically used to: determine the target image display effect; and convert the original image display effect of the first video segment and the original image display effect of the second video segment into the target image display effect.

[0108] In some implementations, the image processing module 604 is specifically used to: use a preset image display effect as the target image display effect; or, determine the target image display effect based on the original image display effect of the first video segment and the original image display effect of the second video segment.

[0109] In some implementations, the display effect includes image quality and display style;

[0110] The image processing module 604 is specifically used to: select one of the original image quality of the first video segment and the original image quality of the second video segment as the target image quality; select one of the original picture style of the first video segment and the original picture style of the second video segment as the target picture style; and determine the target picture display effect based on the target image quality and the target picture style.

[0111] In some implementations, the image effect determination module 604 is specifically used to: select one of the original image quality of the first video segment and the original image quality of the second video segment as the target image quality according to a preset quality selection strategy; wherein, the quality selection strategy includes: quality selection based on user instructions, or quality selection based on the image quality comparison result between the first video segment and the second video segment.

[0112] In some implementations, the image processing module 604 is specifically used to: select one of the original image styles of the first video segment and the second video segment as the target image style according to a preset style selection strategy; wherein, the style selection strategy includes: style selection based on user instructions, style selection based on video source, or style selection based on segment sorting position.

[0113] In some embodiments, the image processing module 604 is specifically used to: determine original image display effects that are inconsistent with the target image display effect based on the original image display effects of the first video segment and the original image display effects of the second video segment, and use the inconsistent original image display effects as the image display effects to be converted; use a preset image quality conversion algorithm to convert the original image quality in the image display effects to be converted into the target image quality in the target image display effect; wherein, the image quality conversion algorithm includes a conversion algorithm between LDR and HDR; use a preset style transfer algorithm to transfer the target image style in the target image display effect to the image display effect to be converted, so as to adjust the original image style of the image display effect to match the target image style.

[0114] In some implementations, the audio processing module 606 is specifically used to: acquire the original background audio of the first video segment and the original background audio of the second video segment; determine the target background audio; and convert both the original background audio of the first video segment and the original background audio of the second video segment into the target background audio.

[0115] In some embodiments, the audio processing module 606 is specifically used to: extract a first specified type of sound contained in the first video segment, and use all other sounds except the first specified type of sound as the original background sound of the first video segment; extract a second specified type of sound contained in the second video segment, and use all other sounds except the second specified type of sound as the original background sound of the second video segment.

[0116] In some implementations, the audio processing module 606 is specifically used to: use a pre-set background sound as the target background sound; or, determine the target background sound based on the original background sound of the first video segment and the original background sound of the second video segment.

[0117] In some implementations, the audio processing module 606 is specifically used to: select one of the original background audio of the first video segment and the original background audio of the second video segment as the target background audio; or, merge the original background audio of the first video segment and the original background audio of the second video segment to obtain the target background audio.

[0118] In some implementations, the audio processing module 606 is specifically used to: delete the original background sound of the first video segment and the original background sound of the second video segment; and uniformly add the target background sound to the first video segment and the second video segment.

[0119] The video splicing device provided in this disclosure can execute the video splicing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0120] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device embodiments can be referred to the corresponding process in the method embodiments, and will not be repeated here.

[0121] This disclosure provides an electronic device, which includes: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement any of the video stitching methods described above. Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Figure 7As shown, the electronic device 700 includes one or more processors 701 and memory 702.

[0122] The processor 701 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 700 to perform desired functions.

[0123] The memory 702 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 701 may execute the program instructions to implement the video stitching method of the embodiments of this disclosure described above and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.

[0124] In one example, the electronic device 700 may also include an input device 703 and an output device 704, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0125] In addition, the input device 703 may also include, for example, a keyboard, a mouse, etc.

[0126] The output device 704 can output various information to the outside, including determined distance information, direction information, etc. The output device 704 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0127] Of course, for the sake of simplicity, Figure 7 Only some of the components of the electronic device 700 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 700 may include any other suitable components depending on the specific application.

[0128] In addition to the methods and devices described above, embodiments of this disclosure may also be computer program products, including computer program instructions that, when executed by a processor, cause the processor to perform the video stitching method provided in the embodiments of this disclosure.

[0129] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0130] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the video stitching method provided in embodiments of this disclosure.

[0131] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0132] This disclosure also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the video stitching method of this disclosure.

[0133] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0134] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video stitching method, characterized in that, include: Obtain the first and second video segments to be spliced; The first video clip and the second video clip are image processed to make the image-processed first video clip and the image-processed second video clip have the same visual display effect; the visual display effect includes image quality and / or visual style. The first video segment and the second video segment are subjected to audio processing so that the audio-processed first video segment and the audio-processed second video segment have the same background sound; The first video segment after image and audio processing and the second video segment after image and audio processing are spliced ​​together; The step of performing audio processing on the first video segment and the second video segment includes: Extract the first specified type of sound contained in the first video segment, and use all other sounds except the first specified type of sound as the original background sound of the first video segment; extract the second specified type of sound contained in the second video segment, and use all other sounds except the second specified type of sound as the original background sound of the second video segment; there is a difference between the original background sound of the first video segment and the original background sound of the second video segment; Determine the target background sound; The original background audio of the first video segment and the original background audio of the second video segment are both converted into the target background audio.

2. The method according to claim 1, characterized in that, The step of performing image processing on the first video segment and the second video segment includes: Determine the desired display effect; The original display effects of the first video segment and the original display effects of the second video segment are both converted into the target display effect.

3. The method according to claim 2, characterized in that, The step of determining the display effect of the target image includes: Use the preset display effect as the target display effect; or, The target display effect is determined based on the original display effect of the first video segment and the original display effect of the second video segment.

4. The method according to claim 3, characterized in that, The display effect includes image quality and image style; The step of determining the target display effect based on the original display effect of the first video segment and the original display effect of the second video segment includes: Choose one of the original image quality of the first video segment and the original image quality of the second video segment as the target image quality; Choose one of the original image styles of the first video segment and the second video segment as the target image style; The target image display effect is determined based on the target image quality and the target image style.

5. The method according to claim 4, characterized in that, The step of selecting one of the original image quality of the first video segment and the original image quality of the second video segment as the target image quality includes: According to a preset quality selection strategy, one of the original image quality of the first video segment and the original image quality of the second video segment is selected as the target image quality; wherein, the quality selection strategy includes: quality selection based on user instructions, or quality selection based on the image quality comparison results between the first video segment and the second video segment.

6. The method according to claim 4, characterized in that, The step of selecting one of the original image styles of the first video segment and the second video segment as the target image style includes: According to a preset style selection strategy, one of the original image styles of the first video segment and the second video segment is selected as the target image style; wherein, the style selection strategy includes: style selection based on user instructions, style selection based on video source, or style selection based on segment sorting position.

7. The method according to claim 2, characterized in that, The step of converting the original display effect of the first video segment and the original display effect of the second video segment into the target display effect includes: Based on the original display effect of the first video segment and the original display effect of the second video segment, an original display effect that is inconsistent with the target display effect is determined, and the inconsistent original display effect is taken as the display effect to be converted. A preset image quality conversion algorithm is used to convert the original image quality in the display effect to be converted into the target image quality in the display effect; wherein, the image quality conversion algorithm includes a conversion algorithm between LDR and HDR; A preset style transfer algorithm is used to transfer the style of the target image in the target image display effect to the image display effect to be converted, so as to adjust the original image style of the image display effect to match the style of the target image.

8. The method according to claim 1, characterized in that, The step of determining the target background sound includes: Use the preset background sound as the target background sound; or, The target background audio is determined based on the original background audio of the first video segment and the original background audio of the second video segment.

9. The method according to claim 8, characterized in that, The step of determining the target background audio based on the original background audio of the first video segment and the original background audio of the second video segment includes: Select one of the original background audio from the first video segment and the original background audio from the second video segment as the target background audio; or, The original background audio of the first video segment and the original background audio of the second video segment are merged to obtain the target background audio.

10. The method according to claim 1, characterized in that, The step of converting the original background audio of the first video segment and the original background audio of the second video segment into the target background audio includes: Delete the original background audio from the first video segment and the original background audio from the second video segment; The target background sound is added uniformly to both the first video segment and the second video segment.

11. A video splicing device, characterized in that, include: The segment acquisition module is used to acquire the first video segment and the second video segment to be spliced. An image processing module is used to perform image processing on the first video clip and the second video clip so that the image-processed first video clip and the image-processed second video clip have the same display effect; the display effect includes image quality and / or image style; An audio processing module is used to perform audio processing on the first video segment and the second video segment so that the first video segment and the second video segment after audio processing have the same background sound; wherein, the background sound is a sound other than a specified type of sound; The segment splicing module is used to splice the first video segment after image processing and audio processing and the second video segment after image processing and audio processing; The audio processing module is specifically used for: extracting a first specified type of sound contained in the first video segment, and using all other sounds besides the first specified type of sound as the original background sound of the first video segment; extracting a second specified type of sound contained in the second video segment, and using all other sounds besides the second specified type of sound as the original background sound of the second video segment; determining a target background sound; converting both the original background sound of the first video segment and the original background sound of the second video segment into the target background sound; wherein there is a difference between the original background sound of the first video segment and the original background sound of the second video segment.

12. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the video stitching method according to any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The storage medium stores a computer program for executing the video stitching method according to any one of claims 1-10.