Video synthesis method and device, computer device and storage medium

By synthesizing dynamic foreground and background materials, this method solves the problem of low video viewership in existing technologies, and enriches video content while increasing viewing time.

CN122340311APending Publication Date: 2026-07-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2020-08-27
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

In existing technologies, users can only perform simple trimming and add music to the materials they want to upload to the network, which cannot effectively improve the video's viewership.

Method used

By acquiring user-selected dynamic foreground and background materials, a video is synthesized and sent to a live broadcast room or electronic conference, allowing viewers or participants to watch the video with the changed background.

Benefits of technology

It enriched the video content, improved the video viewing experience and playback rate, and increased the video viewing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122340311A_ABST
    Figure CN122340311A_ABST
Patent Text Reader

Abstract

This application is a divisional application of Chinese application 202010876955.9. This application relates to a video compositing method, apparatus, computer device, and storage medium. The method includes: entering a dynamic background material acquisition page; selecting a dynamic background material in response to a background material selection operation triggered on the dynamic background material acquisition page; entering a video compositing preview page; playing the dynamic background material; and displaying a dynamic foreground material in an editing state overlaid on the playing dynamic background material; when a video compositing operation is triggered on the video compositing preview page, compositing the dynamic background material and the dynamic foreground material played on the video compositing preview page into a video. This method can merge at least two materials of different types into a video of interest to the user, thereby increasing viewership.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese application No. 202010876955.9, filed on August 27, 2020, entitled "Video Synthesis Method, Apparatus, Computer Equipment and Storage Medium". Technical Field

[0002] This application relates to the field of computer technology, and in particular to a video synthesis method, apparatus, computer device, and storage medium. Background Technology

[0003] With the development of multimedia and internet technologies, users can download various materials (such as videos) from content delivery networks for viewing, and upload their own footage to the network. For self-shot footage, users can edit it using editors (such as cropping and splicing) and add music before uploading it to the network for other viewers. However, the above approach only allows users to perform simple cropping and music additions on the footage they want to upload, thus failing to effectively increase viewership. Summary of the Invention

[0004] Therefore, it is necessary to provide a video synthesis method, apparatus, computer equipment, and storage medium that can merge at least two materials of different types into a video that users are interested in, thereby improving the viewing rate, in order to address the above-mentioned technical problems.

[0005] A video synthesis method, the method comprising: The system acquires user-selected materials, which are either live streaming videos or ongoing electronic conference videos; these user-selected materials are used to provide dynamic foreground materials. Receive background material selection operation, select dynamic background material; the dynamic background material includes presentation; The dynamic background material and the dynamic foreground material are combined into a video; The synthesized video is sent to the live broadcast room or to the user device in the electronic conference, so that the live broadcast audience or conference participants can watch the live broadcast video or conference video with the changed video background.

[0006] A video synthesis apparatus, the apparatus comprising: The acquisition module is used to acquire user-selected materials, which are either live video or video of an ongoing electronic meeting; the user-selected materials are used to provide dynamic foreground materials. The selection module is used to receive background material selection operations and select dynamic background materials; the dynamic background materials include presentations. A compositing module is used to combine the dynamic background material and the dynamic foreground material into a video; The sending module is used to send the synthesized video to the live broadcast room or to the user equipment in the electronic conference, so that the live broadcast audience or conference participants can watch the live broadcast video or conference video after the video background has been changed.

[0007] In one embodiment, the entry module is used to access the dynamic background material acquisition page; The selection module is used to select a dynamic background material in response to a background material selection operation triggered on the dynamic background material acquisition page; The entry module is also used to enter the video synthesis preview page; The playback module is used to play the dynamic background material and overlay the dynamic foreground material in the editing state onto the playing dynamic background material for display; wherein, the dynamic foreground material is extracted from user-selected materials; The compositing module is used to combine the dynamic background material and the dynamic foreground material playing on the video compositing preview page into a video when a video compositing operation is triggered on the video compositing preview page.

[0008] In one embodiment, the user-selected material includes a user-selected target video or moving image; the device further includes: The extraction module is used to extract frame images from the target video or dynamic image; and to perform image segmentation on the dynamic foreground material in the frame image using a machine learning model.

[0009] In one embodiment, the apparatus further includes: The determination module is used to determine the image mask corresponding to the first frame image if the extracted frame image is the first frame image; The separation module is used to perform channel separation on the extracted frame images as the current frame images if the extracted frame image is not the first frame image, so as to obtain the channel images corresponding to each color channel. The extraction module is further configured to input the channel image of the current frame image and the mask of the previous frame image of the current frame image into a machine learning model for processing to obtain a prediction mask corresponding to each frame image; and to segment dynamic foreground material from the corresponding frame image based on the image mask and the prediction mask.

[0010] In one embodiment, the apparatus further includes: The adjustment module is used to adjust the display parameters of the dynamic foreground material; The compositing module is further configured to combine the dynamic background material and the dynamic foreground material, which are played on the video compositing preview page and have had their display parameters adjusted, to obtain a composite video.

[0011] In one embodiment, the adjustment module is further configured to adjust the orientation and size of the dynamic foreground material; Duplicate the dynamic foreground material and adjust the relative positions between at least two of the duplicated dynamic foreground materials.

[0012] In one embodiment, the playback module is further configured to determine the starting composite frames of the dynamic foreground material and the dynamic background material respectively; play the dynamic background material with the starting composite frame of the dynamic background material as the starting playback position, and display the dynamic foreground material in the editing state superimposed on the playing dynamic background material with the starting composite frame of the dynamic background material as the starting playback position; The compositing module is further configured to combine the dynamic background material and the dynamic foreground material that are played on the video compositing preview page and are located after the starting compositing frame to obtain a compositing video.

[0013] In one embodiment, the dynamic background material includes a presentation; the determining module is further configured to segment the presentation page by page into presentation images to be used as backgrounds; and to determine a starting composite frame in the dynamic foreground material and the segmented presentation images respectively; The compositing module is further configured to compose each demonstration image with at least one corresponding frame of dynamic foreground material in the demonstration image and the dynamic foreground material played on the video compositing preview page and located after the starting compositing frame.

[0014] In one embodiment, the compositing module is further configured to use a local video editor to send the dynamic background material and the dynamic foreground material, which are playing on the video compositing preview page and located after the starting compositing frame, to a server, so that the server can perform frame-by-frame compositing of the dynamic foreground material and the dynamic background material to obtain a composite video.

[0015] In one embodiment, the selection module is further configured to select a dynamic background material corresponding to the background material selection operation from a local material library; or download a dynamic background material corresponding to the background material selection operation from an online material library; or determine the video obtained from the current shooting target environment as the dynamic background material.

[0016] In one embodiment, the apparatus further includes: The extraction module is also used to extract frame image samples from the material samples; The separation module is also used to perform channel separation on the extracted frame image samples as the current frame image samples to obtain the channel images corresponding to each color channel. The extraction module is also used to sequentially input the channel image of the current frame image sample and the image mask of the previous frame image sample of the current frame image sample into an untrained machine learning model for processing, so as to obtain the training prediction mask corresponding to each frame image sample. The calculation module is used to calculate the error value between each training prediction mask and the corresponding label; The adjustment module is further configured to adjust the model parameters of the machine learning model according to the error value until the error value between the predicted image mask output by the adjusted machine learning model and the corresponding label is less than the error threshold, and then stop training.

[0017] In one embodiment, the apparatus further includes: The label acquisition module is used to determine the training prediction mask as the image mask of the specified image sample in the extracted frame image samples when the training prediction mask corresponding to the current frame image is obtained and the error value between the training prediction mask and the corresponding label is less than a preset error value.

[0018] In one embodiment, the target video includes live video or conference video; the device further includes: The video acquisition module is used to acquire live video generated when live streaming is conducted through a live streaming room on a social application, or meeting video generated when electronic meetings are conducted in a conference room on the social application. The sending module is used to send the composite video corresponding to the live video or the conference video to the user equipment accessing the live room or the conference room when the composite video is obtained.

[0019] In one embodiment, the compositing module is further configured to identify key information in the dynamic background material; if the dynamic foreground material obscures the key information, the dynamic foreground material obscuring the key information is hidden or faded; and the hidden or faded dynamic foreground material is combined with the dynamic background material to form a video.

[0020] In one embodiment, the compositing module is further configured to identify the facial expressions of the target object in the dynamic foreground material; obtain corresponding scene elements based on the facial expressions; and combine the scene elements, the dynamic background material played on the video compositing preview page, and the dynamic foreground material into a video.

[0021] In one embodiment, the synthesis module is further configured to extract speech from the user-selected material; perform speech recognition on the extracted speech to obtain corresponding text; obtain matching prompt information based on the semantic information of the text; and synthesize the prompt information, the dynamic background material and the dynamic foreground material played on the video synthesis preview page into a video.

[0022] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program performing the following steps: Upon entering the dynamic background material acquisition page, in response to the background material selection operation triggered on the dynamic background material acquisition page, the dynamic background material is selected. Enter the video synthesis preview page, play the dynamic background material, and overlay the dynamic foreground material in editing mode onto the playing dynamic background material for display; wherein, the dynamic foreground material is extracted from user-selected materials; When a video compositing operation is triggered on the video compositing preview page, the dynamic background material and the dynamic foreground material playing on the video compositing preview page will be combined into a video.

[0023] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Upon entering the dynamic background material acquisition page, in response to the background material selection operation triggered on the dynamic background material acquisition page, the dynamic background material is selected. Enter the video synthesis preview page, play the dynamic background material, and overlay the dynamic foreground material in editing mode onto the playing dynamic background material for display; wherein, the dynamic foreground material is extracted from user-selected materials; When a video compositing operation is triggered on the video compositing preview page, the dynamic background material and the dynamic foreground material playing on the video compositing preview page will be combined into a video.

[0024] The aforementioned video compositing method, apparatus, computer equipment, and storage medium, upon entering the dynamic background material acquisition page, respond to a background material selection operation triggered on the dynamic background material acquisition page by selecting the dynamic background material; upon entering the video compositing preview page, the dynamic background material is played, and the dynamic foreground material in editing mode is overlaid on the playing dynamic background material for display; when a video compositing operation is triggered on the video compositing preview page, the dynamic background material and dynamic foreground material playing on the video compositing preview page are combined, allowing users to customize the video background, changing the original background to a video background of interest, thereby enriching the video content, improving the video viewing effect, and helping to increase video playback rate and video viewing time. Attached Figure Description

[0025] Figure 1 This is an application environment diagram of the video synthesis method in one embodiment; Figure 2 This is a flowchart illustrating a video synthesis method in one embodiment; Figure 3a This is a schematic diagram showing a preview of user-selected materials in one embodiment; Figure 3b This is a schematic diagram illustrating the selection of dynamic background material in one embodiment; Figure 4 This is a schematic diagram illustrating the segmentation of target objects using a machine learning model in one embodiment. Figure 5 This is a schematic diagram illustrating the segmentation of a target object in one embodiment; Figure 6 This is a schematic diagram illustrating the process of adjusting the parameters of dynamic foreground materials and compositing them in one embodiment; Figure 7 This is a schematic diagram illustrating the synthesis of video frames from dynamic foreground material and a presentation slide in one embodiment. Figure 8 This is a schematic diagram illustrating the synthesis of video frames from dynamic foreground material and a presentation in another embodiment; Figure 9 This is a flowchart illustrating the steps of combining dynamic foreground material with a presentation to create a video, as shown in one embodiment. Figure 10a This is a flowchart illustrating the model training steps in one embodiment; Figure 10b This is a schematic diagram illustrating the fading of dynamic foreground material in one embodiment; Figure 10c This is an illustration of adding a large smiley face to the background when a target character in a dynamic foreground material starts to smile, as shown in one embodiment. Figure 10d This is a schematic diagram of facial feature points in one embodiment; Figure 10e This is a schematic diagram illustrating the addition of prompt information in the background when a target character in the dynamic foreground material issues a prompt voice in one embodiment. Figure 11 This is a flowchart illustrating the video synthesis method in another embodiment; Figure 12 This is a flowchart illustrating the video synthesis method in another embodiment; Figure 13 This is a structural block diagram of a video synthesis apparatus in one embodiment; Figure 14 This is a structural block diagram of the video synthesis apparatus in another embodiment; Figure 15 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0027] The video synthesis method provided in this application can be applied to, for example... Figure 1 The application environment shown includes terminal 102, server 104, and terminal 106. Terminal 102 can acquire or collect materials locally (i.e., user-selected materials, such as non-live videos, animated images, or live videos), or download corresponding user-selected materials from server 104 according to user download operations. It then extracts all dynamic foreground materials (such as images of people or other target objects) from the user-selected materials and displays a preview of the dynamic foreground materials. Upon entering the dynamic background material acquisition page, a dynamic background material (such as a background video or presentation) is selected from the page based on the triggered background material selection operation. Upon entering the video synthesis preview page, both the dynamic background material and the dynamic foreground material are played simultaneously, with the dynamic foreground material positioned above the image layer of the dynamic background material. When a video synthesis operation is triggered on the video synthesis preview page, the dynamic background material and dynamic foreground material playing on the video synthesis preview page are synthesized into a video, which is then uploaded to server 104 or sent to terminal 106.

[0028] Among them, terminal 102 and terminal 106 can be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, etc., but are not limited to these.

[0029] Server 104 can be a standalone physical server or a server cluster consisting of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud servers, cloud databases, cloud storage, and content delivery networks (CDN).

[0030] Terminal 102 and server 104 can be connected via Bluetooth, USB (Universal Serial Bus) or network, etc., and this application does not impose any restrictions.

[0031] In one embodiment, such as Figure 2 As shown, a video synthesis method is provided, which can be derived from... Figure 1Executed by the terminal in the middle, or by Figure 1 The terminal and server in the process work together to apply this method to Figure 1 Taking terminal 102 as an example, the explanation includes the following steps: S202 displays a preview of the dynamic foreground material extracted from user-selected materials.

[0032] The user-selected material can be a target video or a dynamic image. The target video can be any of the following: non-live video, live video, or the video of an ongoing electronic conference. The dynamic foreground material can be an image sequence of the target object extracted from the user-selected material. The target object can be a real person, a virtual person, or other animals or plants, etc. The preview screen can be any frame image from the dynamic foreground material, an image corresponding to the dynamic foreground material that contains the target object, or an instruction image used to guide video compositing, such as an instruction image showing how to perform video compositing.

[0033] In one embodiment, the video compositing method can be applied to a social application that is configured with a video editor or other video player with video editing capabilities, or to a live streaming room or meeting room created through the social application. Taking a social application with a video editor as an example, when performing video compositing, the terminal can call the social application's video editor, which has a video compositing preview page. Through this preview page, dynamic background and foreground materials can be played, edited, and composited. S202 may specifically include: when playing user-selected materials through the social application's video player, if a video compositing instruction is received, the terminal performs background removal on the user-selected materials to extract the corresponding dynamic foreground material, simultaneously calling the social application's video editor, and displaying a preview of the dynamic foreground material on the video editor's video compositing preview page. Furthermore, this video compositing method can also be applied to other applications, such as video applications.

[0034] The video player, live streaming room, and conference room mentioned above can all have the function of editing and compositing videos.

[0035] In one embodiment, the terminal can display not only previews of dynamic foreground materials but also, based on browsing instructions, various dynamic foreground materials from the user-selected materials, along with their backgrounds. Upon receiving a background deletion instruction, the background is removed from the user-selected materials, thus displaying only the corresponding dynamic foreground material. Furthermore, after deleting the background, a default background can be used to display the dynamic foreground material along with it, such as... Figure 3aAs shown, the original background is replaced with a default white curtain, which is then displayed alongside the portrait image. In this preview, the portrait image is editable, as indicated by the dashed box; users can adjust the size of the portrait image by resizing the dashed box.

[0036] In one embodiment, if the user-selected material is a non-live video, the terminal can select the corresponding target video from a local video library or an online video library as the user-selected material; alternatively, the terminal can select the corresponding dynamic image from a local image library or an online image library as the user-selected material; or, the terminal can capture video of the target object in the environment using its built-in camera as the user-selected material. If the user-selected material is a presentation, the terminal can select the corresponding presentation from the file library, such as... Figure 3b As shown.

[0037] In another embodiment, if the user-selected material is a live video or a conference video, the steps for obtaining the user-selected material include: the terminal obtaining the live video generated when a live broadcast is conducted through a live room on a social application, or the conference video generated when an electronic conference is held in a conference room on a social application.

[0038] In one embodiment, after obtaining user-selected materials, the terminal extracts the dynamic foreground materials corresponding to each frame from the user-selected materials.

[0039] In one embodiment, when the user-selected material includes a user-selected target video or dynamic image, the step of extracting dynamic foreground material includes: the terminal extracting frame images from the target video or dynamic image; and performing image segmentation on the dynamic foreground material in the frame images using a machine learning model.

[0040] The machine learning model can be a convolutional neural network model or a neural network model that can be used for image segmentation. A frame image can refer to the image corresponding to each frame in the target video or moving image. For example, assuming the total number of frames in the target video is n, then a frame image can refer to the images of frames 1 to n in the target video. Here, n is a positive integer greater than 1.

[0041] In one embodiment, the terminal can decode the target video to obtain a series of frame images, and then use a machine learning model to segment the dynamic foreground material in the series of frame images. Furthermore, the terminal can extract frame images containing the target object from the series of frame images, and then use a machine learning model to segment the dynamic foreground material in the extracted frame images.

[0042] In another embodiment, the terminal can extract the corresponding frame images from the dynamic images, and then perform image segmentation on the dynamic foreground material in the frame image using a machine learning model. Furthermore, the terminal can also select frame images containing the target object from the extracted frame images, and then perform image segmentation on the dynamic foreground material in the selected frame images using a machine learning model.

[0043] In one embodiment, if the extracted frame image is the first frame image, the image mask corresponding to the first frame image is determined; if the extracted frame image is not the first frame image, the extracted frame images are sequentially used as the current frame image for channel separation to obtain the channel images corresponding to each color channel; the above-mentioned steps of performing image segmentation on dynamic foreground material in frame images using a machine learning model may specifically include: inputting the channel image of the current frame image and the mask of the previous frame image of the current frame image into the machine learning model for processing to obtain the prediction mask corresponding to each frame image; segmenting the dynamic foreground material from the corresponding frame image according to the image mask and the prediction mask.

[0044] Before performing image segmentation using a machine learning model, the terminal sequentially separates non-first frame images as the current frame image to obtain the channel images corresponding to each color channel (such as RGB channels).

[0045] For example, assuming the total number of frames in a video is n (n is a positive integer greater than 1), if image segmentation is currently performed on the first frame, the terminal determines the image mask corresponding to the first frame, which represents the region of the target object. Then, the image of the target object is segmented from the first frame based on this image mask. If image segmentation is currently performed on the i-th frame (i is a positive integer greater than 1 and less than or equal to n), channel separation is performed on the i-th frame to obtain the R-channel image, G-channel image, and B-channel image of the i-th frame. Then, the R-channel image, G-channel image, and B-channel image of the i-th frame, along with the image mask corresponding to the (i-1)-th frame, are input into a machine learning model. This machine learning model uses the image mask corresponding to the (i-1)-th frame as prior knowledge to process the R-channel image, G-channel image, and B-channel image of the i-th frame, thereby predicting the image mask corresponding to the i-th frame. Figure 5 As shown, the image of the target object is segmented from the i-th frame image based on the image mask, thus obtaining the image sequence containing the target object from the 1st to the nth frame. Figure 4 The image mask prediction process for a kitten is shown. It can also predict the image masks of other target objects (such as people), and then segment the image of the target object based on the image mask. Figure 5 As shown.

[0046] In one embodiment, when the terminal extracts the dynamic foreground material from the user-selected material, it can display one frame of the extracted dynamic foreground material as a preview screen, or obtain an image corresponding to the target object in the dynamic foreground material as a preview screen.

[0047] For example, such as Figure 4 As shown in Figure 5, images of the target object are extracted from frame A of the video. When images of the target object are extracted from each frame of the video, an image sequence of the target object can be obtained.

[0048] S204, enter the dynamic background material acquisition page, and in response to the background material selection operation triggered on the dynamic background material acquisition page, select the dynamic background material.

[0049] The dynamic foreground material is extracted from user-selected materials. The dynamic background material acquisition page can be a page in a social application or other application used to acquire dynamic background materials, and may include one or more of the following: a page for acquiring dynamic background materials from a local material library (i.e., a local dynamic background material acquisition page), a page for acquiring dynamic background materials from an online material library (i.e., an online dynamic background material acquisition page), and a page for capturing the real-world environment to obtain dynamic background materials (i.e., a real-time dynamic background material capture page). It should be noted that S202 and S204 can be performed in any order, and the two steps can be independent of each other or interconnected.

[0050] Dynamic background material refers to the background used as dynamic foreground material, which can be the target video, dynamic image, or presentation (PPT).

[0051] In one embodiment, when the preview screen is displayed for a specified time, or when a dynamic background material selection instruction is received, the user enters the dynamic background material acquisition page, such as the page in a social application used to acquire dynamic background materials from the local material library.

[0052] In one embodiment, after entering the dynamic background material acquisition page, the terminal will detect the background material selection operation triggered on the dynamic background material acquisition page in real time, and then select the dynamic background material corresponding to the background material selection operation through the dynamic background material acquisition page. For example, entering the page of a social application used to obtain dynamic background materials from the local material library, the terminal can obtain the dynamic background material stored in the local material library that corresponds to the background material selection operation through this page.

[0053] In one embodiment, when a background material selection operation triggered on the dynamic background material acquisition page is detected, a dynamic background material corresponding to the background material selection operation is selected from the local material library; or, a dynamic background material corresponding to the background material selection operation is downloaded from the online material library; or, the video obtained from the current shooting target environment is determined as the dynamic background material.

[0054] For example, you can obtain dynamic background materials stored in your local library and corresponding to the background material selection operation through the local dynamic background material acquisition page of a social application. Alternatively, you can obtain dynamic background materials stored in an online library and corresponding to the background material selection operation through the online dynamic background material acquisition page of a social application. Or, you can control the camera to capture video of the target environment through the real-time dynamic background material acquisition page of a social application and identify that video as the dynamic background material.

[0055] S206, enter the video composition preview page, play dynamic background material, and overlay the dynamic foreground material in editing state onto the playing dynamic background material for display.

[0056] The "editing state" refers to the dynamic foreground material being in an editable state, such as adjusting its size and orientation.

[0057] In one embodiment, after selecting a dynamic background material, the terminal automatically enters the video compositing page. Alternatively, after selecting a dynamic background material, the system detects confirmation actions triggered on the dynamic background material acquisition page in real time. When a confirmation action is detected, the terminal enters the video compositing preview page of the video editor. This video editor can be a standalone editor, a video editing function within a video player, or a built-in video editing plugin.

[0058] In one embodiment, before playing the dynamic background material, the terminal aligns the dynamic background material with the dynamic foreground material, and then while playing the dynamic background material, the dynamic foreground material is superimposed on the playing dynamic background material for display, that is, the dynamic background material and the dynamic foreground material are played simultaneously.

[0059] Alignment can refer to aligning the first frame (or first page) of the dynamic background material with the first frame of the dynamic foreground material so that the dynamic background material and dynamic foreground material can be composited starting from the first frame during compositing. Alternatively, alignment can refer to aligning the first frame of the dynamic foreground material with a specified frame (or page) in the dynamic background material so that the compositing can start from the first frame of the dynamic foreground material and the specified frame (or page) in the dynamic background material during compositing.

[0060] S208: When a video compositing operation is triggered on the video compositing preview page, the dynamic background material and dynamic foreground material playing on the video compositing preview page will be combined into a video.

[0061] The video merging preview page can include a video merging button. Triggering this button initiates the video merging process. The video merging operation can be performed by clicking or touching the video merging button on the video merging preview page.

[0062] In one embodiment, the starting composite frame for the dynamic background and dynamic foreground materials is determined, and the dynamic background and dynamic foreground materials played on the video composite preview page are composited into a video starting from the starting frame. If the starting composite frame for both the dynamic background and dynamic foreground materials is the first frame, the dynamic background and dynamic foreground materials are composited into a video starting from the first frame. If the starting composite frame for the dynamic foreground material is the first frame and the starting composite frame for the dynamic background material is not the first frame (assumed to be the i-th frame), the composite is started from the i-th frame of the dynamic background material and the first frame of the dynamic foreground material.

[0063] For example, assuming there are m frames of dynamic background material and n frames of dynamic foreground material, the composite is started from the i-th frame of the background material and the j-th frame of the foreground material, resulting in a video containing dynamic background material from frame i to frame m and dynamic foreground material from frame i to frame n. Here, m and n are both positive integers greater than 1, i and j are both positive integers greater than or equal to 1, and i is less than or equal to m, and j is less than or equal to n.

[0064] In one embodiment, if the user selects live video or conference video as the source material, when a composite video of the live video or conference video is obtained, the terminal sends the composite video to the user device accessing the live broadcast room or the accessing conference room, so that the live broadcast audience can watch the live video with the changed background, or the conference participants can watch the conference video with the changed background.

[0065] For example, when a streamer broadcasts live via a social media app, their image can be extracted from the captured video footage. The original background can then be replaced with video material that interests the streamer or viewers. This extracted image can then be combined with the new video material to create a composite video. As a result, viewers see a live stream with a different background, which can increase interaction between the streamer and viewers, and also increase viewer engagement in the live stream.

[0066] For example, in live-streaming scenarios involving teaching or business promotion, if there's no whiteboard to display the teaching materials or business explanations during the live stream, viewers can only see the host and cannot simultaneously view the materials. The solution in this embodiment allows for the composite image of the host in the live video with the teaching materials or business explanations, resulting in a video with the materials as a background. Viewers then see the host standing in front of a projector explaining the corresponding teaching materials or business explanations, as shown in the image.

[0067] In the above embodiments, upon entering the dynamic background material acquisition page, in response to the background material selection operation triggered on the dynamic background material acquisition page, the dynamic background material is selected; upon entering the video synthesis preview page, the dynamic background material is played, and the dynamic foreground material in the editing state is superimposed on the played dynamic background material for display; when the video synthesis operation is triggered on the video synthesis preview page, the dynamic background material and dynamic foreground material played on the video synthesis preview page are synthesized, so that users can customize the video background, change the original background to the video background of interest, thereby enriching the video content, improving the video viewing effect, and helping to increase the video playback rate and video viewing time.

[0068] In one embodiment, such as Figure 6 As shown, the method may further include: S602 simultaneously plays dynamic foreground and background materials on the video composition preview page.

[0069] In this scenario, dynamic foreground material is overlaid on top of the dynamic background material for playback.

[0070] S604 adjusts the display parameters of dynamic foreground materials.

[0071] When playing dynamic foreground and background materials, if you need to adjust the shape of the dynamic foreground material, you can pause playback and then adjust the display parameters of the dynamic foreground material.

[0072] In one embodiment, S604 may specifically include: the terminal adjusting the orientation and size of the dynamic foreground material; and / or, copying the dynamic foreground material and adjusting the relative position between at least two copied dynamic foreground materials. The orientation may refer to direction and position.

[0073] For example, such as Figure 7As shown, taking the dynamic background material as the anchor's image and the dynamic background material as the presentation (PPT) as an example, the terminal can adjust the direction, position and size of the image displayed on the PPT layer according to the input adjustment command, so that the anchor is in a suitable position so that the live audience can see the content of the PPT.

[0074] For example, such as Figure 8 As shown, taking the dynamic background material as the target person's image and the dynamic background material as a video of watching the sunrise at the beach as an example, the terminal can copy multiple images of the person displayed on the beach sunrise video according to the input adjustment command, and adjust the relative position and size of each image, so that multiple target people can be seen watching the sunrise at the beach.

[0075] S606, When a video compositing operation is triggered on the video compositing preview page, the dynamic background material and dynamic foreground material that are playing on the video compositing preview page and whose display parameters have been adjusted are combined to obtain a composite video.

[0076] The detailed synthesis process of S606 can be found in S208 of the above embodiment.

[0077] In the above embodiments, users can adjust the position and size of dynamic foreground materials, and can copy multiple dynamic foreground materials, thereby giving the video a better visual experience and improving the video viewing effect and viewership.

[0078] In one embodiment, such as Figure 9 As shown, S206 may specifically include: S902, determine the starting composite frame for the dynamic foreground material and the dynamic background material respectively.

[0079] In one embodiment, when the dynamic background material is a presentation, S902 may specifically include: the terminal dividing the presentation into presentation images to be used as backgrounds by page; and determining the starting composite frame in the dynamic foreground material and the divided presentation images respectively.

[0080] S904 plays the dynamic background material starting from the initial composite frame of the dynamic background material, and displays the dynamic foreground material in editing state overlaid on the playing dynamic background material starting from the initial composite frame of the dynamic background material.

[0081] In one embodiment, when the dynamic background material is a presentation, the terminal aligns the presentation image with the dynamic foreground material frame by frame before playing the presentation image frame by frame. Then, while playing the presentation image frame by frame, the dynamic foreground material is superimposed on the playing presentation image for display, that is, the presentation image and the dynamic foreground material are played at the same time.

[0082] S906 will combine the dynamic background and foreground materials playing on the video composition preview page, which are located after the starting composition frame, to obtain the composite video.

[0083] In one embodiment, when the dynamic background material is a presentation, S906 may specifically include: compositing each presentation image with at least one corresponding frame of dynamic foreground material from the presentation image and dynamic foreground material played on the video synthesis preview page and located after the initial synthesis frame. In another embodiment, S906 may specifically include: using a local video editor to send the dynamic background material and dynamic foreground material played on the video synthesis preview page and located after the initial synthesis frame to a server, so that the server can perform frame-by-frame synthesis of the dynamic foreground material and dynamic background material to obtain a synthesized video.

[0084] In one embodiment, when the dynamic background material is a presentation, the starting composite frame of the presentation image and the dynamic foreground material is determined, and the presentation image and dynamic foreground material played on the video composite preview page are composited into a video using the starting frame. If the starting composite frames of both the presentation image and the dynamic foreground material are the first frame, the presentation image and dynamic foreground material are composited into a video starting from the first frame. If the starting composite frame of the dynamic foreground material is the first frame and the starting composite frame of the presentation image is not the first frame (assumed to be the i-th frame), the composite is started from the i-th frame of the presentation image and the first frame of the dynamic foreground material.

[0085] For example, suppose there are m frames of demo images (i.e., m demo images) and n frames of dynamic foreground material. The composite is started with the i-th frame of the background material and the j-th frame of the dynamic foreground material, resulting in a video containing demo images from frame i to frame m and dynamic foreground material from frame i to frame n. Here, m and n are both positive integers greater than 1, i and j are both positive integers greater than or equal to 1, and i is less than or equal to m, and j is less than or equal to n.

[0086] In one embodiment, if the user selects live video or conference video as the source material, when a composite video of the live video or conference video is obtained, the terminal sends the composite video to the user's device that accesses the live room or the conference room, so that the live viewers can watch the live video with the background changed to a presentation, or the conference participants can watch the conference video with the background changed to a presentation.

[0087] For example, in live-streaming scenarios involving teaching or business promotion, if there's no whiteboard to display the teaching materials or business explanations, viewers can only see the host and cannot simultaneously view the materials. The solution in this embodiment combines the host's image with the teaching materials or business explanations in the live-stream video, creating a video with the materials as a background. Viewers then see the host standing in front of a projector explaining the corresponding materials, as shown in the image.

[0088] In the above embodiments, by determining the starting composite frame of the dynamic foreground and dynamic background materials, and using the determined starting composite frame as the seven points for compositing, the dynamic background and dynamic foreground materials are composited. This allows users to select the parts of the dynamic background material they are interested in for compositing, resulting in a more visually appealing composite video. Furthermore, when the dynamic background material is a presentation, the presentation can be used as the background and composited with the dynamic foreground material. This creates a video where the dynamic foreground object interacts with the presentation, allowing the video recorder to composite the presentation and the presenter into the same video even in scenarios without a projector. Viewers can seamlessly watch the presenter and the presentation, enhancing the viewing experience.

[0089] In one embodiment, such as Figure 10a As shown, the training steps for a machine learning model include: S1002, extract frame image samples from material samples.

[0090] The source material samples can be target videos or moving images. The target video can be any of the following: non-live video, live video, or the video of an ongoing electronic conference. Frame image samples can refer to the images of each frame within the source material samples.

[0091] In one embodiment, after extracting frame image samples from the material samples, the terminal also obtains the tags corresponding to each frame image sample.

[0092] S1004, the extracted frame image samples are used as the current frame image samples for channel separation to obtain the channel images corresponding to each color channel.

[0093] In one embodiment, the terminal takes the currently synthesized frame image sample as the current frame image sample, and then performs channel separation on the current frame image sample to obtain the channel image corresponding to each RGB channel.

[0094] For example, taking video as the source material sample, assuming the total number of frames in the video is n (n is a positive integer greater than 1), if we are performing image segmentation on the i-th frame (i is a positive integer greater than 1 and less than or equal to n), then we will perform channel separation on the i-th frame image sample to obtain the R-channel image, G-channel image and B-channel image of the i-th frame image sample.

[0095] S1006, sequentially input the channel image of the current frame image sample and the image mask of the previous frame image sample into the untrained machine learning model for processing, to obtain the training prediction mask corresponding to each frame image sample.

[0096] In one embodiment, if the current frame image sample is the first frame, the terminal can directly calculate the image mask of the current frame image sample and use the image mask of the first frame as the image mask of a specified image sample in the frame image sample. Furthermore, the terminal can replace the specified image sample in the extracted frame image sample with the current frame image sample so that the replaced specified image sample can be used for subsequent training.

[0097] In one embodiment, if the current frame image sample is not the first frame, the channel image of the current frame image sample and the image mask of the previous frame image sample are input into an untrained machine learning model for processing.

[0098] For example, when the source sample is a video, the terminal inputs the R-channel image, G-channel image, and B-channel image of the i-th frame image sample and the image mask corresponding to the (i-1)-th frame image sample into the machine learning model. The machine learning model uses the image mask corresponding to the (i-1)-th frame image sample as prior knowledge to process the R-channel image, G-channel image, and B-channel image of the i-th frame image sample, thereby predicting the training prediction mask corresponding to the i-th frame image sample.

[0099] S1008, calculate the error value between each training prediction mask and the corresponding label.

[0100] In one embodiment, the terminal calculates the error value between each trained prediction mask and its corresponding label based on a loss function. The loss function can be any of the following: Mean Squared Error, Cross-Entropy Loss, L2 Loss, and Focal Loss.

[0101] In one embodiment, to enable the machine learning model to learn that the presence of a target object image in a frame of image samples affects the prediction of the image mask for that frame, the terminal can obtain the label in the following manner: when the training prediction mask corresponding to the current frame image is obtained, and the error value between the training prediction mask and the corresponding label is less than a preset error value, the terminal determines the training prediction mask as the label corresponding to the specified image sample in the extracted frame image samples. Therefore, during the training of the machine learning model, when the specified image sample is used as the current frame image sample for synthesis, the machine learning model predicts the training prediction mask of the specified image sample, compares the training prediction mask with the image mask of the specified image sample, obtains the loss value between the two, and then executes S1010.

[0102] In another embodiment, when the image mask of the first frame is calculated, the terminal can also use the image mask of the first frame as the image mask of a specified image sample in the extracted frame image samples. Thus, during the training of the machine learning model, when the specified image sample is used as the current frame image sample for synthesis, the machine learning model predicts the training prediction mask of the specified image sample, compares the training prediction mask with the image mask of the specified image sample, obtains the loss value between the two, and then executes S1010. Thus, when a target object suddenly appears in front of the camera, the impact of the appearance of the target object on the prediction of the image mask of the subsequent frame image sample can also be effectively resolved.

[0103] S1010 adjusts the model parameters of the machine learning model based on the error value.

[0104] In one embodiment, training is stopped when all frame image samples are input into the machine learning model for training, and the error between the predicted image mask output by the adjusted machine learning model and the corresponding label is less than the error threshold.

[0105] In one embodiment, the terminal backpropagates the loss value to each layer of the machine learning model to obtain the gradients for the model parameters of each layer; and adjusts the model parameters of each layer in the machine learning model according to the gradients.

[0106] In the above embodiments, the channel images of each color channel of the current frame image sample and the image mask of the previous frame image sample are input into the machine learning model. The image mask of the previous frame image sample is used as prior knowledge to predict the training prediction mask of the current frame image sample. The error value between the training prediction mask and the corresponding label is calculated. The model parameters of the machine learning model are adjusted according to the error value. As a result, the trained machine learning model can quickly segment the dynamic foreground material in the user-selected material, which improves the image segmentation speed and makes video synthesis faster, meeting the real-time requirements.

[0107] In one embodiment, S210 may further include: the terminal identifying key information in the dynamic background material; if the dynamic foreground material obscures the key information, the dynamic foreground material obscuring the key information is hidden or faded; and the hidden or faded dynamic foreground material is combined with the dynamic background material to form a video.

[0108] The key information can be important images, text, or animations within the dynamic background material. Hiding can involve extracting the target object from the dynamic foreground material that obscures the key information, or replacing the obscuring dynamic foreground material with blank material. Fading can involve adjusting the transparency of the dynamic foreground material that obscures the key information; additionally, the color can be adjusted simultaneously (e.g., turning a dark tone into a light one). Thus, in the composite video after hiding and fading, the obscured key information can be viewed by the user. It should be noted that the audio corresponding to the dynamic foreground material is preserved during the hiding or fading process.

[0109] For example, such as Figure 10b As shown, when the target person in the dynamic foreground material obscures the content in the PPT, the transparency of the target person can be adjusted.

[0110] In one embodiment, the hiding or fading of the dynamic foreground material will cease once the obscured key information has finished playing, or once the obscured key information has moved to an unobscured area (e.g., when a slide in a PowerPoint presentation changes its position). Therefore, when the previously obscured key information disappears or moves to an unobscured area, the dynamic foreground material will be displayed again.

[0111] In the above embodiments, when important content (i.e. key information) is obscured, the dynamic foreground material that obscures the important content can be hidden or faded. Thus, when watching the composite video, the important content can still be viewed even when it is obscured by the dynamic foreground material, thereby improving the viewing effect of the composite video.

[0112] In one embodiment, S210 may further include: the terminal recognizing the facial expression of the target object in the dynamic foreground material; obtaining the corresponding scene element based on the facial expression; and compositing the scene element, the dynamic background material played on the video synthesis preview page, and the dynamic foreground material into a video.

[0113] Among these, "appropriate elements" can refer to patterns, virtual emoticons, or text that accompany facial expressions, such as a smiling face or a clear sky image. "Face" can broadly refer to the human face, chin, lips, eyes, nose, eyebrows, forehead, and ears. Correspondingly, facial expressions can be composed of various postures of the human face, chin, lips, eyes, nose, eyebrows, forehead, and ears, such as pouting, blinking, and smiling.

[0114] For example, such as Figure 10c As shown, when the target person in the dynamic foreground material is laughing happily, a smiley face (or a clear sky pattern, not shown in the image) can be captured. Then, during video compositing, the captured smiley face (or clear sky pattern) is composited into the corresponding image position. Thus, when the composite video is played, you can see the target person laughing happily while also seeing a smiley face (or clear sky pattern) in the background. Alternatively, when the target person in the dynamic foreground material has a gloomy expression, a dynamic dark cloud can be composited with the dynamic foreground material and the dynamic background material into the video. Thus, when the composite video is played, the dynamic dark cloud will be displayed on the video background.

[0115] In one embodiment, the step of a terminal recognizing the facial expression of a target object in dynamic foreground material may specifically include: the terminal extracting eye feature points from the dynamic foreground material; determining a first distance between an upper eyelid feature point and a lower eyelid feature point, and determining a second distance between a left corner eye feature point and a right corner eye feature point; and determining the eye pose based on the relationship between the ratio of the first distance and the second distance and at least one preset interval.

[0116] The left and right corner feature points refer to the left and right corner feature points of the same eye, respectively. For example, for the left eye, the left corner feature point refers to the left corner feature point of the left eye, and the right corner feature point refers to the right corner feature point of the left eye.

[0117] For example, such as Figure 10d As shown, the terminal calculates the first distance between the upper eyelid feature point 38 and the lower eyelid feature point 42, and the second distance between the left corner feature point 37 and the right corner feature point 40, based on the facial feature points obtained by facial recognition technology. When the ratio between the first distance and the second distance is 0, the test subject is determined to be squinting with the left eye. When the ratio between the first distance and the second distance is less than 0.2, the test subject is determined to be blinking with the left eye. When the ratio between the first distance and the second distance is greater than 0.2 and less than 0.6, the test subject is determined to be staring with the left eye.

[0118] The following methods can be used to identify lip posture: Method 1: Identify lip posture based on the height of lip feature points.

[0119] In one embodiment, the facial expression recognition step further includes: the terminal extracting lip feature points from dynamic foreground material; and determining the lip posture based on the height difference between the lip center feature point and the lip corner feature point among the lip feature points.

[0120] For example, such as Figure 10d As shown, the terminal determines the height of the lip center feature point 63 and the height of the lip corner feature point 49 (or 55), and then calculates the height difference between the lip center feature point 63 and the lip corner feature point 49 (or 55). If the height difference is positive (i.e., the height of the lip center feature point 63 is higher than the height of the lip corner feature point 49 or 55), then the lip posture is determined to be a smile.

[0121] Method 2: Identify lip posture based on the distance between feature points of the upper and lower lips.

[0122] In one embodiment, among the lip feature points, the terminal determines the lip pose based on a third distance between the upper lip feature point and the lower lip feature point.

[0123] For example, the third distance between the upper lip feature points and the lower lip feature points is compared with a distance threshold. When the third distance reaches the distance threshold, the lip pose position (open mouth) is determined. Figure 10d As shown, the third distance between the upper lip feature point 63 and the lower lip feature point 67 is compared with the distance threshold. If it is greater than or equal to the distance threshold, the test object is identified as having an open mouth.

[0124] In another embodiment, the terminal calculates a fourth distance between the feature points of the left and right corners of the lip. When the third distance is greater than or equal to the fourth distance, the lip posture is determined to be an open mouth.

[0125] Method 3 identifies lip posture based on the ratio of the distance between the upper and lower lips to the distance between the feature points of the left and right corners of the lips.

[0126] In one embodiment, the lip pose is determined based on the relationship between the ratio of a third distance to a fourth distance and at least one preset interval among the lip feature points; wherein the fourth distance is the distance between the left corner lip feature point and the right corner lip feature point.

[0127] For example, such as Figure 10dAs shown, the terminal determines a third distance between the upper lip feature point 63 and the lower lip feature point 67, and a fourth distance between the left lip corner feature point 49 and the right lip corner feature point 55. The relationship between the ratio of the third distance and the fourth distance and at least one preset interval determines the lip posture. For example, when the ratio of the third distance and the fourth distance is within a first preset interval, the lip posture is determined to be open; furthermore, when the ratio of the third distance and the fourth distance is within a second preset interval, the lip posture is determined to be closed. The values ​​in the first preset interval are all greater than the values ​​in the second preset interval.

[0128] In one embodiment, the facial expression recognition step further includes: extracting eyebrow feature points and eyelid feature points from dynamic foreground material; determining a fifth distance between the eyebrow feature points and eyelid feature points; and determining the eyebrow pose based on the relationship between the fifth distance and a preset distance.

[0129] For example, such as Figure 10d As shown, the distance between the eyebrow feature point and the eyelid feature point 38 is calculated. When the distance is greater than the preset distance, it is determined that the subject is raising its eyebrows.

[0130] In the above embodiments, by recognizing facial expressions, appropriate elements matching the recognized facial expressions are obtained, and these appropriate elements are combined with dynamic background materials and dynamic foreground materials to synthesize a video, making the video content more vivid and increasing the fun and appeal of the synthesized video.

[0131] In one embodiment, S210 may further include: the terminal extracting speech from user-selected materials; performing speech recognition on the extracted speech to obtain corresponding text; obtaining matching prompt information based on the semantic information of the text; and combining the prompt information, dynamic background material, and dynamic foreground material played on the video synthesis preview page into a video.

[0132] For example, such as Figure 10e As shown, when the target person in the dynamic foreground material says "This is very interesting," the prompt information corresponding to that voice is obtained, such as "Students, look at the PPT, pay attention to the important points." Then, this prompt information is combined with the dynamic background material and the dynamic foreground material to form a video. Thus, when the composite video is played, the background screen dynamically displays "Students, look at the PPT, pay attention to the important points."

[0133] In one embodiment, when the terminal obtains speech, it extracts acoustic features from the speech and obtains the recognized text based on the acoustic features. The acoustic features may include, but are not limited to, PNCC (power-normalized cepstral coefficients) and MFCC (Mel Frequency Cepstrum Coefficient).

[0134] In one embodiment, the step of extracting acoustic features from speech may specifically include: the terminal framing the speech and determining the power spectrum based on the spectrum of each frame of speech; obtaining the logarithmic power spectrum corresponding to the power spectrum; determining the logarithmic power spectrum as a speech feature, or determining the result of the logarithmic power spectrum after discrete cosine transform as a speech feature.

[0135] For example, suppose the signal expression of the collected speech is: The audio after frame splitting and windowing is For the voice after adding a window Performing a Discrete Fourier Transform yields the corresponding spectral signal:

[0136] Where N represents the number of points in the Discrete Fourier Transform.

[0137] When obtaining the spectrum of each frame of speech, the terminal calculates the corresponding power spectrum and obtains the logarithmic power spectrum by calculating the logarithmic value of the power spectrum. The logarithmic power spectrum is then input into a Mel-scale triangular filter, and after discrete cosine transform, Mel-frequency cepstral coefficients are obtained. The obtained Mel-frequency cepstral coefficients are:

[0138] Substituting the logarithmic energy above into the discrete cosine transform, we obtain the L-order Mel frequency cepstral parameters, where L refers to the order of the Mel frequency cepstral coefficients, which can take values ​​from 12 to 16. M refers to the number of triangular filters.

[0139] In one embodiment, before extracting acoustic features, the terminal can first perform speech enhancement (such as noise reduction) on the extracted speech, and then extract acoustic features from the speech after speech enhancement.

[0140] In the above embodiments, the corresponding text is obtained through speech recognition, and the matching prompt information is obtained based on the semantic information of the text. The prompt information is then combined with dynamic background material and dynamic foreground material to form a video, so that users can see the prompt information that matches the speech while hearing the speech, thus focusing the user's attention on the content displayed in the video.

[0141] As an example, when a user selects a target video (including live or non-live videos), the background of the target video can be changed through video compositing, such as... Figure 11 As shown, the video compositing method is described below: S1102, extract the human image from each video frame of the target video.

[0142] S1104, Replace the background in the target video with a background curtain, thereby removing the original background from the target video.

[0143] S1106 allows users to select appropriate video or PPT materials based on their needs.

[0144] For example, you can select video or PPT materials of interest from your local video or file library. Alternatively, you can choose from the video templates provided by the system, or download video or PPT materials of interest from the online video or file library.

[0145] Taking corporate roadshows as an example, during a corporate roadshow, presenters may need to explain using dynamic video or PPT materials, similar to a weather forecaster, or interact with and explain using video or PPT materials. The solution in this application allows for real-time extraction of the human image from the roadshow video, and then replacing the background with other video or PPT materials. Furthermore, the extracted human image can be freely flipped, enlarged, reduced, and tilted onto the new background (i.e., the replaced video or PPT material).

[0146] S1108 combines the extracted human image with selected video or PPT footage to create a composite video with a changed background.

[0147] In one embodiment, during the video compositing process, the extracted human image can be adjusted in position, size, and angle, and then the adjusted human image can be composited with the selected video material or PPT material.

[0148] In another embodiment, multiple portraits can be copied during the video compositing process to create clones. These copied portraits are then combined with selected video or PowerPoint footage.

[0149] As another example, such as Figure 12 As shown, the video compositing method is described below: S1202 extracts human images from each video frame of the target video using a neural network model.

[0150] For example, convolutional neural network models from machine learning can be used to solve semantic segmentation tasks, thereby segmenting human images from each frame of a target video. Specifically, a network architecture and training process suitable for mobile terminals (such as smartphones) are designed by satisfying the following requirements and constraints.

[0151] 1) Constructing the training dataset Frame images with a wide variety of foreground poses (such as human poses) and background environments are acquired as training samples and annotated to provide high-quality data for the machine learning process. These annotations include pixel-level precise localization of foreground elements, such as hair, glasses, neck, skin, and lips; background labels generally achieve cross-validation results that match human-level annotation quality.

[0152] 2) Data Input The segmentation task can refer to calculating binary image masks for frame images of the target video to segment foreground elements from the background environment, and then using the image mask of the previous frame as prior knowledge and the three-channel image of the next frame as input into the neural network model for training.

[0153] Specifically, the current frame image is separated into channels to obtain images of each color channel, namely, the R channel image, G channel image, and B channel image. Then, the R channel image, G channel image, and B channel image of the current frame image, along with the image mask of the previous frame image, are input into a neural network model for training to predict the image mask of the current frame image.

[0154] 3) Training process During segmentation, it's necessary to maintain temporal continuity between frames while also considering temporal discontinuities, such as a person suddenly appearing in front of the camera. To robustly train the model and address these issues, the ground truth labels for each frame can be transformed in several ways and used as a mask for the previous frame: Clear the preceding image mask: Given that the neural network model has correctly processed the first frame image and the new target in the scene, obtain the image mask for that first frame image, and then use that image mask as the label for subsequent frames. This will simulate a scene where someone appears in the camera's viewfinder.

[0155] Affine transformation of the ground truth mask: The Minor transformation is performed on the neural network model to propagate and adjust the image mask of the previous frame; while the Major transformation discards the image mask that the neural network model deems unsuitable.

[0156] The converted image: Thin plate splines smoothing of the original image was implemented to speed up camera movement and rotation.

[0157] S1204, the client will upload the human image extracted from the target video to the server.

[0158] S1206, Obtain video or PPT materials.

[0159] S1208: The client uploads the acquired video or PPT materials to the server.

[0160] Users select video or PPT materials from the client and upload them to the server. After receiving the uploaded video or PPT materials, the server will overlay the video or PPT materials with the cut-out human image.

[0161] During the overlay process, if video footage is overlaid, the extracted portrait and video footage will be played simultaneously according to time. If the user selects to adjust or edit the start timeline of the video footage on the client, the portrait and video footage will be overlaid starting from the selected start timeline.

[0162] If you want to overlay PPT materials, first import the PPT. This will break down each slide into images or, if the PPT has dynamic effects, into image frames. During the overlay process, you can combine a single image or frame from the PPT with a short video clip depicting a person's actions.

[0163] The basic process of real-time overlay or splicing is as follows: 1) Calculate the time parameters, image parameters, and fusion parameters for all videos.

[0164] 2) Use the parameters obtained in the first step to complete real-time overlay or splicing.

[0165] S1210, the client adjusts parameters such as size and orientation of the human image in the target video.

[0166] After overlaying the portrait with the corresponding video footage, the portrait can be shrunk, rotated, and copied. When a user selects a portrait on the client, a selection box appears, and the user can then perform actions such as shrinking, rotating, and copying. The client processes this operation information locally first.

[0167] S1212, the client sends the adjusted parameters.

[0168] S1214, the server will composite the adjusted portrait with video or PPT materials.

[0169] S1216, the server returns the resulting composite video to the client.

[0170] After the user saves and confirms, the client requests the server. The server performs video compositing based on the processing operations on the client and then returns the resulting composite video to the client for display.

[0171] The solution described above allows for real-time extraction of human images from a target video, and the background of the target video can be replaced with other videos or animated PPT presentations. In corporate roadshow scenarios, this enables presentations to be given in conjunction with animated videos or PPT materials, as well as interactive presentations that combine video or PPT content.

[0172] It should be understood that, although Figure 2 , 6 The steps in flowcharts 9-12 are shown sequentially as indicated by the arrows; however, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order requirement for the execution of these steps, and they can be performed in other orders. Furthermore, Figure 2 , 6 At least some of the steps in 9-12 may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.

[0173] In one embodiment, such as Figure 13 As shown, a video compositing device is provided. This device can be a software module, a hardware module, or a combination of both, integrated into a computer device. Specifically, the device includes: an entry module 1302, a selection module 1304, a playback module 1306, and a compositing module 1308, wherein: Enter module 1302 to access the dynamic background material acquisition page; Module 1304 is selected in response to a background material selection operation triggered on the dynamic background material acquisition page to select a dynamic background material; The entry module 1302 is also used to enter the video synthesis preview page; The playback module 1306 is used to play the dynamic background material and overlay the dynamic foreground material in the editing state onto the playing dynamic background material for display; wherein, the dynamic foreground material is extracted from user-selected materials; The compositing module 1308 is used to combine the dynamic background material and the dynamic foreground material playing on the video compositing preview page into a video when a video compositing operation is triggered on the video compositing preview page.

[0174] In the above embodiments, upon entering the dynamic background material acquisition page, in response to the background material selection operation triggered on the dynamic background material acquisition page, the dynamic background material is selected; upon entering the video synthesis preview page, the dynamic background material is played, and the dynamic foreground material in the editing state is superimposed on the played dynamic background material for display; when the video synthesis operation is triggered on the video synthesis preview page, the dynamic background material and dynamic foreground material played on the video synthesis preview page are synthesized, so that users can customize the video background, change the original background to the video background of interest, thereby enriching the video content, improving the video viewing effect, and helping to increase the video playback rate and video viewing time.

[0175] In one embodiment, user-selected materials include user-selected target videos or animated images; such as... Figure 14 As shown, the device also includes: an extraction module 1310; wherein: The extraction module 1310 is used to extract frame images from target videos or dynamic images; and to perform image segmentation on dynamic foreground materials in the frame images using a machine learning model.

[0176] In one embodiment, such as Figure 14 As shown, the device also includes: The determination module 1312 is used to determine the image mask corresponding to the first frame image if the extracted frame image is the first frame image; The separation module 1314 is used to perform channel separation on the extracted frame images as the current frame images if the extracted frame image is not the first frame image, so as to obtain the channel images corresponding to each color channel. The extraction module 1310 is also used to input the channel image of the current frame image and the mask of the previous frame image of the current frame image into the machine learning model for processing, to obtain the prediction mask corresponding to each frame image; and to segment the dynamic foreground material from the corresponding frame image according to the image mask and the prediction mask.

[0177] In one embodiment, such as Figure 14 As shown, the device also includes: Adjustment module 1316 is used to adjust the display parameters of dynamic foreground materials; The compositing module 1308 is also used to composite the dynamic background material and dynamic foreground material that are played on the video compositing preview page and whose display parameters have been adjusted, to obtain a composite video.

[0178] In one embodiment, the adjustment module 1316 is further configured to adjust the orientation and size of the dynamic foreground material; copy the dynamic foreground material; and adjust the relative position between at least two copied dynamic foreground materials.

[0179] In one embodiment, the playback module 1306 is further configured to determine the starting composite frames of the dynamic foreground material and the dynamic background material respectively; play the dynamic background material with the starting composite frame of the dynamic background material as the starting playback position, and display the dynamic foreground material in the editing state superimposed on the playing dynamic background material with the starting composite frame of the dynamic background material as the starting playback position. The compositing module 1308 is also used to composite the dynamic background material and dynamic foreground material that are played on the video compositing preview page and are located after the starting compositing frame to obtain a composite video.

[0180] In the above embodiments, users can adjust the position and size of dynamic foreground materials, and can copy multiple dynamic foreground materials, thereby giving the video a better visual experience and improving the video viewing effect and viewership.

[0181] In one embodiment, the dynamic background material includes a presentation; the determining module 1312 is further configured to segment the presentation page by page into presentation images to be used as backgrounds; and to determine the starting composite frame in the dynamic foreground material and the segmented presentation images respectively; The compositing module 1308 is also used to composite each demonstration image with at least one corresponding frame of dynamic foreground material in the demonstration images and dynamic foreground material played on the video compositing preview page and located after the starting compositing frame.

[0182] In one embodiment, the compositing module 1308 is further configured to use a local video editor to send the dynamic background material and dynamic foreground material that are playing on the video compositing preview page and are located after the starting compositing frame to the server, so that the server can perform frame-by-frame compositing of the dynamic foreground material and dynamic background material to obtain a composite video.

[0183] In one embodiment, the selection module 1304 is further configured to select a dynamic background material corresponding to the background material selection operation from a local material library; or, download a dynamic background material corresponding to the background material selection operation from an online material library; or, determine the video obtained from the current shooting target environment as the dynamic background material.

[0184] In the above embodiments, by determining the starting composite frame of the dynamic foreground and dynamic background materials, and using the determined starting composite frame as the seven points for compositing, the dynamic background and dynamic foreground materials are composited. This allows users to select the parts of the dynamic background material they are interested in for compositing, resulting in a more visually appealing composite video. Furthermore, when the dynamic background material is a presentation, the presentation can be used as the background and composited with the dynamic foreground material. This creates a video where the dynamic foreground object interacts with the presentation, allowing the video recorder to composite the presentation and the presenter into the same video even in scenarios without a projector. Viewers can seamlessly watch the presenter and the presentation, enhancing the viewing experience.

[0185] In one embodiment, such as Figure 14 As shown, the device also includes: Extraction module 1310 is also used to extract frame image samples from material samples; The separation module 1314 is also used to perform channel separation on the extracted frame image samples as the current frame image samples to obtain the channel images corresponding to each color channel; The extraction module 1310 is also used to sequentially input the channel image of the current frame image sample and the image mask of the previous frame image sample of the current frame image sample into the untrained machine learning model for processing, so as to obtain the training prediction mask corresponding to each frame image sample. Calculation module 1318 is used to calculate the error value between each training prediction mask and the corresponding label; The adjustment module 1316 is also used to adjust the model parameters of the machine learning model according to the error value until the error value between the predicted image mask output by the adjusted machine learning model and the corresponding label is less than the error threshold, and then stop training.

[0186] In one embodiment, such as Figure 14 As shown, the device also includes: The label acquisition module 1320 is used to determine the training prediction mask as the image mask of the specified image sample in the extracted frame image sample when the training prediction mask corresponding to the current frame image is obtained and the error value between the training prediction mask and the corresponding label is less than a preset error value.

[0187] In one embodiment, the target video includes live video or conference video; such as Figure 14 As shown, the device also includes: The video acquisition module 1322 is used to acquire live video generated when live streaming is conducted through a live streaming room on a social application, or meeting video generated when electronic meetings are conducted in a meeting room on a social application. The sending module 1324 is used to send the composite video to the user equipment accessing the live broadcast room or the accessing conference room when a composite video corresponding to a live broadcast video or conference video is obtained.

[0188] In the above embodiments, the channel images of each color channel of the current frame image sample and the image mask of the previous frame image sample are input into the machine learning model. The image mask of the previous frame image sample is used as prior knowledge to predict the training prediction mask of the current frame image sample. The error value between the training prediction mask and the corresponding label is calculated. The model parameters of the machine learning model are adjusted according to the error value. As a result, the trained machine learning model can quickly segment the dynamic foreground material in the user-selected material, which improves the image segmentation speed and makes video synthesis faster, meeting the real-time requirements.

[0189] In one embodiment, the compositing module 1308 is further configured to identify key information in the dynamic background material; if the dynamic foreground material obscures the key information, the dynamic foreground material obscuring the key information is hidden or faded; and the dynamic foreground material after being hidden or faded is combined with the dynamic background material to form a video.

[0190] In the above embodiments, when important content (i.e. key information) is obscured, the dynamic foreground material that obscures the important content can be hidden or faded. Thus, when watching the composite video, the important content can still be viewed even when it is obscured by the dynamic foreground material, thereby improving the viewing effect of the composite video.

[0191] In one embodiment, the compositing module 1308 is further configured to identify the facial expressions of the target object in the dynamic foreground material; obtain the corresponding scene elements based on the facial expressions; and compose the scene elements, the dynamic background material played on the video compositing preview page, and the dynamic foreground material into a video.

[0192] In the above embodiments, by recognizing facial expressions, appropriate elements matching the recognized facial expressions are obtained, and these appropriate elements are combined with dynamic background materials and dynamic foreground materials to synthesize a video, making the video content more vivid and increasing the fun and appeal of the synthesized video.

[0193] In one embodiment, the synthesis module 1308 is further configured to extract speech from user-selected materials; perform speech recognition on the extracted speech to obtain corresponding text; obtain matching prompt information based on the semantic information of the text; and synthesize the prompt information, dynamic background material, and dynamic foreground material played on the video synthesis preview page into a video.

[0194] In the above embodiments, the corresponding text is obtained through speech recognition, and the matching prompt information is obtained based on the semantic information of the text. The prompt information is then combined with dynamic background material and dynamic foreground material to form a video, so that users can see the prompt information that matches the speech while hearing the speech, thus focusing the user's attention on the content displayed in the video.

[0195] Specific limitations regarding the video compositing apparatus can be found in the limitations of the video compositing method described above, and will not be repeated here. Each module in the aforementioned video compositing apparatus can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independent of the processor in a computer device, or stored in software in the memory of a computer device, so that the processor can call and execute the corresponding operations of each module.

[0196] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 15 As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a video synthesis method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad located on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0197] Those skilled in the art will understand that Figure 15 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0198] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0199] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0200] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above method embodiments.

[0201] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0202] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0203] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A video synthesis method, characterized in that, The method includes: The system acquires user-selected materials, which are either live streaming videos or ongoing electronic conference videos; these user-selected materials are used to provide dynamic foreground materials. Receive background material selection operation, select dynamic background material; the dynamic background material includes presentation; The dynamic background material and the dynamic foreground material are combined into a video; The synthesized video is sent to the live broadcast room or to the user device in the electronic conference, so that the live broadcast audience or conference participants can watch the live broadcast video or conference video with the changed video background.

2. The method according to claim 1, characterized in that, The step of combining the dynamic background material and the dynamic foreground material into a video includes: Align the presentation image with the dynamic foreground material frame by frame, wherein the presentation image is an image used as a background obtained by dividing the presentation page by page; While playing the demonstration images frame by frame, the dynamic foreground material is superimposed on the playing demonstration images to obtain a composite video.

3. The method according to claim 2, characterized in that, The step of aligning the demonstration image with the dynamic foreground material frame by frame includes: Align each presentation image in the presentation with at least one frame of dynamic foreground material, wherein the presentation images are aligned with the at least one frame of dynamic foreground material based on control operations on the presentation; The dynamic foreground material includes dynamic foreground objects, and the synthesized video includes a video of the dynamic foreground objects interacting and explaining with the presentation.

4. The method according to any one of claims 1 to 3, characterized in that, The step of combining the dynamic background material and the dynamic foreground material into a video includes: Play the dynamic background material on the video synthesis preview page, and overlay the dynamic foreground material in the editing state onto the dynamic background material for display. The editing state is used to indicate that the dynamic foreground material is in an editable state on the video synthesis preview page. Adjust the display parameters of the dynamic foreground material in the editing state; The dynamic background material and the dynamic foreground material, which are played on the video synthesis preview page and have had their display parameters adjusted, are combined to obtain a synthesized video.

5. The method according to claim 4, characterized in that, Adjusting the display parameters of the dynamic foreground material includes one or more of the following: Adjust the orientation and size of the dynamic foreground material; Duplicate the dynamic foreground material and adjust the relative positions between at least two of the duplicated dynamic foreground materials.

6. The method according to claim 4, characterized in that, The video synthesis preview page is a page provided in the social application; The process of obtaining user-selected materials includes: Acquire the live video generated when a live stream is conducted through a live streaming room on the social application, wherein the live streaming room is a live streaming room created in the social application; or; Acquire video of an electronic meeting held in a meeting room on the social application, wherein the electronic meeting is a meeting created in the social application.

7. The method according to any one of claims 1 to 3, characterized in that, The method further includes: When the target person in the dynamic foreground material obscures the content in the dynamic background material, identify the key information in the dynamic background material; When the dynamic foreground material obscures the key information, the target person in the dynamic foreground material that obscures the key information is hidden or faded. Once the obscured key information has finished playing, or once the obscured key information has moved to an unobscured area, the process of hiding or fading the target person in the dynamic foreground material will stop.

8. The method according to any one of claims 1 to 3, characterized in that, The step of combining the dynamic background material and the dynamic foreground material into a video includes: Speech recognition is performed on the speech in the user-selected materials to obtain the corresponding text; Based on the semantic information of the text, obtain matching prompts for marking key content in the presentation. The prompt message, the dynamic background material, and the dynamic foreground material are combined into a video.

9. The method according to any one of claims 1 to 3, characterized in that, The step of combining the dynamic background material and the dynamic foreground material into a video includes: Identify the facial expressions of the target object in the dynamic foreground material; Obtain the corresponding scene elements based on the facial expressions; The appropriate elements, the dynamic background material, and the dynamic foreground material played on the video synthesis preview page are combined into a video.

10. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The user-selected material is video decoded to obtain a series of frame images; Select a frame image containing the target object from the series of frame images; The target object in the selected frame image is segmented using a machine learning model to obtain the dynamic foreground material.

11. The method according to claim 10, characterized in that, The step of segmenting the target object in the selected frame image using a machine learning model to obtain the dynamic foreground material includes: If the extracted frame image is the first frame image, determine the image mask corresponding to the first frame image; If the extracted frame image is not the first frame image, the extracted frame images are used as the current frame image for channel separation to obtain the channel images corresponding to each color channel. The channel image of the current frame and the mask of the previous frame are input into the machine learning model for processing to obtain the prediction mask corresponding to each frame image. The target object is segmented from the corresponding frame image based on the image mask and the prediction mask.

12. A video synthesis apparatus, characterized in that, The device includes: The acquisition module is used to acquire user-selected materials, which are either live video or video of an ongoing electronic meeting; the user-selected materials are used to provide dynamic foreground materials. The selection module is used to receive background material selection operations and select dynamic background materials; the dynamic background materials include presentations. A compositing module is used to combine the dynamic background material and the dynamic foreground material into a video; The sending module is used to send the synthesized video to the live broadcast room or to the user equipment in the electronic conference, so that the live broadcast audience or conference participants can watch the live broadcast video or conference video after the video background has been changed.

13. A computer-readable medium, characterized in that, The medium stores a computer program that, when executed by a processor, implements the video synthesis method as described in any one of claims 1 to 11.

14. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the video compositing method as described in any one of claims 1 to 11.

15. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the video synthesis method as described in any one of claims 1 to 11.