Video synthesis method and device, computer device and storage medium

By selecting and synthesizing dynamic background and foreground materials in a video synthesis method, and using machine learning models for image segmentation and adjustment, the problem of low video viewership in existing technologies has been solved, thereby enriching video content and increasing viewership.

CN112822542BActive Publication Date: 2026-02-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010876955.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-27
Publication Date
2026-02-17
Estimated Expiration
2040-08-27

AI Technical Summary

Technical Problem

In existing technologies, users can only perform simple trimming and add music to the materials they want to upload to the network, which cannot effectively improve the video's viewership.

Method used

By entering the dynamic background material acquisition page, users can select dynamic background materials and overlay the user-selected dynamic foreground materials onto the background materials for display. Finally, the materials are combined into a video on the video synthesis preview page. Machine learning models are used for image segmentation and adjustment of display parameters to achieve customization of the video background.

Benefits of technology

It enriched the video content, improved the video viewing experience and playback rate, and increased the video viewing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112822542B_ABST
    Figure CN112822542B_ABST
Patent Text Reader

Abstract

The application relates to a video synthesis method and device, computer equipment and a storage medium. The method comprises the following steps: entering a dynamic background material acquisition page, selecting a dynamic background material in response to a background material selection operation triggered in the dynamic background material acquisition page; entering a video synthesis preview page, playing the dynamic background material, and superimposing a dynamic foreground material in an editing state on the played dynamic background material to display; and synthesizing the dynamic background material and the dynamic foreground material played in the video synthesis preview page into a video when a video synthesis operation is triggered in the video synthesis preview page. The method can fuse at least two materials of different types into a video of interest to a user, so as to improve the viewing rate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the computer technical field, in particular to a video synthesis method and device, computer equipment and storage medium. BACKGROUND

[0002] With the development of multimedia technology and Internet technology, users can download various favorite materials (such as videos) from a content distribution network for viewing, and upload user's own shooting materials to the network. For some own shooting materials, the user can edit the materials (such as cropping and splicing) and add corresponding music through a corresponding editor, and then upload to the network for other viewers to watch. In the above processing scheme, the user can only perform simple cropping and music addition on the material to be uploaded to the network, so as to effectively improve the viewing rate. SUMMARY

[0003] Therefore, it is necessary to provide a video synthesis method, device, computer equipment and storage medium, which can fuse at least two materials of different types into a video of interest to the user to improve the viewing rate.

[0004] A video synthesis method, the method comprising:

[0005] Entering a dynamic background material acquisition page, in response to a background material selection operation triggered in the dynamic background material acquisition page, selecting a dynamic background material;

[0006] Entering a video synthesis preview page, playing the dynamic background material, and superimposing a dynamic foreground material in an editing state on the played dynamic background material for display; wherein the dynamic foreground material is extracted from user-selected materials;

[0007] When a video synthesis operation is triggered in the video synthesis preview page, the dynamic background material and the dynamic foreground material played in the video synthesis preview page are synthesized into a video.

[0008] A video synthesis device, the device comprising:

[0009] An entering module for entering a dynamic background material acquisition page;

[0010] A selecting module for selecting a dynamic background material in response to a background material selection operation triggered in the dynamic background material acquisition page;

[0011] The entering module is further configured to enter a video synthesis preview page;

[0012] The playing module is configured to play the dynamic background material and superimpose the dynamic foreground material in an editing state on the played dynamic background material to display.

[0013] The synthesizing module is configured to synthesize the dynamic background material and the dynamic foreground material played on the video synthesis preview page into a video when a video synthesis operation is triggered on the video synthesis preview page.

[0014] In one embodiment, the user-selected material includes a target video or a dynamic image selected by a user, and the device further includes:

[0015] The extracting module is configured to extract frame images from the target video or the dynamic image and perform image segmentation on dynamic foreground material in the frame images by using a machine learning model.

[0016] In one embodiment, the device further includes:

[0017] The determining module is configured to determine an image mask corresponding to a first frame image if the extracted frame image is the first frame image.

[0018] The separating module is configured to sequentially perform channel separation on the extracted frame image as a current frame image to obtain a channel image corresponding to each color channel if the extracted frame image is a non-first frame image.

[0019] The extracting module is further configured to input the channel image of the current frame image and a mask of a previous frame image of the current frame image into the machine learning model to obtain a predicted mask corresponding to each frame image, and segment dynamic foreground material from the corresponding frame image according to the image mask and the predicted mask.

[0020] In one embodiment, the device further includes:

[0021] The adjusting module is configured to adjust a display parameter of the dynamic foreground material.

[0022] The synthesizing module is further configured to synthesize the dynamic background material and the dynamic foreground material played on the video synthesis preview page after the display parameter is adjusted to obtain a synthesized video.

[0023] In one embodiment, the adjusting module is further configured to adjust an orientation and a size of the dynamic foreground material.

[0024] The dynamic foreground material is copied, and a relative position between at least two dynamic foreground materials obtained by the copying is adjusted.

[0025] In an embodiment, the playing module is further configured to determine a starting composite frame of the dynamic foreground material and the dynamic background material respectively; play the dynamic background material starting from the starting composite frame of the dynamic background material, and superimpose the dynamic foreground material in an editing state on the played dynamic background material starting from the starting composite frame of the dynamic background material for display.

[0026] The synthesizing module is further configured to synthesize the dynamic background material and the dynamic foreground material played on the video synthesis preview page and located after the starting composite frame to obtain a synthesized video.

[0027] In an embodiment, the dynamic background material comprises a presentation; the determining module is further configured to split the presentation by page into presentation images used as backgrounds; and determine starting composite frames in the dynamic foreground material and the split presentation images respectively.

[0028] The synthesizing module is further configured to synthesize each presentation image and at least one frame of the dynamic foreground material in the presentation image and the dynamic foreground material played on the video synthesis preview page and located after the starting composite frame.

[0029] In an embodiment, the synthesizing module is further configured to synthesize the dynamic background material and the dynamic foreground material played on the video synthesis preview page and located after the starting composite frame by using a local video editor, or send the dynamic background material and the dynamic foreground material located after the starting composite frame to a server to make the server synthesize the dynamic foreground material and the dynamic background material frame by frame to obtain a synthesized video.

[0030] In an embodiment, the selecting module is further configured to select the dynamic background material corresponding to the background material selection operation from a local material library, or download the dynamic background material corresponding to the background material selection operation from an online material library, or determine a video obtained by currently shooting a target environment as the dynamic background material.

[0031] In an embodiment, the apparatus further comprises:

[0032] The extracting module is further configured to extract frame image samples from the material sample;

[0033] The separating module is further configured to separately separate the extracted frame image samples as current frame image samples by channel to obtain channel images corresponding to each color channel;

[0034] The extraction module is further configured to sequentially input a channel image of the current frame image sample and an image mask of a previous frame image sample of the current frame image sample into an untrained machine learning model for processing to obtain a training prediction mask corresponding to each frame image sample.

[0035] The calculation module is configured to calculate an error value between each training prediction mask and a corresponding label.

[0036] The adjustment module is further configured to adjust a model parameter of the machine learning model according to the error value until an error value between a prediction image mask output by the adjusted machine learning model and a corresponding label is less than an error threshold value, and stop training.

[0037] In an embodiment, the device further comprises:

[0038] The label acquisition module is configured to, when a training prediction mask corresponding to the current frame image is obtained and an error value between the training prediction mask and a corresponding label is less than a preset error value, determine the training prediction mask as an image mask of a specified image sample in the extracted frame image sample.

[0039] In an embodiment, the target video comprises a live video or a conference video; and the device further comprises:

[0040] The video acquisition module is configured to acquire a live video generated when a live room on a social application is used for live streaming or a conference video generated when a conference room on the social application is used for electronic conference.

[0041] The sending module is configured to, when a synthesized video corresponding to the live video or the conference video is obtained, send the synthesized video to a user equipment accessing the live room or the conference room.

[0042] In an embodiment, the synthesis module is further configured to identify a key information in the dynamic background material; if the dynamic foreground material shields the key information, perform a hiding or fading processing on the dynamic foreground material shielding the key information; and synthesize the dynamic foreground material after the hiding or fading processing and the dynamic background material into a video.

[0043] In an embodiment, the synthesis module is further configured to identify a facial expression of a target object in the dynamic foreground material; acquire a corresponding situational element according to the facial expression; and synthesize the situational element, the dynamic background material and the dynamic foreground material played on the video synthesis preview page into a video.

[0044] In one embodiment, the synthesizing module is further configured to extract speech from the user-selected material, perform speech recognition on the extracted speech to obtain corresponding text, acquire matching prompt information according to semantic information of the text, and synthesize the prompt information, the dynamic background material and the dynamic foreground material playing on the video synthesis preview page into a video.

[0045] A computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:

[0046] entering a dynamic background material acquisition page, and selecting dynamic background material in response to a background material selection operation triggered on the dynamic background material acquisition page;

[0047] entering a video synthesis preview page, playing the dynamic background material, and superimposing dynamic foreground material in an editing state on the played dynamic background material for display, wherein the dynamic foreground material is extracted from user-selected material;

[0048] when a video synthesis operation is triggered on the video synthesis preview page, synthesizing the dynamic background material and the dynamic foreground material playing on the video synthesis preview page into a video.

[0049] A computer readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the following steps:

[0050] entering a dynamic background material acquisition page, and selecting dynamic background material in response to a background material selection operation triggered on the dynamic background material acquisition page;

[0051] entering a video synthesis preview page, playing the dynamic background material, and superimposing dynamic foreground material in an editing state on the played dynamic background material for display, wherein the dynamic foreground material is extracted from user-selected material;

[0052] when a video synthesis operation is triggered on the video synthesis preview page, synthesizing the dynamic background material and the dynamic foreground material playing on the video synthesis preview page into a video.

[0053] The video synthesis method, device, computer device and storage medium, enter a dynamic background material acquisition page, in response to a background material selection operation triggered in the dynamic background material acquisition page, select a dynamic background material; enter a video synthesis preview page, play the dynamic background material, and superimpose a dynamic foreground material in an editing state on the played dynamic background material for display; when a video synthesis operation is triggered in the video synthesis preview page, synthesize the dynamic background material and the dynamic foreground material played in the video synthesis preview page, so that the user can customize the video background, replace the original background with a video background of interest, thereby enriching the video content, improving the video viewing effect, and helping to improve the video play rate and video viewing time. BRIEF DESCRIPTION OF DRAWINGS

[0054] Figure 1 An application environment diagram of the video synthesis method in one embodiment;

[0055] Figure 2 A flowchart of the video synthesis method in one embodiment;

[0056] Figure 3a A schematic diagram of a preview page for displaying user-selected materials in one embodiment;

[0057] Figure 3b A schematic diagram of selecting a dynamic background material in one embodiment;

[0058] Figure 4 A schematic diagram of segmenting a target object by a machine learning model in one embodiment;

[0059] Figure 5 A schematic diagram of segmenting a target object in one embodiment;

[0060] Figure 6 A flowchart of adjusting parameters of a dynamic foreground material and synthesizing in one embodiment;

[0061] Figure 7 A schematic diagram of synthesizing a dynamic foreground material and a presentation video frame in one embodiment;

[0062] Figure 8 A schematic diagram of synthesizing a dynamic foreground material and a presentation video frame in another embodiment;

[0063] Figure 9 A flowchart of synthesizing a dynamic foreground material and a presentation video in one embodiment;

[0064] Figure 10a A flowchart of a model training step in one embodiment;

[0065] Figure 10bFig. 2 is a schematic diagram of a video synthesis method in an embodiment, in which a dynamic foreground material is faded out;

[0066] Figure 10c Fig. 4 is a schematic diagram of a video synthesis method in an embodiment, in which a laughing face is added in the background when a target person in the dynamic foreground material starts to laugh;

[0067] Figure 10d Fig. 5 is a schematic diagram of a face feature point in an embodiment;

[0068] Figure 10e Fig. 6 is a schematic diagram of a video synthesis method in an embodiment, in which a prompt information is added in the background when a target person in the dynamic foreground material gives a prompt voice;

[0069] Figure 11 Fig. 7 is a schematic diagram of a video synthesis method in another embodiment;

[0070] Figure 12 Fig. 8 is a schematic diagram of a video synthesis method in another embodiment;

[0071] Figure 13 Fig. 9 is a structural block diagram of a video synthesis device in an embodiment;

[0072] Figure 14 Fig. 10 is a structural block diagram of a video synthesis device in another embodiment;

[0073] Figure 15 Fig. 11 is an internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0074] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0075] The video synthesis method provided by the present application can be applied to, for example, Figure 1The application environment is shown. In the application environment, a terminal 102, a server 104 and a terminal 106 are included. The terminal 102 can obtain or collect material (i.e., user-selected material, such as non-live video, dynamic image or live video) locally, or download corresponding user-selected material from the server 104 according to a user download operation, extract all dynamic foreground materials (such as images of a person or other target objects) from the user-selected material, and display a preview screen of the dynamic foreground materials. A dynamic background material acquisition page is entered, and a dynamic background material (such as a background video or a presentation) is selected from the dynamic background material acquisition page according to a triggered background material selection operation. A video synthesis preview page is entered, and the dynamic background material and the dynamic foreground material are played simultaneously, wherein the dynamic foreground material is located above an image layer of the dynamic background material. When a video synthesis operation is triggered in the video synthesis preview page, the dynamic background material and the dynamic foreground material played in the video synthesis preview page are synthesized into a video, and then the synthesized video is uploaded to the server 104 or sent to the terminal 106.

[0076] The terminal 102 and the terminal 106 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but are not limited thereto.

[0077] The server 104 can be a stand-alone physical server, or a server cluster composed of multiple physical servers, and can be a cloud server providing basic cloud computing services such as cloud server, cloud database, cloud storage and content delivery network (CDN).

[0078] The terminal 102 and the server 104 can be connected through a communication connection mode such as Bluetooth, USB (Universal Serial Bus) or network, which is not limited in the present application.

[0079] In one embodiment, as Figure 2 shown, a video synthesis method is provided, which can be executed by a terminal in Figure 1 , or cooperatively executed by a terminal and a server in Figure 1 . The method is applied to the terminal 102 in Figure 1 , for example, and includes the following steps:

[0080] S202, a preview screen of dynamic foreground materials extracted from user-selected materials is displayed.

[0081] The user-selected material can be a target video or a dynamic image. The target video can be any of the following: non-live video, live video, or the video of an ongoing electronic conference. The dynamic foreground material can be an image sequence of the target object extracted from the user-selected material. The target object can be a real person, a virtual person, or other animals or plants, etc. The preview screen can be any frame image from the dynamic foreground material, an image corresponding to the dynamic foreground material that contains the target object, or an instruction image used to guide video compositing, such as an instruction image showing how to perform video compositing.

[0082] In one embodiment, the video compositing method can be applied to a social application that is configured with a video editor or other video player with video editing capabilities, or to a live streaming room or meeting room created through the social application. Taking a social application with a video editor as an example, when performing video compositing, the terminal can call the social application's video editor, which has a video compositing preview page. Through this preview page, dynamic background and foreground materials can be played, edited, and composited. S202 may specifically include: when playing user-selected materials through the social application's video player, if a video compositing instruction is received, the terminal performs background removal on the user-selected materials to extract the corresponding dynamic foreground material, simultaneously calling the social application's video editor, and displaying a preview of the dynamic foreground material on the video editor's video compositing preview page. Furthermore, this video compositing method can also be applied to other applications, such as video applications.

[0083] The video player, live streaming room, and conference room mentioned above can all have the function of editing and compositing videos.

[0084] In one embodiment, the terminal can display not only previews of dynamic foreground materials but also, based on browsing instructions, various dynamic foreground materials from the user-selected materials, along with their backgrounds. Upon receiving a background deletion instruction, the background is removed from the user-selected materials, thus displaying only the corresponding dynamic foreground material. Furthermore, after deleting the background, a default background can be used to display the dynamic foreground material along with it, such as... Figure 3a As shown, the original background is replaced with a default white curtain, which is then displayed alongside the portrait image. In this preview, the portrait image is editable, as indicated by the dashed box; users can adjust the size of the portrait image by resizing the dashed box.

[0085] In an embodiment, if the user-selected material is a non-live type of video, the terminal can select a corresponding target video as the user-selected material from a local video library or an online video library; or the terminal can select a corresponding dynamic image as the user-selected material from a local image library or an online image library; or the terminal can collect a video of a target object in the environment through a built-in camera as the user-selected material. If the user-selected material is a presentation, the terminal can select a corresponding presentation from a file library, as shown in FIG. 8. Figure 3b

[0086] In another embodiment, if the user-selected material is a live video or a conference video, the step of obtaining the user-selected material includes: the terminal obtaining a live video generated when live streaming through a live streaming room on a social application, or a conference video generated when conducting an electronic conference in a conference room on a social application.

[0087] In an embodiment, after the terminal obtains the user-selected material, the terminal extracts dynamic foreground material corresponding to each frame from the user-selected material.

[0088] In an embodiment, when the user-selected material includes a user-selected target video or dynamic image, the step of extracting dynamic foreground material includes: the terminal extracting frame images from the target video or dynamic image; and performing image segmentation on the dynamic foreground material in the frame images through a machine learning model.

[0089] The machine learning model can be a convolutional neural network model or a neural network model that can be used for image segmentation. The frame image can refer to an image corresponding to each frame of the target video or dynamic image. For example, assuming that the total number of frames of the target video is n, the frame image can refer to the images of the 1st to nth frames of the target video. Here, n is a positive integer greater than 1.

[0090] In an embodiment, the terminal can perform video decoding on the target video to obtain a series of frame images, and then perform image segmentation on the dynamic foreground material in the series of frame images through the machine learning model. In addition, the terminal can also extract frame images containing the target object from the series of frame images, and then perform image segmentation on the dynamic foreground material in the extracted frame images through the machine learning model.

[0091] In another embodiment, the terminal can extract frame images corresponding to each frame from the dynamic image, and then perform image segmentation on the dynamic foreground material in the frame images through the machine learning model. In addition, the terminal can also select frame images containing the target object from the extracted frame images, and then perform image segmentation on the dynamic foreground material in the selected frame images through the machine learning model.

[0092] ​In one embodiment, if the extracted frame image is the first frame image, the image mask corresponding to the first frame image is determined; if the extracted frame image is a non-first frame image, the extracted frame image is sequentially taken as a current frame image for channel separation to obtain the channel images corresponding to each color channel. The step of image segmentation of the dynamic foreground material in the frame image by the machine learning model can specifically include: inputting the channel image of the current frame image and the mask of the previous frame image of the current frame image into the machine learning model for processing to obtain the predicted mask corresponding to each frame image; and segmenting the dynamic foreground material from the corresponding frame image according to the image mask and the predicted mask.

[0093] Before image segmentation by the machine learning model, the terminal sequentially takes the frame images of the non-first frame as the current frame image for channel separation to obtain the channel images corresponding to each color channel (such as the RGB channel).

[0094] For example, assuming that the total number of frames of the video is n (n is a positive integer greater than 1), if the image segmentation is currently performed on the first frame image, the terminal determines the image mask corresponding to the first frame image, which can represent the region of the target object; then, the image of the target object is segmented from the first frame image according to the image mask. If the image segmentation is currently performed on the i-th frame image (i is a positive integer greater than 1 and less than or equal to n), the channel separation is performed on the i-th frame image to obtain the R channel image, the G channel image and the B channel image of the i-th frame image, and then the R channel image, the G channel image and the B channel image of the i-th frame image are input into the machine learning model with the image mask corresponding to the i-1-th frame image as the prior knowledge, so that the R channel image, the G channel image and the B channel image of the i-th frame image are processed by the machine learning model to predict the image mask corresponding to the i-th frame image, as shown in FIG. 6. Figure 5 The image of the target object is segmented from the i-th frame image according to the image mask, so as to obtain the image sequence containing the target object from the first frame to the n-th frame. Figure 4 The image mask prediction process of the kitten is shown in FIG. 5, and the image mask of other target objects (such as a person) can also be predicted, and then the image of the target object is segmented according to the image mask, as shown in FIG. 6. Figure 5

[0095] In one embodiment, when the dynamic foreground material extracted from the user-selected material is extracted, one frame of the dynamic foreground material extracted can be taken as a preview picture for display, or the image corresponding to the target object in the dynamic foreground material can be taken as a preview picture for display.

[0096] For example, as shown in FIG. 7, the image of the target object in the dynamic foreground material is taken as a preview picture for display. Figure 4 ​or 5, an image about the target object is extracted from the frame image A of the video. When the image about the target object is extracted from each frame image of the video, a sequence of images about the target object is obtained.

[0097] S204, entering a dynamic background material obtaining page, and selecting the dynamic background material in response to a background material selection operation triggered in the dynamic background material obtaining page.

[0098] The dynamic foreground material is extracted from the user-selected material. The dynamic background material obtaining page can be a page for obtaining dynamic background material of a social application or other application, and can include one or more of the following: a page for obtaining dynamic background material of a local material library (i.e., a local dynamic background material obtaining page), a page for obtaining dynamic background material of an online material library (i.e., an online dynamic background material obtaining page), and a page for collecting a real environment to obtain dynamic background material (i.e., a dynamic background material real-time collection page). It should be noted that S202 and S204 can not be in a specific order, and the two steps can be independent of each other or related to each other.

[0099] The dynamic background material can refer to a background used as a dynamic foreground material, which can be a target video, a dynamic image, or a presentation (PPT).

[0100] In one embodiment, when the preview screen reaches a specified time or a dynamic background material selection instruction is received, the dynamic background material obtaining page is entered, such as a page for obtaining dynamic background material of a local material library of a social application.

[0101] In one embodiment, after entering the dynamic background material obtaining page, the terminal detects a background material selection operation triggered in the dynamic background material obtaining page in real time, and then selects a dynamic background material corresponding to the background material selection operation through the dynamic background material obtaining page. For example, after entering the page for obtaining dynamic background material of a local material library of a social application, a dynamic background material corresponding to the background material selection operation and saved in the local material library is obtained through the page.

[0102] In one embodiment, when a background material selection operation triggered in the dynamic background material obtaining page is detected, a dynamic background material corresponding to the background material selection operation is selected from a local material library; or a dynamic background material corresponding to the background material selection operation is downloaded from an online material library; or a video obtained by currently shooting a target environment is determined as the dynamic background material.

[0103] For example, the dynamic background material corresponding to the background material selection operation and saved in the local material library is obtained through a local dynamic background material obtaining page of the social application. Alternatively, the dynamic background material corresponding to the background material selection operation and saved in the online material library is obtained through an online dynamic background material obtaining page of the social application. Alternatively, the video obtained by the camera collecting the target environment is controlled through a dynamic background material real-time collection page of the social application, and the video is determined as the dynamic background material.

[0104] S206, entering a video synthesis preview page, playing the dynamic background material, and superimposing the dynamic foreground material in the editing state on the played dynamic background material for display.

[0105] The editing state can mean that the dynamic foreground material is in an editable state, for example, the size and position of the dynamic foreground material can be adjusted when in the editing state.

[0106] In an embodiment, after the dynamic background material is selected, the terminal automatically enters the video synthesis page. Alternatively, after the dynamic background material is selected, a confirmation operation triggered in the dynamic background material obtaining page is detected in real time, and when the confirmation operation is detected, the video synthesis preview page of the video editor is entered. The video editor can be a separate editor, or a video editing function or built-in video editing plug-in in a video player.

[0107] In an embodiment, before the dynamic background material is played, the terminal aligns the dynamic background material with the dynamic foreground material, and then plays the dynamic background material while superimposing the dynamic foreground material on the played dynamic background material for display, that is, playing the dynamic background material and the dynamic foreground material at the same time.

[0108] The alignment can mean that the first frame (or the first page) of the dynamic background material is aligned with the first frame of the dynamic foreground material, so that the dynamic background material and the dynamic foreground material are synthesized from the first frame when synthesized. In addition, the alignment can also mean that the first frame of the dynamic foreground material is aligned with the specified frame (or page) of the dynamic background material, so that the dynamic foreground material and the dynamic background material are synthesized from the first frame of the dynamic foreground material and the specified frame (or page) of the dynamic background material when synthesized.

[0109] S208, when a video synthesis operation is triggered in the video synthesis preview page, the dynamic background material and the dynamic foreground material played in the video synthesis preview page are synthesized into a video.

[0110] The video synthesis preview page can be provided with a video synthesis button, and the video synthesis process can be performed by triggering the video synthesis button. The video synthesis operation can be clicking or touching the video synthesis button on the video synthesis preview page.

[0111] In an embodiment, the starting synthesis frame of the dynamic background material and the dynamic foreground material is determined, and the dynamic background material and the dynamic foreground material played on the video synthesis preview page are synthesized into a video with the starting synthesis frame as the starting point of synthesis. If the starting synthesis frame of the dynamic background material and the dynamic foreground material are both the first frame, the dynamic background material and the dynamic foreground material are synthesized into a video starting from the first frame. If the starting synthesis frame of the dynamic foreground material is the first frame and the starting synthesis frame of the dynamic background material is a non-first frame (supposed to be the i-th frame), the synthesis is started with the i-th frame of the dynamic background material and the first frame of the dynamic foreground material.

[0112] For example, supposing that the dynamic background material has m frames in total and the dynamic foreground material has n frames in total, the synthesis is started with the i-th frame of the dynamic background material and the j-th frame of the dynamic foreground material, and a video containing the i-th frame to the m-th frame of the dynamic background material and the first frame to the n-th frame of the dynamic foreground material is obtained. Herein, m and n are both positive integers greater than 1, i and j are both positive integers greater than or equal to 1, and i is less than or equal to m and j is less than or equal to n.

[0113] In an embodiment, if the user-selected material is a live video or a conference video, when a synthesized video about the live video or the conference video is obtained, the terminal sends the synthesized video to the user equipment accessing the live room or the conference room, so that the live audience can watch the live video with a changed background or the conference participant can watch the conference video with a changed video background.

[0114] For example, when a host live streams a video through a live room on a social application, the host can cut out the host's portrait from the collected live video, and then replace the original background with a video material of interest to the host or the live audience as a new background to synthesize the cut-out host's portrait with the video material to obtain a synthesized video. Thus, the live video watched by the live audience is the video with a changed background, which can improve the interaction between the host and the live audience and the live audience's stay time in the live room.

[0115] For another example, for a live streaming scene of teaching or business promotion, if there is no drawing board to place teaching scripts or business description manuscripts during live streaming, at this time, the live audience can only watch the host and cannot watch the teaching scripts or business description manuscripts synchronously. By using the scheme of the embodiment, the host's portrait in the live video can be synthesized with the teaching scripts or business description manuscripts, so that a video with the teaching scripts or business description manuscripts as the background can be obtained, and the live audience can watch the live video as if the host is standing in front of a projector to explain the corresponding teaching scripts or business description manuscripts.

[0116] In the above embodiment, the dynamic background material acquisition page is entered, the dynamic background material is selected in response to the background material selection operation triggered in the dynamic background material acquisition page, the video synthesis preview page is entered, the dynamic background material is played, and the dynamic foreground material in the editing state is superimposed on the played dynamic background material for display; when the video synthesis operation is triggered in the video synthesis preview page, the dynamic background material and the dynamic foreground material played in the video synthesis preview page are synthesized, so that the user can customize the video background, replace the original background with a video background of interest, thereby enriching the video content, improving the video viewing effect, and facilitating the improvement of the video play rate and the video viewing time.

[0117] In one embodiment, as shown in Figure 6 the method can further include:

[0118] S602, playing the dynamic foreground material and the dynamic background material simultaneously in the video synthesis preview page.

[0119] The dynamic foreground material is played by being superimposed on a level where the dynamic background material is located.

[0120] S604, adjusting the display parameters of the dynamic foreground material.

[0121] During the playing of the dynamic foreground material and the dynamic background material, when it is necessary to adjust the form of the dynamic foreground material, the playing can be paused, and then the display parameters of the dynamic foreground material are adjusted.

[0122] In one embodiment, S604 can specifically include: adjusting the position and size of the dynamic foreground material by the terminal; and / or copying the dynamic foreground material and adjusting the relative positions between the at least two dynamic foreground materials obtained by copying. The position can refer to the direction and location.

[0123] For example, as shown in Figure 7 taking a person image of a host with dynamic background material and a PPT with dynamic background material as an example, the terminal can adjust the direction, position and size of the person image superimposed on the PPT level according to the input adjustment instruction, so that the host is in a suitable position to enable the live audience to watch the content of the PPT.

[0124] For example, as shown in Figure 8 taking a person image of a target person with dynamic background material and a sunrise video at the seaside with dynamic background material as an example, the terminal can copy multiple copies of the person image superimposed on the sunrise video at the seaside according to the input adjustment instruction, and adjust the relative positions and sizes of each copy of the person image, so that multiple target persons can be seen watching the sunrise at the seaside.

[0125] S606, when a video synthesis operation is triggered in the video synthesis preview page, the dynamic background material and the dynamic foreground material played in the video synthesis preview page and after the display parameter adjustment are synthesized to obtain a synthesized video.

[0126] The detailed synthesis process of S606 can refer to S208 of the above-mentioned embodiments.

[0127] In the above-mentioned embodiments, the user can adjust the orientation and size of the dynamic foreground material, and can also copy multiple dynamic foreground materials, so that the video has a better visual experience and is beneficial to improve the video watching effect and the video watching rate.

[0128] In one embodiment, as shown in FIG. 6, S206 can specifically include the following steps. Figure 9

[0129] S902, respectively determining the starting synthesis frame of the dynamic foreground material and the dynamic background material.

[0130] In one embodiment, when the dynamic background material is a presentation, S902 can specifically include: the terminal cuts the presentation into presentation images used as backgrounds by pages; and respectively determining the starting synthesis frame in the dynamic foreground material and the cut presentation images.

[0131] S904, playing the dynamic background material from the starting synthesis frame of the dynamic background material as a starting position, and superimposing the dynamic foreground material in the editing state on the played dynamic background material for display.

[0132] In one embodiment, when the dynamic background material is a presentation, the terminal aligns the presentation image with the dynamic foreground material frame by frame before playing the presentation image frame by frame, and then plays the presentation image frame by frame while superimposing the dynamic foreground material on the played presentation image for display, that is, playing the presentation image and the dynamic foreground material simultaneously.

[0133] S906, synthesizing the dynamic background material and the dynamic foreground material played in the video synthesis preview page and after the starting synthesis frame to obtain a synthesized video.

[0134] ​In one embodiment, when the dynamic background material is a presentation, S906 can specifically include that the terminal synthesizes each presentation image with corresponding at least one frame of dynamic foreground material, among the presentation images played on the video synthesis preview page and after the starting synthesis frame. In one embodiment, S906 can specifically include that the dynamic background material and the dynamic foreground material played on the video synthesis preview page and after the starting synthesis frame are synthesized by a local video editor, or the dynamic background material and the dynamic foreground material after the starting synthesis frame are sent to a server to make the server synthesize the dynamic foreground material and the dynamic background material frame by frame to obtain a synthesized video.

[0135] In one embodiment, when the dynamic background material is a presentation, the starting synthesis frame of the presentation image and the dynamic foreground material is determined, and the presentation image and the dynamic foreground material played on the video synthesis preview page are synthesized as a video starting from the starting synthesis frame. If the starting synthesis frame of the presentation image and the dynamic foreground material is the first frame, the presentation image and the dynamic foreground material are synthesized starting from the first frame. If the starting synthesis frame of the dynamic foreground material is the first frame and the starting synthesis frame of the presentation image is a non-first frame (assuming the i-th frame), the synthesis is started from the i-th frame of the presentation image and the first frame of the dynamic foreground material.

[0136] For example, assuming that the presentation image has m frames (i.e., m presentation images) and the dynamic foreground material has n frames, the synthesis is started from the i-th frame of the dynamic background material and the j-th frame of the dynamic foreground material to obtain a video containing the i-th frame to the m-th frame of the presentation image and the first frame to the n-th frame of the dynamic foreground material. Wherein, m and n are positive integers greater than 1, i and j are positive integers greater than or equal to 1, and i is less than or equal to m and j is less than or equal to n.

[0137] In one embodiment, if the user-selected material is a live video or a conference video, when the synthesized video of the live video or the conference video is obtained, the terminal sends the synthesized video to the user equipment accessing the live room or the conference room, so that the live audience can watch the live video with the background replaced by the presentation, or the conference participants can watch the conference video with the background replaced by the presentation.

[0138] For example, for a live scene of teaching or business promotion, if there is no drawing board to place the teaching script or business description manuscript during the live broadcast, at this time, the live audience can only watch the host during the live broadcast and cannot synchronously watch the teaching script or business description manuscript. The scheme of the embodiment can synthesize the host portrait in the live video with the teaching script or business description manuscript, so that a video with the teaching script or business description manuscript as the background can be obtained, and the live audience can watch the live video as if the host is standing in front of the projector and explaining the corresponding teaching script or business description manuscript.

[0139] In the above embodiment, the starting synthesis frame of the dynamic foreground material and the dynamic background material is determined, the dynamic background material and the dynamic foreground material are synthesized with the determined starting synthesis frame as the synthesis seven points, so that the user can select the part of interest in the dynamic background material for synthesis according to personal preference, so that the synthesized video has a better visual effect. In addition, when the dynamic background material is a presentation, the presentation can be synthesized with the dynamic foreground material as the background, so that a video of the dynamic foreground object interacting with the presentation can be obtained, and the video recording party can also synthesize the presentation and the anchor into the same video in a scene without a projector to project the presentation, so that the audience can watch the anchor and the presentation without any sense of strangeness, and the video viewing effect is improved.

[0140] In one embodiment, as shown in FIG. 1, the training step of the machine learning model includes: Figure 10a

[0141] S1002, extracting frame image samples from the material sample.

[0142] The material sample can be a target video or a dynamic image. The target video can be any one of a non-live type video, a live video, and a conference video corresponding to an ongoing electronic conference. The frame image sample can refer to the image of each frame in the material sample.

[0143] In one embodiment, after extracting the frame image sample from the material sample, the terminal also acquires the label corresponding to each frame image sample.

[0144] S1004, respectively performing channel separation on the extracted frame image samples as the current frame image sample to obtain the channel images corresponding to each color channel.

[0145] In one embodiment, the terminal takes the frame image sample currently being synthesized as the current frame image sample, and then performs channel separation on the current frame image sample to obtain the channel images corresponding to the RGB channels.

[0146] For example, taking a video as the material sample, assuming that the total number of frames of the video is n (n is a positive integer greater than 1), and the image segmentation is performed on the i-th frame (i is a positive integer greater than 1 and less than or equal to n) image, the channel separation is performed on the i-th frame image sample to obtain the R channel image, the G channel image and the B channel image of the i-th frame image sample.

[0147] S1006, sequentially inputting the channel images of the current frame image sample and the image mask of the previous frame image sample of the current frame image sample into the untrained machine learning model for processing to obtain the training prediction mask corresponding to each frame image sample.

[0148] ​In one embodiment, if the current frame image sample is the first frame, the terminal can directly calculate the image mask of the current frame image sample, and take the image mask of the first frame as the image mask of the specified image sample in the frame image sample. In addition, the terminal can replace the specified image sample in the extracted frame image sample with the current frame image sample, so as to use the replaced specified image sample for subsequent training.

[0149] In one embodiment, if the current frame image sample is not the first frame, the channel image of the current frame image sample and the image mask of the previous frame image sample of the current frame image sample are input into the untrained machine learning model for processing.

[0150] For example, when the material sample is a video, the terminal inputs the R channel image, the G channel image and the B channel image of the i-th frame image sample and the image mask corresponding to the i-1-th frame image sample into the machine learning model. The machine learning model processes the R channel image, the G channel image and the B channel image of the i-th frame image sample based on the image mask corresponding to the i-1-th frame image sample as prior knowledge, so as to predict the training prediction mask corresponding to the i-th frame image sample.

[0151] S1008, calculating the error value between each training prediction mask and the corresponding label.

[0152] In one embodiment, the terminal calculates the error value between each training prediction mask and the corresponding label according to a loss function. The loss function can be any one of the following: Mean Squared Error, cross-entropy loss function, L2Loss function and Focal Loss function.

[0153] In one embodiment, in order to enable the machine learning model to learn that a certain target object image appears in a certain frame image sample and affects the prediction of the image mask of the frame image sample, the terminal can obtain the label in the following manner: when the training prediction mask corresponding to the current frame image is obtained, and the error value between the training prediction mask and the corresponding label is less than a preset error value, the terminal determines the training prediction mask as the label corresponding to the specified image sample in the extracted frame image sample. Thus, during the training of the machine learning model, when the specified image sample is taken as the current frame image sample for synthesis, the training prediction mask of the specified image sample is predicted by the machine learning model, the training prediction mask and the image mask of the specified image sample are compared, the loss value between the two is obtained, and then S1010 is executed.

[0154] In another embodiment, when the image mask of the first frame is calculated, the terminal can also take the image mask of the first frame as the image mask of the specified image sample in the extracted frame image sample, so that in the process of training the machine learning model, when the specified image sample is synthesized as the current frame image sample, the training prediction mask of the specified image sample is predicted by the machine learning model, the training prediction mask and the image mask of the specified image sample are compared, the loss value between the two is obtained, and then S1010 is executed, so that when a target object suddenly appears in front of the camera, the influence of the appearance of the target object on the prediction of the image mask of the subsequent frame image sample can also be effectively solved.

[0155] S1010, adjusting the model parameters of the machine learning model according to the error value.

[0156] In one embodiment, when all frame image samples are input to the machine learning model training, and the error value between the prediction image mask output by the adjusted machine learning model and the corresponding label is less than the error threshold, the training is stopped.

[0157] In one embodiment, the terminal propagates the loss value back to each layer of the machine learning model to obtain the gradient of each layer model parameter; and adjusts the model parameters of each layer in the machine learning model according to the gradient.

[0158] In the above embodiment, the channel image of each color channel of the current frame image sample is input into the machine learning model together with the image mask of the previous frame image sample, the image mask of the previous frame image sample is used as prior knowledge to predict the training prediction mask of the current frame image sample, the error value between the training prediction mask and the corresponding label is calculated, and the model parameters of the machine learning model are adjusted according to the error value, so that the trained machine learning model can quickly segment the dynamic foreground material in the user-selected material, improve the image segmentation rate, and make the video synthesis faster, meeting the real-time requirement.

[0159] In one embodiment, S210 can further include: the terminal identifies key information in the dynamic background material; if the dynamic foreground material blocks the key information, the dynamic foreground material blocking the key information is hidden or faded; and the dynamic foreground material after the hiding or fading processing and the dynamic background material are synthesized into a video.

[0160] The key information can be an important image, text or animation in the dynamic background material. The hiding processing can refer to cutting out the target object in the dynamic foreground material that hides the key information, or replacing the dynamic foreground material that hides the key information with a blank material. The fading processing can refer to adjusting the transparency of the dynamic foreground material that hides the key information, and in addition, the color can also be adjusted (for example, dark color is adjusted to light color) while adjusting the transparency. Thus, after the hiding and fading processing, the user can watch the hidden key information. It should be noted that during the hiding or fading processing, the audio corresponding to the dynamic foreground material is retained.

[0161] For example, as shown in FIG. 6, when the target person in the dynamic foreground material hides the content in the PPT, the transparency of the target person can be adjusted. Figure 10b

[0162] In an embodiment, if the hidden key information is played or the hidden key information moves to a non-hidden area (for example, when the PPT changes a page, the position of the key information is adjusted), the hiding or fading processing of the dynamic foreground material will be stopped. Therefore, when the originally hidden key information disappears or moves to a non-hidden area, the dynamic foreground material is displayed.

[0163] In the above embodiment, when important content (i.e. key information) is hidden, the dynamic foreground material that hides the important content can be hidden or faded, so that when watching the synthesized video, the important content can still be watched even if it is hidden by the dynamic foreground material, thereby improving the viewing effect of the synthesized video.

[0164] In an embodiment, S210 can further include: identifying a facial expression of a target object in the dynamic foreground material by the terminal; obtaining a corresponding situational element according to the facial expression; and synthesizing the situational element, the dynamic background material and the dynamic foreground material played in the video synthesis preview page into a video.

[0165] The situational element can refer to a pattern, a virtual expression or a text matched with the facial expression, such as a laughing face or a sunny pattern. The face can refer to a human face, a chin, a lip, an eye, a nose, a brow, a forehead and an ear, etc. Correspondingly, the facial expression can be composed of various postures of the face, chin, lip, eye, nose, brow, forehead and ear, etc., such as pouting, winking and smiling, etc.

[0166] For example, as shown in FIG. 7, when the target person in the dynamic foreground material hides the content in the PPT, the target person can be adjusted to a virtual expression matched with the facial expression of the target person. Figure 10c ​As shown, when the target person in the dynamic foreground material is laughing happily, a smiley face (or a clear sky pattern, not shown in the image) can be captured. Then, during video compositing, the captured smiley face (or clear sky pattern) is composited into the corresponding image position. Thus, when the composite video is played, you can see the target person laughing happily while also seeing a smiley face (or clear sky pattern) in the background. Alternatively, when the target person in the dynamic foreground material has a gloomy expression, a dynamic dark cloud can be composited with the dynamic foreground material and the dynamic background material into the video. Thus, when the composite video is played, the dynamic dark cloud will be displayed on the video background.

[0167] In one embodiment, the step of a terminal recognizing the facial expression of a target object in dynamic foreground material may specifically include: the terminal extracting eye feature points from the dynamic foreground material; determining a first distance between an upper eyelid feature point and a lower eyelid feature point, and determining a second distance between a left corner eye feature point and a right corner eye feature point; and determining the eye pose based on the relationship between the ratio of the first distance and the second distance and at least one preset interval.

[0168] The left and right corner feature points refer to the left and right corner feature points of the same eye, respectively. For example, for the left eye, the left corner feature point refers to the left corner feature point of the left eye, and the right corner feature point refers to the right corner feature point of the left eye.

[0169] For example, such as Figure 10d As shown, the terminal calculates the first distance between the upper eyelid feature point 38 and the lower eyelid feature point 42, and the second distance between the left corner feature point 37 and the right corner feature point 40, based on the facial feature points obtained by facial recognition technology. When the ratio between the first distance and the second distance is 0, the test subject is determined to be squinting with the left eye. When the ratio between the first distance and the second distance is less than 0.2, the test subject is determined to be blinking with the left eye. When the ratio between the first distance and the second distance is greater than 0.2 and less than 0.6, the test subject is determined to be staring with the left eye.

[0170] The following methods can be used to identify lip posture:

[0171] Method 1: Identify lip posture based on the height of lip feature points.

[0172] In one embodiment, the facial expression recognition step further includes: the terminal extracting lip feature points from dynamic foreground material; and determining the lip posture based on the height difference between the lip center feature point and the lip corner feature point among the lip feature points.

[0173] For example, such as Figure 10dAs shown, the terminal determines the height of the lip center feature point 63 and the height of the lip corner feature point 49 (or 55), and then calculates the height difference between the lip center feature point 63 and the lip corner feature point 49 (or 55). If the height difference is positive (i.e. the height of the lip center feature point 63 is higher than the height of the lip corner feature point 49 or 55), it is determined that the lip posture is smiling.

[0174] In a second mode, the lip posture is identified according to the distance between the upper and lower lip feature points.

[0175] In one embodiment, among the lip feature points, the terminal determines the lip posture according to a third distance between the upper lip feature point and the lower lip feature point.

[0176] For example, the third distance between the upper lip feature point and the lower lip feature point is compared with a distance threshold value, and when the third distance reaches the distance threshold value, it is determined that the lip posture is open-mouthed. For example, as shown in FIG. 6, the third distance between the upper lip feature point 63 and the lower lip feature point 67 is compared with a distance threshold value, and if it is greater than or equal to the distance threshold value, it is identified that the subject under test is open-mouthed. Figure 10d

[0177] In another embodiment, the terminal calculates a fourth distance between the left lip corner feature point and the right lip corner feature point, and when the third distance is greater than or equal to the fourth distance, it is determined that the lip posture is open-mouthed.

[0178] In a third mode, the lip posture is identified according to the ratio of the distance between the upper and lower lips to the distance between the left and right lip corner feature points.

[0179] In one embodiment, among the lip feature points, the lip posture is determined according to the relationship between the ratio of the third distance to the fourth distance and at least one preset interval; wherein the fourth distance is the distance between the left lip corner feature point and the right lip corner feature point.

[0180] For example, as shown in FIG. 6, the terminal determines the third distance between the upper lip feature point 63 and the lower lip feature point 67, and determines the fourth distance between the left lip corner feature point 49 and the right lip corner feature point 55, and the relationship between the ratio of the third distance to the fourth distance and at least one preset interval determines the lip posture. For example, when the ratio of the third distance to the fourth distance is in a first preset interval, it is determined that the lip posture is open-mouthed; in addition, when the ratio of the third distance to the fourth distance is in a second preset interval, it is determined that the lip posture is closed-mouthed. The values of the first preset interval are all greater than the values of the second preset interval. Figure 10d

[0181] ​​In one embodiment, the step of facial expression recognition further comprises: extracting eyebrow feature points and eyelid feature points from the dynamic foreground material; determining a fifth distance between the eyebrow feature points and the eyelid feature points; and determining the eyebrow posture according to the size relationship between the fifth distance and a preset distance.

[0182] For example, as shown in FIG. 38, the distance between the eyebrow feature points and the eyelid feature points 38 is calculated, and when the distance is greater than a preset distance, it is determined that the subject under test is raising eyebrows. Figure 10d

[0183] In the above embodiments, by recognizing facial expressions, a situational element corresponding to the recognized facial expression is obtained, and the situational element is synthesized into a video together with the dynamic background material and the dynamic foreground material, so that the video content is more vivid and interesting, and the synthesized video is more interesting and infectious.

[0184] In one embodiment, S210 can further specifically comprise: extracting speech from the user-selected material by the terminal; performing speech recognition on the extracted speech to obtain corresponding text; obtaining matching prompt information according to the semantic information of the text; and synthesizing the prompt information, the dynamic background material and the dynamic foreground material played on the video synthesis preview page into a video.

[0185] For example, as shown in FIG. 39, the target person in the dynamic foreground material says "it is very interesting here", at this time, the prompt information corresponding to the speech is obtained, such as "students, please pay attention to the PPT", and then the prompt information is synthesized into a video together with the dynamic background material and the dynamic foreground material, so that when the synthesized video is played, "students, please pay attention to the PPT" is dynamically displayed on the background picture. Figure 10e

[0186] In one embodiment, when the terminal obtains the speech, the terminal extracts acoustic features from the speech, and obtains the recognized text according to the acoustic features. The acoustic features can include but are not limited to PNCC (power-normalized cepstral coefficients), MFCC (Mel Frequency Cepstrum Coefficient).

[0187] In one embodiment, the step of extracting acoustic features from the speech can specifically comprise: the terminal performs frame division on the speech, determines the power spectrum according to the spectrum of each frame of speech; obtains the logarithmic power spectrum corresponding to the power spectrum; determines the logarithmic power spectrum as the speech feature, or determines the result obtained by discrete cosine transform on the logarithmic power spectrum as the speech feature.

[0188] ​​For example, assuming that the signal expression of the collected voice is x(n), the voice after framing and windowing is x'(n)=x(n)×h(n), the discrete Fourier transform is performed on the windowed voice x'(n)=x(n)×h(n), and the corresponding spectrum signal is obtained as:

[0189]

[0190] wherein N represents the number of points of the discrete Fourier transform.

[0191] When the spectrum of each frame of voice is obtained, the terminal calculates the corresponding power spectrum, obtains the logarithmic value of the power spectrum to obtain the logarithmic power spectrum, inputs the logarithmic power spectrum into a mel-scale triangular filter, and obtains mel-frequency cepstral coefficients after discrete cosine transform. The obtained mel-frequency cepstral coefficients are:

[0192]

[0193] The above logarithmic energy is input into the discrete cosine transform, and the L-order mel-frequency cepstral parameters are obtained. L order refers to the order of the mel-frequency cepstral coefficients, which can be 12-16. M refers to the number of triangular filters.

[0194] In one embodiment, before extracting the acoustic features, the terminal can perform voice enhancement (such as noise reduction) processing on the extracted voice, and then perform acoustic feature extraction on the voice after the voice enhancement processing.

[0195] In the above embodiment, the corresponding text is obtained through voice recognition, the matching prompt information is obtained according to the semantic information of the text, the prompt information is synthesized into a video together with the dynamic background material and the dynamic foreground material, so that the user can also watch the prompt information matched with the voice while hearing the voice, and the user's attention is focused on the display content of the video.

[0196] As an example, when the user-selected material is a target video (including a live type video or a non-live type video), the background of the target video is replaced through video synthesis, such as shown in Figure 11 The method of video synthesis is as follows:

[0197] S1102, the portrait is cut out from each video frame of the target video.

[0198] S1104, the background in the target video is replaced with a background curtain, so as to remove the original background in the target video.

[0199] S1106, the corresponding video material or PPT material is selected according to the user's demand.

[0200] For example, the video material or PPT material of interest can be selected from a local video library or file library, in addition to being selected from a video template provided by the system, or downloaded from an online video library or file library.

[0201] Taking the use scenario of enterprise roadshow as an example, it can be necessary to explain or interact with the video material or PPT material in the process of enterprise roadshow, like a weather forecast host explaining to the dynamic video material or PPT material. Using the scheme of the present application, the portrait in the roadshow video can be cut out in real time, and then the background of the roadshow video can be replaced with other video material or PPT material. In addition, the cut-out portrait can be placed on the new background (i.e. the replaced video material or PPT material) by flipping, zooming in and out, and tilting at will.

[0202] S1108, the cut-out portrait is synthesized with the selected video material or PPT material to obtain a synthesized video with a replaced background.

[0203] In an embodiment, during the synthesis of the video, the cut-out portrait can be adjusted in position, size, and angle, and then the adjusted portrait can be synthesized with the selected video material or PPT material.

[0204] In an embodiment, during the synthesis of the video, multiple portraits can also be copied to achieve a split personality. Then the copied multiple portraits can be synthesized with the selected video material or PPT material.

[0205] As another example, as shown in Figure 12 The method of video synthesis is as follows:

[0206] S1202, the portrait of each video frame in the target video is cut out by a neural network model.

[0207] For example, a convolutional neural network model using machine learning is used to solve the semantic segmentation task, so as to segment the portrait of each video frame in the target video. In particular, by meeting the following requirements and constraints, a network architecture and training process suitable for mobile terminals (such as mobile phones) are designed.

[0208] 1) Construct a training data set

[0209] Frame images with a wide range of foreground poses (such as the poses of characters) and background environments are obtained as training samples, and these frame images are labeled to provide high-quality data for the machine learning process. These labels include pixel-level accurate positioning of foreground elements, such as hair, glasses, neck, skin, and lips; and background labels generally achieve cross-validation results of human labeling quality.

[0210] 2) Data Input

[0211] The segmentation task can refer to calculating binary image masks for frame images of the target video to segment foreground elements from the background environment, and then using the image mask of the previous frame as prior knowledge and the three-channel image of the next frame as input into the neural network model for training.

[0212] Specifically, the current frame image is separated into channels to obtain images of each color channel, namely, the R channel image, G channel image, and B channel image. Then, the R channel image, G channel image, and B channel image of the current frame image, along with the image mask of the previous frame image, are input into a neural network model for training to predict the image mask of the current frame image.

[0213] 3) Training process

[0214] During segmentation, it's necessary to maintain temporal continuity between frames while also considering temporal discontinuities, such as a person suddenly appearing in front of the camera. To robustly train the model and address these issues, the ground truth labels for each frame can be transformed in several ways and used as a mask for the previous frame:

[0215] Clear the preceding image mask: Given that the neural network model has correctly processed the first frame image and the new target in the scene, obtain the image mask for that first frame image, and then use that image mask as the label for subsequent frames. This will simulate a scene where someone appears in the camera's viewfinder.

[0216] Affine transformation of the ground truth mask: The Minor transformation is performed on the neural network model to propagate and adjust the image mask of the previous frame; while the Major transformation discards the image mask that the neural network model deems unsuitable.

[0217] The converted image: Thin plate splines smoothing of the original image was implemented to speed up camera movement and rotation.

[0218] S1204, the client will upload the human image extracted from the target video to the server.

[0219] S1206, Obtain video or PPT materials.

[0220] S1208: The client uploads the acquired video or PPT materials to the server.

[0221] Users select video or PPT materials from the client and upload them to the server. After receiving the uploaded video or PPT materials, the server will overlay the video or PPT materials with the cut-out human image.

[0222] If the video material is overlaid, the cut-out portrait and the video material will be played at the same time according to time. If the user selects to adjust or edit the start time line of the video material on the client, the portrait and the video material will be overlaid from the selected start time line.

[0223] If the PPT material is overlaid, the PPT is imported first, and each page of the PPT will be split into pictures or the PPT will be split into image frames if the PPT has dynamic effects. The overlaying process can overlay one picture or one frame of the PPT corresponding to a segment of video of the human action.

[0224] The basic process of real-time overlaying or splicing is as follows:

[0225] 1) Calculate all the time parameters, image parameters, and fusion parameters in the video.

[0226] 2) Use the parameters obtained in the first step to complete real-time overlaying or splicing.

[0227] S1210, the client adjusts the size, orientation, and other parameters of the portrait in the target video.

[0228] After the portrait and the corresponding video material are overlaid and displayed, the portrait can be reduced, rotated, copied, and the like. The user selects the portrait on the client, and a selection box appears on the client. Then the user operates the reduction, rotation, and copying, and the client processes these operation information locally.

[0229] S1212, the client sends the adjusted parameters.

[0230] S1214, the server synthesizes the adjusted portrait and the video material or the PPT material.

[0231] S1216, the server returns the obtained synthesized video to the client.

[0232] After the user saves and confirms, the client requests the server, the server processes the video synthesis according to the processing operation on the client, and then returns the obtained synthesized video to the client for display.

[0233] Through the above-mentioned embodiments, the portrait can be cut out from the target video in real time, and the background of the target video can be replaced with other video or dynamic PPT. For the use scene of enterprise roadshow, the user can explain to the dynamic video or PPT material, and interact with the video or PPT material content.

[0234] It should be understood that, although Figure 2 , 6The steps in the flowcharts of 9-12 are displayed in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, Figure 2 , 6 At least some of the steps in 9-12 can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.

[0235] In one embodiment, as shown in Figure 13 , a video synthesis device is provided, which can be a part of a computer device in the form of a software module or a hardware module, or a combination of the two. The device specifically includes an entering module 1302, a selection module 1304, a playing module 1306, and a synthesis module 1308, wherein:

[0236] The entering module 1302 is configured to enter a dynamic background material acquisition page.

[0237] The selection module 1304 is configured to select a dynamic background material in response to a background material selection operation triggered on the dynamic background material acquisition page.

[0238] The entering module 1302 is further configured to enter a video synthesis preview page.

[0239] The playing module 1306 is configured to play the dynamic background material and superimpose a dynamic foreground material in an editing state on the played dynamic background material for display, wherein the dynamic foreground material is extracted from user-selected materials.

[0240] The synthesis module 1308 is configured to synthesize the dynamic background material and the dynamic foreground material played on the video synthesis preview page into a video when a video synthesis operation is triggered on the video synthesis preview page.

[0241] In the above embodiment, the dynamic background material acquisition page is entered, the dynamic background material is selected in response to the background material selection operation triggered in the dynamic background material acquisition page, the video synthesis preview page is entered, the dynamic background material is played, and the dynamic foreground material in the editing state is superimposed on the played dynamic background material for display; when the video synthesis operation is triggered in the video synthesis preview page, the dynamic background material and the dynamic foreground material played in the video synthesis preview page are synthesized, so that the user can customize the video background, replace the original background with a video background of interest, thereby enriching the video content, improving the video viewing effect, and facilitating the improvement of the video play rate and the video viewing time.

[0242] In one embodiment, the user-selected material includes a target video or a dynamic image selected by the user; as shown in Figure 14 The apparatus further includes an extraction module 1310; wherein:

[0243] The extraction module 1310 is configured to extract frame images from the target video or the dynamic image; and perform image segmentation on the dynamic foreground material in the frame images by using a machine learning model.

[0244] In one embodiment, as shown in Figure 14 The apparatus further includes:

[0245] The determination module 1312 is configured to determine an image mask corresponding to the first frame image if the extracted frame image is the first frame image.

[0246] The separation module 1314 is configured to sequentially perform channel separation on the extracted frame images as current frame images if the extracted frame image is a non-first frame image, to obtain channel images corresponding to each color channel.

[0247] The extraction module 1310 is further configured to input the channel image of the current frame image and the mask of the previous frame image of the current frame image into the machine learning model for processing, to obtain a predicted mask corresponding to each frame image; and segment the dynamic foreground material from the corresponding frame image according to the image mask and the predicted mask.

[0248] In one embodiment, as shown in Figure 14 The apparatus further includes:

[0249] The adjustment module 1316 is configured to adjust the display parameters of the dynamic foreground material.

[0250] The synthesis module 1308 is further configured to synthesize the dynamic background material and the dynamic foreground material played in the video synthesis preview page after the display parameters are adjusted, to obtain a synthesized video.

[0251] In an embodiment, the adjusting module 1316 is further configured to adjust the position and size of the dynamic foreground material; and copy the dynamic foreground material and adjust the relative positions between the at least two copied dynamic foreground materials.

[0252] In an embodiment, the playing module 1306 is further configured to determine the starting synthesis frame of the dynamic foreground material and the dynamic background material respectively; play the dynamic background material starting from the starting synthesis frame of the dynamic background material, and superimpose the dynamic foreground material in the editing state on the played dynamic background material to display, starting from the starting synthesis frame of the dynamic background material.

[0253] The synthesizing module 1308 is further configured to synthesize the dynamic background material and the dynamic foreground material played on the video synthesis preview page and located after the starting synthesis frame to obtain a synthesized video.

[0254] In the above embodiments, the user can adjust the position and size of the dynamic foreground material, and can copy multiple dynamic foreground materials, so that the video has a better visual experience, and the video viewing effect and the video viewing rate are improved.

[0255] In an embodiment, the dynamic background material includes a presentation; the determining module 1312 is further configured to split the presentation into presentation images used as backgrounds page by page; and determine the starting synthesis frame in the dynamic foreground material and the split presentation images respectively.

[0256] The synthesizing module 1308 is further configured to synthesize each presentation image and at least one frame of the dynamic foreground material in the presentation image and the dynamic foreground material played on the video synthesis preview page and located after the starting synthesis frame.

[0257] In an embodiment, the synthesizing module 1308 is further configured to synthesize the dynamic background material and the dynamic foreground material played on the video synthesis preview page and located after the starting synthesis frame by using a local video editor; or send the dynamic background material and the dynamic foreground material located after the starting synthesis frame to a server, so that the server synthesizes the dynamic foreground material and the dynamic background material frame by frame to obtain a synthesized video.

[0258] In an embodiment, the selecting module 1304 is further configured to select the dynamic background material corresponding to the background material selection operation from a local material library; or download the dynamic background material corresponding to the background material selection operation from an online material library; or determine a video obtained by currently shooting a target environment as the dynamic background material.

[0259] In the above embodiment, the starting synthesis frame of the dynamic foreground material and the dynamic background material is determined, the dynamic background material and the dynamic foreground material are synthesized with the determined starting synthesis frame as the synthesis seven points, so that the user can select the part of interest in the dynamic background material for synthesis according to personal preference, so that the synthesized video has a better visual effect. In addition, when the dynamic background material is a presentation, the presentation can be synthesized with the dynamic foreground material as the background, so that a video of the dynamic foreground object interacting with the presentation can be obtained, and the video recording party can also synthesize the presentation and the anchor into the same video in a scene without a projector to project the presentation, so that the audience can watch the anchor and the presentation without any sense of strangeness, and the video watching effect is improved.

[0260] In one embodiment, as shown in FIG. 13, Figure 14 the device further comprises:

[0261] The extraction module 1310 is further configured to extract frame image samples from the material samples.

[0262] The separation module 1314 is further configured to separately perform channel separation on the extracted frame image samples as current frame image samples to obtain channel images corresponding to each color channel.

[0263] The extraction module 1310 is further configured to sequentially input the channel images of the current frame image samples and the image masks of the previous frame image samples of the current frame image samples to the untrained machine learning model for processing to obtain training prediction masks corresponding to each frame image sample.

[0264] The calculation module 1318 is configured to calculate error values between each training prediction mask and the corresponding label.

[0265] The adjustment module 1316 is further configured to adjust the model parameters of the machine learning model according to the error values until the error value between the prediction image mask output by the adjusted machine learning model and the corresponding label is less than the error threshold value, and stop training.

[0266] In one embodiment, as shown in FIG. 13, Figure 14 the device further comprises:

[0267] The label acquisition module 1320 is configured to, when the training prediction mask corresponding to the current frame image is obtained, and the error value between the training prediction mask and the corresponding label is less than a preset error value, determine the training prediction mask as the image mask of the specified image sample in the extracted frame image samples.

[0268] In one embodiment, the target video includes a live video or a conference video; as shown in FIG. 13, Figure 14 the device further comprises:

[0269] The video acquisition module 1322 is configured to acquire a live video generated when a live broadcast is performed through a live broadcast room on a social application, or a conference video generated when an electronic conference is performed in a conference room on the social application.

[0270] The sending module 1324 is configured to send the synthesized video to a user equipment accessing the live broadcast room or the conference room when the synthesized video corresponding to the live video or the conference video is obtained.

[0271] In the above embodiment, the channel image of each color channel of the current frame image sample and the image mask of the previous frame image sample are input into the machine learning model together, the training prediction mask of the current frame image sample is predicted by taking the image mask of the previous frame image sample as prior knowledge, the error value between the training prediction mask and the corresponding label is calculated, and the model parameters of the machine learning model are adjusted according to the error value, so that the trained machine learning model can quickly segment out the dynamic foreground material in the user-selected material, improve the image segmentation rate, and thus make the video synthesis faster and meet the real-time requirement.

[0272] In one embodiment, the synthesizing module 1308 is further configured to identify key information in the dynamic background material, hide or fade the dynamic foreground material that blocks the key information if the dynamic foreground material blocks the key information, and synthesize the dynamic foreground material after the hiding or fading and the dynamic background material into a video.

[0273] In the above embodiment, when important content (i.e., key information) is blocked, the dynamic foreground material that blocks the important content can be hidden or faded, so that when the synthesized video is watched, the important content can still be watched even if it is blocked by the dynamic foreground material, thereby improving the watching effect of the synthesized video.

[0274] In one embodiment, the synthesizing module 1308 is further configured to identify a facial expression of a target object in the dynamic foreground material, acquire a corresponding situational element according to the facial expression, and synthesize the situational element, the dynamic background material, and the dynamic foreground material played on the video synthesis preview page into a video.

[0275] In the above embodiment, by identifying the facial expression, a situational element matched with the identified facial expression is obtained, and the situational element is synthesized into a video together with the dynamic background material and the dynamic foreground material, so that the video content is more vivid and lifelike, and the interestingness and infectivity of the synthesized video are increased.

[0276] In one embodiment, the synthesizing module 1308 is further configured to extract the voice from the user-selected material, perform voice recognition on the extracted voice to obtain corresponding text, acquire matching prompt information according to semantic information of the text, and synthesize the prompt information, the dynamic background material and the dynamic foreground material played on the video synthesis preview page into a video.

[0277] In the above embodiments, the corresponding text is obtained through voice recognition, the matching prompt information is obtained according to semantic information of the text, and the prompt information is synthesized into a video together with the dynamic background material and the dynamic foreground material, so that the user can watch the prompt information matching the voice while hearing the voice, and the user's attention is focused on the display content of the video.

[0278] The specific limitations of the video synthesizing apparatus can refer to the limitations of the video synthesizing method described above, and will not be repeated here. Each module in the above video synthesizing apparatus can be realized by software, hardware and combinations thereof, in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.

[0279] In one embodiment, a computer device is provided, which can be a terminal, and its internal structure diagram can be as shown in FIG. 13. Figure 15 The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The communication interface of the computer device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved through WIFI, operator network, NFC (Near Field Communication) or other technologies. The computer program is executed by the processor to implement a video synthesizing method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0280] Those skilled in the art can understand that Figure 15 the structure shown in FIG. 13 is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0281] In an embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above-mentioned method embodiments when executing the computer program.

[0282] In an embodiment, a computer readable storage medium is provided, storing a computer program, which, when executed by a processor, implements the steps in the above-mentioned method embodiments.

[0283] In an embodiment, a computer program product or computer program is provided, including computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above-mentioned method embodiments.

[0284] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0285] The technical features of the above embodiments can be combined in any manner. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present disclosure.

[0286] The above-described embodiments are merely illustrative of several embodiments of the present application, which are described in more detail and in a specific and detailed manner, but should not be construed as limiting the scope of the patent. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method of video compositing, the method comprising: The method is applied to a social application, and comprises: obtaining user-selected material being played through the social application, the user-selected material being any one of a live video corresponding to a live broadcast being conducted or a conference video corresponding to an electronic conference being conducted; extracting dynamic foreground material from the user-selected material based on a machine learning model; entering a dynamic background material obtaining page of the social application, and selecting dynamic background material in response to a background material selection operation triggered in the dynamic background material obtaining page; the dynamic background material comprises a presentation; a video editor is configured in the social application; entering a video synthesis preview page of the video editor, playing the dynamic background material, and superimposing the dynamic foreground material in an editing state on the played dynamic background material for display; the video synthesis preview page comprises an entry control for selecting dynamic background material; the editing state is that the dynamic foreground material is in an editable state in the video synthesis preview page; identifying key information in the dynamic background material, the key information comprising important images, texts or animations in the dynamic background material; if the dynamic foreground material obscures the key information, performing hiding processing or fading processing on the dynamic foreground material; the hiding processing is to cut out a target person in the dynamic foreground material or replace the dynamic foreground material with blank material; the fading processing is to adjust transparency and color of the dynamic foreground material; during the hiding processing or the fading processing, audio corresponding to the dynamic foreground material is reserved; if the key information originally obscured disappears or moves to a region that is not obscured, stopping the hiding or fading processing on the dynamic foreground material and restoring display of the dynamic foreground material; performing speech recognition on speech in the user-selected material to obtain corresponding text; obtaining matching prompt information according to semantic information of the text; the content of the prompt information is not completely the same as the content of the text; the prompt information is used to attract attention of a user watching a video to display content of the video; identifying a facial expression of a target object in the dynamic foreground material; obtaining corresponding situational elements according to the facial expression; the situational elements comprise patterns, virtual expressions or texts matched with the facial expression; synthesizing the prompt information, the situational elements, the dynamic background material and the dynamic foreground material played in the video synthesis preview page into a video; sending the synthesized video to a user device accessing a live broadcast room or a conference room, so that a live broadcast audience or a conference participant watches the live broadcast video or the conference video after the video background is changed; the training process of the machine learning model comprises: extracting frame image samples from material samples; wherein the material samples are any one of a live video corresponding to a live broadcast being conducted and a conference video corresponding to an electronic conference being conducted; a frame image sample refers to an image of each frame in the material sample; If the current frame image sample is a first frame, an image mask of the first frame image sample is calculated, and the image mask of the first frame is taken as an image mask of a specified image sample in the frame image sample; If the current frame image sample is a non-first frame, a channel image of the current frame image sample and an image mask of a previous frame image sample of the current frame image sample are input into an untrained machine learning model for processing to obtain a training prediction mask of the current frame image sample; When the training prediction mask corresponding to the current frame image sample is obtained, and an error value between the training prediction mask and a corresponding label is less than a preset error value, the training prediction mask is determined as the label corresponding to the specified image sample in the frame image sample; Based on error values between training prediction masks of each frame image sample and corresponding labels, model parameters of the machine learning model are adjusted.

2. The method of claim 1, wherein, The extraction step of the dynamic foreground material includes: extracting frame images from user-selected materials; performing image segmentation on the dynamic foreground material in the frame image through the machine learning model.

3. The method of claim 2, wherein, The method further includes: If the extracted frame image is a first frame image, an image mask corresponding to the first frame image is determined; If the extracted frame image is a non-first frame image, the extracted frame image is sequentially taken as a current frame image for channel separation to obtain channel images corresponding to each color channel; The image segmentation on the dynamic foreground material in the frame image through the machine learning model includes: inputting the channel image of the current frame image and the mask of the previous frame image of the current frame image into the machine learning model for processing to obtain a prediction mask corresponding to each frame image; segmenting the dynamic foreground material from the corresponding frame image according to the image mask and the prediction mask.

4. The method of claim 1, wherein, The method further includes: adjusting display parameters of the dynamic foreground material; synthesizing the dynamic background material and the dynamic foreground material after the display parameters are adjusted, which are played on the video synthesis preview page, to obtain a synthesized video.

5. The method of claim 4, wherein, The adjustment of the display parameters of the dynamic foreground material includes at least one of the following: adjusting the position and size of the dynamic foreground material; copying the dynamic foreground material and adjusting the relative position between at least two copied dynamic foreground materials.

6. The method of claim 1, wherein, The playing of the dynamic background material and the superimposition of the dynamic foreground material in the editing state on the played dynamic background material for display includes: determining the starting synthesis frame of the dynamic foreground material and the dynamic background material, respectively; playing the dynamic background material from the starting synthesis frame of the dynamic background material as a starting position, and superimposing the dynamic foreground material in the editing state on the played dynamic background material from the starting synthesis frame of the dynamic background material as the starting position for display; The method further includes: synthesizing the dynamic background material and the dynamic foreground material after the starting synthesis frame, which are played on the video synthesis preview page, to obtain a synthesized video.

7. The method of claim 6, wherein, The determination of the starting synthesis frame of the dynamic foreground material and the dynamic background material, respectively, includes: cutting the presentation into presentation images used as backgrounds page by page; determine a starting synthesis frame in the dynamic foreground material and the segmented presentation image respectively; the synthesizing the dynamic background material and the dynamic foreground material, which are played in the video synthesis preview page and located after the starting synthesis frame, comprises: synthesizing each presentation image with at least one frame of dynamic foreground material corresponding to the presentation image in the dynamic foreground material and the presentation image, which are played in the video synthesis preview page and located after the starting synthesis frame.

8. The method of claim 6, wherein, the synthesizing the dynamic background material and the dynamic foreground material, which are played in the video synthesis preview page and located after the starting synthesis frame, comprises: synthesizing the dynamic background material and the dynamic foreground material, which are played in the video synthesis preview page and located after the starting synthesis frame, by a local video editor; or sending the dynamic background material and the dynamic foreground material, which are located after the starting synthesis frame, to a server to make the server synthesize the dynamic foreground material and the dynamic background material frame by frame to obtain a synthesized video.

9. The method of claim 1, wherein, the selecting the dynamic background material comprises: selecting a dynamic background material corresponding to the background material selection operation from a local material library; or downloading a dynamic background material corresponding to the background material selection operation from an online material library; or determining a video obtained by currently shooting a target environment as the dynamic background material.

10. The method of claim 2 or 3, wherein, if the current frame image sample is not the first frame, inputting a channel image of the current frame image sample and an image mask of a previous frame image sample of the current frame image sample into an untrained machine learning model for processing, comprising: respectively performing channel separation on the extracted frame image samples as current frame image samples to obtain channel images corresponding to each color channel; sequentially inputting the channel images of the current frame image samples and the image masks of the previous frame image samples of the current frame image samples into the untrained machine learning model for processing to obtain training prediction masks corresponding to the frame image samples; adjusting model parameters of the machine learning model based on error values between the training prediction masks of the frame image samples and corresponding labels, comprising: adjusting the model parameters of the machine learning model according to the error values until an error value between a prediction image mask output by the adjusted machine learning model and a corresponding label is less than an error threshold, and stopping training.

11. The method of claim 2 or 3, wherein, before entering the dynamic background material acquisition page of the social application, the method further comprises: when the dynamic foreground material is extracted, displaying a frame of the dynamic foreground material as a preview picture; or acquiring an image corresponding to a target object in the dynamic foreground material as a preview picture for display.

12. The method according to any one of claims 1 to 9, characterized in that, before playing the dynamic background material, the method further comprises: aligning the dynamic background material with the dynamic foreground material in the editing state; the alignment comprises aligning a first frame of the dynamic background material with a first frame or a specified frame of the dynamic foreground material.

13. A video compositing apparatus characterized by comprising: the device comprises: The video acquisition module is configured to acquire user-selected material being played through a social application, the user-selected material being any one of a live video corresponding to a live broadcast being performed or a conference video corresponding to an electronic conference being performed; The module is configured to extract dynamic foreground material from the user-selected material based on a machine learning model; The entering module is configured to enter a dynamic background material acquisition page of the social application; The selecting module is configured to select dynamic background material in response to a background material selection operation triggered in the dynamic background material acquisition page; the dynamic background material includes a presentation; and the social application is configured with a video editor; The entering module is further configured to enter a video synthesis preview page of the video editor, the video synthesis preview page including an entry control for selecting dynamic background material; The playing module is configured to play the dynamic background material and superimpose the dynamic foreground material in an editing state on the played dynamic background material for display; the editing state is that the dynamic foreground material is in an editable state in the video synthesis preview page; The synthesizing module is configured to identify key information in the dynamic background material, the key information including important images, texts or animations in the dynamic background material; if the dynamic foreground material obscures the key information, the dynamic foreground material is subjected to hiding processing or fading processing; the hiding processing is to cut out a target person in the dynamic foreground material or replace the dynamic foreground material with blank material; the fading processing is to adjust transparency and color of the dynamic foreground material; during the hiding processing or the fading processing, audio corresponding to the dynamic foreground material is retained; if the originally obscured key information disappears or moves to a non-obscured area, the hiding or fading processing of the dynamic foreground material is stopped and the dynamic foreground material is displayed again; The synthesizing module is further configured to perform speech recognition on speech in the user-selected material to obtain corresponding text; obtain matching prompt information according to semantic information of the text, the content of the prompt information being not completely the same as the content of the text, the prompt information being used to attract attention of a user watching a video to display content of the video; The synthesizing module is further configured to identify a facial expression of a target object in the dynamic foreground material; obtain corresponding situational elements according to the facial expression, the situational elements including patterns, virtual expressions or texts matched with the facial expression; The synthesizing module is further configured to synthesize the prompt information, the situational elements, the dynamic background material and the dynamic foreground material played in the video synthesis preview page into a video; The sending module is configured to send the synthesized video to a user device accessing a live broadcast room or a conference room, so that a live broadcast audience or a conference participant watches live broadcast video or conference video with changed video background; The training process of the machine learning model includes: The extraction module is configured to extract frame image samples from a material sample; the material sample is any one of a live video corresponding to a live broadcast being performed and a conference video corresponding to an electronic conference being performed, and the frame image sample refers to an image of each frame in the material sample; The module is configured to perform the following steps: if the current frame image sample is a first frame, calculate an image mask of the first frame image sample, and use the image mask of the first frame as an image mask of a specified image sample in the frame image sample; The extraction module is further configured to, if the current frame image sample is a non-first frame, input a channel image of the current frame image sample and an image mask of a previous frame image sample of the current frame image sample into an untrained machine learning model for processing to obtain a trained prediction mask of the current frame image sample; The label acquisition module is configured to, when the trained prediction mask corresponding to the current frame image sample is obtained and an error value between the trained prediction mask and a corresponding label is less than a preset error value, determine the trained prediction mask as a label corresponding to a specified image sample in the frame image sample; The adjustment module is configured to adjust model parameters of the machine learning model based on error values between trained prediction masks of each frame image sample and corresponding labels.

14. The apparatus of claim 13, wherein, The device further includes: The extraction module is further configured to extract frame images from user-selected materials, and perform image segmentation on dynamic foreground materials in the frame images through the machine learning model.

15. The apparatus of claim 14, wherein, The device further includes: The determination module is configured to, if the extracted frame image is a first frame image, determine an image mask corresponding to the first frame image; The separation module is configured to, if the extracted frame image is a non-first frame image, sequentially perform channel separation on the extracted frame image as a current frame image to obtain channel images corresponding to each color channel; The extraction module is further configured to input the channel image of the current frame image and a mask of a previous frame image of the current frame image into the machine learning model for processing to obtain prediction masks corresponding to each frame image, and segment dynamic foreground materials from the corresponding frame images according to the image mask and the prediction mask.

16. The device of claim 13, wherein The adjustment module is further configured to adjust display parameters of the dynamic foreground materials; The synthesis module is further configured to synthesize the dynamic background materials and the dynamic foreground materials after the display parameters are adjusted, which are played on the video synthesis preview page, to obtain a synthesized video.

17. The apparatus of claim 16, wherein, The adjustment module is further configured to perform at least one of the following: adjust a position and a size of the dynamic foreground materials; and copy the dynamic foreground materials and adjust a relative position between at least two copied dynamic foreground materials.

18. The apparatus of claim 13, wherein, The playing module is further configured to determine a starting synthesis frame of the dynamic foreground materials and the dynamic background materials respectively, play the dynamic background materials from the starting synthesis frame of the dynamic background materials as a starting playing position, and superimpose the dynamic foreground materials in an editing state on the played dynamic background materials for display from the starting synthesis frame of the dynamic background materials as the starting playing position. The synthesis module is further configured to synthesize the dynamic background material and the dynamic foreground material played on the video synthesis preview page and located after the starting synthesis frame to obtain a synthesized video.

19. The apparatus of claim 18, wherein, The device further comprises: The determination module is configured to split the presentation according to pages into presentation images used as backgrounds, and determine starting synthesis frames in the dynamic foreground material and the split presentation images respectively. The synthesis module is further configured to synthesize each presentation image and at least one frame of the dynamic foreground material in the presentation image and the dynamic foreground material played on the video synthesis preview page and located after the starting synthesis frame.

20. The apparatus of claim 18, wherein, The synthesis module is further configured to synthesize the dynamic background material and the dynamic foreground material played on the video synthesis preview page and located after the starting synthesis frame by using a local video editor, or send the dynamic background material and the dynamic foreground material located after the starting synthesis frame to a server to make the server synthesize the dynamic foreground material and the dynamic background material frame by frame to obtain a synthesized video.

21. The apparatus of claim 13, wherein, The selection module is further configured to select the dynamic background material corresponding to the background material selection operation from a local material library, or download the dynamic background material corresponding to the background material selection operation from an online material library, or determine a video obtained by currently shooting a target environment as the dynamic background material.

22. The apparatus of claim 14 or 15, wherein, The device further comprises: The separation module is configured to separately separate the extracted frame image sample as a current frame image sample by channel to obtain channel images corresponding to each color channel. The extraction module is further configured to sequentially input the channel images of the current frame image sample and image masks of a previous frame image sample of the current frame image sample into an untrained machine learning model for processing to obtain training prediction masks corresponding to each frame image sample. The adjustment module is further configured to adjust model parameters of the machine learning model according to the error value until an error value between a prediction image mask output by the adjusted machine learning model and a corresponding label is less than an error threshold, and stop training.

23. The apparatus of any one of claims 14 or 15, wherein, The playing module is further configured to, when the dynamic foreground material is extracted, display one frame of the dynamic foreground material in the dynamic foreground material as a preview picture; or acquire an image corresponding to a target object in the dynamic foreground material as a preview picture for display.

24. The apparatus of any one of claims 14 or 15, wherein, The playing module is further configured to align the dynamic background material with the dynamic foreground material in the editing state; the alignment includes aligning a first frame of the dynamic background material with a first frame or a specified frame of the dynamic foreground material. 25.A computer device, comprising a memory and a processor, wherein the memory stores a computer program. The processor, when executing the computer program, implements the steps of the method of any one of claims 1 to 12.

26. A computer readable storage medium storing a computer program, wherein the computer program comprises the following steps of: The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 12.

Citation Information

Patent Citations

  • Image processing method and device, mobile terminal and computer readable storage medium

    CN108900769A

  • Video processing method and device and storage medium

    CN110290425A