Intelligent video character replacement method and system
By combining a unified algorithm and a deep learning video restoration algorithm with the Rain-Net algorithm for video person replacement, the problems of complex operation and poor fusion effect in existing technologies are solved, and automated video person replacement and high-quality image generation are achieved.
Patent Information
- Application Number
- CN202210094849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2042-01-26
AI Technical Summary
Existing technologies lack intelligence in video character replacement, are complex to operate, cannot achieve good integration of characters and backgrounds, affect the realism of the synthesized image, and cannot automatically segment and fill in lost information.
A unified algorithm is used for automatic segmentation of the target person, combined with a deep learning video restoration algorithm to supplement the loss information, and the Rain-Net algorithm is used for video harmonization. Transparency masking is used for compositing and lighting unification.
It achieves automated video character replacement, saving manual time, generating well-blended and natural images, automatically segmenting and completing the background, and unifying the color tone and lighting.
Smart Images

Figure CN114627404B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of media post-processing, and more particularly, to an intelligent video character replacement method and system. BACKGROUND
[0002] In recent years, the emergence of short videos has driven the development of the video production industry in pursuit of better user perception and experience. With the development of artificial intelligence, there are currently software and systems on the market that assist in automatic image cutting and other video production, but the effect is often not as refined as that produced by hand, and there are problems such as the inability to remove specified characters from the video, the inability to fill in missing background information, and the inability to implement a complete character replacement process. This type of software can only perform partial operations for video character replacement and it is difficult to achieve the desired effect, so in the process of replacing characters in a video, in order to achieve an excellent effect, image processing software and video processing software are still needed to manually cut out images, fill in background videos, and unify lighting and color tones for each frame of the video, which consumes a large amount of labor and time.
[0003] The prior art uses traditional manual processing of images and videos to achieve satisfactory results, lacks intelligence, and consumes time and effort to manually replace characters in a video. The video image cutting and synthesis technology currently available on the market cannot achieve a good fusion effect, and the image cutting technology only removes the characters and directly pastes other characters on the background with missing information. Since no video completion is performed, when the characters move, the missing parts in the background are exposed, and the sense of realism of the character replacement is not achieved. In the character and background synthesis technology, the replaced characters have a large color difference from the original background, do not have a strong sense of realism, cannot achieve fusion of the characters and the original video, have poor effects, and have obvious splicing marks. Or only one module of the video character replacement is implemented, such as camera acquisition and synthesis, without integration.
[0004] For example, a method for replacing characters in a video can play the fusion of the replaced characters and the original video, and play each frame in turn. However, this solution has the following problems that have not been solved: on the one hand, the method is a non-intelligent character replacement method, and the operation is relatively complex; on the other hand, the method cannot achieve a good fusion effect of the characters, affects the realism of the synthesized picture, and lacks authenticity.
[0005] Therefore, there is an urgent need for an intelligent video character replacement method and system that can automatically segment, automatically fill in missing information, automatically perform video synthesis, and adjust the lighting of the synthesized video to be consistent, so that the foreground character video is completely integrated into the background video. SUMMARY
[0006] In view of the above problems, the purpose of the present application is to provide an intelligent video character replacement method and system to solve the problems in the prior art that the existing non-intelligent character replacement method is relatively complex to operate, and cannot achieve good fusion effect of characters, affecting the realism of the synthesized picture and lacking in authenticity.
[0007] The intelligent video character replacement method provided by the present application comprises:
[0008] The target character in the video is automatically segmented by a preset unified algorithm to remove the target character to form a loss area; wherein the unified algorithm is an algorithm that unifies visual target tracking and video target segmentation;
[0009] Loss information about the loss area is obtained, and the video repair algorithm based on deep learning is used to supplement the loss information with adjacent frames and distant frames of the loss area in the video to form a video background;
[0010] The portrait extracted in the green screen is called to synthesize the portrait with the video background to form an interactive video, and the black screen white foreground formed by removing the green screen is exported to the local as a transparent channel mask;
[0011] The Rain-Net algorithm is used to perform video reconciliation on the interactive video based on the transparent channel mask to form an updated video.
[0012] Preferably, the unified algorithm adds a transparent channel mask branch based on the visual target tracking; wherein,
[0013] The unified algorithm comprises two layers of neural networks h φ and learning parameters φ to predict the transparent channel mask.
[0014] Preferably, the process of obtaining loss information about the loss area and supplementing the loss information with adjacent frames and distant frames of the loss area in the video by using the video repair algorithm based on deep learning to form a video background comprises:
[0015] A video repair algorithm combining time dimension and space dimension is constructed by a deep learning algorithm;
[0016] Adjacent frames and distant frames of the loss area in the video are used as inputs of the video repair algorithm;
[0017] The video repair algorithm extracts detection blocks from the adjacent frame and the remote frame, compares the similarity between the detection blocks to obtain the closest part to the loss area, and supplements the closest part to the loss area to complete the loss information to form a video background.
[0018] Preferably, the process of comparing the similarity between the detection blocks extracted from the adjacent frame and the remote frame includes:
[0019] The features of each frame of the adjacent frame and the remote frame are mapped as query features and memory features, and the query features and the memory features are searched through i ∈R h ×w×c wherein,
[0020] The f i represents the features required by the frame-level encoder, R represents all features; h and w are the height and width of the image of the adjacent frame or the remote frame, respectively, and c represents the number of channels embedding the query features and the memory features into the multi-scale-based attention module.
[0021] The spatial blocks with a shape of r1*r2*c are extracted from the query features and the memory features as detection blocks, and the similarity between the detection blocks is calculated using matrix multiplication; wherein r1 and r2 are the length and width of the detection blocks; and c represents the number of channels embedding the query features and the memory features into the multi-scale-based attention module.
[0022] Preferably, the process of calling the person image extracted from the green screen, synthesizing the person image with the video background to form an interactive video, and exporting the black screen white foreground formed by removing the green screen to the local as a transparent channel mask includes:
[0023] The person image in front of the green screen is extracted as a foreground by calling a capture card through the UE4 platform;
[0024] The video background and the black screen are simultaneously used as inputs for video synthesis; the person image is synthesized with the video background as a foreground to generate an interactive video; the processed image of the person image is synthesized with the black screen as a pure white foreground to generate a black screen white foreground, and the black screen white foreground is exported to the local as a transparent channel mask.
[0025] Preferably, the process of obtaining the processed image of the person image includes:
[0026] The luminance value of the portrait is adjusted to the maximum to generate a pure white foreground.
[0027] Preferably, the process of video harmonization of the interactive video by the Rain-Net algorithm based on the transparent channel mask to form an updated video, comprises:
[0028] The transparent channel mask is pre-processed to form a standard channel mask;
[0029] The standard channel mask is harmonized by the Rain-Net algorithm to generate an updated video to complete the video character replacement.
[0030] Preferably, in the Rain-Net algorithm, the foreground video about the portrait is represented by I f , the background video is represented by I b , the transparent channel mask is represented by M; the image of the synthesized interactive video is represented by I c =M·I f +(1-M)·I b .
[0031] The harmonization model for harmonizing the interactive video is a self-learning model, defined as G, and the image generated after harmonization by the harmonization model is
[0032] In another aspect, the present application also provides an intelligent video character replacement system, which performs video character replacement processing based on the intelligent video character replacement method as described above, comprising:
[0033] A target segmentation unit is configured to automatically segment a target character in a video by a preset unified algorithm to remove the target character to form a loss area; wherein the unified algorithm is an algorithm that unifies visual target tracking and video target segmentation;
[0034] A loss filling unit is configured to obtain loss information about the loss area, and fill the loss information completely by using adjacent frames and distant frames of the loss area in the video by a video repair algorithm based on deep learning to form a video background;
[0035] A portrait synthesis unit is configured to call a portrait extracted from a green screen, synthesize the portrait with the video background to form an interactive video, and export a black screen white foreground formed by removing the green screen to a local as a transparent channel mask;
[0036] A video harmonization unit is configured to perform video harmonization of the interactive video by the Rain-Net algorithm based on the transparent channel mask to form an updated video.
[0037] Preferably, the video repair algorithm is an algorithm combining time dimension and space dimension constructed by a deep learning algorithm; the adjacent frames and the far frames of the loss region in the video are input into the video repair algorithm, so that the video repair algorithm extracts detection blocks from the adjacent frames and the far frames, compares the similarity between the detection blocks, obtains the closest part to the loss region, and supplements the closest part to the loss region in the loss region to complete the loss information, thereby forming a video background.
[0038] From the above technical solution, the intelligent video character replacement method provided by the application first automatically segments the target character in the video by a preset unified algorithm to form a loss region by removing the target character; wherein the unified algorithm is an algorithm that unifies visual target tracking and video target segmentation; then loss information about the loss region is obtained, and the video repair algorithm based on deep learning is used to supplement the loss information to form a video background by using the adjacent frames and the far frames of the loss region in the video; then the portrait extracted in the green screen is called, the portrait is synthesized with the video background to form an interactive video, and the black screen white foreground formed by removing the green screen is exported to the local as a transparent channel mask; then the Rain-Net algorithm is used to perform video reconciliation on the interactive video based on the transparent channel mask to form an updated video, thereby saving a lot of time for manually producing videos; only the video to be modified needs to be input and the character needs to be selected to automatically generate character removal and background video completion; and the synthesized video is unified in tone and illumination, so that the video shows a good fusion and more natural picture. BRIEF DESCRIPTION OF DRAWINGS
[0039] Other objects and results of the application will become more apparent and easily understood with reference to the following description of the application taken in conjunction with the accompanying drawings, and with the more complete understanding of the application. In the drawings:
[0040] Figure 1 A flowchart of the intelligent video character replacement method according to the embodiment of the application;
[0041] Figure 2 A system block diagram of the intelligent video character replacement system according to the embodiment of the application. DETAILED DESCRIPTION
[0042] In the prior art, in order to achieve a satisfactory effect, a traditional manual image and video processing is generally adopted, which lacks intelligence, and manual video character replacement is time-consuming and laborious; the video cutout and synthesis technologies on the market cannot achieve a good effect of fusion, and the cutout technology only removes the character and directly pastes other characters on the background with information missing, without video completion, so that when the character moves, the missing part in the background is exposed, and the real sense of character replacement is not achieved. In the technology of character and background synthesis, the replaced character and the original background have a large color difference, do not have a strong sense of reality, cannot achieve the fusion of the character and the original video, the effect is not good, and there are obvious splicing traces; or only one module of video character replacement is achieved, such as camera collection, synthesis, etc., without integration.
[0043] In view of the above problems, the present application provides an intelligent video character replacement method and system, which will be described in detail below in combination with the drawings.
[0044] In order to illustrate the intelligent video character replacement method provided by the present application, Figure 1 The intelligent video character replacement method of the embodiment of the present application is exemplarily indicated; Figure 2 The intelligent video character replacement system of the embodiment of the present application is exemplarily indicated.
[0045] The following exemplary embodiment description is actually merely illustrative, and by no means as any limitation on the present application and its application or use. The technology and equipment known to those skilled in the relevant art can not be discussed in detail, but in appropriate cases, the technology and equipment should be regarded as part of the specification.
[0046] As Figure 1 indicated, the intelligent video character replacement method of the embodiment of the present application provided by the present application shows the processing through the following steps, the image change of the image in the video, including:
[0047] S1: automatically segmenting a target character in a video by a preset unified algorithm to remove the target character to form a loss area; wherein the unified algorithm is an algorithm that unifies visual target tracking and video target segmentation;
[0048] S2: obtaining loss information about the loss area, and using a video repair algorithm based on deep learning to utilize adjacent frames and distant frames of the loss area in the video to supplement the loss information to form a video background;
[0049] S3: Call the human image extracted from the green screen, and composite the human image with the video background to form an interactive video. At the same time, export the black screen and white foreground formed by removing the green screen to the local machine as a transparency channel mask.
[0050] S4: The interactive video is harmonized based on the transparency channel mask using the Rain-Net algorithm to form an updated video.
[0051] exist Figure 1 In the embodiment shown, step S1 is a process of automatically segmenting the target person in the video using a preset unified algorithm to remove the target person to form a loss region; wherein, the unified algorithm is an algorithm that unifies visual target tracking and video target segmentation.
[0052] Specifically, this unified algorithm adds a transparency channel mask branch to the visual target tracking algorithm; the unified algorithm includes two layers of neural network h. φ And learn the parameter φ to predict the alpha channel mask;
[0053] The network framework of this unified algorithm is as follows:
[0054]
[0055] Where, m n The alpha channel mask representing the prediction, h φ This represents a two-layer neural network. express The similarity between the nth candidate window of sample z and x is encoded; in this unified algorithm, the tracking and segmentation of the target person can be achieved by relying only on the target person selected in the first frame of the video. After the algorithm runs, the transparent channel mask of the target person in each frame of the video can be obtained.
[0056] exist Figure 1 In the illustrated embodiment, step S2 involves obtaining loss information about the lost region and using a deep learning-based video inpainting algorithm to complete the loss information using adjacent and distant frames of the lost region in the video to form a video background; this includes:
[0057] S21: Construct a video restoration algorithm that combines temporal and spatial dimensions using deep learning algorithms;
[0058] S22: Use the adjacent frames and distant frames of the lost region in the video as input to the video restoration algorithm;
[0059] S23: the video repair algorithm extracts detection blocks from the adjacent frames and the remote frames, compares the similarity between the detection blocks to obtain the most similar part to the loss area, and supplements the most similar part to the loss area to complete the loss information to form a video background;
[0060] wherein the process of comparing the similarity between the detection blocks extracted from the adjacent frames and the remote frames by the video repair algorithm comprises:
[0061] The features of each frame of the adjacent frames and the remote frames are mapped as query features and memory features, and the query features and the memory features are searched by f i ∈R h ×w×c wherein,
[0062] f i represents the features required by the frame-level encoder, R represents all features, i.e. i f i is the value obtained from R (all features); h and w are the height and width of the image of the adjacent frames or the remote frames, respectively, and c represents the number of channels for embedding the query features and the memory features into the multi-scale-based attention module.
[0063] The spatial blocks with a shape of r1*r2*c are extracted from the query features and the memory features as detection blocks, and the similarity between the detection blocks is calculated by matrix multiplication; wherein r1 and r2 are the length and width of the detection blocks, and c represents the number of channels for embedding the query features and the memory features into the multi-scale-based attention module.
[0064] Specifically, in the present embodiment, the original video has been removed the original character (target character) in the previous part, and the loss information generated by removing the character is completed in step S2. In step S2, a deep learning algorithm is used to construct a video repair algorithm combining time dimension and space dimension. The video repair algorithm uses frames far apart (remote frames) and adjacent frames (adjacent frames) as input, extracts different detection blocks from the frames, compares the similarity between the blocks (detection blocks) to detect the most similar part to the missing part, and then uses the similar part to complete the video. The algorithm is divided into three processes. In the encoding process, the features of each frame are mapped as query and memory for further searching, using wherein f i ∈Rh×w×c where h, w are the height and width of the input image respectively, c represents the number of channels embedding the query and memory into the multi-scale based attention module, which can model the corresponding relationship of each region in different semantic spaces; in the matching process, spatial blocks (detection blocks) with the shape of R1*R2*c are extracted from the query features and memory features of each frame of image, obtaining N=T*h / r1*w / r2. The features in each frame are mapped to obtain query features and key value features, and then the blocks in the query and the key module are respectively reshaped into one-dimensional vectors, and matrix multiplication is used to calculate the similarity between the blocks. After modeling the corresponding relationship of all spatial blocks in the final attention stage, the query value of each block can be obtained by the weighted sum of the correlation values, the most similar region of the missing part in each frame is detected for matching, and the completed video is output.
[0065] In Figure 1 In the embodiment shown, step S3 is a process of calling the person image cut out in the green screen, synthesizing the person image with the video background to form an interactive video, and exporting the black screen white foreground formed by removing the green screen to the local as a transparent channel mask; wherein, the process includes:
[0066] S31: calling a capture card through a UE4 platform to cut out the person image in front of the green screen as a foreground;
[0067] S32: simultaneously taking the video background and the black screen as inputs for video synthesis; synthesizing the person image as a foreground with the video background to generate an interactive video; synthesizing the processed image of the person image as a pure white foreground with the black screen to generate a black screen white foreground, and exporting the black screen white foreground to the local as a transparent channel mask;
[0068] wherein, the process of obtaining the processed image of the person image includes:
[0069] adjusting the brightness value of the person image to the maximum to generate a pure white foreground.
[0070] Specifically, in the embodiment, the green screen portrait is cut out as foreground by the UE4 platform calling the capture card, while two backgrounds are input, one is the video background already completed in the previous step, which is combined with the cut-out foreground; the other is a black screen background, and a copy of the previous foreground is obtained by adjusting the brightness value to the maximum to obtain a pure white foreground. The white foreground and black background here form a group. The "camera-foreground-background" sequence is arranged horizontally in space, and the perspective relationship can obtain the lens of the two groups of foreground and background synthesis. The portrait foreground in one group is combined with the video background to generate an interactive video for real-time return and interaction with the user, and the interactive video is output to the local. The other group of white foreground and black background is used for video synthesis to generate a black screen and white foreground, and the black screen and white foreground are exported to the local as a transparent channel mask of the return image (image of the interactive video) for image blending in the next step.
[0071] In the embodiment shown in the figure, step S4 is a process of video blending of the interactive video based on the transparent channel mask by Rain-Net algorithm to form an updated video; wherein, it includes:
[0072] S41: preprocessing the transparent channel mask to form a standard channel mask;
[0073] S42: blending the standard channel mask with the interactive video by the Rain-Net algorithm to generate an updated video, completing the video character replacement.
[0074] In the Rain-Net algorithm, the foreground video of the portrait is represented by I f , the background video is represented by I b , and the transparent channel mask is represented by M. The image of the synthesized interactive video is represented by I c =M·I f +(1-M)·I b ; wherein, "·" is an operator symbol, called Hadamard product, Hadamard product, a class of operations next to the true operation, or called basic product; the blending model for blending the interactive video is a self-learning model, defined as G, and the image generated after blending by the blending model is
[0075] More specifically, in the embodiment, the transparent channel mask exported from UE4 is actually a grayscale image due to the lighting effect of the UE4 environment, and can be used only after a binaryzation preprocessing, i.e. the standard channel mask Then the synthesized interactive video is harmonized based on the standard channel mask using the Rain-Net algorithm. Specifically, the color of the white area (i.e., the foreground part in the video) in the standard channel mask is harmonized to the color of the black area (i.e., the background part in the video) so as to be unified and harmonious. In the Rain-Net algorithm, the foreground video of the person is represented by I f , the background video is represented by I b , and the foreground transparent channel mask is represented by M. The composition of the image is represented by I c =M·I f +(1-M)·I b , the harmonization model is defined as G, and the image obtained after harmonization is represented by I . G is a self-learning model, and according to self-learning, the illumination and color tone of the foreground image can be made closer to the background through ||G(I c ,M)-I||1, so as to realize the unification of the person and the background video in color tone and illumination, and make the video show a good fusion and a more natural picture.
[0076] Since video production relies on image and video processing software to perform complex operations on each frame of the video, by using the intelligent video person replacement system and the deep learning algorithm, only the original video and the person video need to be input, the manual frame-by-frame operation can be saved, and thus the video with unified color tone style after person replacement can be directly obtained, the operation time is saved, and the threshold of video operation is reduced.
[0077] In summary, the intelligent video person replacement method provided by the present application first automatically segments the target person in the video by using a preset unified algorithm to remove the target person and form a loss area. The unified algorithm is an algorithm that unifies visual object tracking and video object segmentation. Then, loss information about the loss area is obtained, and the video repair algorithm based on deep learning is used to supplement the loss information using adjacent frames and distant frames of the loss area in the video to form a video background. Then, the portrait extracted in the green screen is called, the portrait and the video background are synthesized to form an interactive video, and the black screen white foreground formed by removing the green screen is exported to the local as a transparent channel mask. Then, the Rain-Net algorithm is used to harmonize the interactive video based on the transparent channel mask to form an updated video. In this way, a lot of time for manually producing the video is saved. Only the video to be modified and the person to be selected need to be input, and the person extraction and background video completion can be automatically generated. The synthesized video is unified in color tone and illumination, and the video shows a good fusion and a more natural picture.
[0078] As Figure 2As shown, the application also provides an intelligent video character replacement system 100, which performs video character replacement processing based on the intelligent video character replacement method as described above, comprising:
[0079] a target segmentation unit 101 configured to automatically segment a target character in a video by a preset unified algorithm to remove the target character to form a loss area; wherein the unified algorithm is an algorithm that unifies visual target tracking and video target segmentation;
[0080] a loss filling unit 102 configured to obtain loss information about the loss area, and fill the loss information completely by using adjacent frames and distant frames of the loss area in the video based on a video repair algorithm obtained by deep learning to form a video background;
[0081] a portrait synthesis unit 103 configured to call a portrait extracted in a green screen, synthesize the portrait with the video background to form an interactive video, and export a black screen white foreground formed by removing the green screen to a local as a transparent channel mask;
[0082] a video harmonization unit 104 configured to perform video harmonization on the interactive video based on the transparent channel mask by a Rain-Net algorithm to form an updated video.
[0083] wherein the video repair algorithm is an algorithm that combines time dimension and space dimension constructed by deep learning algorithm; the adjacent frames and the distant frames of the loss area in the video are input to the video repair algorithm, so that the video repair algorithm extracts detection blocks from the adjacent frames and the distant frames, compares the similarity between the detection blocks, obtains the most similar part to the loss area, and fills the most similar part to the loss area in the loss area to completely fill the loss information and form a video background.
[0084] As described above, the intelligent video character replacement system provided by the present application utilizes the target segmentation unit 101 to automatically segment the target character in the video through a preset unified algorithm to remove the target character to form a loss area; wherein the unified algorithm is an algorithm that unifies visual target tracking and video target segmentation; the loss filling unit 102 obtains loss information about the loss area, and utilizes the video repair algorithm based on deep learning to supplement the loss information with adjacent frames and distant frames of the loss area in the video to form a video background; the portrait synthesis unit 103 calls the portrait extracted in the green screen, synthesizes the portrait with the video background to form an interactive video, and exports the black screen white foreground formed by the green screen extraction to the local as a transparent channel mask; the video reconciliation unit 104 utilizes the Rain-Net algorithm to reconcile the interactive video based on the transparent channel mask to form an updated video, thus saving a lot of time for manually producing videos; only the video to be modified needs to be input and the character needs to be selected to automatically generate character extraction and background completion video; and the synthesized video is unified in tone and lighting, so that the video shows a good fusion and a more natural picture.
[0085] The intelligent video character replacement method and system according to the present application are described above with reference to the accompanying drawings by way of example. However, those skilled in the art should understand that various improvements can be made to the intelligent video character replacement method and system according to the present application described above without departing from the content of the present application. Therefore, the protection scope of the present application should be determined by the content of the appended claims.
Claims
1. An intelligent video character replacement method, characterized in that, The method comprises the following steps: automatically segmenting a target person in a video by a preset unified algorithm to remove the target person and form a loss area; wherein the unified algorithm is an algorithm that unifies visual target tracking and video target segmentation; obtaining loss information about the loss area, and using a video repair algorithm based on deep learning to supplement the loss information with adjacent frames and distant frames of the loss area in the video to form a video background; wherein the method comprises: constructing a video repair algorithm combining time dimension and space dimension by a deep learning algorithm; taking the adjacent frames and distant frames of the loss area in the video as inputs of the video repair algorithm; causing the video repair algorithm to extract detection blocks from the adjacent frames and distant frames, compare the similarity between the detection blocks, obtain the part most similar to the loss area, and supplement the part most similar to the loss area in the loss area to supplement the loss information and form a video background; calling the portrait extracted in the green screen, synthesizing the portrait with the video background to form an interactive video, and exporting the black screen white foreground formed by removing the green screen to the local as a transparent channel mask; performing video reconciliation on the interactive video based on the transparent channel mask by a Rain-Net algorithm to form an updated video.
2. The intelligent video character replacement method of claim 1, wherein: the unified algorithm adds a transparent channel mask branch based on the visual target tracking; and The unified algorithm includes a two-layer neural network h φ and a learning parameter φ to predict a transparency pass mask.
3. The intelligent video character replacement method of claim 2, wherein: the process of comparing the similarity between the detection blocks extracted from the adjacent frames and distant frames by the video repair algorithm comprises: Features of each frame of the adjacent frames and the far frames are mapped as query features and memory features, and are searched by The query features and the memory features are searched; wherein, f i ∈R h×w×c , i = 1, 2, 3…T; wherein, The f i R represents all the features; h, w are the height and width of the image of the adjacent frame or the remote frame, respectively, and c represents the number of channels embedding the query feature and the memory feature into the multi-scale-based attention module. extracting spatial blocks with a shape of r1*r2*c from the query features and memory features as detection blocks, and calculating the similarity between the detection blocks by matrix multiplication; wherein r1 and r2 are the length and width of the detection blocks, and c represents the number of channels embedding the query features and memory features into a multi-scale based attention module.
4. The intelligent video character replacement method of claim 1, wherein, The process of calling the portrait extracted in the green screen, synthesizing the portrait with the video background to form an interactive video, and exporting the black screen white foreground formed by removing the green screen to the local as a transparent channel mask comprises: calling a capture card through a UE4 platform to extract the portrait in front of the green screen as a foreground; simultaneously taking the video background and a black screen as inputs of video synthesis; causing the portrait to be synthesized with the video background as a foreground to generate an interactive video; and causing the processed image of the portrait to be synthesized with the black screen as a pure white foreground to generate a black screen white foreground, and exporting the black screen white foreground to the local as a transparent channel mask.
5. The intelligent video character replacement method of claim 4, wherein, The process of obtaining the processed image of the portrait comprises: adjusting the brightness value of the portrait to the maximum to generate a pure white foreground.
6. The intelligent video character replacement method of claim 1, wherein, The process of performing video reconciliation on the interactive video based on the transparent channel mask by a Rain-Net algorithm to form an updated video comprises: preprocessing the transparent channel mask to form a standard channel mask; harmonizing the standard channel mask to the interactive video through the Rain-Net algorithm to generate an updated video, completing the video character replacement. 7.The intelligent video character replacement method of claim 6, wherein In the Rain-Net algorithm, the foreground video about the human image is represented by I f , the background video is represented by I b , the transparent channel mask is represented by M; the image of the synthesized interactive video is represented by I c =M·I f +(1-M)·I b ; The harmonization model for harmonizing the interactive video is a self-learning model, defined as G, and the image generated after harmonization by the harmonization model is 8. An intelligent video character replacement system, characterized by, the video character replacement processing based on the intelligent video character replacement method of any one of claims 1-7, comprising: a target segmentation unit configured to automatically segment a target character in a video through a preset unified algorithm to remove the target character to form a loss area; wherein the unified algorithm is an algorithm that unifies visual target tracking and video target segmentation; a loss filling unit configured to obtain loss information about the loss area, and fill the loss information completely through a video repair algorithm based on deep learning using adjacent frames and distant frames of the loss area in the video to form a video background; wherein the video repair algorithm is an algorithm that combines time dimension and space dimension through deep learning algorithm; the adjacent frames and the distant frames of the loss area in the video are input to the video repair algorithm, so that the video repair algorithm extracts detection blocks from the adjacent frames and the distant frames, compares the similarity between the detection blocks, obtains the closest part to the loss area, and fills the closest part to the loss area in the loss area to completely fill the loss information and form a video background; a portrait synthesis unit configured to call a portrait extracted in a green screen, synthesize the portrait with the video background to form an interactive video, and export a black screen white foreground formed by removing the green screen to a local as a transparent channel mask; a video harmonization unit configured to harmonize the interactive video through the Rain-Net algorithm based on the transparent channel mask to form an updated video.
Citation Information
Patent Citations
Face replacement method, device and electronic device
CN107316020A