Video character replacement method and device, equipment and storage medium

CN120034699APending Publication Date: 2025-05-23BEIJING 58 INFORMATION TTECH CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510174465.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

During the replacement process, traditional video character replacement schemes can easily lead to unnatural transition between the background image and the replaced characters, and the replaced characters are not smooth, making it difficult to achieve a natural and realistic replacement effect.

Method used

By obtaining the bone key points and background video frames of the characters to be replaced from the original video frame sequence, using the background generation model to fill the background blank area, using the character generation model to generate action-matched target character images, and combining the generated background and character images to generate a natural transition target video frame.

Benefits of technology

The natural transition between the characters and the background image in the target video frame, as well as the coordinated presentation of the actions and background image of the target video frame, achieving a natural and realistic replacement effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034699A_ABST
    Figure CN120034699A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video character replacement method and device, equipment and a storage medium. The method comprises the following steps: acquiring a first target video frame containing a to-be-replaced character from a video frame sequence corresponding to an original video; acquiring skeleton key points of a to-be-replaced character in the first target video frame and a first background video frame obtained after the to-be-replaced character is removed from the first target video frame; inputting the first background video frame into a background generation model so as to perform background filling on a blank area corresponding to the character to be replaced in the first background video frame, and generating a second background video frame; inputting the skeleton key points and the visual information of the target person into a person generation model to generate a target person image matched with the action described by the skeleton key points; and according to the second background video frame and the target person image, generating a second target video frame, and realizing natural transition between the target person and the person action and the background picture, so that a target video generated based on the second target video frame has a natural and real replacement effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and storage medium for replacing characters in a video. Background Art

[0002] In the video processing fields such as film and television production, advertising, and video editing, there is usually a need to replace a specific object (such as a person) in a video. However, when replacing a person in a video, traditional video character replacement solutions are prone to problems such as unnatural transitions between the background image in the video and the replaced person, and unsmooth movements of the replaced person, making it difficult to achieve a natural and realistic replacement effect. Summary of the invention

[0003] Embodiments of the present application provide a method, apparatus, device and storage medium for replacing characters in a video, so as to achieve natural replacement of characters in a video.

[0004] In a first aspect, an embodiment of the present application provides a method for replacing a character in a video, the method comprising:

[0005] Acquire multiple video frames containing the person to be replaced from a video frame sequence corresponding to the original video;

[0006] Acquire skeleton key points of the character to be replaced in a first target video frame, and a first background video frame obtained by removing the character to be replaced from the first target video frame, wherein the first target video frame is any one of the multiple video frames;

[0007] Inputting the first background video frame into a background generation model, so that the background generation model performs background filling on a blank area corresponding to the to-be-replaced person in the first background video frame, so as to generate a second background video frame;

[0008] Inputting the skeleton key points and the visual information of the target person into a person generation model to generate an image of the target person that matches the action described by the skeleton key points;

[0009] Generate a second target video frame according to the second background video frame and the target person image;

[0010] Generate a target video according to the second target video frames respectively corresponding to the multiple video frames.

[0011] Optionally, obtaining a plurality of video frames containing the person to be replaced from a video frame sequence corresponding to the original video includes:

[0012] According to the frame rate of the original video and the frame rate of the target video to be generated, a plurality of initial video frames are obtained from a video frame sequence corresponding to the original video at a preset time interval, wherein the target video is a video corresponding to the original video after the character is replaced;

[0013] A plurality of video frames including the person to be replaced are obtained from the plurality of initial video frames.

[0014] Optionally, the step of acquiring the skeleton key points of the person to be replaced in the first target video frame, and a first background video frame obtained by removing the person to be replaced from the first target video frame, comprises:

[0015] Obtaining character outline information corresponding to the character to be replaced in the first target video frame;

[0016] According to the character outline information, the skeleton key points of the character to be replaced in the first target video frame are identified, and the character to be replaced is removed from the first target video frame to obtain a first background video frame.

[0017] Optionally, the step of inputting the skeleton key points and the visual information of the target person into a person generation model to generate an image of the target person matching the actions described by the skeleton key points comprises:

[0018] The character contour information, the skeleton key points and the visual information of the target character are input into a character generation model to generate a target character image that matches the action described by the skeleton key points, wherein the character contour of the target character in the target character image matches the character contour described by the character contour information.

[0019] Optionally, generating a second target video frame according to the second background video frame and the target person image includes:

[0020] Determining a target area corresponding to the target person image in the second background video frame according to the person outline information;

[0021] The target person image is added to the target area to generate a second target video frame.

[0022] Optionally, generating a target video according to the second target video frames respectively corresponding to the multiple video frames includes:

[0023] Performing frame interpolation processing on adjacent second target video frames to obtain interpolation video frames between adjacent second target video frames;

[0024] A target video is generated according to the second target video frames and the interpolated video frames respectively corresponding to the multiple video frames.

[0025] Optionally, performing interpolation processing on adjacent second target video frames to obtain interpolation video frames between adjacent second target video frames includes:

[0026] In response to determining that the movements of the target person in the adjacent second target video frames are incoherent, the adjacent second target video frames are input into a video frame interpolation model so that the video frame interpolation model outputs interpolation video frames between the adjacent second target video frames.

[0027] In a second aspect, an embodiment of the present application provides a video character replacement device, the device comprising:

[0028] An acquisition module is used to acquire multiple video frames containing the person to be replaced from a video frame sequence corresponding to the original video; acquire the skeleton key points of the person to be replaced in a first target video frame, and a first background video frame obtained by removing the person to be replaced from the first target video frame, wherein the first target video frame is any one of the multiple video frames;

[0029] A generation module is used to input the first background video frame into a background generation model so that the background generation model fills the blank area corresponding to the to-be-replaced person in the first background video frame with the background to generate a second background video frame; input the skeleton key points and the visual information of the target person into the person generation model to generate an image of the target person that matches the action described by the skeleton key points;

[0030] The processing module is used to generate a second target video frame according to the second background video frame and the target person image; and generate a target video according to the second target video frames respectively corresponding to the multiple video frames.

[0031] Optionally, the acquisition module is specifically used to acquire multiple initial video frames from a video frame sequence corresponding to the original video at preset time intervals according to the frame rate of the original video and the frame rate of a target video to be generated, wherein the target video is a video corresponding to the original video after character replacement; and acquire multiple video frames containing the characters to be replaced from the multiple initial video frames.

[0032] Optionally, the acquisition module is also specifically used to obtain character contour information corresponding to the character to be replaced in the first target video frame; based on the character contour information, identify the skeletal key points of the character to be replaced in the first target video frame, and remove the character to be replaced from the first target video frame to obtain a first background video frame.

[0033] Optionally, the acquisition module is also specifically used to input the character contour information, the skeletal key points and the visual information of the target character into a character generation model to generate a target character image that matches the action described by the skeletal key points, wherein the character contour of the target character in the target character image matches the character contour described by the character contour information.

[0034] Optionally, the processing module is specifically used to determine, based on the character outline information, a target area corresponding to the target character image in the second background video frame; and add the target character image to the target area to generate a second target video frame.

[0035] Optionally, the processing module is also used to perform frame interpolation processing on adjacent second target video frames to obtain interpolated video frames between adjacent second target video frames; and generate a target video according to the second target video frames and the interpolated video frames respectively corresponding to the multiple video frames.

[0036] Optionally, the processing module is also specifically used to, in response to determining that the movements of the target person in adjacent second target video frames are incoherent, input adjacent second target video frames into a video frame interpolation model so that the video frame interpolation model outputs interpolation video frames between adjacent second target video frames.

[0037] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a communication interface; wherein a computer program is stored on the memory, and when the computer program is executed by the processor, the processor can at least implement the video character replacement method described in the first aspect.

[0038] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor of an electronic device, the processor can at least implement the video character replacement method as described in the first aspect.

[0039] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising: a computer program or instructions, which, when executed by a processor of an electronic device, enables the processor to at least implement the video character replacement method as described in the first aspect.

[0040] The video character replacement method provided by the embodiment of the present application replaces the character to be replaced in the original video with the target character and generates a target video containing the target character. First, multiple video frames containing the character to be replaced are obtained from the video frame sequence corresponding to the original video. After that, the skeleton key points of the character to be replaced in the first target video frame and the first background video frame obtained after the character to be replaced is removed from the first target video frame are obtained, wherein the first target video frame is any one of the multiple video frames. Next, the first background video frame is input into the background generation model so that the background generation model fills the blank area corresponding to the character to be replaced in the first background video frame to generate a second background video frame; the skeleton key points and the visual information of the target character are input into the character generation model to generate a target character image that matches the action described by the skeleton key points. Finally, a second target video frame is generated according to the second background video frame and the target character image; and a target video is generated according to the second target video frames corresponding to the multiple video frames. In this solution, for the first target video frame containing the person to be replaced, a second background video frame with a complete background is generated through the background generation model, and a target person image matching the action of the person to be replaced is generated through the person generation model. After that, the second background video frame is combined with the target task image to generate a second target video frame, so as to achieve a natural transition between the target person and the background image in the second target video frame, and a coordinated presentation of the target person's action and the background image. Thus, the target video generated based on the second target video frame has a natural and realistic replacement effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0042] Figure 1 A flowchart of a method for replacing a video character provided in an embodiment of the present application;

[0043] Figure 2 A schematic diagram of a scene of video character replacement provided in an embodiment of the present application;

[0044] Figure 3 A flowchart of another method for replacing a video character provided in an embodiment of the present application;

[0045] Figure 4 A flowchart of another method for replacing a video character provided in an embodiment of the present application;

[0046] Figure 5A flowchart of another method for replacing a video character provided in an embodiment of the present application;

[0047] Figure 6 A schematic diagram of the structure of a video character replacement device provided in an embodiment of the present application;

[0048] Figure 7 For Figure 6 A schematic structural diagram of an electronic device corresponding to the video character replacement device provided by the illustrated embodiment. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0050] It should be noted that, in the case where the embodiments of the present application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) are in compliance with relevant laws and standards.

[0051] In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.

[0052] Before introducing the video task replacement method provided in the embodiment of the present application, the relevant concepts involved in the embodiment of the present application are first explained.

[0053] The original video is the video in which the video character needs to be replaced. In the embodiment of the present application, for the convenience of description, the character to be replaced in the original video is referred to as a character to be replaced. Optionally, the number of the characters to be replaced can be one or more.

[0054] The target video is a video obtained after the original video is replaced with a character. The target video contains a target character, and the target character is a character used to replace the character to be replaced in the original video.

[0055] The background generation model is used to fill in the missing parts of the background image to generate a complete background image with natural picture transition.

[0056] The character generation model is used to generate a character image with the visual features based on the input character skeleton key points and the character's visual features (such as appearance, clothing, etc.), and the character in the character image presents the actions described by the skeleton key points.

[0057] The video frame interpolation model is used to generate a new video frame between two video frames. It is understandable that if the same person in the two video frames input to the video frame interpolation model corresponds to two different actions, then in the new video frame generated by the video frame interpolation model, the action presented by the person is a transition action between the two different actions. Optionally, the number of new video frames generated by the video frame interpolation model can be one or more, which can be customized.

[0058] Optionally, in the embodiment of the present application, the background generation model, the character generation model and the video frame filling model can be implemented as the same or different large models (Large Mode 1, referred to as LM), and different large models can have different parameter scales. In the specific implementation process, they can be flexibly selected based on actual needs. Among them, the large model refers to a deep learning model with a large number of parameters, capable of processing large-scale data and having strong computing power, including but not limited to a large language model (LargeLanguage Mode 1, referred to as LLM).

[0059] In the video processing fields such as film and television production, advertising, and video editing, there is usually a need to replace a specific object (such as a person) in a video. However, when replacing a person in a video, traditional video character replacement solutions are prone to problems such as unnatural transitions between the background image in the video and the replaced person, and unsmooth movements of the replaced person, making it difficult to achieve a natural and realistic replacement effect.

[0060] In view of at least one of the above technical problems, an embodiment of the present application provides a solution, and the basic idea is: first, obtain multiple video frames containing the person to be replaced from the video frame sequence corresponding to the original video. After that, obtain the skeleton key points of the person to be replaced in the first target video frame, and the first background video frame obtained after the person to be replaced is removed from the first target video frame, wherein the first target video frame is any one of the multiple video frames. Next, input the first background video frame into the background generation model so that the background generation model fills the blank area corresponding to the person to be replaced in the first background video frame with the background to generate a second background video frame; input the skeleton key points and the visual information of the target person into the person generation model to generate a target person image that matches the action described by the skeleton key points. Finally, generate a second target video frame based on the second background video frame and the target person image; and generate a target video based on the second target video frames corresponding to the multiple video frames. In this solution, for the first target video frame containing the person to be replaced, a second background video frame with a complete background is generated through the background generation model, and a target person image matching the action of the person to be replaced is generated through the person generation model. Furthermore, the second background video frame is combined with the target task image to generate a second target video frame, so as to achieve a natural transition between the target person and the background image in the second target video frame, and a coordinated presentation of the target person's action and the background image. Thus, the target video generated based on the second target video frame has a natural and realistic replacement effect.

[0061] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0062] The video character replacement method provided in the embodiment of the present application can be performed by an electronic device, which can be a terminal device such as a PC, a laptop, a smart phone, or a server. The server can be a physical server including an independent host, or a virtual server, or a cloud server or a server cluster.

[0063] Figure 1 A flowchart of a method for replacing a video character provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the following steps may be included:

[0064] 101. Acquire multiple video frames containing the person to be replaced from a video frame sequence corresponding to the original video.

[0065] 102. Obtain skeleton key points of the person to be replaced in a first target video frame, and a first background video frame obtained by removing the person to be replaced from the first target video frame, wherein the first target video frame is any one of the multiple video frames.

[0066] 103. Input the first background video frame into the background generation model, so that the background generation model fills the blank area corresponding to the person to be replaced in the first background video frame with the background, so as to generate a second background video frame.

[0067] 104. Input the skeleton key points and the visual information of the target person into the person generation model to generate an image of the target person that matches the action described by the skeleton key points.

[0068] 105. Generate a second target video frame according to the second background video frame and the target person image.

[0069] 106. Generate a target video according to second target video frames corresponding to the multiple video frames.

[0070] The video character replacement method provided in the embodiment of the present application is used to replace the character to be replaced in the original video with the target character. For ease of description, the video corresponding to the character to be replaced in the original video is called the target video.

[0071] To facilitate understanding of the video character replacement method provided in the embodiment of the present application, the following Figure 2 To explain, Figure 2 A schematic diagram of a scene for replacing a character in a video provided in an embodiment of the present application.

[0072] like Figure 2 As shown, the original video is composed of a sequence of video frames. It is easy to understand that in the sequence of video frames of the original video, each video frame may contain the person to be replaced, or only some video frames may contain the person to be replaced.

[0073] In the scenario of replacing a character to be replaced in an original video, first, it is necessary to obtain a plurality of video frames containing the character to be replaced from a video frame sequence corresponding to the original video.

[0074] In an optional embodiment, the person to be replaced may be directly identified on the video frames included in the video frame sequence of the original video, thereby obtaining a plurality of video frames including the person to be replaced.

[0075] However, when the number of video frames included in the video frame sequence is large, performing the identification of the person to be replaced frame by frame may consume a lot of time and hardware resources, and it is difficult to ensure the efficiency of video person replacement. In addition, when the number of multiple video frames containing the person to be replaced identified is large, performing the person replacement frame by frame will also reduce the efficiency of video person replacement.

[0076] Therefore, in another optional embodiment, another method for obtaining multiple video frames containing the person to be replaced from the video frame sequence corresponding to the original video is provided: according to the frame rate of the original video and the frame rate of the target video to be generated, multiple initial video frames are obtained from the video frame sequence corresponding to the original video at preset time intervals; and multiple video frames containing the person to be replaced are obtained from the multiple initial video frames.

[0077] In a specific implementation process, when the frame rate of the target video is less than or equal to the frame rate of the original video, multiple initial video frames are obtained from the video frame sequence corresponding to the original video at a preset first time interval; when the frame rate of the target video is greater than the frame rate of the original video, multiple initial video frames are obtained from the video frame sequence corresponding to the original video at a preset second time interval. The first time interval is greater than the second time interval, and the second time interval is greater than or equal to the time interval between two adjacent video frames in the video frame sequence of the original video. Among the multiple initial video frames, the time interval between any two adjacent initial video frames is the preset time interval (i.e., the first time interval or the second time interval).

[0078] It is understandable that, among the multiple initial video frames obtained from the video frame sequence corresponding to the original video based on the time dimension, there may be video frames that do not contain the person to be replaced. Therefore, further, it is necessary to obtain multiple video frames containing the person to be replaced from the multiple initial video frames.

[0079] In this solution, based on the frame rate of the original video and the frame rate of the target video to be generated, by controlling the time length corresponding to the preset time interval, the number of initial video frames that match the video character replacement requirements can be obtained, and further, the number of acquired video frames containing the characters to be replaced can be controlled. For example, when the frame rate of the target video is high and more video frames containing the characters to be replaced need to be obtained, more initial video frames can be obtained from the video frame sequence corresponding to the original video at a smaller preset time interval; when the frame rate of the target video is low and fewer video frames containing the characters to be replaced need to be obtained, fewer initial video frames can be obtained from the video frame sequence corresponding to the original video at a larger preset time interval.

[0080] For ease of description, in the embodiment of the present application, any one video frame from a plurality of video frames containing a person to be replaced obtained from a video frame sequence corresponding to the original video is referred to as a first target video frame.

[0081] In order to replace the character to be replaced in the first target video frame after obtaining the first target video frame, Figure 2As shown, on the one hand, the skeleton key points of the person to be replaced in the first target video frame are obtained, such as the key points corresponding to the skeleton positions of the head, shoulder, elbow, wrist, hip, knee, ankle, etc. of the person to be replaced, and the skeleton key points of the person to be replaced and the visual information of the target person to be replaced (such as appearance information, clothing information, etc.) are input into the character generation model to generate the target person image. Among them, in the target person image, the posture presented by the target person matches the action described by the skeleton key points of the person to be replaced. In other words, the relative position relationship between the skeleton key points of the target person in the target person image is the same as the relative position relationship between the skeleton key points of the person to be replaced.

[0082] Optionally, the visual information of the target person may be expressed in the form of text or picture.

[0083] On the other hand, a first background video frame is obtained after the person to be replaced is removed from the first target video frame, and the first background video frame is input into a background generation model so that the background generation model fills the blank area corresponding to the person to be replaced in the first background video frame with the background to generate a second background video frame.

[0084] Among them, the background generation model has learned a large amount of image data during the training stage and has the ability to generate and repair images. It can fill in the blank area left by removing the character to be replaced in the first background video frame based on the existing background image information, that is, generate image content that is coordinated with the original background image information in the blank area.

[0085] It is easy to understand that after the person to be replaced is removed from the first target video frame, the area originally corresponding to the person to be replaced in the first target video frame is a blank area. If the target person image used to replace the person to be replaced is directly filled into the blank area, then the target person image may not completely match the blank area, resulting in certain blank areas in the filled image, which is what is referred to as an unnatural transition between the target person and the background image in the relevant technology.

[0086] In an embodiment of the present application, in order to ensure a natural transition between the target person and the background image in the second target video frame obtained after replacing the video person in the first target video frame, the background generation model is used to fill the blank area corresponding to the person to be replaced in the first background video frame to generate a second background video frame. Thereby, the target person image is combined with the second background video frame. In the generated second target video frame, the target person image and the background image have a natural transition, and the target person and the replacement person have the same action.

[0087] After generating second target video frames corresponding to multiple video frames containing the person to be replaced in the above manner, the multiple second target video frames are combined into a video frame sequence in a corresponding order to generate a target video, which is the video person replacement result of the original video.

[0088] In summary, in the embodiment of the present application, for the first target video frame containing the person to be replaced in the original video, a second background video frame with a complete background is generated by the background generation model, and a target person image matching the action of the person to be replaced is generated by the person generation model. Furthermore, the second background video frame is combined with the target task image to generate a second target video frame, so as to achieve a natural transition between the target person and the background image in the second target video frame, and a coordinated presentation of the action of the target person and the background image. Thus, the target video generated based on the second target video frame has a natural and realistic replacement effect.

[0089] Figure 3 A flowchart of another method for replacing a video character provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the following steps may be included:

[0090] 301. Acquire multiple video frames containing a person to be replaced from a video frame sequence corresponding to an original video.

[0091] 302. Obtain character contour information corresponding to the character to be replaced in the first target video frame, identify skeletal key points of the character to be replaced in the first target video frame based on the character contour information, and remove the character to be replaced from the first target video frame to obtain a first background video frame, where the first target video frame is any one of the multiple video frames.

[0092] 303. Input the first background video frame into the background generation model, so that the background generation model fills the blank area corresponding to the person to be replaced in the first background video frame with the background, so as to generate a second background video frame.

[0093] 304. Input the skeleton key points and the visual information of the target person into the person generation model to generate an image of the target person that matches the action described by the skeleton key points.

[0094] 305. Generate a second target video frame according to the second background video frame and the target person image.

[0095] 306. Generate a target video according to second target video frames corresponding to the multiple video frames.

[0096] The specific implementation process of step 301 and step 303 to step 306 may refer to the aforementioned embodiment and will not be described again here.

[0097] As an optional method of obtaining the skeleton key points of the person to be replaced in the first target video frame and the first background video frame, the character contour information corresponding to the person to be replaced in the first target video frame can be first obtained; then, based on the character contour information, the skeleton key points of the person to be replaced in the first target video frame are identified, and the person to be replaced is removed from the first target video frame to obtain the first background video frame.

[0098] Specifically, for the first target video frame, a mask image of the person to be replaced can be generated first, and then the person contour information of the person to be replaced can be determined through a deep learning model such as a semantic segmentation model. Afterwards, based on the person contour information, on the one hand, the position of the skeleton key points of the person to be replaced can be identified through a key point detection algorithm; on the other hand, the person to be replaced can be cut out from the first target video frame based on the person contour information of the person to be replaced.

[0099] In this solution, the accuracy of the skeleton key points of the person to be replaced and the accuracy of the removal of the person to be replaced can be guaranteed through the person outline information. It can be understood that the higher the accuracy of the skeleton key points identified, the higher the matching degree between the target person's actions and the actions of the person to be replaced in the target person image generated based on the skeleton key points; the more accurate the removal of the person to be replaced, the more natural the background filling in the second background video frame generated based on the first background video frame after the removal of the person to be replaced; finally, the more natural the transition between the target person and the background picture based on the target person image and the second background video frame, and the more coordinated the presentation of the target person's actions and the background picture.

[0100] Figure 4 A flowchart of another method for replacing a video character provided in an embodiment of the present application is shown in FIG. Figure 4 As shown, the following steps may be included:

[0101] 401. Acquire multiple video frames containing the person to be replaced from a video frame sequence corresponding to the original video.

[0102] 402. Obtain character contour information corresponding to the character to be replaced in the first target video frame, identify skeletal key points of the character to be replaced in the first target video frame based on the character contour information, and remove the character to be replaced from the first target video frame to obtain a first background video frame, where the first target video frame is any one of the multiple video frames.

[0103] 403. Input the character contour information, the skeleton key points and the visual information of the target character into the character generation model to generate a target character image that matches the action described by the skeleton key points, and the character contour of the target character in the target character image matches the character contour described by the character contour information.

[0104] 404. Input the skeleton key points and the visual information of the target person into the person generation model to generate an image of the target person that matches the action described by the skeleton key points.

[0105] 405. Determine a target area corresponding to the target person image in the second background video frame according to the person outline information, and add the target person image to the target area to generate a second target video frame.

[0106] 406. Generate a target video according to second target video frames corresponding to the multiple video frames.

[0107] The specific implementation process of step 401, step 402, step 404, and step 406 may refer to the aforementioned embodiment and will not be described in detail here.

[0108] In an embodiment of the present application, in order to further enhance the naturalness of the combination of the target person image and the second background image and generate the fusion of the target person and the background image in the second target video frame, when generating the target task image, the person contour information is also used as one of the input information of the person generation model, and is input into the person generation model together with the skeletal key points and the visual information of the target person. Thus, the generated target person image matches the action described by the skeletal key points, and the person contour of the target person in the target person image matches the person contour described by the person contour information.

[0109] The character outline of the target person in the target person image matches the character outline described by the character outline information (i.e., the character outline of the person to be replaced), which can be understood as the degree of overlap between the character outline of the target person and the character outline of the person to be replaced reaches a preset ratio, which can be set by user. By ensuring the consistency of the character outlines of the target person and the person to be replaced, the naturalness of the fusion of the target person and the background image in the generated second target video frame can be effectively improved.

[0110] In an optional embodiment, a second target video frame is generated based on the second background video frame and the target person image, including: determining the target area corresponding to the target person image in the second background video frame based on the person contour information, and adding the target person image to the target area to generate the second target video frame.

[0111] In practical applications, the size of the image area corresponding to the target person in the target person image may be inconsistent with the image area corresponding to the person to be replaced in the first target video frame. Optionally, the size of the target person image can be adjusted according to the person profile information of the person to be replaced, so that the image area corresponding to the target person in the target person image is consistent with the image area corresponding to the person to be replaced in the first target video frame. When adding the target person image to the target area to generate the second target video frame, the position and angle of the target person image can be further adjusted so that the target person image can be better integrated with the second background video frame and naturally integrated into the background picture in the second background video frame.

[0112] Figure 5 A flowchart of another method for replacing a video character provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the following steps may be included:

[0113] 501. Acquire multiple video frames containing a person to be replaced from a video frame sequence corresponding to an original video.

[0114] 502. Obtain skeleton key points of the person to be replaced in a first target video frame, and a first background video frame obtained by removing the person to be replaced from the first target video frame, wherein the first target video frame is any one of the multiple video frames.

[0115] 503. Input the first background video frame into the background generation model, so that the background generation model fills the blank area corresponding to the person to be replaced in the first background video frame with the background, so as to generate a second background video frame.

[0116] 504. Input the skeleton key points and the visual information of the target person into the person generation model to generate an image of the target person that matches the action described by the skeleton key points.

[0117] 505. Generate a second target video frame according to the second background video frame and the target person image.

[0118] 506. Perform frame interpolation processing on adjacent second target video frames to obtain interpolation video frames between adjacent second target video frames, and generate a target video according to the second target video frames and the interpolation video frames respectively corresponding to a plurality of video frames.

[0119] The specific implementation process of step 501 to step 505 may refer to the aforementioned embodiment and will not be described in detail here.

[0120] In practical applications, there are situations where the second target video frames generated based on multiple video frames containing the character to be replaced are insufficient to generate the target video. For example, more video frames are needed to generate the target video, or there are problems with the continuity of the target character's movements in adjacent second target video frames. In this case, it is necessary to perform frame interpolation processing on adjacent second target video frames to obtain interpolated video frames between adjacent second target video frames. Finally, the target video is generated based on the second target video frames and interpolated video frames corresponding to the multiple video frames.

[0121] During the specific implementation process, taking the incoherent movements of the target character in adjacent second target video frames as an example, optionally, in response to determining that the movements of the target character in adjacent second target video frames are incoherent, the adjacent second target video frames can be input into a video frame interpolation model so that the video frame interpolation model outputs interpolation video frames between adjacent second target video frames.

[0122] Optionally, the number of interpolated video frames output by the video frame interpolation model can be custom set.

[0123] In this scheme, the corresponding interpolation video frames are generated through the video interpolation model. On the one hand, a corresponding number of video frames can be generated for the generation of the target video. On the other hand, the action gaps between adjacent second target video frames can be filled to achieve smooth action transition.

[0124] The following will describe in detail one or more embodiments of the video character replacement device of the present application. Those skilled in the art will appreciate that these devices can be configured using commercially available hardware components through the steps taught in this solution.

[0125] Figure 6 A schematic diagram of the structure of a video character replacement device provided in an embodiment of the present application is shown in FIG. Figure 6 As shown, the device includes: an acquisition module 11, a generation module 12, and a processing module 13.

[0126] The acquisition module 11 is used to acquire multiple video frames containing the person to be replaced from the video frame sequence corresponding to the original video; acquire the skeleton key points of the person to be replaced in the first target video frame, and the first background video frame obtained by removing the person to be replaced from the first target video frame, wherein the first target video frame is any one of the multiple video frames;

[0127] The generating module 12 is used to input the first background video frame into the background generating model, so that the background generating model fills the blank area corresponding to the character to be replaced in the first background video frame with the background, so as to generate a second background video frame; input the skeleton key points and the visual information of the target character into the character generating model, so as to generate an image of the target character matching the action described by the skeleton key points;

[0128] The processing module 13 is used to generate a second target video frame according to the second background video frame and the target person image; and generate a target video according to the second target video frames respectively corresponding to the multiple video frames.

[0129] Optionally, the acquisition module 11 is specifically used to acquire multiple initial video frames from a video frame sequence corresponding to the original video at preset time intervals according to the frame rate of the original video and the frame rate of a target video to be generated, wherein the target video is a video corresponding to the original video after character replacement; and acquire multiple video frames containing the characters to be replaced from the multiple initial video frames.

[0130] Optionally, the acquisition module 11 is also specifically used to obtain the character contour information corresponding to the character to be replaced in the first target video frame; based on the character contour information, identify the skeletal key points of the character to be replaced in the first target video frame, and remove the character to be replaced from the first target video frame to obtain a first background video frame.

[0131] Optionally, the acquisition module 11 is also specifically used to input the character contour information, the skeletal key points and the visual information of the target character into a character generation model to generate a target character image that matches the action described by the skeletal key points, and the character contour of the target character in the target character image matches the character contour described by the character contour information.

[0132] Optionally, the processing module 13 is specifically configured to determine, according to the character outline information, a target area corresponding to the target character image in the second background video frame; and add the target character image to the target area to generate a second target video frame.

[0133] Optionally, the processing module 13 is further used to perform frame interpolation processing on adjacent second target video frames to obtain interpolated video frames between adjacent second target video frames; and generate a target video according to the second target video frames and the interpolated video frames respectively corresponding to the multiple video frames.

[0134] Optionally, the processing module 13 is also specifically used to, in response to determining that the movements of the target person in adjacent second target video frames are incoherent, input adjacent second target video frames into a video frame interpolation model so that the video frame interpolation model outputs interpolation video frames between adjacent second target video frames.

[0135] Figure 6 The device shown can execute the steps introduced in the aforementioned embodiments. For detailed execution process and technical effects, please refer to the description in the aforementioned embodiments, which will not be repeated here.

[0136] In one possible design, the above Figure 6 The structure of the video character replacement device shown can be implemented as an electronic device, such as Figure 7 As shown, the electronic device may include: a memory 21, a processor 22, and a communication interface 23. The memory 21 stores a computer program, and when the computer program is executed by the processor 22, the processor 22 can at least implement the video character replacement method provided in the above embodiment.

[0137] The above-mentioned memory 21 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0138] Accordingly, the embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor is enabled to implement each step in the above method embodiment. Among them, the computer-readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (Phase-change Random Access Memory, PRAM), static random access memory (SRAM), dynamic random access memory (Dynamic Random Access Memory, DRAM), other types of random access memory (Random-Access Memory, RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (Digital Video Disc, DVD) or other optical storage, magnetic cassette, tape disk storage or other magnetic storage device or any other non-transmission medium

[0139] Accordingly, the embodiment of the present application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the processor is enabled to implement each step in the above method embodiment. It should be understood that each process or a combination of multiple processes in the above method flow can be implemented by a computer program or instruction. In addition, these computer programs or instructions can be applied to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable video character replacement device, so that the processor of the general-purpose computer, the special-purpose computer, the embedded processor, or other programmable video character replacement device can be implemented as a device to implement the corresponding functions in the above method embodiment.

[0140] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Those of ordinary skill in the art may understand and implement the present invention without creative effort.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by adding a necessary general hardware platform, and of course can also be implemented by combining hardware and software. Based on such an understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a computer product, and the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0142] Finally, it should be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device that includes the elements.

[0143] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for replacing a person in a video, characterized in that: include: Acquire multiple video frames containing the person to be replaced from a video frame sequence corresponding to the original video; Acquire skeleton key points of the character to be replaced in a first target video frame, and a first background video frame obtained by removing the character to be replaced from the first target video frame, wherein the first target video frame is any one of the multiple video frames; Inputting the first background video frame into a background generation model, so that the background generation model performs background filling on a blank area corresponding to the to-be-replaced person in the first background video frame, so as to generate a second background video frame; Inputting the skeleton key points and the visual information of the target person into a person generation model to generate an image of the target person that matches the action described by the skeleton key points; Generate a second target video frame according to the second background video frame and the target person image; Generate a target video according to the second target video frames respectively corresponding to the multiple video frames.

2. The method according to claim 1, characterized in that The step of obtaining a plurality of video frames containing the person to be replaced from a video frame sequence corresponding to the original video includes: According to the frame rate of the original video and the frame rate of the target video to be generated, a plurality of initial video frames are obtained from a video frame sequence corresponding to the original video at a preset time interval, wherein the target video is a video corresponding to the original video after the character is replaced; A plurality of video frames including the person to be replaced are obtained from the plurality of initial video frames.

3. The method according to claim 1, characterized in that: The step of obtaining the skeleton key points of the person to be replaced in the first target video frame and the first background video frame obtained by removing the person to be replaced from the first target video frame includes: Acquire character outline information corresponding to the character to be replaced in the first target video frame; According to the character outline information, the skeleton key points of the character to be replaced in the first target video frame are identified, and the character to be replaced is removed from the first target video frame to obtain a first background video frame.

4. The method according to claim 3, characterized in that The step of inputting the skeleton key points and the visual information of the target person into a person generation model to generate an image of the target person matching the action described by the skeleton key points comprises: The character contour information, the skeleton key points and the visual information of the target character are input into a character generation model to generate a target character image that matches the action described by the skeleton key points, wherein the character contour of the target character in the target character image matches the character contour described by the character contour information.

5. The method according to claim 4, characterized in that The step of generating a second target video frame according to the second background video frame and the target person image includes: Determining a target area corresponding to the target person image in the second background video frame according to the person outline information; The target person image is added to the target area to generate a second target video frame.

6. The method according to any one of claims 1 to 5, characterized in that The generating a target video according to the second target video frames respectively corresponding to the plurality of video frames comprises: Performing frame interpolation processing on adjacent second target video frames to obtain interpolation video frames between adjacent second target video frames; A target video is generated according to the second target video frames and the interpolated video frames respectively corresponding to the multiple video frames.

7. The method according to claim 6, characterized in that The performing frame interpolation processing on adjacent second target video frames to obtain interpolated frame video frames between adjacent second target video frames includes: In response to determining that the movements of the target person in the adjacent second target video frames are incoherent, the adjacent second target video frames are input into a video frame interpolation model so that the video frame interpolation model outputs interpolation video frames between the adjacent second target video frames.

8. A video character replacement device, characterized in that: include: An acquisition module, used for acquiring a plurality of video frames containing the person to be replaced from a video frame sequence corresponding to the original video; Acquire skeleton key points of the character to be replaced in a first target video frame, and a first background video frame obtained by removing the character to be replaced from the first target video frame, wherein the first target video frame is any one of the multiple video frames; A generation module is used to input the first background video frame into a background generation model so that the background generation model fills the blank area corresponding to the to-be-replaced person in the first background video frame with the background to generate a second background video frame; input the skeleton key points and the visual information of the target person into the person generation model to generate an image of the target person that matches the action described by the skeleton key points; The processing module is used to generate a second target video frame according to the second background video frame and the target person image; and generate a target video according to the second target video frames respectively corresponding to the multiple video frames.

9. An electronic device, characterized in that: include: A memory, a processor, and a communication interface; wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the video character replacement method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, the processor is caused to execute the video character replacement method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Video person changing method and device and storage medium

    CN121194004A

  • Video face swapping method, device and storage medium

    CN121194004B

  • Video content replacement method and system

    CN121665084A

  • A method and system for video content replacement

    CN121665084B