Video processing method and related device

By using the characteristic portrait of the target person as a reference during the video style redrawing process and adjusting the redrawing of its characteristic areas, the problem of inconsistent character features was solved, and the quality of the redrawn video was improved.

WO2025260722A1PCT designated stage Publication Date: 2025-12-26HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/071484
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-20
Filing Date
2025-01-09
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing technologies cannot maintain consistency in character features during video style redrawing, resulting in low-quality redrawn videos and a poor user viewing experience.

Method used

The feature profile of the target person is obtained from the feature database and input into the redraw model along with the feature region of the target person. As a reference, the redrawing process of the feature region is adjusted to ensure consistency in different video frames.

Benefits of technology

It improves the consistency of the feature regions of the target person in the redrawn video, thereby enhancing the quality of the redrawn video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025071484_26122025_PF_FP_ABST
    Figure CN2025071484_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides a video processing method and a related device, used for improving the consistency of a target person in a video that has undergone style transfer, and improving the quality of a redrawn video. The method comprises: acquiring a video to be processed, the video to be processed comprising a target person; acquiring a feature portrait of the target person from a person feature library, the feature portrait of the target person corresponding to a feature region of the target person, and the person feature library comprising a feature portrait of at least one person; and inputting the feature portrait of the target person and the feature region of the target person in the video to be processed into a first redrawing model to obtain a first redrawn video.
Need to check novelty before this filing date? Find Prior Art

Description

Method and related device for video processing

[0001] The present application claims priority to the Chinese Patent Application No. 202410799827.7, filed on June 19, 2024, entitled "Video Redrawing Method and Device for Keeping Consistency of Object ID in Different Scenes", and to the Chinese Patent Application No. 202411147449.0, filed on August 20, 2024, entitled "Method and Related Device for Video Processing", both of which are incorporated by reference in their entirety. TECHNICAL FIELD

[0002] The present application relates to the field of communication, and in particular to a method and related device for video processing. BACKGROUND

[0003] With the development of artificial intelligence (AI) technology, the way of video processing is becoming more diversified. Among them, the style redrawing of video has attracted widespread attention. The style redrawing of video, which can be referred to as video redrawing, is a technology for generating a new video by changing the style of a character in an original video. The character image in the new video is different in style from the character image in the original video, for example, the character image in the original video is a real task, and the character image in the new video is a virtual image, such as an animation image, a simulation image, etc.

[0004] In the related technical solution, the AI model is used to redraw the video input by the user, thereby changing the style of the video. In this technical solution, since the character in the video can be in a dynamic state, the state of the character before and after the video frame changes, so that the AI tool cannot keep the consistency of the character features before and after the video is redrew, resulting in low quality of the redrew video and poor user viewing experience. SUMMARY

[0005] The present application provides a method and related device for video processing. In the video processing method, a feature portrait of a target character is obtained from a character feature library, and the feature portrait corresponds to a feature region of the target character. The feature portrait of the target character and the feature region of the target character are input into a first redrawing model, and the feature portrait of the target character is used as a reference for the feature region to adjust the feature region of the target character during the redrawing process, thereby providing a more accurate reference for the redrawing of the feature region. Therefore, even in different video frames, the redrawing of the feature region of the same target character is referenced to the same feature portrait, thereby improving the consistency of the feature region of the target character before and after the first redrawing video output by the first redrawing model, and improving the quality of the first redrawing video.

[0006] In a first aspect, a method for video processing is provided, comprising:

[0007] The video processing apparatus obtains a to-be-processed video, the to-be-processed video comprising a target person. That is, the to-be-processed video comprises an image of the target person in a video frame. In this application, the to-be-processed video can also be referred to as a video to be redrawn. The video processing apparatus obtains a feature portrait of the target person from a person feature library, the feature portrait of the target person corresponding to a feature region of the target person, the person feature library comprising feature portraits of at least one person. The feature region of the target person is also the region that needs to be redrawn. The video processing apparatus inputs the feature portrait of the target person and the feature region of the target person in the to-be-processed video into a first redrawing model to obtain a first redrawn video. Specifically, the foregoing process can be understood as adjusting the feature region of the target person with the feature portrait of the target person as a reference to obtain the first redrawn video. In the first redrawn video, the style of the feature region of the target person is different from that of the feature region of the target person in the to-be-processed video.

[0008] In this application, the feature portrait of the target person is obtained from the person feature library, which corresponds to the feature region of the target person. The feature portrait of the target person and the feature region of the target person are input into the first redrawing model, and the feature portrait of the target person is used as a reference for the feature region of the target person in the redrawing process to provide a more accurate reference for the redrawing of the feature region. Therefore, even in different video frames, the redrawing of the feature region of the same target person is always based on the same feature portrait, thereby improving the consistency of the feature region of the target person in the first redrawn video output by the first redrawing model and improving the quality of the first redrawn video.

[0009] In some optional implementations of the first aspect, the input of the first redrawing model can further include a target redrawing style in addition to the foregoing feature portrait of the target person and the feature region of the target person in the to-be-processed video. The target redrawing style is used to indicate the style of the feature region of the target person in the first redrawn video. It can also be understood that the target redrawing style indicates the style of the target person to be redrawn in the to-be-processed video. The target redrawing style can be input into the first redrawing model at the same time as the feature portrait of the target person and the feature region of the target person in the to-be-processed video, or before or after the feature portrait of the target person and the feature region of the target person in the to-be-processed video, which is not limited herein.

[0010] In the present application, the input of the first redraw model further includes a target redraw style, which is used to indicate the style of the feature region of the target character to be redrawn, that is, the target character feature region is to be redrawn into which style. The target redraw style can be selected according to the actual application needs, personalized services are provided for users, and the flexibility of the technical solution of the present application is improved.

[0011] In some optional implementations of the first aspect, the target redraw style includes a two-dimensional image style and / or a three-dimensional image style. The two-dimensional image style includes a pixel style, a painting style, and the like. The painting style includes a comic style, an ink painting style, an oil painting style, a meticulous painting style, a print style, a gouache painting style, a sketch style, and the like. The three-dimensional image style includes a clay style, a paper folding style, a building block style, a doll style, a ceramic style, a bronze style, and the like.

[0012] In the present application, the target redraw style has multiple possibilities and can be flexibly selected, further enriching the application scenarios and implementation manners of the technical solution of the present application, and improving the flexibility of the technical solution of the present application.

[0013] In some optional implementations of the first aspect, the video processing apparatus can determine the target character based on the to-be-processed video. Specifically, after obtaining the to-be-processed video, the video processing apparatus can further determine the target character in response to a first operation instruction for a target video frame in the to-be-processed video. The first operation instruction is used to specify the target character included in the target video frame. That is, after the video processing apparatus obtains the to-be-processed video, the target video frame is displayed. The user selects the target character in the target video frame based on a touch operation or through an external device. The video processing apparatus further stores the feature portrait of the target character in the character feature library, and the feature portrait of the target character is contained in the first video frame included in the to-be-processed video. The confidence of the target character in the first image is higher than the confidence of the target character in other video frames.

[0014] In the present application, the video processing apparatus can determine the target character based on the to-be-processed video, which also indicates that the data processed by the video processing apparatus is derived from the to-be-processed video, and there is no need to obtain other content, which is simple to operate. In addition, the feature portrait of the target character is contained in the first image included in the to-be-processed video, and the confidence of the target character in the first image is higher than the confidence of the target character in other video frames. Therefore, when the feature region of the target character is redrawn, the reference of the feature portrait of the target character is more accurate, the similarity of the feature region of the target character after the redraw and before the redraw is improved, and the redraw quality of the feature region of the target character is further improved.

[0015] In some optional implementation of the first aspect, the video processing apparatus can further determine the target person in other manners. Specifically, the video processing apparatus further acquires a target image, and in response to a second operation instruction for the target image, determines the target person, the second operation instruction being used to specify a target object included in the target image. That is, the target person is a person selected in the target image. The video processing apparatus stores the feature image of the target person into the person feature library, the feature image of the target person being contained in the second image, the confidence of the target person in the second image being higher than the confidence of the target person in other images, the other images being the plurality of video frames included in the video to be processed and the images in the target image except for the second image.

[0016] In the present application, the video processing apparatus can further determine the target person from the target image, which enriches the application scenarios of the technical solution of the present application. In addition, the feature image of the target person is contained in the second image, and the confidence of the target person in the second image is higher than the confidence of the target person in other images, the other images being the plurality of video frames included in the video to be processed and the images in the target image except for the second image. This means that the feature image of the target person is contained in the image in which the confidence of the target person is the highest among the plurality of video frames included in the video to be processed and the images in the target image. Therefore, when the feature region of the target person is redrawn, the reference of the feature image of the target person is more accurate, which improves the similarity between the feature region of the target person after redrawing and the feature region of the target person before redrawing, and further improves the redrawing quality of the feature region of the target person.

[0017] In some optional implementation of the first aspect, the person feature library further includes an identifier of at least one person, the identifier of the at least one person corresponding to the feature image of the at least one person in a one-to-one manner. Then, the video processing apparatus acquires the feature image of the target person from the person feature library, specifically based on the identifier of the target person. Specifically, after the video processing apparatus determines the target person, the video processing apparatus can acquire the identifier of the target person, determine the identifier in the person feature library that matches the identifier of the target person, find the feature image of the person corresponding to the identifier, and thus find the feature image of the target person.

[0018] In the present application, in the person feature library, the identifier of a person corresponds to the feature image of the person in a one-to-one manner. After the identifier of the target person is determined, the feature image of the target person can be uniquely determined, which improves the efficiency of the technical solution of the present application.

[0019] In some optional implementation of the first aspect, the video to be processed can be an initial video. Then, the video processing apparatus can further input the video to be processed into the second redraw model to obtain a second redrawn video. The style of the second redrawn video can be the same as or different from the target redrawn style, which is not limited here. The video processing apparatus can further splice the first redrawn video and the second redrawn video to obtain a third redrawn video. That is, each video frame in the third redrawn video is redrawn compared with the video to be processed.

[0020] In some optional implementation of the first aspect, the input of the second redraw model includes the first redrawn style in addition to the video to be processed, which is used to indicate the style of the video to be processed. The first redrawn style can be input into the second redraw model at the same time as the video to be processed, or before or after the video to be processed, which is not limited here. The first redrawn style is similar to the target redrawn style, which has multiple possibilities and will not be repeated here. The first redrawn style can be the same as or different from the target redrawn style, which is not limited here.

[0021] In some optional implementation of the first aspect, the video to be processed can be a redrawn video. That is, the video processing apparatus obtains the video to be processed by inputting an initial video into the second redraw model. The video processing apparatus further splices the first redrawn video and the video to be processed to obtain a fourth redrawn video.

[0022] In some optional implementation of the first aspect, when the initial video is input into the second redraw model, the second redrawn style is also input into the second redraw model, which is used to indicate the style of the initial video to be redrawn. The second redrawn style can be input into the second redraw model at the same time as the initial video, or before or after the initial video, which is not limited here. The second redrawn style is similar to the target redrawn style, which has multiple possibilities and will not be repeated here. The second redrawn style can be the same as or different from the target redrawn style, which is not limited here.

[0023] In the present application, the video to be processed has multiple possibilities, and the feature region of the target person obtained from the video to be processed has multiple possibilities, which enriches the implementation modes and application scenarios of the technical scheme of the present application and further improves the flexibility of the technical scheme of the present application. In addition, the video processing apparatus splices the first redrawn video and the redrawn video (i.e., the second redrawn video or the third redrawn video) obtained by redrawing other regions in the video to be processed or the initial video, so that each video frame in the spliced video is a redrawn video frame, that is, the entire video is redrawn.

[0024] In some optional implementation forms of the first aspect, after obtaining the video to be processed, the video processing apparatus further splits the video to be processed into a plurality of video frames. Then starting from a first frame of the plurality of video frames, all persons included in the plurality of video frames are determined based on the person feature detection algorithm. Feature portraits of all the persons are stored in the person feature library, wherein each feature portrait of each person is contained in each third image of the video to be processed, and a confidence of each person in each third image is higher than a confidence of each person in other video frames. For example, assuming that the video to be processed has 600 frames in total, wherein the first frame to the 400th frame all include a person 1, and the 200th frame to the 600th frame all include a person 2. If the confidence of the person 1 in the 300th frame is higher than the confidence of the person 1 in other video frames from the first frame to the 400th frame, the third image of the person 1 is the 300th frame. If the confidence of the person 2 in the 505th frame is higher than the confidence of the person 2 in other video frames from the 200th frame to the 600th frame, the third image of the person 2 is the 505th frame.

[0025] In the present application, the person feature library can also include the feature portrait of each person in the video to be processed. Therefore, when the style of a person other than the target person is redrawn, the corresponding feature portrait can be found in the person feature library, thereby simplifying the operation process.

[0026] In some optional implementation forms of the first aspect, the feature region of the target person includes a facial feature region of the target person and / or a clothing feature region of the target person.

[0027] In the present application, the feature region of the target person has multiple possibilities, which enriches the application scenarios of the technical solution of the present application and further improves the flexibility of the technical solution of the present application.

[0028] In a second aspect, the present application provides a video processing apparatus, comprising:

[0029] An obtaining unit is configured to obtain a video to be processed, wherein the video to be processed includes a target person. The obtaining unit is further configured to obtain a feature portrait of the target person from a person feature library, wherein the feature portrait of the target person corresponds to a feature region of the target person, and the person feature library includes feature portraits of at least one person.

[0030] A processing unit is configured to input the feature portrait of the target person and the feature region of the target person in the video to be processed into a first redrawing model to obtain a first redrawing video.

[0031] The video processing apparatus is configured to implement the first aspect or any possible implementation form of the first aspect, which will not be described herein again.

[0032] In a third aspect, the present application provides a video processing apparatus, comprising a processor and a memory. The processor stores instructions which, when executed by the processor, implement the method of the first aspect or any possible implementation of the first aspect.

[0033] In a fourth aspect, the present application provides a computing device, comprising a processor and a memory. The processor of the computing device is configured to execute instructions stored in the memory, so that the computing device implements the method of the first aspect or any possible implementation of the first aspect.

[0034] In a fifth aspect, the present application provides a computing device cluster, comprising at least one computing device, each computing device comprising a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster implements the method of the first aspect or any possible implementation of the first aspect.

[0035] In a sixth aspect, the present application provides a computer program product comprising instructions which, when executed by a processor, implement the method of the first aspect or any possible implementation of the first aspect, or, when executed by a computing device cluster, cause the computing device cluster to implement the method of the first aspect or any possible implementation of the first aspect.

[0036] In a seventh aspect, the present application provides a computer-readable storage medium having computer program instructions stored therein, which, when executed by a processor, implement the method of the first aspect or any possible implementation of the first aspect, or, when executed by a computing device cluster, cause the computing device cluster to implement the method of the first aspect or any possible implementation of the first aspect.

[0037] The advantageous effects of any one of the second aspect to the seventh aspect are similar to those of the first aspect or any possible implementation of the first aspect, and will not be described here. BRIEF DESCRIPTION OF DRAWINGS

[0038] FIG. 1 is a schematic diagram of a system architecture provided by an embodiment of the present application;

[0039] FIG. 2 is another schematic diagram of a system architecture provided by an embodiment of the present application;

[0040] FIG. 3 is a flow diagram of a video processing method provided by an embodiment of the present application;

[0041] FIG. 4 is another flow diagram of a video processing method provided by an embodiment of the present application;

[0042] FIG. 5 is another flow diagram of a video processing method according to an embodiment of the present application;

[0043] FIG. 6 is a structural diagram of a video processing apparatus according to an embodiment of the present application;

[0044] FIG. 7 is a structural diagram of a computing device according to an embodiment of the present application;

[0045] FIG. 8 is a structural diagram of a computing device cluster according to an embodiment of the present application;

[0046] FIG. 9 is another structural diagram of a computing device cluster according to an embodiment of the present application. DETAILED DESCRIPTION

[0047] The embodiments of the present application provide a video processing method and related devices. In the video processing method, a feature portrait of a target person is obtained from a feature portrait library, and the feature portrait corresponds to a feature region of the target person. The feature portrait of the target person and the feature region of the target person are input into a first redrawing model, and the feature portrait of the target person is used as a reference for the feature region to adjust the feature region of the target person in a redrawing process, so as to provide a more accurate reference for redrawing of the feature region. Therefore, even in different video frames, the redrawing of the feature region of the same target person is based on the same feature portrait, so that the consistency of the feature region of the target person in a first redrawing video output by the first redrawing model is improved, and the quality of the first redrawing video is improved.

[0048] The embodiments of the present application are described below with reference to the drawings. Those skilled in the art can know that, as technology develops and new scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0049] The terms "first", "second", and the like in the description and in the claims of the present application and above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the terms used in this way can be interchanged, as appropriate, and are merely a way of distinguishing between objects of the same attribute in the description of the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device that includes a list of elements does not necessarily limit to those elements, but can include other elements not clearly listed or inherent to such a process, method, product or device. In addition, "one or more" means one or more, and "multiple" means two or more. "And / or" describes the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or the like means any combination of the items, including any combination of single or multiple items. For example, at least one of a, b, or c can represent: a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.

[0050] First, refer to FIG. 1 and FIG. 2, both of which are system architecture diagrams provided by embodiments of the present application.

[0051] In the embodiment shown in FIG. 1, an application program is running on the client 101, which is used to implement the video processing method provided by the embodiments of the present application. That is, the user selects to start the application program, inputs a video to the client 101, the application program processes the video, obtains a redrawn video, and displays the redrawn video on the screen of the client 101. Optionally, the computing power of the application program can be provided by the server 102.

[0052] The video processing method provided by the embodiments of the present application can also be applied to a cloud scenario. In the cloud scenario, the AI model capability used for video processing can be deployed in the cloud, and the high computing power of the cloud platform can be used to process long videos better.

[0053] As shown in FIG. 2, a user logs in to the cloud service platform 203 via the Internet 202 through the client 201 by an account and a password registered in the cloud service platform 203. The cloud service platform 203 manages the infrastructure, which includes multiple data centers arranged in different regions, for example, region 1 includes cloud data center 1 and cloud data center 2, and region 2 includes cloud data center 3 and cloud data center 4 as shown in FIG. 1. Each cloud data center is provided with multiple servers on which business instances (including at least one of virtual machines, containers, and dedicated hosts) are operated.

[0054] In the embodiments of the present application, a video service is deployed in the business instance, and a user purchases a cloud service in the cloud service platform 203 through the client 201. The user sends a calling request to the cloud service platform 203, and the calling request is used to request the cloud service from the cloud service platform 203. The specific content of the cloud service includes providing a video processing service introduced below for tenants and the like.

[0055] Next, refer to FIG. 3, which is a flowchart of a video processing method provided by the embodiments of the present application.

[0056] 301. Obtain a to-be-processed video, the to-be-processed video including a target person.

[0057] The to-be-processed video obtained by the video processing apparatus has multiple possibilities, which can be an initial video or a video after redrawing. Regardless of which possibility, the to-be-processed video includes a target person, which means that the to-be-processed video includes an image of the target person in at least two video frames.

[0058] In the scheme in which the to-be-processed video is an initial video, the video processing apparatus obtains the to-be-processed video in multiple possible ways, which can be obtained through an external device, obtained from internal storage or a database, or obtained through other ways, which are not limited here. Optionally, the to-be-processed video is obtained through an external device in multiple possible ways, for example, the to-be-processed video stored in the external device is stored in the video processing apparatus through a transmission medium, or the to-be-processed video uploaded by a user is received. The initial video refers to a video including a real person, or a video taken of a real person.

[0059] In the scheme in which the to-be-processed video is a video after redrawing, the video processing apparatus inputs the initial video to the second redrawing model to obtain the to-be-processed video after redrawing. The style of the to-be-processed video is different from that of the initial video. For example, the initial video is a video of a real person, and the image of the person in the to-be-processed video is a redrawing style, such as a two-dimensional portrait style or a three-dimensional portrait style.

[0060] 302. Obtain a feature image of the target person from a person feature library, the feature image of the target person corresponding to a feature region of the target person, the person feature library comprising feature images of at least one person.

[0061] The feature image of each person in the person feature library corresponds to a feature region of each person, that is, the feature image of the person is an image of the feature region of the person. For example, the feature image of person 1 is an image of the feature region of person 1. In the embodiments of the present application, the feature region of the person includes the face region of the person and / or the clothing region of the person. That is, the feature region of the person includes any one of the following: the face region of the person, the clothing region of the person, the face region of the person and the clothing region of the person. Similarly, the feature region of the target person also includes the face feature region of the target person and / or the clothing feature region of the target person.

[0062] Correspondingly, the feature image of the person includes an image of the face region of the person and / or the clothing region of the person.

[0063] It should be noted that the face region of the person can include the face part of the person, specifically including the eyebrows, glasses, nose and mouth of the person. Optionally, on the basis of the foregoing face part, the face region of the person can also include the ears of the person.

[0064] In the embodiments of the present application, the feature region of the target person has multiple possibilities, which enriches the application scenarios of the technical solutions of the present application and further improves the flexibility of the technical solutions of the present application.

[0065] In addition to storing the feature images of at least one person in the person feature library, the person feature library also stores the identities of at least one person. The identities of the at least one person correspond one-to-one to the feature images of the at least one person. That is, the identity of a person uniquely corresponds to the feature image of the person, and the feature image of the person uniquely corresponds to the identity of the person.

[0066] Therefore, the video processing apparatus obtains the feature image of the target person from the person feature library, specifically by determining the feature image of the target person from the feature images of the at least one person in the person feature library according to the identity of the target person. The specific implementation process is described in detail below, which is not expanded here.

[0067] In the embodiments of the present application, the identities of the persons in the person feature library correspond one-to-one to the feature images of the persons, so after the identity of the target person is determined, the feature image of the target person can be uniquely determined, which improves the efficiency of the technical solutions of the present application.

[0068] 303. Input the feature image of the target person and the feature region of the target person in the to-be-processed video into the first redrawing model to obtain a first redrawing video.

[0069] The video processing apparatus inputs the feature image of the target person and the feature region of the target person in the video to be processed into a first redrawing model, and the first redrawing model is used to redraw the feature region of the target person. It can also be understood that the feature image of the target person is used as a reference to adjust the feature region of the target person, and finally a first redrawing video is obtained.

[0070] When the video processing apparatus redraws the feature region of the target person using the first redrawing model, it also needs to obtain a target redrawing style. The target redrawing style indicates the style of the feature region of the target person, that is, the style of the feature region of the target person to be redrawn, or the style of the feature region of the target person in the first redrawing video. Then, it means that the input of the first redrawing model includes the target redrawing style in addition to the feature image of the target person and the feature region of the target person in the video to be processed. The target redrawing style can be input into the first redrawing model at the same time as the feature image of the target person and the feature region of the target person in the video to be processed, or before or after the feature image of the target person and the feature region of the target person in the video to be processed. The specific place is not limited here.

[0071] The video processing apparatus can obtain the target redrawing style in various ways, such as in response to an operation instruction from the user, or by obtaining a default redrawing style, or by other means, such as by default using the redrawing style for the initial video as the target redrawing style. The specific place is not limited here.

[0072] In the embodiments of the present application, the target redrawing style includes various possibilities, which can be summarized as two-dimensional image style and / or three-dimensional image style. The two-dimensional image style refers to a two-dimensional image style, including pixel style, painting style, etc. The painting style includes animation style, cartoon style, ink painting style, oil painting style, detailed painting style, print style, watercolor painting style, sketch style, etc. The three-dimensional image style refers to a three-dimensional image style, including clay style, origami style, building block style, doll style, ceramic style, bronze style, etc.

[0073] In addition, it should be noted that in the embodiments of the present application, the target person can be one or more. In the case of multiple target persons, the target redrawing style of each target person can be the same or different, and the specific place is not limited here. In the case of different redrawing styles for multiple target persons, the video processing apparatus obtains the target redrawing style corresponding to each target person.

[0074] In the embodiment of the present application, the input of the first redrawing model further includes a target redrawing style, which is used to indicate the style of the feature region of the target character to be redrawn. Since there are multiple possible target redrawing styles, it means that the target character feature region can be redrawn into any style according to the actual application needs, thereby providing personalized services for users and improving the flexibility of the technical solution of the present application.

[0075] As can be known from the foregoing description of FIG. 3, in the embodiment of the present application, the feature image of the target character is obtained from the character feature library, and the feature image corresponds to the feature region of the target character. The feature image of the target character and the feature region of the target character are input into the first redrawing model, and the feature image of the target character is used as a reference for the feature region to adjust the feature region of the target character during the redrawing process, thereby providing a more accurate reference for the redrawing of the feature region. Therefore, even in different video frames, the redrawing of the feature region of the same target character is all based on the same feature image, thereby improving the consistency of the feature region of the target character in the first redrawing video output by the first redrawing model, and improving the quality of the first redrawing video.

[0076] Next, the process of determining the target character by the video processing apparatus will be described in detail. The target character can be understood as a character that needs to be continuously tracked in the video to be processed, or a character whose consistency is focused on during video redrawing. In the embodiment of the present application, there are multiple possible ways for the video processing apparatus to determine the target character, which will be described as follows:

[0077] In some optional embodiments, the target character can be determined from the video to be processed. In this technical solution, the target character can be specified by a user or set by a system, which will be described as follows.

[0078] Optionally, after obtaining the video to be processed, the video processing apparatus determines the target character in response to a first operation instruction for a target video frame in the video to be processed, and the first operation instruction is used to specify the target character included in the target video frame. The target video frame is any image frame including the target character in the video to be processed.

[0079] Optionally, the target video frame can be a video frame selected by a user or a video frame set by default. The video frame set by default has multiple possibilities, which can be a video frame including the most characters, a video frame including characters of the same gender, or other types of video frames, such as a video frame including characters all older or all younger than a threshold, etc. The default setting can be set according to the actual application needs, and the specific implementation is not limited herein.

[0080] The video processing apparatus can determine the target person in response to a first operation instruction for a target video frame in the video to be processed. In other words, the video processing apparatus displays the target video frame, and a user performs a touch operation on the display screen or inputs the first operation instruction to the video processing apparatus through an external device (e.g., a mouse, a keyboard, etc.) so that the video processing apparatus determines the target person.

[0081] Optionally, after obtaining the video to be processed, the video processing apparatus can determine the target person based on a default setting. The default setting can be all persons in the video to be processed, or persons of a female gender in the video to be processed, or persons with an age greater than or less than a threshold in the video to be processed, etc. The default setting can have various possibilities, and can be set according to actual application requirements, and is not limited herein.

[0082] It should be noted that the number of target persons can be one or more, and is not limited herein. For example, when a user selects one person, the video processing apparatus can display a prompt box to prompt the user whether to continue to select persons, until the user no longer selects new persons, and the video processing apparatus determines the number of target persons.

[0083] After determining the target person, the video processing apparatus stores a feature portrait of the target person in a person feature library. The feature portrait of the target person is included in a first image included in the video to be processed, and the confidence of the target person in the first image is higher than the confidence of the target person in other video frames. For example, assume that the video to be processed includes 600 video frames, and the target person is included in 400 video frames. The confidence of the target person in the 300th frame is higher than the confidence of the target person in the remaining 399 frames. In this case, the 300th frame is the first image, and the feature portrait of the target person is a portrait of a feature region of the target person in the first image.

[0084] It can be understood that the feature portrait of the target person is included in the first image included in the video to be processed, and the confidence of the target person in the first image is higher than the confidence of the target person in other video frames. This means that the target person in the video frame in the video to be processed is closest to the appearance of the target person itself, and therefore, the feature portrait of the target person provides a more accurate reference when the video is redrawn.

[0085] It should be noted that in the case of multiple target persons, the feature image of each target person is determined from the video frames including each target person. For example, assume that the target persons include person 1 and person 2, and the video frames to be processed include 600 video frames, 500 of which include person 1 and 300 of which include person 2. Then, the feature image of person 1 is determined from the 500 video frames including person 1, and the feature image of person 2 is determined from the 300 video frames including person 2.

[0086] It should also be noted that the feature region of the target person includes the face region and / or the clothing region of the target person. Accordingly, the feature image of the target person can be the image of the face region of the target person, the image of the clothing region of the target person, or the image of the face region and the clothing region of the target person. Alternatively, the feature image of the face region and the clothing region of the target person can be directly used as the feature image of the target person, so that the feature image of the target person can be used as a reference to provide prior experience regardless of the feature region of the target person.

[0087] In the embodiments of the present application, the video processing apparatus can determine the target person based on the video to be processed, which indicates that the data processed by the video processing apparatus is derived from the video to be processed, and no other content needs to be acquired, which is simple to operate. In addition, the feature image of the target person is included in the first image included in the video to be processed, and the confidence of the target person in the first image is higher than that of the target person in other video frames. Therefore, the reference of the feature image of the target person is more accurate when the feature region of the target person is redrawn, which improves the similarity of the feature region of the target person before and after redrawing, and further improves the redrawing quality of the feature region of the target person.

[0088] In some optional embodiments, the target person can be determined from a target image. Specifically, the video processing apparatus acquires not only the video to be processed but also a target image. Then, in response to a second operation instruction for the target image, the target person is determined, and the second operation instruction is used to specify the target person included in the target image. The feature image of the target person is stored in the person feature library, and the feature image of the target person is included in the multiple video frames included in the video to be processed and the target image, and the image with the highest confidence of the target person.

[0089] The target image can be an image of a real person or a redrawn image, which is not limited here.

[0090] The video processing apparatus determines the target person in response to a second operation instruction for the target image. It can be understood that the video processing apparatus displays the target image, and a user performs a touch operation on the display screen or inputs the second operation instruction to the video processing apparatus through an external device (such as a mouse, a keyboard, etc.), so that the video processing apparatus determines the target person. Optionally, the target person can be part or all of the persons in the target image, that is, the number of target persons can be one or more. For example, when a user selects a person, the video processing apparatus can display a prompt box to prompt the user whether to continue to select a person, until the user no longer selects a new person, and the video processing apparatus determines the number of target persons.

[0091] After the video processing apparatus determines the target person, the video processing apparatus also stores a feature image of the target person in the person feature library. The feature image of the target person is contained in the second image, and the confidence of the target person in the second image is higher than that of the target person in other images. The other images refer to the target image and all video frames containing the target person in the to-be-processed video except the second image. For example, it is assumed that the to-be-processed video includes 600 video frames, 400 video frames all include the target person, and the confidence of the target person in the 305th video frame is higher than that of the target person in the other 399 video frames and that of the target person in the target image. Then, the 305th video frame is the first image, the feature image of the target person, and the image of the feature region of the target person in the first image.

[0092] In addition, it should be noted that in the case where the target person is multiple, the feature image of each person is determined based on the confidence of each person. For example, it is assumed that the target person includes person 3 and person 4, and the to-be-processed video frame includes 100 video frames, 50 video frames include person 3, and 40 video frames include person 4. Then, the feature image of person 3 is determined from the target image and the 50 video frames including person 3; and the feature image of person 4 is determined from the target image and the 40 video frames including person 4.

[0093] In the embodiment of the present application, the video processing apparatus can also determine the target person from the target image, which enriches the application scenarios of the technical scheme of the present application. In addition, the feature image of the target person is contained in the first image, and the confidence of the target person in the first image is higher than that of the target person in other images. The other images refer to the target image and all video frames containing the target person in the to-be-processed video except the first image. Therefore, when the feature region of the target person is redrawn, the reference of the feature image of the target person is more accurate, the similarity of the feature region of the target person after redrawing is improved, and the redrawing quality of the feature region of the target person is further improved.

[0094] In the embodiments of the present application, the target person can be one or more. When the target person is more than one, the target redraw style of different target persons can be the same or different. The manner of determining the target redraw style of the target person is described below.

[0095] In some optional embodiments, after determining a target person, the video processing apparatus can display a candidate box near the display area of the target person, and the candidate box includes at least one candidate redraw style. The user can slide the options in the candidate box to browse the candidate redraw styles provided by the video processing apparatus. The video processing apparatus determines the candidate redraw style as the target redraw style corresponding to the target person in response to a hit operation on the candidate redraw style.

[0096] In some optional embodiments, after determining all target persons, the video processing apparatus displays a redraw style matching interface, which includes a person selection area and a redraw style selection area. The person selection area displays candidate persons, and the user can slide the controls of the person selection area to browse all target persons. The redraw style selection area displays candidate redraw styles, and the user can slide the controls of the redraw style selection area to browse all candidate redraw styles. The video processing apparatus selects a target redraw style for a target person in response to a hit operation on the target person in the person selection area. On this basis, the video processing apparatus determines the redraw style as the target redraw style corresponding to the target person in response to a hit operation on a redraw style in the redraw style selection area.

[0097] In some optional embodiments, the target redraw style of the target person can also be input by the user.

[0098] In the foregoing embodiments, the redraw of the feature area of the target person is emphasized. In actual applications, other areas in the initial video can also be redrawn, which are described below.

[0099] In the scheme in which the video to be processed is the initial video, the video processing apparatus can also input the video to be processed into the second redraw model to obtain a second redraw video. The person style in the second redraw video is different from the person style in the video to be processed, and all areas in the second redraw video are redrawn. The video processing apparatus splices the first redraw video and the second redraw video to obtain a third redraw video. Each video frame in the third redraw video is redrawn.

[0100] Optionally, the input of the second redrawing model further includes a first redrawing style in addition to the video to be processed. The first redrawing style can be input to the second redrawing model at the same time as the video to be processed, or can be input to the first redrawing model before or after the video to be processed, which is not limited here. The first redrawing style is used to indicate the style of the video to be processed. The first redrawing style is similar to the target redrawing style described above, which is not repeated here. The first redrawing style can be the same as the target redrawing style, or can be different, which is not limited here.

[0101] In the case where the video to be processed is a video that has been redrawn, it means that the video processing device obtains the video to be processed by inputting the initial video to the second redrawing model. Among them, the character style in the video to be processed is different from the character style in the initial video, and all regions in the video to be processed are redrawn. The video processing device can also splice the first redrawn video and the video to be processed to obtain a fourth redrawn video.

[0102] Optionally, the input of the second redrawing model further includes a second redrawing style in addition to the initial video. The second redrawing style can be input to the second redrawing model at the same time as the initial video, or can be input to the second redrawing model before or after the initial video, which is not limited here. The second redrawing style is used to indicate the style of the initial video being redrawn, which is similar to the target redrawing style described above. There are many possibilities for the second redrawing style, which can be the same as the target redrawing style or different, which is not limited here.

[0103] In the embodiments of the present application, the video to be processed has many possibilities, and the feature region of the target character obtained from the video to be processed has many possibilities, which enriches the implementation mode and application scene of the technical scheme of the present application, and further improves the flexibility of the technical scheme of the present application. In addition, the video processing device will splice the first redrawn video and the redrawn video (i.e. the second redrawn video or the third redrawn video) obtained by redrawing other regions in the video to be processed or the initial video, so that each video frame in the spliced video is a redrawn video frame, that is, the entire video is redrawn.

[0104] Similarly, the video processing device determines the first redrawing style and the second redrawing style in a manner similar to the manner of determining the target redrawing style of the target character, as described above, which is not repeated here.

[0105] In the foregoing embodiment, the feature image of the target person is stored in the person feature library. In actual application, the person feature library can store the feature images of all persons in the video to be processed. Specifically, after obtaining the video to be processed, the video processing device can split the video to be processed into a plurality of video frames. Starting from the first frame of the plurality of video frames, all persons included in the plurality of video frames are determined based on a person feature detection algorithm. The feature images of all persons are stored in the person feature library, wherein the feature image of each person is included in each third image included in the video to be processed, and the confidence of each person in each third image is higher than the confidence of each person in other video frames, that is, the feature image of the person 1 is included in the video frame in which the confidence of the person 1 is highest among the plurality of video frames; and the feature image of the person 2 is included in the video frame in which the confidence of the person 2 is highest among the plurality of video frames.

[0106] In the present application, the person feature library can also include the feature images of each person in the video to be processed. Therefore, when the style of a person other than the target person is redrawn, the corresponding feature image can be found in the person feature library, thereby simplifying the operation process.

[0107] Next, the video processing method provided by the embodiments of the present application will be described in detail in combination with examples

[0108] For example, refer to FIG. 4, which is a flowchart of the video processing method provided by the embodiments of the present application. It should be noted that in the embodiment shown in FIG. 4, the video to be processed is taken as the video to be redrawn.

[0109] As shown in FIG. 4, the video processing device obtains an initial video, and splits the initial video into a plurality of video frames according to the frame rate of the initial video itself. Then starting from the first frame, the plurality of video frames are subjected to person detection based on a person feature detection algorithm. Specifically, the person feature detection algorithm calculates all person features appearing in each video frame. If the difference between two person features exceeds a specified threshold, it means that the two person features correspond to different persons, which are respectively denoted as person 1 and person 2. If the difference between two person features is less than the specified threshold, it means that the two person features correspond to the same person, and the identification of the person is the same. If the difference between a third person feature and the already identified person features is greater than the specified threshold, it means that the third person feature corresponds to a new person, which is denoted as person 3. In this way, N persons in the plurality of video frames are identified. The difference between the person features can be calculated based on an algorithm as a certain distance.

[0110] Based on the character detection algorithm, N characters included in the initial video are identified, and the identities of the N characters and the feature portraits of the N characters are stored in the character feature library. The feature portrait of each character is included in each third image included in the video to be processed, and the confidence of each character in each third image is higher than the confidence of each character in other video frames.

[0111] In the video redrawing stage, based on the aforementioned character detection, after determining the target character, the identity of the target character can be determined. Specifically, the features of the target character are queried and matched with the character features in the character feature library, and when the difference between the features of the two characters is less than a threshold, it is considered that the matching is successful. The identity of the target character is matched successfully, and the feature portrait of the target character is obtained.

[0112] The plurality of video frames obtained by splitting the initial video are input into the second redrawing model for redrawing to obtain the video to be processed. The style of the video to be processed is the second redrawing style.

[0113] Based on the character detection algorithm, the feature region of the target character is obtained from the video to be processed, and the feature region of the target character and the feature portrait of the target character are input into the first redrawing model to obtain the first redrawing video. In the first redrawing video, the redrawing style of the feature region of the target character is the target redrawing style. The target redrawing style and the second redrawing style can be the same or different, and the specific implementation is not limited here.

[0114] Finally, the first redrawing video and the video to be processed are spliced to obtain the fourth redrawing video, that is, the output video. The splicing operation is specifically to paste the first redrawing video to the feature region of the target character in the video to be processed.

[0115] For example, refer to FIG. 5, which is a flowchart of a video processing method provided by an embodiment of the present application. It should be noted that in the embodiment shown in FIG. 5, the video to be processed is taken as the initial video as an example.

[0116] As shown in FIG. 5, the video processing apparatus obtains the initial video, and splits the initial video into a plurality of video frames according to the frame rate of the initial video itself. Then, starting from the first frame, the character detection algorithm is used to detect the characters in the plurality of video frames to identify all characters included in the video to be processed. The specific implementation process is similar to that of the embodiment shown in FIG. 4, and will not be described here.

[0117] In the embodiment shown in FIG. 5, the video processing apparatus also obtains a target image, and determines the target characters in the target image in response to a second operation instruction for the target image. Assuming that the target characters are character 1 and character 2. Then, the video processing apparatus compares all the characters identified above with the target characters to determine the feature portraits of the target characters.

[0118] In the embodiment shown in FIG. 5, the target person is person 1 and person 2, and the identity of person 1 and the feature image of person 1, and the identity of person 2 and the feature image of person 2 are stored in the person feature library. Among them, the feature image of each person is contained in each third image included in the video to be processed, and the confidence of each person in each third image is higher than the confidence of each person in other video frames.

[0119] The person detection algorithm of the video processing device determines the identity of the target person, which is used to obtain the feature image of the target person from the person feature library when redrawing the feature region of the target person. In addition, in the video redrawing stage, the feature region of the target person is obtained from the video to be processed based on the person detection algorithm. The feature region of the target person and the feature image of the target person are input into the first redrawing model to obtain the first redrawing video. In the first redrawing video, the redrawing style of the feature region of the target person is the target redrawing style.

[0120] The video processing device can also input the plurality of video frames obtained by splitting the initial video into the second redrawing model for redrawing to obtain a second redrawing video. The style of the second redrawing video is the first redrawing style. Among them, the target redrawing style and the first redrawing style can be the same or different, which is not limited here.

[0121] Finally, the first redrawing video and the second redrawing video are spliced to obtain a third redrawing video, that is, an output video. Among them, the splicing operation is specifically pasting the first redrawing video into the feature region of the target person in the second redrawing video.

[0122] It should be noted that in the embodiments of the present application, the specific person feature detection algorithm is not limited, as long as it can recognize the face region or clothing region of the person. The algorithm used by the first redrawing model is also not limited, as long as it can maintain the feature region of the person, such as roop algorithm, instant ID algorithm, image prompt adapter (IP adapter) algorithm, person ID low-rank adapter (LoRA) algorithm, etc. The types of the first redrawing model and the second redrawing model are also not limited, as long as they are models used for video redrawing, such as diffusion model, autoregressive model, etc. The specific embodiments are not limited here.

[0123] Please refer to FIG. 6, which is a structural schematic diagram of a video processing device provided by an embodiment of the present application. As shown in FIG. 6, the video processing device 600 includes an acquisition unit 601 and a processing unit 602.

[0124] In some optional embodiments, the obtaining unit 601 is configured to obtain a to-be-processed video, the to-be-processed video comprising a target person. A feature image of the target person is obtained from a person feature library, the feature image of the target person corresponding to a feature region of the target person, the person feature library comprising feature images of at least one person.

[0125] The processing unit 602 is configured to input the feature image of the target person and the feature region of the target person in the to-be-processed video into a first redraw model to obtain a first redrawn video.

[0126] In some optional embodiments, the processing unit 602 is specifically configured to input the feature image of the target person, the feature region of the target person in the to-be-processed video, and a target redraw style into the first redraw model to obtain the first redrawn video, the redraw style being used to indicate a style of the feature region of the target person in the first redrawn video.

[0127] In some optional embodiments, the target redraw style comprises a two-dimensional image style and / or a three-dimensional image style.

[0128] In some optional embodiments, the processing unit 602 is further configured to, in response to a first operation instruction for a target video frame in the to-be-processed video, determine the target person, the first operation instruction being used to specify the target person included in the target video frame. The feature image of the target person is stored in the person feature library, the feature image of the target person being contained in a first image included in the to-be-processed video, and a confidence of the target person in the first image being higher than a confidence of the target person in other video frames.

[0129] In some optional embodiments, the obtaining unit 601 is further configured to obtain a target image.

[0130] The processing unit 602 is further configured to, in response to a second operation instruction for the target image, determine the target person, the second operation instruction being used to specify the target person included in the target image. The feature image of the target person is stored in the person feature library, the feature image of the target person being contained in a second image, and a confidence of the target person in the second image being higher than a confidence of the target person in other images, the other images being a plurality of video frames included in the to-be-processed video and images other than the second image in the target image.

[0131] In some optional embodiments, the person feature library further comprises identifications of at least one person, the identifications of the at least one person corresponding to the feature images of the at least one person in a one-to-one manner. The obtaining unit 601 is specifically configured to obtain the feature image of the target person based on the identification of the target person.

[0132] In some optional embodiments, the processing unit 602 is further configured to input the to-be-processed video into the second redrawing model to obtain a second redrawing video. The first redrawing video and the second redrawing video are spliced to obtain a third redrawing video.

[0133] In some optional embodiments, the processing unit 602 is specifically configured to input the first redrawing style and the to-be-processed video into the second redrawing model, and the first redrawing style is used to indicate a style in which the to-be-processed video is redrawn.

[0134] In some optional embodiments, the obtaining unit 601 is specifically configured to input the initial video into the second redrawing model to obtain the to-be-processed video.

[0135] The processing unit 602 is further configured to splice the first redrawing video and the to-be-processed video to obtain a fourth redrawing video.

[0136] In some optional embodiments, the obtaining unit 601 is specifically configured to input the initial video and a second redrawing style into the second redrawing model to obtain the to-be-processed video. The second redrawing style is used to indicate a style in which the initial video is redrawn.

[0137] In some optional embodiments, the processing unit 602 is further configured to split the to-be-processed video into a plurality of video frames. Starting from a first frame of the plurality of video frames, all characters included in the plurality of video frames are determined based on a character feature detection algorithm. Feature portraits of all the characters are stored in a character feature library, wherein each feature portrait of each character is contained in each third image included in the to-be-processed video, and a confidence of each character in each third image is higher than a confidence of each character in other video frames.

[0138] In some optional embodiments, the feature region of the target character includes a facial feature region of the target character and / or a clothing feature region of the target character.

[0139] The obtaining unit 601 and the processing unit 602 can be implemented by software or by hardware. For example, the implementation of the processing unit 602 is described below. Similarly, the implementation of the obtaining unit 601 can refer to the implementation of the processing unit 602.

[0140] As an example of a software functional unit, the processing unit 602 can include code running on a compute instance. Among others, the compute instance can include at least one of a physical host (a computing device), a virtual machine, a container. Further, the compute instance can be one or more. For example, the processing unit 602 can include code running on multiple hosts / virtual machines / containers. It is noted that the multiple hosts / virtual machines / containers for running the code can be distributed in the same region, or in different regions. Further, the multiple hosts / virtual machines / containers for running the code can be distributed in the same availability zone (AZ), or in different AZs, each of which includes one data center or multiple data centers in close geographical proximity. Typically, a region can include multiple AZs.

[0141] Similarly, the multiple hosts / virtual machines / containers for running the code can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Typically, a VPC is set up within a region, and communication between two VPCs in the same region, or between VPCs in different regions, requires a communication gateway to be set up in each VPC, and the interconnection between VPCs is achieved via the communication gateway.

[0142] As an example of a hardware functional unit, the processing unit 602 can include at least one computing device, such as a server, etc. Alternatively, the processing unit 602 can also be a device implemented with an application-specific integrated circuit (ASIC), or a programmable logic device (PLD), etc. Among others, the PLD can be implemented with a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0143] The plurality of computing devices included in the processing unit 602 can be distributed in the same region or in different regions. The plurality of computing devices included in the processing unit 602 can be distributed in the same AZ or in different AZs. Similarly, the plurality of computing devices included in the processing unit 602 can be distributed in the same VPC or in multiple VPCs. The plurality of computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0144] It should be noted that the acquisition unit 601 and the processing unit 602 respectively implement different steps in the video processing method to realize the overall function of the video processing apparatus 600. The video processing apparatus 600 is configured to perform the operations performed by the client in the embodiment shown in FIG. 1, the cloud service platform in the embodiment shown in FIG. 2, or the video processing apparatus in the embodiments shown in FIGS. 3 to 5, which will not be repeated here.

[0145] Referring to FIG. 7, FIG. 7 is a structural schematic diagram of a computing device provided in an embodiment of the present application. The computing device 700 includes a processor 701, a communication interface 702, a bus 703, and a memory 704. The processor 701, the communication interface 702, and the memory 704 communicate with each other through the bus 703, and in actual application, communication can also be realized through wireless transmission or other means, which is not limited here.

[0146] The computing device 700 can be a server or a terminal device, and it should be understood that the number of processors and memories in the computing device 700 is not limited in the present application.

[0147] The processor 701 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0148] The communication interface 702 uses a transceiver module such as, but not limited to, a network interface card and a transceiver to realize communication between the computing device 700 and other devices or communication networks.

[0149] The bus 703 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is shown in FIG. 7, but it does not mean that there is only one bus or only one type of bus. The bus 703 can include a path for transmitting information between various components (e.g., the memory 704, the processor 701, the communication interface 702) of the computing device 700.

[0150] The memory 704 can include a volatile memory, such as a random access memory (RAM). The memory 704 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0151] The memory 704 stores executable program code, and the processor 701 executes the executable program code to implement the functions of the foregoing acquisition unit 601 and processing unit 602, respectively, so as to implement the video processing method. That is, the memory 704 stores instructions for executing the video processing method.

[0152] The embodiments of the present application also provide a computing device cluster, which includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some optional embodiments, the computing device can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.

[0153] Please refer to FIG. 8 and FIG. 9, both of which are structural schematic diagrams of a computing device cluster provided by the embodiments of the present application.

[0154] As shown in FIG. 8, the computing device cluster includes at least one computing device 700. The memory 704 in one or more computing devices 700 in the computing device cluster can store the same instructions for executing the video processing method provided by the embodiments of the present application.

[0155] In some possible implementation manners, the memory 704 of one or more of the computing devices 700 in the computing device cluster can also respectively store partial instructions for performing the video processing method. In other words, the combination of the memories 704 of the one or more computing devices can collectively perform the instructions for performing the video processing method.

[0156] It should be noted that the memories 704 in different computing devices 700 in the computing device cluster can store different instructions, respectively used to perform partial functions of the video processing apparatus. That is, the instructions stored in the memories 704 in different computing devices 700 can implement the functions of one or more of the obtaining unit 601 and the processing unit 602.

[0157] In some possible implementation manners, one or more of the computing devices in the computing device cluster can be connected through a network. The network can be a wide area network, a local area network, or the like. FIG. 9 shows a possible implementation manner. As shown in FIG. 9, two computing devices 700A and 700B are connected through a network. Specifically, the computing devices are connected to the network through communication interfaces in the computing devices. In this type of possible implementation manner, the memory 704 in the computing device 700A stores instructions for performing the functions of the obtaining unit 601. Meanwhile, the memory 704 in the computing device 700B stores instructions for performing the functions of the processing unit 602.

[0158] The connection manner between the computing device cluster shown in FIG. 9 can be that, in the video processing method provided in the present application, the processing operation and the operation other than the processing operation are performed separately, that is, the functions of the obtaining unit 601 are thus performed by the computing device 700A, and the functions of the processing unit 602 are thus performed by the computing device 700B.

[0159] It should be understood that the functions of the computing device 700A shown in FIG. 9 can also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B can also be completed by multiple computing devices 700.

[0160] The embodiments of the present application further provide another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to the connection manners of the computing device clusters described with reference to FIG. 8 and FIG. 9, which will not be described herein again.

[0161] The embodiments of the present application further provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions, which can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device is caused to perform the video processing method described above.

[0162] The embodiments of the present application further provide a computer readable storage medium. The computer readable storage medium can be any available medium or data storage device that can be used to store instructions that can be executed by a computing device, or a data center containing one or more available media or data storage devices. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium includes instructions that instruct the computing device to execute the video processing method described above.

[0163] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.

[0164] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A video processing method, characterized in that, include: Acquire a video to be processed, the video including the target person; The feature profile of the target person is obtained from the feature database, and the feature profile of the target person corresponds to the feature region of the target person. The feature database includes feature profiles of at least one person. The feature profile of the target person and the feature region of the target person in the video to be processed are input into the first redraw model to obtain the first redraw video.

2. The method according to claim 1, characterized in that, The step of inputting the feature profile of the target person and the feature region of the target person in the video to be processed into the first redraw model to obtain the first redraw video includes: The feature profile of the target person, the feature region of the target person in the video to be processed, and the target redrawing style are input into the first redrawing model to obtain the first redrawing video. The redrawing style is used to indicate the style of the feature region of the target person in the first redrawing video.

3. The method according to claim 2, characterized in that, The target redrawing style includes a two-dimensional image style and / or a three-dimensional image style.

4. The method according to any one of claims 1 to 3, characterized in that, After acquiring the video to be processed, the method further includes: In response to a first operation instruction for a target video frame in the video to be processed, the target person is determined, wherein the first operation instruction is used to specify the target person included in the target video frame; The feature profile of the target person is stored in the feature database. The feature profile of the target person is included in the first image of the video to be processed. The confidence level of the target person in the first image is higher than that of the target person in other video frames.

5. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Acquire the target image; In response to a second operation instruction for the target image, the target person is determined, wherein the second operation instruction is used to specify the target person included in the target image; The feature profile of the target person is stored in the feature database. The feature profile of the target person is included in the second image. The confidence level of the target person in the second image is higher than that of the target person in other images. The other images are multiple video frames included in the video to be processed and images in the target image other than the second image.

6. The method according to any one of claims 1 to 5, characterized in that, The character feature database also includes the identifier of at least one character, and the identifier of the at least one character corresponds one-to-one with the feature portrait of the at least one character; The step of obtaining the feature profile of the target person from the feature database includes: Based on the identifier of the target person, obtain a feature profile of the target person.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: The video to be processed is input into the second redraw model to obtain the second redraw video; The first redraw video and the second redraw video are spliced ​​together to obtain the third redraw video.

8. The method according to any one of claims 1 to 6, characterized in that, The step of obtaining the video to be processed includes: inputting the initial video into the second redraw model to obtain the video to be processed; The method further includes: splicing the first redrawn video and the video to be processed to obtain a fourth redrawn video.

9. The method according to any one of claims 1 to 7, characterized in that, After acquiring the video to be processed, the method further includes: The video to be processed is split into multiple video frames; Starting from the first frame of the plurality of video frames, all the people included in the plurality of video frames are determined based on a person feature detection algorithm; The feature profiles of all the individuals are stored in the individual feature database, wherein the feature profile of each individual is included in each third image of the video to be processed, and the confidence level of each individual in each third image is higher than the confidence level of each individual in other video frames.

10. The method according to any one of claims 1 to 9, characterized in that, The characteristic regions of the target person include the facial characteristic regions of the target person and / or the clothing characteristic regions of the target person.

11. A video processing apparatus, characterized in that, include: An acquisition unit is used to acquire a video to be processed, wherein the video to be processed includes the target person; The acquisition unit is further configured to acquire the feature portrait of the target person from the feature database, wherein the feature portrait of the target person corresponds to the feature region of the target person, and the feature database includes the feature portrait of at least one person. The processing unit is used to input the feature portrait of the target person and the feature region of the target person in the video to be processed into the first redraw model to obtain the first redraw video.

12. The apparatus according to claim 11, characterized in that, The processing unit is specifically used for: The feature profile of the target person, the feature region of the target person in the video to be processed, and the target redrawing style are input into the first redrawing model to obtain the first redrawing video. The redrawing style is used to indicate the style of the feature region of the target person in the first redrawing video.

13. The apparatus according to claim 12, characterized in that, The target redrawing style includes a two-dimensional image style and / or a three-dimensional image style.

14. The apparatus according to any one of claims 11 to 13, characterized in that, The processing unit is further configured to: In response to a first operation instruction for a target video frame in the video to be processed, the target person is determined, wherein the first operation instruction is used to specify the target person included in the target video frame; The feature profile of the target person is stored in the person feature database. The feature profile of the target person is included in the first image of the video to be processed. The confidence level of the target person in the first image is higher than that of the target person in the other video frames.

15. The apparatus according to any one of claims 11 to 13, characterized in that, The acquisition unit is also used to acquire the target image; The processing unit is further configured to respond to a second operation instruction for the target image to determine the target person, wherein the second operation instruction is configured to specify the target person included in the target image; The processing unit is further configured to store the feature portrait of the target person in the person feature database. The feature portrait of the target person is contained in the second image. The confidence level of the target person in the second image is higher than that of the target person in other images. The other images are multiple video frames included in the video to be processed and images in the target image other than the second image.

16. The apparatus according to any one of claims 11 to 15, characterized in that, The character feature database also includes the identifier of at least one character, and the identifier of the at least one character corresponds one-to-one with the feature portrait of the at least one character; The acquisition unit is specifically used to acquire a feature profile of the target person based on the target person's identifier.

17. The apparatus according to any one of claims 11 to 16, characterized in that, The processing unit is further configured to: The video to be processed is input into the second redraw model to obtain the second redraw video; The first redraw video and the second redraw video are spliced ​​together to obtain the third redraw video.

18. The apparatus according to any one of claims 11 to 16, characterized in that, The acquisition unit is specifically used to input the initial video into the second redraw model to obtain the video to be processed. The processing unit is also used to stitch together the first redrawn video and the video to be processed to obtain a fourth redrawn video.

19. The apparatus according to any one of claims 11 to 17, characterized in that, The processing unit is further configured to: The video to be processed is split into multiple video frames; Starting from the first frame of the plurality of video frames, all the people included in the plurality of video frames are determined based on a person feature detection algorithm; The feature profiles of all the individuals are stored in the individual feature database, wherein the feature profile of each individual is included in each third image of the video to be processed, and the confidence level of each individual in each third image is higher than the confidence level of each individual in other video frames.

20. A video processing apparatus, characterized in that, Includes a processor, which is coupled to a memory; The memory stores instructions that, when executed on the processor, cause the video processing apparatus to perform the method of any one of claims 1 to 10.

21. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 10.

22. A computer program product containing instructions, characterized in that, When the instruction is executed on the communication device, it causes the communication device to perform the method as described in any one of claims 1 to 10; or, When the instruction is executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1 to 10.

23. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer program instructions that, when executed on a communication device, cause the communication device to perform the method as described in any one of claims 1 to 10; or, When the computer program instructions are executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and device for extracting human face from video

    CN110738117A

  • Method for redrawing planar object in video, electronic equipment and medium

    CN115880167A

  • Image redrawing model training method, image redrawing method and device

    CN116664719A

  • Animation generation optimization method and system

    CN116843799A

  • Design support device, computer-readable recording medium, design support method and integrated circuit

    US20120198370A1