Video stylized processing method and device, equipment and medium
By acquiring the original video, style guidance information, and noise data, and using the target model to perform multi-step denoising operations, the problem of large differences in frame image style in video stylization processing is solved, achieving high-quality video stylization processing and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies for stylizing videos result in significant stylistic differences between frames, leading to poor video quality and a subpar user experience.
By acquiring the original video, style guidance information, and noise data, multi-step denoising operations are performed using the target model, and stylization processing is carried out based on the target feature tensor to generate the target video.
It enables stylization processing of existing videos, improves the effect of video stylization processing, and enhances the user experience.
Smart Images

Figure CN121860844A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to the field of video processing technology, and more particularly to methods, apparatus, devices, and media for stylizing video. Background Technology
[0002] With the continuous development of internet and digital media technologies, digital media is increasingly widely used in people's work and lives, providing numerous conveniences. Currently, many platforms have emerged that provide video services, not only promoting the dissemination of videos on the internet but also offering people more ways to edit and process them. To make videos more engaging, people want to stylize existing videos, giving them one or more user-specified styles. Therefore, a solution for stylizing videos is needed. Summary of the Invention
[0003] Embodiments of this disclosure describe a method, apparatus, device, and medium for stylizing video.
[0004] According to a first aspect, a method for stylizing a video is provided, the method comprising: acquiring an original video to be processed, style guidance information, and noise data; wherein the style guidance information includes at least a first reference image for indicating a first style; determining a target feature tensor based at least on the original video, the first reference image, and the noise data; performing a multi-step denoising operation using a target model based on the target feature tensor to obtain a target video; wherein the target video has the same content as the original video, and the target video has at least the first style.
[0005] According to a second aspect, a video stylization processing apparatus is provided, the apparatus comprising: an acquisition unit configured to acquire an original video to be processed, style guidance information, and noise data; wherein the style guidance information includes at least a first reference image for indicating a first style; a determination unit configured to determine a target feature tensor based at least on the original video, the first reference image, and the noise data; and a denoising unit configured to perform multi-step denoising operations using a target model based on the target feature tensor to obtain a target video; wherein the target video has the same content as the original video, and the target video has at least the first style.
[0006] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed in a computer, causes the computer to perform any of the methods described in the first aspect.
[0007] According to a fourth aspect, an electronic device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in any one of the first aspects.
[0008] According to the video stylization processing scheme provided in this disclosure, the original video to be processed, style guidance information, and noise data are obtained. The style guidance information includes at least a first reference image indicating a first style. A target feature tensor is determined based on at least the original video, the first reference image, and the noise data. Based on the target feature tensor, a multi-step denoising operation is performed using a target model to obtain a target video. The target video has the same content as the original video, and the target video at least possesses the first style. This achieves the purpose of stylizing existing videos, improves the effect of stylizing existing videos, and enhances the user experience. Attached Figure Description
[0009] Figure 1 This is a schematic diagram illustrating a scene of stylized video processing according to an exemplary embodiment of the present disclosure;
[0010] Figure 2 This is a schematic diagram of an exemplary system architecture for applying embodiments of this disclosure;
[0011] Figure 3 This is a flowchart illustrating a video stylization processing method according to an exemplary embodiment of the present disclosure;
[0012] Figure 4 This is a schematic diagram illustrating another scenario of stylized video processing according to an exemplary embodiment of this disclosure;
[0013] Figure 5 This is a schematic diagram illustrating the effect of stylized processing of a video according to an exemplary embodiment of the present disclosure;
[0014] Figure 6 This is a block diagram of a video stylization processing apparatus according to an exemplary embodiment of the present disclosure;
[0015] Figure 7 This is a schematic block diagram of an electronic device provided in some embodiments of this disclosure. Detailed Implementation
[0016] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0017] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations of the technical solutions disclosed herein, based on the prompt message.
[0018] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0019] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0020] The technical solutions provided in this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the relevant invention and not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0021] With the continuous development of internet and digital media technologies, digital media is increasingly widely used in people's work and lives, providing numerous conveniences. Currently, many platforms have emerged that provide video services, not only promoting the dissemination of videos on the internet but also offering people more ways to edit and process them. To make videos more engaging, people want to stylize existing videos, giving them one or more styles specified by the user.
[0022] Currently, image stylization is feasible. Since videos consist of multiple frames, theoretically, each frame can be stylized, and the resulting stylized images can then be combined to create a stylized video. However, in practical applications, the stylization results for each frame are highly random, leading to significant differences in style between different images and poor video quality.
[0023] This disclosure provides a video stylization processing scheme. It involves acquiring the original video to be processed, style guidance information, and noise data. The style guidance information includes at least a first reference image indicating a first style. A target feature tensor is determined based on at least the original video, the first reference image, and the noise data. Based on the target feature tensor, a multi-step denoising operation is performed using a target model to obtain a target video. The target video has the same content as the original video and at least possesses the first style. This achieves the purpose of stylizing existing videos, improves the effect of stylizing existing videos, and enhances the user experience.
[0024] See Figure 1 This is a schematic diagram illustrating a scene of stylized processing of a video according to an exemplary embodiment.
[0025] like Figure 1 As shown, firstly, the user can input the original video to be stylized into the video processing client via a terminal device. For example, the original video may include n video frames. The video processing client can extract the n video frames from the original video to obtain a video image sequence, and input the video image sequence into encoder A for encoding processing. The user can also input reference images P1 and P2 into encoder A for encoding processing. Reference image P1 can be an image with a preset style, serving as a stylization reference, while reference image P2 can be an image obtained by stylizing the first frame of the original video according to a specified style. The styles corresponding to reference image P1 and reference image P2 can be different. After encoding processing by encoder A, encoder A can output a conditional tensor and merge the noise tensor and the conditional tensor to obtain the target tensor.
[0026] Additionally, users can input text guidance information into the video processing client via their terminal devices. This text guidance information can describe a specific style. Encoder C can be used to extract semantic features from the text guidance information to obtain text style semantic information. The reference image P1 is then input into encoder B for semantic feature extraction to obtain image style semantic information.
[0027] Finally, the target tensor, text style semantic information, and image style semantic information can be input into model M. Model M can be a diffusion model. Model M can perform multi-step denoising processing on the noisy part of the target tensor based on the target tensor, text style semantic information, and image style semantic information to obtain the target video.
[0028] It should be noted that, in Figure 1In the example of video stylization processing, the stylization process is described using the video processing client directly performing stylization. In other embodiments, the video processing client can transmit the original video and style guidance information over the network to the video processing server deployed on the service platform. The video processing server then generates a target video based on the original video and style guidance information and transmits the target video over the network to the video processing client to provide the target video to the user. See details below. Figure 2 Example.
[0029] Figure 2 This is a schematic diagram of an exemplary system architecture for applying embodiments of this disclosure.
[0030] like Figure 2 As shown, system architecture 200 may include terminal device 202, network 203, and server 204. It should be understood that... Figure 2 The number or type of terminal devices, networks, and servers shown in the diagram is merely illustrative. Any number or type of terminal devices, networks, and servers can be included depending on actual needs.
[0031] Network 203 is a medium used to provide communication links between terminal devices and servers. Network 203 can include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0032] The terminal device 202 is equipped with a video processing client. The terminal device 202 can interact with the server via network 203 to receive or send requests or information. The terminal device 202 can be various electronic devices, including but not limited to smartphones, tablets, laptops, desktop computers, and smart wearable devices.
[0033] Server 204 houses a video processing server. Server 204 can store, analyze, and process received data, and can also send control commands or requests to terminal devices or other servers. The server can provide video processing services in response to user service requests. It is understood that a single server can provide one or more services, and the same service can be provided by multiple servers.
[0034] based on Figure 2In the system architecture shown in this embodiment, user 201 can input the original video and style guidance information to be stylized via terminal device 202. Terminal device 202 can transmit the original video and style guidance information to server 204 via network 203. After receiving the original video and style guidance information, server 204 can generate a target video based on the original video and style guidance information. Finally, server 204 can return the target video to terminal device 202 via network 203, allowing user 201 to view and save the target video via terminal device 202.
[0035] The present disclosure will now be described in detail with reference to specific embodiments.
[0036] Figure 3 This is a flowchart illustrating a video stylization processing method according to an exemplary embodiment. The method can be applied to a video processing client or a video processing server. In this embodiment, the video processing client is installed on a terminal device, which may include, but is not limited to, mobile terminal devices such as smartphones, smart wearable devices, tablets, laptops, and desktop computers. The video processing server is deployed in a service platform, which can be any device, server, or device cluster with computing and processing capabilities. The method may include the following steps:
[0037] like Figure 3 As shown, in step 301, the original video to be processed, style guidance information, and noise data are obtained.
[0038] In this embodiment, the original video to be processed can be a video requiring stylization. The style guidance information can at least include a first reference image indicating a first style, meaning the first reference image can be any image with the first style. The style guidance information can also include a second reference image indicating a second style, which can be an image with the second style obtained by stylizing the first frame of the original video. The first style of the first reference image and the second style of the second reference image can be different styles. The style guidance information can also include text guidance information indicating a third style.
[0039] In this embodiment, upon triggering the first event, the video processing client can obtain the original video to be processed. The first event here specifically refers to the event triggered when a user performs a preset operation on the original video or the video processing client on their terminal device. The original video, i.e., the video material requiring stylization processing, has a flexible source; it can be a local video downloaded or filmed by the user beforehand on their terminal device, or a cloud video uploaded to an online photo album by the user in advance. Specifically, the methods for obtaining the original video can be categorized as follows:
[0040] Firstly, the video processing client can provide users with a video input interface. After the user triggers this interface, the client will display a list of videos in the terminal device's local album or the user's online album. The user can then select the target video and transfer it to the video processing client.
[0041] Secondly, users can also directly open the local album or personal online album on their terminal device, select the original video, and then perform preset operations to trigger the stylization function on the video (such as clicking the stylization button, completing a specified gesture operation, etc.), thereby transmitting the original video to the video processing client.
[0042] Thirdly, when a user posts a video on a social media platform, the posting page will display the video to be posted. The user can then perform preset operations on the video to trigger a stylization process. The system will automatically transmit the video as the original video to the video processing client and simultaneously open the client's stylization page for further processing. Therefore, this embodiment does not limit the specific method of obtaining the original video.
[0043] In this embodiment, the video processing client can acquire style guidance information upon triggering the second event. The second event refers to the event triggered when a user performs a specified operation on the user interface provided by the video processing client via a terminal device. The style guidance information can take various forms, including reference images carrying a preset style, text guidance information describing the preset style, or a combination of reference images and text guidance information.
[0044] Style reference information can be obtained through one or a combination of the following methods:
[0045] In the first method, the video processing client can provide users with a style image selection interface. After the user triggers this interface, the client will display a list of style images to select from, and the user can choose their favorite style image as the first reference image.
[0046] In the second method, the video processing client can also provide a text input interface. After the user triggers this interface, the client will pop up a text input box. The user can enter text content to describe the preset style in the box, and this text can be used as style reference information.
[0047] The third approach is that the video processing client can also provide users with a style selection interface. Users can select their preferred style through the style selection interface, and the video processing client can stylize the first frame of the original video according to the user's preferred style to obtain a second reference image.
[0048] When the three methods mentioned above are combined, the first reference image, the second reference image, and the text guidance information can complement each other and more accurately define the required video style.
[0049] In step 302, the target feature tensor is determined based at least on the original video, the first reference image, and the noise data.
[0050] In this embodiment, the style guidance information may include, but is not limited to, any one or more of the following: a first reference image indicating a first style, a second reference image indicating a second style, and text guidance information indicating a third style. The target feature tensor can be determined based on the original video, the first reference image, the second reference image, the text guidance information, and noise data.
[0051] Specifically, a preset encoder can be used to encode the original video, the first reference image, and the second reference image respectively, to obtain a first tensor, a second tensor, and a third tensor. The preset encoder can be, for example, a variational autoencoder. Encoding the original video using the preset encoder yields the first tensor; encoding the first reference image using the preset encoder yields the second tensor; and encoding the second reference image using the preset encoder yields the third tensor.
[0052] Next, the first, second, and third tensors can be merged to obtain a conditional tensor, where the second tensor follows the first and the third tensor precedes it. Additionally, a mask tensor is generated for each conditional tensor, such that each token of the conditional tensor corresponds to a token of the mask tensor, and each mask tensor token corresponds to a value of 0 or 1. The mask tensor token value corresponding to the token of the first tensor is 0, and the mask tensor token values corresponding to the tokens of the second and third tensors are 1. A noise tensor can also be generated based on the noise data, such that each noise tensor token corresponds to a token of the conditional tensor. Finally, the conditional tensor, mask tensor, and noise tensor are merged to obtain the target feature tensor.
[0053] In step 303, based on the target feature tensor, a multi-step denoising operation is performed using the target model to obtain the target video.
[0054] In this embodiment, a multi-step denoising operation can be performed using a target model based on the target feature tensor to obtain the target video. The target video has the same content as the original video, and the target video has at least a first style. If the style guidance information also includes a second reference image, the target video also has a second style. If the style guidance information also includes style guidance information, the target video also has a third style.
[0055] In this embodiment, the target feature tensor can be injected into the target model through a global attention mechanism, enabling the target model to perform multi-step denoising operations based on the target feature tensor to obtain the target video. Specifically, the target feature tensor can first be dimensionality-reduced and projected using a dimensionality-reduction matrix to obtain multiple category tensors. Different category tensors correspond to different categories of image input information. For example, the different categories of input information may include the first category corresponding to the original video, the second category corresponding to the first reference image, and the third category corresponding to the second reference image. After inputting the different category tensors into different dimensionality-reduction matrices for dimensionality-reduction projection, they are merged to obtain a merged result, which is then injected into the target model through a global attention mechanism.
[0056] like Figure 4 As shown, tensor 401 is the target feature tensor. Tensor 401 can be input into the dimensionality reduction matrix Wdone. The dimensionality reduction matrix Wdone projects tensor 401, resulting in three category tensors, corresponding to the third, first, and second categories, respectively. Then, these three category tensors are input into the dimensionality increase matrices Wup1, Wup2, and Wup3 for dimensionality increase projection, yielding tensor 402 (corresponding to the third category of the second reference image), tensor 403 (corresponding to the first category of the original image), and tensor 404 (corresponding to the second category of the first reference image). Tensor 402, tensor 403, and tensor 404 are merged, and the merged result is injected into the target model through a global attention mechanism.
[0057] Additionally, a semantic extraction model can be used to extract image style semantic information corresponding to the first reference image, and this image style semantic information can be injected into the target model through a cross-attention mechanism. If the style guidance information also includes text guidance information, a text encoder can be used to extract the text style semantic information corresponding to the text guidance information, and this text style semantic information can be injected into the target model through a cross-attention mechanism.
[0058] Figure 5 The effect diagram provided for this implementation is as follows: Figure 5 As shown, the user can input a first reference image 501 and text guidance information 502, namely "colored pencil style", through the client, and input the original video. The client can generate a target video 503 based on the original video, the first reference image 501, and the text guidance information 502. The target video 503 has the style indicated by the first reference image 501 and the style indicated by the text guidance information 502.
[0059] This disclosure provides a video stylization processing method. It involves acquiring the original video to be processed, style guidance information, and noise data. The style guidance information includes at least a first reference image indicating a first style. A target feature tensor is determined based on at least the original video, the first reference image, and the noise data. Based on the target feature tensor, a multi-step denoising operation is performed using a target model to obtain a target video. The target video has the same content as the original video and at least possesses the first style. This achieves the purpose of stylizing existing videos, improves the effect of stylizing existing videos, and enhances the user experience.
[0060] It should be noted that although the operations of the methods of this disclosure embodiment are described in a specific order in the above embodiments, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0061] Corresponding to the aforementioned video stylization processing method embodiments, this disclosure also provides embodiments of video stylization processing apparatus.
[0062] like Figure 6 As shown, Figure 6 This is a block diagram of a video stylization processing apparatus according to an exemplary embodiment of the present disclosure. The apparatus may include: an acquisition unit 601, a determination unit 602, and a noise reduction unit 603.
[0063] The acquisition unit is configured to acquire the original video to be processed, style guidance information, and noise data, wherein the style guidance information includes at least a first reference image for indicating a first style.
[0064] The determination unit 602 is configured to determine the target feature tensor based at least on the original video, the first reference image, and the noise data.
[0065] The denoising unit 603 is configured to perform multi-step denoising operations based on the target feature tensor and the target model to obtain the target video. The target video has the same content as the original video and has at least a first style.
[0066] In some implementations, the denoising unit 603 is configured to: inject the target feature tensor into the target model through a global attention mechanism, so that the target model performs multi-step denoising operations based on the target feature tensor to obtain the target video.
[0067] In other embodiments, the apparatus further includes a first extraction unit and a first injection unit (not shown in the figure).
[0068] The first extraction unit is configured to extract image style semantic information corresponding to the first reference image using a semantic extraction model.
[0069] The first injection unit is configured to inject image style semantic information into the target model through a cross-attention mechanism.
[0070] In other embodiments, the style guidance information may further include text guidance information indicating a third style. Wherein, the target video has at least a first style, including: the target video has both a first style and a third style.
[0071] In other embodiments, the device further includes a second extraction unit and a second injection unit (not shown in the figure).
[0072] The second extraction unit is configured to use a text encoder to extract text style semantic information corresponding to the text guidance information.
[0073] The second injection unit is configured to inject text style semantic information into the target model through a cross-attention mechanism.
[0074] In other embodiments, the denoising unit 603 injects the target feature tensor into the target model via a global attention mechanism as follows: The target feature tensor is dimensionality-reduced and projected using a dimensionality-reducing matrix to obtain multiple category tensors, each corresponding to a different category of image input information. The different category tensors are then input into different dimensionality-increasing matrices for dimensionality-increasing projection and merged to obtain a merged result. This merged result is then injected into the target model via a global attention mechanism.
[0075] In other embodiments, the style guidance information further includes a second reference image for indicating the second style, the second reference image being an image with the second style obtained by stylizing the first frame of the original video. Wherein, the target video has at least the first style, including: the target video having both a first style and a second style.
[0076] In other embodiments, the determining unit 602 is configured to: determine the target feature tensor based on the original video, the first reference image, the second reference image, and the noise data.
[0077] In other embodiments, the determining unit 602 determines the target feature tensor based on the original video, the first reference image, the second reference image, and noise data in the following manner: Using a preset encoder, the original video, the first reference image, and the second reference image are encoded respectively to obtain a first tensor, a second tensor, and a third tensor. The first tensor, the second tensor, and the third tensor are merged to obtain a conditional tensor, such that the second tensor in the conditional tensor follows the first tensor and the third tensor precedes the first tensor. Based on the noise data, a noise tensor is obtained. At least the conditional tensor and the noise tensor are merged to obtain the target feature tensor.
[0078] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the embodiments of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0079] The following is for reference. Figure 7 , Figure 7 This is a schematic block diagram of an electronic device provided for some embodiments of the present disclosure. The electronic device 920 is, for example, suitable for implementing the video stylization processing method provided in the embodiments of the present disclosure. The electronic device 920 can be a terminal device, etc., and can be used to implement a client or server. The electronic device 920 can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. It should be noted that... Figure 7 The illustrated electronic device 920 is merely an example and does not impose any limitation on the functionality and scope of use of the embodiments of this disclosure.
[0080] like Figure 7As shown, the electronic device 920 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 921, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 922 or a program loaded from a storage device 928 into a random access memory (RAM) 923. The RAM 923 also stores various programs and data required for the operation of the electronic device 920. The processing unit 921, ROM 922, and RAM 923 are interconnected via a bus 924. An input / output (I / O) interface 925 is also connected to the bus 924.
[0081] Typically, the following devices can be connected to I / O interface 925: input devices 926 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 927 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 928 including, for example, magnetic tapes, hard disks, etc.; and communication devices 929. Communication device 929 allows electronic device 920 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 7 An electronic device 920 with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown, and the electronic device 920 may alternatively implement or have more or fewer devices. Figure 7 Each box shown can represent a device or multiple devices as needed.
[0082] According to embodiments of this disclosure, the video stylization processing method described above can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the video stylization processing method described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 929, or installed from a storage device 928, or installed from a ROM 922. When the computer program is executed by the processing device 921, the functions defined in the video stylization processing method provided by embodiments of this disclosure can be implemented.
[0083] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the methods provided in this disclosure.
[0084] It should be noted that the computer-readable medium described in the embodiments of this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In the embodiments of this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the embodiments of this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (Radio Frequency), etc., or any suitable combination thereof.
[0085] Computer program code for performing the operations of embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0086] The various embodiments in this disclosure are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for storage media and computing devices are basically similar to the method embodiments, so they are described more simply; relevant parts can be referred to the descriptions of the method embodiments.
[0087] Those skilled in the art will recognize that the functions described in the embodiments of this disclosure in one or more of the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.
[0088] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the embodiments of this disclosure. It should be understood that the above descriptions are merely specific implementations of the embodiments of this disclosure and are not intended to limit the scope of protection of this invention. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of this disclosure should be included within the scope of protection of this invention.
Claims
1. A method for stylizing a video, the method comprising: The process involves acquiring the original video to be processed, style guidance information, and noise data; wherein the style guidance information includes at least a first reference image for indicating a first style. The target feature tensor is determined based at least on the original video, the first reference image, and the noise data; Based on the target feature tensor, a multi-step denoising operation is performed using the target model to obtain the target video; the target video has the same content as the original video, and the target video has at least the first style.
2. The method according to claim 1, wherein, The step of performing multi-step denoising operations based on the target feature tensor and using the target model to obtain the target video includes: The target feature tensor is injected into the target model through a global attention mechanism, enabling the target model to perform multi-step denoising operations based on the target feature tensor to obtain the target video.
3. The method according to claim 2, wherein, The method further includes: The semantic extraction model is used to extract the image style semantic information corresponding to the first reference image; The image style semantic information is injected into the target model through a cross-attention mechanism.
4. The method according to claim 2, wherein, The style guidance information also includes text guidance information for indicating a third style; wherein the target video has at least the first style, including: the target video has both the first style and the third style.
5. The method according to claim 4, wherein, The method further includes: The text style semantic information corresponding to the text guidance information is extracted using a text encoder; The text style semantic information is injected into the target model through a cross-attention mechanism.
6. The method according to claim 2, wherein, The step of injecting the target feature tensor into the target model through a global attention mechanism includes: The target feature tensor is projected into a dimension reduction matrix to obtain multiple category tensors; different category tensors correspond to different categories of image input information. The tensors of different categories are input into different up-dimensional matrices, projected into up-dimensional matrices, and then merged to obtain the merged result. The merged result is injected into the target model through a global attention mechanism.
7. The method according to claim 1, wherein, The style guidance information also includes a second reference image for indicating the second style; the second reference image is an image with the second style obtained by stylizing the first frame of the original video; wherein the target video has at least the first style, including: the target video has both the first style and the second style.
8. The method according to claim 7, wherein, Determining the target feature tensor based at least on the original video, the first reference image, and the noise data includes: The target feature tensor is determined based on the original video, the first reference image, the second reference image, and the noise data.
9. The method according to claim 8, wherein, The step of determining the target feature tensor based on the original video, the first reference image, the second reference image, and the noise data includes: Using a preset encoder, the original video, the first reference image, and the second reference image are encoded respectively to obtain a first tensor, a second tensor, and a third tensor; The first tensor, the second tensor, and the third tensor are merged to obtain a conditional tensor, such that the second tensor is after the first tensor and the third tensor is before the first tensor in the conditional tensor. Based on the noise data, obtain the noise tensor; The condition tensor and the noise tensor are at least merged to obtain the target feature tensor.
10. A video stylization processing apparatus, the apparatus comprising: The acquisition unit is configured to acquire the original video to be processed, style guidance information, and noise data; wherein the style guidance information includes at least a first reference image for indicating a first style; The determining unit is configured to determine a target feature tensor based at least on the original video, the first reference image, and the noise data; The denoising unit is configured to perform multi-step denoising operations based on the target feature tensor using the target model to obtain the target video; the target video has the same content as the original video, and the target video has at least the first style.
11. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-9.
12. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-9.